
Google Research has introduced a new framework called knowledge profiling to better understand factual errors in large language models (LLMs). The study finds that models like GPT-5 and Gemini-3 encode most facts but often fail to recall them, especially when the query context changes. This suggests that the bottleneck in LLMs is shifting from knowledge acquisition to utilization. The research also highlights that 'thinking' processes can help recover inaccessible knowledge, though they incur computational costs. This insight could guide future improvements in LLMs by focusing on recall rather than just scaling.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
AI Explained · May 20, 2026 · Background
TechCrunch AI · June 10, 2026 · Background
Google DeepMind · June 10, 2026 · Background
Google Research Blog · June 24, 2026 · Same story
Microsoft Research · June 29, 2026 · Related
Cole Medin · July 2, 2026 · Background
MIT Technology Review AI · July 30, 2026 · Related
Hugging Face Blog · August 10, 2026 · Related
llama.cpp Releases · August 16, 2026 · Background
Together AI Blog · August 17, 2026 · Background
MIT Technology Review AI · August 24, 2026 · Related
Hugging Face Blog · August 26, 2026 · Related
Google Research Blog · September 10, 2026 · Related
© WIRED AIAnthropic’s Claude identified a novel reverse transcriptase system in jumbo phages that resembles CRISPR, but the scientific community remains skeptical. While the speed of discovery is impressive, experts note the finding lacks wet-lab validation and may simply be pattern recognition on known data. The real story isn't a new gene-editing tool, but the opaque nature of how an AI model sifts through genomic databases to propose hypotheses that humans must still verify.
© Hugging Face BlogMost fact-checkers for AI agents only check if a claim is true in the evidence pool, ignoring where it came from. ProvenanceGuard fixes this by tracking source identity through every step of verification, catching cases where a true fact is wrongly attributed to the wrong tool or document. In medical agent tests, it caught 138 out of 139 incorrect attributions that standard verifiers missed, proving that provenance matters as much as truth in multi-tool environments. This shifts the focus from simple RAG retrieval to rigorous source-aware auditing for high-stakes applications.
© MIT News AIA new MIT study dismantles the alarmist narrative that widespread adoption of a single AI algorithm inevitably leads to systemic exclusion. By modeling hiring scenarios, researchers prove that while monoculture reduces individual discovery, it can actually increase candidate bargaining power and overall hiring volume. The real risk is informational stagnation, which the paper suggests can be mitigated through ensemble methods or injected randomness. This shifts the debate from moral panic to technical optimization of algorithmic diversity.