
Google Research has introduced a new framework called knowledge profiling to better understand factual errors in large language models (LLMs). The study finds that models like GPT-5 and Gemini-3 encode most facts but often fail to recall them, especially when the query context changes. This suggests that the bottleneck in LLMs is shifting from knowledge acquisition to utilization. The research also highlights that 'thinking' processes can help recover inaccessible knowledge, though they incur computational costs. This insight could guide future improvements in LLMs by focusing on recall rather than just scaling.
Read original