
Researchers have found a way to extract hidden reasoning from AI models, potentially exposing vulnerabilities in models from OpenAI, Anthropic, and Google. This method suggests that some Chinese models might have been trained using reasoning patterns from US models, sparking concerns about distillation practices. The technique also revealed a vulnerability that allowed the extraction of sensitive information, which has been addressed by the companies. This discovery underscores the geopolitical tensions and ethical considerations surrounding AI model development and distillation.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
© AI ExplainedOpenAI has released a new research paper exploring the potential for AI systems to recursively improve themselves, leading to rapid intelligence growth.
© TechCrunch AIThe Verge AI · June 14, 2026 · Related
The AI Daily Brief · June 25, 2026 · Related
MIT News AI · July 13, 2026 · Related
The Rundown AI · July 23, 2026 · Related
TechCrunch AI · July 23, 2026 · Related
WIRED AI · July 24, 2026 · Related
WIRED AI · July 25, 2026 · Related
MIT Technology Review AI · August 3, 2026 · Related
Google DeepMind · August 27, 2026 · Related
Hugging Face Blog · September 8, 2026 · Related
TechCrunch AI · September 10, 2026 · Related
The Verge AI · September 11, 2026 · Related
MIT Technology Review AI · September 23, 2026 · Related

A new analysis by Graphite proves that frontier models are failing to shed their robotic DNA. Despite labs claiming natural prose, Claude Opus 5.5 still uses "this matters" 116 times more often than humans, while OpenAI's Astra relies on hedging phrases like "may provide." The study of 13,000 phrases shows that as models eliminate old habits like em-dashes, they simply adopt new ones. This suggests labs cannot fully control the statistical quirks inherent in billions of parameters.