
OpenAI's internal model, Astra, has reportedly solved 10 longstanding problems in mathematics and computer science, including some that have remained unsolved for nearly 30 years. Among the achievements, Astra proved the existence of non-sofic groups and solved Alain Connes’s rigidity conjecture. The solutions were verified using Lean, and the computational cost was approximately $2,000. This development highlights the potential of AI to tackle complex problems at a fraction of the traditional cost, sparking discussions about the future role of AI in scientific discovery.
Read originalTopicOpenAI Mathematics ControversyCooling
Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
OpenAI · May 20, 2026 · Same story
TechCrunch AI · May 20, 2026 · Same story
The Rundown AI · May 21, 2026 · Same story
The AI Daily Brief · May 24, 2026 · Related
OpenAI · August 1, 2026 · Related
The Verge AI · August 20, 2026 · Same story
WIRED AI · September 8, 2026 · Same story
Wes Roth · September 9, 2026 · Related
The Rundown AI · September 9, 2026 · Same story
The Verge AI · September 12, 2026 · Same story
OpenAI's Astra Solves 10 Long-Standing Math Problems
4 developments
© Hugging Face BlogMicrosoft and Hugging Face’s ThinkingBox benchmark exposes a critical flaw in AI agents: they often execute tool calls correctly while leaving the database in the wrong state. Testing 507 workflows across 12 models showed that nearly two-thirds of failures involved clean execution but incorrect final side effects. The data proves that capability does not equal consistency; Kimi-K3 solved more tasks initially, but Claude Opus 5.5 was far more reliable on repeated attempts. This shifts the evaluation metric from single-shot success to terminal state verification.
© MIT News AICathy Wu’s team at MIT has cracked a persistent bottleneck in reinforcement learning: its notorious sensitivity to specific problem setups. By identifying that RL models train effectively on only about 10 percent of related problems, they developed an algorithm to select those high-yield training cases. This approach boosts training efficiency by up to 30 times, allowing researchers to generalize solutions across complex transportation networks without retraining from scratch. The method transforms RL from a fragile proof-of-concept into a viable tool for evidence-based policy design, specifically showing eco-driving could cut emissions by 11-22 percent.
© MIT Technology Review AIA former Google DeepMind researcher argues that current LLMs lack genuine reasoning capabilities, relying instead on fast pattern matching rather than the deliberative search mechanisms seen in AlphaGo. The core issue is that LLMs maintain no persistent, inspectable epistemic state, meaning they cannot track hypotheses or evidence systematically. This architectural flaw makes them unreliable for high-stakes fields like medicine and science where auditability is critical. True machine intelligence requires a separation between knowledge representation and manipulation, moving beyond next-token prediction to auditable inference.