
MIT associate professor Cathy Wu and her team have published research addressing the instability of reinforcement learning (RL) in complex optimization tasks. Their work identifies that RL algorithms typically succeed on only 10% of related problem variants, leading to a new selection algorithm that improves training efficiency by up to 30 times. Applied to transportation, this method demonstrates that intelligent eco-driving controls could reduce vehicle emissions by 11-22%. The findings offer a scalable framework for using RL in logistics and supply chain optimization.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
MIT News AI · June 17, 2026 · Related
Matt Wolfe · June 25, 2026 · Background
Hugging Face Blog · August 18, 2026 · Background
AI News · August 25, 2026 · Background
MIT News AI · September 2, 2026 · Related
AI News · September 2, 2026 · Related
Matt Wolfe · September 3, 2026 · Background
Hugging Face Blog · September 8, 2026 · Background
© Hugging Face BlogMicrosoft and Hugging Face’s ThinkingBox benchmark exposes a critical flaw in AI agents: they often execute tool calls correctly while leaving the database in the wrong state. Testing 507 workflows across 12 models showed that nearly two-thirds of failures involved clean execution but incorrect final side effects. The data proves that capability does not equal consistency; Kimi-K3 solved more tasks initially, but Claude Opus 5.5 was far more reliable on repeated attempts. This shifts the evaluation metric from single-shot success to terminal state verification.
© MIT Technology Review AIA former Google DeepMind researcher argues that current LLMs lack genuine reasoning capabilities, relying instead on fast pattern matching rather than the deliberative search mechanisms seen in AlphaGo. The core issue is that LLMs maintain no persistent, inspectable epistemic state, meaning they cannot track hypotheses or evidence systematically. This architectural flaw makes them unreliable for high-stakes fields like medicine and science where auditability is critical. True machine intelligence requires a separation between knowledge representation and manipulation, moving beyond next-token prediction to auditable inference.
© AI ExplainedOpenAI has released a new research paper exploring the potential for AI systems to recursively improve themselves, leading to rapid intelligence growth.