
Researchers from Axiom successfully drove a Toyota Corolla using OpenAI’s GPT-6 Astra, connecting the model to the car's power steering via windscreen-mounted cameras. The team, consisting of Aditya Ramabadran, Simon Mahns, and Tobias Gessler, used prompt engineering to bypass safety refusals and navigate an In-N-Out Burger drive-thru without prior fine-tuning for driving tasks. They also released 'DrivingBench,' a benchmark showing that while Astra completed the course, other models like Claude Fable 5.1 and Grok performed significantly worse. The experiment highlights emergent physical reasoning in large multimodal models but underscores significant safety risks and performance gaps compared to dedicated autonomous systems.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
MIT News AI · June 17, 2026 · Related
AI News · September 2, 2026 · Related
OpenAI · September 3, 2026 · Related
Matthew Berman · September 3, 2026 · Related
The Rundown AI · September 4, 2026 · Related
The AI Advantage · September 4, 2026 · Related
Skill Leap AI · September 9, 2026 · Related
TechCrunch AI · September 29, 2026 · Related
© WIRED AIOpenAI is shifting ChatGPT from static text to dynamic, interactive interfaces powered by the new GPT-6 model. The update generates custom tools like calculators and clickable diagrams directly within the chat, moving beyond simple image insertion. This generative UI approach mirrors Google’s recent Search updates but targets OpenAI’s massive user base immediately. It marks a tangible step toward AI-generated software components rather than just content generation. Users can now interact with data through sliders and maps instead of reading about them.
© WIRED AIThe Pentagon’s Tradewinds program is bypassing months of red tape by letting companies pitch AI tools in five-minute videos for immediate 'post-competitive' status. This shift allows the military to award contracts in under a week, targeting OpenAI, Anthropic, and Google directly rather than traditional defense primes. The move signals a desperate attempt to inject competition into a market dominated by Silicon Valley giants who now hold more leverage than the government itself. While it accelerates deployment for lethal AI agents, it also raises serious transparency concerns about unreported 'other transaction' spending.
© WIRED AIOpenAI’s Dots agents are no longer just a demo; they are live for ChatGPT subscribers and actively browsing the web to execute complex tasks like shopping. The experience is undeniably janky—misnaming users, failing captchas, and hallucinating emotional connections—but it proves that autonomous agents can navigate real e-commerce flows with surprising competence. This marks a shift from chatbots that talk about actions to bots that actually perform them, even if the execution is currently prone to embarrassing errors. The $100/month price tag signals OpenAI’s confidence in this utility, betting that reliability will catch up to ambition over time.
© Hugging Face BlogNVIDIA’s Nemotron models just crossed the gold-medal threshold in both the International Olympiad in Informatics and Mathematics. This isn't a new foundation model; it’s proof that specialized fine-tuning combined with iterative generate-verify-refine inference loops can push existing architectures to world-class levels. The Ultra-CC variant scored 535.4/600 on IOI, while the IMO system solved complex proofs without external tools or formal provers. By releasing the datasets and pipelines, NVIDIA is shifting the narrative from raw parameter count to reproducible specialization recipes.
© The Verge AIOpenAI has published 722 manuscripts covering 372 result families, marking a significant escalation in AI-driven mathematical discovery. This release, guided by the AGMAI advisory group's ethical guidelines, includes solutions to hundreds of open questions and details on compute usage, such as an average of three hours of ChatGPT Pro thinking per result. The move shifts the conversation from speculative claims to verifiable data, forcing the academic community to confront the reality of AI-generated proofs. It underscores a growing tension between rapid corporate output and traditional peer review standards. Mathematicians now have concrete artifacts to audit rather than vague promises. The transparency around compute costs sets a precedent for future frontier model releases in scientific domains.
© MIT News AIThe Lincoln Laboratory Supercomputing Center has published its sixth annual survey of commercial AI accelerators, tracking a landscape that has grown from 57 to over 120 distinct devices since 2018. This longitudinal analysis provides rare, unbiased data on peak performance versus power consumption across CPUs, GPUs, ASICs, and emerging dataflow architectures. By aggregating public specs from dozens of startups and incumbents, the team offers a critical reference for government sponsors navigating a saturated but rapidly innovating market. The findings clarify how architectural shifts like lower numerical precision drive efficiency gains, helping buyers cut through vendor hype to make informed acquisition decisions.