
OpenAI has released a batch of 722 mathematical manuscripts generated by an unreleased frontier model, covering 372 distinct result families. The release includes solutions to hundreds of open questions, with the company providing details on reasoning summaries and compute costs, noting an average usage of three hours of ChatGPT Pro thinking per result. This publication follows recommendations from the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) for responsible disclosure and academic transparency. The move adds to a growing body of AI-generated mathematical results that are currently being assessed by the broader scientific community.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
OpenAI · August 1, 2026 · Same story
The Verge AI · August 11, 2026 · Related
WIRED AI · September 8, 2026 · Related
Wes Roth · September 9, 2026 · Related
The Rundown AI · September 9, 2026 · Related
TechCrunch AI · September 11, 2026 · Related
The Verge AI · September 12, 2026 · Same story
TechCrunch AI · September 21, 2026 · Related
© The Verge AIOpenAI is shifting ChatGPT from a static text stream to an interactive interface by integrating visuals and tools directly into responses. This Intelligent UI feature, powered by the new GPT-6 models, allows users to engage with diagrams, calculators, and games without leaving the chat window. The update prioritizes utility over pure conversation, enabling tasks like retirement planning or learning Mahjong through in-line interactive elements. While Plus and Pro users get the more capable Sol model, free users receive the efficient Luna variant, making this a broad consumer-facing upgrade rather than a niche developer tool.
© The Verge AIThe Surface Laptop Ultra marks the commercial debut of Nvidia’s RTX Spark, an Arm-based chip designed to bring serious local AI inference to Windows laptops. Starting at $2,599, this device signals a shift toward premium hardware capable of running large models locally, moving beyond cloud dependency for enterprise and creative workflows. Microsoft pairs this with 'Hybrid Intelligence' features in Copilot, allowing the agent to access local files and take OS-level actions like filing taxes or managing emails. This isn't just a new laptop; it's the first concrete proof that Arm-based PC silicon can handle the thermal and memory demands of 120B+ parameter models on-device.
© The Verge AIMicrosoft is shifting Copilot from a chat interface to an active agent capable of manipulating local files and system settings. The new 'Hybrid Intelligence' approach combines cloud reasoning with local execution, allowing the AI to search folders, rename documents, and draft emails without leaving the desktop environment. This moves beyond simple text generation into tangible OS-level automation, effectively turning Copilot into a personal assistant that can handle multi-step workflows like tax filing preparation. The capability arrives over the next few months, marking a significant step toward autonomous desktop agents.
© WIRED AIThree engineers at Axiom proved that general-purpose language models can control physical hardware without task-specific training. By linking OpenAI’s GPT-6 Astra to a Toyota Corolla’s steering system, they navigated the vehicle through an In-N-Out drive-thru using only prompt engineering and camera input. While the car moved slowly and required a safety driver, the experiment reveals that multimodal models are developing emergent spatial reasoning capabilities previously thought to require dedicated robotics stacks. This blurs the line between digital assistants and physical agents, suggesting that scaling text-and-image training yields unexpected real-world utility.
© Hugging Face BlogNVIDIA’s Nemotron models just crossed the gold-medal threshold in both the International Olympiad in Informatics and Mathematics. This isn't a new foundation model; it’s proof that specialized fine-tuning combined with iterative generate-verify-refine inference loops can push existing architectures to world-class levels. The Ultra-CC variant scored 535.4/600 on IOI, while the IMO system solved complex proofs without external tools or formal provers. By releasing the datasets and pipelines, NVIDIA is shifting the narrative from raw parameter count to reproducible specialization recipes.
© MIT News AIThe Lincoln Laboratory Supercomputing Center has published its sixth annual survey of commercial AI accelerators, tracking a landscape that has grown from 57 to over 120 distinct devices since 2018. This longitudinal analysis provides rare, unbiased data on peak performance versus power consumption across CPUs, GPUs, ASICs, and emerging dataflow architectures. By aggregating public specs from dozens of startups and incumbents, the team offers a critical reference for government sponsors navigating a saturated but rapidly innovating market. The findings clarify how architectural shifts like lower numerical precision drive efficiency gains, helping buyers cut through vendor hype to make informed acquisition decisions.