
NVIDIA has announced a series of new open-source models and tools aimed at enhancing local AI capabilities. Key releases include the Cosmos 3 Edge model for robotics and the MiniMax-H3 model for video and audio generation, both optimized for NVIDIA GPUs. Additionally, the launch of Unsloth Desktop offers a fully open-source solution for AI model training and inference on personal devices. These advancements highlight NVIDIA's commitment to empowering developers with the tools needed to run sophisticated AI tasks locally, minimizing the need for cloud-based resources.
Read originalTopicNvidia Product And Business UpdatesCooling
Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Together AI Blog · April 28, 2026 · Related
Hugging Face Blog · June 1, 2026 · Same story
NVIDIA Blog · June 1, 2026 · Same story
Sam Witteveen · June 1, 2026 · Same story
The Rundown AI · June 2, 2026 · Related
Ollama Blog · June 4, 2026 · Related
Hugging Face Blog · July 20, 2026 · Same story
Ollama Blog · August 11, 2026 · Related
TechCrunch AI · August 21, 2026 · Related
The Verge AI · September 3, 2026 · Related
NVIDIA Blog · September 3, 2026 · Same story
Sam Witteveen · September 6, 2026 · Same story
WIRED AI · September 28, 2026 · Related
NVIDIA Cosmos 3 Advances Open World Models for Physical AI
2 developments
© The Verge AIMeta is handing the keys to its Muse AI agent by open-sourcing the software needed to run it on custom hardware. Developers can now hook Muse into ESP32 boards or Raspberry Pi setups, effectively turning workbench scraps into personalized AI terminals. This moves Muse beyond a cloud-only interface into tangible, local devices like E Ink displays or HDMI sticks. It signals a shift toward decentralized, user-owned AI interactions rather than relying solely on proprietary apps.
© Hugging Face BlogAllen Institute for AI has released AstaBrief 8B, an open-weight model designed specifically for generating cited scientific literature reviews. Built on Qwen3-8B and trained with supervised fine-tuning and direct preference optimization, it prioritizes speed and grounding over complex multi-step reasoning. The model generates full reports in a single pass, cutting generation time to roughly 51 seconds compared to the 178 seconds required by proprietary alternatives like Claude. This release offers researchers a faster, locally deployable option for synthesizing evidence without relying on external APIs.
IBM is finally getting first-class CI support in llama.cpp with the addition of the ZDNN backend for s390x architecture. This isn't just a minor tweak; it enables efficient inference on mainframe hardware, bridging a gap for enterprise environments that rely on IBM Z systems. While currently limited to build pipelines without automated testing, this signals a serious commitment to supporting non-x86/ARM infrastructure in the local LLM ecosystem. It’s a quiet but necessary expansion for anyone running models on legacy or specialized enterprise silicon.