16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

NVIDIA Unveils Nemotron 3.5 Lightning and NeMo Switchyard

NVIDIA Blog·August 11, 2026·high confidence

Why it matters

  • →Nemotron 3.5 Lightning offers significant speed improvements for agentic AI tasks.
  • →NeMo Switchyard optimizes model routing, reducing costs and improving efficiency.
  • →These tools enhance control over AI deployment across various platforms.
NVIDIA Unveils Nemotron 3.5 Lightning and NeMo Switchyard
©NVIDIA Blog

NVIDIA has introduced Nemotron 3.5 Lightning, a high-efficiency model for agentic AI workloads, and NeMo Switchyard, an open-source library for smart model routing. Nemotron 3.5 Lightning, a 30-billion-parameter model, offers significant speed improvements and customization options for specialized tasks. NeMo Switchyard allows enterprises to route AI tasks to the most suitable models, optimizing performance and cost. These releases enhance AI deployment flexibility across different environments, from local systems to cloud platforms.

Read original

The story around this

TopicNvidia Product And Business UpdatesCooling

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

NVIDIA Launches Nemotron 3 Nano Omni Model — NVIDIA Blog1NVIDIA Launches Nemotron 3 Nano Omni — Matt Wolfe2Nvidia Unveils AI Agent-Centric Tech at COMPUTEX — The Rundown AI3NVIDIA Launches Nemotron 3 Ultra Model — Ollama Blog4NVIDIA Unveils Nemotron 3 Ultra Model — Sam Witteveen5NVIDIA Unveils Nemotron 3 Ultra for AI Agents — Matt Wolfe6NVIDIA Nemotron 3 Ultra Boosts AI Agent Performance — NVIDIA Blog7NVIDIA's Nemotron Enhances AI with Open Synthetic Data — Hugging Face Blog8NVIDIA Unveils Nemotron 3.5 Lightning and NeMo SwitchyardNVIDIA Jetson Orin Nano 2 targets edge AI for robots — AI News9NVIDIA Automates Supply Chain with AI Tools — AI News10NVIDIA Nemotron 3 Diarization Model Released — Sam Witteveen11Apr 28You are hereSep 23

How we got here

  1. 1
    NVIDIA Launches Nemotron 3 Nano Omni Model

    NVIDIA Blog · April 28, 2026 · Same story

  2. 2
    NVIDIA Launches Nemotron 3 Nano Omni

    Matt Wolfe · May 1, 2026 · Same story

  3. 3
    Nvidia Unveils AI Agent-Centric Tech at COMPUTEX

    The Rundown AI · June 2, 2026 · Same story

  4. 4
    NVIDIA Launches Nemotron 3 Ultra Model

    Ollama Blog · June 4, 2026 · Same story

  5. 5
    NVIDIA Unveils Nemotron 3 Ultra Model

    Sam Witteveen · June 4, 2026 · Same story

  6. 6
    NVIDIA Unveils Nemotron 3 Ultra for AI Agents

    Matt Wolfe · June 5, 2026 · Same story

  7. 7
    NVIDIA Nemotron 3 Ultra Boosts AI Agent Performance

    NVIDIA Blog · July 8, 2026 · Same story

  8. 8
    NVIDIA's Nemotron Enhances AI with Open Synthetic Data

    Hugging Face Blog · July 8, 2026 · Same story

What happened next

  1. 9
    NVIDIA Jetson Orin Nano 2 targets edge AI for robots

    AI News · August 26, 2026 · Related

  2. 10
    NVIDIA Automates Supply Chain with AI Tools

    AI News · September 11, 2026 · Related

  3. 11
    NVIDIA Nemotron 3 Diarization Model Released

    Sam Witteveen · September 23, 2026 · Related

Follow this story

Open the full story →

NVIDIA Nemotron 3.5 Lightning Launches for Local AI Agents

6 developments

  1. Aug 11 · Ollama Blog
    NVIDIA Nemotron 3.5 Lightning Launches for Local AI Agents
  2. Aug 11 · NVIDIA Blog
    NVIDIA Unveils Nemotron 3.5 Lightning and NeMo Switchyard (This article)
  3. Aug 11 · Sam Witteveen
    NVIDIA Launches NeMo Switchyard for AI Agents↳ NeMo Switchyard is an open-source library that acts as a router to select the best model for each task, optimizing efficiency and token usag
  4. Aug 11 · Sam Witteveen
    NVIDIA Unveils Nemotron Lightning Model
  5. Aug 14 · Matt Wolfe
    NVIDIA Launches Nemotron 3.5 Lightning↳ The model is designed for specialized task execution.
  6. Aug 14 · Lev Selector
    NVIDIA Unveils Nemotron 3.5 Lightning

More in Models & Labs

Models & Labscoding

vLLM v0.31.0: SM100 defaults and fast restarts

vLLM is quietly becoming the definitive runtime for NVIDIA's latest hardware, making NVFP4 compressed KV caches the default for DeepSeek-V4.1-Flash on SM100 GPUs. This isn't just a performance tweak; it fundamentally changes how enterprise inference scales by keeping post-quantized weights resident in GPU memory across engine restarts via the new preload daemon. The release also hardens speculative decoding with Model Runner V2, fixing OOMs that previously plagued wide expert deployments. For builders, this means lower latency and higher throughput on next-gen hardware without manual configuration overhead.

vLLM Releases·Oct 6, 2026
Models & Labsmodels

llama.cpp v0.6.0 adds decision model API and Apple Silicon speedups

This release shifts llama.cpp from a pure text engine to a multimodal inference runtime capable of handling 'decision models' like Clef and GLM-5.3-Flash. The new /v1/systemone server endpoint standardizes how these non-autoregressive models are queried, while the extended batch API allows mixed token and embedding inputs for complex architectures. Apple Silicon users get a tangible performance boost with new Metal MMA kernels that accelerate speculative decoding by up to 3x. It’s a significant step toward supporting the next generation of hybrid reasoning models locally.

llama.cpp Releases·Oct 6, 2026
Falcon-Emirati-7B targets dialect nuance© Hugging Face Blog
Models & Labsmodels

Falcon-Emirati-7B targets dialect nuance

Most Arabic models treat the language as a monolith, missing the cultural and linguistic depth of specific dialects. Falcon-Emirati-7B closes this gap by fine-tuning on native Emirati text, synthetic data constrained by strict glossaries, and cultural heritage knowledge. It tops the new Alyah benchmark with 84.83%, proving that scale alone doesn't buy dialect competence. This release underscores a critical shift: true multilingual capability requires targeted adaptation, not just larger parameter counts.

Hugging Face Blog·Oct 6, 2026