16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Open Source
Open Source

NVIDIA Launches NeMo Switchyard for AI Agents

Sam Witteveen·August 11, 2026·high confidence

Why it matters

  • →NeMo Switchyard optimizes AI agent workflows by selecting the best model for each task.
  • →It enhances efficiency in long-running AI agents, improving response times and token usage.
  • →Being open-source, it invites community engagement and innovation in AI agent development.
NVIDIA Launches NeMo Switchyard for AI Agents
©Sam Witteveen

NVIDIA has introduced NeMo Switchyard, an open-source library aimed at improving AI agent workflows. The library functions as a router, selecting the appropriate model for each step in an agent's process, thereby enhancing response times and token efficiency. This development is particularly relevant for long-running AI agents, offering a more streamlined and effective approach to managing workloads. As an open-source project, NeMo Switchyard is accessible to developers looking to innovate in the field of large language model agents.

Read original

The story around this

TopicNvidia Product And Business UpdatesCooling

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

Together AI's Inference Engine Outperforms Competitors — Together AI Blog1NVIDIA Jetson Advances Agentic AI in Robotics — NVIDIA Blog2Nvidia Unveils AI Agent-Centric Tech at COMPUTEX — The Rundown AI3NVIDIA Unveils Nemotron 3 Ultra Model — Sam Witteveen4NVIDIA Unveils Nemotron 3 Ultra for AI Agents — Matt Wolfe5NVIDIA NeMo AutoModel Boosts Transformers Fine-Tuning — Hugging Face Blog6NVIDIA Nemotron 3 Ultra Boosts AI Agent Performance — NVIDIA Blog7NVIDIA's Nemotron Enhances AI with Open Synthetic Data — Hugging Face Blog8NVIDIA Launches NeMo Switchyard for AI AgentsNvidia Highlights Importance of AI Harness Over Model — TechCrunch AI9Nvidia Unveils PAIR for Home AI Computing — The Verge AI10NVIDIA Launches PAIR Local AI Router — Sam Witteveen11Jev: A Decision-Only AI Model for Agents — Cole Medin12Nvidia Launches Open-Source AI Agent Security Platform — WIRED AI13May 19You are hereSep 28

How we got here

  1. 1
    Together AI's Inference Engine Outperforms Competitors

    Together AI Blog · May 19, 2026 · Background

  2. 2
    NVIDIA Jetson Advances Agentic AI in Robotics

    NVIDIA Blog · June 2, 2026 · Related

  3. 3
    Nvidia Unveils AI Agent-Centric Tech at COMPUTEX

    The Rundown AI · June 2, 2026 · Related

  4. 4
    NVIDIA Unveils Nemotron 3 Ultra Model

    Sam Witteveen · June 4, 2026 · Background

  5. 5
    NVIDIA Unveils Nemotron 3 Ultra for AI Agents

    Matt Wolfe · June 5, 2026 · Background

  6. 6
    NVIDIA NeMo AutoModel Boosts Transformers Fine-Tuning

    Hugging Face Blog · June 24, 2026 · Related

  7. 7
    NVIDIA Nemotron 3 Ultra Boosts AI Agent Performance

    NVIDIA Blog · July 8, 2026 · Related

  8. 8
    NVIDIA's Nemotron Enhances AI with Open Synthetic Data

    Hugging Face Blog · July 8, 2026 · Related

What happened next

  1. 9
    Nvidia Highlights Importance of AI Harness Over Model

    TechCrunch AI · August 21, 2026 · Background

  2. 10
    Nvidia Unveils PAIR for Home AI Computing

    The Verge AI · September 3, 2026 · Related

  3. 11
    NVIDIA Launches PAIR Local AI Router

    Sam Witteveen · September 6, 2026 · Related

  4. 12
    Jev: A Decision-Only AI Model for Agents

    Cole Medin · September 21, 2026 · Background

  5. 13
    Nvidia Launches Open-Source AI Agent Security Platform

    WIRED AI · September 28, 2026 · Related

Follow this story

Open the full story →

NVIDIA Nemotron 3.5 Lightning Launches for Local AI Agents

6 developments

  1. Aug 11 · Ollama Blog
    NVIDIA Nemotron 3.5 Lightning Launches for Local AI Agents
  2. Aug 11 · NVIDIA Blog
    NVIDIA Unveils Nemotron 3.5 Lightning and NeMo Switchyard
  3. Aug 11 · Sam Witteveen
    NVIDIA Launches NeMo Switchyard for AI Agents (This article)↳ NeMo Switchyard is an open-source library that acts as a router to select the best model for each task, optimizing efficiency and token usag
  4. Aug 11 · Sam Witteveen
    NVIDIA Unveils Nemotron Lightning Model
  5. Aug 14 · Matt Wolfe
    NVIDIA Launches Nemotron 3.5 Lightning↳ The model is designed for specialized task execution.
  6. Aug 14 · Lev Selector
    NVIDIA Unveils Nemotron 3.5 Lightning

More from Sam Witteveen

Fine-tunes reduce Qwen3.8-27B reasoning tokens© Sam Witteveen
Coding Toolscoding

Fine-tunes reduce Qwen3.8-27B reasoning tokens

Reasoning models are notoriously slow and expensive because they generate excessive internal thought traces before answering. This analysis benchmarks three specific fine-tunes of Qwen3.8-27B—ThinkingCap, Swift 1.5, and QwenPi—that aggressively prune these tokens while maintaining accuracy. The results show a tangible trade-off: significantly faster inference and lower costs for practical tasks without the bloat of full chain-of-thought. For builders running local agents, this offers a viable path to deploy reasoning-capable models that actually feel responsive.

Sam Witteveen·Oct 4, 2026
Specialized Image Models for RPA Decision Making© Sam Witteveen
Coding Toolsagents

Specialized Image Models for RPA Decision Making

RPA has long struggled with unstructured visual inputs like forms and screenshots, often relying on brittle rule-based systems. This video explores two open models, ImaJev-4B and Jev-Omni, designed specifically to handle these image-based decisions. By focusing on confidence scores and conditional logic, these tools aim to bridge the gap between simple automation and true cognitive processing in document workflows. The approach moves beyond generic vision-language models to offer targeted accuracy for enterprise tasks like form inspection.

Sam Witteveen·Oct 2, 2026

More in Open Source

Meta open sources Muse AI agent SDK for custom hardware© The Verge AI
Open Sourceagents

Meta open sources Muse AI agent SDK for custom hardware

Meta is handing the keys to its Muse AI agent by open-sourcing the software needed to run it on custom hardware. Developers can now hook Muse into ESP32 boards or Raspberry Pi setups, effectively turning workbench scraps into personalized AI terminals. This moves Muse beyond a cloud-only interface into tangible, local devices like E Ink displays or HDMI sticks. It signals a shift toward decentralized, user-owned AI interactions rather than relying solely on proprietary apps.

The Verge AI·Oct 2, 2026
Ai2 open-sources AstaBrief 8B for scientific reports© Hugging Face Blog
Open Sourcewriting

Ai2 open-sources AstaBrief 8B for scientific reports

Allen Institute for AI has released AstaBrief 8B, an open-weight model designed specifically for generating cited scientific literature reviews. Built on Qwen3-8B and trained with supervised fine-tuning and direct preference optimization, it prioritizes speed and grounding over complex multi-step reasoning. The model generates full reports in a single pass, cutting generation time to roughly 51 seconds compared to the 178 seconds required by proprietary alternatives like Claude. This release offers researchers a faster, locally deployable option for synthesizing evidence without relying on external APIs.

Hugging Face Blog·Oct 2, 2026
Open Sourceother

llama.cpp adds IBM ZDNN backend support

IBM is finally getting first-class CI support in llama.cpp with the addition of the ZDNN backend for s390x architecture. This isn't just a minor tweak; it enables efficient inference on mainframe hardware, bridging a gap for enterprise environments that rely on IBM Z systems. While currently limited to build pipelines without automated testing, this signals a serious commitment to supporting non-x86/ARM infrastructure in the local LLM ecosystem. It’s a quiet but necessary expansion for anyone running models on legacy or specialized enterprise silicon.

llama.cpp Releases·Sep 30, 2026