16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

Together AI Enables A/B Testing for LLMs

Together AI Blog·August 17, 2026·high confidence

Why it matters

  • →Enables precise measurement of model performance with real user data.
  • →Simplifies traffic routing and experiment management at the endpoint level.
  • →Facilitates data-driven decisions for optimizing AI model deployments.
Together AI Enables A/B Testing for LLMs
©Together AI Blog

Together AI has launched a new A/B testing feature for large language models in production environments. This tool allows developers to split live traffic between a control model and up to 20 variants, enabling real-world performance comparisons based on user engagement metrics. The platform handles traffic routing at the endpoint level, simplifying the process and avoiding complex client-side configurations. This advancement provides a streamlined approach for teams to assess and improve AI models based on actual user interactions, ensuring more effective deployment decisions.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

TML Unveils Real-Time AI Interaction Models — The Rundown AI1OpenAI Unveils Deployment Simulation for AI Models — OpenAI2OpenAI and Broadcom unveil LLM-optimized chip — OpenAI3vLLM v0.25.0rc2 Release Fixes Key Issues — vLLM Releases4Together AI Enhances Model Inference Configuration — Together AI Blog5Together AI Introduces Autoscaling for LLM Inference — Together AI Blog6LFM2.5-2.6B: Efficient AI Agents for Local Deployment — Hugging Face Blog7Startups Innovate Beyond Transformers in LLMs — MIT Technology Review AI8Together AI Enables A/B Testing for LLMsQwen3.8-27B Model Explored and Optimized — Sam Witteveen9Granite 4.2 LLMs Introduced by IBM and Hugging Face — Hugging Face Blog10Google's ToolGrad Enhances Tool-Use Dataset Generation — Google Research Blog11Top AI Leaders Call for Slowdown in LLM Development — MIT Technology Review AI12llama.cpp v0.6.0 adds decision model API and Apple Silicon speedups — llama.cpp Releases13May 12You are hereOct 6

How we got here

  1. 1
    TML Unveils Real-Time AI Interaction Models

    The Rundown AI · May 12, 2026 · Related

  2. 2
    OpenAI Unveils Deployment Simulation for AI Models

    OpenAI · June 16, 2026 · Background

  3. 3
    OpenAI and Broadcom unveil LLM-optimized chip

    OpenAI · June 24, 2026 · Related

  4. 4
    vLLM v0.25.0rc2 Release Fixes Key Issues

    vLLM Releases · July 9, 2026 · Related

  5. 5
    Together AI Enhances Model Inference Configuration

    Together AI Blog · July 29, 2026 · Same story

  6. 6
    Together AI Introduces Autoscaling for LLM Inference

    Together AI Blog · July 31, 2026 · Same story

  7. 7
    LFM2.5-2.6B: Efficient AI Agents for Local Deployment

    Hugging Face Blog · August 4, 2026 · Related

  8. 8
    Startups Innovate Beyond Transformers in LLMs

    MIT Technology Review AI · August 10, 2026 · Background

What happened next

  1. 9
    Qwen3.8-27B Model Explored and Optimized

    Sam Witteveen · August 18, 2026 · Background

  2. 10
    Granite 4.2 LLMs Introduced by IBM and Hugging Face

    Hugging Face Blog · August 25, 2026 · Related

  3. 11
    Google's ToolGrad Enhances Tool-Use Dataset Generation

    Google Research Blog · September 10, 2026 · Background

  4. 12
    Top AI Leaders Call for Slowdown in LLM Development

    MIT Technology Review AI · September 14, 2026 · Background

  5. 13
    llama.cpp v0.6.0 adds decision model API and Apple Silicon speedups

    llama.cpp Releases · October 6, 2026 · Background

More from Together AI Blog

Together Link bridges open models to coding agents© Together AI Blog
Coding Toolscoding

Together Link bridges open models to coding agents

Together AI is solving the cost explosion of coding agents by routing tasks through its serverless platform instead of burning cash on premium closed models. The tool integrates directly into Claude Code, Codex, and OpenCode, allowing teams to swap expensive Opus calls for cheaper alternatives like GLM 5.3 or Kimi K3 without changing their workflow. It’s not just a proxy; the router dynamically assigns hard problems to frontier-capable open weights while offloading routine fixes to low-cost flash models. This turns open inference into a drop-in cost-saving layer for existing engineering stacks, making high-performance coding significantly cheaper.

Together AI Blog·Oct 5, 2026

More in Models & Labs

Microsoft launches Nvidia Spark AI PCs with agent-ready Windows 11© TechCrunch AI
Models & Labsagents

Microsoft launches Nvidia Spark AI PCs with agent-ready Windows 11

Microsoft is finally shipping the hardware promised alongside Nvidia’s RTX Spark chip, targeting developers who want to run local AI agents without cloud dependency. The Surface Laptop Ultra and Dev Box feature unified memory and specialized cooling designed specifically for on-device model orchestration. More importantly, Windows 11 now includes Execution Containers, a sandboxing layer that lets multiple AI agents operate safely alongside user data. This moves the platform from just hosting models to actively managing agent context and actions locally. It signals a shift where the OS itself becomes the runtime environment for autonomous software.

TechCrunch AI·Oct 7, 2026
Claude Haiku 5.5 arrives in GitHub Copilot© GitHub Changelog
Models & Labscoding

Claude Haiku 5.5 arrives in GitHub Copilot

Anthropic’s latest lightweight model is now live inside GitHub Copilot, targeting high-volume coding tasks like subagents and terminal work. Early benchmarks suggest it matches Sonnet 5 on many coding challenges while consuming significantly fewer tokens and steps. This availability across VS Code, JetBrains, and mobile apps gives developers a faster, cheaper option for routine edits without sacrificing quality. The gradual rollout means most users will see it soon, with admin controls allowing enterprises to manage access via model policies.

GitHub Changelog·Oct 7, 2026
Microsoft Surface Laptop Ultra launches with Nvidia RTX Spark© The Verge AI
Models & Labsother

Microsoft Surface Laptop Ultra launches with Nvidia RTX Spark

The Surface Laptop Ultra marks the commercial debut of Nvidia’s RTX Spark, an Arm-based chip designed to bring serious local AI inference to Windows laptops. Starting at $2,599, this device signals a shift toward premium hardware capable of running large models locally, moving beyond cloud dependency for enterprise and creative workflows. Microsoft pairs this with 'Hybrid Intelligence' features in Copilot, allowing the agent to access local files and take OS-level actions like filing taxes or managing emails. This isn't just a new laptop; it's the first concrete proof that Arm-based PC silicon can handle the thermal and memory demands of 120B+ parameter models on-device.

The Verge AI·Oct 7, 2026