16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

Hugging Face launches objective Open TTS Leaderboard

Hugging Face Blog·September 30, 2026·high confidence

Why it matters

  • →Provides standardized, objective benchmarks for the fragmented open-source TTS market.
  • →Enables rapid comparison of multilingual performance and voice cloning fidelity without human voting delays.
  • →Highlights open-weight models that are often underrepresented in commercial API-focused leaderboards.
Hugging Face launches objective Open TTS Leaderboard
©Hugging Face Blog

Hugging Face has launched the Open TTS Leaderboard, a new evaluation platform designed to standardize benchmarking for multilingual text-to-speech and voice cloning models. The leaderboard utilizes objective metrics including word error rate (WER), speaker similarity scores via WavLM embeddings, and real-time factors on H200 GPUs to rank over 8,000 open-source models. This approach addresses the scalability limitations of previous arena-style leaderboards, which relied on slow human preference voting and heavily favored commercial API providers. The platform also includes a 'Listen' tab for subjective verification and streaming latency comparisons, with evaluation scripts planned for open-source release.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

Hugging Face Launches FFASR Leaderboard for ASR Models — Hugging Face Blog1OpenAI Unveils GPT-Live for Voice Interaction — OpenAI2OpenAI Unveils New Voice Models for ChatGPT — TechCrunch AI3Hugging Face Launches Real World VoiceEQ Benchmark — Hugging Face Blog4llama.cpp b10311 Release Enhances TTS Generation — llama.cpp Releases5BreezeTTS2: Real-time local voice synthesis — Sam Witteveen6Base Labs partners with Hugging Face on open-weight AI safety — TechCrunch AI7Google Gemini 3.8 Flash TTS with Voice Cloning — Sam Witteveen8Hugging Face launches objective Open TTS LeaderboardJun 24You are here

How we got here

  1. 1
    Hugging Face Launches FFASR Leaderboard for ASR Models

    Hugging Face Blog · June 24, 2026 · Same story

  2. 2
    OpenAI Unveils GPT-Live for Voice Interaction

    OpenAI · July 8, 2026 · Background

  3. 3
    OpenAI Unveils New Voice Models for ChatGPT

    TechCrunch AI · July 8, 2026 · Background

  4. 4
    Hugging Face Launches Real World VoiceEQ Benchmark

    Hugging Face Blog · July 15, 2026 · Same story

  5. 5
    llama.cpp b10311 Release Enhances TTS Generation

    llama.cpp Releases · August 8, 2026 · Background

  6. 6
    BreezeTTS2: Real-time local voice synthesis

    Sam Witteveen · August 31, 2026 · Related

  7. 7
    Base Labs partners with Hugging Face on open-weight AI safety

    TechCrunch AI · September 17, 2026 · Background

  8. 8
    Google Gemini 3.8 Flash TTS with Voice Cloning

    Sam Witteveen · September 24, 2026 · Related

More from Hugging Face Blog

ServiceNow AutoSynthData automates agent training data© Hugging Face Blog
Agentsagents

ServiceNow AutoSynthData automates agent training data

ServiceNow CoreAI released AutoSynthData, a pipeline that turns enterprise agent failures into targeted training datasets. By using a stronger teacher model to identify capability gaps and generate feasible, realistic tasks, it solves the bottleneck of creating high-quality synthetic data for specific environments. The system validates every generated task against strict verifiers before adding it to the curriculum, ensuring the model learns from actual weaknesses rather than noise. This approach shifts agent training from manual curation to automated, continuous improvement loops grounded in real-world constraints.

Hugging Face Blog·Oct 2, 2026
Hugging Face Olmo-core 3 scales MoE training to trillions of parameters© Hugging Face Blog
Researchmodels

Hugging Face Olmo-core 3 scales MoE training to trillions of parameters

Hugging Face’s Olmo-core 3 rewrites the rules for open-source Mixture-of-Experts training by shifting from FSDP to DDP, keeping experts resident on GPUs to slash communication overhead. This architectural pivot yields a 2.7x throughput jump on NVIDIA B300s and enables stable training of models with over one trillion total parameters while keeping active compute fixed. By integrating MXFP8 precision and optimized routing, the framework closes the efficiency gap with proprietary stacks like Megatron-Core. Researchers now have an open, battle-tested infrastructure to build massive sparse models without relying on closed-source enterprise tools.

Hugging Face Blog·Oct 1, 2026
NVIDIA Kumo Tabular beats GBDT on benchmarks© Hugging Face Blog
Models & Labsother

NVIDIA Kumo Tabular beats GBDT on benchmarks

NVIDIA has released Kumo Tabular, an open foundation model that challenges the two-decade dominance of gradient-boosted trees in enterprise data. By training exclusively on synthetic tables generated via causal models, it achieves zero-shot classification and regression without feature engineering or hyperparameter tuning. It currently tops the TabArena leaderboard, offering a significant speed advantage over competitors like LimiX-2 while maintaining state-of-the-art accuracy. This shifts tabular ML from manual pipeline construction to direct inference, fundamentally changing how structured data is processed.

Hugging Face Blog·Sep 29, 2026

More in Models & Labs

Google launches Guided Vision in Gemini Live© The Verge AI
Models & Labsother

Google launches Guided Vision in Gemini Live

Google is bringing real-time audio scene description to Android via Gemini Live, directly challenging Apple’s VoiceOver Live Recognition. This feature targets users with low vision by providing immediate audio cues and follow-up Q&A capabilities for physical objects. It integrates deeply into the accessibility ecosystem through TalkBack, moving beyond simple text reading to contextual environmental awareness. The move signals a shift toward multimodal AI as a standard utility for daily navigation rather than just a novelty.

The Verge AI·Oct 1, 2026
AWS releases open-source Strands Decider 2B© TechCrunch AI
Models & Labsagents

AWS releases open-source Strands Decider 2B

Amazon’s Strands Decider 2B joins the growing wave of decision models designed to replace heavy LLMs for simple routing tasks. Built on Qwen3.5-2B, it outputs calibrated choices with confidence scores rather than generating text, offering a cheaper, faster alternative for agentic workflows. The release signals AWS’s push into specialized agent infrastructure, aiming to solve the latency and cost bottlenecks of general-purpose models. While TypeSafe’s Jev pioneered this space, Amazon’s entry brings enterprise-grade credibility and open-source accessibility to a niche that is rapidly filling with experimental clones.

TechCrunch AI·Oct 1, 2026
Google Announces Gemini 4 Argon with 1M Token Output© Sam Witteveen
Models & Labsmodels

Google Announces Gemini 4 Argon with 1M Token Output

Google is pushing the boundaries of context windows with Gemini 4 Argon, a new model capable of generating up to one million tokens in a single response. This isn't just about reading long documents; it's designed for complex agentic workflows where the AI must produce extensive codebases or detailed reports without truncation. Early benchmarks suggest it aims to reclaim top-tier intelligence status against competitors like GPT-6, specifically targeting tasks that require sustained reasoning and massive output generation. The shift from 64K caps to a million-token horizon fundamentally changes how developers might architect multi-step autonomous systems.

Sam Witteveen·Oct 1, 2026