16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Research
Research

Hugging Face Olmo-core 3 scales MoE training to trillions of parameters

Hugging Face Blog·October 1, 2026·high confidence

Why it matters

  • →Provides an open-source alternative to proprietary stacks like Megatron-Core for training trillion-parameter MoE models.
  • →Achieves significant throughput gains (2.7x) by optimizing GPU resident expert routing and reducing communication overhead.
  • →Enables academic and smaller labs to experiment with massive sparse architectures that were previously cost-prohibitive or technically inaccessible.
Hugging Face Olmo-core 3 scales MoE training to trillions of parameters
©Hugging Face Blog

Hugging Face has released Olmo-core 3, an open-source training framework designed for large-scale Mixture-of-Experts (MoE) language models. The update replaces the previous Fully Sharded Data Parallelism (FSDP) approach with Distributed Data Parallelism (DDP), keeping expert weights on GPUs to reduce data movement and improve throughput by approximately 2.7x on NVIDIA B300 hardware. Benchmarks demonstrate the system's ability to handle models exceeding one trillion total parameters, with specific tests reaching 858 TFLOP/s/GPU on 512 GPUs. The release includes a technical report detailing optimizations like MXFP8 precision support and insights into routing efficiency, positioning Olmo-core 3 as a scalable foundation for the next generation of open MoE models.

Read original

The story around this

TopicNvidia Acquires Hugging Face

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

NVIDIA Blackwell Dominates MLPerf Training 6.0 — NVIDIA Blog1ParallelKernelBench Reveals Gaps in Multi-GPU Kernel Generation — Together AI Blog2NVIDIA and Hugging Face Enhance Robotics with New Models — NVIDIA Blog3NVIDIA and Hugging Face Enhance Diffusion Model Training — Hugging Face Blog4Together AI Introduces Autoscaling for LLM Inference — Together AI Blog5Hugging Face Launches 200+ WebGPU Kernels — Hugging Face Blog6Base Labs partners with Hugging Face on open-weight AI safety — TechCrunch AI7llama.cpp fixes MoE dispatch inefficiency on Vulkan — llama.cpp Releases8Hugging Face Olmo-core 3 scales MoE training to trillions of parametersllama.cpp optimizes NVIDIA V100 inference performance — llama.cpp Releases9Jun 16You are hereOct 2

How we got here

  1. 1
    NVIDIA Blackwell Dominates MLPerf Training 6.0

    NVIDIA Blog · June 16, 2026 · Background

  2. 2
    ParallelKernelBench Reveals Gaps in Multi-GPU Kernel Generation

    Together AI Blog · June 23, 2026 · Background

  3. 3
    NVIDIA and Hugging Face Enhance Robotics with New Models

    NVIDIA Blog · July 7, 2026 · Background

  4. 4
    NVIDIA and Hugging Face Enhance Diffusion Model Training

    Hugging Face Blog · July 17, 2026 · Related

  5. 5
    Together AI Introduces Autoscaling for LLM Inference

    Together AI Blog · July 31, 2026 · Background

  6. 6
    Hugging Face Launches 200+ WebGPU Kernels

    Hugging Face Blog · September 1, 2026 · Related

  7. 7
    Base Labs partners with Hugging Face on open-weight AI safety

    TechCrunch AI · September 17, 2026 · Background

  8. 8
    llama.cpp fixes MoE dispatch inefficiency on Vulkan

    llama.cpp Releases · September 30, 2026 · Background

What happened next

  1. 9
    llama.cpp optimizes NVIDIA V100 inference performance

    llama.cpp Releases · October 2, 2026 · Background

More from Hugging Face Blog

ServiceNow AutoSynthData automates agent training data© Hugging Face Blog
Agentsagents

ServiceNow AutoSynthData automates agent training data

ServiceNow CoreAI released AutoSynthData, a pipeline that turns enterprise agent failures into targeted training datasets. By using a stronger teacher model to identify capability gaps and generate feasible, realistic tasks, it solves the bottleneck of creating high-quality synthetic data for specific environments. The system validates every generated task against strict verifiers before adding it to the curriculum, ensuring the model learns from actual weaknesses rather than noise. This approach shifts agent training from manual curation to automated, continuous improvement loops grounded in real-world constraints.

Hugging Face Blog·Oct 2, 2026
Hugging Face launches objective Open TTS Leaderboard© Hugging Face Blog
Models & Labsother

Hugging Face launches objective Open TTS Leaderboard

The open-source text-to-speech ecosystem is drowning in models but starving for reliable benchmarks. Hugging Face’s new Open TTS Leaderboard solves this by replacing slow, subjective human voting with automated metrics: word error rate for intelligibility, speaker similarity for voice cloning fidelity, and latency measurements for streaming viability. This shift allows developers to instantly compare 8,000+ models on objective performance rather than waiting weeks for community votes. It finally gives open-weight TTS models the standardized visibility they’ve lacked compared to commercial APIs.

Hugging Face Blog·Sep 30, 2026
NVIDIA Kumo Tabular beats GBDT on benchmarks© Hugging Face Blog
Models & Labsother

NVIDIA Kumo Tabular beats GBDT on benchmarks

NVIDIA has released Kumo Tabular, an open foundation model that challenges the two-decade dominance of gradient-boosted trees in enterprise data. By training exclusively on synthetic tables generated via causal models, it achieves zero-shot classification and regression without feature engineering or hyperparameter tuning. It currently tops the TabArena leaderboard, offering a significant speed advantage over competitors like LimiX-2 while maintaining state-of-the-art accuracy. This shifts tabular ML from manual pipeline construction to direct inference, fundamentally changing how structured data is processed.

Hugging Face Blog·Sep 29, 2026

More in Research

OpenAI Publishes Research on AI-Driven Intelligence Explosions© AI Explained
Researchresearch

OpenAI Publishes Research on AI-Driven Intelligence Explosions

OpenAI has released a new research paper exploring the potential for AI systems to recursively improve themselves, leading to rapid intelligence growth.

AI Explained·Oct 1, 2026
Graphite study reveals persistent AI writing tells© TechCrunch AI
Researchwriting

Graphite study reveals persistent AI writing tells

A new analysis by Graphite proves that frontier models are failing to shed their robotic DNA. Despite labs claiming natural prose, Claude Opus 5.5 still uses "this matters" 116 times more often than humans, while OpenAI's Astra relies on hedging phrases like "may provide." The study of 13,000 phrases shows that as models eliminate old habits like em-dashes, they simply adopt new ones. This suggests labs cannot fully control the statistical quirks inherent in billions of parameters.

TechCrunch AI·Oct 1, 2026
Ataraxos AI beats top human Stratego players© MIT News AI
Researchresearch

Ataraxos AI beats top human Stratego players

A multi-university team has released Ataraxos, an AI that dominates the imperfect-information game Stratego by a massive margin. Unlike previous attempts that burned millions in compute, this system uses efficient self-play and decision-time planning to outperform top humans with less than one-thirtieth of the training effort. It secured a 39-2 record against world-class players, proving that calculated risk assessment can scale beyond poker-like games. This marks a significant leap in handling hidden information without prohibitive computational costs.

MIT News AI·Sep 30, 2026