16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

LFM2.5-Encoders Boost Long-Context Inference on CPU

Hugging Face Blog·July 28, 2026·high confidence

Why it matters

  • →LFM2.5-Encoders offer significant speed improvements for long-context tasks on CPU.
  • →They provide a cost-effective alternative to larger models without sacrificing performance.
  • →The open-source nature allows for easy adaptation and fine-tuning for specific applications.
LFM2.5-Encoders Boost Long-Context Inference on CPU
©Hugging Face Blog

Hugging Face has introduced LFM2.5-Encoders, designed to enhance long-context inference, especially on CPU. These models are faster than ModernBERT-base, handling up to 8,192-token contexts efficiently. They are particularly suited for tasks like classification and routing, offering a cost-effective solution without sacrificing speed. The encoders are open-source and ready for developers to fine-tune for various applications. This development highlights a move towards more efficient NLP models that maintain high performance on standard hardware.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

OpenAI launches GPT-5.5 Instant for ChatGPT — TechCrunch AI1DeepSeek-V4 Tackles Million-Token Context Challenge — Together AI Blog2OpenAI Unveils GPT-5.5 Instant for ChatGPT — The AI Advantage3llama.cpp b9271 Release Enhances Efficiency — llama.cpp Releases4vLLM v0.23.0 Release Enhances Model Support — vLLM Releases5Subquadratic claims breakthrough in LLM efficiency — MIT Technology Review AI6OpenAI and Broadcom unveil LLM-optimized chip — OpenAI7vLLM v0.25.0rc2 Release Fixes Key Issues — vLLM Releases8LFM2.5-Encoders Boost Long-Context Inference on CPURecall Bottleneck in LLMs: New Google Research Insights — Google Research Blog9LFM2.5-VL-3B Enhances Vision-Language Capabilities — Hugging Face Blog10Hugging Face tokenizers v1 delivers massive speed gains — Hugging Face Blog11Together AI: Fine-tune Jev-like classifier for $17 — Together AI Blog12llama.cpp fixes LFM2 audio transcription accuracy — llama.cpp Releases13May 5You are hereSep 26

How we got here

  1. 1
    OpenAI launches GPT-5.5 Instant for ChatGPT

    TechCrunch AI · May 5, 2026 · Background

  2. 2
    DeepSeek-V4 Tackles Million-Token Context Challenge

    Together AI Blog · May 8, 2026 · Background

  3. 3
    OpenAI Unveils GPT-5.5 Instant for ChatGPT

    The AI Advantage · May 8, 2026 · Background

  4. 4
    llama.cpp b9271 Release Enhances Efficiency

    llama.cpp Releases · May 22, 2026 · Background

  5. 5
    vLLM v0.23.0 Release Enhances Model Support

    vLLM Releases · June 14, 2026 · Background

  6. 6
    Subquadratic claims breakthrough in LLM efficiency

    MIT Technology Review AI · June 19, 2026 · Background

  7. 7
    OpenAI and Broadcom unveil LLM-optimized chip

    OpenAI · June 24, 2026 · Background

  8. 8
    vLLM v0.25.0rc2 Release Fixes Key Issues

    vLLM Releases · July 9, 2026 · Background

What happened next

  1. 9
    Recall Bottleneck in LLMs: New Google Research Insights

    Google Research Blog · August 12, 2026 · Related

  2. 10
    LFM2.5-VL-3B Enhances Vision-Language Capabilities

    Hugging Face Blog · August 12, 2026 · Same story

  3. 11
    Hugging Face tokenizers v1 delivers massive speed gains

    Hugging Face Blog · September 21, 2026 · Related

  4. 12
    Together AI: Fine-tune Jev-like classifier for $17

    Together AI Blog · September 23, 2026 · Background

  5. 13
    llama.cpp fixes LFM2 audio transcription accuracy

    llama.cpp Releases · September 26, 2026 · Related

More from Hugging Face Blog

NVIDIA Kumo Tabular beats GBDT on benchmarks© Hugging Face Blog
Models & Labsother

NVIDIA Kumo Tabular beats GBDT on benchmarks

NVIDIA has released Kumo Tabular, an open foundation model that challenges the two-decade dominance of gradient-boosted trees in enterprise data. By training exclusively on synthetic tables generated via causal models, it achieves zero-shot classification and regression without feature engineering or hyperparameter tuning. It currently tops the TabArena leaderboard, offering a significant speed advantage over competitors like LimiX-2 while maintaining state-of-the-art accuracy. This shifts tabular ML from manual pipeline construction to direct inference, fundamentally changing how structured data is processed.

Hugging Face Blog·Sep 29, 2026
ProvenanceGuard verifies MCP agent source attribution© Hugging Face Blog
Researchresearch

ProvenanceGuard verifies MCP agent source attribution

Most fact-checkers for AI agents only check if a claim is true in the evidence pool, ignoring where it came from. ProvenanceGuard fixes this by tracking source identity through every step of verification, catching cases where a true fact is wrongly attributed to the wrong tool or document. In medical agent tests, it caught 138 out of 139 incorrect attributions that standard verifiers missed, proving that provenance matters as much as truth in multi-tool environments. This shifts the focus from simple RAG retrieval to rigorous source-aware auditing for high-stakes applications.

Hugging Face Blog·Sep 29, 2026
Holo4: Generalist Agentic Models for GUI and Code© Hugging Face Blog
Models & Labsagents

Holo4: Generalist Agentic Models for GUI and Code

H company released Holo4, a new series of agentic models designed to navigate software through any interface—GUIs, code, MCP, or APIs. The 27B dense and 35B-A3B MoE variants significantly outperform their Qwen bases on complex workflows like building 3D models in FreeCAD or coding games in Godot. While trailing Opus 5 on OSWorld benchmarks, Holo4 achieves this with orders of magnitude fewer parameters and lower cost. This release marks a shift toward unified agents that don't need separate models for different interaction modes.

Hugging Face Blog·Sep 28, 2026

More in Models & Labs

OpenAI launches collaborative office suite features© TechCrunch AI
Models & Labsproductivity

OpenAI launches collaborative office suite features

OpenAI is directly challenging Microsoft’s dominance by embedding a full office suite into ChatGPT. The new 'Space' feature acts as a shared workspace where users and AI agents collaborate in real-time, while 'Pages' serves as an agent-native word processor. Collaborative slides allow teams to co-edit presentations generated from conversation. This marks a strategic pivot from pure chatbot utility to comprehensive workplace infrastructure, forcing incumbents to defend their core business models against an AI-native competitor.

TechCrunch AI·Sep 29, 2026
OpenAI delays GPT-6.1 Astra over safety risks© Wes Roth
Investment · $42 billion net loss
Models & Labsmodels

OpenAI delays GPT-6.1 Astra over safety risks

OpenAI has paused the release of GPT-6.1 Astra, citing significant safety concerns that outweigh its performance gains. This decision underscores a growing tension in the industry: as models become more capable, they also become harder to constrain within safe operational boundaries. While Anthropic pushes forward with Claude Sonnet 5.5 and reports massive financial losses ahead of an IPO, OpenAI is choosing caution over speed. The delay signals that safety evaluations are becoming a primary bottleneck for next-generation model releases, potentially shifting the competitive landscape toward more conservative development cycles. By halting Astra, OpenAI acknowledges that raw capability alone no longer guarantees a viable product launch. This move forces competitors to reconsider their own release timelines and safety protocols. The industry now watches closely to see if this caution becomes the new standard for frontier AI development.

Wes Roth·Sep 29, 2026
OpenAI expands ChatGPT plugins with app-like interfaces© TechCrunch AI
Models & Labsproductivity

OpenAI expands ChatGPT plugins with app-like interfaces

OpenAI is shifting ChatGPT from a chat interface to an application platform by introducing dedicated sidebar panels and interactive tools for third-party developers. This move transforms static integrations into persistent, app-like experiences where users can view files and manage data without leaving the conversation. By supporting the MCP Events specification, OpenAI enables plugins to trigger automations based on external events, bridging the gap between conversational AI and workflow execution. The redesign of the plugin directory and creator tools aims to lower barriers for developers while giving users finer control over permissions. This effectively positions ChatGPT as a central hub for enterprise productivity rather than just a search or writing assistant.

TechCrunch AI·Sep 29, 2026