16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & Labs

Liquid AI releases DSpark drafter for vision models

Hugging Face Blog·September 24, 2026·high confidence

Why it matters

  • →Speculative decoding drafter for vision-language models reduces latency significantly on edge hardware.
  • →Day-one support for major inference engines (llama.cpp, MLX, SGLang) lowers integration friction.
  • →Minimal parameter overhead (280M) makes it practical to deploy alongside existing VLMs.
Liquid AI releases DSpark drafter for vision models
©Hugging Face Blog

Liquid AI has released LFM2.5-VL-DSpark, a speculative decoding drafter designed to accelerate vision-language models on both edge devices and datacenter GPUs. The open-weight drafter adds just 280M parameters (8.9% overhead) to the target model while achieving up to 3.13x faster decoding on Apple Silicon M5 Max chips and 2.66x speedups on NVIDIA H100s. It ships with immediate integration support for llama.cpp, MLX-VLM, and SGLang, allowing developers to deploy faster VLM inference without rewriting their inference stacks. The release underscores a growing industry shift toward optimizing inference efficiency through specialized drafters rather than solely increasing model scale.

Read original

More from Hugging Face Blog

NVIDIA Warp and MjWarp guide for robotics simulation© Hugging Face Blog
Coding Toolscoding

NVIDIA Warp and MjWarp guide for robotics simulation

NVIDIA is pushing hard to make GPU-accelerated physics the standard for robot learning. This deep dive into MuJoCo Warp (MJWarp) shows how to scale a single SO-101 arm simulation to 2,048 parallel environments on CUDA hardware. The real value isn't faster single-step latency, but massive aggregate throughput for reinforcement learning data collection. By leveraging Warp's kernel compilation and CUDA graph capture, developers can batch thousands of physics steps simultaneously, turning the GPU into a high-throughput experience generator rather than just a fast simulator.

Hugging Face Blog·Sep 23, 2026

More in Models & Labs

Models & Labsother

llama.cpp b11193 adds Snapdragon Hexagon and ROCm 10

This release quietly expands llama.cpp's hardware support to include Qualcomm's Hexagon NPU on Linux arm64, a significant step for local inference on Snapdragon devices. It also updates CUDA builds to version 13.4 and introduces ROCm 10.0 binaries, keeping the project aligned with the latest NVIDIA and AMD driver ecosystems. KleidiAI on Apple Silicon is temporarily disabled in this build, likely due to stability checks rather than a feature rollback. For developers targeting edge AI or diverse GPU stacks, this update ensures broader compatibility without requiring custom compilation.

llama.cpp Releases·Sep 26, 2026
NVIDIA NVFP4 and SoL-Pi Token Efficiency© Lev Selector
Models & Labsmodels

NVIDIA NVFP4 and SoL-Pi Token Efficiency

NVIDIA introduced the NVFP4 4-bit format and SoL-Pi technology, which uses 2x fewer tokens for improved efficiency.

Lev Selector·Sep 25, 2026
Claude Opus 5.5 and GPT-6 Sol Released© Lev Selector
Models & Labsmodels

Claude Opus 5.5 and GPT-6 Sol Released

Anthropic released Claude Opus 5.5 on September 22, while OpenAI launched GPT-6 Sol and Luna variants, advancing frontier model capabilities.

Lev Selector·Sep 25, 2026