
Liquid AI has released open-weight versions of its d1 decision models, specifically designed for single-pass inference rather than token generation. The flagship d1-3B model, built on the LFM2.5-VL-3B backbone, supports text and image inputs and ranks as the best-performing model under 10B parameters on the Decision Index v0.2.1. A smaller variant, d1-omni-600M, handles multimodal inputs including audio with only 600 million parameters. Benchmarks show d1-3B completing inference in 16ms on an NVIDIA Jetson AGX Thor and under 10ms on RTX 4090 GPUs. Both models are available on Hugging Face for immediate integration into edge computing workflows.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
TechCrunch AI · July 15, 2026 · Background
WIRED AI · July 15, 2026 · Related
Hugging Face Blog · August 4, 2026 · Related
MIT Technology Review AI · August 10, 2026 · Related
NVIDIA Blog · August 24, 2026 · Background
Sam Witteveen · August 30, 2026 · Background
Sam Witteveen · September 18, 2026 · Background
Hugging Face Blog · September 24, 2026 · Same story
© Hugging Face BlogAllen Institute for AI solved the 'tragedy of the commons' in its H100 and B200 clusters by abandoning priority queues for a budget-based system. Researchers now spend allocated GPU time rather than hoarding it, turning resource allocation into a transparent administrative process. This shift eliminates squatting and priority inflation while keeping occupancy high through hierarchical fair-share scheduling. It proves that treating compute as a financial asset works better than treating it as a shared utility.
Hugging Face’s ML-Intern agent proves that autonomous model training is no longer theoretical. By handling dataset curation, hyperparameter tuning, and cost management with a single prompt, it produced six distinct fine-tuned models in days for under $50 total. This shifts the barrier from engineering complexity to prompt precision, allowing developers to iterate on specialized capabilities like camera-angle LoRAs or domain-specific vision without manual infrastructure overhead. The real shift is the democratization of custom model creation, turning what used to be a week-long engineering sprint into a low-cost, automated workflow.
© Hugging Face BlogTII’s Falcon-ASR finally gives the UAE a homegrown speech model that actually understands local accents. With a 20.92% WER on Arabic benchmarks and beating Qwen3-Omni by over four points on internal Emirati tests, it solves the dialect gap that plagues most multilingual ASR systems. The single-weight architecture handles five languages without flags, making deployment trivial for developers who previously had to juggle separate models or accept poor accuracy on Gulf speech.
This release quietly cements llama.cpp as the universal inference runtime by adding default support for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD GPU users finally get parity with NVIDIA's latest driver stack without manual configuration, while Apple Silicon KleidiAI builds are temporarily disabled to resolve stability issues. The inclusion of Snapdragon NPU support on Linux signals a serious push into edge AI hardware beyond just x86 and ARM CPUs. It is less about new features and more about ensuring the toolchain keeps pace with the rapidly evolving GPU landscape.
This release quietly closes the hardware gap for local inference by adding default builds for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD users finally get parity with NVIDIA in the binary distribution, while CUDA 13 support future-proofs setups on newer drivers. The inclusion of Snapdragon and OpenVINO binaries further broadens the hardware surface area without requiring custom compilation. It is a pragmatic update that makes llama.cpp the most accessible runtime for diverse local AI hardware.
© Lev SelectorMistral releases Large 4 'Le Chonk' while Anthropic launches Claude Haiku 5.5, continuing the trend of cheaper, faster frontier models.