llama.cpp has released version b11380, expanding hardware support to include ROCm 10.0 for AMD GPUs on Linux and Windows, as well as native binaries for Linux ARM64 devices like Snapdragon laptops. The update also adds CUDA 13.4 builds across multiple platforms and includes OpenVINO and SYCL optimizations. Notably, macOS KleidiAI support is currently disabled in this specific build. This release significantly broadens the range of consumer and enterprise hardware capable of running local large language models efficiently.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 28, 2026 · Same story
llama.cpp Releases · October 4, 2026 · Same story
This release tackles the notorious memory hunger of long-context inference for Qwen4-exp models by halving indexer score memory. The optimization works by computing head scores in place rather than materializing separate tensors, a change that significantly reduces VRAM pressure during heavy workloads. Beyond memory efficiency, b11372 expands hardware coverage with CUDA 13 support and Vulkan tiling for the lightning indexer. It also adds ROCm 10.0 binaries, keeping AMD users in step with NVIDIA's latest driver ecosystem. The result is a leaner runtime that handles extended contexts without hitting out-of-memory errors as quickly.
This release significantly tightens llama.cpp’s integration with Intel’s OpenVINO backend, specifically targeting Mixture of Experts (MoE) models like Qwen3.5 and Gemma-4. By fusing MoE routing and GDN normalization operations, prefill throughput on Arc GPUs jumps from 66 to over 1,600 tokens per second, effectively removing a major bottleneck for local inference on Intel hardware. The update also fixes critical stateful execution bugs that previously caused crashes or incorrect axis handling during decoding. This makes OpenVINO a far more viable option for running complex MoE architectures on consumer-grade Intel GPUs without relying on NVIDIA CUDA.
This release resolves a critical crash in Mamba SSM inference when batch cells aren't contiguous. By gathering recurrent states into a single reserve that covers every split, the engine avoids illegal graph reallocations under strict scheduling modes. It’s a quiet but essential fix for anyone running non-standard sequence lengths or complex batching logic with stateful models.
© The AI Daily BriefMeta's stock price surged following positive market reaction to its new Muse model capabilities.
© The Verge AIGoogle is bringing real-time audio scene description to Android via Gemini Live, directly challenging Apple’s VoiceOver Live Recognition. This feature targets users with low vision by providing immediate audio cues and follow-up Q&A capabilities for physical objects. It integrates deeply into the accessibility ecosystem through TalkBack, moving beyond simple text reading to contextual environmental awareness. The move signals a shift toward multimodal AI as a standard utility for daily navigation rather than just a novelty.
© TechCrunch AIAmazon’s Strands Decider 2B joins the growing wave of decision models designed to replace heavy LLMs for simple routing tasks. Built on Qwen3.5-2B, it outputs calibrated choices with confidence scores rather than generating text, offering a cheaper, faster alternative for agentic workflows. The release signals AWS’s push into specialized agent infrastructure, aiming to solve the latency and cost bottlenecks of general-purpose models. While TypeSafe’s Jev pioneered this space, Amazon’s entry brings enterprise-grade credibility and open-source accessibility to a niche that is rapidly filling with experimental clones.