The b10419 release of llama.cpp introduces several enhancements to the OpenVINO backend, aimed at improving memory management and operational accuracy. Key updates include fixing mixed-rank broadcast issues in GPU plugins, which significantly improved perplexity scores. Additionally, a new mode reduces memory usage by releasing host weight buffers after model compilation, effectively halving the steady-state RSS. These updates enhance the efficiency and reliability of the OpenVINO backend, particularly for GPU-based inference tasks.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · June 28, 2026 · Same story
llama.cpp Releases · September 16, 2026 · Same story
This release tackles the notorious memory hunger of long-context inference for Qwen4-exp models by halving indexer score memory. The optimization works by computing head scores in place rather than materializing separate tensors, a change that significantly reduces VRAM pressure during heavy workloads. Beyond memory efficiency, b11372 expands hardware coverage with CUDA 13 support and Vulkan tiling for the lightning indexer. It also adds ROCm 10.0 binaries, keeping AMD users in step with NVIDIA's latest driver ecosystem. The result is a leaner runtime that handles extended contexts without hitting out-of-memory errors as quickly.
This release significantly tightens llama.cpp’s integration with Intel’s OpenVINO backend, specifically targeting Mixture of Experts (MoE) models like Qwen3.5 and Gemma-4. By fusing MoE routing and GDN normalization operations, prefill throughput on Arc GPUs jumps from 66 to over 1,600 tokens per second, effectively removing a major bottleneck for local inference on Intel hardware. The update also fixes critical stateful execution bugs that previously caused crashes or incorrect axis handling during decoding. This makes OpenVINO a far more viable option for running complex MoE architectures on consumer-grade Intel GPUs without relying on NVIDIA CUDA.
Google is bringing real-time audio scene description to Android via Gemini Live, directly challenging Apple’s VoiceOver Live Recognition. This feature targets users with low vision by providing immediate audio cues and follow-up Q&A capabilities for physical objects. It integrates deeply into the accessibility ecosystem through TalkBack, moving beyond simple text reading to contextual environmental awareness. The move signals a shift toward multimodal AI as a standard utility for daily navigation rather than just a novelty.