llama.cpp version b11375 addresses a memory reservation issue in Mamba SSM inference that caused crashes when batch cells were not contiguous. The update modifies the graph construction to gather recurrent states into a single reserve, preventing illegal reallocations under GGML_SCHED_NO_REALLOC. This fix ensures stable execution for complex batching scenarios involving stateful models.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 20, 2026 · Related
llama.cpp Releases · October 2, 2026 · Same story
This release tackles the notorious memory hunger of long-context inference for Qwen4-exp models by halving indexer score memory. The optimization works by computing head scores in place rather than materializing separate tensors, a change that significantly reduces VRAM pressure during heavy workloads. Beyond memory efficiency, b11372 expands hardware coverage with CUDA 13 support and Vulkan tiling for the lightning indexer. It also adds ROCm 10.0 binaries, keeping AMD users in step with NVIDIA's latest driver ecosystem. The result is a leaner runtime that handles extended contexts without hitting out-of-memory errors as quickly.
This release significantly tightens llama.cpp’s integration with Intel’s OpenVINO backend, specifically targeting Mixture of Experts (MoE) models like Qwen3.5 and Gemma-4. By fusing MoE routing and GDN normalization operations, prefill throughput on Arc GPUs jumps from 66 to over 1,600 tokens per second, effectively removing a major bottleneck for local inference on Intel hardware. The update also fixes critical stateful execution bugs that previously caused crashes or incorrect axis handling during decoding. This makes OpenVINO a far more viable option for running complex MoE architectures on consumer-grade Intel GPUs without relying on NVIDIA CUDA.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with CUDA. Equally notable is the new Snapdragon binary for Linux, which unlocks local LLM execution on ARM-based mobile chips via CPU, Adreno GPU, and Hexagon NPU paths. While KleidiAI builds are temporarily disabled, the expansion to Windows ROCm and mobile silicon signals a strategic shift toward hardware-agnostic accessibility.
This release stabilizes the core session management of Claude Code, specifically targeting the fragile state of resumed conversations where context or thinking traces were previously lost. It also patches critical reliability issues in the Model Context Protocol (MCP) integration, ensuring tool calls don't duplicate or hang indefinitely when remote servers misbehave. The addition of $.ui.selection() for mods and better GitHub CLI handling in cloud sessions shows a focus on developer workflow friction rather than new capabilities. These are necessary maintenance updates that make the tool more robust for heavy daily use.
This release is a classic maintenance patch for Claude Code, focusing on stabilizing the terminal interface and tightening security rules. It fixes critical bugs where deny/ask rules were bypassed in nested shell commands or via symlinks, ensuring sandbox policies actually hold. The update also resolves numerous UI freezes caused by malformed HTML tags and plugin rendering errors, making the agent feel less brittle during complex coding sessions.
© GitHub ChangelogGitHub finally exposes Copilot code review to external automation via REST and GraphQL APIs, moving it from a manual UI action to an integrable pipeline step. This allows developers to trigger reviews directly from scripts or internal tools rather than relying on the web interface. Simultaneously, the default effort level shifts to Balanced, striking a middle ground between speed and depth for most repositories. While Lite remains available for those prioritizing raw throughput, the API access is the real win here, enabling true CI/CD integration for automated code quality checks.