llama.cpp has released build b11333, expanding hardware support for local inference across AMD, Qualcomm, and Intel platforms. Key additions include ROCm 10.0 binaries for Linux and Windows, CUDA 13.4 libraries (13.4) for x64 and arm64 architectures, and native Snapdragon support on Linux arm64 utilizing CPU, Adreno GPU, and Hexagon NPU acceleration. The release also maintains OpenVINO and SYCL support while disabling KleidiAI on macOS Apple Silicon and openEuler builds in this specific iteration. This update broadens the accessibility of llama.cpp beyond NVIDIA-centric environments, providing pre-compiled binaries for a wider range of consumer and enterprise hardware configurations.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 26, 2026 · Same story
llama.cpp Releases · October 2, 2026 · Same story
This release quietly expands llama.cpp's hardware reach with two major additions: ROCm 10.0 for AMD GPUs and native support for Linux arm64 Snapdragon devices. The inclusion of ROCm 10 is significant, as it brings AMD users closer to parity with CUDA in terms of supported versions, reducing the friction for local inference on non-NVIDIA hardware. Meanwhile, Snapdragon support opens up a new class of mobile AI acceleration, allowing developers to leverage Adreno GPUs and Hexagon NPUs directly. While Apple Silicon builds have KleidiAI disabled by default, the core value here is the broadening of accessible compute backends without requiring complex custom compilation.
This release addresses a memory management issue in the Metal backend by properly releasing temporary private transfer buffers. While not a feature upgrade, it stabilizes performance on Apple Silicon devices where previous builds might have suffered from memory pressure or fragmentation during inference. The update also includes standard platform binaries for CUDA 13 and ROCm 10.0, ensuring compatibility with the latest NVIDIA and AMD driver stacks. For local LLM users on macOS, this is a quiet but necessary maintenance step to keep inference smooth.
NVIDIA's older Volta architecture has long been an afterthought in local LLM inference, but this patch finally gives Tesla V100 users a tangible speed boost. By routing sm_70 to the Turing MMVQ nwarps table, llama.cpp reduces warp overhead for K-quant batch-1 decoding, delivering a measurable 3% throughput increase on Qwen3.8-27B models without touching perplexity. This isn't just code cleanup; it validates that legacy hardware can still compete efficiently when the runtime respects its specific kernel tuning. The change merges community work from the V100-focused anyei fork, proving that niche optimizations can become mainstream defaults.
Anthropic quietly fixed a critical credential leakage bug where MCP error messages were exposing raw API keys in plaintext logs. Beyond the security patch, this release stabilizes the notoriously fragile background agent system by fixing subagent hand-offs and connection stalls that previously caused silent failures. The update also tightens session management for cloud environments, ensuring large transcripts actually load instead of hanging indefinitely. It’s a maintenance-heavy release, but essential for anyone running complex, multi-step automated workflows.
This update shifts Claude Code from a simple CLI wrapper to a more extensible platform by introducing 'Claude Mods,' allowing plugins to modify deeper behavior rather than just adding tools. The inclusion of a built-in 'You should know' side agent that flags potential oversights is a notable step toward autonomous oversight within the coding workflow. Beyond features, the release addresses critical stability issues in remote sessions and significantly improves accessibility for screen reader users, making the tool more robust for enterprise and diverse developer environments.
The v0.31.0rc3 release of vLLM brings a critical infrastructure tweak to the new Model Runner V2: support for randomized dummy inputs. This isn't a feature for end-users but a developer-facing fix that stabilizes how the runner handles initial tensor shapes during compilation and warm-up phases. By allowing randomized inputs, it reduces the likelihood of shape-mismatch errors when tracing models with dynamic dimensions. For builders running large-scale inference workloads, this means fewer silent failures and more robust model loading sequences in production environments.