llama.cpp has released build b11213, expanding hardware support to include ROCm 10.0 for AMD GPUs on Linux and Windows, alongside new arm64 builds for Linux Snapdragon devices. The release also provides CUDA 13.4 libraries for both NVIDIA platforms and maintains OpenVINO and SYCL options for Intel hardware. Notably, KleidiAI optimizations for Apple Silicon are currently disabled in this build, while openEuler support remains restricted to specific Huawei Ascend configurations. This update broadens the accessibility of local LLM inference across diverse non-NVIDIA architectures.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 24, 2026 · Same story
llama.cpp Releases · September 24, 2026 · Same story
This release quietly sharpens llama.cpp’s edge on AMD hardware by enabling the fattn-mma kernel for large query dimensions on CDNA architectures. It specifically targets high-throughput scenarios where batch sizes push dkq beyond 256, a common bottleneck in serving workloads. By optimizing these specific matrix multiplication paths, the update reduces latency and improves throughput for enterprise-style inference without requiring code changes. This is another step in making AMD GPUs competitive with NVIDIA for heavy lifting.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Snapdragon binary for Linux, which unlocks local AI on ARM-based laptops using Adreno GPUs and Hexagon NPUs. While KleidiAI on macOS has been disabled in this build, the expansion to AMD and Qualcomm hardware makes this one of the most significant platform broadening efforts yet.
Intel GPU users running llama.cpp finally get O(n log n) performance for large Hadamard transforms instead of falling back to slow dense matrix multiplication. This PR extends the Fast Walsh-Hadamard Transform kernel to handle block widths up to 8192, a critical optimization for certain quantization and attention mechanisms on SYCL hardware. The implementation uses work-group local memory to shuffle data efficiently, avoiding the quadratic cost that previously bottlenecked these operations. While it doesn't touch CUDA or ROCm, it significantly closes the performance gap for Intel Arc and Data Center GPU users who were left behind by previous narrow-kernel limits.
This release stabilizes Claude Code by fixing a cascade of session-breaking errors that previously caused silent data loss or API drops. The most significant fix addresses resumed conversations re-sending messages in altered forms, which was corrupting reasoning traces and breaking extended thinking workflows. It also resolves persistent login refresh loops and managed setting parsing failures that plagued enterprise deployments. While the changelog is dense with UI tweaks like scrollbar fixes and vim mode corrections, the core value lies in restoring reliability for long-running agent sessions.
This release quietly solves a major pain point for enterprise AI workflows by adding gateway hint headers, allowing LLM gateways to correctly group requests per user prompt instead of treating them as isolated events. The new managed settings for availableModelsMatch and deniedModels give organizations precise control over model access, blocking specific versions even when broader allowances exist. Beyond governance, the update stabilizes the plugin ecosystem with rigorous validation checks that prevent silent failures from broken or misconfigured extensions. These changes shift Claude Code from a developer tool to a manageable enterprise component.
© Matt WolfeOpenAI launches GPT-Live 1 and a new Agent API, enabling real-time voice interactions and autonomous agent development.