llama.cpp has released build b11539, expanding hardware support to include ROCm 10.0 and CUDA 13.4 binaries for Linux and Windows. The update adds specific builds for AMD GPUs via ROCm 10.0 and NVIDIA GPUs using the newer CUDA 13.4 libraries, alongside existing CUDA 12.x options. Apple Silicon builds remain available but with KleidiAI optimizations disabled in this iteration. The release also includes support for OpenVINO, SYCL, and various CPU architectures including ARM64 and s390x.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 20, 2026 · Same story
llama.cpp Releases · September 22, 2026 · Same story
This release quietly fixes a critical accuracy gap for ModernBERT encoders by implementing exact GELU activation, ensuring semantic embeddings match the original PyTorch models rather than approximations. It also brings native support for CUDA 13.4 across Linux and Windows, closing the driver compatibility lag that has plagued NVIDIA users on newer hardware stacks. While KleidiAI builds are temporarily disabled on Apple Silicon, the broader expansion to ROCm 10.0 and Snapdragon NPU keeps llama.cpp as the most versatile local inference runtime available today.
This release targets a specific but painful stability issue for Android users running llama.cpp on Qualcomm Adreno A6X GPUs. The kernel compiler was crashing due to argument limits in the iot device backend, effectively breaking local inference on those chips. By skipping the problematic kernel and adding explicit detection for the Adreno 623, the team restores functionality where it previously failed hard. It’s a narrow fix, but essential for anyone trying to run models on mid-range Android hardware without hitting compiler errors.
This release quietly sharpens llama.cpp’s performance on NVIDIA GPUs by fusing state snapshot copies into the recurrent cache during SSM scans. It also removes redundant CUDA copies in specific non-speculative decoding scenarios, shaving off latency where it counts. On the AMD side, ROCm 10.0 support arrives alongside stable builds for CUDA 12.8 and 13.4, keeping the library competitive across hardware vendors. KleidiAI on Apple Silicon is temporarily disabled, a minor setback for Mac users until that integration is stabilized. The net result is faster inference for SSM-based models without changing the user experience.
This release tightens the leash on Claude Code's autonomous capabilities while fixing critical sandbox escapes. The new effort parameter for Agent tools lets developers explicitly control sub-agent depth, a necessary guardrail as these systems grow more complex. Security fixes are prominent, addressing how plugins handle network paths and how file permissions persist during session resumption. It’s a stability patch that ensures the tool remains usable in enterprise environments without compromising on the new agent features.
Anthropic quietly shipped a significant model update alongside routine maintenance. Claude Haiku 5.5 is now the default on the API, offering a 1M context window at $0.10 per million tokens, which lowers the cost floor for high-volume coding tasks. The release also patches critical stability issues in the local agent runtime, specifically fixing memory leaks in HTTP MCP connections and resolving session state corruption during context compaction. These fixes matter because they stabilize the autonomous coding workflow that developers rely on daily. With Haiku 5.5 now standard, teams can deploy cheaper, faster iterations without manual configuration.
Anthropic quietly patched a frustrating edge case in Claude Code’s agent hooks. Previously, instructions like 'Block commands that...' were often ignored because the model didn't recognize them as valid blocking criteria. This update ensures those prompts are properly interpreted, while also refining how stop conditions are judged to prevent premature termination. It’s a small but necessary fix for anyone relying on strict guardrails in automated coding workflows.