Anthropic has released Claude Code v2.1.281, a patch focused on stability and enterprise integration. Key updates include Bedrock upstream support for IAM role assumption and guardrails, alongside fixes for session resumption bugs that previously caused prompt cache loss or infinite retry loops. The release also resolves issues with proxy stream handling, oversized tool calls, and permission dialog edge cases on macOS.
Read originalThis release targets a specific bottleneck in long-context inference by optimizing the sparse flash attention prefill step for NVIDIA GPUs. By templating kernels to unroll loops at compile time, batched sparse operations drop from 586 microseconds to 244 microseconds on 49k context windows. This isn't just a generic speed bump; it makes handling very long documents significantly more efficient for users relying on sparse attention mechanisms. The change is already baked into the standard CUDA builds, requiring no special flags.
The latest llama.cpp build brings immediate relevance to users on bleeding-edge NVIDIA hardware with native CUDA 13.4 support across Linux and Windows, closing the gap for those testing next-gen GPU architectures. More notably, it finally addresses the mobile inference landscape by including a dedicated build for Linux arm64 Snapdragon devices, covering CPU, Adreno GPU, and Hexagon NPU paths. This moves local AI beyond just desktop GPUs into the realm of high-performance edge computing on Qualcomm silicon. While Apple Silicon builds have KleidiAI disabled in this specific release, the expansion to ARM-based mobile NPUs marks a significant shift in where llama.cpp can run efficiently.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Linux arm64 build targeting Snapdragon chips, which unlocks local LLM execution on high-performance mobile hardware via CPU, Adreno GPU, and Hexagon NPU acceleration. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the broader platform expansion signals a shift toward heterogeneous computing that extends well beyond traditional desktop GPUs.