llama.cpp has released build b11425, expanding hardware support to include ROCm 10.0 and Snapdragon AI stack binaries. The update adds ROCm 10.0 builds for Ubuntu x64 and Windows x64, addressing a significant gap in AMD GPU compatibility on Linux. Additionally, new Linux arm64 binaries now support the Snapdragon CPU, Adreno GPU, and Hexagon NPU, targeting edge AI deployment. The release also introduces CUDA 13 (13.4) builds for Ubuntu and Windows across x64 and arm64 architectures. macOS KleidiAI builds have been disabled in this specific release.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 22, 2026 · Same story
llama.cpp Releases · September 28, 2026 · Same story
This release quietly closes the hardware gap for local inference by adding default support for CUDA 13 and ROCm 10.0 alongside existing CUDA 12 builds. NVIDIA users can now leverage newer driver stacks without manual configuration, while AMD GPU owners finally get first-class parity with the same ease of use previously reserved for CUDA. Apple Silicon KleidiAI is disabled in this specific build, a notable regression for Mac users who rely on that optimization. The inclusion of Snapdragon and OpenVINO binaries further broadens the reach to edge devices and Intel hardware. It’s less about new features and more about llama.cpp solidifying its position as the universal runtime for every major accelerator.
The llama.cpp 0.6.0 release quietly expands hardware coverage where it counts most: next-gen NVIDIA GPUs and mobile silicon. By shipping native builds for CUDA 13.4 alongside the existing CUDA 12 binaries, users can finally leverage newer GPU architectures without compiling from source. The inclusion of Linux arm64 support for Snapdragon chips with Adreno GPU and Hexagon NPU acceleration signals a serious push into on-device inference beyond Apple Silicon. While KleidiAI on macOS is temporarily disabled, the broader platform expansion makes this one of the most versatile local inference releases in recent memory.
This release shifts llama.cpp from a pure text engine to a multimodal inference runtime capable of handling 'decision models' like Clef and GLM-5.3-Flash. The new /v1/systemone server endpoint standardizes how these non-autoregressive models are queried, while the extended batch API allows mixed token and embedding inputs for complex architectures. Apple Silicon users get a tangible performance boost with new Metal MMA kernels that accelerate speculative decoding by up to 3x. It’s a significant step toward supporting the next generation of hybrid reasoning models locally.
This release significantly tightens the security model for Claude Code plugins by exposing server tool IDs and approval ceilings to hook functions, allowing developers to build more granular permission checks. It also stabilizes long-running agent sessions by fixing critical bugs in subagent resume logic and scheduled task persistence after compaction. For plugin authors, the new validation flags ensure gating hooks are properly configured before deployment. These changes make the platform safer for enterprise use while reducing friction for complex automated workflows.
© GitHub ChangelogGitHub quietly expanded its secret scanning partnership program to include Lovable Labs, Pydantic Services, and Supabase. This update means credentials from these popular development platforms are now automatically detected in public repositories, allowing the providers to revoke or rotate compromised keys before abuse occurs. For developers using Supabase or Lovable, this adds a critical layer of automated security hygiene without requiring manual configuration. It reflects GitHub's ongoing effort to integrate directly with the modern AI and database tooling stack that dominates current development workflows.
© The Verge AIApple removed the simple toggle to disable AI features in macOS 27, leaving 12GB of models lingering on disk even when turned off. RemoveMacAI solves this by automating the cleanup: it disables Siri, Writing Tools, and Genmoji, deletes the underlying models, and blocks future downloads. This restores user control over storage and privacy without requiring manual navigation through scattered settings panes. It’s a practical patch for an ecosystem change that prioritized feature retention over user choice.