llama.cpp has released build b11535, updating its binary distribution to support ROCm 10.0 and CUDA 13.4 libraries on both Linux and Windows platforms. The release also adds native support for Linux arm64 devices with Snapdragon processors, including CPU, Adreno GPU, and Hexagon NPU acceleration. Conversely, KleidiAI builds for macOS Apple Silicon are currently disabled pending stability fixes. This update ensures compatibility with the latest driver stacks from NVIDIA and AMD, maintaining llama.cpp's position as a leading local inference engine.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · September 22, 2026 · Same story
llama.cpp Releases · October 6, 2026 · Same story
This release quietly fixes a critical accuracy gap for ModernBERT encoders by implementing exact GELU activation, ensuring semantic embeddings match the original PyTorch models rather than approximations. It also brings native support for CUDA 13.4 across Linux and Windows, closing the driver compatibility lag that has plagued NVIDIA users on newer hardware stacks. While KleidiAI builds are temporarily disabled on Apple Silicon, the broader expansion to ROCm 10.0 and Snapdragon NPU keeps llama.cpp as the most versatile local inference runtime available today.
This release targets a specific but painful stability issue for Android users running llama.cpp on Qualcomm Adreno A6X GPUs. The kernel compiler was crashing due to argument limits in the iot device backend, effectively breaking local inference on those chips. By skipping the problematic kernel and adding explicit detection for the Adreno 623, the team restores functionality where it previously failed hard. It’s a narrow fix, but essential for anyone trying to run models on mid-range Android hardware without hitting compiler errors.
This release quietly sharpens llama.cpp’s performance on NVIDIA GPUs by fusing state snapshot copies into the recurrent cache during SSM scans. It also removes redundant CUDA copies in specific non-speculative decoding scenarios, shaving off latency where it counts. On the AMD side, ROCm 10.0 support arrives alongside stable builds for CUDA 12.8 and 13.4, keeping the library competitive across hardware vendors. KleidiAI on Apple Silicon is temporarily disabled, a minor setback for Mac users until that integration is stabilized. The net result is faster inference for SSM-based models without changing the user experience.
© Lev SelectorMistral releases Large 4 'Le Chonk' while Anthropic launches Claude Haiku 5.5, continuing the trend of cheaper, faster frontier models.
© Matt WolfeOpenAI announced GPT-6 for everyone, featuring an 'Intelligent UI' that adapts to user context and workflow needs.
© Matt WolfeFrench AI startup Mistral has released Mistral Large 4, its latest flagship model competing with top-tier US counterparts.