Llama.cpp has released version b10171, which resolves a critical issue in the Adreno KQ/KQV image kernels. The bug, which ignored certain dimensions in multi-stream batches, led to incorrect outputs on devices such as the Adreno 740. The update reroutes operations to a path that properly handles these dimensions, ensuring accurate processing. This fix is particularly important for developers using llama.cpp on devices without flash attention, as it restores expected performance and reliability.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
llama.cpp Releases · June 12, 2026 · Same story
llama.cpp Releases · July 6, 2026 · Same story
This release quietly fixes a critical accuracy gap for ModernBERT encoders by implementing exact GELU activation, ensuring semantic embeddings match the original PyTorch models rather than approximations. It also brings native support for CUDA 13.4 across Linux and Windows, closing the driver compatibility lag that has plagued NVIDIA users on newer hardware stacks. While KleidiAI builds are temporarily disabled on Apple Silicon, the broader expansion to ROCm 10.0 and Snapdragon NPU keeps llama.cpp as the most versatile local inference runtime available today.
This release targets a specific but painful stability issue for Android users running llama.cpp on Qualcomm Adreno A6X GPUs. The kernel compiler was crashing due to argument limits in the iot device backend, effectively breaking local inference on those chips. By skipping the problematic kernel and adding explicit detection for the Adreno 623, the team restores functionality where it previously failed hard. It’s a narrow fix, but essential for anyone trying to run models on mid-range Android hardware without hitting compiler errors.
OpenAI announced GPT-6 for everyone, featuring an 'Intelligent UI' that adapts to user context and workflow needs.