16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Coding Tools
Coding Tools

llama.cpp b11215 adds ROCm 10 and Snapdragon support

llama.cpp Releases·September 28, 2026·high confidence

Why it matters

  • →ROCm 10.0 support removes AMD GPU users from being second-class citizens in local inference.
  • →New Snapdragon binaries enable efficient AI on ARM-based laptops with Hexagon NPUs.
  • →CUDA 13.4 builds ensure compatibility with the latest NVIDIA driver stacks.

llama.cpp has released version b11215, expanding hardware support to include ROCm 10.0 for AMD GPUs on Linux and Windows, as well as native binaries for Linux ARM64 devices like Snapdragon laptops. The update also adds CUDA 13.4 builds across Ubuntu and Windows platforms alongside existing Vulkan and OpenVINO options. Notably, macOS KleidiAI support is currently disabled in this specific build. This release significantly widens the range of consumer hardware capable of running local large language models efficiently.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

llama.cpp b11153 adds ROCm 10 and Snapdragon support — llama.cpp Releases1llama.cpp b11217 adds ROCm 10 and Snapdragon support — llama.cpp Releases2llama.cpp b11215 adds ROCm 10 and Snapdragon supportSep 24You are here

How we got here

  1. 1
    llama.cpp b11153 adds ROCm 10 and Snapdragon support

    llama.cpp Releases · September 24, 2026 · Same story

  2. 2
    llama.cpp b11217 adds ROCm 10 and Snapdragon support

    llama.cpp Releases · September 28, 2026 · Same story

More from llama.cpp Releases

Coding Toolscoding

llama.cpp b11213 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds for modern hardware, closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Linux arm64 build targeting Snapdragon chips, which unlocks local LLM inference on high-performance mobile SoCs via CPU, Adreno GPU, and Hexagon NPU acceleration. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the broader platform coverage signals a strategic shift toward hardware agnosticism that benefits anyone running models outside of standard NVIDIA setups.

llama.cpp Releases·Sep 28, 2026
Coding Toolscoding

llama.cpp b11214 boosts AMD CDNA large batch performance

This release quietly sharpens llama.cpp’s edge on AMD hardware by enabling the fattn-mma kernel for large query dimensions on CDNA architectures. It specifically targets high-throughput scenarios where batch sizes push dkq beyond 256, a common bottleneck in serving workloads. By optimizing these specific matrix multiplication paths, the update reduces latency and improves throughput for enterprise-style inference without requiring code changes. This is another step in making AMD GPUs competitive with NVIDIA for heavy lifting.

llama.cpp Releases·Sep 28, 2026
Coding Toolscoding

llama.cpp adds wide SYCL FWHT kernels for Intel GPUs

Intel GPU users running llama.cpp finally get O(n log n) performance for large Hadamard transforms instead of falling back to slow dense matrix multiplication. This PR extends the Fast Walsh-Hadamard Transform kernel to handle block widths up to 8192, a critical optimization for certain quantization and attention mechanisms on SYCL hardware. The implementation uses work-group local memory to shuffle data efficiently, avoiding the quadratic cost that previously bottlenecked these operations. While it doesn't touch CUDA or ROCm, it significantly closes the performance gap for Intel Arc and Data Center GPU users who were left behind by previous narrow-kernel limits.

llama.cpp Releases·Sep 28, 2026

More in Coding Tools

Coding Toolscoding

Claude Code v2.1.282 patches critical session and UX bugs

This release stabilizes Claude Code by fixing a cascade of session-breaking errors that previously caused silent data loss or API drops. The most significant fix addresses resumed conversations re-sending messages in altered forms, which was corrupting reasoning traces and breaking extended thinking workflows. It also resolves persistent login refresh loops and managed setting parsing failures that plagued enterprise deployments. While the changelog is dense with UI tweaks like scrollbar fixes and vim mode corrections, the core value lies in restoring reliability for long-running agent sessions.

Claude Code Releases·Sep 26, 2026
Coding Toolscoding

Claude Code v2.1.283: Gateway hints and plugin fixes

This release quietly solves a major pain point for enterprise AI workflows by adding gateway hint headers, allowing LLM gateways to correctly group requests per user prompt instead of treating them as isolated events. The new managed settings for availableModelsMatch and deniedModels give organizations precise control over model access, blocking specific versions even when broader allowances exist. Beyond governance, the update stabilizes the plugin ecosystem with rigorous validation checks that prevent silent failures from broken or misconfigured extensions. These changes shift Claude Code from a developer tool to a manageable enterprise component.

Claude Code Releases·Sep 26, 2026
OpenAI Introduces GPT-Live and Agent API© Matt Wolfe
Coding Toolsagents

OpenAI Introduces GPT-Live and Agent API

OpenAI launches GPT-Live 1 and a new Agent API, enabling real-time voice interactions and autonomous agent development.

Matt Wolfe·Sep 26, 2026