16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Coding Tools
Coding Tools

llama.cpp b11424 adds CUDA 13 and ROCm 10.0

llama.cpp Releases·October 6, 2026·high confidence

Why it matters

  • →CUDA 13 support allows users with newer NVIDIA drivers to run inference without downgrading or manual setup.
  • →ROCm 10.0 adds parity for AMD GPU users, making them first-class citizens alongside CUDA.
  • →Snapdragon and OpenVINO binaries extend local AI capabilities to mobile and Intel edge devices.

llama.cpp has released build b11424, expanding hardware compatibility with native support for CUDA 13 (versions 12.4 and 13.4) and ROCm 10.0 across Linux and Windows platforms. The update also adds binaries for Ubuntu arm64 with CUDA 13, Windows arm64 with OpenCL Adreno, and Linux arm64 targeting Snapdragon devices with CPU, GPU, and NPU support. Notably, macOS Apple Silicon KleidiAI builds are disabled in this release, while openEuler support remains limited to specific Huawei Ascend configurations. This update ensures developers can run local models on the latest NVIDIA and AMD driver stacks without custom compilation.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

llama.cpp b11027 adds CUDA 13 and ROCm 10 builds — llama.cpp Releases1llama.cpp b11036 adds CUDA 13 and ROCm 10.0 — llama.cpp Releases2llama.cpp b11424 adds CUDA 13 and ROCm 10.0Sep 18You are here

How we got here

  1. 1
    llama.cpp b11027 adds CUDA 13 and ROCm 10 builds

    llama.cpp Releases · September 18, 2026 · Same story

  2. 2
    llama.cpp b11036 adds CUDA 13 and ROCm 10.0

    llama.cpp Releases · September 19, 2026 · Same story

More from llama.cpp Releases

Coding Toolscoding

llama.cpp b11425 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. The inclusion of Snapdragon AI stack binaries for Linux marks a strategic push into ARM-based edge devices, while the simultaneous addition of CUDA 13 builds ensures compatibility with the latest driver stacks. By standardizing these hardware backends across major operating systems, the project removes friction for developers deploying models on diverse non-NVIDIA hardware.

llama.cpp Releases·Oct 6, 2026
Coding Toolscoding

llama.cpp 0.6.0 adds CUDA 13 and Snapdragon support

The llama.cpp 0.6.0 release quietly expands hardware coverage where it counts most: next-gen NVIDIA GPUs and mobile silicon. By shipping native builds for CUDA 13.4 alongside the existing CUDA 12 binaries, users can finally leverage newer GPU architectures without compiling from source. The inclusion of Linux arm64 support for Snapdragon chips with Adreno GPU and Hexagon NPU acceleration signals a serious push into on-device inference beyond Apple Silicon. While KleidiAI on macOS is temporarily disabled, the broader platform expansion makes this one of the most versatile local inference releases in recent memory.

llama.cpp Releases·Oct 6, 2026
Models & Labsmodels

llama.cpp v0.6.0 adds decision model API and Apple Silicon speedups

This release shifts llama.cpp from a pure text engine to a multimodal inference runtime capable of handling 'decision models' like Clef and GLM-5.3-Flash. The new /v1/systemone server endpoint standardizes how these non-autoregressive models are queried, while the extended batch API allows mixed token and embedding inputs for complex architectures. Apple Silicon users get a tangible performance boost with new Metal MMA kernels that accelerate speculative decoding by up to 3x. It’s a significant step toward supporting the next generation of hybrid reasoning models locally.

llama.cpp Releases·Oct 6, 2026

More in Coding Tools

Coding Toolscoding

Claude Code v2.1.290: Plugin hooks and agent stability fixes

This release significantly tightens the security model for Claude Code plugins by exposing server tool IDs and approval ceilings to hook functions, allowing developers to build more granular permission checks. It also stabilizes long-running agent sessions by fixing critical bugs in subagent resume logic and scheduled task persistence after compaction. For plugin authors, the new validation flags ensure gating hooks are properly configured before deployment. These changes make the platform safer for enterprise use while reducing friction for complex automated workflows.

Claude Code Releases·Oct 6, 2026
GitHub Secret Scanning adds Lovable and Supabase detectors© GitHub Changelog
Coding Toolscoding

GitHub Secret Scanning adds Lovable and Supabase detectors

GitHub quietly expanded its secret scanning partnership program to include Lovable Labs, Pydantic Services, and Supabase. This update means credentials from these popular development platforms are now automatically detected in public repositories, allowing the providers to revoke or rotate compromised keys before abuse occurs. For developers using Supabase or Lovable, this adds a critical layer of automated security hygiene without requiring manual configuration. It reflects GitHub's ongoing effort to integrate directly with the modern AI and database tooling stack that dominates current development workflows.

GitHub Changelog·Oct 5, 2026
Open-source tool deletes Apple Intelligence data on macOS© The Verge AI
Coding Toolsother

Open-source tool deletes Apple Intelligence data on macOS

Apple removed the simple toggle to disable AI features in macOS 27, leaving 12GB of models lingering on disk even when turned off. RemoveMacAI solves this by automating the cleanup: it disables Siri, Writing Tools, and Genmoji, deletes the underlying models, and blocks future downloads. This restores user control over storage and privacy without requiring manual navigation through scattered settings panes. It’s a practical patch for an ecosystem change that prioritized feature retention over user choice.

The Verge AI·Oct 5, 2026