16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

llama.cpp b11382 adds ROCm 10 and Snapdragon support

llama.cpp Releases·October 4, 2026·high confidence

Why it matters

  • →ROCm 10.0 support removes the last major barrier for AMD GPU users running local LLMs.
  • →Snapdragon Linux binaries enable AI inference on ARM-based laptops and emerging hardware platforms.
  • →Continued CUDA 13 support ensures compatibility with the latest NVIDIA driver stacks.

llama.cpp has released version b11382, expanding hardware compatibility with a focus on non-NVIDIA accelerators. The update adds ROCm 10.0 support for Linux x64 and Windows x64, allowing AMD Radeon GPUs to run local models natively. Additionally, new binaries for Linux arm64 enable inference on Snapdragon devices utilizing Adreno GPUs and Hexagon NPUs. Existing builds for CUDA 12/13, Vulkan, and OpenVINO remain available, though KleidiAI support on macOS Apple Silicon has been temporarily disabled.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

llama.cpp b11215 adds ROCm 10 and Snapdragon support — llama.cpp Releases1llama.cpp b11382 adds ROCm 10 and Snapdragon supportllama.cpp b11380 adds ROCm 10 and Snapdragon support — llama.cpp Releases2Sep 28You are hereOct 4

How we got here

  1. 1
    llama.cpp b11215 adds ROCm 10 and Snapdragon support

    llama.cpp Releases · September 28, 2026 · Same story

What happened next

  1. 2
    llama.cpp b11380 adds ROCm 10 and Snapdragon support

    llama.cpp Releases · October 4, 2026 · Same story

More from llama.cpp Releases

Coding Toolscoding

llama.cpp b11372 optimizes Qwen4-Exp memory and GPU support

This release tackles the notorious memory hunger of long-context inference for Qwen4-exp models by halving indexer score memory. The optimization works by computing head scores in place rather than materializing separate tensors, a change that significantly reduces VRAM pressure during heavy workloads. Beyond memory efficiency, b11372 expands hardware coverage with CUDA 13 support and Vulkan tiling for the lightning indexer. It also adds ROCm 10.0 binaries, keeping AMD users in step with NVIDIA's latest driver ecosystem. The result is a leaner runtime that handles extended contexts without hitting out-of-memory errors as quickly.

llama.cpp Releases·Oct 4, 2026
Coding Toolscoding

llama.cpp b11374 optimizes OpenVINO MoE performance

This release significantly tightens llama.cpp’s integration with Intel’s OpenVINO backend, specifically targeting Mixture of Experts (MoE) models like Qwen3.5 and Gemma-4. By fusing MoE routing and GDN normalization operations, prefill throughput on Arc GPUs jumps from 66 to over 1,600 tokens per second, effectively removing a major bottleneck for local inference on Intel hardware. The update also fixes critical stateful execution bugs that previously caused crashes or incorrect axis handling during decoding. This makes OpenVINO a far more viable option for running complex MoE architectures on consumer-grade Intel GPUs without relying on NVIDIA CUDA.

llama.cpp Releases·Oct 4, 2026
Coding Toolscoding

llama.cpp b11375 fixes Mamba state graph reallocation

This release resolves a critical crash in Mamba SSM inference when batch cells aren't contiguous. By gathering recurrent states into a single reserve that covers every split, the engine avoids illegal graph reallocations under strict scheduling modes. It’s a quiet but essential fix for anyone running non-standard sequence lengths or complex batching logic with stateful models.

llama.cpp Releases·Oct 4, 2026

More in Models & Labs

Meta Stock Rallies on Muse Model Performance© The AI Daily Brief
Models & Labsmodels

Meta Stock Rallies on Muse Model Performance

Meta's stock price surged following positive market reaction to its new Muse model capabilities.

The AI Daily Brief·Oct 3, 2026
Google launches Guided Vision in Gemini Live© The Verge AI
Models & Labsother

Google launches Guided Vision in Gemini Live

Google is bringing real-time audio scene description to Android via Gemini Live, directly challenging Apple’s VoiceOver Live Recognition. This feature targets users with low vision by providing immediate audio cues and follow-up Q&A capabilities for physical objects. It integrates deeply into the accessibility ecosystem through TalkBack, moving beyond simple text reading to contextual environmental awareness. The move signals a shift toward multimodal AI as a standard utility for daily navigation rather than just a novelty.

The Verge AI·Oct 1, 2026
AWS releases open-source Strands Decider 2B© TechCrunch AI
Models & Labsagents

AWS releases open-source Strands Decider 2B

Amazon’s Strands Decider 2B joins the growing wave of decision models designed to replace heavy LLMs for simple routing tasks. Built on Qwen3.5-2B, it outputs calibrated choices with confidence scores rather than generating text, offering a cheaper, faster alternative for agentic workflows. The release signals AWS’s push into specialized agent infrastructure, aiming to solve the latency and cost bottlenecks of general-purpose models. While TypeSafe’s Jev pioneered this space, Amazon’s entry brings enterprise-grade credibility and open-source accessibility to a niche that is rapidly filling with experimental clones.

TechCrunch AI·Oct 1, 2026