16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Coding Tools
Coding Tools

llama.cpp b11372 optimizes Qwen4-Exp memory and GPU support

llama.cpp Releases·October 4, 2026·high confidence

Why it matters

  • →Reduces VRAM usage for Qwen4-exp long-context inference by optimizing tensor allocation.
  • →Adds CUDA 13 and ROCm 10.0 support, keeping pace with latest GPU driver releases.
  • →Improves Vulkan performance via tiling optimizations for the lightning indexer.

llama.cpp has released version b11372, focusing on memory optimization for Qwen4-exp models and expanded hardware support. Key changes include halving indexer score memory usage by computing head scores in place, reducing VRAM consumption during long-context inference. The update adds CUDA 13 compatibility (13.4 libraries) and ROCm 10.0 binaries, alongside Vulkan tiling improvements for the lightning indexer. macOS KleidiAI builds are currently disabled in this release. This version is available for download on GitHub.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

llama.cpp b11018 adds CUDA 13 and ROCm 10 builds — llama.cpp Releases1llama.cpp b11081 release with CUDA 13 and ROCm 10 support — llama.cpp Releases2llama.cpp b11372 optimizes Qwen4-Exp memory and GPU supportSep 18You are here

How we got here

  1. 1
    llama.cpp b11018 adds CUDA 13 and ROCm 10 builds

    llama.cpp Releases · September 18, 2026 · Same story

  2. 2
    llama.cpp b11081 release with CUDA 13 and ROCm 10 support

    llama.cpp Releases · September 22, 2026 · Same story

More from llama.cpp Releases

Coding Toolscoding

llama.cpp b11374 optimizes OpenVINO MoE performance

This release significantly tightens llama.cpp’s integration with Intel’s OpenVINO backend, specifically targeting Mixture of Experts (MoE) models like Qwen3.5 and Gemma-4. By fusing MoE routing and GDN normalization operations, prefill throughput on Arc GPUs jumps from 66 to over 1,600 tokens per second, effectively removing a major bottleneck for local inference on Intel hardware. The update also fixes critical stateful execution bugs that previously caused crashes or incorrect axis handling during decoding. This makes OpenVINO a far more viable option for running complex MoE architectures on consumer-grade Intel GPUs without relying on NVIDIA CUDA.

llama.cpp Releases·Oct 4, 2026
Coding Toolscoding

llama.cpp b11375 fixes Mamba state graph reallocation

This release resolves a critical crash in Mamba SSM inference when batch cells aren't contiguous. By gathering recurrent states into a single reserve that covers every split, the engine avoids illegal graph reallocations under strict scheduling modes. It’s a quiet but essential fix for anyone running non-standard sequence lengths or complex batching logic with stateful models.

llama.cpp Releases·Oct 4, 2026
Coding Toolscoding

llama.cpp b11376 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with CUDA. Equally notable is the new Snapdragon binary for Linux, which unlocks local LLM execution on ARM-based mobile chips via CPU, Adreno GPU, and Hexagon NPU paths. While KleidiAI builds are temporarily disabled, the expansion to Windows ROCm and mobile silicon signals a strategic shift toward hardware-agnostic accessibility.

llama.cpp Releases·Oct 4, 2026

More in Coding Tools

Coding Toolscoding

Claude Code v2.1.288 fixes resume and MCP bugs

This release stabilizes the core session management of Claude Code, specifically targeting the fragile state of resumed conversations where context or thinking traces were previously lost. It also patches critical reliability issues in the Model Context Protocol (MCP) integration, ensuring tool calls don't duplicate or hang indefinitely when remote servers misbehave. The addition of $.ui.selection() for mods and better GitHub CLI handling in cloud sessions shows a focus on developer workflow friction rather than new capabilities. These are necessary maintenance updates that make the tool more robust for heavy daily use.

Claude Code Releases·Oct 4, 2026
Coding Toolscoding

Claude Code v2.1.289 patches sandbox and plugin stability

This release is a classic maintenance patch for Claude Code, focusing on stabilizing the terminal interface and tightening security rules. It fixes critical bugs where deny/ask rules were bypassed in nested shell commands or via symlinks, ensuring sandbox policies actually hold. The update also resolves numerous UI freezes caused by malformed HTML tags and plugin rendering errors, making the agent feel less brittle during complex coding sessions.

Claude Code Releases·Oct 4, 2026
GitHub Copilot Code Review API and Default Effort Change© GitHub Changelog
Coding Toolscoding

GitHub Copilot Code Review API and Default Effort Change

GitHub finally exposes Copilot code review to external automation via REST and GraphQL APIs, moving it from a manual UI action to an integrable pipeline step. This allows developers to trigger reviews directly from scripts or internal tools rather than relying on the web interface. Simultaneously, the default effort level shifts to Balanced, striking a middle ground between speed and depth for most repositories. While Lite remains available for those prioritizing raw throughput, the API access is the real win here, enabling true CI/CD integration for automated code quality checks.

GitHub Changelog·Oct 2, 2026