16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Coding Tools
Coding Tools

llama.cpp b11272 optimizes Hexagon NPU and adds CUDA 13

llama.cpp Releases·September 30, 2026·high confidence

Why it matters

  • →Hexagon NPU optimizations improve efficiency for AI inference on Snapdragon devices.
  • →CUDA 13 support ensures compatibility with the latest NVIDIA driver stacks.
  • →ROCm 10.0 inclusion maintains viable AMD GPU support for local inference.

llama.cpp has released build b11272, focusing on performance optimizations for Qualcomm Hexagon NPUs and expanded GPU driver support. Key changes include optimizing the concat operation by reducing packet overhead in gather/transpose hot loops and replacing software division with fast paths. The release adds official support for CUDA 13 (13.4 libraries) on Linux and Windows, alongside continued ROCm 10.0 support for AMD GPUs. Notably, macOS Apple Silicon builds with KleidiAI are currently disabled in this version.

Read original

More from llama.cpp Releases

Coding Toolscoding

llama.cpp fixes MoE dispatch inefficiency on Vulkan

This release targets a specific but costly bottleneck in Mixture-of-Experts inference on GPUs. The previous tile selection logic wasted significant compute time by misjudging the active workload per expert during dispatch. By correcting how matmul tiles are assigned, the patch ensures workers stay busy instead of idling. This is a quiet optimization that directly improves throughput for large MoE models running on Vulkan backends.

llama.cpp Releases·Sep 30, 2026
Coding Toolscoding

llama.cpp Vulkan perf boost on Intel Arc

Intel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.

More in Coding Tools

Coding Toolscoding

Claude Code v2.1.284 adds Sonnet 5.5

Anthropic quietly upgrades its local coding agent with Sonnet 5.5 as the new default, bringing a massive 1M context window to developers' terminals. This isn't just a model swap; it fundamentally changes how much codebase history you can keep in memory without manual chunking. The release also patches critical stability issues like malformed image crashes and broken MCP reconnections, making the tool significantly more reliable for complex workflows. For builders, this means deeper context awareness and fewer interruptions during long coding sessions.

Claude Code Releases·Sep 30, 2026
Coding Toolscoding

Claude Code v2.1.285 patches security and stability

This release prioritizes security hygiene and session reliability over new features. The addition of CLAUDE_CODE_DISABLE_WEB_FETCH is a critical control for enterprise environments needing to restrict external data access. Bug fixes address subtle race conditions in cloud sessions and artifact publishing that could lead to data loss or incorrect state. SSH and plugin installation issues are resolved, ensuring smoother remote workflows. It’s a maintenance update that tightens the tool's operational boundaries.

llama.cpp Releases
·
Sep 30, 2026
Coding Toolscoding

llama.cpp b11267 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds for modern hardware, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Snapdragon build for Linux, which targets the emerging AI PC market by leveraging Adreno GPUs and Hexagon NPUs directly. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the expansion into non-NVIDIA silicon signals a strategic shift toward hardware agnosticism that benefits anyone running local models outside of standard data centers.

llama.cpp Releases·Sep 30, 2026
Claude Code Releases·Sep 30, 2026
GitHub adds repo-level Dependabot runner config© GitHub Changelog
Coding Toolscoding

GitHub adds repo-level Dependabot runner config

Dependabot finally gets granular control over where its jobs run. Repository admins can now override the organization-level default to target specific self-hosted runners or larger GitHub-hosted environments using custom labels and groups. This closes a long-standing gap for teams managing private registries or specialized build environments, allowing security updates to run in isolated contexts without affecting other workflows. While public repos and Enterprise Server are currently excluded, this is a critical step toward secure, compliant dependency scanning at scale.

GitHub Changelog·Sep 29, 2026