16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Open Source
Open Source

llama.cpp adds IBM ZDNN backend support

llama.cpp Releases·September 30, 2026·high confidence

Why it matters

  • →Enables efficient LLM inference on IBM Z mainframes via ZDNN.
  • →Expands llama.cpp's hardware support to s390x architecture.
  • →Signals growing enterprise interest in diverse AI infrastructure.

The llama.cpp project has merged a pull request adding continuous integration builds for the IBM ZDNN backend, targeting s390x architecture. Signed off by IBM engineer Aaron Teo, the update includes compiler fixes and CI configuration adjustments to support Ubuntu-based builds on mainframe hardware. Although automated testing is currently disabled, the inclusion of this backend in the official release pipeline marks a significant step toward broader hardware compatibility for local LLM inference beyond standard x86 and ARM platforms.

Read original

More from llama.cpp Releases

Coding Toolscoding

llama.cpp fixes MoE dispatch inefficiency on Vulkan

This release targets a specific but costly bottleneck in Mixture-of-Experts inference on GPUs. The previous tile selection logic wasted significant compute time by misjudging the active workload per expert during dispatch. By correcting how matmul tiles are assigned, the patch ensures workers stay busy instead of idling. This is a quiet optimization that directly improves throughput for large MoE models running on Vulkan backends.

llama.cpp Releases·Sep 30, 2026
Coding Toolscoding

llama.cpp Vulkan perf boost on Intel Arc

Intel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.

More in Open Source

Bonsai 2 and Needle: Tiny Local AI Models© Lev Selector
Open Sourcemodels

Bonsai 2 and Needle: Tiny Local AI Models

New tiny local models Bonsai 2 and Needle (8-29 MB) demonstrate that small, offline-capable AI can make fast, useful decisions.

Lev Selector·Sep 25, 2026
Exploring the Open Source AI Stack© Together AI Blog
Open Sourcecoding

Exploring the Open Source AI Stack

llama.cpp Releases
·
Sep 30, 2026
Coding Toolscoding

llama.cpp b11267 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds for modern hardware, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Snapdragon build for Linux, which targets the emerging AI PC market by leveraging Adreno GPUs and Hexagon NPUs directly. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the expansion into non-NVIDIA silicon signals a strategic shift toward hardware agnosticism that benefits anyone running local models outside of standard data centers.

llama.cpp Releases·Sep 30, 2026

The open source AI stack is becoming increasingly attractive to developers and organizations seeking more control and cost efficiency compared to closed-source models. This stack, known as the MIGHT Stack, consists of independent layers including models, inference, gateways, harnesses, and tools, allowing for flexible and customizable development workflows. Large models like Kimi K3 offer robust capabilities for complex tasks, while smaller models like GLM 5.3 Flash provide cost-effective solutions for well-defined tasks. This modular approach enables developers to experiment with new models quickly, adapting to the fast-paced evolution of AI technology.

Together AI Blog·Sep 9, 2026
AI Software Factory: Open Source Alpha Released© Cole Medin
Open Sourcecoding

AI Software Factory: Open Source Alpha Released

Cole Medin is pioneering a new approach to AI-driven software development with his AI Software Factory. This open-source project automates the entire coding process from a product requirements document to shipped code, without human intervention in the coding phase. Medin's experiment has already produced Dynachat, an AI tutor, demonstrating the potential of this method. Now, he's inviting developers to explore and contribute to the early alpha version, aiming to make this transformative tool accessible to all. This could redefine how software is developed, offering a glimpse into a future where AI handles the heavy lifting of coding.

Cole Medin·Sep 3, 2026