The llama.cpp project has merged a pull request adding continuous integration builds for the IBM ZDNN backend, targeting s390x architecture. Signed off by IBM engineer Aaron Teo, the update includes compiler fixes and CI configuration adjustments to support Ubuntu-based builds on mainframe hardware. Although automated testing is currently disabled, the inclusion of this backend in the official release pipeline marks a significant step toward broader hardware compatibility for local LLM inference beyond standard x86 and ARM platforms.
Read originalThis release targets a specific but costly bottleneck in Mixture-of-Experts inference on GPUs. The previous tile selection logic wasted significant compute time by misjudging the active workload per expert during dispatch. By correcting how matmul tiles are assigned, the patch ensures workers stay busy instead of idling. This is a quiet optimization that directly improves throughput for large MoE models running on Vulkan backends.
Intel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.
The open source AI stack is becoming increasingly attractive to developers and organizations seeking more control and cost efficiency compared to closed-source models. This stack, known as the MIGHT Stack, consists of independent layers including models, inference, gateways, harnesses, and tools, allowing for flexible and customizable development workflows. Large models like Kimi K3 offer robust capabilities for complex tasks, while smaller models like GLM 5.3 Flash provide cost-effective solutions for well-defined tasks. This modular approach enables developers to experiment with new models quickly, adapting to the fast-paced evolution of AI technology.