16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & Labs

Llama.cpp b10164 Release Enhances CUDA Performance

llama.cpp Releases·July 29, 2026·high confidence

Why it matters

  • →Enhances CUDA performance, crucial for developers relying on GPU acceleration.
  • →Addresses technical issues, improving reliability and efficiency.
  • →Strengthens llama.cpp's position as a robust tool for CUDA-based development.

Llama.cpp has released its b10164 update, which primarily enhances CUDA performance for Mamba-2 prefill acceleration. The update introduces chunked SSD matrix multiplication, aiming to improve efficiency and memory coalescing. It also addresses a read-write race condition in CUDA operations, along with several other technical fixes. These improvements make llama.cpp a more reliable tool for developers using CUDA and related technologies.

Read original

More from llama.cpp Releases

Models & Labsmodels

llama.cpp b10156 Release Expands Platform Support

The latest b10156 release of llama.cpp continues its trend of broadening platform compatibility, notably adding support for ROCm 7.2 on Ubuntu x64. This update ensures that AMD GPU users can leverage llama.cpp more effectively, narrowing the gap with NVIDIA's CUDA. The release also includes Vulkan support for both Ubuntu and Windows, enhancing the versatility of the software for developers. While no new models or quantization methods are introduced, this update solidifies llama.cpp's position as a versatile inference runtime across diverse hardware configurations.

llama.cpp Releases·Jul 29, 2026
Open Sourcemodels

llama.cpp b10158 Release Expands Platform Support

The latest b10158 release of llama.cpp continues its trend of broadening platform compatibility, though without major new features. Notably, the release includes support for ROCm 7.2 on Ubuntu x64, which is significant for AMD GPU users seeking alternatives to NVIDIA's CUDA. While KleidiAI support for Apple Silicon remains disabled, the release still covers a wide array of platforms, including Windows and openEuler. This update demonstrates llama.cpp's commitment to being a versatile inference runtime across diverse hardware configurations.

llama.cpp Releases·Jul 29, 2026
Open Sourcemodels

llama.cpp b10159 release enhances Metal backend

The latest b10159 release of llama.cpp introduces a new FWHT kernel for the Metal backend, significantly boosting performance for Apple Silicon users. This update, co-authored by YiChen Lv and Georgi Gerganov, also resolves a narrowing issue and refines formatting and style. Although the KleidiAI feature for macOS Apple Silicon is still disabled, the release maintains compatibility with platforms like Ubuntu, Windows, and openEuler. With ROCm 7.2 and CUDA 12 and 13 support, llama.cpp continues to evolve as a robust inference runtime, catering to diverse hardware configurations.

llama.cpp Releases·Jul 29, 2026

More in Models & Labs

Together AI Enhances Model Inference Configuration© Together AI Blog
Models & Labsmodels

Together AI Enhances Model Inference Configuration

Together AI has introduced a sophisticated architecture for model inference that integrates endpoints, deployments, and configurations with capacity-aware traffic splitting. This system allows for seamless rollouts, A/B testing, and zero-downtime updates, making it easier for developers to manage and optimize AI models. By using immutable configurations and a weight-based traffic split, the platform ensures efficient resource allocation and scaling. This development simplifies the deployment process and enhances the reliability of AI applications by ensuring consistent performance and easy rollback options.

Together AI Blog·Jul 29, 2026
Grok 4.5 Now Integrated with GitHub Copilot© GitHub Changelog
Models & Labscoding

Grok 4.5 Now Integrated with GitHub Copilot

Grok 4.5, xAI's latest reasoning model, is now integrated into GitHub Copilot, enhancing its capabilities for complex coding tasks. With a massive context window of up to 500,000 tokens and support for both text and image inputs, Grok 4.5 is designed for fast, agentic coding and multi-step workflows. It excels in terminal-based coding tasks, particularly in Visual Studio Code and Copilot CLI, making it ideal for time-sensitive and complex coding challenges. This integration marks a significant step in improving the efficiency and capability of GitHub Copilot for developers.

GitHub Changelog·Jul 28, 2026
OlmoEarth Platform Enables Large-Scale Geospatial Inference© Hugging Face Blog
Models & Labsmodels

OlmoEarth Platform Enables Large-Scale Geospatial Inference

The OlmoEarth Platform is a significant advancement in geospatial inference, designed to handle the massive scale of Earth observation data. By processing terabytes of satellite imagery efficiently, it enables organizations to generate continent-scale maps in a day, at minimal cost. This platform addresses the challenges of data acquisition, processing, and inference, making it accessible even to organizations without extensive engineering resources. With its ability to run large-scale inference jobs using thousands of CPUs and GPUs, OlmoEarth is poised to transform how environmental data is utilized for applications like wildfire risk mapping and deforestation monitoring.

Hugging Face Blog·Jul 28, 2026