16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & Labs

llama.cpp b10144 Release Fixes Stream Routes

llama.cpp Releases·July 27, 2026·high confidence

Why it matters

  • →Fixes ensure reliable session management for complex model names.
  • →Enhancements improve user experience during model loading and page reloads.
  • →Developers benefit from more robust streaming operations.

The b10144 release of llama.cpp introduces fixes for stream routes, particularly for model names containing slashes. This update ensures that stop and resume functions match the session correctly, improving the reliability of streaming operations. Additionally, the release addresses issues with pending requests during model loading, allowing sessions to persist through page reloads. These enhancements are aimed at improving the user experience and reliability of the platform for developers.

Read original

More from llama.cpp Releases

Open Sourcemodels

llama.cpp b10141 Release Expands Platform Support

The latest b10141 release of llama.cpp continues its trend of broadening platform compatibility, though without major new features. Notably, it includes support for ROCm 7.2 on Ubuntu x64, which is significant for AMD GPU users seeking alternatives to NVIDIA's CUDA. The release also maintains a wide array of builds across macOS, Windows, and Linux, ensuring that developers have the flexibility to deploy on various hardware configurations. While the update doesn't introduce groundbreaking changes, it solidifies llama.cpp's position as a versatile tool for AI inference across diverse systems.

llama.cpp Releases·Jul 27, 2026
Models & Labsmodels

Llama.cpp Adds Vision Support with MiniMax-M3

Llama.cpp's latest update introduces preliminary support for the MiniMax-M3 model, marking a significant step towards integrating vision capabilities. This release reuses existing components from MiniMax-M2, incorporating advanced features like per-head QK-norm and partial rotary, while also optimizing performance with GPU and CPU operations. Although sparse attention isn't supported yet, the update promises a substantial speedup in processing long contexts. This development positions llama.cpp to better handle vision tasks, expanding its utility beyond text-only applications.

llama.cpp Releases·Jul 27, 2026
Models & Labsmodels

Llama.cpp b10093 Release Fixes DeepSeek4 Template

The b10093 release of llama.cpp focuses on refining the DeepSeek4 template to ensure it behaves consistently with reference standards. This update introduces support for the DeepSeekv4 flag and integrates the DS3.2 parser for DS4, enhancing its functionality. Developers working on macOS, Linux, and Windows can benefit from improved performance, especially with Vulkan, ROCm, and CUDA technologies. The release also addresses tool result reordering and post-merge fixes, contributing to a more stable and reliable development environment. While not revolutionary, these enhancements make llama.cpp a more dependable choice for developers seeking robust AI model support.

llama.cpp Releases·Jul 25, 2026

More in Models & Labs

Models & Labsmodels

vLLM v0.26.0 Release Enhances Model Support

The vLLM v0.26.0 release marks a significant update with the introduction of the new Inkling model family, which includes advanced features like piecewise CUDA graph support and speculative decoding. This release also enhances performance across various hardware platforms, including AMD and XPU, with optimizations like the DeepSeek-V4 performance push. Additionally, the update brings improvements in attention mechanisms and KV offloading, offering more flexibility and efficiency for hybrid models. These advancements make vLLM a more robust and versatile tool for developers working with large-scale AI models.

vLLM Releases·Jul 27, 2026
NVIDIA Uses Vera CPU to Enhance Chip Design© NVIDIA Blog
Models & Labsmodels

NVIDIA Uses Vera CPU to Enhance Chip Design

NVIDIA is deploying its Vera CPU to accelerate the design of its next-generation CPUs and GPUs, working with Cadence and Synopsys to optimize electronic design automation (EDA) applications. This initiative underscores the critical role of high-performance CPU architecture in expediting engineering workloads such as logic simulation and formal verification. Early testing reveals up to 1.5x performance improvements in key EDA applications, demonstrating Vera's potential to boost productivity in chip design. By integrating Vera into its workflows, NVIDIA aims to create a feedback loop that continuously refines its silicon design process, paving the way for more efficient future developments.

NVIDIA Blog·Jul 27, 2026
Kimi K3 vs GPT-5.6 Sol: Cost and Performance© Together AI Blog
Models & Labsmodels

Kimi K3 vs GPT-5.6 Sol: Cost and Performance

In the latest comparison on the DeepSWE benchmark, Kimi K3 and GPT-5.6 Sol showcase distinct strengths. GPT-5.6 Sol excels in single-shot quality with a pass@1 score of 72.7%, while Kimi K3 shines in multi-attempt scenarios, achieving a pass@4 score of 89.4%. Kimi K3 is also significantly more cost-effective, offering 2.8 times more solved tasks per dollar than Sol. The models' divergent strengths suggest that a routing strategy, where tasks are first attempted by Kimi K3 and escalated to Sol if necessary, could maximize coverage and efficiency. This approach leverages Kimi's cost advantage and Sol's reliability, covering 108 of 113 tasks effectively.

Together AI Blog·Jul 26, 2026