The b10211 release of llama.cpp has been announced, featuring expanded support across multiple platforms. This update includes ROCm 7.2 for Ubuntu, enhancing AMD GPU performance, and CUDA 12 and 13 DLLs for Windows, ensuring compatibility with NVIDIA's latest technologies. The release does not introduce new model architectures but focuses on broadening platform compatibility, making llama.cpp a versatile tool for developers. This update reinforces llama.cpp's role as a flexible inference runtime across various hardware configurations.
Read originalThe latest b10208 release of llama.cpp introduces significant improvements in SYCL performance, particularly with the addition of oneMKL GEMM flash attention for XMX-accelerated prompt processing. This update addresses previous issues with interleaved destination layouts in the normalize kernel, ensuring more accurate attention outputs across models. By removing redundant stream waits and refining MKL FA dispatch gates, the release optimizes processing speeds, nearly doubling performance in some cases. These enhancements make llama.cpp a more robust and efficient tool for developers working with large language models.
The b10212 release of llama.cpp brings a significant efficiency boost by ensuring MTP tensors are loaded only when necessary. This optimization, co-authored by Stanisław Szymczyk, targets models that support MTP, reducing unnecessary resource usage. The update is applicable across environments like macOS, Linux, Windows, and openEuler, making it widely relevant. While there are no new models or architectures introduced, the focus on performance and resource management makes llama.cpp more effective for developers. This release quietly enhances the runtime experience, particularly for those leveraging MTP-supported models.
The latest b10213 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across various systems. Notably, this update includes support for ROCm 7.2 on Ubuntu x64, which is significant for AMD GPU users seeking alternatives to NVIDIA's CUDA. The release also maintains its comprehensive support for Windows, macOS, and Linux, ensuring that developers can leverage llama.cpp's capabilities regardless of their hardware preferences. While no groundbreaking new features are introduced, the consistent expansion of platform support solidifies llama.cpp's position as a flexible inference runtime.
© Lev SelectorDeepSeek has released version 4 of its Flash model, offering improved performance and capabilities.
© Lev SelectorAnthropic has streamlined its Claude Code by removing 80% of internal prompts to improve model performance.
© Lev SelectorOpenAI has significantly reduced the price of its GPT-5.6 Luna model, making it more affordable for users.