The b10227 release of llama.cpp focuses on enhancing its parsing capabilities with the introduction of a specialized parser for Qwen3. This update includes a tagged thinking tool parser and improvements like a permute helper and support for omitting <tool_call>. These enhancements aim to streamline the parsing process, making it more efficient for developers. The release does not feature new models but strengthens the existing framework, offering a more robust tool for complex parsing tasks.
Read originalThe latest b10226 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across diverse systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. The release also maintains its comprehensive support for Windows, macOS, and Linux, ensuring that developers can leverage llama.cpp's capabilities regardless of their hardware setup. While there are no groundbreaking new features, this update solidifies llama.cpp's position as a flexible and accessible inference runtime for multiple environments.
The b10228 release of llama.cpp focuses on enhancing platform compatibility without introducing major new features. This update includes ROCm 7.2 support for Ubuntu x64, providing AMD GPU users with a viable alternative to NVIDIA's CUDA. Developers can now access llama.cpp across various systems, including macOS, Windows, and Linux, ensuring its utility in diverse environments. While the release doesn't bring groundbreaking changes, it reinforces llama.cpp's role as a flexible tool for AI inference, accommodating different hardware configurations and developer needs.
The b10231 release of llama.cpp brings significant improvements to DSpark sidecar resolution, making it the default choice over DFlash due to its additional Markov head. This update allows for resolving sidecars without needing a full model at the tag, and provides an option to disable discovery with explicit -md selection. While no new models are introduced, the release extends platform support across macOS, Linux, Windows, and openEuler, enhancing its adaptability for developers. With these changes, llama.cpp continues to evolve as a flexible inference runtime, catering to a wide range of system configurations.
© Together AI BlogMoonshot AI has unveiled Kimi K3, a groundbreaking 2.8-trillion-parameter model, marking it as the largest open-weight model available. This model is designed for complex tasks such as long-horizon coding and deep reasoning, competing with top-tier proprietary models like GPT 5.6 Sol. Kimi K3 introduces innovative architectural features like Kimi Delta Attention and Attention Residuals, enhancing its ability to handle extensive context lengths and complex reasoning. This release signifies a major step in open-source AI, offering developers a powerful tool for advanced AI applications.
© Lev SelectorDeepSeek has released version 4 of its Flash model, offering improved performance and capabilities.
© Lev SelectorAnthropic has streamlined its Claude Code by removing 80% of internal prompts to improve model performance.