The b10231 release of llama.cpp focuses on improving DSpark sidecar resolution, making it the preferred choice over DFlash due to its extra Markov head. This update allows sidecar resolution without a full model at the tag and offers an option to disable discovery with explicit -md selection. The release also broadens platform support, including macOS, Linux, Windows, and openEuler, enhancing its utility for developers. This update solidifies llama.cpp's role as a versatile tool for AI inference across various systems.
Read originalThe latest b10226 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across diverse systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. The release also maintains its comprehensive support for Windows, macOS, and Linux, ensuring that developers can leverage llama.cpp's capabilities regardless of their hardware setup. While there are no groundbreaking new features, this update solidifies llama.cpp's position as a flexible and accessible inference runtime for multiple environments.
The latest b10227 release of llama.cpp introduces a specialized parser for Qwen3, enhancing its tool parsing capabilities. This update includes a tagged thinking tool parser and refactoring efforts to improve functionality, such as the addition of a permute helper and support for omitting <tool_call>. These changes aim to streamline the parsing process and improve the overall efficiency of the system. While the release doesn't introduce new models, it strengthens the existing framework, making it more robust for developers working with complex parsing tasks.
The b10228 release of llama.cpp focuses on enhancing platform compatibility without introducing major new features. This update includes ROCm 7.2 support for Ubuntu x64, providing AMD GPU users with a viable alternative to NVIDIA's CUDA. Developers can now access llama.cpp across various systems, including macOS, Windows, and Linux, ensuring its utility in diverse environments. While the release doesn't bring groundbreaking changes, it reinforces llama.cpp's role as a flexible tool for AI inference, accommodating different hardware configurations and developer needs.
© Together AI BlogMoonshot AI has unveiled Kimi K3, a groundbreaking 2.8-trillion-parameter model, marking it as the largest open-weight model available. This model is designed for complex tasks such as long-horizon coding and deep reasoning, competing with top-tier proprietary models like GPT 5.6 Sol. Kimi K3 introduces innovative architectural features like Kimi Delta Attention and Attention Residuals, enhancing its ability to handle extensive context lengths and complex reasoning. This release signifies a major step in open-source AI, offering developers a powerful tool for advanced AI applications.
© Lev SelectorDeepSeek has released version 4 of its Flash model, offering improved performance and capabilities.
© Lev SelectorAnthropic has streamlined its Claude Code by removing 80% of internal prompts to improve model performance.