Llama.cpp's b10089 release introduces significant improvements to CUDA support, particularly in handling quantized data. The update adds k-quant and i-quant support to the GET_ROWS function, enabling more efficient device-side embedding lookups. This reduces the need for fallback to the host, enhancing performance in single-device graphs. The release also refines the handling of super-block dequantizers, ensuring comprehensive coverage for all quantized GGML types. These enhancements make CUDA more robust and efficient in processing quantized data.
Read originalThe latest b10083 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile choice for developers across different systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. Windows users benefit from updated CUDA support, with DLLs for both CUDA 12.4 and 13.3, ensuring compatibility with the latest NVIDIA technologies. While no groundbreaking new features are introduced, the release solidifies llama.cpp's position as a flexible inference runtime across diverse hardware setups.
The latest b10084 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across various systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users, and expands Vulkan support across multiple operating systems. While KleidiAI support for macOS Apple Silicon is disabled, the release still offers a comprehensive range of builds for Windows, Linux, and openEuler. This update solidifies llama.cpp's position as a go-to runtime for diverse hardware configurations, though it doesn't introduce new model architectures.
The latest b10085 release of llama.cpp addresses a key issue with the Qwen3-VL vision model's position embedding interpolation. By aligning the interpolation method with the transformers reference, the update ensures more accurate grounding coordinates, particularly for larger and non-square images. This change is crucial for developers working with image processing tasks, as it reduces discrepancies in image scaling. While the update doesn't introduce new models, it enhances the precision of existing functionalities, making llama.cpp a more reliable tool for AI developers.
© NVIDIA BlogNVIDIA has commissioned its DGX GB300 supercomputer at the Naval Postgraduate School, marking a significant step in integrating advanced AI capabilities into military education. This powerful AI platform will enable students and faculty to engage in large-scale AI computing, enhancing research in areas like weather prediction and cybersecurity. The collaboration aims to modernize military education by providing hands-on experience with cutting-edge AI tools. This deployment not only enriches academic programs but also prepares military leaders to leverage AI in real-world scenarios.
Chinese AI labs are making significant strides with open-source models that are beginning to rival the best from Silicon Valley. Moonshot AI's Kimi K3 model, in particular, has drawn attention for its impressive performance in web development and agentic tasks, challenging the belief that only closed-source models can achieve top-tier results. This development marks a growing divergence in strategy between Chinese and American AI companies, with the former embracing openness to attract users and collaborators. As these models gain traction, they are prompting a reevaluation of the value of paying for Western alternatives, suggesting a potential shift in the AI landscape.
© FireshipMoonshot has unveiled Kimi K3, an open-weight AI model boasting an impressive 2.8 trillion parameters. This release marks a significant leap in the scale of AI models, potentially offering enhanced capabilities in processing and understanding complex data. While the sheer size of Kimi K3 is noteworthy, the real test will be in its practical applications and performance compared to existing models. This development could pave the way for more advanced AI systems, but its true impact will depend on how effectively it can be utilized in real-world scenarios.