The b10661 release of llama.cpp has been announced, featuring expanded support across multiple platforms including Windows, macOS, Linux, and openEuler. This update introduces compatibility with Vulkan and ROCm 7.14 on Ubuntu, and CUDA 12 and 13 on Windows, catering to a variety of hardware setups. Although KleidiAI support on macOS Apple Silicon is disabled, the release enhances llama.cpp's adaptability for AI inference tasks. This update reinforces llama.cpp's role as a versatile tool for developers working outside the NVIDIA ecosystem.
Read originalThe latest b10656 release of llama.cpp continues its trend of broadening platform compatibility, now supporting a wide array of systems including macOS, Linux, Windows, and openEuler. Notably, this update includes support for Vulkan and ROCm 7.14 on Ubuntu, as well as CUDA 13 on Windows, which enhances performance on AMD and NVIDIA GPUs. While KleidiAI support for Apple Silicon is disabled, the release still marks a significant step in making llama.cpp a versatile tool across diverse hardware configurations. This update doesn't introduce new models but solidifies llama.cpp's position as a flexible inference runtime for developers.
The b10657 release of llama.cpp brings new OpenCL binary kernels, enhancing performance and compatibility across a wide range of systems. This update includes specific improvements for Apple Silicon, with KleidiAI support, and Vulkan on Ubuntu, making it more accessible for developers using these platforms. While no new model architectures are introduced, the release focuses on strengthening llama.cpp's capabilities as an inference runtime, particularly for those not using NVIDIA hardware. With ROCm 7.14 support on Ubuntu and CUDA 12 and 13 DLLs for Windows, llama.cpp continues to evolve as a versatile tool for AI model deployment. This release underscores the commitment to broadening hardware compatibility and optimizing performance across different environments.
The v0.28.0 release of vLLM introduces substantial improvements in performance and functionality, particularly for the Kimi-K3 model. With the addition of Decode Context Parallel support and fused FlashKDA decode kernels, the update significantly enhances processing speed and efficiency. DeepSeek V4 now includes sparse MLA support and advances in speculative decoding, offering better execution on both NVIDIA and AMD hardware. These updates make vLLM more robust and adaptable, providing developers with enhanced tools for deploying and executing models on a broader range of hardware configurations.
© TechCrunch AIHugging Face has introduced the Microduck, a $399 open-source robot designed to democratize physical AI. This duck-like robot, equipped with a camera, lidar sensors, and IMUs, can perform tasks like waddling, picking up objects, and roller skating. The Microduck's behaviors can be trained in simulation and deployed directly, offering developers a platform to experiment with reinforcement learning. This launch marks Hugging Face's continued expansion into affordable AI hardware, following their acquisition of Pollen Robotics and the release of the Reachy Mini robots.