Llama.cpp has released an update that adds F16 support for unary operations in its Hexagon backend, specifically including the ABS function. This enhancement builds on the existing support for operations like NORM and SQRT, and merges execution paths for F32 and F16 to improve efficiency. The update has been verified on-device, ensuring reliable performance without relying on CPU fallback. This development marks a significant improvement in AI model execution across various hardware platforms.
Read originalThe b10764 release of llama.cpp marks another step in its evolution, enhancing its reach across different computing environments. With new support for Ubuntu systems using Vulkan and ROCm 7.14, and Windows systems equipped with CUDA 13, developers gain more flexibility in deploying AI models. This update doesn't bring new features but reinforces llama.cpp's adaptability, making it a reliable choice for developers working with diverse hardware setups. By broadening its compatibility, llama.cpp continues to be a preferred runtime for those seeking performance optimization across a spectrum of platforms.
Llama.cpp's b10766 release marks a notable enhancement by enabling input vision capabilities for the deepseek4 model. This update expands the framework's reach across a variety of platforms, including macOS, Linux, and Windows, with integration for Vulkan, ROCm, and CUDA technologies. While no new models are added, the focus is on strengthening the existing infrastructure, making it more adaptable for developers using different hardware setups. This quiet yet impactful update ensures that llama.cpp remains a versatile and reliable tool for AI developers, enhancing its utility without altering its core model offerings.
The latest b10767 release of llama.cpp continues its trend of broadening platform compatibility, now supporting a wide array of systems including macOS, Linux, Windows, and openEuler. Notably, this update includes support for Vulkan and ROCm 10.0 on Ubuntu, as well as CUDA 13 on Windows, which enhances performance options for developers using these platforms. While the release doesn't introduce new model architectures, it solidifies llama.cpp's position as a versatile inference runtime across diverse hardware configurations. This update is a testament to llama.cpp's commitment to making AI more accessible and efficient for developers working on various systems.
© The Verge AIGoogle's new Gemini 3.8 Flash model promises enhanced performance by executing more reasoning steps and iteratively calling tools, potentially increasing token usage and costs. Despite maintaining the same introductory pricing as its predecessor, the model's efficiency in handling complex tasks could lead to higher expenses for users. It has shown significant improvements in software engineering and autonomous AI agents, outperforming competitors on various benchmarks. The model is available to consumers and developers, with a special version, Gemini 3.8 Flash Cyber, offered to governments and trusted partners through Google's Fairwind Program.
© Google DeepMindGoogle DeepMind has unveiled Gemini 3.8 Flash and 3.8 Flash Cyber, marking a significant step forward in AI-driven reasoning and cybersecurity. Gemini 3.8 Flash is designed for complex software engineering and agentic tasks, outperforming larger models at a fraction of the cost. Meanwhile, Gemini 3.8 Flash Cyber excels in vulnerability detection and automated patching, offering a decisive edge in cybersecurity. These models are powered by advanced reasoning capabilities and are available to developers and enterprises, with the Cyber variant accessible through the Fairwind Program for trusted defenders.
© Hugging Face BlogIBM and Confluent have teamed up to integrate IBM's Granite Time Series Models with Confluent's data streaming platform, enabling real-time intelligence directly on streaming data. This collaboration allows businesses to perform forecasting and anomaly detection without the need for separate ML platforms, reducing the time and complexity typically involved in deploying such models. By leveraging Confluent Cloud's native inference capabilities, users can seamlessly run these models within Apache Flink, ensuring that data remains secure and cost-efficient. This integration empowers domain experts to make informed decisions quickly, transforming live business events into actionable insights.