Llama.cpp has released version b10766, which now supports input vision for the deepseek4 model. This update enhances compatibility across multiple platforms, including macOS, Linux, and Windows, with support for technologies such as Vulkan, ROCm, and CUDA. The release does not introduce new models but focuses on improving the existing framework's versatility. This makes llama.cpp a more robust tool for developers working with various hardware configurations.
Read originalThe b10764 release of llama.cpp marks another step in its evolution, enhancing its reach across different computing environments. With new support for Ubuntu systems using Vulkan and ROCm 7.14, and Windows systems equipped with CUDA 13, developers gain more flexibility in deploying AI models. This update doesn't bring new features but reinforces llama.cpp's adaptability, making it a reliable choice for developers working with diverse hardware setups. By broadening its compatibility, llama.cpp continues to be a preferred runtime for those seeking performance optimization across a spectrum of platforms.
The latest b10767 release of llama.cpp continues its trend of broadening platform compatibility, now supporting a wide array of systems including macOS, Linux, Windows, and openEuler. Notably, this update includes support for Vulkan and ROCm 10.0 on Ubuntu, as well as CUDA 13 on Windows, which enhances performance options for developers using these platforms. While the release doesn't introduce new model architectures, it solidifies llama.cpp's position as a versatile inference runtime across diverse hardware configurations. This update is a testament to llama.cpp's commitment to making AI more accessible and efficient for developers working on various systems.
The latest b10770 release of llama.cpp continues its trend of broadening platform compatibility, now including support for ROCm 10.0 on Ubuntu and Windows. This update also introduces Vulkan support across multiple operating systems, enhancing the flexibility for developers working with diverse hardware configurations. While KleidiAI support on macOS Apple Silicon remains disabled, the release still marks a significant step in making llama.cpp a versatile tool for AI inference across various environments. The focus remains on expanding accessibility rather than introducing new model architectures.
© The Verge AIGoogle's new Gemini 3.8 Flash model promises enhanced performance by executing more reasoning steps and iteratively calling tools, potentially increasing token usage and costs. Despite maintaining the same introductory pricing as its predecessor, the model's efficiency in handling complex tasks could lead to higher expenses for users. It has shown significant improvements in software engineering and autonomous AI agents, outperforming competitors on various benchmarks. The model is available to consumers and developers, with a special version, Gemini 3.8 Flash Cyber, offered to governments and trusted partners through Google's Fairwind Program.
© Google DeepMindGoogle DeepMind has unveiled Gemini 3.8 Flash and 3.8 Flash Cyber, marking a significant step forward in AI-driven reasoning and cybersecurity. Gemini 3.8 Flash is designed for complex software engineering and agentic tasks, outperforming larger models at a fraction of the cost. Meanwhile, Gemini 3.8 Flash Cyber excels in vulnerability detection and automated patching, offering a decisive edge in cybersecurity. These models are powered by advanced reasoning capabilities and are available to developers and enterprises, with the Cyber variant accessible through the Fairwind Program for trusted defenders.
© Hugging Face BlogIBM and Confluent have teamed up to integrate IBM's Granite Time Series Models with Confluent's data streaming platform, enabling real-time intelligence directly on streaming data. This collaboration allows businesses to perform forecasting and anomaly detection without the need for separate ML platforms, reducing the time and complexity typically involved in deploying such models. By leveraging Confluent Cloud's native inference capabilities, users can seamlessly run these models within Apache Flink, ensuring that data remains secure and cost-efficient. This integration empowers domain experts to make informed decisions quickly, transforming live business events into actionable insights.