
IBM and Confluent have partnered to bring IBM's Granite Time Series Models to Confluent's data streaming platform, enhancing real-time data processing capabilities. This integration allows businesses to perform forecasting and anomaly detection directly on streaming data, eliminating the need for separate machine learning platforms. With native inference capabilities in Confluent Cloud, users can run these models within Apache Flink, ensuring secure and cost-efficient data processing. This collaboration aims to streamline the deployment of time series models, enabling faster decision-making and improved operational efficiency.
Read original
© Hugging Face BlogHugging Face has unveiled BenchMIRT, a novel tool designed to dissect LLM benchmarks at the level of individual prompts. By leveraging multidimensional Item Response Theory, BenchMIRT can distinguish between different capabilities like safety and general reasoning within a single benchmark. This allows researchers to better understand what specific skills are being measured and how they contribute to a model's overall score. The tool's ability to predict model performance on unseen questions further enhances its utility, offering a more nuanced view of model capabilities beyond aggregate scores.
Hugging Face has unveiled @huggingface/kernels, a library offering over 200 optimized WebGPU kernels for local AI processing. This release aims to enhance browser-based AI inference by providing a collection of kernels that are individually versioned and benchmarked. The accompanying Fleet tool allows users to benchmark these kernels on their own hardware, contributing to a community-driven performance database. This initiative promises to improve the efficiency of AI operations in browsers, making them more accessible and performant across different devices and configurations.
Llama.cpp's b10766 release marks a notable enhancement by enabling input vision capabilities for the deepseek4 model. This update expands the framework's reach across a variety of platforms, including macOS, Linux, and Windows, with integration for Vulkan, ROCm, and CUDA technologies. While no new models are added, the focus is on strengthening the existing infrastructure, making it more adaptable for developers using different hardware setups. This quiet yet impactful update ensures that llama.cpp remains a versatile and reliable tool for AI developers, enhancing its utility without altering its core model offerings.
Llama.cpp's latest update enhances its Hexagon backend by adding F16 support for unary operations, including the ABS function. This development extends the existing capabilities of the HTP backend, which already supports operations like NORM and SQRT. By merging F32 and F16 execution paths, the update streamlines processing and avoids redundancy, ensuring efficient operation across different data types. This change is verified on-device, indicating robust performance without CPU fallback. The update signifies a step forward in optimizing AI model execution on diverse hardware platforms.
© The Verge AIGoogle's new Gemini 3.8 Flash model promises enhanced performance by executing more reasoning steps and iteratively calling tools, potentially increasing token usage and costs. Despite maintaining the same introductory pricing as its predecessor, the model's efficiency in handling complex tasks could lead to higher expenses for users. It has shown significant improvements in software engineering and autonomous AI agents, outperforming competitors on various benchmarks. The model is available to consumers and developers, with a special version, Gemini 3.8 Flash Cyber, offered to governments and trusted partners through Google's Fairwind Program.