Google has launched Gemini 3.6 Flash and 3.5 Flash-Lite, designed to reduce latency and token costs for enterprise AI agents. The 3.6 Flash model cuts output tokens by 17%, improving efficiency in reasoning tasks, while 3.5 Flash-Lite is optimized for high-volume document processing. These models are integrated into platforms like Figma and Harvey, enhancing design and document workflows. Google's move aims to make AI tools more efficient and cost-effective for large-scale enterprise applications, with the models available through the Gemini API and other Google platforms.
Read originalSenseTime's Galaxy Project is a bold move to scale domestic AI chip infrastructure in China, aiming to create a closed-loop system that integrates chip technology, ecosystem partnerships, and commercial deployment. By collaborating with nearly 20 partners, including major domestic chip vendors, SenseTime seeks to enhance the efficiency and adaptability of AI computing power. The project also introduces a new metric, Tokens Per Watt, to measure data center efficiency, highlighting the company's focus on energy optimization. While the ambitious forecasts and claims of increased token throughput and cost-effectiveness are promising, they remain unverified by third parties, leaving room for skepticism until proven in real-world applications.
Bristol Myers Squibb is making a significant leap in its AI capabilities by acquiring an Nvidia DGX SuperPOD built on the Vera Rubin architecture. This move positions BMS as the first life sciences company to adopt this advanced system, which promises to enhance their drug discovery and development processes. The new infrastructure will allow BMS to train proprietary models and run predictions more efficiently, potentially reducing the time required for drug candidate evaluation. By integrating this system, BMS aims to streamline its research operations, enabling scientists to focus on the most promising compounds and accelerate the development of new treatments.
The latest b10083 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile choice for developers across different systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. Windows users benefit from updated CUDA support, with DLLs for both CUDA 12.4 and 13.3, ensuring compatibility with the latest NVIDIA technologies. While no groundbreaking new features are introduced, the release solidifies llama.cpp's position as a flexible inference runtime across diverse hardware setups.
The latest b10085 release of llama.cpp addresses a key issue with the Qwen3-VL vision model's position embedding interpolation. By aligning the interpolation method with the transformers reference, the update ensures more accurate grounding coordinates, particularly for larger and non-square images. This change is crucial for developers working with image processing tasks, as it reduces discrepancies in image scaling. While the update doesn't introduce new models, it enhances the precision of existing functionalities, making llama.cpp a more reliable tool for AI developers.