
Google has unveiled Gemini Robotics 2, a significant advancement in AI software for humanoid robots. This release enhances coordination from feet to fingertips, allowing robots to balance, walk, and handle delicate tasks like sealing Ziploc bags and tying knots. The software includes 22 degrees of hand movement and can run locally without an internet connection. It adapts to new robot bodies within hours and powers hardware from partners such as Apptronik's Apollo 2. Developers can access the reasoning model 'Gemini Robotics ER 2' on Google AI Studio, while hardware integrators can access the full Vision-Language-Action and On-Device models through Google DeepMind's early access program.
Read originalThe b10311 release of llama.cpp tackles inefficiencies in text-to-speech (TTS) generation by refining how text streams are processed. Previously, the system would redundantly handle utterances, causing them to be read twice before completion. This update aligns the streaming overlay with the non-streaming prefill, effectively eliminating the duplication. Developers working with TTS systems will find this change streamlines the generation process and boosts efficiency. The update is accessible on macOS, Linux, and Windows, ensuring that a broad range of users can benefit from these improvements.
The b10313 release of llama.cpp introduces an LRU scheduler, significantly enhancing task management efficiency. This update includes improvements in handling coalescing, optimizing the waiting queue, and fixes for stream cases to ensure smoother operations. The release also expands platform-specific builds, such as Vulkan and ROCm 7.2 support on Ubuntu, and CUDA 12 and 13 on Windows. While there are no new model architectures, these updates demonstrate a commitment to refining performance and compatibility across various systems.