
Google has announced the general availability of Gemini Live Avatars, a multimodal feature within its Gemini API that enables real-time video interaction with AI agents. The service supports 97 languages and includes an Avatar Studio for customization, allowing developers to create branded digital humans for customer service or education. Pricing details have been released alongside the launch, signaling a shift from experimental preview to a commercial product ready for enterprise integration. This move expands Google's multimodal capabilities beyond text and audio into synchronized video generation.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Google AI Blog · June 5, 2026 · Related
Google DeepMind · June 9, 2026 · Related
Matt Wolfe · June 12, 2026 · Related
TechCrunch AI · June 29, 2026 · Related
Google AI Blog · July 16, 2026 · Related
TechCrunch AI · July 16, 2026 · Same story
The Verge AI · August 11, 2026 · Related
Sam Witteveen · September 24, 2026 · Related
Wes Roth · October 1, 2026 · Background
Sam Witteveen · October 1, 2026 · Background
Google launches Gemini 3.8 Live Avatar for enterprise
4 developments
© The Verge AIGoogle is bringing real-time audio scene description to Android via Gemini Live, directly challenging Apple’s VoiceOver Live Recognition. This feature targets users with low vision by providing immediate audio cues and follow-up Q&A capabilities for physical objects. It integrates deeply into the accessibility ecosystem through TalkBack, moving beyond simple text reading to contextual environmental awareness. The move signals a shift toward multimodal AI as a standard utility for daily navigation rather than just a novelty.
© TechCrunch AIAmazon’s Strands Decider 2B joins the growing wave of decision models designed to replace heavy LLMs for simple routing tasks. Built on Qwen3.5-2B, it outputs calibrated choices with confidence scores rather than generating text, offering a cheaper, faster alternative for agentic workflows. The release signals AWS’s push into specialized agent infrastructure, aiming to solve the latency and cost bottlenecks of general-purpose models. While TypeSafe’s Jev pioneered this space, Amazon’s entry brings enterprise-grade credibility and open-source accessibility to a niche that is rapidly filling with experimental clones.
© Wes RothGoogle is positioning Gemini 4 Argon as a strategic comeback model, aiming to reclaim ground in the competitive landscape. Simultaneously, OpenAI’s DevDay shifts focus from pure chat interfaces toward persistent agents with Dots and broader developer tooling. These moves signal an industry-wide pivot where AI transitions from conversational novelty to embedded utility. The real test lies in whether these new architectures deliver tangible reliability or just incremental benchmark gains.