
Google DeepMind has unveiled Gemini Omni 1.1 Flash, a new suite of tools aimed at enhancing generative video capabilities for developers. The update allows for scene extensions, keyframe specification, and video upscaling to 4K, making it suitable for professional production. By analyzing up to 10 seconds of prior video context, the model improves narrative consistency. Additionally, it offers faster and cheaper 360p previews for efficient prototyping. This release is available through Google AI Studio and the Gemini app, providing developers with advanced tools for creative video workflows.
Read originalGoogle DeepMind is pioneering a new approach to AI model evaluation with the introduction of double-blind testing. This method ensures that AI models are evaluated without prior exposure to test questions, addressing the issue of benchmark contamination. By partnering with organizations like the Singapore AI Safety Institute and OpenMined, DeepMind aims to enhance the integrity of AI assessments. This initiative marks a significant step in building trust in AI benchmarks, ensuring they accurately reflect a model's capabilities without artificial score inflation.
© Google DeepMindGoogle DeepMind has unveiled Gemini 3.5 Transcribe, a cutting-edge speech-to-text model that excels in handling background noise, complex jargon, and disfluency cleanup. This model is integrated into various Google products, offering developers the ability to build advanced voice capabilities through the Gemini API. With impressive word error rates and support for over 85 languages, Gemini 3.5 Transcribe sets a new standard for transcription accuracy and latency. This release marks a significant improvement over previous models, making voice interactions more natural and intuitive across Google's ecosystem.
© The Verge AIGoogle's Gemini Notebook has introduced a new 'Expert Intelligence' feature that allows users to interact with books purchased from Google Play Books. This integration enables users to ask questions, generate content like infographics and podcasts, and even view full texts within the app. With support from over 100,000 books and collaborations with authors like Michael Pollan and Kim Scott, this feature aims to enhance the note-taking experience by providing personalized insights and content creation tools. This move not only enriches user engagement but also opens new revenue streams for authors and publishers.
© The Verge AIHugging Face's Pollen Robotics has introduced the Microduck, a charming AI robot that combines playful design with functional capabilities. Standing under 10 inches tall, this rollerskating duck can pick up objects, follow a laser pointer, and even sing with a unique voice generated for each unit. The open-source software allows users to customize and retrain the robot for new tasks, making it a versatile addition to any tech enthusiast's collection. With its blend of whimsy and utility, the Microduck represents a playful yet practical step forward in consumer robotics.
© The Verge AIAdobe's latest Photoshop update introduces a new interface that brings all AI tools together, making them more accessible to users. The 'AI Assisted Editor' view offers a consolidated toolbar featuring capabilities like prompt-based image editing and background removal. A standout addition is the 'markup' feature, which allows users to draw directly on images to guide AI adjustments, eliminating the need for text prompts. This update, powered by Adobe's Firefly Image 5 model, enhances the editing process by understanding the full context of images, making complex tasks like opening closed eyes or adding elements more intuitive and efficient.