
Google has announced that users will now have the option to remove visible watermarks from AI-generated content, including images, videos, and songs. This feature will be available for the Nano Banana, Omni, and Lyria models, and can be toggled in Gemini and Google's video editor, Flow. Despite this change, invisible SynthID watermarks and C2PA metadata will remain to ensure transparency. The update is part of Google's effort to balance creative control with the need to identify AI-generated content. The feature will be available in the coming days.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
TechCrunch AI · May 19, 2026 · Related
WIRED AI · May 19, 2026 · Related
The Verge AI · May 20, 2026 · Related
Matt Wolfe · May 26, 2026 · Related
The Verge AI · July 9, 2026 · Related
The Rundown AI · August 12, 2026 · Related
The Rundown AI · August 12, 2026 · Related
TechCrunch AI · August 12, 2026 · Related
Lev Selector · August 14, 2026 · Related
MIT News AI · August 18, 2026 · Background
WIRED AI · August 19, 2026 · Related
Matt Wolfe · August 20, 2026 · Related
Google Allows Removal of Visible Watermarks in AI Outputs
2 developments
© TechCrunch AIReflection AI is challenging the Chinese dominance in open-weight models with Beam, a 501B-parameter MoE model that claims to match Z.ai’s GLM-5.2 on reasoning benchmarks while using significantly less inference compute. Backed by $4.7 billion and secured GPU deals worth over $7 billion, this two-year-old startup is positioning itself as the Western alternative to DeepSeek and Qwen for enterprise and sovereign AI deployments. The model targets developers and institutions needing cost-effective, localizable infrastructure rather than just raw API access. With weights releasing this month, Beam offers a tangible option for those looking to reduce reliance on closed labs or Chinese open-source ecosystems.
© TechCrunch AIInstinct is pushing consumer AI agents into the messy reality of group dynamics by allowing them to join chats with friends who don't even have accounts. This moves beyond solo productivity tools into collaborative coordination for travel, events, and logistics, directly challenging Meta's ecosystem dominance. The architecture keeps personal data siloed from the group agent, requiring explicit permission before any action is taken, which addresses a major friction point in multi-user AI adoption. It signals that the next battleground for agents isn't just capability, but social integration.
© TechCrunch AITikTok is closing the loop on social commerce by embedding a conversational AI agent directly into its For You feed. This isn't just a chatbot; it remembers user preferences to guide discovery and pairs with one-click checkout via Stripe and Shopify partners. The move shifts TikTok from a passive discovery engine to an active transactional platform, aiming to keep users within the app rather than sending them to external search or AI tools. With $15.8 billion in estimated U.S. sales last year, this integration targets higher conversion rates by removing friction between impulse and purchase.
vLLM is quietly becoming the definitive runtime for NVIDIA's latest hardware, making NVFP4 compressed KV caches the default for DeepSeek-V4.1-Flash on SM100 GPUs. This isn't just a performance tweak; it fundamentally changes how enterprise inference scales by keeping post-quantized weights resident in GPU memory across engine restarts via the new preload daemon. The release also hardens speculative decoding with Model Runner V2, fixing OOMs that previously plagued wide expert deployments. For builders, this means lower latency and higher throughput on next-gen hardware without manual configuration overhead.
This release shifts llama.cpp from a pure text engine to a multimodal inference runtime capable of handling 'decision models' like Clef and GLM-5.3-Flash. The new /v1/systemone server endpoint standardizes how these non-autoregressive models are queried, while the extended batch API allows mixed token and embedding inputs for complex architectures. Apple Silicon users get a tangible performance boost with new Metal MMA kernels that accelerate speculative decoding by up to 3x. It’s a significant step toward supporting the next generation of hybrid reasoning models locally.
© Hugging Face BlogMost Arabic models treat the language as a monolith, missing the cultural and linguistic depth of specific dialects. Falcon-Emirati-7B closes this gap by fine-tuning on native Emirati text, synthetic data constrained by strict glossaries, and cultural heritage knowledge. It tops the new Alyah benchmark with 84.83%, proving that scale alone doesn't buy dialect competence. This release underscores a critical shift: true multilingual capability requires targeted adaptation, not just larger parameter counts.