
Z.ai has introduced GLM-5.2, a new open-source model featuring 1 million tokens and available under the MIT license. The model is designed to be cost-effective compared to frontier AI models. It can be accessed via a hosted web app, API and agent harness, or self-hosted infrastructure. GLM-5.2 is particularly suited for workflows that are long, code-heavy, or document-heavy, offering a significant cost advantage.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
© Lev SelectorNew tiny local models Bonsai 2 and Needle (8-29 MB) demonstrate that small, offline-capable AI can make fast, useful decisions.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Linux arm64 build targeting Snapdragon chips, which unlocks local AI on high-performance mobile hardware via CPU, Adreno GPU, and Hexagon NPU acceleration. While KleidiAI on Apple Silicon has been disabled in this specific binary, the broader platform expansion signals that llama.cpp is aggressively standardizing how models run across every major silicon architecture.
Lev Selector · March 20, 2026 · Related
Fireship · April 8, 2026 · Background
Hugging Face Blog · June 17, 2026 · Related
Sam Witteveen · June 17, 2026 · Same story
Matt Wolfe · June 19, 2026 · Same story
The AI Daily Brief · June 21, 2026 · Same story
The AI Daily Brief · June 22, 2026 · Related
The Verge AI · June 28, 2026 · Related
TechCrunch AI · August 4, 2026 · Same story
WIRED AI · August 18, 2026 · Same story
Lev Selector · August 21, 2026 · Related
Matt Wolfe · August 28, 2026 · Related
Sam Witteveen · August 30, 2026 · Related
Google updates Gemini 3.8 with Live Avatar technology and advanced Text-to-Speech capabilities.