
Reflection AI, the $25 billion startup founded by former DeepMind researchers, has released Beam, its first public open-weight model for coding and agents. The company claims Beam achieves performance comparable to Z.ai’s GLM-5.2 while requiring significantly less computing power. Weights will be available under an Apache 2.0 license later this month, allowing organizations to self-host the model. Reflection is positioning Beam as a solution for enterprise 'AI factories' seeking to reduce dependency on Chinese open-source alternatives.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Sam Witteveen · June 17, 2026 · Related
TechCrunch AI · June 22, 2026 · Related
TechCrunch AI · July 14, 2026 · Related
The Verge AI · July 20, 2026 · Related
AI News · July 21, 2026 · Related
WIRED AI · July 22, 2026 · Related
The Verge AI · July 27, 2026 · Related
Hugging Face Blog · August 14, 2026 · Related
Sifted · October 6, 2026 · Background
MIT News AI · October 6, 2026 · Background
Hugging Face Blog · October 8, 2026 · Background
Reflection debuts Beam open-weight model
4 developments
This release quietly cements llama.cpp as the universal inference runtime by adding default support for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD GPU users finally get parity with NVIDIA's latest driver stack without manual configuration, while Apple Silicon KleidiAI builds are temporarily disabled to resolve stability issues. The inclusion of Snapdragon NPU support on Linux signals a serious push into edge AI hardware beyond just x86 and ARM CPUs. It is less about new features and more about ensuring the toolchain keeps pace with the rapidly evolving GPU landscape.
This release quietly closes the hardware gap for local inference by adding default builds for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD users finally get parity with NVIDIA in the binary distribution, while CUDA 13 support future-proofs setups on newer drivers. The inclusion of Snapdragon and OpenVINO binaries further broadens the hardware surface area without requiring custom compilation. It is a pragmatic update that makes llama.cpp the most accessible runtime for diverse local AI hardware.
© Lev SelectorMistral releases Large 4 'Le Chonk' while Anthropic launches Claude Haiku 5.5, continuing the trend of cheaper, faster frontier models.