
OpenRouter has launched Fusion, a tool designed to enhance model routing capabilities in AI systems. This development is part of a larger trend towards more strategic and fragmented AI ecosystems. Fusion aims to provide users with greater flexibility and control over how AI models are deployed and managed, reflecting a shift away from reliance on single, monolithic systems.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
OpenAI · May 5, 2026 · Background
Lev Selector · May 8, 2026 · Background
TechCrunch AI · May 26, 2026 · Related
The Rundown AI · June 3, 2026 · Background
Hugging Face Blog · July 15, 2026 · Related
Together AI Blog · July 29, 2026 · Background
AI News · August 20, 2026 · Related
TechCrunch AI · August 20, 2026 · Related
Lev Selector · August 21, 2026 · Background
OpenRouter Fusion Combines Models to Reduce Costs
2 developments
© Lev SelectorNew tiny local models Bonsai 2 and Needle (8-29 MB) demonstrate that small, offline-capable AI can make fast, useful decisions.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Linux arm64 build targeting Snapdragon chips, which unlocks local AI on high-performance mobile hardware via CPU, Adreno GPU, and Hexagon NPU acceleration. While KleidiAI on Apple Silicon has been disabled in this specific binary, the broader platform expansion signals that llama.cpp is aggressively standardizing how models run across every major silicon architecture.
The latest b10989 release of llama.cpp continues its trend of broadening platform compatibility, now supporting a wide array of systems including macOS, Linux, Windows, and openEuler. Notably, this update includes support for CUDA 13.3 libraries on Ubuntu and Windows, as well as ROCm 10.0, enhancing its utility for AMD GPU users. While KleidiAI support is disabled on macOS, the release still marks a significant step in making llama.cpp a versatile tool for developers across different hardware configurations. This update doesn't introduce new models but focuses on expanding the accessibility of existing capabilities.