
French startup Kog is pushing the boundaries of AI inference by optimizing existing GPUs for faster performance. Their recent tech preview showcased the ability to achieve rapid single-request decoding on standard datacenter GPUs, such as AMD MI300X and Nvidia H200. This approach could significantly reduce inference times, addressing a critical bottleneck in AI workflows. While Kog faces the challenge of scaling its solution to larger models, its focus on software optimization rather than new hardware could offer a cost-effective alternative for enterprises. The startup aims to demonstrate its approach on large language models by September to secure further funding.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
This release quietly cements llama.cpp as the universal inference runtime by adding default support for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD GPU users finally get parity with NVIDIA's latest driver stack without manual configuration, while Apple Silicon KleidiAI builds are temporarily disabled to resolve stability issues. The inclusion of Snapdragon NPU support on Linux signals a serious push into edge AI hardware beyond just x86 and ARM CPUs. It is less about new features and more about ensuring the toolchain keeps pace with the rapidly evolving GPU landscape.
This release quietly closes the hardware gap for local inference by adding default builds for ROCm 10.0 and CUDA 13.4 across Linux and Windows. AMD users finally get parity with NVIDIA in the binary distribution, while CUDA 13 support future-proofs setups on newer drivers. The inclusion of Snapdragon and OpenVINO binaries further broadens the hardware surface area without requiring custom compilation. It is a pragmatic update that makes llama.cpp the most accessible runtime for diverse local AI hardware.
Together AI Blog · May 19, 2026 · Related
The AI Daily Brief · May 27, 2026 · Background
TechCrunch AI · May 28, 2026 · Background
NVIDIA Blog · June 10, 2026 · Related
Google DeepMind · June 10, 2026 · Related
NVIDIA Blog · June 30, 2026 · Related
TechCrunch AI · July 8, 2026 · Background
Hugging Face Blog · July 30, 2026 · Related
OpenAI · August 18, 2026 · Related
Hugging Face Blog · August 20, 2026 · Background
Together AI Blog · September 10, 2026 · Background
llama.cpp Releases · September 24, 2026 · Background
llama.cpp Releases · September 30, 2026 · Background