
Anthropic released Claude Opus 5.5, achieving a score of 58 on the AI Intelligence Index and reducing prices by 40% compared to its predecessor. OpenAI countered with GPT-6 Sol and Luna, offering models that are 50% cheaper than previous versions while maintaining near-frontier performance. These moves mark a strategic pivot toward cost-efficiency in the frontier model market, challenging existing pricing structures.
Read originalvLLM is quietly closing the hardware gap for AMD users with this release candidate. By adding dense NVFP4 and MoRI kernel mirrors for the new MI355 GPU, they are enabling high-efficiency inference on hardware that previously lacked first-class support. This isn't just a driver update; it's a critical infrastructure patch that allows enterprises to deploy advanced quantization formats on AMD silicon without waiting for upstream integration. The inclusion of OpenAI Codex in the commit history suggests automated testing is helping maintain this parity, making AMD a more viable option for cost-sensitive inference workloads.
This release quietly solidifies llama.cpp’s position as the universal inference runtime by adding explicit ROCm 10.0 builds for both Linux and Windows. The inclusion of CUDA 13.4 alongside the existing 12.x variants ensures compatibility with the latest NVIDIA driver stacks without forcing users to stick to older libraries. More importantly, the new backend testing infrastructure means these diverse hardware configurations are now validated systematically rather than left to chance. This reduces fragmentation for developers running on AMD or newer NVIDIA cards who previously had to troubleshoot build issues manually.
This release refines the internal test suite for better visibility, but the real signal is platform expansion. CUDA 13 builds are now available across Linux and Windows, giving developers early access to the latest NVIDIA stack without waiting for stable drivers. AMD ROCm support also advances with version 10.0 binaries, keeping pace with hardware shifts. While KleidiAI on Apple Silicon is currently disabled, the core inference runtime remains robust across major architectures. This is a maintenance-heavy update that ensures compatibility with cutting-edge GPU libraries.