llama.cpp has released an update addressing performance inefficiencies in Mixture-of-Experts (MoE) inference on Vulkan. The fix corrects the matmul tile selection logic, which previously miscalculated the number of active rows per expert, leading to significant idle time among GPU workers. This optimization is particularly relevant for large MoE models like Sarvam 30B, where dispatch overhead can dominate processing time. The update includes standard binaries for various platforms including CUDA, ROCm, and Apple Silicon.
Read originalIntel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds for modern hardware, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Snapdragon build for Linux, which targets the emerging AI PC market by leveraging Adreno GPUs and Hexagon NPUs directly. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the expansion into non-NVIDIA silicon signals a strategic shift toward hardware agnosticism that benefits anyone running local models outside of standard data centers.
Anthropic quietly upgrades its local coding agent with Sonnet 5.5 as the new default, bringing a massive 1M context window to developers' terminals. This isn't just a model swap; it fundamentally changes how much codebase history you can keep in memory without manual chunking. The release also patches critical stability issues like malformed image crashes and broken MCP reconnections, making the tool significantly more reliable for complex workflows. For builders, this means deeper context awareness and fewer interruptions during long coding sessions.
This release prioritizes security hygiene and session reliability over new features. The addition of CLAUDE_CODE_DISABLE_WEB_FETCH is a critical control for enterprise environments needing to restrict external data access. Bug fixes address subtle race conditions in cloud sessions and artifact publishing that could lead to data loss or incorrect state. SSH and plugin installation issues are resolved, ensuring smoother remote workflows. It’s a maintenance update that tightens the tool's operational boundaries.