Claude Code has released version 2.1.225, bringing several enhancements and bug fixes. This update introduces gateway spend-limit support, providing detailed usage warnings. It also resolves issues with OAuth token errors on macOS and improves session handling in headless modes. The Remote Control feature now allows users to start conversations across machines by name, improving communication efficiency. These updates aim to enhance the overall user experience and reliability of Claude Code.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Claude Code Releases · May 15, 2026 · Same story
Claude Code Releases · May 19, 2026 · Same story
This release tightens the leash on Claude Code's autonomous capabilities while fixing critical sandbox escapes. The new effort parameter for Agent tools lets developers explicitly control sub-agent depth, a necessary guardrail as these systems grow more complex. Security fixes are prominent, addressing how plugins handle network paths and how file permissions persist during session resumption. It’s a stability patch that ensures the tool remains usable in enterprise environments without compromising on the new agent features.
Anthropic quietly shipped a significant model update alongside routine maintenance. Claude Haiku 5.5 is now the default on the API, offering a 1M context window at $0.10 per million tokens, which lowers the cost floor for high-volume coding tasks. The release also patches critical stability issues in the local agent runtime, specifically fixing memory leaks in HTTP MCP connections and resolving session state corruption during context compaction. These fixes matter because they stabilize the autonomous coding workflow that developers rely on daily. With Haiku 5.5 now standard, teams can deploy cheaper, faster iterations without manual configuration.
This release adds a critical observability layer to vLLM's KV cache system by exposing prompt token counts broken down by cache tier. For operators running large-scale inference, seeing exactly how much data is served from the GPU versus CPU or disk is essential for tuning memory allocation and cost efficiency. It transforms the cache from a black box into a measurable metric, allowing teams to validate whether their prefill strategies are actually hitting the intended storage layers. This granularity helps prevent over-provisioning and identifies bottlenecks in prompt processing pipelines.
This release quietly fixes a critical accuracy gap for ModernBERT encoders by implementing exact GELU activation, ensuring semantic embeddings match the original PyTorch models rather than approximations. It also brings native support for CUDA 13.4 across Linux and Windows, closing the driver compatibility lag that has plagued NVIDIA users on newer hardware stacks. While KleidiAI builds are temporarily disabled on Apple Silicon, the broader expansion to ROCm 10.0 and Snapdragon NPU keeps llama.cpp as the most versatile local inference runtime available today.