Anthropic has released Claude Code v2.1.293, updating the default Haiku model to version 5.5 with a 1M context window and reduced pricing of $0.10 per million input tokens. The update addresses several stability issues in the local agent environment, including memory leaks in HTTP MCP connections and bugs causing lost messages during session backgrounding. Additional fixes resolve path-scoped rule loading errors and improve error handling for the 'claude purge' command. The release also includes minor improvements to startup times for enterprise organizations and better handling of Slack workspace configurations.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Anthropic · June 30, 2026 · Related
TechCrunch AI · June 30, 2026 · Related
Wes Roth · September 1, 2026 · Related
Claude Code Releases · September 22, 2026 · Same story
Claude Code Releases · September 30, 2026 · Same story
GitHub Changelog · October 7, 2026 · Same story
The AI Daily Brief · October 9, 2026 · Same story
GitHub Changelog · October 9, 2026 · Related
This release tightens the leash on Claude Code's autonomous capabilities while fixing critical sandbox escapes. The new effort parameter for Agent tools lets developers explicitly control sub-agent depth, a necessary guardrail as these systems grow more complex. Security fixes are prominent, addressing how plugins handle network paths and how file permissions persist during session resumption. It’s a stability patch that ensures the tool remains usable in enterprise environments without compromising on the new agent features.
Anthropic quietly patched a frustrating edge case in Claude Code’s agent hooks. Previously, instructions like 'Block commands that...' were often ignored because the model didn't recognize them as valid blocking criteria. This update ensures those prompts are properly interpreted, while also refining how stop conditions are judged to prevent premature termination. It’s a small but necessary fix for anyone relying on strict guardrails in automated coding workflows.
This release prioritizes stability over new features, addressing critical reliability issues in Claude Code's backgrounding and MCP integration. The most significant change is the addition of an 'onFailure: block' for hooks, which prevents silent failures from bypassing safety checks—a crucial update for developers relying on automated workflows. Gateway connectivity also sees major improvements with upstream timeout controls and better error handling for Bedrock and Vertex integrations. While there are no headline-grabbing capabilities, these fixes make the tool significantly more robust for heavy daily use.
This release adds a critical observability layer to vLLM's KV cache system by exposing prompt token counts broken down by cache tier. For operators running large-scale inference, seeing exactly how much data is served from the GPU versus CPU or disk is essential for tuning memory allocation and cost efficiency. It transforms the cache from a black box into a measurable metric, allowing teams to validate whether their prefill strategies are actually hitting the intended storage layers. This granularity helps prevent over-provisioning and identifies bottlenecks in prompt processing pipelines.
This release quietly fixes a critical accuracy gap for ModernBERT encoders by implementing exact GELU activation, ensuring semantic embeddings match the original PyTorch models rather than approximations. It also brings native support for CUDA 13.4 across Linux and Windows, closing the driver compatibility lag that has plagued NVIDIA users on newer hardware stacks. While KleidiAI builds are temporarily disabled on Apple Silicon, the broader expansion to ROCm 10.0 and Snapdragon NPU keeps llama.cpp as the most versatile local inference runtime available today.
This release targets a specific but painful stability issue for Android users running llama.cpp on Qualcomm Adreno A6X GPUs. The kernel compiler was crashing due to argument limits in the iot device backend, effectively breaking local inference on those chips. By skipping the problematic kernel and adding explicit detection for the Adreno 623, the team restores functionality where it previously failed hard. It’s a narrow fix, but essential for anyone trying to run models on mid-range Android hardware without hitting compiler errors.