Anthropic has released Claude Code version 2.1.284, updating the default model to Claude Sonnet 5.5 with a 1 million token context window. The update addresses several stability issues, including fixes for malformed image handling in Agent SDK sessions and improved error reporting for MCP server connections. Additional changes include better usage limit warnings for Max plan users and fixes for vim mode keybindings. The release also enhances the Claude apps gateway with certificate client authentication and updated telemetry forwarding options.
Read originalThis release prioritizes security hygiene and session reliability over new features. The addition of CLAUDE_CODE_DISABLE_WEB_FETCH is a critical control for enterprise environments needing to restrict external data access. Bug fixes address subtle race conditions in cloud sessions and artifact publishing that could lead to data loss or incorrect state. SSH and plugin installation issues are resolved, ensuring smoother remote workflows. It’s a maintenance update that tightens the tool's operational boundaries.
This release stabilizes Claude Code by fixing a cascade of session-breaking errors that previously caused silent data loss or API drops. The most significant fix addresses resumed conversations re-sending messages in altered forms, which was corrupting reasoning traces and breaking extended thinking workflows. It also resolves persistent login refresh loops and managed setting parsing failures that plagued enterprise deployments. While the changelog is dense with UI tweaks like scrollbar fixes and vim mode corrections, the core value lies in restoring reliability for long-running agent sessions.
This release targets a specific but costly bottleneck in Mixture-of-Experts inference on GPUs. The previous tile selection logic wasted significant compute time by misjudging the active workload per expert during dispatch. By correcting how matmul tiles are assigned, the patch ensures workers stay busy instead of idling. This is a quiet optimization that directly improves throughput for large MoE models running on Vulkan backends.
Intel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.