vLLM has released version v0.31.1rc0, introducing metrics that expose cached prompt tokens by cache tier. The update allows users to monitor how much of the input context is served from different storage levels, such as GPU memory or CPU offload. This feature provides deeper visibility into KV cache utilization for inference workloads. The release is a pre-release candidate aimed at improving operational transparency for large language model serving.
Read originalEarlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.
Together AI Blog · June 23, 2026 · Background
Hugging Face Blog · June 26, 2026 · Background
Hugging Face Blog · July 8, 2026 · Background
llama.cpp Releases · July 30, 2026 · Background
Together AI Blog · July 31, 2026 · Background
llama.cpp Releases · August 28, 2026 · Related
vLLM Releases · September 16, 2026 · Related
vLLM Releases · September 22, 2026 · Related
This release tightens the leash on Claude Code's autonomous capabilities while fixing critical sandbox escapes. The new effort parameter for Agent tools lets developers explicitly control sub-agent depth, a necessary guardrail as these systems grow more complex. Security fixes are prominent, addressing how plugins handle network paths and how file permissions persist during session resumption. It’s a stability patch that ensures the tool remains usable in enterprise environments without compromising on the new agent features.
Anthropic quietly shipped a significant model update alongside routine maintenance. Claude Haiku 5.5 is now the default on the API, offering a 1M context window at $0.10 per million tokens, which lowers the cost floor for high-volume coding tasks. The release also patches critical stability issues in the local agent runtime, specifically fixing memory leaks in HTTP MCP connections and resolving session state corruption during context compaction. These fixes matter because they stabilize the autonomous coding workflow that developers rely on daily. With Haiku 5.5 now standard, teams can deploy cheaper, faster iterations without manual configuration.
Anthropic quietly patched a frustrating edge case in Claude Code’s agent hooks. Previously, instructions like 'Block commands that...' were often ignored because the model didn't recognize them as valid blocking criteria. This update ensures those prompts are properly interpreted, while also refining how stop conditions are judged to prevent premature termination. It’s a small but necessary fix for anyone relying on strict guardrails in automated coding workflows.