16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Coding Tools
Coding Tools

vLLM v0.31.1rc0 exposes cached prompt tokens by tier

vLLM Releases·October 10, 2026·high confidence

Why it matters

  • →Enables precise monitoring of KV cache hit rates across different storage tiers.
  • →Helps operators optimize memory allocation and reduce inference costs.
  • →Improves debugging capabilities for prompt processing bottlenecks.

vLLM has released version v0.31.1rc0, introducing metrics that expose cached prompt tokens by cache tier. The update allows users to monitor how much of the input context is served from different storage levels, such as GPU memory or CPU offload. This feature provides deeper visibility into KV cache utilization for inference workloads. The release is a pre-release candidate aimed at improving operational transparency for large language model serving.

Read original

The story around this

TopicvLLM Software Updates

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

ParallelKernelBench Reveals Gaps in Multi-GPU Kernel Generation — Together AI Blog1Run vLLM Server on HF Jobs with One Command — Hugging Face Blog2vLLM Boosts Transformers Backend for Native-Speed Inference — Hugging Face Blog3Llama.cpp b10178 Release Adds Trace Logging — llama.cpp Releases4Together AI Introduces Autoscaling for LLM Inference — Together AI Blog5llama.cpp b10662 release refines KV cache handling — llama.cpp Releases6vLLM v0.29.0rc6 fixes hybrid model caching — vLLM Releases7vLLM v0.30.0: DeepSeek V4.1 and Fast Start — vLLM Releases8vLLM v0.31.1rc0 exposes cached prompt tokens by tierJun 23You are here

How we got here

  1. 1
    ParallelKernelBench Reveals Gaps in Multi-GPU Kernel Generation

    Together AI Blog · June 23, 2026 · Background

  2. 2
    Run vLLM Server on HF Jobs with One Command

    Hugging Face Blog · June 26, 2026 · Background

  3. 3
    vLLM Boosts Transformers Backend for Native-Speed Inference

    Hugging Face Blog · July 8, 2026 · Background

  4. 4
    Llama.cpp b10178 Release Adds Trace Logging

    llama.cpp Releases · July 30, 2026 · Background

  5. 5
    Together AI Introduces Autoscaling for LLM Inference

    Together AI Blog · July 31, 2026 · Background

  6. 6
    llama.cpp b10662 release refines KV cache handling

    llama.cpp Releases · August 28, 2026 · Related

  7. 7
    vLLM v0.29.0rc6 fixes hybrid model caching

    vLLM Releases · September 16, 2026 · Related

  8. 8
    vLLM v0.30.0: DeepSeek V4.1 and Fast Start

    vLLM Releases · September 22, 2026 · Related

More in Coding Tools

Coding Toolscoding

Claude Code v2.1.292 patches security and agent logic

This release tightens the leash on Claude Code's autonomous capabilities while fixing critical sandbox escapes. The new effort parameter for Agent tools lets developers explicitly control sub-agent depth, a necessary guardrail as these systems grow more complex. Security fixes are prominent, addressing how plugins handle network paths and how file permissions persist during session resumption. It’s a stability patch that ensures the tool remains usable in enterprise environments without compromising on the new agent features.

Claude Code Releases·Oct 10, 2026
Coding Toolscoding

Claude Code v2.1.293 updates Haiku and fixes agent bugs

Anthropic quietly shipped a significant model update alongside routine maintenance. Claude Haiku 5.5 is now the default on the API, offering a 1M context window at $0.10 per million tokens, which lowers the cost floor for high-volume coding tasks. The release also patches critical stability issues in the local agent runtime, specifically fixing memory leaks in HTTP MCP connections and resolving session state corruption during context compaction. These fixes matter because they stabilize the autonomous coding workflow that developers rely on daily. With Haiku 5.5 now standard, teams can deploy cheaper, faster iterations without manual configuration.

Claude Code Releases·Oct 10, 2026
Coding Toolscoding

Claude Code v2.1.294 fixes hook logic

Anthropic quietly patched a frustrating edge case in Claude Code’s agent hooks. Previously, instructions like 'Block commands that...' were often ignored because the model didn't recognize them as valid blocking criteria. This update ensures those prompts are properly interpreted, while also refining how stop conditions are judged to prevent premature termination. It’s a small but necessary fix for anyone relying on strict guardrails in automated coding workflows.

Claude Code Releases·Oct 10, 2026