16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

v0.28.0rc2: DFlash2 Local Convolution Update

vLLM Releases·August 22, 2026·medium confidence

Why it matters

  • →Enhances local convolution capabilities, potentially improving model efficiency.
  • →Introduces a candidate selector, offering more precise operations.
  • →Provides developers with tools to optimize AI models further.

vLLM has released version 0.28.0rc2, featuring the DFlash2 update which includes local convolution and a candidate selector. This update, derived from commit b389ac2, is signed off by developer khluu. The focus on local convolution indicates an enhancement in processing capabilities, potentially improving model efficiency. This release is particularly relevant for developers seeking to refine AI model performance.

Read original

The story around this

TopicvLLM Software Updates

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

vLLM v0.22.0 Release Enhances Model Performance — vLLM Releases1llama.cpp b9831 release adds DFlash support — llama.cpp Releases2vLLM Boosts Transformers Backend for Native-Speed Inference — Hugging Face Blog3vLLM v0.27.0 Release — vLLM Releases4v0.28.0rc2: DFlash2 Local Convolution Updatellama.cpp b10658 release adds DFlash2 support — llama.cpp Releases5GLM 5.3 Flash: Z.AI's Cost-Efficient Model Variant — Sam Witteveen6Liquid AI releases DSpark drafter for vision models — Hugging Face Blog7May 31You are hereSep 24

How we got here

  1. 1
    vLLM v0.22.0 Release Enhances Model Performance

    vLLM Releases · May 31, 2026 · Related

  2. 2
    llama.cpp b9831 release adds DFlash support

    llama.cpp Releases · June 30, 2026 · Related

  3. 3
    vLLM Boosts Transformers Backend for Native-Speed Inference

    Hugging Face Blog · July 8, 2026 · Related

  4. 4
    vLLM v0.27.0 Release

    vLLM Releases · August 12, 2026 · Related

What happened next

  1. 5
    llama.cpp b10658 release adds DFlash2 support

    llama.cpp Releases · August 28, 2026 · Related

  2. 6
    GLM 5.3 Flash: Z.AI's Cost-Efficient Model Variant

    Sam Witteveen · August 30, 2026 · Background

  3. 7
    Liquid AI releases DSpark drafter for vision models

    Hugging Face Blog · September 24, 2026 · Related

More from vLLM Releases

Models & Labscoding

vLLM v0.31.0: SM100 defaults and fast restarts

vLLM is quietly becoming the definitive runtime for NVIDIA's latest hardware, making NVFP4 compressed KV caches the default for DeepSeek-V4.1-Flash on SM100 GPUs. This isn't just a performance tweak; it fundamentally changes how enterprise inference scales by keeping post-quantized weights resident in GPU memory across engine restarts via the new preload daemon. The release also hardens speculative decoding with Model Runner V2, fixing OOMs that previously plagued wide expert deployments. For builders, this means lower latency and higher throughput on next-gen hardware without manual configuration overhead.

vLLM Releases·Oct 6, 2026
Coding Toolscoding

vLLM v0.31.0rc3 adds Model Runner V2 dummy inputs

The v0.31.0rc3 release of vLLM brings a critical infrastructure tweak to the new Model Runner V2: support for randomized dummy inputs. This isn't a feature for end-users but a developer-facing fix that stabilizes how the runner handles initial tensor shapes during compilation and warm-up phases. By allowing randomized inputs, it reduces the likelihood of shape-mismatch errors when tracing models with dynamic dimensions. For builders running large-scale inference workloads, this means fewer silent failures and more robust model loading sequences in production environments.

vLLM Releases·Oct 2, 2026

More in Models & Labs

Models & Labsmodels

llama.cpp v0.6.0 adds decision model API and Apple Silicon speedups

This release shifts llama.cpp from a pure text engine to a multimodal inference runtime capable of handling 'decision models' like Clef and GLM-5.3-Flash. The new /v1/systemone server endpoint standardizes how these non-autoregressive models are queried, while the extended batch API allows mixed token and embedding inputs for complex architectures. Apple Silicon users get a tangible performance boost with new Metal MMA kernels that accelerate speculative decoding by up to 3x. It’s a significant step toward supporting the next generation of hybrid reasoning models locally.

llama.cpp Releases·Oct 6, 2026
Falcon-Emirati-7B targets dialect nuance© Hugging Face Blog
Models & Labsmodels

Falcon-Emirati-7B targets dialect nuance

Most Arabic models treat the language as a monolith, missing the cultural and linguistic depth of specific dialects. Falcon-Emirati-7B closes this gap by fine-tuning on native Emirati text, synthetic data constrained by strict glossaries, and cultural heritage knowledge. It tops the new Alyah benchmark with 84.83%, proving that scale alone doesn't buy dialect competence. This release underscores a critical shift: true multilingual capability requires targeted adaptation, not just larger parameter counts.

Hugging Face Blog·Oct 6, 2026
Reflection debuts Beam open-weight model© TechCrunch AI
Models & Labsmodels

Reflection debuts Beam open-weight model

Reflection AI is challenging the Chinese dominance in open-weight models with Beam, a 501B-parameter MoE model that claims to match Z.ai’s GLM-5.2 on reasoning benchmarks while using significantly less inference compute. Backed by $4.7 billion and secured GPU deals worth over $7 billion, this two-year-old startup is positioning itself as the Western alternative to DeepSeek and Qwen for enterprise and sovereign AI deployments. The model targets developers and institutions needing cost-effective, localizable infrastructure rather than just raw API access. With weights releasing this month, Beam offers a tangible option for those looking to reduce reliance on closed labs or Chinese open-source ecosystems.

TechCrunch AI·Oct 5, 2026