16 × AIAI signal, amplified
AI newsTopicsAboutSources
TelegramFollow on Telegram
AI newsTopicsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Newsletter

Used only to send this newsletter. Privacy

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.A new issue every two days.
Home/Models & Labs
Models & Labs

llama.cpp Adds WebGPU Depthwise Conv2D Kernel

llama.cpp Releases·July 23, 2026·high confidence

Why it matters

  • →Expands llama.cpp's capabilities with WebGPU, enhancing performance for AI tasks.
  • →Provides developers with more flexibility in choosing computational backends.
  • →Strengthens llama.cpp's position as a versatile AI framework across platforms.

Llama.cpp has released an update adding a depthwise convolutional 2D kernel to its WebGPU backend. This kernel, originally from the Vulkan backend, enhances the framework's computational capabilities. The update also includes minor code cleanups and updates to supported operations tables. This development is significant for developers using WebGPU, as it expands the framework's versatility and performance in handling AI tasks.

Read original

The story around this

Earlier coverage that leads up to this article, and what followed. Lines connect each piece to the closest one after it, converging here.

llama.cpp b9565 Release Enhances WebGPU Support — llama.cpp Releases1llama.cpp b9684 Release Adds 3D Convolution — llama.cpp Releases2llama.cpp Adds WebGPU Depthwise Conv2D KernelJun 9You are here

How we got here

  1. 1
    llama.cpp b9565 Release Enhances WebGPU Support

    llama.cpp Releases · June 9, 2026 · Same story

  2. 2
    llama.cpp b9684 Release Adds 3D Convolution

    llama.cpp Releases · June 18, 2026 · Same story

Follow this story

Open the full story →

llama.cpp b10058 release enhances Vulkan support

3 developments

  1. Jul 19 · llama.cpp Releases
    llama.cpp b10058 release enhances Vulkan support
  2. Jul 23 · llama.cpp Releases
    llama.cpp Adds WebGPU Depthwise Conv2D Kernel (This article)↳ llama.cpp adds a WebGPU Depthwise Conv2D kernel ported from Vulkan
  3. Jul 27 · llama.cpp Releases
    Llama.cpp Adds Vision Support with MiniMax-M3↳ llama.cpp adds preliminary MiniMax-M3 vision support with per-head QK-norm and partial rotary

More from llama.cpp Releases

Coding Toolscoding

llama.cpp fixes MoE dispatch inefficiency on Vulkan

This release targets a specific but costly bottleneck in Mixture-of-Experts inference on GPUs. The previous tile selection logic wasted significant compute time by misjudging the active workload per expert during dispatch. By correcting how matmul tiles are assigned, the patch ensures workers stay busy instead of idling. This is a quiet optimization that directly improves throughput for large MoE models running on Vulkan backends.

llama.cpp Releases·Sep 30, 2026
Coding Toolscoding

llama.cpp Vulkan perf boost on Intel Arc

Intel's discrete GPUs have long been second-class citizens in local inference due to inefficient memory access patterns. This patch fixes that by batching F32 matrix loads two at a time, squeezing significant throughput out of the B60 architecture. Benchmarks show raw GFLOPS jumping from 153 to 221 on specific shapes, proving that driver-level optimizations matter as much as model architecture. It’s a quiet but necessary fix for anyone running llama.cpp on AMD or Intel hardware.

llama.cpp Releases·Sep 30, 2026
Coding Toolscoding

llama.cpp b11267 adds ROCm 10 and Snapdragon support

This release quietly cements llama.cpp as the universal inference runtime by finally bringing first-class ROCm 10.0 support to both Linux and Windows. AMD GPU users no longer need workarounds for modern hardware, effectively closing a long-standing parity gap with NVIDIA's CUDA ecosystem. Equally notable is the new Snapdragon build for Linux, which targets the emerging AI PC market by leveraging Adreno GPUs and Hexagon NPUs directly. While KleidiAI on Apple Silicon has been disabled in this specific binary set, the expansion into non-NVIDIA silicon signals a strategic shift toward hardware agnosticism that benefits anyone running local models outside of standard data centers.

llama.cpp Releases·Sep 30, 2026

More in Models & Labs

OpenAI launches collaborative office suite features© TechCrunch AI
Models & Labsproductivity

OpenAI launches collaborative office suite features

OpenAI is directly challenging Microsoft’s dominance by embedding a full office suite into ChatGPT. The new 'Space' feature acts as a shared workspace where users and AI agents collaborate in real-time, while 'Pages' serves as an agent-native word processor. Collaborative slides allow teams to co-edit presentations generated from conversation. This marks a strategic pivot from pure chatbot utility to comprehensive workplace infrastructure, forcing incumbents to defend their core business models against an AI-native competitor.

TechCrunch AI·Sep 29, 2026
OpenAI delays GPT-6.1 Astra over safety risks© Wes Roth
Investment · $42 billion net loss
Models & Labsmodels

OpenAI delays GPT-6.1 Astra over safety risks

OpenAI has paused the release of GPT-6.1 Astra, citing significant safety concerns that outweigh its performance gains. This decision underscores a growing tension in the industry: as models become more capable, they also become harder to constrain within safe operational boundaries. While Anthropic pushes forward with Claude Sonnet 5.5 and reports massive financial losses ahead of an IPO, OpenAI is choosing caution over speed. The delay signals that safety evaluations are becoming a primary bottleneck for next-generation model releases, potentially shifting the competitive landscape toward more conservative development cycles. By halting Astra, OpenAI acknowledges that raw capability alone no longer guarantees a viable product launch. This move forces competitors to reconsider their own release timelines and safety protocols. The industry now watches closely to see if this caution becomes the new standard for frontier AI development.

Wes Roth·Sep 29, 2026
OpenAI expands ChatGPT plugins with app-like interfaces© TechCrunch AI
Models & Labsproductivity

OpenAI expands ChatGPT plugins with app-like interfaces

OpenAI is shifting ChatGPT from a chat interface to an application platform by introducing dedicated sidebar panels and interactive tools for third-party developers. This move transforms static integrations into persistent, app-like experiences where users can view files and manage data without leaving the conversation. By supporting the MCP Events specification, OpenAI enables plugins to trigger automations based on external events, bridging the gap between conversational AI and workflow execution. The redesign of the plugin directory and creator tools aims to lower barriers for developers while giving users finer control over permissions. This effectively positions ChatGPT as a central hub for enterprise productivity rather than just a search or writing assistant.

TechCrunch AI·Sep 29, 2026