16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & Labs

OpenAI previews Ultrafast API with Cerebras

The Rundown AI·August 14, 2026·high confidence

Why it matters

  • →The Ultrafast tier dramatically increases AI processing speeds, potentially transforming workflows.
  • →It showcases the potential of hardware partnerships in enhancing AI capabilities.
  • →The invite-only nature suggests a strategic rollout, with broader implications for AI accessibility.
OpenAI previews Ultrafast API with Cerebras
©The Rundown AI

OpenAI has introduced a new 'Ultrafast' API tier in collaboration with Cerebras, enhancing the speed of its GPT-5.6 Sol model by up to 14 times. This advancement allows the model to process up to 750 tokens per second, significantly reducing the time required for complex tasks. The Ultrafast tier is currently available as an invite-only preview, with broader access expected as more Cerebras capacity becomes available. This development could revolutionize AI-driven workflows, although pricing information has not yet been released.

Read original

More from The Rundown AI

SpaceXAI launches Grok 4.6 model© The Rundown AI
Models & Labsmodels

SpaceXAI launches Grok 4.6 model

SpaceXAI's Grok 4.6 is making waves in the AI landscape by offering competitive performance at a significantly lower cost compared to its rivals. With benchmark scores that challenge models like Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol, Grok 4.6 is priced at just $2/$6 per million tokens, making it an attractive option for cost-conscious users. Elon Musk's bold claim that Grok is 'objectively No. 1' in terms of intelligence, speed, and cost adds to the intrigue. This release not only elevates SpaceXAI's standing but also pressures competitors to innovate and potentially adjust their pricing strategies.

The Rundown AI·Aug 13, 2026

More in Models & Labs

Models & Labsmodels

llama.cpp adds MiniMax model support

The latest llama.cpp update expands its functionality by integrating the MiniMax-Text-01 and MiniMaxM1ForCausalLM models, enhancing its role in causal language modeling. This release focuses on refining the MiniMax-Text-01 model by eliminating state transpose operations and implementing a logits mask to manage zero-valued embeddings. These adjustments aim to streamline the token sampling process and boost model efficiency. While no new model architectures are introduced, the update significantly refines existing processes, making llama.cpp more robust and efficient for developers working with these specific models.

llama.cpp Releases·Aug 16, 2026
Models & Labsmodels

llama.cpp b10441 Release Updates Load Mode

The latest release of llama.cpp, version b10441, introduces a significant change by replacing deprecated flags with a unified --load-mode argument. This update simplifies the configuration process across scripts, examples, and documentation, making it easier for developers to manage memory mapping and loading options. The release also includes updates to internal warning messages and environment variable documentation, ensuring clarity and consistency. While this update doesn't introduce new features, it streamlines the user experience and reduces potential confusion for developers working with llama.cpp.

llama.cpp Releases·Aug 16, 2026
Models & Labsmodels

b10442 Release Enhances Vulkan Support for Intel Xe

The b10442 release of llama.cpp brings notable improvements to Vulkan support, specifically targeting Intel Xe platforms. By adding SHMEM_STRIDE_PAD and APPLY_SLM_A_RESHAPE for cooperative matrix operations, this update aims to optimize performance on Intel hardware. Additionally, it addresses a critical out-of-bounds read issue in kvalues_mxfp4 initialization, enhancing stability. While these changes are technical, they signify a focused effort to refine performance and compatibility for developers working with Intel's Vulkan drivers. This release doesn't introduce new models but strengthens the existing infrastructure for better efficiency.

llama.cpp Releases·Aug 16, 2026