16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & Labs

Kimi K3: Largest Open-Weight AI Model Released

Together AI Blog·August 1, 2026·high confidence

Why it matters

  • →Kimi K3 sets a new benchmark as the largest open-weight AI model, expanding possibilities for open-source AI development.
  • →Its advanced architecture supports complex tasks, making it a competitive alternative to proprietary models.
  • →The model's release enhances accessibility to cutting-edge AI capabilities for developers and researchers.
Kimi K3: Largest Open-Weight AI Model Released
©Together AI Blog

Moonshot AI has released Kimi K3, a 2.8-trillion-parameter model, making it the largest open-weight AI model to date. Designed for advanced tasks such as long-horizon coding and deep reasoning, Kimi K3 competes with leading proprietary models like GPT 5.6 Sol. It features new architectural innovations, including Kimi Delta Attention and Attention Residuals, which improve its ability to process long sequences and complex tasks. This release represents a significant advancement in open-source AI, providing developers with a powerful new tool for sophisticated AI applications.

Read original

More from Together AI Blog

Together AI Introduces Autoscaling for LLM Inference© Together AI Blog
Models & Labsmodels

Together AI Introduces Autoscaling for LLM Inference

Together AI has unveiled a sophisticated autoscaling feature for LLM inference on its platform, allowing deployments to scale based on metrics like in-flight requests and GPU utilization. This approach addresses the unique challenges of LLM serving, where traditional autoscaling methods fall short due to misleading CPU-style metrics and lengthy cold starts. By offering a catalog of inference-native metrics, Together AI enables users to fine-tune their autoscaling policies, balancing cost and performance. This development marks a significant step in optimizing resource allocation for AI deployments, particularly in GPU-constrained environments.

Together AI Blog·Jul 31, 2026

More in Models & Labs

Models & Labsmodels

llama.cpp b10227 release enhances tool parsing

The latest b10227 release of llama.cpp introduces a specialized parser for Qwen3, enhancing its tool parsing capabilities. This update includes a tagged thinking tool parser and refactoring efforts to improve functionality, such as the addition of a permute helper and support for omitting <tool_call>. These changes aim to streamline the parsing process and improve the overall efficiency of the system. While the release doesn't introduce new models, it strengthens the existing framework, making it more robust for developers working with complex parsing tasks.

llama.cpp Releases·Aug 3, 2026
Models & Labsmodels

llama.cpp b10231 Release Enhances DSpark Support

The b10231 release of llama.cpp brings significant improvements to DSpark sidecar resolution, making it the default choice over DFlash due to its additional Markov head. This update allows for resolving sidecars without needing a full model at the tag, and provides an option to disable discovery with explicit -md selection. While no new models are introduced, the release extends platform support across macOS, Linux, Windows, and openEuler, enhancing its adaptability for developers. With these changes, llama.cpp continues to evolve as a flexible inference runtime, catering to a wide range of system configurations.

llama.cpp Releases·Aug 3, 2026
Models & Labsmodels

llama.cpp b10232 release enhances Metal support

The b10232 release of llama.cpp brings notable improvements with the introduction of DeepSeek V4 hyper-connections, specifically optimized for Metal. This update features new SIMDgroup register and shuffle optimized kernels, enhancing performance on macOS and iOS devices. The release also extends support to platforms like Ubuntu and Windows, with targeted optimizations for Vulkan, ROCm, and CUDA environments. While no new models are introduced, these enhancements reinforce llama.cpp's capability as a flexible inference runtime, catering to a wide range of hardware setups.

llama.cpp Releases·Aug 3, 2026