16 × AIAI signal, amplified
AI newsAboutSources
TelegramFollow on Telegram
AI newsAboutSources
16 × AIAI signal, amplified

An AI news engine that ingests trusted sources, scores with Claude, and posts only what clears the bar.

Follow on Telegram →

Subscribe

  • Telegram
  • RSS
  • All channels

Legal

  • Privacy
  • Imprint
© 2026 16 × AI. All rights reserved.Curated by Claude. Posts every 6 hours. No newsletter, no funnel.
Home/Models & Labs
Models & LabsInvestment · $22.25M

PrismML compresses LLMs to run locally with minimal loss

TechCrunch AI·September 17, 2026·high confidence

Why it matters

  • →Ternary weight compression achieves near-parity with uncompressed models, challenging the assumption that small models are weak.
  • →A 5.9 GB footprint makes advanced reasoning models viable for consumer devices without cloud infrastructure.
  • →Strong academic backing and early traction suggest this compression method could become a standard for local AI.
PrismML compresses LLMs to run locally with minimal loss
©TechCrunch AI

PrismML, a startup founded by Caltech researchers, has released Bonsai 2 27B, a highly compressed version of Alibaba's Qwen3.8 model that fits into 5.9 GB of memory. The company achieved this using ternary weight compression, reducing the model size by up to 10x while retaining 98% of the original benchmark performance. PrismML recently closed a $22.25 million seed round led by Khosla Ventures and Cerberus Capital to further develop its compression technology. The startup plans to apply this technique to larger models in the coming months, aiming for local deployment on PCs and smartphones.

Read original

More from TechCrunch AI

Crusoe raises $3.9B for modular AI data centers© TechCrunch AI
Investment · $3.9B
Market & Regulationbusiness

Crusoe raises $3.9B for modular AI data centers

The AI infrastructure race just got significantly more expensive. Crusoe’s $3.9 billion Series F round values the company at nearly $31 billion, signaling that capital is flowing aggressively into physical compute capacity rather than just model weights. The funds target 'Spark' modular factories—truck-deployable data centers designed to bypass local zoning battles and accelerate deployment. With major backers like Nvidia and Mubadala, this bet on hardware logistics suggests the bottleneck for AI growth is shifting from algorithms to electricity and real estate.

TechCrunch AI·Sep 17, 2026
DeepMind launches institute for AGI debate© TechCrunch AI
General AIother

DeepMind launches institute for AGI debate

Google DeepMind is formalizing the industry's safety anxiety with a new institute dedicated to debating AGI risks. The move signals a shift from vague concerns to concrete governance proposals, including Demis Hassabis’s call for a U.S.-led standards body that could eventually mandate pre-release model evaluations. Simultaneously, researchers argue against opaque architectures, pushing for limits on 'serial depth' to preserve interpretability. This isn't just PR; it's an attempt to set the regulatory and technical guardrails before the technology outpaces human oversight.

TechCrunch AI·Sep 17, 2026
FAA awards $875M AI contract for air traffic control© TechCrunch AI
Market & Regulationother

FAA awards $875M AI contract for air traffic control

The FAA is betting $875 million over twelve years on an AI system called SMART to manage airspace. Developed by Air Space Intelligence, the cloud-based platform uses machine learning to predict traffic flows and identify conflicts before they happen. This massive investment signals a shift from manual coordination to algorithmic management in critical infrastructure. The rollout begins in the Washington D.C. metro area, marking one of the largest government contracts for operational AI to date.

TechCrunch AI·Sep 17, 2026

More in Models & Labs

Models & Labscoding

llama.cpp b11025 adds CUDA 13 and ROCm 10.0 support

This release quietly closes the hardware gap for local inference by adding native support for CUDA 13 and ROCm 10.0 alongside existing CUDA 12 builds. Users with newer NVIDIA GPUs or AMD accelerators no longer need to compile from source to get hardware acceleration, as pre-built binaries now include the necessary libraries. Apple Silicon builds have reverted KleidiAI to disabled by default, likely due to stability concerns, though it remains available. The inclusion of openEuler support for Huawei's Ascend 910b chips further expands the ecosystem beyond standard x86 and ARM architectures. This is a critical infrastructure update that ensures llama.cpp remains viable on the latest AI hardware without requiring developer intervention.

llama.cpp Releases·Sep 18, 2026
TypeSafe Launches Jev Judgment Model© The AI Daily Brief
Models & Labsmodels

TypeSafe Launches Jev Judgment Model

TypeSafe introduces Jev, a model that outputs calibrated probabilities instead of text, claiming 20-200x speed improvements over traditional LLMs for decision tasks.

The AI Daily Brief·Sep 17, 2026
OpenAI Accelerates Antibiotic Discovery with ChatGPT© AI Explained
Models & Labsother

OpenAI Accelerates Antibiotic Discovery with ChatGPT

OpenAI announced a partnership or tool update leveraging ChatGPT to accelerate the discovery of new antibiotics, addressing critical bottlenecks in drug development.

AI Explained·Sep 16, 2026