
PrismML, a startup founded by Caltech researchers, has released Bonsai 2 27B, a highly compressed version of Alibaba's Qwen3.8 model that fits into 5.9 GB of memory. The company achieved this using ternary weight compression, reducing the model size by up to 10x while retaining 98% of the original benchmark performance. PrismML recently closed a $22.25 million seed round led by Khosla Ventures and Cerberus Capital to further develop its compression technology. The startup plans to apply this technique to larger models in the coming months, aiming for local deployment on PCs and smartphones.
Read original
© TechCrunch AIThe AI infrastructure race just got significantly more expensive. Crusoe’s $3.9 billion Series F round values the company at nearly $31 billion, signaling that capital is flowing aggressively into physical compute capacity rather than just model weights. The funds target 'Spark' modular factories—truck-deployable data centers designed to bypass local zoning battles and accelerate deployment. With major backers like Nvidia and Mubadala, this bet on hardware logistics suggests the bottleneck for AI growth is shifting from algorithms to electricity and real estate.
© TechCrunch AIGoogle DeepMind is formalizing the industry's safety anxiety with a new institute dedicated to debating AGI risks. The move signals a shift from vague concerns to concrete governance proposals, including Demis Hassabis’s call for a U.S.-led standards body that could eventually mandate pre-release model evaluations. Simultaneously, researchers argue against opaque architectures, pushing for limits on 'serial depth' to preserve interpretability. This isn't just PR; it's an attempt to set the regulatory and technical guardrails before the technology outpaces human oversight.
© TechCrunch AIThe FAA is betting $875 million over twelve years on an AI system called SMART to manage airspace. Developed by Air Space Intelligence, the cloud-based platform uses machine learning to predict traffic flows and identify conflicts before they happen. This massive investment signals a shift from manual coordination to algorithmic management in critical infrastructure. The rollout begins in the Washington D.C. metro area, marking one of the largest government contracts for operational AI to date.
This release quietly closes the hardware gap for local inference by adding native support for CUDA 13 and ROCm 10.0 alongside existing CUDA 12 builds. Users with newer NVIDIA GPUs or AMD accelerators no longer need to compile from source to get hardware acceleration, as pre-built binaries now include the necessary libraries. Apple Silicon builds have reverted KleidiAI to disabled by default, likely due to stability concerns, though it remains available. The inclusion of openEuler support for Huawei's Ascend 910b chips further expands the ecosystem beyond standard x86 and ARM architectures. This is a critical infrastructure update that ensures llama.cpp remains viable on the latest AI hardware without requiring developer intervention.
© The AI Daily BriefTypeSafe introduces Jev, a model that outputs calibrated probabilities instead of text, claiming 20-200x speed improvements over traditional LLMs for decision tasks.
© AI ExplainedOpenAI announced a partnership or tool update leveraging ChatGPT to accelerate the discovery of new antibiotics, addressing critical bottlenecks in drug development.