
TypeSafe has launched Jev, a new class of AI 'judgment model' designed to skip text generation entirely in favor of returning calibrated probabilities. The company claims this approach delivers speeds 20 to 200 times faster than standard large language models (LLMs) at a fraction of the cost. This release highlights a growing industry shift toward specialized models that handle small, discrete business decisions rather than generative tasks, positioning Jev as a component within a broader model stack alongside traditional LLMs.
Read original
© The AI Daily BriefSalesforce has released its first in-house AI model in years, marking a strategic shift from relying solely on third-party providers to developing proprietary foundation models.
© The AI Daily BriefDonald Trump posts multiple times calling AI extinction risk a hoax comparable to global warming debates.
© The AI Daily BriefZAI secures $5 billion in funding to develop a self-training loop for AI models.
This release quietly closes the hardware gap for local inference by adding native support for CUDA 13 and ROCm 10.0 alongside existing CUDA 12 builds. Users with newer NVIDIA GPUs or AMD accelerators no longer need to compile from source to get hardware acceleration, as pre-built binaries now include the necessary libraries. Apple Silicon builds have reverted KleidiAI to disabled by default, likely due to stability concerns, though it remains available. The inclusion of openEuler support for Huawei's Ascend 910b chips further expands the ecosystem beyond standard x86 and ARM architectures. This is a critical infrastructure update that ensures llama.cpp remains viable on the latest AI hardware without requiring developer intervention.
© TechCrunch AIPrismML is proving that extreme model compression doesn't have to mean dumb models. Their Bonsai 2 27B model shrinks Alibaba's Qwen3.8 down to just 5.9 GB using ternary weights, hitting 98% of the original benchmark scores. This isn't just a technical curiosity; it means high-performance reasoning can finally run on consumer hardware without cloud dependency. With $22.25M in seed funding and backing from Khosla Ventures, they are positioning themselves as the bridge between massive lab models and private, local inference.
© TechCrunch AIHuawei is accelerating its next-generation Ascend 960DT chip to Q1 2027, cutting the original timeline by six months. This move signals a push to scale its Peerium Computing Architecture, which aims to link hundreds of thousands of accelerators into a single massive system. While the chip release is earlier than expected, there are conflicting reports about the scaling capacity of their Atlas SuperCluster, with some analysts noting a reduction in claimed node counts. The shift highlights Huawei's determination to build domestic AI infrastructure despite U.S. export restrictions. It marks a tangible step in China's effort to reduce reliance on Nvidia for large-scale training and inference workloads.