
OpenAI has published the first results from its proprietary Jalapeño inference chip, designed to optimize the cost and speed of running large language models. The data indicates improvements in inference efficiency, which is crucial for scaling AI services profitably. This move represents OpenAI's continued effort to reduce dependency on third-party hardware providers like NVIDIA by developing custom silicon tailored to its specific model architectures. The release provides early insights into how custom chips might reshape the economics of AI deployment.
Read original
© AI ExplainedOpenAI announced a partnership or tool update leveraging ChatGPT to accelerate the discovery of new antibiotics, addressing critical bottlenecks in drug development.
© AI ExplainedAnthropic released its threat intelligence report for September 2026, detailing new patterns of AI misuse and strategies to counter them.
© AI ExplainedOpenAI announced that its AI systems have successfully solved the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems.
This release quietly closes the hardware gap for local inference by adding native support for CUDA 13 and ROCm 10.0 alongside existing CUDA 12 builds. Users with newer NVIDIA GPUs or AMD accelerators no longer need to compile from source to get hardware acceleration, as pre-built binaries now include the necessary libraries. Apple Silicon builds have reverted KleidiAI to disabled by default, likely due to stability concerns, though it remains available. The inclusion of openEuler support for Huawei's Ascend 910b chips further expands the ecosystem beyond standard x86 and ARM architectures. This is a critical infrastructure update that ensures llama.cpp remains viable on the latest AI hardware without requiring developer intervention.
© TechCrunch AIPrismML is proving that extreme model compression doesn't have to mean dumb models. Their Bonsai 2 27B model shrinks Alibaba's Qwen3.8 down to just 5.9 GB using ternary weights, hitting 98% of the original benchmark scores. This isn't just a technical curiosity; it means high-performance reasoning can finally run on consumer hardware without cloud dependency. With $22.25M in seed funding and backing from Khosla Ventures, they are positioning themselves as the bridge between massive lab models and private, local inference.
© TechCrunch AIHuawei is accelerating its next-generation Ascend 960DT chip to Q1 2027, cutting the original timeline by six months. This move signals a push to scale its Peerium Computing Architecture, which aims to link hundreds of thousands of accelerators into a single massive system. While the chip release is earlier than expected, there are conflicting reports about the scaling capacity of their Atlas SuperCluster, with some analysts noting a reduction in claimed node counts. The shift highlights Huawei's determination to build domestic AI infrastructure despite U.S. export restrictions. It marks a tangible step in China's effort to reduce reliance on Nvidia for large-scale training and inference workloads.