
OpenAI has announced its first custom-built inference processor, Jalapeño, created in partnership with Broadcom. This chip is designed to enhance the efficiency of OpenAI's AI models, offering improved performance-per-watt over current alternatives. The development marks a strategic move to reduce dependence on Nvidia GPUs, aligning with similar efforts by tech giants like Google and Amazon. Jalapeño is specifically optimized for inference tasks, which could significantly lower operational costs for OpenAI's real-time coding models. This development underscores OpenAI's commitment to optimizing its entire technology stack for better performance and affordability.
Read original
© TechCrunch AIOpenAI's acquisition of Glass Imaging for over $300 million signals a strategic move into hardware, leveraging AI to enhance smartphone camera capabilities. Glass Imaging, founded by former Apple engineers, specializes in using neural networks to improve image quality at the moment of capture, rather than post-processing. This acquisition aligns with rumors of OpenAI's interest in developing its own hardware, potentially including smartphones and AI companion devices. The move could position OpenAI to integrate advanced AI-driven imaging technology into future products, expanding its influence beyond software.
© TechCrunch AIThe latest release of llama.cpp, b10955, tackles a critical issue of heap corruption by disabling the ggml-cpu precompiled header and fixing CACHE_LINE_SIZE ambiguity. This update ensures consistent CACHE_LINE_SIZE values across C++ kernels and C work-buffer sizing code, preventing heap-buffer-overflow and subsequent crashes. By restoring the natural include order and removing the std::hardware_destructive_interference_size branch, the update makes the value deterministic and include-order independent. This release is a technical fix that stabilizes the runtime environment for developers using llama.cpp.
The latest llama.cpp release, b10956, introduces significant improvements to the SYCL backend, particularly for handling large k values in TOP_K operations. By implementing a radix select method, the update allows for efficient GPU-resident processing, avoiding previous limitations that forced operations to fall back to the CPU. This change enhances performance, especially in scenarios requiring large k values, such as qwen4exp's sparse-attention indexer. The update ensures that operations are more efficient and scalable, providing a notable boost in processing speed without regressing any measured shapes.
iOS 27 marks a significant leap for Siri, transforming it from a basic assistant into a more sophisticated AI tool. Built on Google's Gemini models, Siri now handles complex requests and contextual tasks, making it a more integral part of the iOS experience. Users can ask Siri to perform multistep actions, fetch information from emails, and even interact with the Camera app for real-time insights. This update positions Siri as a more reliable and versatile assistant, encouraging users to rely on it for more than just simple tasks. The integration with third-party apps could further enhance its utility as developers adapt to the new capabilities.