
ZML, a French AI startup, has released ZML/LLMD, a free software designed to enhance inference performance across various AI chips, including Nvidia, AMD, and Google's TPU. The software aims to break vendor lock-in and optimize chip usage, potentially reducing costs and energy consumption for enterprises. Although not open source, ZML/LLMD is available for free, allowing the company to learn from user interactions. This launch positions ZML as a key player in the AI inference market, challenging existing paradigms and promoting chip diversity.
Read original
© TechCrunch AIOpenAI's acquisition of Glass Imaging for over $300 million signals a strategic move into hardware, leveraging AI to enhance smartphone camera capabilities. Glass Imaging, founded by former Apple engineers, specializes in using neural networks to improve image quality at the moment of capture, rather than post-processing. This acquisition aligns with rumors of OpenAI's interest in developing its own hardware, potentially including smartphones and AI companion devices. The move could position OpenAI to integrate advanced AI-driven imaging technology into future products, expanding its influence beyond software.
© TechCrunch AIiOS 27 marks a significant leap for Siri, transforming it from a basic assistant into a more sophisticated AI tool. Built on Google's Gemini models, Siri now handles complex requests and contextual tasks, making it a more integral part of the iOS experience. Users can ask Siri to perform multistep actions, fetch information from emails, and even interact with the Camera app for real-time insights. This update positions Siri as a more reliable and versatile assistant, encouraging users to rely on it for more than just simple tasks. The integration with third-party apps could further enhance its utility as developers adapt to the new capabilities.
© TechCrunch AIMicrosoft has introduced a new AI code of conduct aimed at ensuring AI models adhere to safety and ethical standards. This document outlines principles to prevent AI from engaging in harmful activities like cyberattacks or creating deepfakes. It emphasizes the importance of AI supporting human endeavors rather than replacing them, and includes strict guidelines to maintain human oversight. This move reflects the growing industry focus on AI safety, aligning Microsoft with other major players like OpenAI and Anthropic in prioritizing responsible AI development.
The latest release of llama.cpp, b10955, tackles a critical issue of heap corruption by disabling the ggml-cpu precompiled header and fixing CACHE_LINE_SIZE ambiguity. This update ensures consistent CACHE_LINE_SIZE values across C++ kernels and C work-buffer sizing code, preventing heap-buffer-overflow and subsequent crashes. By restoring the natural include order and removing the std::hardware_destructive_interference_size branch, the update makes the value deterministic and include-order independent. This release is a technical fix that stabilizes the runtime environment for developers using llama.cpp.
The latest llama.cpp release, b10956, introduces significant improvements to the SYCL backend, particularly for handling large k values in TOP_K operations. By implementing a radix select method, the update allows for efficient GPU-resident processing, avoiding previous limitations that forced operations to fall back to the CPU. This change enhances performance, especially in scenarios requiring large k values, such as qwen4exp's sparse-attention indexer. The update ensures that operations are more efficient and scalable, providing a notable boost in processing speed without regressing any measured shapes.
The b10970 release of llama.cpp enhances its reach by incorporating fp32 accumulators in fattn-mma on CDNA devices, boosting performance on specific hardware. This update extends compatibility across macOS, Linux, Windows, and openEuler, with particular attention to CUDA and ROCm libraries. Although there are no new models introduced, the release reinforces llama.cpp's role as a flexible inference runtime, accommodating a wide array of hardware setups. Developers can now enjoy improved performance and broader deployment options, making it easier to integrate AI models into different environments.