The b10059 release of llama.cpp has been announced, focusing on expanding platform support and optimizing performance. The update defaults the Hadamard multiplication to a CPU routine, potentially improving performance consistency. While KleidiAI support for Apple Silicon is disabled, the release supports a wide range of platforms, including macOS, Windows, and Linux, with various configurations like Vulkan and ROCm 7.2. This update reinforces llama.cpp's role as a versatile tool for AI inference across multiple hardware setups.
Read originalThe b10069 release of llama.cpp brings notable improvements to OpenCL support, particularly targeting Adreno GPUs. By enabling broadcast for Adreno MUL_MAT and respecting view offsets, this update aims to boost performance for multi-stream operations on llama-server. The release also extends general GEMM/GEMV support for broadcast, which could optimize operations across different hardware setups. Although there are no revolutionary new features, these updates represent a consistent enhancement in compatibility and performance, especially for developers working with a range of hardware configurations.
The b10075 release of llama.cpp marks a significant step in enhancing its compatibility across diverse hardware setups. With the addition of ROCm 7.2 support on Ubuntu, AMD GPU users can now enjoy improved performance. Windows users benefit from the inclusion of CUDA 13.3, ensuring better integration with NVIDIA GPUs. The update also brings Vulkan support, which optimizes GPU utilization for developers. Although no new model architectures are introduced, this release reinforces llama.cpp's role as a flexible and adaptable inference runtime for developers working in varied environments.
The b10056 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across various systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. The release also maintains its commitment to diverse hardware by supporting both Intel and Apple Silicon on macOS, as well as Vulkan and OpenVINO on Windows. While no groundbreaking new features are introduced, the steady expansion of supported environments ensures that llama.cpp remains a go-to choice for developers seeking flexibility in AI model deployment.
© Lev SelectorNous Research, an open-source AI lab, has raised $75 million at a $1.5 billion valuation.
© Lev SelectorTencent has released Hy3, an open-source mixture of experts (MoE) large language model.
© TechCrunch AIClem Delangue, CEO of Hugging Face, underscores the critical role of open source AI, comparing the platform to a GitHub for AI models and datasets. He observes that as companies expand, they often move from expensive proprietary APIs to more affordable open source options, which he believes is essential for democratizing AI technology. Delangue voices concerns about the risk of a few large companies dominating the AI landscape, advocating for openness and transparency, particularly in the field of robotics. This approach is reflected in Hugging Face's decision to focus on capital efficiency rather than traditional fundraising, even declining a significant investment offer from Nvidia to stay true to its open source principles.