The b10069 release of llama.cpp introduces several enhancements to OpenCL support, particularly for Adreno GPUs. This update includes support for broadcast in Adreno MUL_MAT and adjustments to honor view offsets, aimed at improving multi-stream operations on llama-server. Additionally, the release provides general GEMM/GEMV support for broadcast, removing unnecessary tests and comments to streamline the codebase. These changes are part of ongoing efforts to enhance performance and compatibility across different hardware platforms.
Read originalThe b10059 release of llama.cpp enhances its platform compatibility, now supporting numerous operating systems and architectures. A key change is the defaulting of Hadamard multiplication to a CPU routine, which may lead to more consistent performance across setups. Although KleidiAI support for Apple Silicon is currently disabled, the release still accommodates platforms like macOS, Windows, and Linux, with configurations such as Vulkan and ROCm 7.2. While no new models are introduced, this update solidifies llama.cpp's role as a flexible inference runtime across diverse hardware environments.
The b10075 release of llama.cpp marks a significant step in enhancing its compatibility across diverse hardware setups. With the addition of ROCm 7.2 support on Ubuntu, AMD GPU users can now enjoy improved performance. Windows users benefit from the inclusion of CUDA 13.3, ensuring better integration with NVIDIA GPUs. The update also brings Vulkan support, which optimizes GPU utilization for developers. Although no new model architectures are introduced, this release reinforces llama.cpp's role as a flexible and adaptable inference runtime for developers working in varied environments.
The b10056 release of llama.cpp continues its trend of broadening platform compatibility, making it a versatile tool for developers across various systems. Notably, this update includes support for Ubuntu with ROCm 7.2, enhancing performance for AMD GPU users. The release also maintains its commitment to diverse hardware by supporting both Intel and Apple Silicon on macOS, as well as Vulkan and OpenVINO on Windows. While no groundbreaking new features are introduced, the steady expansion of supported environments ensures that llama.cpp remains a go-to choice for developers seeking flexibility in AI model deployment.
© TechCrunch AIGoogle is reportedly working on a new AI chip, dubbed 'Frozen v2', aimed at significantly enhancing the efficiency of its Gemini models. Expected to be released by 2028, this chip could be six to ten times more efficient than current AI chips, potentially transforming Google's AI capabilities. This move aligns with a broader industry trend where tech giants are developing custom chips to reduce reliance on Nvidia and address AI computing capacity shortages. The anticipation of this chip has already positively impacted Google's stock, reflecting investor confidence in the company's strategic direction.
© TechCrunch AIThe Model Context Protocol (MCP) is undergoing an update that aims to simplify its application in large-scale AI deployments. By adopting a stateless approach to session IDs, the update addresses the complexities faced by companies operating MCP servers across multiple machines. This change is expected to streamline operations and potentially reduce costs for businesses integrating AI agents. Although end users might not notice the difference, this development is a crucial step in refining AI infrastructure, which is necessary for deploying AI models effectively in real-world scenarios.
© FireshipThinking Machines has unveiled Inkling, a new open-weights model boasting 975 billion parameters. While the model is described as 'deliberately mid,' its release marks a significant step in the ongoing evolution of large language models. This development could provide new opportunities for developers seeking to leverage massive AI models with open access. The introduction of Inkling suggests a shift towards more accessible and customizable AI tools, potentially democratizing the use of advanced AI capabilities.