
Thinking Machines has released Inkling, a large language model with 975 billion parameters. The model is open-weights, allowing developers to access and customize it freely. Despite being described as 'deliberately mid,' Inkling represents a significant addition to the landscape of large AI models. This release could enhance the accessibility of advanced AI tools for developers, promoting further innovation in the field.
Read originalThe b10069 release of llama.cpp brings notable improvements to OpenCL support, particularly targeting Adreno GPUs. By enabling broadcast for Adreno MUL_MAT and respecting view offsets, this update aims to boost performance for multi-stream operations on llama-server. The release also extends general GEMM/GEMV support for broadcast, which could optimize operations across different hardware setups. Although there are no revolutionary new features, these updates represent a consistent enhancement in compatibility and performance, especially for developers working with a range of hardware configurations.
The b10075 release of llama.cpp marks a significant step in enhancing its compatibility across diverse hardware setups. With the addition of ROCm 7.2 support on Ubuntu, AMD GPU users can now enjoy improved performance. Windows users benefit from the inclusion of CUDA 13.3, ensuring better integration with NVIDIA GPUs. The update also brings Vulkan support, which optimizes GPU utilization for developers. Although no new model architectures are introduced, this release reinforces llama.cpp's role as a flexible and adaptable inference runtime for developers working in varied environments.
© TechCrunch AIGoogle is reportedly working on a new AI chip, dubbed 'Frozen v2', aimed at significantly enhancing the efficiency of its Gemini models. Expected to be released by 2028, this chip could be six to ten times more efficient than current AI chips, potentially transforming Google's AI capabilities. This move aligns with a broader industry trend where tech giants are developing custom chips to reduce reliance on Nvidia and address AI computing capacity shortages. The anticipation of this chip has already positively impacted Google's stock, reflecting investor confidence in the company's strategic direction.