
NVIDIA has launched Cosmos 3 Edge, a 4-billion-parameter model designed for edge devices, available on Hugging Face. This model is tailored for real-time reasoning and action generation in robotics and vision AI, operating efficiently on NVIDIA's edge computing platforms. Cosmos 3 Edge integrates two transformer towers to process multimodal data, enabling robots to understand and interact with their environment. This release enhances the capability of AI systems to perform complex tasks directly on-device, improving their adaptability and efficiency in various settings.
Read original
© Hugging Face BlogNVIDIA and Hugging Face have teamed up to streamline the training and fine-tuning of diffusion models using the NeMo Automodel library. This collaboration allows developers to train models at scale without needing to convert checkpoints or rewrite models, making it easier to integrate new models from the Hugging Face Hub. The integration supports both full and parameter-efficient fine-tuning, offering flexibility in training approaches. This development significantly simplifies the workflow for those working with diffusion models, enabling more efficient and scalable model training.
© Hugging Face BlogNVIDIA's Nemotron 3 Embed models have set a new benchmark in retrieval quality, with the 8B model ranking #1 on the RTEB leaderboard. This collection of embedding models is designed for production-scale retrieval tasks, offering open weights and datasets for customization. The models support multilingual and code retrieval, and are optimized for high-throughput deployment with NVIDIA's NVFP4 technology. This release provides developers with powerful tools for efficient and accurate retrieval, enhancing capabilities in agentic retrieval and reducing operational costs.
The b10069 release of llama.cpp brings notable improvements to OpenCL support, particularly targeting Adreno GPUs. By enabling broadcast for Adreno MUL_MAT and respecting view offsets, this update aims to boost performance for multi-stream operations on llama-server. The release also extends general GEMM/GEMV support for broadcast, which could optimize operations across different hardware setups. Although there are no revolutionary new features, these updates represent a consistent enhancement in compatibility and performance, especially for developers working with a range of hardware configurations.
The b10075 release of llama.cpp marks a significant step in enhancing its compatibility across diverse hardware setups. With the addition of ROCm 7.2 support on Ubuntu, AMD GPU users can now enjoy improved performance. Windows users benefit from the inclusion of CUDA 13.3, ensuring better integration with NVIDIA GPUs. The update also brings Vulkan support, which optimizes GPU utilization for developers. Although no new model architectures are introduced, this release reinforces llama.cpp's role as a flexible and adaptable inference runtime for developers working in varied environments.
© TechCrunch AIGoogle is reportedly working on a new AI chip, dubbed 'Frozen v2', aimed at significantly enhancing the efficiency of its Gemini models. Expected to be released by 2028, this chip could be six to ten times more efficient than current AI chips, potentially transforming Google's AI capabilities. This move aligns with a broader industry trend where tech giants are developing custom chips to reduce reliance on Nvidia and address AI computing capacity shortages. The anticipation of this chip has already positively impacted Google's stock, reflecting investor confidence in the company's strategic direction.