
At the SIGGRAPH conference, NVIDIA unveiled its latest advancements in AI technologies for graphics and simulation. Key highlights include new neural rendering techniques and the Cosmos 3 Edge model, which supports real-time processing for physical AI systems. These developments are set to enhance digital content creation, offering more realistic and efficient virtual worlds. NVIDIA's innovations are poised to impact various industries, from gaming to autonomous systems, by providing more integrated and powerful AI tools.
Read original
© NVIDIA BlogBristol Myers Squibb is significantly enhancing its AI capabilities by deploying a second NVIDIA DGX SuperPOD, dubbed the 'SuperDuperPOD', built on eight DGX Vera Rubin NVL72 systems. This move aims to democratize access to powerful AI tools across its research teams, enabling faster drug discovery cycles and broader exploration of chemical spaces. By integrating NVIDIA's BioNeMo Agent Toolkit, BMS is poised to streamline its drug discovery pipeline, allowing scientists to focus more on scientific innovation rather than logistical challenges. This expansion represents a major step in leveraging AI to accelerate therapeutic discoveries and improve patient outcomes.
© NVIDIA BlogNVIDIA's Vera Rubin platform is redefining the economics of AI post-training by maximizing intelligence per dollar. This platform allows for more efficient use of resources, requiring only a quarter of the GPUs compared to previous generations, making continuous post-training economically viable. By integrating with tools like NeMo RL and Nemotron 3 Ultra, Vera Rubin supports large-scale reinforcement learning environments, enhancing the model's ability to adapt and improve in real-time. This shift means AI models can continuously refine their capabilities, offering more value per token served and making AI more adaptable to changing environments.
The b10069 release of llama.cpp brings notable improvements to OpenCL support, particularly targeting Adreno GPUs. By enabling broadcast for Adreno MUL_MAT and respecting view offsets, this update aims to boost performance for multi-stream operations on llama-server. The release also extends general GEMM/GEMV support for broadcast, which could optimize operations across different hardware setups. Although there are no revolutionary new features, these updates represent a consistent enhancement in compatibility and performance, especially for developers working with a range of hardware configurations.
The b10075 release of llama.cpp marks a significant step in enhancing its compatibility across diverse hardware setups. With the addition of ROCm 7.2 support on Ubuntu, AMD GPU users can now enjoy improved performance. Windows users benefit from the inclusion of CUDA 13.3, ensuring better integration with NVIDIA GPUs. The update also brings Vulkan support, which optimizes GPU utilization for developers. Although no new model architectures are introduced, this release reinforces llama.cpp's role as a flexible and adaptable inference runtime for developers working in varied environments.
© TechCrunch AIGoogle is reportedly working on a new AI chip, dubbed 'Frozen v2', aimed at significantly enhancing the efficiency of its Gemini models. Expected to be released by 2028, this chip could be six to ten times more efficient than current AI chips, potentially transforming Google's AI capabilities. This move aligns with a broader industry trend where tech giants are developing custom chips to reduce reliance on Nvidia and address AI computing capacity shortages. The anticipation of this chip has already positively impacted Google's stock, reflecting investor confidence in the company's strategic direction.