NVIDIA Unveils Rubin Platform to Drastically Lower AI Training and Inference Costs

NVIDIA launches its next-generation Rubin platform, an AI supercomputer built from six co-designed chips that cuts inference token costs by up to 10 times. The new system aims to accelerate agentic AI and massive-scale models while significantly reducing the number of required GPUs.

NVIDIA launches the Rubin platform at CES, introducing a new AI supercomputer built from six co-designed chips that work together to slash training time and inference token costs. Compared to the previous Blackwell platform, Rubin delivers up to a 10x reduction in inference token cost and a 4x reduction in the number of GPUs required to train mixture-of-experts models. Founder and CEO Jensen Huang states that this annual generation of AI supercomputers arrives exactly as global AI computing demand skyrockets.

The platform relies on extreme codesign across the NVIDIA Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch. This integrated approach introduces five major innovations, including upgraded NVLink interconnect technology, a new Transformer Engine, Confidential Computing, and a RAS Engine. Additionally, the new NVIDIA Inference Context Memory Storage Platform utilizes the BlueField-4 processor to specifically accelerate advanced agentic AI reasoning.

Major technology partners are already aligning with the new architecture to scale their AI infrastructure. Microsoft plans to deploy NVIDIA Vera Rubin NVL72 rack-scale systems in its next-generation Fairwater AI superfactories, scaling to hundreds of thousands of superchips. CoreWeave becomes one of the first providers to offer Rubin systems through its Mission Control platform, while Red Hat collaborates with NVIDIA to deliver a complete AI software stack optimized for the new hardware using Red Hat Enterprise Linux, OpenShift, and Red Hat AI.

Read More at the original source →