Supermicro has begun shipping Nvidia's Vera Rubin NVL72 rack-scale systems, the first such systems to reach customers, moving the GPU maker's next-generation platform from reference design into production deployment.
Each rack pairs 72 Rubin GPUs with 36 Vera CPUs as a single liquid-cooled machine. Eighteen 1U compute trays, each holding four GPUs and two CPUs, connect through nine sixth-generation NVLink switch trays that provide 216 TB/s of scale-up bandwidth. A single rack carries 20.7 TB of HBM4 memory and up to 54 TB of LPDDR5X.
Supermicro supplies the full direct-liquid-cooling stack, from cold plates and manifolds to in-row cooling distribution units rated at 1.8 MW each and deployed with N+1 redundancy, plus optional rear-door heat exchangers. Its DCBBS blueprints define balanced bills of materials from 5 MW up to gigawatt scale; one blueprint built around a single Vera Rubin NVL72 scalable unit spans 16 compute racks with 1,152 Rubin GPUs and 331 TB of HBM4. The company says one team handles site survey, design, integration, testing, delivery and support.
'We have spent years building the liquid-cooling stack, the manufacturing capacity, and the deployment teams for exactly this moment,' said CEO Charles Liang. Because the platform's power density makes liquid cooling mandatory rather than optional, the rack, the coolant loop and the facility are sold as one package. Nvidia has said it aims to begin production shipments of Vera Rubin systems in the second half of the year.




