Just three months after CEO Jensen Huang personally unveiled the Kyber NVL144 on stage at GTC, Nvidia's most ambitious AI rack system has hit a serious roadblock. The next-generation architecture has been delayed by more than a year to 2028 due to manufacturing problems with its critical midplane printed circuit board (PCB), according to semiconductor analyst firm SemiAnalysis.
The Kyber NVL144 was designed to be Nvidia's densest rack-scale system, housing 144 GPUs with an all-copper NVLink interconnect. The architecture was central to Nvidia's roadmap for scaling AI compute for hyperscalers and enterprises pushing the limits of large model training and inference.
"The Kyber NVL144 rack architecture has been delayed to 2028 as the PCB midplane remains challenging from a manufacturability standpoint," SemiAnalysis reported. The midplane is the component responsible for enabling the dense, high-bandwidth NVLink connections between GPUs — essentially the backbone of the entire system.
Making matters worse, the associated NVL72x2 back-to-back rack design has been outright canceled. That configuration was meant to serve as a bridge to Kyber by linking two NVL72 racks together for expanded scalability. Its cancellation effectively caps the scaling ceiling of Nvidia's current-generation systems.
The Rubin Ultra architecture is also now on hold, with its progress dependent on the maturation of co-packaged optics — a technology that integrates optical components directly into chip packages. Reports indicate Rubin Ultra has been scaled back from quad-chip to dual-chip variants, further constraining what Nvidia can deliver next.
Nvidia continues shipping its existing Oberon and Rubin rack systems to meet current demand. For now, hyperscalers and enterprises betting on Nvidia's next leap in AI compute density will have to wait — and the entire industry's scaling roadmap may need rewriting.
Analysts say the delay could create opportunities for competitors like AMD and custom ASIC players, while also potentially cooling the pace of frontier AI model training progress, which depends on ever-denser compute clusters.




