Dell, CoreWeave Push AI Reliability Beyond the Rack to Whole Data Centers
Dell and CoreWeave say AI reliability now requires data-center-wide coordination of compute, power and cooling.
Sarat Krishnan, director of PowerEdge AI architecture and systems development engineering at Dell, and Jacob Yundt, vice president of engineering, compute architecture, at CoreWeave, spoke with theCUBE Research's Dave Vellante and John Furrier. They described joint engineering, manufacturing diagnostics, liquid cooling and data center coordination.
Krishnan said data centers differ in how cooling enters—from the top or bottom—and in power whip sizes. He said Dell has learned to build modular rack-scale infrastructure that can adapt quickly to those requirements. Dell's partnership with CoreWeave has helped it adapt successive generations of rack-scale systems to different facilities.
For CoreWeave and Dell, design work begins months or years before new systems enter production. Yundt said firmware settings, mechanical changes and deployment requirements are worked through before the first rack. He described the partnership as tight co-engineering, with CoreWeave embedded at the factory to improve the product, scale it and deploy it, while leveraging Dell's supply chain and engineering experience.
That collaboration also informs validation of the rack as an integrated production system. Krishnan said Dell uses operating feedback from CoreWeave to refine diagnostics across server and rack assembly. He said the goal is to "push your risk to the left," because remediating problems becomes more expensive later in manufacturing and extremely expensive if a part must be replaced in a data center. Over the past two years, he said, Dell has shifted its most complex diagnostics—those meant to capture true hardware failures—further left.
Liquid cooling adds another dimension. CoreWeave's Racky rack manager brings power, cooling and environmental telemetry into a unified control interface. Krishnan said Dell works with CoreWeave to integrate Racky in Dell's L11 factories, so systems are already tested when they ship. He said the partnership inspired better leak detection mechanisms because leaks can be catastrophic.
The operating challenge for Nvidia Corp.'s Vera Rubin systems now extends across compute trays, switches, data processing units and network fabrics. Yundt said coordinating those elements with facility power and liquid cooling requires management across the full infrastructure stack. For Vera Rubin and beyond, he said, "it's no longer the rack [that] is the computer" but rather the row, the data hall or the data center. That means tight systems integration across the entire stack.
SiliconANGLE noted that theCUBE is a paid media partner for the Fully Connected event and that neither CoreWeave, the sponsor of theCUBE's coverage, nor other sponsors had editorial control over theCUBE or SiliconANGLE content.
Editor's Summary Dell and CoreWeave are moving AI infrastructure reliability beyond the rack by co-engineering modular rack-scale systems, shifting hardware diagnostics earlier in manufacturing and integrating power, cooling and leak detection. Their work on Nvidia Vera Rubin systems shows the unit of reliability expanding to rows, data halls and entire data centers. The partnership reflects a broader need for coordination across engineering, manufacturing and operations.