NetApp and Nvidia split data and metadata for AI factory storage
NetApp and Nvidia are co-engineering Novus, a storage architecture for AI factories that separates data and metadata to scale independently. Executives say agentic AI and enterprise inference are raising concurrency and throughput demands.
The work builds on a partnership that dates back more than a decade and now includes co-engineering around AI infrastructure. NetApp’s Novus architecture is designed to separate data and metadata functions so each can scale on different axes according to workload demands, according to Arindam Banerjee, NetApp’s chief platform and technology officer. He said metadata transactions and large sequential operations such as checkpoints often occur at the same time. Previous architectures could handle one or the other well but not both at once, he said.
Banerjee described the split as a way to prevent smaller transactional operations from competing with large data transfers for the same resources. If metadata and data share the same media and network, he said, the system becomes inefficient because data operations queue behind metadata operations. That can leave costly GPUs underutilized while they continue to consume power. The architectural shift is intended to restore efficiency and keep GPUs fed.
Jason Hardy, vice president of storage technology at Nvidia, said AI infrastructure also needs to scale without forcing enterprises to redesign the system each time a new use case comes online. The goal is greater flexibility across fine-tuning, inference, capacity and AI-ready data as production workloads change and demand shifts across different parts of the infrastructure. He said enterprises are moving past experimentation into mainstream production. AI agents are starting to show up, and enterprise-scale inferencing is happening, Hardy said.
Agentic workloads add another kind of pressure because thousands of agents may need to access data at the same time. Their permissions may also be temporary and narrowly scoped, increasing metadata operations alongside the throughput demands already placed on storage systems. Banerjee said concurrency is becoming as important as throughput. He described thousands of agents hitting data simultaneously and authorizing themselves, creating concurrency requirements that differ from traditional storage demands. Metadata access through concurrent agents is redefining how these systems will be built, he said.
The operational shift also changes who interacts with the storage layer. AI teams may need to provision and consume infrastructure without becoming storage specialists, making APIs, SDKs and agent-friendly controls more important as these environments move into production and become part of everyday enterprise operations. Banerjee said the aim is to remove complexity for AI teams, which are not storage engineering teams. They should be able to consume storage through an API or code to an API or SDK, and in the future agents may do that work, he said. NetApp is designing a new control plane that allows consumption of Novus through APIs, according to Banerjee.
Banerjee and Hardy spoke with Christophe Bertrand and Rebecca Knight at NetApp INSIGHT. The interview was part of SiliconANGLE’s and theCUBE’s coverage of the event. TheCUBE is a paid media partner, according to the disclosure accompanying the interview.
Editor's Summary
NetApp and Nvidia are co-engineering Novus, a storage architecture for AI factories that separates data and metadata to scale independently and avoid leaving GPUs idle. Executives said agentic AI and enterprise inference are raising concurrency and throughput demands, requiring API-driven, agent-friendly storage controls. The discussion took place at NetApp INSIGHT on theCUBE, with TheCUBE disclosed as a paid media partner.