NetApp and Iterate.ai Pitch Private AI Appliance; CoreWeave Flags GPU Networking Impact
NetApp and Iterate.ai detailed a private AI appliance that runs on corporate storage, while CoreWeave said GPU networking and power can swing AI latency by orders of magnitude.
At NetApp INSIGHT, Ashish Dhawan, senior vice president, general manager and chief revenue officer of NetApp's Cloud Business Unit, said private AI usage has matured. Open-weight and closed models are getting close to the frontier, he said, and that is where NetApp's AIPod Mini with Iterate can work well.
The partnership centers on AIPod Mini, a turnkey appliance that pairs NetApp infrastructure with Iterate.ai's Generate platform. Brian Sathianathan, co-founder and chief technology officer of Iterate.ai, said the appliance connects to NetApp storage and communicates using the ONTAP protocol. Generate software, with an embedded large language model, runs on top of AIPod Mini so customers can run AI queries locally and privately within their own walls, he said.
Jon Nordmark, co-founder and chief executive officer of Iterate.ai, said the system ships with more than 200 agent templates, more than 200 skills and access to more than 800 tools. In one insurance deployment, accident-report analysis that normally took an analyst about eight hours was completed in eight minutes. Iterate.ai describes its approach as outcome-based AI, judged by business results rather than technical specifications.
Sathianathan cited healthcare as another example. Hospitals typically collect 80% to 85% of what they bill insurers, and Iterate.ai's forensic revenue cycle agent identified $17.4 million in denied claims for a hospital with $150 million in annual billing, he said. Nordmark added that the workflow chains three agents that read payer contracts, examine incoming denials and draft resubmissions. As autonomous agents take on such work, Sathianathan said governance and permission become critical, and that is where the NetApp platform is powerful.
NetApp also used the event to launch Novus, a storage architecture engineered for zettabyte-scale capacity in AI factories. Nordmark said capable open-weight models pay off only when they can reach the institutional knowledge held in enterprise storage. 'Memory plus context equals knowledge,' he said. 'Memory is what makes AI work. So much of a company's memory is in storage today, and that's what NetApp's here to unleash.'
Separately, at Fully Connected, CoreWeave and LlamaIndex described a neocloud market moving past its origins as a stopgap for scarce GPUs. AI-native startups now choose infrastructure on latency, burst capacity and openness, not just chip availability, according to the discussion. CoreWeave is expanding beyond GPU compute into networking, storage and software as inference demand grows, while LlamaIndex has evolved from an open-source retrieval-augmented generation framework into a model builder that rents compute rather than owning it.
Jerry Liu, co-founder and chief executive officer of LlamaIndex, said the company is effectively a specialized AI lab focused on building models for document parsing and extraction. It post-trains open-weight models, gathers its own datasets and aims to tailor its work at the Pareto frontier of performance, cost and latency, he said. LlamaIndex's workload is about 75% inference and 25% training, processing millions of document pages per day for finance, legal and insurance customers whose paperwork arrives in bursts. The company owns no GPU cluster, so guaranteed capacity matters more than hardware ownership.
Lukas Biewald, senior vice president of AI initiatives at CoreWeave, said capacity alone is not the differentiator. Biewald joined CoreWeave through its acquisition of Weights & Biases, the AI observability startup he co-founded. CoreWeave used the event to launch CoreWeave Forge, a development layer that runs training, inference, evaluation and agent development in one connected environment. Coming from software, Biewald said he initially questioned how much chip configuration mattered, but concluded the difference is massive. 'I'm talking orders of magnitude difference depending on how you do the networking for the chips [and] how you do the power distribution,' he said.
Biewald also positioned openness as a distinction from hyperscalers. He said providers such as Amazon Web Services rely on proprietary application programming interfaces that make workloads hard to move, while CoreWeave follows standard networking protocols recommended by Nvidia and brings broader open-source support. CoreWeave knows customers will host their web services on AWS or Google Cloud rather than CoreWeave, he said, and is comfortable playing with other clouds. Liu added that abundant neocloud capacity, combined with fast-improving coding agents, could open post-training of small open-weight models to more people beyond today's narrow group of specialists.