Vast Pitches Tiered Storage to Ease AI Agent Memory Pressure
Vast CTO Alon Horev said long-running AI agents are straining GPU memory with KV caches, and tiered storage can offload sessions to avoid repeated inference computation.
AI agent memory is creating new demands on infrastructure as agents run longer sessions and spread across the enterprise. Retaining that context and making it available when needed puts pressure on memory capacity and data movement. Horev said those demands extend beyond the context held during an individual interaction, because enterprise agents also need shared knowledge that persists across sessions.
“Memory for agents is a bit different,” Horev said. “First of all, there are multiple types of memory. There’s long-term memory where an agent can see past conversations and past interactions and look back and learn from its past experiences.”
The pressure shows up first in inference. Each long-running session holds its KV cache in graphics processing unit memory, and a session of half a million tokens can take up one-tenth to one-twentieth of a GPU’s memory, Horev explained.
“It’s also possible the agent would stop talking to the [large language model] because it’s compiling code, it’s testing software, or, as a human, I want to have a cup of coffee,” he said. “What you see is that if you could stretch that memory wall and basically offload those sessions to storage, you can avoid that repeat recalculation.”
Vast’s approach uses memory in tiers. GPU memory is used first, then central processing unit memory on the same machine, then persistent media that can hold petabytes of KV cache, with Nvidia Corp.’s Dynamo software orchestrating the process, Horev said.
“You can move a session from one busy GPU to one less busy GPU and move KV cache either over the network or read it from Vast,” he said. “Once you look at inference as a distributed problem where you have the opportunity to use GPU memory, CPU memory and Vast across a fleet of machines, you have more optionality and you have more optimized scheduling.”
The stakes rise as companies deploy thousands of agents that handle sensitive data and act on customers’ behalf. Those enterprises need to record everything their agents do and retain it for a set period, which makes AI agent memory both a governance and performance asset, according to Horev. Vast has also launched a confidential computing service for sensitive workloads.
“These conversations that the agent is doing, it’s also gold,” he said. “It’s the same information that it can use for fine-tuning or training or creating purpose-built models.”
The interview was part of SiliconANGLE’s and theCUBE’s coverage of Fully Connected 2026. TheCUBE is a paid media partner for the event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.