SurrealDB CEO Warns Memory Is the Hidden Cost of Enterprise AI Agents
In a TechRadar article, SurrealDB's CEO says multi-agent AI systems can cost far more than chatbots because they continuously write, share and retain memory. He outlines concurrency, shared state and active-versus-historical storage as key design principles.
As tasks evolve, agents juggle plan development, outbound tool calls, updating records and recording results. When that activity is multiplied across hundreds or thousands of agents, the data layer starts behaving differently, with significant implications for both cost and accuracy. Those implications are becoming major considerations as businesses move agentic AI projects from pilot to production, the article said.
Early pilots tend to be narrow by design. Typically, one team works with a limited dataset and relatively simple, short-lived interactions, such as requesting a summary of a document. The model completes a request and the task is done. In that setting, model access, tokens and context retrieval appear to account for most of the bill. When shifting from passive chatbots to multi-agent systems, however, the economics flip, according to the article. Traditional generative AI costs scale linearly with users, while the cost of running multi-agent deployments tends to escalate with task complexity. Proof-of-concept expenditure typically multiplies as tasks become more complex.
The article cited AWS estimates: a text-based proof of concept handling about 100 interactions a day costs around $40 a month, while an agent-based proof of concept using a knowledge base and guardrails at roughly the same volume costs around $840 a month. Anthropic's write-up on its multi-agent research supports that finding, the article said. The data shows multi-agent systems use 15 times more tokens than chat interactions, and economic viability requires the value of the task to be high enough to justify the expense.
In multi-agent systems, costs accrue because agents use dynamic loops, tool calls, retries and context transfers to complete a task. Agents also need to remember and reason across increasing volumes of data. They have longer operational lives than a typical chatbot session. They must remember yesterday's actions, preserve the evidence behind a decision and make that information available to other agents in the fleet. Every completed task becomes a new state. As the agent fleet grows, the business pays not only for the next answer but also for the expanding operational provenance required to operate efficiently and safely.
Early generative AI systems often treated memory as context, retrieved when required. Persistent agents create a different requirement: memory becomes a live operational system that is constantly updated. That matters for both performance and cost. A narrow write path can create contention as more agents update records. Separate copies of the same context increase storage and synchronization work, and keeping every historical record on the fastest storage tier makes rarely used information unnecessarily expensive.
When designing the data layer to support agentic fleets, the service level dictates how memory is accessed and managed, the article said. Leaders should establish how quickly each kind of memory must be available, who can update it, which agents should share it, and how long it must remain immediately accessible.
The article outlined three principles. The first is concurrency: leaders should ask whether write capacity can grow as the number of agents grows, because infrastructure designed mainly for read-heavy applications may behave very differently once large numbers of autonomous processes start updating state at the same time. The second is shared memory: if several agents work on the same customer, asset or process, creating separate copies of the same context can introduce cost and inconsistency, while a common source of state can make collaboration easier if access controls and provenance remain clear. The third is separating active memory from historical memory. An agent may need the last few minutes of a workflow immediately, while records from six months ago might only be required for an audit or an unusual query. Treating those categories identically is an expensive default, the article said.
Frequently accessed state can remain close to the compute layer, while older information moves to lower-cost durable storage. Compute can then scale according to current activity rather than the total memory accumulated since the system was launched. The article said how historical changes are stored poses a similar question.