Futurum Report Warns Agentic AI Is Breaking Per-Token Pricing for Enterprises
Futurum report: agentic AI uses 10 to 100 times more tokens per task, pressuring enterprises to rethink pricing and compute.
Per-token pricing’s appeal is straightforward. A developer can call an application programming interface and have a working prototype by the afternoon without capacity planning or a procurement cycle. The problem is that the meter does not distinguish between a pilot and a production system serving 20,000 employees. Agents are token machines. A chatbot answers a question, while an agent plans, calls tools, checks its work, retries, hands off and summarizes, generating tokens at each step.
Futurum forecasts that agent and reasoning inference will grow by 219% this year, and total inference spending will rise from $120 billion in 2025 to $885 billion by 2030. With a pricing model that scales linearly with consumption, the result is a budget line that grows faster than the value it creates. SiliconANGLE reported that one organization budgeted $1 million for the year and spent it in three months because the initiative was so successful. CIOs and CFOs increasingly describe the same pattern: the most successful AI projects end up costing the most, often with budget estimates far off.
Mazda Marvasti, co-founder and chief executive of Amberd.ai, described the pattern in the report. “When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said. “It starts getting the attention of the CFO and the CIO in terms of how much I’m exactly spending to run this tool, and whether it’s worth it.” Marvasti noted that some customers abandoned internally built automation tools because they could not forecast or justify the costs. The risk is not only a high bill; unpredictable bills can kill useful projects. That is a governance failure disguised as a pricing problem, and it will slow AI adoption more than any model limitation.
Enterprises have already moved beyond the assumption that all AI workloads belong in the public cloud. According to Futurum’s survey of 824 AI decision-makers, reserved and owned infrastructure account for 66% of AI compute consumption, compared with 19% for on-demand cloud. About 59% of respondents primarily run AI workloads outside hyperscaler public clouds, in their own data centers, colocation facilities or with bare-metal providers. Much of that owned capacity reflects GPU purchases made when on-demand capacity was unavailable, so it does not necessarily signal a retreat from hyperscalers. It does show that enterprises are comfortable making capacity commitments for AI. The question is no longer whether to commit, but which workloads justify such commitments.
The pattern mirrors the adoption cycle information technology went through with cloud computing. Companies start on demand, discover that steady-state workloads are cheaper on reserved capacity, and end up hybrid. AI is compressing that curve from years to quarters, and agents are the accelerant. The real economics are about utilization, not price.
The report points to Amberd.ai’s deployment on QumulusAI bare metal as an example. The company partitions an eight-GPU Nvidia H200 server into four virtual environments, each with two GPUs, and tiers customers across them based on latency tolerance. “With one 8x H200 server, two customers pay for the entire server, and I can probably have about 30 to 35 customers running on that one server,” Marvasti said. “After the second customer, the server is free to me, and any customer that comes after that is profit.” The lesson is not that bare metal is cheap. Amberd.ai built a custom virtualization layer and tiered pricing to drive utilization. Reserved infrastructure turns a variable cost into a fixed one, and fixed costs only pay off when kept busy. An idle reserved GPU is the most expensive GPU there is.
Futurum recommends reserved bare metal for sustained workloads with predictable utilization above roughly 60%. The report also acknowledges that these environments require more custom engineering.
Editor's Summary
The Futurum report says agentic AI can consume 10 to 100 times more tokens per task than simple inference, making per-token pricing difficult to sustain in production. It reports that enterprises are already shifting toward reserved and owned compute for AI workloads, while cautioning that utilization and workload predictability determine whether those commitments pay off. The findings point to a need for pricing and governance models that can handle agent-driven token consumption.