AI News Feed
Market watch
Products & Applications

SiliconANGLE Reports Warn Enterprise AI Budgets Miss Data Gravity and Agent Testing Costs

Two SiliconANGLE reports say enterprise AI programs face hidden data movement and testing costs as agentic workloads scale, with 95% of gen AI pilots showing little P&L impact.

The first report describes a hidden tax on enterprise AI. Large companies have spent years investing in AI infrastructure, software and implementation, with boards approving plans, finance building business cases and procurement negotiating compute capacity. The unresolved question was how they would manage the data underneath those systems. The failure is not a major outage, breach or recall, according to the report, but a quiet and compounding drag that appears as unplanned headcount and slipping timelines. Infrastructure leaders repeatedly cite data gravity.

A 2025 report from MIT's NANDA initiative found that 95% of enterprise generative AI pilots produced little or no measurable impact on the profit-and-loss statement. The MIT report points to gaps in how companies integrate AI into operations, not simply model quality, according to SiliconANGLE. A successful pilot does not prove that an organization can run the same system across its full data estate.

Most enterprise AI architectures still assume that data will move to wherever the new platform lives. That assumption is rarely true at enterprise scale. Regulations dictate where some records may live, and sovereignty rules keep other datasets inside a jurisdiction even when compute is cheaper elsewhere. Business units that spent a decade building governance around a dataset have little incentive to hand it to a new central store, and sometimes no legal path to do so. Applications built around their data can also break in expensive ways when the data and application are separated. The problem is typically hidden in a proof of concept that runs on a small, pre-cleaned slice of data. Costs appear later, when the program must handle the messy majority left out of the demonstration.

Data gravity means large accumulations of data pull applications toward them, not the other way around. The bigger the dataset, the more expensive it is to move. Transfer fees are only the beginning; latency, bandwidth, security controls and operational effort add to the bill before a model produces a useful answer. The cost usually does not show up as one line item. It appears as dozens of small demands on teams that already have too much to manage: building and maintaining pipelines to move data out of systems that were never designed for that traffic, reconciling each copy with the original when the source system changes, applying the same governance controls wherever data is stored, and tracking extra datasets created for development, testing, analytics and training. None of these tasks looks disastrous alone. Together they become a tax that gets more expensive as the organization tries to scale.

The timing is the most dangerous part. The invoice often appears 12 to 24 months after the platform is procured and the team is staffed, when initial success metrics have already been reported upward. By then, contracts are signed, teams are hired and the direction is publicly committed. Unwinding a centralization-first architecture is far more expensive than designing around data gravity from the start. Compute is easy to identify and assign to a budget, but pipeline maintenance, reconciliation, governance, security reviews and the staff needed to keep the system running are spread across different organizations.

SiliconANGLE says many postmortems stop at the diagnosis that data was not ready, which is accurate but incomplete. The deeper issue is that organizations are trying to solve a context problem with a storage strategy. A model does not need raw data dumped in front of it; it needs context, meaning the ability to find the right record, cross-reference it against policy, respect access rules and work from current information rather than a snapshot from the week the project began. That does not come from a bigger warehouse, according to the report, but from an enterprise-wide context layer that can discover, connect and retrieve information across systems and locations without requiring every source to hand its data to a central store. Records remain constrained by regulation, sovereignty, ownership and the applications built around them, while metadata and vectors can help AI find and reason about those records without inheriting every constraint.

The second SiliconANGLE report describes how agentic workloads break assumptions about software testing. Most traditional enterprise systems were built around three assumptions: jobs finish quickly, retrying one is free, and the same input always produces the same output. Agents violate all three. An agentic task can run long enough to exceed timeout thresholds that nothing in the stack has encountered before. Retries now cost money because every attempt against a metered model consumes compute resources whether the result is usable or not. Cloud and model API usage is billed asynchronously on monthly cycles, so compounded retry costs quietly accumulate and become visible only when the invoice arrives weeks later. Failures also cannot be reproduced without a step-by-step trace of agent decisions, tool calls and API runs.

Pilots hide these issues because humans act as error handlers when working interactively. People read each result, notice problems and try again. Costs are visible because attempts are counted and fixes are applied by hand. Automated agent workflows run headlessly, with the potential for unrecorded failures to break downstream systems. A workflow that behaved reliably when a person checked each response can behave differently when a scheduler fires it hundreds of times overnight with nobody watching. Before moving to automated production with agents, engineering leaders must identify every task the human was performing manually and specify which automated check or system will take over that responsibility.

The report highlights three risks. Invisible costs arise because autonomous AI loops retry failed tasks, regenerate responses and hit APIs without human intervention or approval. A developer-set threshold can move the monthly bill more than a procurement negotiation. The report recommends asking for the cost per completed unit rather than cost per API call, since the second number excludes discarded attempts. Failures that report success are the most expensive defects, because probabilistic agents can produce correctly formatted outputs with bad data that appears structurally valid. Downstream automated systems may accept and process such outputs without triggering alerts, and conventional monitoring misses these errors because nothing failed. AI output quality must be measured with predefined programmatic criteria such as assertion checks, LLM-as-a-judge rules or semantic benchmarks. Incidents nobody can explain occur when a team cannot say which model version, inputs and settings produced a specific output. That information is cheap to capture while work runs and close to impossible to reconstruct later.

Among the questions the report says teams should answer are what defines acceptable output and whether it is written down before the work runs; how many attempts a typical completed unit takes and whether that is capped; and what is recorded for every run. Job data should include the exact input payload, prompt or model version, timestamp, execution latency, retry counts, token cost and the final output artifact.

Editor's Summary SiliconANGLE reports say enterprise AI programs are underestimating data movement, governance and agent-testing costs. A 2025 MIT NANDA report found that 95% of generative AI pilots had little or no measurable P&L impact, and agentic workloads add retry expenses and reproducibility problems when they move from interactive pilots to automated production. The reports advise budgeting for context layers, completed-unit costs, traceability and automated acceptance criteria before scaling.