60% of agentic AI costs go to response refinement, and most enterprises are already over budget
A McKinsey study reveals that 93% of enterprises exceed their AI budgets as agentic AI systems expand. A significant portion, 60%, of AI costs are directed towards refining responses. The cost structures for these systems are often not fully understood by many operators.
This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.
Key facts, context, and what it means, in one minute.
Key takeaways
93% of enterprises exceed AI budgets as agentic systems scale.
60% of agentic AI costs are allocated to response refinement.
Many operators have not fully mapped the cost structures of agentic systems.
Token prices have collapsed. Inference on GPT-3.5-level capability cost $20 per million tokens in early 2024 and fell to $0.07 by year's end, according to the Stanford HAI 2025 AI Index cited in McKinsey's new report. Enterprise AI spending is accelerating anyway. LLM expenditures tripled over a 12-month period by the end of 2025, according to a Menlo Ventures study also cited by McKinsey, and 93 percent of organizations surveyed by McKinsey in May 2026 reported exceeding their AI budgets. Cheaper tokens, it turns out, do not automatically mean cheaper AI programs.
The explanation sits inside the cost structure of agentic AI, which differs fundamentally from the simple prompt-and-response model most budget forecasts were built around. McKinsey's July 2026 report, published in the McKinsey Quarterly and authored by researchers from QuantumBlack, AI by McKinsey, pinpoints response refinement as the dominant cost driver: 60 percent of total agentic AI spend goes not to the first inference call but to the iterative cycles of checking, correcting, and improving that agents run before delivering a usable output.
Why the cost structure of agents caught enterprises off guard
The shift from subscription to consumption pricing by major LLM providers is one structural cause McKinsey identifies. Under consumption models, answer length directly affects the bill, creating an incentive architecture that rewards verbosity. Enterprises are also routinely routing straightforward tasks through frontier models that are priced for complexity, compounding the overspend. Neither dynamic was fully visible when organizations set their 2025 and 2026 AI budgets.
The scale-up effect amplifies both problems. Companies that began with pilots at controlled token volumes are now running agents across production workflows, and the cost curves are non-linear. A single agent orchestrating multiple sub-agents, each refining its outputs before passing results upstream, can generate a token bill that is orders of magnitude larger than a direct LLM query producing the same end result.
Sixty percent of agentic AI costs sit in response refinement, a cost center that most enterprise budget models never built a line item for.
The breadth of the budget problem is striking. McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 with 75 qualified respondents across five major industries, found that 93 percent have already blown past their AI budgets. Separately, one in five participants in McKinsey's forthcoming 2026 State of AI survey, which gathered responses from 1,719 participants between May and June 2026, said their organizations have actively constrained AI use because of operating costs. That is a notable reversal: AI adoption being slowed not by capability gaps or organizational resistance but by economics.
Token cost is the wrong target metric
McKinsey's central argument is that obsessing over token price reduction is a category error. The report cites David Tepper, CEO of Pay-i, making the distinction plainly: tokens are not value, tokens are the bill. The relevant question for enterprise operators is whether the output an agent produces is worth more than the full cost of producing it, including refinement cycles, human supervision time, and the cost of correcting errors downstream.
That framing shifts the evaluation criteria considerably. An agent that costs three times as much per task but requires no human review and produces zero rework may be the cheaper option in total. Conversely, an agent running on a cheaper model but requiring frequent human correction may cost more in fully loaded terms. Neither the token price nor the model tier alone tells the story, according to McKinsey's analysis.
The report lays out several variables that determine real agent value: output correctness, human supervision and repair burden, compute consumed during reasoning, and whether the completed work exceeds the full operational cost of generation. McKinsey frames these as the levers a CEO needs to understand as AI systems evolve from agentic coworkers toward more autonomous multi-agent architectures.
What this means for enterprise FinOps and procurement teams
For operations and IT leaders, the practical implication is that current AI cost models are almost certainly undercounting. If 60 percent of the bill sits in refinement loops rather than primary inference, and most budget frameworks were built around inference costs alone, the gap between forecast and actual spend will widen as agentic deployments scale. The McKinsey findings, reported by the Economic Times, suggest that the next phase of enterprise GenAI adoption will be shaped less by which models organizations choose and more by how well they instrument and govern agent behavior end to end.
Model selection, routing logic, and output validation are no longer just engineering decisions. They carry direct budget consequences that procurement and finance teams need to co-own with their technology counterparts. Organizations that build FinOps disciplines around agent-level cost attribution, rather than aggregate LLM spend, will be better positioned to scale without the budget overruns that are already affecting the majority of enterprise AI programs.
McKinsey's next data point to watch is its full 2026 State of AI report, which will carry responses from 1,719 global participants and is expected to detail how constrained AI budgets are reshaping deployment priorities across industries. The findings from its May 2026 FinOps survey already suggest that the budget conversation has moved from the IT team to the C-suite, and it is unlikely to move back.
Sources
Featured companies
About the author
The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.