60% of agentic AI costs go to response refinement, and most enterprises are already over budget
McKinsey's research indicates that 93% of enterprise AI teams are over budget, primarily due to costs related to response refinement in agentic AI, which accounts for 60% of the total AI expenditure. This highlights the financial strain organizations face while trying to enhance the effectiveness of their AI systems.
This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.
Key facts, context, and what it means, in one minute.
Key takeaways
93% of enterprise AI teams exceed their budgets.
Response refinement consumes 60% of agentic AI spending.
Agentic AI cost overruns are a common issue among enterprises.
Token prices have fallen more than 99% in roughly two years. Enterprise AI bills have tripled anyway. That contradiction sits at the center of a July 2026 McKinsey report that maps, in unusually precise terms, where agentic AI money actually goes and why most organizations cannot yet account for it.
The numbers are stark. According to McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 across 75 qualified respondents spanning five major industries, 93% of enterprise participants report exceeding their AI budgets. Separately, data from Menlo Ventures cited in the report showed that enterprise large language model spending tripled over a 12-month period by the end of 2025. Meanwhile, Stanford's HAI 2025 AI Index documented that inference cost for GPT-3.5-level capability collapsed from $20 per million tokens to $0.07 through 2024, a drop of more than 99%.
The divergence between falling unit costs and rising total bills is not a paradox. It is a volume and architecture problem, and McKinsey's QuantumBlack team lays out why.
Where the money actually goes
The single largest cost center in an agentic AI deployment is response refinement, the iterative loop in which an agent checks, revises, and regenerates its own outputs before returning a final answer. According to McKinsey, that process consumes 60% of total agentic AI costs. For enterprise teams that assumed their spend was split evenly across retrieval, reasoning, and generation, that concentration is a significant recalibration.
Three structural forces compound the problem, according to the McKinsey report. First, enterprises are scaling AI efforts at pace, so more workloads are generating token consumption. Second, LLM providers have broadly shifted from flat subscription pricing to consumption-based models, which creates an incentive for longer, more elaborate outputs. Third, and perhaps most correctable, expensive frontier models are frequently deployed for routine tasks that cheaper, smaller models could handle.
The Economic Times also reported on the McKinsey findings, noting that enterprise leaders are increasingly focused on the economics of operating AI agents rather than the underlying technology, a shift that reflects how quickly agentic deployments have moved from pilot to production.
The question is no longer whether you can deploy an AI agent. It is whether the value of what that agent produces justifies every dollar it costs to run it.
The constraint is already real
Budget pressure is not a future concern. McKinsey's forthcoming 2026 State of AI global survey, fielded between May 4 and June 8, 2026, with 1,719 participants, found that one in five organizations has already constrained AI use specifically because of AI-related operating costs. That is a material brake on adoption, one that shows up in deployment decisions, model selection, and the scope of workflows that get assigned to agents.
The pattern reflects a broader tension in enterprise AI right now. Boards and executive teams approved AI investment expecting that declining model costs would keep total spend manageable. What they underestimated was the multiplicative effect of scale and architecture: more agents, more tasks, more refinement loops, and consumption pricing that rewards output volume.
McKinsey describes the core executive question as deceptively simple: are the AI agent capabilities being built and run worth the value being extracted from them? But answering it requires instrumentation most enterprises do not yet have, specifically the ability to track whether agent output is correct, how much human supervision or repair it requires, and whether the completed work's value actually exceeds its full operational cost.
Tokens are the bill, not the value
The McKinsey report draws a pointed distinction between token spend and business value, attributing a framing to David Tepper, CEO of Pay-i: tokens are not value, tokens are the bill. That reframe matters operationally. Cost-reduction conversations anchored to per-token pricing miss the actual lever, which is the ratio of agent output quality and business impact to total operational cost.
For a VP of Operations or a CIO evaluating agentic AI portfolios, this means the relevant metric is not cost per million tokens. It is cost per completed, accurate, human-review-free task. An agent that runs more refinement loops but delivers output that never requires correction may be cheaper in practice than a cheaper-per-token agent whose outputs routinely need human repair.
The McKinsey authors, including Lari Hämäläinen, Mark Patel, Sven Blumberg, Tanguy Catlin, and Wasim Lala from QuantumBlack, position agentic economics as the defining operational challenge for enterprise AI in this phase of adoption. With 93% of surveyed enterprises already over budget and one-fifth actively pulling back on deployment scope, the window to build proper cost-value instrumentation is narrowing.
What this means for your team
- Audit where your agentic token spend actually concentrates. If you cannot attribute costs by workload phase, including retrieval, reasoning, and refinement, you cannot target the 60% that McKinsey identifies as the dominant cost driver.
- Evaluate model-task fit across your agent stack. Deploying frontier models on routine tasks is a documented cost driver; map each agent use case to the minimum capable model tier and quantify the savings before your next budget cycle.
- Build output-quality tracking alongside cost tracking. Measuring tokens spent without measuring accuracy, human correction rates, and task completion rates makes cost-value comparison impossible.
- Use the McKinsey framing as an executive forcing function: for each active agent deployment, require a documented answer to whether the value of completed work exceeds the full operational cost of generating it.
Sources
Featured companies
About the author
The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.