Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

A McKinsey study reveals that 93% of enterprises exceed their AI budgets as agentic AI systems expand. A significant portion, 60%, of AI costs are directed towards refining responses. The cost structures for these systems are often not fully understood by many operators.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Learn this in 60 seconds

Key facts, context, and what it means, in one minute.

:60
0:001:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprises exceed AI budgets as agentic systems scale.

02

60% of agentic AI costs are allocated to response refinement.

03

Many operators have not fully mapped the cost structures of agentic systems.

Get featured

Want to get featured in MarketScale Software & Technology?

Create a free MarketScale workspace and get your company's expertise featured across our Software & Technology coverage. No credit card, no demo required.

Request an invite

Token prices have collapsed. Inference on GPT-3.5-level capability cost $20 per million tokens in early 2024 and fell to $0.07 by year's end, according to the Stanford HAI 2025 AI Index cited in McKinsey's new report. Enterprise AI spending is accelerating anyway. LLM expenditures tripled over a 12-month period by the end of 2025, according to a Menlo Ventures study also cited by McKinsey, and 93 percent of organizations surveyed by McKinsey in May 2026 reported exceeding their AI budgets. Cheaper tokens, it turns out, do not automatically mean cheaper AI programs.

The explanation sits inside the cost structure of agentic AI, which differs fundamentally from the simple prompt-and-response model most budget forecasts were built around. McKinsey's July 2026 report, published in the McKinsey Quarterly and authored by researchers from QuantumBlack, AI by McKinsey, pinpoints response refinement as the dominant cost driver: 60 percent of total agentic AI spend goes not to the first inference call but to the iterative cycles of checking, correcting, and improving that agents run before delivering a usable output.

Why the cost structure of agents caught enterprises off guard

The shift from subscription to consumption pricing by major LLM providers is one structural cause McKinsey identifies. Under consumption models, answer length directly affects the bill, creating an incentive architecture that rewards verbosity. Enterprises are also routinely routing straightforward tasks through frontier models that are priced for complexity, compounding the overspend. Neither dynamic was fully visible when organizations set their 2025 and 2026 AI budgets.

The scale-up effect amplifies both problems. Companies that began with pilots at controlled token volumes are now running agents across production workflows, and the cost curves are non-linear. A single agent orchestrating multiple sub-agents, each refining its outputs before passing results upstream, can generate a token bill that is orders of magnitude larger than a direct LLM query producing the same end result.

Sixty percent of agentic AI costs sit in response refinement, a cost center that most enterprise budget models never built a line item for.

The breadth of the budget problem is striking. McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 with 75 qualified respondents across five major industries, found that 93 percent have already blown past their AI budgets. Separately, one in five participants in McKinsey's forthcoming 2026 State of AI survey, which gathered responses from 1,719 participants between May and June 2026, said their organizations have actively constrained AI use because of operating costs. That is a notable reversal: AI adoption being slowed not by capability gaps or organizational resistance but by economics.

Token cost is the wrong target metric

McKinsey's central argument is that obsessing over token price reduction is a category error. The report cites David Tepper, CEO of Pay-i, making the distinction plainly: tokens are not value, tokens are the bill. The relevant question for enterprise operators is whether the output an agent produces is worth more than the full cost of producing it, including refinement cycles, human supervision time, and the cost of correcting errors downstream.

That framing shifts the evaluation criteria considerably. An agent that costs three times as much per task but requires no human review and produces zero rework may be the cheaper option in total. Conversely, an agent running on a cheaper model but requiring frequent human correction may cost more in fully loaded terms. Neither the token price nor the model tier alone tells the story, according to McKinsey's analysis.

The report lays out several variables that determine real agent value: output correctness, human supervision and repair burden, compute consumed during reasoning, and whether the completed work exceeds the full operational cost of generation. McKinsey frames these as the levers a CEO needs to understand as AI systems evolve from agentic coworkers toward more autonomous multi-agent architectures.

What this means for enterprise FinOps and procurement teams

For operations and IT leaders, the practical implication is that current AI cost models are almost certainly undercounting. If 60 percent of the bill sits in refinement loops rather than primary inference, and most budget frameworks were built around inference costs alone, the gap between forecast and actual spend will widen as agentic deployments scale. The McKinsey findings, reported by the Economic Times, suggest that the next phase of enterprise GenAI adoption will be shaped less by which models organizations choose and more by how well they instrument and govern agent behavior end to end.

Model selection, routing logic, and output validation are no longer just engineering decisions. They carry direct budget consequences that procurement and finance teams need to co-own with their technology counterparts. Organizations that build FinOps disciplines around agent-level cost attribution, rather than aggregate LLM spend, will be better positioned to scale without the budget overruns that are already affecting the majority of enterprise AI programs.

Where agentic AI costs go
McKinsey & Company, July 2026 · © MarketScaleDownload chart

McKinsey's next data point to watch is its full 2026 State of AI report, which will carry responses from 1,719 global participants and is expected to detail how constrained AI budgets are reshaping deployment priorities across industries. The findings from its May 2026 FinOps survey already suggest that the budget conversation has moved from the IT team to the C-suite, and it is unlikely to move back.

Featured companies

Your experts belong here

Every story in MarketScale Software & Technology starts with a company putting its solutions engineers, product teams, and customer engineers on the record. Buyers are already reading this topic. The only question is whose experts they find.

Buyers ask AI engines who to consider, and published expert answers are what those engines cite.

Get your team featuredSee how it works15 minutes, straight to a calendar.

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Follow Software & Technology Insights

Get new expert content in your inbox.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. See how AI describes your company today, and where competitors show up instead.

Free workspace

You just read one Software & Technology expert. Your company is full of them.

This article was produced through MarketScale. The same platform turns your solutions engineers, product teams, and customer engineers into the articles, video, and social content Software & Technology buyers are searching for. Create a free workspace and see it with your own people. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale Studio workspace
One video edit a month, on us
AI writing, editing, and publishing tools
In-platform coaching to learn the system

More Software & Technology Insights

Vantage’s 1.4GW Texas campus makes grid contracts the real data center schedule

Vantage’s 1.4GW Texas campus makes grid contracts the real data center schedule

Vantage Data Centers is targeting a 1.4GW “Frontier” campus in Texas, with first delivery slated for H2 2026. Power procurement and cooling design land first on operators. Emissions accounting follows, alongside carbon-removal contracting.

  • 01For large AI campuses, the interconnect and power-delivery agreement is becoming the long pole, it now sets when IT can arrive.
  • 02Carbon-removal offtake is shifting from pilot-scale buys to 8–10 year contracts that support final investment decisions, useful for sustainability procurement playbooks.
  • 03250kW-plus racks and liquid cooling are moving from special requests to baseline specs for new AI capacity, changing mechanical and service vendor selection.

Sep 7, 2026

Dreamforce 2026 goes all-in on AI agents, but ROI numbers are still missing

Pre-event materials cited include no customer-reported ROI, adoption metrics, or cost-to-run figures for Agentforce. The main keynote is Sept. 15, 2026. UC Today says Dreamforce runs Sept. 15-17 at Moscone, with a free Salesforce+ virtual program Sept. 15-18.

  • 01The sources set an expectation gap: Dreamforce 2026 messaging leans on “agentic” adoption, but the pre-event materials cited here include no customer ROI figures or cost-to-run numbers for Agentforce, so procurement and operations teams should arrive with measurement and cost-accounting questions ready (per UC Today).
  • 02UC Today lists Dreamforce 2026’s published scale as 1,600+ breakout sessions, 50+ keynotes, 150+ hands-on trainings and demos, and 240+ community roundtables, plus one-to-one sessions with Agentforce and Slack product experts.
  • 03The pass price gap, $1,899 “Last Chance” vs $2,299 full price, is a practical benchmark for budgeting onsite attendance against free Salesforce+ virtual access (per UC Today).

Sep 6, 2026

AI could raise enterprise IT costs by as much as 75% in less than a decade

AI could raise enterprise IT costs by as much as 75% in less than a decade

Bain & Company projects AI could raise enterprise IT costs by as much as 75% in less than a decade. Procurement and IT teams will feel it first. The impact shows up in vendor contracts, capacity planning, and governance workflows.

  • 01A 75% IT cost lift is no longer a scare number, it is becoming a budgeting baseline once security, data movement, and talent are counted (Bain via CIO Dive).
  • 02For firms standardizing on AI agents, contract language is shifting toward reliability and control artifacts, not model brand names (KPMG certification coverage via CIO Dive).
  • 03Infrastructure availability is turning into a scheduling problem, not a procurement event, with Dell citing a $95B AI backlog that can push deployments into future quarters (CIO).

Sep 5, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512