Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

A McKinsey study reveals that 93% of enterprises exceed their AI budgets as agentic AI systems expand. A significant portion, 60%, of AI costs are directed towards refining responses. The cost structures for these systems are often not fully understood by many operators.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Learn this in 60 seconds

Key facts, context, and what it means, in one minute.

:60
0:001:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprises exceed AI budgets as agentic systems scale.

02

60% of agentic AI costs are allocated to response refinement.

03

Many operators have not fully mapped the cost structures of agentic systems.

Get featured

Want to get featured in MarketScale Software & Technology?

Create a free MarketScale workspace and get your company's expertise featured across our Software & Technology coverage. No credit card, no demo required.

Request an invite

Token prices have collapsed. Inference on GPT-3.5-level capability cost $20 per million tokens in early 2024 and fell to $0.07 by year's end, according to the Stanford HAI 2025 AI Index cited in McKinsey's new report. Enterprise AI spending is accelerating anyway. LLM expenditures tripled over a 12-month period by the end of 2025, according to a Menlo Ventures study also cited by McKinsey, and 93 percent of organizations surveyed by McKinsey in May 2026 reported exceeding their AI budgets. Cheaper tokens, it turns out, do not automatically mean cheaper AI programs.

The explanation sits inside the cost structure of agentic AI, which differs fundamentally from the simple prompt-and-response model most budget forecasts were built around. McKinsey's July 2026 report, published in the McKinsey Quarterly and authored by researchers from QuantumBlack, AI by McKinsey, pinpoints response refinement as the dominant cost driver: 60 percent of total agentic AI spend goes not to the first inference call but to the iterative cycles of checking, correcting, and improving that agents run before delivering a usable output.

Why the cost structure of agents caught enterprises off guard

The shift from subscription to consumption pricing by major LLM providers is one structural cause McKinsey identifies. Under consumption models, answer length directly affects the bill, creating an incentive architecture that rewards verbosity. Enterprises are also routinely routing straightforward tasks through frontier models that are priced for complexity, compounding the overspend. Neither dynamic was fully visible when organizations set their 2025 and 2026 AI budgets.

The scale-up effect amplifies both problems. Companies that began with pilots at controlled token volumes are now running agents across production workflows, and the cost curves are non-linear. A single agent orchestrating multiple sub-agents, each refining its outputs before passing results upstream, can generate a token bill that is orders of magnitude larger than a direct LLM query producing the same end result.

Sixty percent of agentic AI costs sit in response refinement, a cost center that most enterprise budget models never built a line item for.

The breadth of the budget problem is striking. McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 with 75 qualified respondents across five major industries, found that 93 percent have already blown past their AI budgets. Separately, one in five participants in McKinsey's forthcoming 2026 State of AI survey, which gathered responses from 1,719 participants between May and June 2026, said their organizations have actively constrained AI use because of operating costs. That is a notable reversal: AI adoption being slowed not by capability gaps or organizational resistance but by economics.

Token cost is the wrong target metric

McKinsey's central argument is that obsessing over token price reduction is a category error. The report cites David Tepper, CEO of Pay-i, making the distinction plainly: tokens are not value, tokens are the bill. The relevant question for enterprise operators is whether the output an agent produces is worth more than the full cost of producing it, including refinement cycles, human supervision time, and the cost of correcting errors downstream.

That framing shifts the evaluation criteria considerably. An agent that costs three times as much per task but requires no human review and produces zero rework may be the cheaper option in total. Conversely, an agent running on a cheaper model but requiring frequent human correction may cost more in fully loaded terms. Neither the token price nor the model tier alone tells the story, according to McKinsey's analysis.

The report lays out several variables that determine real agent value: output correctness, human supervision and repair burden, compute consumed during reasoning, and whether the completed work exceeds the full operational cost of generation. McKinsey frames these as the levers a CEO needs to understand as AI systems evolve from agentic coworkers toward more autonomous multi-agent architectures.

What this means for enterprise FinOps and procurement teams

For operations and IT leaders, the practical implication is that current AI cost models are almost certainly undercounting. If 60 percent of the bill sits in refinement loops rather than primary inference, and most budget frameworks were built around inference costs alone, the gap between forecast and actual spend will widen as agentic deployments scale. The McKinsey findings, reported by the Economic Times, suggest that the next phase of enterprise GenAI adoption will be shaped less by which models organizations choose and more by how well they instrument and govern agent behavior end to end.

Model selection, routing logic, and output validation are no longer just engineering decisions. They carry direct budget consequences that procurement and finance teams need to co-own with their technology counterparts. Organizations that build FinOps disciplines around agent-level cost attribution, rather than aggregate LLM spend, will be better positioned to scale without the budget overruns that are already affecting the majority of enterprise AI programs.

Where agentic AI costs go
McKinsey & Company, July 2026 · © MarketScaleDownload chart

McKinsey's next data point to watch is its full 2026 State of AI report, which will carry responses from 1,719 global participants and is expected to detail how constrained AI budgets are reshaping deployment priorities across industries. The findings from its May 2026 FinOps survey already suggest that the budget conversation has moved from the IT team to the C-suite, and it is unlikely to move back.

Featured companies

Your experts belong here

Every story in MarketScale Software & Technology starts with a company putting its solutions engineers, product teams, and customer engineers on the record. Buyers are already reading this topic. The only question is whose experts they find.

Buyers ask AI engines who to consider, and published expert answers are what those engines cite.

Get your team featuredSee how it works15 minutes, straight to a calendar.

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Follow Software & Technology Insights

Get new expert content in your inbox.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. See how AI describes your company today, and where competitors show up instead.

Free workspace

You just read one Software & Technology expert. Your company is full of them.

This article was produced through MarketScale. The same platform turns your solutions engineers, product teams, and customer engineers into the articles, video, and social content Software & Technology buyers are searching for. Create a free workspace and see it with your own people. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale Studio workspace
One video edit a month, on us
AI writing, editing, and publishing tools
In-platform coaching to learn the system

More Software & Technology Insights

Bending Spoons acquires Airtable at 2.7x ARR, an 89% collapse from its 2021 peak valuation

Bending Spoons acquires Airtable at 2.7x ARR, an 89% collapse from its 2021 peak valuation

Bending Spoons has purchased Airtable for $1.29 billion, with the valuation showing a significant drop from Airtable's $11.7 billion peak in 2021. This acquisition reflects the challenges and changes within the B2B SaaS industry. B2B SaaS operators may need to adjust expectations and strategies in light of evolving market conditions.

  • 01Bending Spoons acquired Airtable at an enterprise value of $1.29 billion.
  • 02Airtable's current valuation represents an 89% drop from its peak valuation in 2021.
  • 03The acquisition sends a signal about changing conditions in the B2B SaaS market.

Aug 14, 2026

Samsung chip profit soared 250-fold this earnings season, and AI infrastructure spending shows no sign of slowing

Samsung chip profit soared 250-fold this earnings season, and AI infrastructure spending shows no sign of slowing

Samsung's chip division experienced a 250-fold increase in profits, largely driven by high demand for AI hardware. Other companies like Palantir, Cloudflare, and Hon Hai also reported significant financial growth attributed to similar market demands.

  • 01Samsung's chip profits increased 250 times due to AI hardware demand.
  • 02Companies such as Palantir, Cloudflare, and Hon Hai also saw significant financial gains.
  • 03AI infrastructure spending continues to rise with no signs of slowing.

Aug 14, 2026

Fiserv and Stuut bring agentic AI to enterprise order-to-cash, with $2B in invoices already processed

Fiserv and Stuut bring agentic AI to enterprise order-to-cash, with $2B in invoices already processed

Fiserv's Commerce Hub and SnapPay are partnering with Stuut's AI agent to streamline enterprise receivables. This integration has already resulted in the processing of over $2 billion in B2B invoices. The collaboration aims to automate and enhance the order-to-cash cycle for businesses.

  • 01Fiserv's AI integration has processed over $2 billion in B2B invoices.
  • 02The collaboration aims to automate the enterprise receivables process.
  • 03Fiserv's Commerce Hub and SnapPay are key components in this integration.

Aug 14, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512