Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

McKinsey's research indicates that 93% of enterprise AI teams are over budget, primarily due to costs related to response refinement in agentic AI, which accounts for 60% of the total AI expenditure. This highlights the financial strain organizations face while trying to enhance the effectiveness of their AI systems.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Learn this in 60 seconds

Key facts, context, and what it means, in one minute.

:60
0:001:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprise AI teams exceed their budgets.

02

Response refinement consumes 60% of agentic AI spending.

03

Agentic AI cost overruns are a common issue among enterprises.

Token prices have fallen more than 99% in roughly two years. Enterprise AI bills have tripled anyway. That contradiction sits at the center of a July 2026 McKinsey report that maps, in unusually precise terms, where agentic AI money actually goes and why most organizations cannot yet account for it.

The numbers are stark. According to McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 across 75 qualified respondents spanning five major industries, 93% of enterprise participants report exceeding their AI budgets. Separately, data from Menlo Ventures cited in the report showed that enterprise large language model spending tripled over a 12-month period by the end of 2025. Meanwhile, Stanford's HAI 2025 AI Index documented that inference cost for GPT-3.5-level capability collapsed from $20 per million tokens to $0.07 through 2024, a drop of more than 99%.

The divergence between falling unit costs and rising total bills is not a paradox. It is a volume and architecture problem, and McKinsey's QuantumBlack team lays out why.

Where the money actually goes

The single largest cost center in an agentic AI deployment is response refinement, the iterative loop in which an agent checks, revises, and regenerates its own outputs before returning a final answer. According to McKinsey, that process consumes 60% of total agentic AI costs. For enterprise teams that assumed their spend was split evenly across retrieval, reasoning, and generation, that concentration is a significant recalibration.

Share of agentic AI costs by category60Response refinement40Other agentic operations
McKinsey & Company, July 2026 · © MarketScaleDownload chart

Three structural forces compound the problem, according to the McKinsey report. First, enterprises are scaling AI efforts at pace, so more workloads are generating token consumption. Second, LLM providers have broadly shifted from flat subscription pricing to consumption-based models, which creates an incentive for longer, more elaborate outputs. Third, and perhaps most correctable, expensive frontier models are frequently deployed for routine tasks that cheaper, smaller models could handle.

The Economic Times also reported on the McKinsey findings, noting that enterprise leaders are increasingly focused on the economics of operating AI agents rather than the underlying technology, a shift that reflects how quickly agentic deployments have moved from pilot to production.

The question is no longer whether you can deploy an AI agent. It is whether the value of what that agent produces justifies every dollar it costs to run it.

The constraint is already real

Budget pressure is not a future concern. McKinsey's forthcoming 2026 State of AI global survey, fielded between May 4 and June 8, 2026, with 1,719 participants, found that one in five organizations has already constrained AI use specifically because of AI-related operating costs. That is a material brake on adoption, one that shows up in deployment decisions, model selection, and the scope of workflows that get assigned to agents.

The pattern reflects a broader tension in enterprise AI right now. Boards and executive teams approved AI investment expecting that declining model costs would keep total spend manageable. What they underestimated was the multiplicative effect of scale and architecture: more agents, more tasks, more refinement loops, and consumption pricing that rewards output volume.

McKinsey describes the core executive question as deceptively simple: are the AI agent capabilities being built and run worth the value being extracted from them? But answering it requires instrumentation most enterprises do not yet have, specifically the ability to track whether agent output is correct, how much human supervision or repair it requires, and whether the completed work's value actually exceeds its full operational cost.

Tokens are the bill, not the value

The McKinsey report draws a pointed distinction between token spend and business value, attributing a framing to David Tepper, CEO of Pay-i: tokens are not value, tokens are the bill. That reframe matters operationally. Cost-reduction conversations anchored to per-token pricing miss the actual lever, which is the ratio of agent output quality and business impact to total operational cost.

For a VP of Operations or a CIO evaluating agentic AI portfolios, this means the relevant metric is not cost per million tokens. It is cost per completed, accurate, human-review-free task. An agent that runs more refinement loops but delivers output that never requires correction may be cheaper in practice than a cheaper-per-token agent whose outputs routinely need human repair.

The McKinsey authors, including Lari Hämäläinen, Mark Patel, Sven Blumberg, Tanguy Catlin, and Wasim Lala from QuantumBlack, position agentic economics as the defining operational challenge for enterprise AI in this phase of adoption. With 93% of surveyed enterprises already over budget and one-fifth actively pulling back on deployment scope, the window to build proper cost-value instrumentation is narrowing.

What this means for your team

  • Audit where your agentic token spend actually concentrates. If you cannot attribute costs by workload phase, including retrieval, reasoning, and refinement, you cannot target the 60% that McKinsey identifies as the dominant cost driver.
  • Evaluate model-task fit across your agent stack. Deploying frontier models on routine tasks is a documented cost driver; map each agent use case to the minimum capable model tier and quantify the savings before your next budget cycle.
  • Build output-quality tracking alongside cost tracking. Measuring tokens spent without measuring accuracy, human correction rates, and task completion rates makes cost-value comparison impossible.
  • Use the McKinsey framing as an executive forcing function: for each active agent deployment, require a documented answer to whether the value of completed work exceeds the full operational cost of generating it.

Featured companies

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. See how AI describes your company today, and where competitors show up instead.

Free workspace

You just read one expert. Imagine publishing your whole team.

This article was produced through MarketScale. Create a free workspace and turn your own team's expertise into articles, video, and social posts. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale Studio workspace
One video edit a month, on us
AI writing, editing, and publishing tools
In-platform coaching to learn the system

More Software & Technology Insights

Why enterprise AI programs stall before scaling, and what the roadmap actually requires

Why enterprise AI programs stall before scaling, and what the roadmap actually requires

Many enterprise AI initiatives stall during the pilot phase and do not progress to production. A structured roadmap consisting of five key steps can aid in successful adoption. Critical factors include ensuring readiness, implementing proper governance, and aligning projects with business KPIs.

  • 01Most enterprise AI budgets are spent on pilots that fail to reach production.
  • 02A successful AI adoption roadmap should focus on readiness, governance, and alignment with business KPIs.
  • 03Structured steps are essential for transitioning AI initiatives from pilot stages to full production.

Jul 20, 2026

Accenture Edge and Google Cloud target midmarket AI gap with pre-built agentic tools

Accenture Edge and Google Cloud target midmarket AI gap with pre-built agentic tools

Accenture Edge, in collaboration with Google Cloud, is targeting midmarket businesses by providing pre-built AI tools using the Gemini platform. These tools are specifically designed for companies with revenue under $3 billion, aiming to bridge the AI adoption gap in this segment.

  • 01Accenture Edge and Google Cloud offer pre-configured AI tools for midmarket companies.
  • 02The tools are built on the Gemini platform, targeting firms with revenue under $3 billion.
  • 03The collaboration aims to facilitate AI adoption among midmarket businesses.

Jul 20, 2026

Cognizant expands Google Cloud partnership to move 100,000 associates onto Gemini Enterprise

Cognizant expands Google Cloud partnership to move 100,000 associates onto Gemini Enterprise

Cognizant and Google Cloud are enhancing their Gemini Enterprise collaboration by aiming to transition 100,000 associates onto their platform. This partnership focuses on expanding joint market offerings in essential industries.

  • 01Cognizant is working with Google Cloud to transition 100,000 associates onto the Gemini Enterprise platform.
  • 02The partnership aims to enhance go-to-market offerings across various key industries.
  • 03Cognizant is focusing on deepening its collaboration with Google Cloud for strategic growth.

Jul 19, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512