Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

McKinsey's research indicates that 93% of enterprise AI teams are over budget, primarily due to costs related to response refinement in agentic AI, which accounts for 60% of the total AI expenditure. This highlights the financial strain organizations face while trying to enhance the effectiveness of their AI systems.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Learn this in 60 seconds

Key facts, context, and what it means, in one minute.

:60
0:001:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprise AI teams exceed their budgets.

02

Response refinement consumes 60% of agentic AI spending.

03

Agentic AI cost overruns are a common issue among enterprises.

Get featured

Want to get featured in MarketScale Software & Technology?

Create a free MarketScale workspace and get your company's expertise featured across our Software & Technology coverage. No credit card, no demo required.

Request an invite

Token prices have fallen more than 99% in roughly two years. Enterprise AI bills have tripled anyway. That contradiction sits at the center of a July 2026 McKinsey report that maps, in unusually precise terms, where agentic AI money actually goes and why most organizations cannot yet account for it.

The numbers are stark. According to McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 across 75 qualified respondents spanning five major industries, 93% of enterprise participants report exceeding their AI budgets. Separately, data from Menlo Ventures cited in the report showed that enterprise large language model spending tripled over a 12-month period by the end of 2025. Meanwhile, Stanford's HAI 2025 AI Index documented that inference cost for GPT-3.5-level capability collapsed from $20 per million tokens to $0.07 through 2024, a drop of more than 99%.

The divergence between falling unit costs and rising total bills is not a paradox. It is a volume and architecture problem, and McKinsey's QuantumBlack team lays out why.

Where the money actually goes

The single largest cost center in an agentic AI deployment is response refinement, the iterative loop in which an agent checks, revises, and regenerates its own outputs before returning a final answer. According to McKinsey, that process consumes 60% of total agentic AI costs. For enterprise teams that assumed their spend was split evenly across retrieval, reasoning, and generation, that concentration is a significant recalibration.

Share of agentic AI costs by category
Response refinement60%
Other agentic operations40%
McKinsey & Company, July 2026 · © MarketScaleDownload chart

Three structural forces compound the problem, according to the McKinsey report. First, enterprises are scaling AI efforts at pace, so more workloads are generating token consumption. Second, LLM providers have broadly shifted from flat subscription pricing to consumption-based models, which creates an incentive for longer, more elaborate outputs. Third, and perhaps most correctable, expensive frontier models are frequently deployed for routine tasks that cheaper, smaller models could handle.

The Economic Times also reported on the McKinsey findings, noting that enterprise leaders are increasingly focused on the economics of operating AI agents rather than the underlying technology, a shift that reflects how quickly agentic deployments have moved from pilot to production.

The question is no longer whether you can deploy an AI agent. It is whether the value of what that agent produces justifies every dollar it costs to run it.

The constraint is already real

Budget pressure is not a future concern. McKinsey's forthcoming 2026 State of AI global survey, fielded between May 4 and June 8, 2026, with 1,719 participants, found that one in five organizations has already constrained AI use specifically because of AI-related operating costs. That is a material brake on adoption, one that shows up in deployment decisions, model selection, and the scope of workflows that get assigned to agents.

The pattern reflects a broader tension in enterprise AI right now. Boards and executive teams approved AI investment expecting that declining model costs would keep total spend manageable. What they underestimated was the multiplicative effect of scale and architecture: more agents, more tasks, more refinement loops, and consumption pricing that rewards output volume.

McKinsey describes the core executive question as deceptively simple: are the AI agent capabilities being built and run worth the value being extracted from them? But answering it requires instrumentation most enterprises do not yet have, specifically the ability to track whether agent output is correct, how much human supervision or repair it requires, and whether the completed work's value actually exceeds its full operational cost.

Tokens are the bill, not the value

The McKinsey report draws a pointed distinction between token spend and business value, attributing a framing to David Tepper, CEO of Pay-i: tokens are not value, tokens are the bill. That reframe matters operationally. Cost-reduction conversations anchored to per-token pricing miss the actual lever, which is the ratio of agent output quality and business impact to total operational cost.

For a VP of Operations or a CIO evaluating agentic AI portfolios, this means the relevant metric is not cost per million tokens. It is cost per completed, accurate, human-review-free task. An agent that runs more refinement loops but delivers output that never requires correction may be cheaper in practice than a cheaper-per-token agent whose outputs routinely need human repair.

The McKinsey authors, including Lari Hämäläinen, Mark Patel, Sven Blumberg, Tanguy Catlin, and Wasim Lala from QuantumBlack, position agentic economics as the defining operational challenge for enterprise AI in this phase of adoption. With 93% of surveyed enterprises already over budget and one-fifth actively pulling back on deployment scope, the window to build proper cost-value instrumentation is narrowing.

What this means for your team

  • Audit where your agentic token spend actually concentrates. If you cannot attribute costs by workload phase, including retrieval, reasoning, and refinement, you cannot target the 60% that McKinsey identifies as the dominant cost driver.
  • Evaluate model-task fit across your agent stack. Deploying frontier models on routine tasks is a documented cost driver; map each agent use case to the minimum capable model tier and quantify the savings before your next budget cycle.
  • Build output-quality tracking alongside cost tracking. Measuring tokens spent without measuring accuracy, human correction rates, and task completion rates makes cost-value comparison impossible.
  • Use the McKinsey framing as an executive forcing function: for each active agent deployment, require a documented answer to whether the value of completed work exceeds the full operational cost of generating it.

Featured companies

Your experts belong here

Every story in MarketScale Software & Technology starts with a company putting its solutions engineers, product teams, and customer engineers on the record. Buyers are already reading this topic. The only question is whose experts they find.

Buyers ask AI engines who to consider, and published expert answers are what those engines cite.

Get your team featuredSee how it works15 minutes, straight to a calendar.

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Follow Software & Technology Insights

Get new expert content in your inbox.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. See how AI describes your company today, and where competitors show up instead.

Free workspace

You just read one Software & Technology expert. Your company is full of them.

This article was produced through MarketScale. The same platform turns your solutions engineers, product teams, and customer engineers into the articles, video, and social content Software & Technology buyers are searching for. Create a free workspace and see it with your own people. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale Studio workspace
One video edit a month, on us
AI writing, editing, and publishing tools
In-platform coaching to learn the system

More Software & Technology Insights

OpenAI's GPT-6 Astra pitch is to skip integrations and run the software UI itself

OpenAI's GPT-6 Astra pitch is to skip integrations and run the software UI itself

OpenAI shipped GPT-6 Astra on Sept. 3 with a "computer use" capability that lets the model operate existing software UIs directly instead of requiring custom API integrations. The staged rollout through Daybreak, ChatGPT tiers, and AWS shifts the automation bottleneck from building connectors to governing UI-driven sessions, with speed measured in minutes per task as the cost input.

  • 01Astra operates software via pixels, keyboard, and mouse interactions to bypass API integration work on the long tail of internal tools without clean API access
  • 02OpenAI reported Astra at 40 minutes per task (47% faster than GPT-5.6 Sol at 75 minutes), making task time the practical proxy for compute cost modeling and throughput evaluation
  • 03Enterprises must define governance before broad rollout: eligible workflows for UI automation, audit logging systems, identity and secrets handling, and fallback procedures when UIs change or sessions break

Sep 3, 2026

Cloud is set to take 26% of IT budgets, and hiring is shifting toward platforms

Cloud is set to take 26% of IT budgets, and hiring is shifting toward platforms

Foundry’s 2026 Cloud Computing Study, cited by CIO, reports IT leaders expect 26% of IT budgets to go to cloud computing in the next year and that 74% accelerated cloud migrations in the last 12 months. At the same time, CIO’s coverage of Robert Half Technology’s 2026 IT salary report shows AI/ML engineers at a $170,750 median salary and a striking capability gap, with only 7% of leaders saying they have the capabilities to complete prioritized projects and 65% expecting to upskill existing staff. Separate reporting from Nextgov on senior appointments in the Pentagon CIO office, and Government Technology’s account of Illinois’ multi-agency data-sharing MOU, indicate large organizations are staffing for organizational change, governance, and cross-domain data sharing, not just for lift-and-shift migrations. For enterprise operators, the signal is that cloud programs increasingly depend on internal platform engineering, adoption enablement, and finance-aligned governance, because those functions connect migration speed to day-two operations and business results.

  • 01A useful benchmark for 2026 planning: Foundry’s survey puts cloud at 26% of IT budget next year, so showback/FinOps and workload-level chargeback governance can’t stay a side project (per CIO citing Foundry).
  • 02Talent pricing has become an input to architecture decisions. A CIO, citing Robert Half, compared the median pay for AI/ML engineers ($170,750) with systems administrators ($98,000), arguing that platform standardization and self-service guardrails are as much about labor strategy as technology.
  • 03If 65% of leaders expect to upskill to close gaps, the differentiator becomes your internal adoption system, in-app guidance, runbooks, and training telemetry, not the third new tool in the stack (per CIO citing Robert Half; illustrated by Ferring’s Whatfix rollout reported by BankInfoSecurity).

Sep 2, 2026

AI agents are pushing access controls and testing into the data layer

AI agents are pushing access controls and testing into the data layer

Snowflake is warning that dashboard-era access controls do not hold up once AI agents can query and act across datasets, pushing governance closer to the data layer, according to TechTarget’s Computer Weekly. In parallel, TechTarget reported that enterprise AI-agent testing needs to expand beyond pre-deployment checks into continuous monitoring so agents don’t drift beyond prescribed instructions. InformationWeek’s reporting on CISOs at Intuit, Smartsheet and ETS adds the operational risk: unmanaged “AI orphans” and identity sprawl as agent count grows, which shifts near-term workload onto IAM, data governance and platform engineering teams building the guardrails.

  • 01If an AI agent can reach multiple tools, the real control plane becomes the data layer and identity, not the BI dashboard permissions model.
  • 02Agent rollouts that stop at pre-production testing are likely to miss the failure mode operators actually see, post-deploy tool changes that alter what the agent can do.
  • 03“AI orphans” is a practical inventory problem: if teams can’t enumerate agents, they can’t set ownership, secrets rotation, or access reviews on a schedule.

Sep 2, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512