Palo Alto Networks CEO puts a number on the AI cost problem: 90% token price drop needed
Nikesh Arora, CEO of Palo Alto Networks, stated that for enterprise AI to scale, token costs must decrease by 90% within two years. He highlighted that high costs have already impacted companies like Uber, which spent its full-year AI budget by April.
This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.
Key facts, context, and what it means, in one minute.
Key takeaways
Token costs for AI need to decline by 90% in two years for scalability.
Uber exhausted its annual AI budget by April due to high costs.
Palo Alto Networks CEO Nikesh Arora put precise numbers on the enterprise AI cost problem July 9, telling CNBC's Squawk on the Street that token prices need to fall 20% within a year and 90% the year after before companies can realistically scale AI workloads. The remarks land as mounting evidence shows enterprises are already pulling back spending they committed to earlier in 2026.
The cost ceiling is real and already being hit
Uber is the clearest data point. According to PYMNTS, the company burned through its entire 2026 AI budget by April. Chief Operating Officer Andrew Macdonald said Uber would weigh token costs directly against the cost of hiring engineers, a comparison that would have seemed far-fetched two years ago. CTO Praveen Neppalli Naga described the situation as being "back to the drawing board."
Uber's situation is not isolated. PYMNTS reported in June that companies that once encouraged broad internal AI tool adoption, when costs were lower, are now rationing access through usage caps, nudging employees toward task-appropriate models, and routing lower-stakes work to older, cheaper options. The economics shifted faster than most IT and procurement teams planned for.
When Arora was asked about OpenAI CEO Sam Altman's claim that OpenAI's latest model is 54% more efficient for coding, Arora said the improvement is a good start but not sufficient. "I think we probably need another turn at it," he said, per CNBC. The comment signals that even headline efficiency gains from frontier model vendors are not closing the gap fast enough for enterprise buyers.
Agentic tools amplify the exposure
Standard chatbot interactions generate a single inference call per exchange. Agentic coding tools, which complete multi-step tasks autonomously, generate many inference calls per session. That structural difference means enterprises that deployed agentic tools based on chatbot-era cost assumptions are seeing usage bills that scale non-linearly with adoption, according to PYMNTS.
For operations and IT leaders, this is a procurement design problem. Budgets built on per-seat or per-user assumptions break down when the actual unit of cost is inference volume, which varies sharply by use case, user behavior, and model selection.
Cheaper alternatives are gaining ground
The cost pressure is creating an opening for lower-priced alternatives. PYMNTS reported in June that Chinese AI labs are attracting attention from enterprise buyers because their more efficient models and China's lower energy costs let them undercut U.S. providers on price. Procurement teams evaluating AI vendors in 2026 are now treating price per token as a primary selection criterion alongside capability benchmarks.
Open-source models are also seeing renewed interest. Companies are deploying them for internal or lower-risk tasks where a frontier model's performance advantage does not justify the cost premium. That tiered-model approach is becoming standard practice for cost-conscious AI programs.
Budget discipline is replacing blank-check experimentation
The PYMNTS Intelligence Enterprise AI Benchmark Report found that enterprises across financial services, insurance, healthcare, and media and advertising are continuing to increase AI budgets in 2026. But the report also noted a meaningful shift in posture: companies are becoming more selective, deciding which projects warrant real capital and which still need to prove their value before receiving it.
That selectivity is the direct operational consequence of token shock. Arora's 90% cost-reduction benchmark gives procurement and IT leaders a concrete yardstick: at current prices, broad deployment is financially constrained. At prices 90% lower, the economics of many use cases flip.
What this means for your team
- Audit your AI cost structure by use case now. Separate agentic workloads from single-turn interactions in your tracking; they have fundamentally different cost profiles and need separate budget lines.
- Build model-tiering into your AI procurement policy. Define which tasks require frontier models and which can run on older, open-source, or lower-cost alternatives. Cost governance should be a design requirement, not an afterthought.
- Add price-per-token to your vendor evaluation scorecard. Capability benchmarks alone no longer tell the full story. Efficiency metrics and pricing trajectories are equally material for multi-year contracts.
- Establish a cost-reduction trigger in your AI roadmap. Arora's 20%/90% timeline gives you a concrete signal to watch. If token prices hit those thresholds on schedule, use cases that are marginal today may become viable, and your deployment plan should account for that shift.
Sources
About the author
The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.