Chinese AI models now handle nearly a third of enterprise tokens at one-tenth the cost of US rivals
Chinese AI models now manage nearly a third of enterprise tokens, achieving this at one-tenth the cost of their U.S. counterparts. Companies like DeepSeek, Z.ai, and ByteDance are rapidly capturing production workloads. The significant price advantage is a major driver behind this shift towards Chinese AI solutions.
This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.
Key facts, context, and what it means, in one minute.
Key takeaways
Chinese AI models handle nearly a third of enterprise tokens.
These models operate at one-tenth the cost of US competitors.
DeepSeek, Z.ai, and ByteDance are significant players in this market shift.
Chinese-built open-weight AI models processed 29% of all tokens flowing through Vercel's production AI gateway in June 2026, up from roughly one-ninth of total volume in April, according to Vercel's AI Gateway Production Index. The cost arithmetic behind that shift is stark: those models consumed less than 4% of total spending, priced at approximately one-tenth the average token rate of US frontier systems on the platform.
The same trend is visible across other developer infrastructure. On OpenRouter, a platform that routes developer traffic across a wide range of AI models, US companies directed more than 30% of their tokens to Chinese models in every week since February 8, 2026, according to CNBC. That figure has reached as high as 46%. The prior 12-month average was 11%, and the first-half 2025 figure was just 4.5%, making the acceleration within 2026 alone notable for any team tracking vendor concentration risk.
Price pressure is reshaping production routing decisions
The shift is not simply about experimentation. Vercel's index, which tracks tens of trillions of tokens monthly between production applications and model providers, reflects live workload routing in deployed systems, not developer sandboxes. When a task does not require frontier-level accuracy, teams are increasingly sending it to the cheapest model that clears a quality threshold.
When a task doesn't need the best model, teams are beginning to route it to the cheapest one that's good enough, and the recent wave of models coming out of China is winning that trade.
Harpreet Arora, Vercel's head of agentic infrastructure, made that point to both Computing and CNBC, noting that privacy protections and data residency requirements remain the primary gating factors for customers deciding whether open-weight models can enter their production stacks. For organizations operating under strict data governance rules, those requirements determine whether the cost savings are accessible at all.
The dynamic is playing out concretely at the company level. AI startup Lindy moved 100% of its traffic from Anthropic's Claude to DeepSeek in June 2026, according to CNBC. CEO Flo Crivello told CNBC the cost curve dropped sharply immediately after the switch, and the move is projected to save the company millions of dollars within months. Lindy's case is an early signal of what cost-conscious procurement decisions look like in practice when cheaper alternatives clear the quality bar.
Kyle Chan, a fellow at the Brookings Institution's John L. Thornton China Center, told CNBC that enterprise cost consciousness is a direct contributor to adoption. Where companies previously prioritized getting AI deployed regardless of which model they used, rising token prices from leading US labs have made model selection a budget-line decision rather than a purely technical one.
DeepSeek, Z.ai, and ByteDance each carve out distinct niches
DeepSeek is the furthest advanced in production volume. It accounted for 22.6% of token volume on Vercel's gateway in June 2026, ranking third overall behind Anthropic and Google, and within two percentage points of Google's 24% share, according to Computing. Google's share itself declined from a surge in April, narrowing the gap.
Z.ai's GLM 5.2, released in June 2026, posted the fastest adoption rate of any model Vercel tracked in 2026, according to Arora's comments to CNBC. That speed of uptake suggests developer teams are actively evaluating new Chinese model releases on release day rather than waiting for broader industry validation.
Video generation is a separate competitive zone. ByteDance's Seedance captured nearly half of all video-related spending on Vercel's gateway despite producing roughly one-third of video output, according to Computing. That spending premium indicates the model is being used for higher-value or longer video tasks rather than simple previews. In image generation, OpenAI retained the largest share with more than half of all images created through the gateway, followed by Google's Nano Banana model at nearly 39%.
US frontier providers retain revenue dominance, but volume tells a different story
The four leading US frontier AI companies, Anthropic, Google, OpenAI, and one other, collectively held 95% of total spending through Vercel's AI Gateway in June 2026, according to Computing. Anthropic alone captured 61% of spending despite processing only 32% of tokens. The company's concentration in high-stakes workloads explains the gap: coding assistants, back-office automation, and application generation are categories where enterprises pay a premium for accuracy and tend not to route traffic to unproven or lower-cost alternatives.
Back-office AI agents were the single most expensive workload category relative to token count, consuming 14% of total spending while representing only 5% of token volume, per the Vercel index. That ratio reflects the higher per-token complexity of agentic tasks, where model errors carry real downstream business cost.
Token volume on Vercel's gateway grew 29% in June while spending grew 27%, according to Computing. Average token prices held flat, the result of two forces canceling out: cheaper open-weight workloads pulling prices down, and a roughly 12% price increase from leading closed-weight frontier providers pushing them back up. The net effect is a market where the volume of AI work is expanding fast, but the cost per unit is no longer falling.
What this means for your team
- Audit your current model routing logic. If your AI gateway or orchestration layer routes all workloads to the same frontier model by default, you may be paying frontier prices for tasks that a lower-cost open-weight model can handle. Segment workloads by accuracy requirement and evaluate whether Chinese open-weight models clear your quality bar for the lower-stakes tier.
- Validate data residency before routing. Vercel's Arora specifically flagged privacy protections and data residency requirements as the primary blockers for open-weight adoption in production. Before any procurement decision involving a non-US model host, confirm your legal and compliance team has reviewed jurisdictional data handling obligations.
- Treat model selection as a recurring budget decision. With frontier token prices rising roughly 12% in a single month and Chinese alternatives priced at one-tenth that rate, model pricing is no longer stable enough to set once and forget. Build a quarterly review of cost-per-token into your AI operations cadence.
- Watch Z.ai GLM and DeepSeek for agentic use cases. Both are in active production on major developer platforms now. If your team is building or evaluating AI agents, the performance-to-cost ratio of these models warrants a direct benchmark against your current provider before your next contract renewal.
Sources
About the author
The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.