📑 Table of contents

Here's the English translation: **Title:** Open models already serve 56% of tokens on Vercel: the silent migration of enterprises away from OpenAI and Anthropic

Actu IA 🟢 Beginner ⏱️ 14 min read 📅 2026-09-28

Open models already run 56% of Vercel's tokens: the silent migration of enterprises away from OpenAI and Anthropic

🔎 The crossover happened, without an announcement

There was no keynote, no viral tweet, no dramatic price cut. And yet something tipped this summer in American AI infrastructure: in August 2026, open-weight models executed 56% of token volume on Vercel's AI Gateway, up from less than 10% in December 2025. The peak reached 62% on August 22, according to the Production Index of Vercel's AI Gateway published on September 17.

The Financial Times confirmed the shift on the enterprise side. Mentions of "open" or "open weight" models in US earnings calls and investor conferences multiplied sixfold in one year (AlphaSense data, August–September 2026). And these aren't bootstrapped startups: AT&T runs roughly 40% of its AI workloads on open models, with a target of 70% within a year.

What's at stake here isn't an ideological debate about open source. It's a blunt economic trade-off. A closed token costs on average 7.8 times more than an open token. CIOs have figured this out, and they're rewriting their architectures accordingly.


Key takeaways

  • 56% of production tokens on Vercel's AI Gateway are now run by open models (August 2026), up from 7% of volume a year earlier and under 10% in December 2025.
  • AT&T runs ~40% of its AI workloads on open models (Llama, Gemma, Nemotron), with savings of 80–90% on some applications, according to the WSJ.
  • Smart routing pays off: at AT&T, routing to open models cut code costs by up to 56% for only a ~2% drop in quality.
  • China is riding the wave: Llama, Qwen (Alibaba), GLM (Z.ai), and DeepSeek are capturing a growing share of volume. The capability gap with closed models has narrowed to ~3.3% (Mozilla, 2026).
  • But the money stays closed: open models account for just 14% of estimated spend. Anthropic still captures 64% of August's dollars.
  • Price signal: OpenAI cut GPT-5.6 Luna pricing by up to 80%, and Anthropic is pushing Opus 5 at half price. The pricing pressure is only just beginning.

Tool Main use Price (September 2026, check the official website) Ideal for
Vercel AI Gateway Multi-model routing, production index Pay-as-you-go, no markup on tokens Teams that want to test open models without changing their infra
Hugging Face Hosting and deployment of open weights Free + paid Inference plans Self-hosting Llama, Qwen, GLM
LiteLLM Open-source proxy/router between models Open source, free The router used by AT&T to score requests
OpenRouter Unified API access to hundreds of models Pay-per-token, ~5% fee Comparing open and closed models in production
Hostinger VPS hosting to self-host your open models From a few dollars/month (September 2026) Developers who want to keep full control of their data

Why now? The number that changes the conversation

The shift didn't happen because open models became "good enough" one morning. It happened because three curves crossed.

First curve: volume. According to Vercel's historical data, open models were serving 11% of tokens in April 2026, 29% in June, 36% in July, 56% in August. This is no longer a trend, it's an exponential trajectory. Saanya Ojha notes that the live leaderboard on September 18 even showed ~78% open tokens, and that Moonshot and DeepSeek climbed to 3rd and 4th place in spend — ahead of OpenAI if you add Z.ai to the group.

Second curve: quality. Mozilla's "State of Open-Source AI 2026" report measures the capability gap between the best open and closed models at roughly 3.3%. Near-parity on coding, instruction following, and general knowledge. When the gap was 20 points, the closed-model premium was justified. At 3 points, it becomes indefensible for 90% of use cases.

Third curve: price. The closed token costs ~7.8× the open token on average on the gateway. The average token price fell 23.2% in August alone. Chinese models run 60 to 90% cheaper than OpenAI or Anthropic, according to StartupFortune's analysis.

These figures also line up with our investigation into the 46% of US companies' AI tokens going to Chinese models, published in late September: the share of US tokens running on Chinese models via OpenRouter has stayed above 30% every week since February 8, peaking at 46%.


AT&T, the case study everyone is watching — Direct answer

AT&T demonstrates that you can move 40% of your AI workloads to open source without degrading the user experience — and that the secret isn't the model, it's the routing.

The telecom giant isn't replacing Claude or GPT with a single model. It uses LiteLLM, an open source router, which scores the difficulty of each request. Simple questions (which make up the vast majority) go to open models: Meta's Llama, Google's Gemma, Nvidia's Nemotron. Complex code stays on a frontier model.

The results, reported by the Wall Street Journal:

Metric Result at AT&T
Share of AI workloads on open models ~40% (target 70%)
Employees covered by Ask AT&T ~100,000
Savings on certain applications 80 to 90% (WSJ source)
Cost reduction on code Up to 56%
Measured quality drop ~2%

Andy Markus, AT&T's chief data & AI officer, sums up the mindset: open models fine-tuned with the company's proprietary data perform "as well or better" than closed alternatives on specific tasks. And he adds, in the FT: "The open models are getting better and better."


The Volume Paradox: 56% of Tokens, 14% of Dollars

Here is the most misunderstood point of the whole story, and the one that will help avoid bad strategic decisions.

Open models are winning in volume, but not (yet) in value. In August 2026, they account for 56% of production tokens but only 14% of estimated spend. The four big American frontier labs were still capturing most of the dollars, with Anthropic at 64% of the month's spend.

Why? Three structural reasons.

The frontier keeps a premium. For very high-value tasks — complex agentic reasoning, critical code — companies are still willing to pay for the best model available. The 2026 agentic leaderboard shows it: OpenAI's GPT-5.5 (98.2) and Google's Gemini 3 Pro Deep Think (95.4) remain ahead of the best self-hosted models like Moonshot's Kimi K2.6 (88.1) or Z.AI's GLM-5 Reasoning (82).

Usage is not homogeneous. 80% of requests are simple and routable to open models. The remaining 20% concentrate the value — and a large share of the budget.

Infrastructure is lagging behind. Guillermo Rauch, CEO of Vercel, sums it up well: "This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc. need to be adapted to be model-agnostic." Tools still need to become model-agnostic.

But the internal dynamics within closed labs are already brutal. At Anthropic, the spend share of Fable 5 (the most capable model) fell from 13.2% to 4.9% in one month, while Opus 5 — roughly half the price — rose to 22.5%. Nine out of ten teams using Fable reduced their usage of it. Even among loyal customers, downgrading happens as soon as quality allows it.


The Chinese wave: Llama is no longer the only horse

The open-weight movement is no longer a Meta phenomenon. Inference spending is structurally shifting toward three Chinese families: Qwen (Alibaba), GLM (Z.ai), and DeepSeek.

The figures in Vercel's July report already showed it: DeepSeek reached 22.6% of token volume on the gateway (3rd place, within two points of Google), and GLM 5.2 — MIT-licensed, at about one-fifth the price of Opus 4.8 — saw its daily volume multiplied by ~50 in the two weeks after its API opened.

The ecosystem is following suit. Chinese open models have amassed 3.2 billion downloads on Hugging Face, roughly double the American total. Traffic for open models on OpenRouter went from ~1 trillion tokens per week a year ago to ~80 trillion.

It's no accident that GLM and Qwen released two flash models on the same day: Chinese labs have understood that the battle is fought on the price/performance ratio in production, not on demonstration benchmarks. In the same vein, DeepSeek V4.1 Flash and its architecture optimized to cut the KV cache by a quarter illustrates the efficiency race that makes these models so competitive on inference cost. And MiniMax M3, with its 1M-token context, goes head-on after the long-context use cases where proprietary models charge the most.

To be factual: this shift toward Chinese models raises real sovereignty and compliance questions for regulated companies. That's precisely what our CNBC investigation into American tokens flowing to China documents — and it's also why OpenAI, Google, Microsoft, Anthropic, and more than 100 companies signed an unprecedented collective appeal warning of an imminent wave of AI cyberattacks: securing models and the AI supply chain is becoming a board-level issue, not just an IT one.


What this changes for your API pricing — Direct answer

Price pressure on closed models is already visible, and it's going to intensify. If you're negotiating a frontier API contract without leveraging open competition, you're overpaying.

Pricing reactions are documented:

  • OpenAI cut GPT-5.6 Luna prices by up to 80% depending on token type — a direct response to the volume hemorrhage.
  • Anthropic now sells Opus 5 at half the price of the previous generation, after watching Fable 5 collapse in its own spending figures.
  • The average token price on the Vercel gateway dropped 23.2% in August alone. For teams exceeding 10 million tokens, the median fell 7.6% in one month (versus 2.9% in July).

For teams that want the best of both worlds, the winning architecture in 2026 looks like AT&T's: a router that sends trivial requests to low-cost open models, and reserves the frontier for what deserves it. That's exactly what the 2026 agentic leaderboard does with models like OpenAI's GPT-5.3 Codex or Claude Sonnet 4.6 for code, while letting Llama and Qwen absorb the volume.

Guillermo Rauch sums up the moment: "This is very likely just the start, because enterprise adoption is still early." Translated: current closed API pricing isn't a floor, it's a provisional ceiling.


How to test the migration without breaking everything

No need to bet everything on open models overnight. The method used by successful companies comes down to four steps.

1. Classify your requests. Log a month of traffic and categorize it: simple questions, RAG, code, complex reasoning. For most applications, trivial requests dominate by far.

2. Benchmark on YOUR data. Public benchmarks aren't enough. AT&T tunes its open models with its proprietary data and finds they perform "as well or even better" than closed ones on its specific tasks. Your mileage will depend on your domain.

3. Route, don't replace. Use a router (LiteLLM, or a gateway like Vercel) that scores the difficulty of each request. Keep a fallback path to a frontier model.

4. Measure the quality/price ratio, not quality alone. A 2% drop in quality for 56% in savings is a trade most products can accept. A 2% drop on a critical legal workflow — maybe not. It's a product decision, not an architecture decision.

For teams that want to self-host (compliance, sensitive data, avoiding vendor lock-in), models like Kimi K2.6 in self-host or GLM-5 Reasoning from Z.AI are in the top 15 of the 2026 agentic leaderboard — a VPS-style hosting offer is enough to get started on test volumes.

❌ Common mistakes

Mistake 1: Confusing tokens and spending

The headline "56% of tokens" suggests open models dominate the market. False: they account for only 14% of estimated spend. If you're managing an AI budget, always look at both curves. A closed token costs ~7.8× an open one — converting volume → value isn't automatic.

Mistake 2: Switching 100% of traffic to open "because it's cheaper"

AT&T's savings come from routing, not replacement. Complex code stays on a frontier model. Teams that switch everything lose quality where it matters, then backtrack at the cost of internal credibility.

Mistake 3: Benchmarking on public datasets

Your application isn't MMLU. An open model can be excellent on product RAG and poor at your JSON output format. Test on your real data, with your evaluation criteria, before deciding.

Mistake 4: Neglecting governance of Chinese models

Qwen, GLM, and DeepSeek are excellent and cheap, but they raise compliance questions (data, jurisdiction, licenses) depending on your sector. The sound approach: model the risk as you would for any third-party vendor, and keep the ability to swap.


❓ Frequently Asked Questions

Are open models really on par with closed models?

On common tasks, almost yes: the gap measured by Mozilla in 2026 is ~3.3%, with parity on code and instruction following. Frontier models retain an edge on very complex agentic reasoning, but that edge narrows every quarter.

Should you leave OpenAI and Anthropic now?

No, you need to route. Keep a frontier model for high-value tasks and send the trivial volume to open models. That's exactly what AT&T and most mature teams do — and that's where the 50–90% savings lie.

Are Chinese open models safe for enterprise use?

It's a question of governance, not dogma. Check the license (MIT for GLM, for example), the hosting jurisdiction, and the possibility of self-hosting. Open weights let you host on your own infrastructure, which closed models don't allow.

What does this shift mean for API pricing?

Structural downward pressure. OpenAI has already cut GPT-5.6 Luna by up to 80%, Anthropic halved Opus 5. The average token price on Vercel dropped 23.2% in August. The frontier will become a high-margin premium product, the rest a commodity.

Is Vercel AI Gateway representative of the market?

It's a proxy, not a census. But the tens of trillions of monthly tokens, the corroboration from earnings calls (FT/AlphaSense), and the AT&T/Coinbase cases make the trend hard to dispute.


✅ Conclusion

The shift toward open models is no longer a thesis, it's a measured fact: 56% of production tokens on Vercel, 6x in earnings calls, 40% of AT&T's workloads — and savings that run into the tens of percent. The real question for 2026 isn't "open or closed?" but "what percentage of my traffic truly deserves to pay 7.8× more?". If you haven't yet set up multi-model routing, that's the project to kick off this quarter — and if the topic interests you, our breakdown of the July Vercel report details the full trajectory.