TL;DR — Key Takeaways
– Open-weight AI models processed 56% of production tokens through Vercel’s AI Gateway in August, up from 36% in July and less than 10% in December 2025.
– Despite dominating token volume, open-weight models accounted for just 14% of estimated spending, reflecting their substantially lower average inference costs.
– Closed-weight tokens cost approximately 7.8 times as much as open-weight tokens on average, based on Vercel’s usage and estimated spending data.
Vercel data shows downloadable-weight models processing 56% of production AI Gateway tokens in August, while capturing only 14% of estimated spending.
For a while now, developers have loved open-weight and open-source Large Language Models (LLMs). Businesses? Not so much. Until now. According to Vercel’s AI Gateway, an AI LLM gateway for developers, open-weight AI models have the lead in token volume for the first time.
Open-weight models processed 56% of all tokens routed through the gateway in August. That’s a big jump from July’s open-weight share of 36% and less than 10% in December 2025. It’s only going to get higher.
According to Vercel CEO Guillermo Rauch in a LinkedIn post, August 22 was a “record day for open-weight share of tokens,” with 62% of traffic. He added, “This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc. need to be adapted to be model-agnostic.” Since Vercel’s gateway, the company claims, has tens of trillions of tokens transferring through it, those are significant numbers.
At the same time, though, the open models accounted for only 14% of estimated customer spending. Why the difference? Easy. Lower-cost open models have become a volume choice while closed vendors retain the market’s high end.
Specifically, based on my back-of-the-envelope calculations of Vercel’s data, closed-weight tokens cost about 7.8× as much as open-weight tokens. Open-weight tokens cost roughly 87% less on average.
At the same time, open-weight models are catching up with closed models in terms of productivity. According to Mozilla’s State of Open-Source AI 2026 report, the capability gap between leading open-weight and closed models has shrunk to about 3.3%. In particular, open-weight models are at or near parity on coding, instruction-following, and general-knowledge tasks. So is it worth spending 7.8 times more for a frontier model for day-in, day-out AI work? I don’t think so, and neither are companies.
It’s not just developers. Businesses are moving large production workloads to open-weight models. Enterprises are reserving more costly proprietary systems for work that justifies the premium. In its report, Vercel said, “Teams can now get more inference from the same budget and reserve frontier models only for the tasks that justify the premium.”
The rise of open-weight systems has also helped push the average price per token on AI Gateway down 23.2% in August. Among teams using more than 10 million tokens in both July and August, the median price per token dropped 7.6%, more than twice July’s 2.9% decrease.
There’s still some wiggle room in those numbers. Vercel’s figures cover anonymized, aggregate traffic through its own gateway. The company doesn’t capture direct API use, private self-hosted deployments, cloud-provider endpoints, or all enterprise contracts. Nor do the spending figures represent actual invoices: Vercel estimates them from published list prices.
On the proprietary side, Anthropic remained the largest recipient of spending on Vercel’s gateway even as open-weight traffic surged. Its models accounted for 64% of estimated August spending and have taken at least 61 cents of every dollar spent through the gateway every month since December.
The report also suggests that users are price-sensitive even inside proprietary-model lineups. Fable 5, which Vercel describes as Anthropic’s most capable model, saw its share of gateway spending fall from 13.2% in July to 4.9% in August. Over the same period, the less-expensive Opus 5 increased its spending share to 22.5%.
Vercel said Opus 5 costs about half as much per token as Fable 5, and that nine in 10 teams using Fable reduced their usage. More of those customers moved to Opus than any other model. In short, customers decided that Fable’s added capability wasn’t worth twice the price for most production workloads.
So, it’s not only the rise of open-weight that’s driving token prices down. It’s also coming from customers willing to “step down” within a vendor’s own product range when a less costly model meets the job’s requirements.
For Big AI, that’s concerning. Pricing, rather than who has the most impressive frontier model, is clearly becoming the real driver of which AI models get adopted. For AI companies whose leaders dream of becoming trillionaires, that’s unnerving.

