56% of tokens, 14% of spend: the split that should decide your AI stack
Your tokens doubled and your bill didn’t. That’s a market split, not a pricing glitch.
I wrote that the open-weight wave is good enough now because Qwen3.8-Max and GLM 5.3 run well outside a frontier API. This week Vercel’s September Production Index showing open-weight models taking 56% of gateway tokens makes it measurable: 56% of tokens flowed through open-weight models in August, but only 14 cents of every dollar did.
The split is simple: open-weight does most of the work for a small share of the money, while Anthropic still takes 64% of spend — so route by task price, not by model loyalty.
What does Vercel’s September gateway data actually show?
Open-weight models crossed a majority of token volume for the first time, and the price per token kept falling.
In August they handled 56% of all gateway tokens, up from 7% in December and 36% in July. Spend didn’t follow — open-weight was 14% of dollars.
Anthropic still took 64 cents of every dollar. That share hasn’t dropped below 61% since December, and its models held the top two spend slots every month. Inside it, teams shifted from Fable 5 (13.2% to 4.9% of spend in one month) to the cheaper Opus 5 (22.5%) at roughly half the price.
Average price per token fell 23.2% in August, the third straight monthly drop. Median team cost for 10M+ tokens fell 7.6%. Volume up, unit cost down, dollars still on frontier judgment.
Why do tokens and dollars diverge so sharply?
Cheap tokens and expensive judgment solve different jobs.
Open-weight models from DeepSeek, Moonshot AI, and Z.ai handle bulk drafts, extraction, and tool calls at a much lower price. The New Stack’s breakdown of token volume versus spend on Vercel’s AI Gateway frames it directly: 56% of tokens for 14% of spend means they absorb the high-volume, low-price work by design.
Frontier models keep the high-stakes calls. Long context and ambiguity where a wrong answer costs a user — that’s where teams still pay Anthropic prices. You don’t rent the best model to rewrite 200 descriptions. You rent it to decide what the product should say.
What should a solo founder do with this split?
Stop picking a model. Start picking a price per task.
If you call a model 50,000 times a day, you can’t pay frontier prices for all of it. You also can’t afford open-weight errors where a user will notice. Founders who ship treat the gateway as a router, not a favorite.
Try this split this week:
- Bulk and repeatable goes open-weight. Drafts, extraction, classification, and internal tool calls. Quantized 7B and 27B models are fast and cheap — your eval set catches misses before users do.
- Judgment and edit goes frontier. Final copy and ambiguous decisions where you’d want a paper trail. Pay more per token because fixing it later costs more.
- Everything gets graded. Twenty to fifty real inputs, a golden answer, and a score you run before you ship. Your AI eval set is your taste, made measurable — it’s how you know which model earned its keep.
How do you build a stack that benefits either way?
Make the model swappable and the workflow sticky.
Vercel’s gateway exists because routing between providers is now normal. If your product breaks when you swap the model, the model was the product — as I argued in your product is one model update away from being a free feature, that product doesn’t survive the next announcement.
Build what a gateway can’t commoditize:
- Put taste in a file, not a prompt. Palette, type scale, spacing, radius, and two things you never do. One page agents read every session.
- Keep data that compounds. At month twelve, does your product know something about this customer it didn’t at month one? If yes, cheaper tokens help you. If no, they help the clone.
- Design the interface as the moat. Generation is cheap. Trust at the handoff and recovery from failure aren’t. That’s where solo founders win.
Price per token will keep falling. Don’t bet on next month’s leader. Bet on a system that routes by task, grades itself, and stays consistent while everything underneath changes.
Frequently asked questions
In August, open-weight models handled 56% of tokens routed through Vercel AI Gateway for the first time, up from 7% in December, but accounted for only 14% of estimated spend. Anthropic models took 64% of spend and have held the top two spend slots every month since December.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- You don't have a prompting problem. You have a vocabulary problem
Ben Tossell shipped a style picker because founders can't name what they want — and Google's DESIGN.md proves a shared vocabulary beats a longer prompt every time.