All notes

56% of tokens, 14% of spend: the split that should decide your AI stack

Your tokens doubled and your bill didn’t. That’s a market split, not a pricing glitch.

I wrote that the open-weight wave is good enough now because Qwen3.8-Max and GLM 5.3 run well outside a frontier API. This week Vercel’s September Production Index showing open-weight models taking 56% of gateway tokens makes it measurable: 56% of tokens flowed through open-weight models in August, but only 14 cents of every dollar did.

The split is simple: open-weight does most of the work for a small share of the money, while Anthropic still takes 64% of spend — so route by task price, not by model loyalty.

What does Vercel’s September gateway data actually show?

Open-weight models crossed a majority of token volume for the first time, and the price per token kept falling.

In August they handled 56% of all gateway tokens, up from 7% in December and 36% in July. Spend didn’t follow — open-weight was 14% of dollars.

Anthropic still took 64 cents of every dollar. That share hasn’t dropped below 61% since December, and its models held the top two spend slots every month. Inside it, teams shifted from Fable 5 (13.2% to 4.9% of spend in one month) to the cheaper Opus 5 (22.5%) at roughly half the price.

Average price per token fell 23.2% in August, the third straight monthly drop. Median team cost for 10M+ tokens fell 7.6%. Volume up, unit cost down, dollars still on frontier judgment.

Why do tokens and dollars diverge so sharply?

Cheap tokens and expensive judgment solve different jobs.

Open-weight models from DeepSeek, Moonshot AI, and Z.ai handle bulk drafts, extraction, and tool calls at a much lower price. The New Stack’s breakdown of token volume versus spend on Vercel’s AI Gateway frames it directly: 56% of tokens for 14% of spend means they absorb the high-volume, low-price work by design.

Frontier models keep the high-stakes calls. Long context and ambiguity where a wrong answer costs a user — that’s where teams still pay Anthropic prices. You don’t rent the best model to rewrite 200 descriptions. You rent it to decide what the product should say.

What should a solo founder do with this split?

Stop picking a model. Start picking a price per task.

If you call a model 50,000 times a day, you can’t pay frontier prices for all of it. You also can’t afford open-weight errors where a user will notice. Founders who ship treat the gateway as a router, not a favorite.

Try this split this week:

  • Bulk and repeatable goes open-weight. Drafts, extraction, classification, and internal tool calls. Quantized 7B and 27B models are fast and cheap — your eval set catches misses before users do.
  • Judgment and edit goes frontier. Final copy and ambiguous decisions where you’d want a paper trail. Pay more per token because fixing it later costs more.
  • Everything gets graded. Twenty to fifty real inputs, a golden answer, and a score you run before you ship. Your AI eval set is your taste, made measurable — it’s how you know which model earned its keep.

How do you build a stack that benefits either way?

Make the model swappable and the workflow sticky.

Vercel’s gateway exists because routing between providers is now normal. If your product breaks when you swap the model, the model was the product — as I argued in your product is one model update away from being a free feature, that product doesn’t survive the next announcement.

Build what a gateway can’t commoditize:

  • Put taste in a file, not a prompt. Palette, type scale, spacing, radius, and two things you never do. One page agents read every session.
  • Keep data that compounds. At month twelve, does your product know something about this customer it didn’t at month one? If yes, cheaper tokens help you. If no, they help the clone.
  • Design the interface as the moat. Generation is cheap. Trust at the handoff and recovery from failure aren’t. That’s where solo founders win.

Price per token will keep falling. Don’t bet on next month’s leader. Bet on a system that routes by task, grades itself, and stays consistent while everything underneath changes.

Frequently asked questions

  • In August, open-weight models handled 56% of tokens routed through Vercel AI Gateway for the first time, up from 7% in December, but accounted for only 14% of estimated spend. Anthropic models took 64% of spend and have held the top two spend slots every month since December.

About the author

mosh

mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.

Keep reading