Your AI stack just shrank to three layers. Cut the rest
This week your solo stack got smaller whether you noticed or not. Vercel put five models behind one endpoint, turned builds and sandboxes and functions into one shape, and Claude stopped making you re-explain yourself between chat and Cowork. I felt it — fewer tabs, fewer re-briefs. I wrote that you don't need another AI tool, you need fewer — this week the platforms did the cutting for you.
Your AI stack is now three layers — a gateway for models, unified compute, and scoped memory — and everything else has to earn its keep or go.
Why did your stack shrink this week?
It shrank because the platforms ate the glue you used to manage by hand.
On September 1, Vercel launched Vercel's Fluid compute launch that unifies builds, sandboxes, and functions — one system that spins up the shape you need on demand instead of three separate primitives. Same week, the Vercel AI Gateway overview with zero-markup model routing added five models in six days — Claude Fable 5.1, Muse Spark 1.3, Gemini 3.8 Flash, GPT 6 Astra, and Qwen 3.8 Max — all behind one base URL with failover built in.
At the same time Anthropic merged memory. The Anthropic announcement unifying Claude memory across chat and Cowork moved topics from post-chat summaries to live updates, shared between chat and Cowork with an editable Topics list. You explain a metric once and every Cowork deck after uses it.
What are the three layers that actually remain?
Three verbs cover the loop solo founders repeat weekly.
1. Route — the gateway. You write ai-gateway.vercel.sh once and swap models with a string. No SDK rewrite, no extra dashboard. Vercel handles bring-your-own-keys, pay-as-you-go with no markup, and failover.
2. Run — unified compute. Fluid replaces "build machine vs function vs sandbox" with one system. Push to main, run an agent in a sandbox, or schedule a workflow — same env, same primitives.
3. Remember — scoped memory and taste. Claude's Topics plus your own AGENTS.md and design tokens. Gateway and compute are boring by design. Memory is where you put the taste file and spacing scale so agents show up briefed.
Everything outside those three is either a real fourth verb (payments, database, evals) or duplication.
How do you cut tools without breaking shipping?
Don't cut by category. Cut by verb and by proof.
I do a 20-minute audit every first Monday. Open billing, list each tool, and ask: what did this replace this month that I would have done by hand?
Two tools route models? Keep the gateway, cut the wrapper. Two places run code? Keep Fluid, cut the extra sandbox. Two memories? Keep Topics + one repo file, cut the doc you re-paste.
Use a simple bar: a $20 tool needs to save an hour you can point to in the last 30 days. If you can't name the hour, cancel. You can resubscribe — most times you won't miss it.
Keep the eval habit even as you shrink the stack. Five inputs with known good outputs, scored after every prompt change. When I cut a review tool last month the eval set still caught drift in empty states that week.
When should you add a tool back?
Add only when a missing verb blocks you, not when a new model tops a chart.
If you have no way to take money, add Stripe. If you have no way to test the second session, add that test. If you have no way to reach users, pick one distribution channel and commit to it for 30 days.
What isn't a gap: a second gateway that does what AI Gateway already does, a second memory app that copies Topics, a new model that is 2% better on a benchmark you don't ship against. I keep one rotation slot — try for a week, measure against a real task, keep only if you feel the win in your day.
September is a decision month. Q4 budgets get set and teams pick the gateway they will run next year. For solo founders the win isn't the next model. It's three layers you trust, taste where agents can read it, and the hours you get back for the work AI can't do — deciding what good looks like and holding it.
Frequently asked questions
Vercel collapsed compute and model routing into two surfaces — Fluid for builds, functions, and sandboxes, and AI Gateway for model calls — while Anthropic merged Claude memory across chat and Cowork. Three layers now do what used to take seven tools and manual wiring.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- ScoreAi: match search redesign that lifted engagement
ScoreAi had demand but its match search leaked engagement. A conversion-first rebuild of hierarchy, scanning, and onboarding fixed discovery.