All notes

Vercel ran it 200 times to make taste repeatable

Vercel ships pages that have to feel like Vercel. When agents started building reports and proposals outside the codebase, the output drifted.

I've felt that drift — taste was clear in my head and invisible to the agent. I wrote about how your agent can't see your taste until it's readable — Vercel just showed what it takes to fix that. More than 200 runs.

Vercel made taste repeatable by splitting it into three layers — guidance for judgment, a stylesheet for mechanics, and checks for proof — and testing every rule against real pages.

Why didn't a longer prompt fix taste drift?

A longer prompt can't carry context that isn't there.

Inside Vercel's repositories, agents read the real system. The Vercel's product-design skill that lives inside the repository sits next to components and shipped pull requests.

Outside the repo that grounding disappears. Vercel first merged the skill into one public prompt. Models still diverged — "clean" meant different pages to Claude and Codex with no components to anchor it. Some models invented rules that were never written.

They threw that prompt away and rebuilt against outputs. Every new line had to improve a real page or it didn't get in. Taste isn't a description you make more precise. It's a decision you can show in a before and after.

What lives where: guidance, stylesheet, or check?

Three buckets, three jobs.

Guidance in design.md is judgment — who the page is for, what they came to decide, how evidence earns it. Vercel's Vercel's write-up on how design.md keeps agent output on-brand starts with reader and job, then order, then copy honesty. Example: put the recommendation and both prices on the same scale at the top. That's composition, not color.

They also named what to avoid: gradients that add nothing, cards inside cards, every page as hero plus grid. Agents default to decoration unless you tell them not to.

Mechanics move to a stylesheet. Type scale, spacing, table width, chart styles were already decided. Vercel published them as CSS with class names. No invented spacing, no tokens in the context window.

Proof is checks and review. Code catches width bugs or missing labels. People judge whether the page gives the reader what they came for. If a rule never changes, make it a component. If a machine can test it, make it a check. Only judgment belongs in guidance.

How do you test if your taste file works?

Freeze everything except the file.

Vercel locked seven tasks — renewal proposal, performance report, benchmark, and four more — with the same prompt, fake data, and viewport. Each round generated all seven on Claude Opus 4.8 and Codex with GPT-5.5 and saved inputs, model version, and screenshots. Then blind A/B with shuffled outputs.

To see if corrections stuck, they compared three scenarios, six first-attempt pages, no re-rolls. Code checks counted known failures — 39 with design.md, 91 without, 57% fewer in that test.

Useful and limited. Checks only see failures already named. Six pages don't prove reliability. And every page still had a blocker that stopped shipping. The honest read: repeat mistakes showed up less, so humans spent less time on the same corrections.

Weekly, feedback from Slack @design-agent threads, GitHub reviews, and Figma gets grouped and routed to the right layer. One-off quirks stay in the backlog until they repeat.

You don't need 200 runs. You need one frozen scenario and the discipline to keep first attempts.

What should solo founders copy this week?

You don't have a review team. Your last ten corrections are the team.

Pick one artifact you ship — a report, proposal, or changelog page. Save the unassisted first draft as your baseline.

Open your last ten fixes. Real diffs where you moved the recommendation up, cut a box, fixed a table. Route each:

  • Needs judgment? One sentence — "lead with recommendation and comparison on the same scale."
  • Never changes? Make it a component and require that class.
  • Testable? Add a tiny check — grep for a banned pattern or a width assertion.

Run the same prompt once with guidance and once without. Keep first attempts, shuffle, review blind. Did you spend less time on repeats? If not, the rule was prose, not progress.

Taste turns into engineering when every line has a place, a test, and an owner. You can't out-review a team, but you can out-system one artifact at a time.

Frequently asked questions

  • Vercel first merged its internal product-design skill into a public prompt, but models read words like “clean” and “layered” differently without real components or shipped examples to ground them. Different models turned the same sentence into different layouts, so adding more adjectives made drift worse, not better.

About the author

mosh

mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.

Keep reading