All notes

You don't have a prompting problem. You have a vocabulary problem

My agent kept handing me the same polite dashboard. Light gray, blue accent, Inter, 16px radius — competent and forgettable. I rewrote the prompt longer. It got worse. I've written that taste is judgment, not preference — and AI can't fake it — and this week proved the gap isn't effort. It's language. I didn't need a better prompt. I needed words for what I actually wanted.

Generic output isn't a model failure — it's a vocabulary failure, and a short file of named choices fixes it faster than any prompt trick.

Why do prompts keep producing the same generic design?

Because vague words have one median output.

Tell an agent "clean and modern" and it picks the average of every clean and modern site it has seen — centered hero, grid of cards, blue gradient, Inter. Vercel's testing showed that same sentence produced different layouts in Claude and Codex. "Clean" means nothing until you name what clean does.

Specificity is not length. "Use 12px radius, editorial hierarchy, dense tables, no shadows" is short and precise. "Premium, trustworthy, innovative" is long and empty.

If you can't name radius, density, or hierarchy, the model fills the gap with its default.

What does Ben Tossell's Design Words actually fix?

It fixes the moment before you prompt.

On September 11, Ben Tossell shipped Ben Tossell's breakdown of why telling AI to design is hard and a tool called Design Words at bensbites.com/design-words. The idea is simple: browse themes and UI styles, see them applied to components, shuffle, then copy the words that made the preview you liked.

It's a vocabulary picker, not a generator. You pick swiss or editorial or brutalist, rounded or sharp, flat or shadowed, and the tool turns those choices into a sentence your agent can execute. For a non-designer, that's the missing step — you know what looks right when you see it, but you don't have the nouns to ask for it.

Tossell notes the previews still converge — that's honest. Early tools name the problem before they solve it perfectly. Spend 20 minutes picking words that match a reference you like, paste them into your agent, and compare against your last "clean and modern" attempt. The gap is obvious.

Why does DESIGN.md work when longer prompts don't?

Because a file is memory and a prompt is a wish.

A prompt lives in one chat and dies there. Every session you re-explain. Google Stitch docs on DESIGN.md as a portable design system treat the opposite as default: colors, type, spacing, grid, and what to never do — written once in markdown that Claude, Codex, and Cursor all read as constraints.

Vercel's write-up on how design.md keeps agent output on-brand is the receipt. Vercel split taste into three layers — judgment in design.md, mechanics in a stylesheet, proof in checks — and tested it frozen. Same prompts, same data, same viewport. Checks found 39 failures with design.md versus 91 without. Longer prompts never did that. A readable file did.

The win isn't the format. It's deciding once and making it stick. Write "lead with recommendation and both prices on the same scale" instead of "be premium." Name banned moves: no neon, no centered hero on every page. Agents follow what they can read.

How do you build your own design vocabulary this week?

Start with three screenshots and three sentences.

Pick one artifact you ship often — a landing, a dashboard, a report. Find three references closer to what you want. For each, write one sentence of vocabulary:

  • "Swiss grid, 8px radius, tight density, table uses full width"
  • "Editorial hierarchy, large serif for numbers, muted palette, no shadows"
  • "Brutalist, sharp corners, high contrast, single accent only"

Save those to a DESIGN.md at your repo root. Add two anti-patterns: what the page must never do. Keep it to one page.

Then generate once. Don't re-roll. Screenshot the first draft and write one correction as vocabulary: "hierarchy flat — lead with price comparison above the fold." That's your next line. After ten diffs, your file knows your taste better than your memory does.

Taste looks mysterious from outside. Inside it's named choices saved where agents can read them.

Frequently asked questions

  • They default to the median of their training data when your brief lacks specific vocabulary. Words like clean, modern, and premium map to the same blue gradient and thin icons every time because you did not name radius, density, palette, or layout intent.

About the author

mosh

mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.

Keep reading