Task-specific agents ship. General agents demo
Everyone's agent demo looks the same right now: "just tell it what you want and it figures it out." I tried that. It figured it out once, then guessed the next three times.
The data caught up this month. McKinsey's State of AI report finds 62% of organizations now experiment with AI agents but only 23% have scaled one in a single function — most are stuck between pilot and production. I've been arguing that AI ideas are compressions of work that already exists, not inventions of new work. Agents prove it. The ones that ship compress one workflow you already do. The ones that stall promise to do everything.
Task-specific agents ship because you can define, grade, and trust them. General agents demo well and fail quietly.
Why does the everything agent fail after the demo?
It has no contract. You give it a big goal — "handle my inbox" or "manage my files" — and it has to invent scope, permissions, and done-ness on the fly.
Anthropic's own building advice says the same. Anthropic's guide to building effective agents warns that agents need narrow tool sets, clear handoffs, and a human in the loop for judgment. A general agent violates all three. It touches too many tools, decides when it is done, and hides where it guessed. You notice the win once. You pay for the guesses the other four times.
Scaling needs a permission boundary, an eval set, and a team willing to hand real work over. You don't scale "do anything." You scale "do this one thing the same way ten times."
What does a one-workflow agent actually do?
One input, one outcome, one owner. That's it.
My shippable agents all look boring: take this type of note, pull three fields, draft the same update format, save it where I already save it, and leave a diff I can approve in one tap. The output has a shape I wrote down before I automated anything.
That constraint makes everything else cheaper. Permissions are simple because the scope is two folders, not your drive. Review is fast because I'm checking format and facts, not intent. And evals work — ten real inputs, ten golden outputs, scored on every change. Nielsen Norman Group's research on AI agent UX frames it as trust through legibility: show what the agent did, why, and how to undo it. Narrow scope is how you get that without building a second product.
If you can't write the done sentence — "a weekly changelog draft in this doc from these three sources" — you don't have a workflow.
How do you pick the workflow worth compressing?
Pick the one you hate doing by hand. I use three filters:
1. Repetition with variation. You do it weekly, but it is never copy-paste. Changelogs, triage notes, proposal first drafts, image alt text. If it's identical each time, it's a template. If it needs judgment each time, it's a compression candidate.
2. Visible done. You know it is finished without asking someone else. That gives you a pass rate. I picked "draft the release note from merged PRs" before "answer any question about the repo" because I can grade the first in seconds and the second not at all.
3. Contained blast radius. If it fails, the fix is reverting one file, not unwinding five systems. Solo founders live here. Ship the agent that writes a draft over the one that sends the email.
Start with the most annoying repetition that passes all three. Not the most impressive. The most annoying.
Where does taste decide if it ships?
At the review gate, not the prompt.
Two agents can use the same model and produce different products. The difference is the rubric you hold against the output before it reaches a user. I learned this after watching agents drift between sessions — they don't remember what I found good last week unless I wrote it where they can read it. A short taste file with format, tone, and edge-case rules keeps the agent inside the same box I would use myself.
That gate is also where solo founders have an edge. A team needs alignment to say no. You just need five checks: does it match the type scale, does it use only our components, does it keep copy plain, does it handle empty state, does it leave an undo path? Ship the one you would sign.
Agents that do everything ask you to trust the model. Agents that do one workflow ask you to trust the workflow. The second earns a second week of use.
Frequently asked questions
They try to handle every task with one prompt and one permission set. The result is vague output, too many approval asks, and no clear way to grade success. A broad agent looks good in a demo but breaks on edge cases because there is no defined workflow to test against.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- ScoreAi: match search redesign that lifted engagement
ScoreAi had demand but its match search leaked engagement. A conversion-first rebuild of hierarchy, scanning, and onboarding fixed discovery.