Your AI prototype passes the demo. The second session decides if it ships
Your demo wows on the first run. Fresh data, perfect prompt, no history to get in the way. The second time someone opens it — empty workspace, expired permission, stale context — is where most AI products quietly fall apart.
I've felt this gap myself, and I wrote that your prototype works once but determinism is what ships it. That first-time magic isn't the product. What happens when the user comes back is.
The second-session test says your AI product isn't done until a returning user gets a clear empty state, correct old data, and a one-step recovery from every error — without re-teaching the tool who they are.
Why does every AI demo feel finished until the second session?
Every AI demo feels finished because it only exercises the happy path with new data. The model generates a polished first screen, the data is fresh, and there's no memory to get wrong.
That hides the real work. Designer Fund's 2026 AI in Design report found weekly AI use jumped from 54% to 91% in a year, but unreliable output quality stayed the top complaint. Figma's 2026 report on how product teams use AI showed stacks jumping from three tools to seven.
What does the second-session test actually check?
The second-session test checks the four places where returning users actually quit.
Empty. What does a new or cleared workspace look like after the demo data is gone? If it's a blank canvas with no action, users assume the product broke. A good empty state names the next step in one sentence and shows a single button that does it.
Populated but stale. What does last week's data look like today? Does the old AI output still carry its reasoning, can the user edit or rerun it without starting over? If your product only looks right with fresh data, it fails here.
Error and permission. What happens when the API key expired or the model timed out? Most prototypes show a red toast and stop. The test asks: can the user fix it in one step right where they are, with the original input still saved?
Memory that helps. What did you keep from last time, and can the user see and delete it? Keep taste and intent, forget pasted secrets and one-off chats.
Run those four checks with a real account that has history. If any makes you re-paste or re-explain, you haven't shipped yet.
How do you build for the second session without slowing to a crawl?
You don't slow down. You add guardrails so the fast part stays fast.
Lock the set before you scale it. Pick your buttons, radii, type scale, and spacing once — the same advice from why curation is the job now. When the set is locked, every AI pass has something to fail against.
Add three screens to every flow, not one. For each happy-path screen, generate its empty, its error, and its populated-aged version in the same run. It takes ten extra minutes and catches most second-session bugs before users do.
Write a short rubric and make it the gate. Anthropic's guide to building reliable evals for agents frames this well: score the output against what you decided was good, not how much context you carried. My version is three lines: does it match intent, does it stay consistent, does it handle an edge without falling apart? If it fails one, it doesn't ship.
Test returns, not just fresh prompts. Open the product tomorrow. Use the account that already has data. Paste the wrong file on purpose. The demo tests if you can generate. The second session tests if you can recover.
When have you passed the second-session test?
You've passed when a returning user can do three things without thinking: recognize the product, pick up where they left off, and recover from failure in one step.
A demo proves you can generate a good first impression. The second session proves you respected the user's time when they came back. Ship when the return feels as considered as the reveal.
Frequently asked questions
It is the check that your product works when a user returns: empty states are designed, past data loads correctly, errors have a way forward, and memory from last time helps instead of confusing. If it only works on first-run demo data, it fails the test.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- ScoreAi: match search redesign that lifted engagement
ScoreAi had demand but its match search leaked engagement. A conversion-first rebuild of hierarchy, scanning, and onboarding fixed discovery.