The Empathize Stage of Design Thinking: How to Stop Confirming What You Already Believed
The empathize stage is where design thinking is most often faked: a week of talking to users that ends up ratifying the roadmap. What the stage is for, which methods belong in it, why empathy maps are a synthesis tool, and how to tell when you're done.
A team blocks out a week to empathize with users before rebuilding onboarding. They run five calls. Every call opens with a walkthrough of the current flow and the question where did that feel confusing? Participants point at the third screen, because you have to point somewhere. The team ships a redesigned third screen. Activation doesn't move, and nobody can say why, because the study never asked what people were trying to do before they arrived — only how they felt about the thing that already existed.
The week was thorough. Well-attended, well-documented, and it produced a deck. It still failed, because the research took the product as its subject, and in design thinking the empathize stage is the one place where the product isn't supposed to be the subject at all.
What the stage is for
The five stages — empathize, define, ideate, prototype, test — get drawn as a pipeline, which makes the first one look like a warm-up before the real work. Each stage produces an input the next one can't manufacture on its own. What empathize produces is an account of the user's situation that doesn't mention your product.
That constraint sounds arbitrary until you try to write one. A finding like users find the invite step confusing reviews a screen. An account of a situation reads more like this: the person setting up the workspace is rarely the person who'll use it daily, and they're doing it in a fifteen-minute gap before a meeting, from an email thread that already has four people arguing in it. There's no interface anywhere in that sentence. It still rules out about half of what a normal onboarding flow asks for.
When the stage goes well it hands over three things:
- The job as they understand it, in their vocabulary. Not your feature names, and not the category name you use in sales calls.
- The context around the moment your product appears. What happened before, who else was involved, what they'd already tried, what deadline they were under.
- The cost of the current workaround. What it costs them to keep doing it, in time or money or standing with a colleague. That figure is the closest thing you'll get to a predictor of whether they'll switch.
Define turns those into a problem statement. Hand it screen-level feedback instead and it has nothing to work with, so it quietly reverts to the assumption everyone walked in with.
Empathize vs sympathize, and why the difference shows up in the transcript
The distinction gets taught as a matter of attitude: sympathy is feeling for someone, empathy is feeling with them. That framing is hard to act on once you're in the room. The operational difference is what you do with a complaint.
Sympathy accepts the complaint and moves to the fix. A participant says the export is frustrating, you say that makes sense, we've heard that a lot, and you write down "improve export." Empathy treats the complaint as a symptom with a history behind it: when was the last time you exported something? Walk me through that. Two minutes later it turns out they export because their finance lead won't open a shared link, and the real constraint has nothing to do with the export button.
This is hardest for people who know the product well. Expertise makes you fast at mapping a complaint onto a known backlog item, and that speed is what stops you hearing the part you didn't already have a ticket for. NN/g's guidance on leading questions covers the mechanics. The discipline that matters more is noticing when an answer confirms your prior, and slowing down at exactly that point. Confirmation is the default failure of this stage, the same pattern catalogued in confirmation bias in UX design.
Which methods belong in this stage
Empathize is a phase. The methods that fit inside it share one property: they surface the user's world.
Semi-structured interviews are the backbone — eight to ten open questions about past behavior, asked in whatever order the conversation wants, with follow-ups that chase specifics. Ask about the last time something happened, never about whether something would be useful. Contextual observation, watching someone do the work where they normally do it, costs more and is worth doing roughly once per problem space, because people omit the parts of their process they've stopped noticing: the spreadsheet on the second monitor, the colleague they Slack mid-task. Diary studies cover behavior spread across days, where nobody can reconstruct it accurately in a one-hour call. Onboarding, renewal, anything seasonal.
Then there's the evidence already sitting in the building. Support tickets, sales-call recordings, search logs, churn interviews someone else ran last year. It's the cheapest input available and it gets skipped, because reading two hundred tickets feels less like research than booking calls does.
Two things get scheduled into this stage constantly and shouldn't be. Usability testing evaluates an artifact, which is the test stage doing its job; run it during empathize and you re-anchor the whole team on the design you were trying to get distance from. Surveys measure how common a thing is, which is worth knowing once you know what to ask about. We don't schedule either one before the interviews. Focus groups sit in between: strong for language and for surfacing disagreement inside a team, thin on individual behavioral detail, because the loudest participant sets the frame in the first four minutes.
Empathy maps, and when they help
The empathy map — says, thinks, does, feels — is the artifact most associated with this stage, and the one most often used as a substitute for it. A room of people filling in the quadrants from memory hasn't empathized with anyone. They've formatted their assumptions into four boxes and given them the visual authority of research output. We don't count that session as research.
The map earns its place after the sessions, as a way to hold one participant's account in a single view and make the contradictions visible. The useful cell is almost always the gap between says and does: the person who says the approval step is critical and has never once used it. That gap is worth a finding, and it only shows up when the map is built from a transcript.
Personas get drafted here too, and the same thing decides whether they're worth anything: whether they came out of the sessions or out of the room. A persona assembled from imagination is a mission statement with a stock photo. NN/g's personas study guide is good on the difference. Both artifacts sit downstream of the interviews.
How much is enough
You're done when new sessions stop changing your account of the situation — when you can predict the next participant's answer and you're right. That arrives faster than most teams expect. Five to eight participants per distinct user type is the usual working range, and the five-users heuristic is a reasonable floor, with one caveat that gets dropped in the retelling: it counts per segment. Two segments, two sets.
Signs you stopped too early: every participant surprised you, which means saturation hasn't happened yet and you should keep going. Or you can restate the findings but not the disagreements between participants — if everyone in your sample agreed, you probably recruited one segment.
The sign you've gone too long is different in kind. The study is still running, and the team has started making the decision anyway, in Slack, off the first two calls.
Where teams get the empathize stage wrong
Running it after the solution is chosen. The most common version by a distance. The roadmap is set, research gets scheduled to de-risk it, and the questions can only come back confirming. If the study can't change what gets built, it's a rehearsal.
Treating it as a one-time phase. The stages are named in an order and drawn in a loop for a reason. Prototypes generate new questions about context, and those go back here. A team that empathizes once at kickoff is still working from that first picture nine months later, after the situation has moved.
Delegating it. When one researcher runs the sessions and reports back, the team gets the conclusions without the texture, and texture is what changes minds. A conclusion like "users don't trust the auto-import" persuades nobody who has an incentive to disagree. The people who sat in the session watched someone stop mid-flow and check the number by hand on a calculator.
Confusing thoroughness with rigor. Twelve sessions of the wrong question just gives you a bigger transcript. What turns volume into a claim you can defend is synthesis discipline. We run thematic analysis across the transcripts before anyone opens a deck.
FAQ
What is the empathize stage in design thinking? The first of the five stages, where you build an evidence-based account of the user's situation, motives, and constraints, deliberately without reference to your product. Its output is the raw material the define stage turns into a problem statement.
How long should it take? For a focused problem, one to two weeks including synthesis. Longer than that usually means the question is too broad to answer in a single study.
What's the difference between empathize and sympathize? Sympathy responds to the feeling and moves to a fix; empathy treats the feeling as a symptom and goes looking for the situation that produced it. In a session, sympathy sounds like agreement and empathy sounds like walk me through the last time that happened.
Can you empathize without talking to users? Partly. Support tickets, sales recordings, and search logs are real evidence and badly underused. They tell you what goes wrong and rarely why, and they never tell you what the person was trying to do before they got there. Use them to narrow the interviews.
Do you need an empathy map? No. It's one format for synthesis, useful because the says/does gap is easy to see in it. Built from memory instead of transcripts it actively hurts, because it makes assumptions look like findings.
Where to go next
This is one stage of a longer loop. The complete guide to design thinking covers how it hands off to define and why the sequence isn't a checklist. When a product's polish has drifted away from what its users understand, that gap is what the UX Clarity framework scores, and what a Full UX Audit is built to find.
Sources: NN/g — Leading Questions · NN/g — Personas Study Guide · NN/g — Thematic Analysis · NN/g — Why You Only Need to Test with 5 Users · NN/g — UX Research Cheat Sheet.
Not sure whether your last research week changed anyone's mind? Apply for a Full UX Audit →
Related
Information Architecture
Wireframing: From Lo-Fi to Hi-Fi (and What Each Wireframe Is For)
Fidelity isn't a quality ladder you climb. What a wireframe is supposed to settle, what lo-fi, mid-fi, and hi-fi each buy you, when a mood board is the right artifact instead, and the failure mode that costs teams a sprint.
TYPENORMLabs · 8 min · September 5, 2026
Information Architecture
WIRED UX Teardown: One Category Template, Three Different Jobs
A UX teardown of WIRED's category pages: the same template runs Business, Science and Reviews, and what each one puts in its first rail gives away what the section is actually for.
TYPENORMLabs · 5 min · August 31, 2026
Information Architecture
Webflow UX Teardown: $15, $25, $2,500, and a Footnote That Changes All Three
A UX teardown of Webflow's public pages: the homepage argues entirely in revenue, the product page won't let you self-serve, the marketplace shows no price at all. Then the pricing page hides its real variable in a two-word footnote.
TYPENORMLabs · 6 min · September 3, 2026