How to Run User Interviews Without Getting the Answer You Wanted
User interviews fail quietly: the transcript is full, everyone was polite, and the finding was already in the room before the first session. How to write the script, moderate against your own bias, and turn eight conversations into a claim you can defend.
Here is a shape that repeats often enough to plan around. A PM books six calls to find out whether teams want a shared dashboard. Each call ends with the same question: would a shared dashboard be useful to your team? Five say yes, one says probably, the deck reads 83% of interviewed customers want a shared dashboard, and a quarter gets planned around it. Nine months later four people open the thing each week.
Nothing went wrong in the room. The calls were friendly, the notes were thorough, the participants were real customers. What went wrong is that the question could not have returned a no. "Would this be useful?" asks a person to imagine a future version of themselves, in a hypothetical week, with time they don't have, and then to say it out loud to the face of the person building it. Everyone in that study was being helpful. Helpfulness is the noise floor of this method.
User interviews are the highest-return research you can run and the easiest to run into a wall, because the failure is invisible from inside. A bad usability test looks bad. The participant is stuck, the screen isn't moving, everyone in the observation room can see it. A bad interview looks like a great conversation.
What the method actually answers
An interview is a retrospective, attitudinal method. You are asking a person to narrate their own past and interpret it for you. Done well, that gives you three things almost nothing else does.
The shape of the problem in their words. Their vocabulary, not your feature names. This is what makes interviews the best available input to naming things, writing copy, and choosing which problem is worth solving at all.
Sequence and context. What happened before the moment your product enters the story. Who else was involved. What they tried first. Products fail in the seam before the first click far more often than inside the flow.
Motive and constraint. Why they kept doing the workaround. What they'd lose by switching. Why the obvious fix isn't obvious to them.
The list of things interviews cannot give you is longer, and the length is the point: this method is narrow, and most bad studies come from stretching it.
- Whether people can use your interface. That's a usability test. Asking "was that screen clear?" gets you a review of your feelings.
- How common the pattern is. Eight people can tell you a problem exists and is severe. They cannot tell you it's 40% of the base. That's a survey question, and it comes second.
- What they will do next. Stated intention is the weakest signal in research. People aren't lying when they predict their own behavior. They're just bad at it, under conditions they haven't met yet.
- What they'd pay. Same failure, higher stakes. Price sensitivity stated in a friendly call has almost no relationship to price sensitivity at a checkout.
The whole discipline of the method is keeping your questions inside the first list. NN/g's Interviewing Users draws the same line: interviews are for understanding, not for validation.
Decide the question before you decide the method
Most weak studies begin with "let's talk to some users," which is a decision to spend twelve hours without a decision attached to it.
Before you recruit anyone, write one sentence in this shape:
We are running user interviews because we need to decide X, and right now we're guessing about Y.
Then check two things. First: could the answer change what we do? If every plausible finding leads to the same roadmap, you're running research as a ritual. Ship instead. Second: is Y a why or a how many? A how-many question handed to interviews will come back looking answered, and you'll read six anecdotes as a prevalence estimate. That split is the qualitative/quantitative line on the UX research methods hub, and getting it wrong is the most expensive mistake in this area.
Write the decision rule down before the first call. If four or more of eight describe doing the workaround manually every week, we build it. If it's two or fewer, we don't. A rule written in advance is the only real defense against a study that agrees with whoever talks loudest in the readout.
Recruit the people who are hard to get on a call
The most common recruiting failure is interviewing whoever answers the phone. Your power users answer the phone. They also survived every rough edge you're trying to find.
Recruit against the axis your question turns on. If the question is about adoption, you need people who didn't adopt: churned accounts, trial abandoners, the team that bought a license and never rolled it out. If it's about a workflow, you need the range of roles inside it, including the person who receives the output rather than only the one who produces it. If it's about switching, you need someone who chose a competitor.
Six well-chosen strangers beat twenty friendly customers.
Screen with a behavioral question. "How many reports did you send last month?" tells you something. "Do you do a lot of reporting?" tells you how they'd like to be described.
The script: three layers, thirty minutes
A good interview script is short. Ten to twelve questions, most of which you won't ask, because the useful material comes from following what they said.
Layer one, warm-up and context (5 min). Their role, who they work with, what a normal week looks like. This isn't filler. It's the map you'll need to interpret everything after it, and it teaches them that the session is about their work rather than about your product.
Layer two, the last time (15 min). The core of the method, and it's one move: tell me about the last time you did X. Not "how do you usually." Not "what would you do if." One actual instance, walked through in order.
"Walk me through the last time you had to pull numbers for the Monday review. Start from when you found out you needed them."
Specific past episodes are grounded in memory. Generalizations are grounded in self-image. When someone slips into the general — "usually I'd just export it" — pull them back: and last Monday, what did you actually do? Half the time the answer is different, and that gap is the finding.
Layer three, the seams (10 min). Now probe where it broke. What did you try first. Who did you ask. What happened when it didn't work. What did you do instead. This is where workarounds live, and a workaround is the most valuable object in a transcript. It's a user who has already designed your feature for you, badly, at their own cost.
Leave the last two minutes for the only future-facing question worth asking: is there anything I should have asked you about?
The questions that poison the data
Four question shapes ruin user interviews, and three of them sound professional.
The hypothetical. "Would you use a shared dashboard?" You are collecting imagination. Ask about history: "when did you last need to show these numbers to someone else, and what did you do?"
The leading question. "How frustrating was the export step?" You've told them the answer's shape and they will fill it in politely. The neutral form is "walk me through the export step." If it was frustrating, that shows up on its own, in their word rather than yours.
The feature request. "What would you like us to build?" Users are experts in their problem and amateurs in your solution space. The answer will be a faster horse and you'll feel obligated to it once it's in the transcript. Take the problem they described and throw away the solution they proposed.
The compound question. "How do you handle reporting, and is that different for the quarterly stuff?" They'll answer the second half and you'll file it under the first.
One instinct fixes all four: ask about behavior, in the past, in a single clause.
Moderating: the work is the silence
The script is maybe a third of the outcome.
Wait longer than is comfortable. Three seconds of silence after an answer does more work than anything else in the method. Most people finish their polite answer, then, if nobody rushes in, finish the real one. New interviewers fill the gap because silence feels like failure.
Ladder. When something interesting surfaces, go down: "why did you do it that way?" Then again. Then once more. Three whys gets you from I export to a spreadsheet to because finance won't accept anything that isn't a spreadsheet and I've been burned before. The third answer is the one that changes a roadmap.
Echo, don't paraphrase. Repeating their last few words back — "…so it was already too late?" — reopens a thread without steering it. Paraphrasing swaps their vocabulary for yours, which quietly destroys the main thing you came for.
Never defend the product. The instant you explain that the export button is actually right there, the session is over. They now know the correct answers and will supply them for the remaining twenty minutes. If they're wrong about your own product, write it down and move on.
Don't take the notes yourself if you can avoid it. A second person on notes lets you keep eye contact and chase tangents. If you're solo, record with consent and note only timestamps and follow-ups.
How many user interviews is enough
The honest answer is: until new sessions stop surprising you.
The practical answer, for one reasonably homogeneous group, is five to eight. Nielsen's five-user finding is about usability testing rather than interviews, but the underlying curve carries. The first few sessions surface most of what a homogeneous group has to say, and each session after that returns less.
Two adjustments matter more than the number. Count per segment: if enterprise admins and solo users are different populations, five each. Eight sessions split across four unrelated segments is four studies of two people. And watch the curve rather than the target. If session six still produces something nobody had heard, the population is more varied than you assumed, which is itself worth reporting.
Needing twenty interviews to feel confident usually means the question was a how-many question wearing a why costume. Run six to learn the language, then size it with a survey.
Moderated, unmoderated, remote
Remote video is the default now and costs less than people assume. You lose peripheral context, meaning what's on their other monitor and who interrupts them. Almost everything else survives. Two things are worth protecting: ask them to share their screen while they narrate the last-time story, because people point at things they would never think to describe, and keep it one-to-one.
Unmoderated tools don't really do interviews. They do prompted monologues, which are useful for volume and useless for laddering, since the value of the method sits in the follow-up question you didn't plan. And an interview with several participants at once isn't an interview. It's a focus group, a different method with a different failure mode, where the most confident person in the room sets the answer for everyone else.
From transcripts to a claim
This is the step teams skip. Eight recordings in a folder is raw material, and what usually happens to raw material is that the deck quotes whichever line the loudest reader remembered.
A workable synthesis pass:
- Extract observations. One line per thing that actually happened or was actually said, tagged with who said it. "P4 rebuilds the report manually every Monday," not "P4 finds reporting painful."
- Cluster bottom-up. Group the lines by what they have in common and let the group names come last. Naming the buckets first means sorting the evidence into the answer you walked in with.
- Count, and label the count honestly. "6 of 8" is fair. "75% of users" is not: eight people are not a percentage of anything. Prevalence needs a scaled instrument.
- Separate finding from recommendation. The finding is what you saw. The recommendation is your judgment about it. Keeping them in different columns lets someone argue with your judgment without attacking your evidence, and lets you notice when four recommendations are resting on one participant.
- Write the disconfirming line. One sentence: what did we hear that argues against what we're about to do?
Personas and journey maps are downstream artifacts of this pass. The persona and empathy map templates are worth filling in once the observations exist to fill them with, and not before.
When interviews are the wrong tool
Anything that needs a number belongs somewhere else. Sizing, prioritization, before-and-after: surveys and analytics. If the question is whether someone can complete a task, run a usability test, because that question asked in an interview gets you a compliment instead of a completion rate. Habits are the hardest case, and not because people hide them. Nobody can narrate a default they've stopped noticing, so the export ritual they've done every Monday for two years comes back as "I just pull the numbers," with the four minutes of manual cleanup missing. Observe or instrument.
The fourth case isn't methodological. If the decision is already made, research becomes a procurement exercise: it returns exactly the evidence requested, at full cost, with the added expense of everyone believing it afterward.
The line nobody writes
Step five above is the one worth defending when the schedule gets tight, and it's the first thing cut.
Every study produces something inconvenient. A participant who solved the problem without you. A workaround that was faster than the feature. A segment that didn't recognize the problem at all. That material doesn't fit the narrative arc of a readout, so it drifts into the appendix, then into nobody's memory, and the report ships as a clean story about a problem you already believed in.
Writing one sentence — here is what we heard that argues against this — is what separates a study from a briefing. It's also the cheapest credibility you will ever buy in a roadmap review, because the person across the table is already thinking it, and hearing it from you first changes what the rest of your evidence is worth.
If there's nothing to write in that line, you didn't listen. You collected.
FAQ
How many user interviews do I need?
Five to eight per distinct segment, and stop when new sessions stop surprising you. Fewer than five and one unusual participant dominates. More than eight in a homogeneous group returns very little. If what you need is a percentage rather than a pattern, no number of interviews will get you there.
What questions should I ask in a user interview?
Ask about specific past behavior in single clauses: "walk me through the last time you did X," "what did you try first," "what happened then," "why that way?" Avoid anything hypothetical ("would you use…"), anything leading ("how frustrating was…"), and anything that asks them to design ("what should we build?").
How long should a user interview be?
Thirty to forty-five minutes for most product questions. Under twenty and you never get past the polite answers. Over an hour and both of you are tired, and the last section is worth less than the time it cost. Book the full hour and plan to finish early.
What's the difference between user interviews and usability testing?
Interviews ask people to narrate and explain their own past behavior. Usability tests watch them attempt a task in the present. Interviews answer why, and in what context. Usability tests answer whether they can, and where they get stuck. Asking interview questions mid-test ("was that clear?") contaminates both.
Should I interview users before or after building the prototype?
Before, for the problem: interviews are the cheapest way to find out you're solving the wrong thing. Once a prototype exists, switch methods. Put it in front of them and watch. An opinion about a prototype is worth much less than a recording of someone trying to use it.
Can I run user interviews with existing customers only?
You can, as long as you know what it costs. Existing customers are survivors of every barrier you're trying to see. Mix in people who churned, trialled and left, or chose a competitor. If that's impossible, say so in the write-up. A study that names its own sampling limit is still defensible; the damage comes from the limit nobody mentioned.
Related
Information Architecture
Wireframing: From Lo-Fi to Hi-Fi (and What Each Wireframe Is For)
Fidelity isn't a quality ladder you climb. What a wireframe is supposed to settle, what lo-fi, mid-fi, and hi-fi each buy you, when a mood board is the right artifact instead, and the failure mode that costs teams a sprint.
TYPENORMLabs · 8 min · September 5, 2026
Information Architecture
WIRED UX Teardown: One Category Template, Three Different Jobs
A UX teardown of WIRED's category pages: the same template runs Business, Science and Reviews, and what each one puts in its first rail gives away what the section is actually for.
TYPENORMLabs · 5 min · August 31, 2026
Information Architecture
Webflow UX Teardown: $15, $25, $2,500, and a Footnote That Changes All Three
A UX teardown of Webflow's public pages: the homepage argues entirely in revenue, the product page won't let you self-serve, the marketplace shows no price at all. Then the pricing page hides its real variable in a two-word footnote.
TYPENORMLabs · 6 min · September 3, 2026