UX Research Methodologies: Your Methodology Is Not Your Method
A method is the thing you ran. A methodology is the chain of reasoning that makes its output count as evidence. What the three research methodology families actually commit you to, the five decisions every methodology makes, and where they fail quietly.
A team runs eight moderated sessions on a new onboarding flow. Clean script, good facilitation, careful notes. Seven of the eight sail through. The finding goes in the deck as onboarding is working, the flow ships, and activation doesn't move.
Nothing about the method was wrong. The sessions were competent usability tests. What was missing sat one level up: the participants came from the existing customer list, which meant everyone tested had already survived onboarding once. The study measured whether people who had completed onboarding could complete onboarding. It could not have returned any other answer.
That gap is what the word methodology is for: the distance between a method executed well and a study that could actually have come out the other way. In most product teams the word gets used as a longer synonym for "method," and the two failures that follow from the substitution are the ones in this article.
Method is what you did. Methodology is why it counts.
A method is a procedure. Moderated usability test. Diary study. Five-second test. Survey. Card sort. You can describe one in a sentence and someone else can run it.
A methodology is the reasoning that connects a procedure to a claim. It answers a harder question: given what we ran, on whom, analyzed how — what are we now entitled to say?
The distinction is easiest to see when the method is fixed and the methodology varies. Run the same eight usability sessions three ways:
| Same method, different methodology | What it can support |
|---|---|
| Recruited from current customers, findings written up after the sessions | Which parts of the flow confuse existing users |
| Recruited from the segment that abandons, sessions coded against a pre-written list of failure types | Which failure types dominate in the population that actually churns |
| Recruited from the abandoning segment, split across two prototype variants, coded blind | Whether variant B removes a failure type variant A produces |
Same script, same facilitator, same eight people's worth of effort. The first row supports a claim about people who already got through the flow; the second supports a claim about the people leaving. Only one of those is on the roadmap.
Which is why "what research methodology should we use?" is usually malformed. You don't pick a methodology off a shelf, you assemble one out of five decisions, and the research method is only the third of them.
The three research methodology families
The textbook split runs qualitative, quantitative, mixed. Most people meet it in Creswell's Research Design. It's a real distinction and it's worth reading as three sets of constraints you take on, since what each one forbids matters more day to day than what it permits.
Qualitative methodology is organized around mechanism: what this person was trying to do, what they believed the system would do, where those diverged. Sampling is purposive, so you recruit for the property you're studying. Analysis is interpretive, and the standard technique is thematic coding, whose canonical account is still Braun & Clarke (2006). What makes a theme valid is that it holds across the whole corpus. Sample size has nothing to do with it.
Quantitative methodology buys comparison and charges for it up front. You name the population before you look, specify the analysis before the data exists, and the study is valid to the degree the number would have come out differently if your hypothesis were false. That last clause does all the work, and it's what separates a well-run A/B test from a dashboard someone stared at until a story appeared.
Mixed methodology is the one worth spending words on, because in practice it usually means nothing at all.
What it means in most decks: a qual study and a quant study were both run, roughly in parallel, by different people, and stapled together at the readout. Two independent findings sit in adjacent sections, agree in tone, and neither was ever in a position to contradict the other. The presence of both is offered as the rigor.
What it has to mean to earn the name: one of them runs first and hands the other something specific. Qual→quant, sessions generate a list of candidate mechanisms and a survey estimates how widespread each one is; the survey items are written from the codes, so the survey could return "none of these" and that would be a result. Quant→qual, analytics locates the step that leaks and sessions recruited from that step explain why; the recruiting criterion comes out of the funnel query, so the sessions could fail to reproduce the drop-off and that would also be a result.
The direction is the design decision. Write it down and the second study is falsifiable by the first. Leave it unstated and the second study is decoration.
The trap in the taxonomy is treating the three as a menu. Picking qualitative doesn't just buy you fewer participants, it costs you the ability to say how many, and any percentage in your report is now in violation of your own methodology. Read the qualitative vs quantitative distinction as a question-selection problem. The methodology is where that choice becomes binding.
A second cut is often more useful in product work than the qual/quant one: attitudinal vs behavioral, what people say against what they do. NN/g's map of when to use which UX research method plots the common methods across both axes at once. The quadrants your team never uses are the interesting part of that chart.
The five decisions every methodology makes
Strip the vocabulary away and a research methodology is five commitments, made in order. Skip one and you don't get a worse study. You get a study that can't be wrong, which is the same thing as a study that can't be informative.
1. The question, stated so it could fail. Is onboarding good? is not a research question; no result contradicts it. Do first-time users complete account setup without external help? is. The test is whether you can describe, in one sentence, a finding that would make you abandon the belief.
2. The sampling frame. Who, drawn from where. Not how many. This is where the eight-session study at the top of this article died. Write the frame as a sentence with an exclusion in it: "users who created an account in the last 14 days and did not complete setup, excluding anyone who has contacted support." The exclusions are the part that does work. If you can't state your frame in one sentence, you inherited it from whichever list was easiest to recruit from.
3. The instrument. The script, the task, the survey items, the event you're logging. Instruments carry assumptions the way a lens carries distortion. A Likert scale with no neutral option forces a direction. A task written as "buy a pair of running shoes" tests something different from "find out whether this store sells your size." For quantitative work this is also where you name the dependent variable precisely. Every measure has a degenerate way to improve (time-on-task drops when people give up), and a vaguely named measure is the one that gets optimized that way.
4. The analysis plan. How raw material becomes a finding. For qualitative work: who codes, against what scheme, and what counts as a theme. For quantitative: which comparison, which threshold, what happens to outliers. Write it before collection.
5. The decision rule. What each possible outcome causes you to do. "If completion is under 60% we rebuild the step; if it's over 80% we move on; between those we run a second round on the specific failure." Research that gets called inconclusive was often conclusive and simply had no rule attached, so every possible result mapped to "interesting, let's discuss."
Only decision 3 is what most people mean by "the method." The other four are the methodology, and they're the ones that get skipped under deadline.
Write the methodology down before the data arrives
The five decisions have to be fixed before collection, and the reason isn't bureaucratic tidiness. It's that a study with unfixed decisions has far more ways to produce a positive result than anyone intuits.
Simmons, Nelson and Simonsohn demonstrated this with a deliberately absurd experiment. Ordinary flexibility, each instance individually defensible (dropping a condition, adding participants until it works, choosing among measures after seeing them), was enough to produce statistically significant evidence for a false claim at will (False-Positive Psychology, 2011). Their term for that flexibility is researcher degrees of freedom. Product research has the same ones and one fewer check on them, since nobody is peer-reviewing your funnel query.
The companion failure has a name too: HARKing, hypothesizing after the results are known, which means presenting something you found in the data as though it were what you set out to test (Kerr, 1998). It is almost never dishonest in intent. It's what happens when a study has no written question and the report gets written from whatever was most interesting in the data.
The fix is preregistration: write the five decisions down, timestamp them somewhere you can't quietly edit, then collect. Nosek and colleagues make the case for it as a general research reform, and the lightweight tooling built for academia — AsPredicted, the OSF preregistration templates — works fine for product work. So does a dated document in your own repo, provided nobody can rewrite it after the readout.
This is the procedural counter to confirmation bias, which does its worst work in exactly this window. When a result matches what the team expected, it's accepted and the analysis stops. When it doesn't, the instrumentation gets audited. Both responses feel like diligence, and only a pre-written rule makes them symmetric.
Exploratory work is still legitimate. Most good product research starts there. The requirement is that you label it. A study that generated a hypothesis and a study that tested one are both useful; a study that did the first while claiming the second is how a team ships confidently into an activation number that never moves.
A worked methodology example
Concrete beats abstract. Here is the same investigation written twice.
As a method: "We'll run some user interviews about checkout."
As a methodology:
Question. Do returning customers abandon checkout because of the shipping cost itself, or because of when it appears?
Frame. Customers with ≥2 prior orders who reached the shipping step and did not complete, in the last 30 days. Excluding orders over $500 (different decision process) and anyone who has contacted support about shipping (already primed).
Design. Quant→qual, sequenced. First, funnel analysis segmenting abandonment by whether the cart total crossed the free-shipping threshold. Then eight moderated sessions recruited from the larger of the two segments, replaying their own abandoned cart.
Instrument. Sessions use the participant's real cart. Facilitator does not mention shipping until the participant does; timestamp when they raise it. Coding scheme fixed in advance: cost too high, cost unexpected at this point, cost unclear earlier, unrelated.
Analysis. Two coders, independently, on the four codes. Disagreements resolved by a third pass, not by discussion. A code counts as dominant if it's the primary code in ≥5 of 8 sessions.
Decision. Cost unexpected dominant → move the estimate to the cart page and re-measure. Cost too high dominant → this is a pricing problem, not a UX one; hand it to commercial and stop. Neither dominant → the segmentation was wrong; return to the funnel data before running more sessions.
The second version is about 200 words longer. Look at the third branch of the decision rule in particular: writing down in advance what a null result means is what keeps a team from re-running the same eight sessions with a slightly different script.
One property of this format is worth noticing. The sessions could be focus groups, diary entries, or unmoderated tests, and four of the five sections wouldn't change a word. That's the tell that what you wrote is a methodology.
Where methodologies fail quietly
Methods fail loudly: a broken prototype, a no-show participant, a survey with a typo. Methodologies fail silently and produce a clean-looking report.
The sampling frame drifted. Recruiting got hard, the panel was substituted, the criteria loosened by one clause. Frame drift is invisible by construction, because the report describes the frame you intended rather than the one you recruited. When the documented decisions and the actual study diverge, the document has to change, visibly, before the readout.
The construct isn't the thing. You measured task completion and claimed the design is clear. You measured satisfaction scores and claimed usability. Every leap from a measure to a concept is a separate claim needing its own defense. ISO 9241-11 defines usability as effectiveness, efficiency and satisfaction for a specified user, goal and context, and that string of qualifiers is pedantic on purpose. The specification is what stops the construct from floating.
The sample carried an assumption you didn't make on purpose. Faulkner's re-sampling study is the standard illustration: testing five users catches about 85% of usability problems on average, but her worst random draw of five caught 55% (Faulkner, 2003). The five-user rule comes out of a model that assumes you're iterating and testing again (Nielsen & Landauer, 1993). Import the number without the iteration and you've adopted an assumption you never evaluated.
The reasoning direction was never chosen. Generalizing from observations and testing a stated prediction are different operations with different failure modes. A study that slides between them mid-analysis gets the protections of neither. Being explicit about inductive versus deductive reasoning at planning time costs a paragraph and prevents a whole category of unfixable report.
The variables were entangled. You changed the copy and the layout together, the number moved, and now you own a win nobody can attribute. Keeping independent and dependent variables separated is the cheapest methodological hygiene available, and it goes first under release pressure.
Writing the methodology section
If your research output goes to anyone who wasn't in the room, it needs a methodology section, and it should be short. Five headings, matching the five decisions: Question · Frame · Design · Instrument · Analysis, plus the decision rule if one existed. Half a page.
Three rules for it:
Write it before the study. The section starts as a plan and becomes a record. Drafting it up front is what surfaces a missing sampling frame while fixing it is still free.
Include what you excluded. Who you didn't recruit, what data you dropped, which sessions didn't count. Exclusions are where most of a study's hidden assumptions live, and a reader can only evaluate the ones you name.
Say what the study can't support. One sentence: "this tells us which failure modes exist in the abandoning segment; it does not tell us their relative frequency." Readers supply their own boundary otherwise, and theirs will be more generous than yours.
This is also the section that makes a study reusable six months later, when someone asks whether the old finding still applies to the redesigned flow. Without it the honest answer is always no.
Frequently asked questions
What is a research methodology? The reasoning that connects a procedure to a claim: the question, the sampling frame, the design, the instrument, the analysis plan, and the decision rule. A methodology is what makes a method's output count as evidence. Without one you have an anecdote with a sample size attached.
What's the difference between a method and a methodology? A method is a procedure you can run — a usability test, a survey, a card sort. A methodology is the set of decisions around it that determines what its results are entitled to say. Two teams can run the identical method and end up with claims of completely different strength.
What are the main types of research methodology? Qualitative (mechanism and meaning, purposive sampling, interpretive analysis), quantitative (measurement and comparison, defined population, pre-specified analysis), and mixed (the two in a stated sequence, each handing something specific to the other). What separates them is what you're allowed to conclude. Whether the data is words or numbers is downstream of that.
Which research methodology should I use? Start from the decision the study is meant to unblock, then work backwards. If the decision needs a mechanism — why is this happening — that's qualitative. If it needs a magnitude — how many, how much, is it bigger than noise — that's quantitative. If you can't name the decision, the methodology question is premature. The research methods overview maps the common methods against the questions they answer.
How many participants does a methodology require? The methodology determines it. Qualitative work runs until new sessions stop producing new mechanisms. Quantitative work needs enough participants to detect the smallest difference that would change your decision, a threshold you can set before collecting anything. Pick the count before the question and you've set a budget.
Do I need a formal methodology for a small product study? The five decisions take about twenty minutes to write for a study that takes two weeks to run, and skipping them is what produces the confident, unusable report. Formality isn't the point. Being written down before collection is. If the study only informs a reversible decision, keep it to a paragraph, and keep the paragraph.
Research in product work usually fails before anyone runs it, at the point where nobody decided in advance and in writing what the results would be allowed to mean. That decision is the methodology. It costs twenty minutes and it is unrecoverable afterward. If you'd rather have an expert pass over an existing product than stand up a study, a full UX audit is the summative version of the same discipline: fixed criteria, written before the review, applied to what you already shipped.
Related
Information Architecture
Wireframing: From Lo-Fi to Hi-Fi (and What Each Wireframe Is For)
Fidelity isn't a quality ladder you climb. What a wireframe is supposed to settle, what lo-fi, mid-fi, and hi-fi each buy you, when a mood board is the right artifact instead, and the failure mode that costs teams a sprint.
TYPENORMLabs · 8 min · September 5, 2026
Information Architecture
WIRED UX Teardown: One Category Template, Three Different Jobs
A UX teardown of WIRED's category pages: the same template runs Business, Science and Reviews, and what each one puts in its first rail gives away what the section is actually for.
TYPENORMLabs · 5 min · August 31, 2026
Information Architecture
Webflow UX Teardown: $15, $25, $2,500, and a Footnote That Changes All Three
A UX teardown of Webflow's public pages: the homepage argues entirely in revenue, the product page won't let you self-serve, the marketplace shows no price at all. Then the pricing page hides its real variable in a two-word footnote.
TYPENORMLabs · 6 min · September 3, 2026