The Likert Scale: Design, Examples, and How to Read the Data
A Likert scale is cheap to write and easy to get wrong. How many points to use, whether to keep the neutral midpoint, why agree/disagree is the weakest format available, and what the numbers can honestly support.
Rensis Likert's 1932 monograph was a cost-cutting exercise. The dominant way to measure an attitude at the time was Thurstone's method of equal-appearing intervals, which required assembling a panel of judges to sort statements and assign each one a scale value before you could ask a single respondent anything. Likert's proposal, in A Technique for the Measurement of Attitudes (Archives of Psychology, No. 140), was to skip the judges: write a set of statements, let people mark their agreement on a five-point response set, and add the marks up.
It worked well enough that ninety years later it is the default response format for most of the survey questions anyone in product ever writes. It is also the format most likely to be written in four minutes between meetings, which is where the trouble starts.
Every parameter in a Likert scale bends the distribution you get back: the number of points, the presence of a midpoint, the labels, the agree/disagree framing itself. This is what each of those decisions buys, and what the resulting numbers can honestly support.
An item is not a scale
The vocabulary confuses this from the start, and the confusion has consequences at analysis time.
A Likert item is one statement plus its response options: "The checkout process was confusing" · Strongly disagree → Strongly agree. A Likert scale, in Likert's original sense, is several such items measuring the same underlying construct, whose responses get summed or averaged into one score. That summing is the entire point of the technique. Likert's method is a summated ratings method; the single item was never meant to carry a construct on its own.
Almost everything written in product work is a Likert item, called a scale. That's usually fine, as long as you don't then borrow the analytical privileges that belong only to the summated version. Hold the distinction now; it decides, further down, whether reporting a mean is defensible.
How many points should a Likert scale have?
The 5-point and 7-point versions dominate for reasons that turn out to be better than "everyone does it."
Preston and Colman tested rating scales with 2 through 11 response categories on the same task and found that reliability and validity were clearly worst at 2, 3, and 4 categories, rose sharply, and then flattened out from roughly 7 categories upward (Preston & Colman, 2000). Respondents also reported preferring the longer scales while completing the shorter ones faster. More points means more information per answer and more time spent per answer.
The practical read:
- 5-point is the safe default for a survey where people will answer many items in a row. It fits on a phone, every point can be labeled in words, and fatigue accumulates slowly.
- 7-point buys real discriminating power when the item is the measurement — a single post-task ease question, a benchmark you intend to track across releases. The Single Ease Question, asked immediately after a task in usability testing, is 7-point for exactly this reason.
- 10 or 11 points belongs to instruments designed around it, like NPS. Building your own eleven-point item usually means the extra granularity is imaginary: nobody can reliably distinguish their own 6 from their own 7 on a construct like "confusing."
- 4-point is a forced-choice decision, not a precision decision. See the midpoint question below.
One constraint outranks the count. If you're comparing results over time, the scale is frozen: changing a 5-point Likert scale to a 7-point one mid-programme ends the time series, and rescaling the old data to match is arithmetic dressed as continuity.
Should you keep the neutral midpoint?
Removing the midpoint is often framed as forcing people off the fence. What it actually does is remove the only honest option for three different respondents: the one with a genuinely balanced view, the one with no opinion, and the one who doesn't understand the item. Drop the midpoint and all three get pushed into a direction they don't hold, and you can no longer tell them apart from people who do.
Keep the midpoint by default, and label it for what you mean. "Neither agree nor disagree" is a position. "Don't know" is an absence of one. They are different answers and collapsing them into one button loses the distinction permanently. If both are plausible for your item, offer the midpoint on the scale and a separate, visually detached Not applicable, kept off the scale so it can't be read as a scale position.
Drop the midpoint only when a forced direction is the actual decision you're making — a go/no-go gate, a preference test between two designs — and accept that the resulting split is a choice you manufactured, not a preference you found.
Label every point you show
A scale with labeled endpoints and bare numbers in between asks each respondent to invent the meaning of 3. They invent different meanings, and the variance that produces is indistinguishable from real disagreement.
Fully verbal labels are the stronger default: Strongly disagree · Disagree · Neither agree nor disagree · Agree · Strongly agree. Krosnick's review of the survey methodology literature comes down on the side of labeling every point (Krosnick, 1999). The reasoning matches what makes a labeled axis beat a bare one on a chart. All of the respondent's effort should be going into the judgment; decoding the instrument is overhead.
Two things follow. Keep the spacing between adjacent labels psychologically even — "Excellent · Good · Fair · Poor · Terrible" is lopsided, with three positive-ish steps and a cliff at the end. And keep the direction consistent for the whole survey; a single item that runs the other way to catch inattentive respondents catches attentive ones too.
Agree/disagree is the weakest format available
This is the least intuitive part, and the one worth doing first.
"The navigation was easy to use — Strongly disagree → Strongly agree" invites acquiescence bias: a measurable tendency to agree with a statement regardless of its content, stronger under fatigue, time pressure, or any suggestion that agreement is the cooperative answer. Krosnick's account of survey satisficing describes the mechanism — respondents facing cognitive load stop optimizing and start taking the first defensible route through the question, and "yes" is almost always the cheapest route available (Krosnick, 1991).
The fix is to move the content out of the statement and into the response options. Instead of asking someone to agree with "the navigation was easy," ask:
How easy or difficult was it to find what you were looking for? Very difficult · Difficult · Neither easy nor difficult · Easy · Very easy
The respondent still marks one of five points, so it still looks like a Likert scale and still analyzes like one. But there is no proposition to nod along to, and both directions are named in the question stem, so neither reads as the expected answer.
Reverse-worded items are the traditional patch for acquiescence — mix in some negatives so agreement bias cancels out. It half works. It also generates its own failure mode, where respondents who miss the negation produce responses that look like real disagreement, and the item behaves oddly in the resulting factor structure. Item-specific response options are the better instrument; keep reverse-wording for cases where you're stuck with an established scale that already uses it.
Likert scale examples that do specific jobs
Some standard instruments are worth borrowing whole, because their wording has been tested at a scale your survey never will be.
| Instrument | Format | What it measures | When to use it |
|---|---|---|---|
| SUS (System Usability Scale) | 10 items, 5-point agree/disagree, alternating tone | Perceived usability of a whole system, scored 0–100 | End of a session, or as a release-over-release benchmark |
| SEQ (Single Ease Question) | 1 item, 7-point | How difficult a specific task felt | Immediately after each task in a usability test |
| CSAT | 1 item, 5-point satisfaction | Satisfaction with one interaction | After a support contact or a completed job |
| Custom item-specific set | 3–5 items, 5-point, no agree/disagree | One narrow construct you actually care about | When no standard instrument covers your question |
SUS is the interesting case, because it violates the advice above and survives. It is a pure agree/disagree instrument, and its alternating positive/negative items are precisely the reverse-wording pattern just described as fragile. It works anyway because the item set never changes and the score is read against a large accumulated body of other studies. That benchmark is what makes a raw 68 legible: it sits around the middle of everything ever measured with the instrument, so a score is a percentile in disguise. Later analyses suggest an all-positive variant performs comparably (MeasuringU on rating scale practice), which supports the general point. The alternation is not what makes SUS work.
The rule that falls out: borrow a standard instrument whole, or write item-specific questions of your own. What fails is a home-made agree/disagree battery with no benchmark to compare against. If you want a starting point, our survey question bank collects open, closed, Likert, and NPS questions by answer type, and the UX survey template is a ready 1–5 Likert set worded to avoid the traps above.
How to analyze Likert scale data
Here is where the item/scale distinction gets paid.
A single Likert item is ordinal. The options are ordered, but nothing guarantees the psychological distance from "disagree" to "neither" equals the distance from "agree" to "strongly agree." Averaging ordinal codes produces a number, and the number will have decimals, and the decimals will not mean what the decimals in a load-time measurement mean. For a single item, report the frequency distribution first, and the median or mode alongside it. A count against a denominator in both directions — how many landed in the top two boxes, how many in the bottom — survives a challenge. "Mean 3.8" does not. A 3.8 is consistent with wildly different distributions, including a bimodal one where half your users are delighted and half are stuck. That split is the most actionable pattern a survey can surface and the one a mean is guaranteed to hide.
A summated multi-item scale is treated as interval by convention, and that convention is defensible when the items genuinely measure one construct. Check that. Internal consistency (Cronbach's alpha, or better, the item–total correlations underneath it) tells you whether the items are pulling together. If they aren't, the sum is a mixture of two things and its mean is a number about nothing.
What keeps the reporting honest:
- Show the distribution. A stacked bar of the five categories communicates more than any summary statistic, and it makes bimodality visible instantly.
- Report n next to every percentage. "62%" from 13 respondents is four opinions and a rounding artifact.
- Decide the comparison before you look. If a Likert scale is your dependent variable, the threshold that counts as a real change belongs on paper before the data arrives. It's the same discipline that keeps an A/B test honest.
Statistical significance on a Likert item deserves one caution. With a large enough sample, a shift of 0.1 on a five-point scale will reach significance and mean nothing anyone can act on. The question is never whether the difference is detectable; it's whether the difference is large enough to justify the work it implies.
The biases that come with the format
Worth naming, because each one has a specific countermeasure.
Acquiescence is the one that costs the most and is hardest to spot afterward: a tendency to agree with a statement independent of its content, which turns a battery of positively-worded items into a machine for manufacturing approval. It is also the only bias on this list you can design out completely, by moving the content into the response options as described above. Everything else here you manage.
The rest, briefly. Central tendency — avoiding the extremes — is countered by labeling all points, so the ends read as usable positions. Extreme responding is the opposite habit and varies systematically across populations, which is why cross-country comparisons of absolute scores mislead and within-group trends don't. Straightlining, marking the same column down a whole grid, is one of the satisficing shortcuts Krosnick documents under cognitive load; it responds to shorter batteries, and better items don't fix it. Social desirability is countered by asking about behavior in the past tense and by never signposting which direction the sponsor prefers.
All five get worse as the survey gets longer.
What to change on the next survey
Three edits, in order of return:
- Rewrite every agree/disagree item as an item-specific question. A mechanical pass over the draft that removes the largest known bias in the format.
- Label every response point in words, and check that the steps are evenly spaced in meaning rather than just evenly spaced on screen.
- Replace the mean in your reporting template with a distribution plus n. One change to one artifact, and it survives everyone forgetting why.
None of this makes a survey a substitute for watching someone use the thing. A Likert scale measures what people report about an experience, which is a different quantity from what happened during it — the attitudinal side of the research methods map, not the behavioral one. Its best use is to establish how widely a pattern holds after other work has told you what the pattern is.
Frequently asked questions
What is a Likert scale?
A statement plus an ordered set of agreement options, classically five. In Likert's 1932 original, several such items measuring one construct are summed into a single score.
Is a 5-point or 7-point Likert scale better?
Both are defensible; 7-point extracts slightly more information per item, and the gain flattens out past that. Use 5-point when respondents face many items in sequence or will answer on a phone. Use 7-point for a single item that carries real measurement weight, such as a post-task ease question or a tracked benchmark. Whichever you pick, freeze it — changing the number of points breaks comparability with everything you've already collected.
Should a Likert scale have a neutral midpoint?
Usually yes. Removing it doesn't eliminate ambivalence, it just relabels ambivalent, uninformed, and confused respondents as leaning in a direction they don't hold. Keep the midpoint, and offer a separate, off-scale "not applicable" option if the item might not apply.
Can you calculate a mean of Likert scale data?
This is the contested one, and the honest answer has two halves.
You can always compute a mean; the question is whether the arithmetic corresponds to anything. A single Likert item is ordinal. The options have an order but no guaranteed unit, so the gap between "disagree" and "neither" may be nothing like the gap between "agree" and "strongly agree," and averaging codes that sit at unequal spacings produces a figure with no interpretable step size. That's why a single item should be reported as a distribution with a median, not as a mean to one decimal place.
A summated multi-item scale is different in practice if not in theory. The purists' position is that summing ordinals doesn't manufacture an interval scale. The applied position, and the one nearly all published survey work follows, is that a sum of several items measuring one construct behaves well enough that parametric methods on it hold up. The condition attached is the part usually skipped: the items have to demonstrably measure one thing, which is what internal consistency checks are for. Run that check and the mean is standard practice. Skip it and you're averaging a mixture.
Are Likert scale data ordinal or interval?
Ordinal at the item level, treated as interval at the summated-scale level by convention. See above — the distinction decides which statistics are available to you.
What's the difference between a Likert scale and a semantic differential?
A Likert item presents a statement and measures agreement with it. A semantic differential presents a pair of opposed adjectives — confusing / clear — and asks the respondent to place the object between them. Because there's no proposition to agree with, the semantic differential sidesteps acquiescence bias by construction, which is the main reason to reach for it. NN/g compares the two formats directly in its guide to rating scales in UX research.
How many Likert items should a survey have?
Fewer than the draft contains. For each item, write down the decision that would change based on the answer; the ones that fail that test are the ones to cut.
Do Likert scales work for measuring usability?
For perceived usability, yes — that's what SUS and the SEQ measure. They do not tell you whether anyone completed the task, how long it took, or where they got stuck; those are behavioral measures that require observing the usability test itself. When a participant fails a task and then rates it a 5 for ease, the gap between the two measures is usually the finding.
Take it further
A Likert scale tells you how widely a problem is felt. Locating where in the interface the problem lives is an expert-review question, and it's the one the UX Clarity framework is built to answer. A Full UX Audit applies it end to end. For choosing between survey work and the other options in the first place, start from the research methods map, or the CSAT template if you already know satisfaction is the thing you need to track.
Sources: Preston & Colman, 2000 — Optimal number of response categories in rating scales · Krosnick, 1999 — Survey Research, Annual Review of Psychology · Krosnick, 1991 — Response strategies for coping with the cognitive demands of attitude measures in surveys · NN/g — Rating Scales in UX Research · MeasuringU — Rating Scale Best Practices · NN/g — Usability Metrics.
Sitting on survey results that everyone reads differently? Apply for a Full UX Audit →
Related
Navigation Design
Zara UX Teardown: The Homepage That Doesn't Scroll
A UX teardown of Zara's public store: a homepage one screen tall, navigation reduced to grey hairlines, and a catalog that won't quote a price until you type into the search box.
TYPENORMLabs · 7 min · August 16, 2026
Interaction Design
Whimsical UX Teardown: Free Until You Share It
A UX teardown of Whimsical's product pages: a whiteboard that sells speed by removing the blank canvas, and a free plan that gives away unlimited private boards while capping shared ones at three.
TYPENORMLabs · 6 min · August 6, 2026
Research Methods
Writing Closed Questions in Research: Getting Answers You Can Count
A closed question fixes the answer set before anyone reads it, which is what makes it countable and what makes it fragile. The forms, the five ways the wording breaks, and how to pretest before you send.
TYPENORMLabs · 9 min · August 15, 2026