The Likert scale is the workhorse of psychology and the social sciences. “I feel confident in my abilities: strongly disagree, disagree, neutral, agree, strongly agree.” Assign the numbers 1 through 5, collect responses, compute means, run correlations, publish. It’s so familiar that most researchers never stop to think about what they’re actually doing.
They should. Because underneath that familiar five-option format sits a stack of measurement questions that are genuinely complicated, and the answers you choose affect what your data can and cannot tell you. This post walks through five of them: the measurement level of Likert responses, the choice between Pearson and Spearman correlations, sum scores versus latent variables, the number of response categories, and response styles.
Are Likert Responses Ordinal?
To answer this, we need to start with measurement levels themselves.
The simplest level is nominal. Nominal variables discriminate between values qualitatively, with no numerical relationship between them. Where were you born? Which university do you study at? The answers are categories – Tbilisi, Batumi, Kutaisi – and there’s no sense in which one is “more” than another. There is no numerical difference.
At the other end sits the interval/ratio level, where differences between values are genuinely mathematical. Age is the classic example: someone who is 40 is exactly twice as old as someone who is 20. And the difference between 20 and 40 is exactly the same as the difference between 40 and 60. The numbers behave like numbers.
Between these two sits the ordinal level, and this is where Likert responses live. Ordinal data gives us more information than nominal: we can order the values by strength. “Strongly agree” expresses more agreement than “somewhat agree.” That ordering is real and meaningful.
But two crucial properties of interval data are missing. First, the distances between response options are not equal. The psychological gap between “neutral” and “agree” is not necessarily the same as the gap between “agree” and “strongly agree” – unlike the equal steps of 20-40-60 in age. Second, each person interprets the options a bit differently. Your “strongly agree” and my “strongly agree” may not represent the same intensity of attitude. Assigning the numbers 1 through 5 makes the options look evenly spaced and objective, and it may even help respondents discriminate between them, but in real life, the numbers don’t behave that way.
So the honest answer is: yes, Likert responses are ordinal. And understanding that is important, because everything else in this post follows from it.
So Why Does Everyone Use Means and t-Tests?
If Likert data is ordinal, then strictly speaking, means, standard deviations, t-tests, ANOVAs, and Pearson correlations – all designed for interval data – shouldn’t apply. Yet open any psychology journal and you’ll find them everywhere. Is the whole field doing it wrong?
Not exactly. There are two honest reasons for the current practice, one pragmatic and one empirical.
The pragmatic reason is that ordinal data is statistically limiting. Suppose I collect Likert responses about attitudes toward different branches of government. Which branch do people like most, and which least? With non-parametric tests alone, getting to a clear, interpretable answer is genuinely difficult. Non-parametric methods exist, but they answer narrower questions, are harder to interpret, and don’t extend gracefully to the complex models researchers actually need. There’s also a simpler factor: general statistical training. Most researchers know parametric methods well and non-parametric or ordinal methods poorly, so the parametric route is the one they can actually walk.
The empirical reason is more reassuring. A substantial body of methodological research suggests that parametric statistics are surprisingly robust when applied to Likert-type data. Geoff Norman’s much-cited 2010 paper in Advances in Health Sciences Education reviewed evidence going back to the 1930s and concluded that parametric methods can be used on Likert data without serious risk of reaching wrong conclusions, even with skewed distributions. Earlier work by Hsu and Feldt (1969) and later reviews by Carifio and Perla (2007) point in the same direction, particularly when items are aggregated into multi-item scale scores rather than analyzed individually. And crucially, the robustness improves as the number of response categories increases – data from items with more options behaves more and more like interval data.
So the practice isn’t indefensible. But it is a compromise, and knowing it’s a compromise matters, especially for single items, small samples, and heavily skewed distributions, where the ordinal nature of the data bites hardest.
Pearson vs Spearman
The correlation question is a concrete version of the same dilemma. Pearson’s correlation assumes interval-level data and measures the strength of a linear relationship between two variables. Spearman’s correlation works on ranks – it converts each variable to rank order and asks whether the two orderings agree. Spearman is, by construction, the ordinal-appropriate choice.
In practice, the two often produce very similar results with Likert data. Studies that compute both routinely find nearly identical significance patterns and coefficient magnitudes. When the distributions are roughly symmetric and the relationship roughly monotonic, the choice barely matters.
But they diverge in exactly the situations where Likert data is most problematic: heavily skewed distributions, few response categories, floor and ceiling effects, and outliers. A single extreme respondent can distort a Pearson coefficient; Spearman, working on ranks, is far more resistant. My practical guidance: for correlations between single Likert items, Spearman is the defensible default. For correlations between multi-item scale scores – sums or means of many items, which approximate continuous distributions – Pearson is usually fine. And when in doubt, compute both. If they agree, report either with confidence. If they disagree substantially, that disagreement is itself telling you something about your distributions worth investigating.
Sum Scores vs Latent Variables
The most common way to use a Likert scale is to add up the items. Ten items, each scored 1-5, produces a total between 10 and 50, and that total becomes “the score” – anxiety, satisfaction, self-efficacy, whatever the scale measures.
Sum scores are convenient, but they smuggle in strong assumptions that are rarely examined. Adding items with equal weight assumes every item is an equally good indicator of the construct – that item 3 measures anxiety exactly as well as item 7. It assumes the response categories function as equal intervals, which we’ve already established they don’t. And it treats the resulting total as if it were measured without error.
Latent variable approaches – confirmatory factor analysis for the factor-analytic tradition, or IRT models like the graded response model for ordinal items – replace those assumptions with estimates. Each item gets its own loading or discrimination parameter: the model learns which items are strong indicators of the construct and which are weak, and weights them accordingly. The category thresholds are estimated too, respecting the ordinal structure instead of imposing equal intervals. The model separates true score variance from measurement error rather than pretending error doesn’t exist. And latent variable models unlock analyses sum scores can’t support at all – most importantly measurement invariance testing, which asks whether the scale works the same way across groups, languages, or time points. For anyone doing cross-cultural adaptation work, as I do with Georgian versions of established instruments, invariance testing isn’t optional – and it requires the latent variable framework.
None of this means sum scores are useless. When a scale is well-constructed, unidimensional, and items load roughly equally, sum scores and latent scores correlate very highly, and the simple approach loses little. But that’s a conclusion you can only reach after fitting the latent model and checking. The sum score should be an informed simplification, not a default you never questioned.
How Many Response Categories?
Four? Five? Seven? Ten? The choice looks cosmetic but shapes both the respondent’s experience and the data’s statistical behavior.
My own practice: five categories is the easiest format for respondents. It’s familiar, quick to process, and the labels come naturally: strongly disagree through strongly agree with a neutral midpoint. This is my default.
When I don’t want respondents to have a neutral option, I move to four or six categories. Removing the midpoint forces a direction: the respondent must lean at least slightly toward agreement or disagreement. I do this when the research question requires people to take a side – when a neutral answer would dodge exactly the question I’m asking. The trade-off is real, though. Some respondents genuinely are neutral, and forcing them to choose introduces its own noise. The midpoint also serves as a resting place for satisficers – people answering carelessly – so removing it can push careless responses outward instead of concentrating them safely in the middle. The decision should follow from the research question, not from habit.
I’ve used seven categories, but labeling becomes genuinely difficult – finding seven verbal anchors that respondents interpret as roughly evenly spaced is hard in any language, and harder still when the instrument needs to work in Georgian and English simultaneously. Beyond seven, fully labeled scales become practically impossible. For longer scales – ten points, say – I use polar labeling: anchor only the endpoints (1 = totally disagree, 10 = totally agree) and leave the numbers in between unlabeled. Respondents treat the unlabeled numbers as evenly spaced steps, which, interestingly, may bring the data closer to interval behavior than fully labeled formats do.
The research literature broadly supports this pragmatism. Dawes (2008) compared 5, 7, and 10-point formats and found rescaled scores were largely comparable across them. More categories yield somewhat finer information and, as noted earlier, data that better approximates interval properties, but the gains flatten out quickly, and respondent burden rises. There is no magic number. There are trade-offs, and the best format depends on the construct, the population, and what you plan to do with the data.
Response Styles – The Human Element
Everything so far assumed respondents use the scale the way we intend. They don’t. or at least, not all of them, not always. People bring systematic habits to rating scales that have nothing to do with the construct being measured. These are response styles, and they contaminate Likert data in ways that are invisible in any single response.
Acquiescence is the tendency to agree – to say yes regardless of content. An acquiescent respondent drifts toward “agree” on every item, inflating scores on positively-worded scales. The classic remedy is reverse-coded items: word half the items in the opposite direction so agreement sometimes raises and sometimes lowers the score. But reverse-coded items carry their own problems – respondents misread negations, and reversed items often form spurious method factors in factor analyses – so the cure is only partial.
Extreme responding is the habit of living at the endpoints – everything is 1 or 5, strongly this or strongly that. Its mirror image, central tendency responding, avoids the endpoints entirely and hovers around the middle. Two respondents with identical attitudes but different styles will produce visibly different data. Critically for anyone doing cross-cultural work, these styles vary systematically between cultures: some populations use endpoints far more freely than others. When comparing Likert-based scores across countries or language versions – precisely the business of instrument adaptation – apparent group differences can be partly or wholly artifacts of differing response styles. This is yet another reason measurement invariance testing matters before any cross-group comparison is taken at face value.
Social desirability is the pull toward answers that look good: under-reporting stigmatized attitudes and behaviors, over-reporting virtuous ones. It’s strongest for sensitive topics and in non-anonymous settings.
What can be done? Guarantee and emphasize anonymity. Word items neutrally. Use balanced keying with care. Consider anchoring vignettes or forced-choice formats for high-stakes cross-cultural comparisons. And model response styles statistically when the data and design allow. None of these eliminates the problem; together, they shrink it.
Use Them Carefully
After all this, should researchers abandon Likert scales? No. The Likert format survives because it works: it’s fast, flexible, understood by respondents everywhere, and – handled properly – supports genuinely rigorous measurement. Nearly a century of psychological science has been built on it, and much of that science replicates just fine.
The point is not to stop using Likert scales. The point is to stop using them thoughtlessly. Know that your data is ordinal, and know what that costs you. Choose your correlation coefficient deliberately. Fit a latent variable model before trusting the sum score. Pick your number of response categories to fit the research question, not out of habit. And remember that behind every rating sits a human being with their own habits of using scales, which your numbers absorb along with the construct you care about.
A Likert scale looks like the simplest thing in research methods. It is actually a chain of measurement decisions, each with consequences. The researchers who understand those decisions get data they can defend. The ones who don’t get numbers that look precise and mean less than they appear to.
References mentioned: Norman, G. (2010). Likert scales, levels of measurement and the “laws” of statistics. Advances in Health Sciences Education, 15(5), 625-632. · Carifio, J., & Perla, R. (2007). Ten common misunderstandings, misconceptions, persistent myths and urban legends about Likert scales and Likert response formats. Journal of Social Sciences, 3(3), 106-116. · Hsu, T. C., & Feldt, L. S. (1969). The effect of limitations on the number of criterion score values on the significance level of the F-test. American Educational Research Journal, 6(4), 515-527. · Dawes, J. (2008). Do data characteristics change according to the number of scale points used? International Journal of Market Research, 50(1), 61-77.
—
Giorgi Tchumburidze
July, 2026