What a p-Value Actually Means (and What It Doesn’t)

The p-value is the most used and most misunderstood number in all of statistics. It appears in nearly every empirical paper in psychology, medicine, and the social sciences. Researchers chase it, reviewers demand it, and careers are built on getting it below a certain threshold. And yet, if you ask most people who use p-values every day to define one precisely, they can’t. Or worse, they define it wrong.

Let’s fix that. And the best way to understand a p-value is to go back to the story where it was born.

The Lady Tasting Tea

In the 1920s, the statistician Ronald Fisher was at a gathering where a woman claimed she could tell, just by tasting, whether the milk or the tea had been poured into the cup first. Most people would either believe her or dismiss her. Fisher, being Fisher, decided to test her.

Here’s what he did. He prepared eight cups of tea, some with milk poured first, some with tea poured first, and asked her to identify which was which. The question Fisher asked himself was simple: if she’s just guessing, how likely is it that she’d get them all right by chance?

This is the heart of the p-value. We start by assuming she has no special ability, that she’s purely guessing. This assumption is called the null hypothesis. Then we ask: under that assumption, how probable is the result we observed?

With eight cups arranged as four of each type, there’s exactly one way to identify all of them correctly out of the many possible arrangements of guesses. The probability of getting all eight right purely by chance works out to about 1.4%, less than one in seventy. So if she does get them all right, we’re faced with a choice: either something very unlikely happened by pure luck, or she isn’t actually guessing. The rarer that “by pure luck” explanation is, the more we’re inclined to believe she genuinely can taste the difference.

That probability – the chance of getting a result at least as impressive as the one we observed, assuming she was only guessing – is the p-value. If it’s small enough (Fisher’s convention was below 5%, or 0.05), we call the result statistically significant and conclude that guessing alone probably doesn’t explain what we saw.

So What Is a p-Value, Precisely?

Let’s state it carefully, because the precise wording matters: a p-value is the probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is true.

In the tea example, the null hypothesis is “she’s guessing.” The p-value is the probability that pure guessing would produce a performance as good as (or better than) what she actually achieved. A small p-value means the observed result would be surprising if the null hypothesis were true – which gives us reason to doubt the null hypothesis.

Notice what the p-value is doing. It’s not telling us the probability that she has the ability. It’s telling us how compatible our data is with the assumption that she doesn’t. That distinction is subtle, and it’s the source of nearly every misconception about p-values.

What a p-Value Is NOT

Here’s where almost everyone goes wrong. Four misconceptions are so common that they’ve become part of how researchers casually (and incorrectly) talk about their results.

Misconception 1: The p-value is the probability that the null hypothesis is true. It is not. A p-value of 0.02 does not mean “there’s a 2% chance she’s just guessing.” The p-value is computed assuming the null is true – it can’t also tell you the probability that the null is true. That would be circular. The p-value answers “how surprising is my data if the null holds?” not “how likely is the null given my data?” Those are different questions with different answers.

Misconception 2: The p-value is the probability your result was due to chance. Closely related, and equally wrong. The p-value already assumes chance is the only thing operating (that’s the null). It can’t then tell you the probability that chance was responsible – it presupposes it. What it tells you is how often chance alone would produce a result this extreme.

Misconception 3: A small p-value means a large or important effect. Absolutely not. Statistical significance is not the same as practical importance. With a large enough sample, a trivially small, meaningless effect can produce a tiny p-value. The p-value tells you whether an effect is detectable, not whether it’s big or whether it matters. A p-value says nothing about the size of what you found.

Misconception 4: A p-value above 0.05 proves the null hypothesis is true. No. Failing to reject the null is not the same as confirming it. If the lady gets six of eight right, we might not have enough evidence to rule out guessing, but that doesn’t prove she was guessing. Absence of evidence is not evidence of absence. A non-significant result means “we didn’t find enough evidence,” not “there is nothing there.”

If you internalize these four points, you already understand p-values better than a large share of people who use them professionally.

The Arbitrary 0.05

Why 0.05? Why is that the magic line between “significant” and “not significant”?

The honest answer: it’s arbitrary. There’s no deep mathematical meaning to 0.05. Fisher simply proposed it as a reasonable convention, and over the decades it hardened into something treated almost like a law of nature. But we could just as easily use any other threshold. If we wanted to be stricter, we could demand 0.01, or 0.001.

Coming back to the tea: a stricter threshold simply means demanding more evidence. If we wanted to be more confident before declaring the lady a genuine taster, we’d give her more cups. Eight cups can get us to around 1.4%; more cups, done perfectly, would push the probability of a lucky guess even lower. The threshold isn’t sacred, it’s a choice about how much evidence you require before you’re willing to abandon the null hypothesis. Fisher picked a number, the field adopted it, and now it’s a magic number that far too many people treat as the boundary between truth and noise.

The Danger of Chasing Significance

Treating 0.05 as a sacred line creates a serious problem, because it turns a continuous measure of evidence into a binary verdict with enormous consequences. A result at p = 0.049 is “significant” and publishable. A result at p = 0.051 is “non-significant” and often unpublishable. But statistically, these two results are almost identical. The difference between them is trivial, yet the professional consequences are night and day.

This creates a perverse incentive. Imagine a researcher whose analysis returns p = 0.051. So close. The result they hoped for is just barely out of reach. What happens if they remove a single inconvenient data point – an “outlier” – and the p-value drops to 0.049? Suddenly the result is significant, publishable, career-advancing. One data point, changed or dropped, flips the entire conclusion.

This is the essence of p-hacking: consciously or unconsciously nudging an analysis until it crosses the magic threshold. Removing outliers selectively, trying different statistical tests until one works, collecting a few more participants and re-testing, splitting the data in different ways – each of these can push a p-value across the line. And because the threshold is arbitrary but the rewards for crossing it are real, the temptation is enormous. Multiply this across thousands of researchers under pressure to publish, and you get a literature full of results that sit suspiciously close to 0.05 and often fail to replicate. This is a major driver of the replication crisis that has shaken psychology and other fields.

The problem isn’t the p-value itself. The problem is the binary, all-or-nothing way we’ve chosen to use it.

What to Report Instead, Or Alongside

So what should researchers do? The answer isn’t to abandon p-values, but to stop letting them carry all the weight. Three additions make results far more informative and far more honest.

Effect sizes. An effect size tells you how big the phenomenon is, not just whether it’s detectable. Where a p-value gives a yes/no verdict on statistical significance, an effect size (like Cohen’s d, or a correlation coefficient, or an odds ratio) quantifies the magnitude of what you found. Two studies can both have p < 0.05 while one has a large, meaningful effect and the other has a tiny, trivial one. Reporting the effect size makes that difference visible. It answers the question the p-value can’t: does this actually matter?

Confidence intervals. A confidence interval gives you a range of plausible values for the true effect, rather than a single point estimate and a binary decision. Instead of “the effect is significant (p < 0.05),” you say “the effect is estimated at this value, and the true value is plausibly somewhere in this range.” A wide interval signals a lot of uncertainty; a narrow one signals precision. Confidence intervals carry all the information a p-value does – if the interval excludes zero, the result is significant – plus a great deal more about magnitude and precision. They shift the focus from “is there an effect?” to “how big, and how sure are we?”

Bootstrapping. Bootstrapping is a resampling technique that estimates uncertainty by repeatedly drawing samples, with replacement, from your own data. It’s a powerful option because it doesn’t rely on the distributional assumptions that classical tests require – it builds an empirical picture of how your statistic varies, directly from the data itself. When your data violates the assumptions behind traditional tests, or when you want a robust estimate of a confidence interval for some complex statistic, bootstrapping often gives you a more trustworthy answer. Modern computing makes it easy, and it’s an excellent complement to (or replacement for) formula-based inference.

Report an effect size so readers know how big. Report a confidence interval so readers know how precise. Use bootstrapping when your assumptions are shaky. Together, these give a far richer and more honest picture than a lone p-value ever could.

The Bottom Line

The p-value has taken a lot of criticism in recent years, and some of it is deserved. But the p-value isn’t evil. It’s a genuinely useful tool that answers a specific, well-defined question: how compatible is my data with the null hypothesis? The trouble comes when people misunderstand what it says, treat its arbitrary threshold as sacred, and let a single number make decisions it was never designed to make.

The p-value isn’t the villain of statistics. You just have to know what it actually tells you, and use it accordingly – as one piece of evidence among several, not as the final word. Understand it correctly, pair it with effect sizes and confidence intervals, resist the pull of the magic threshold, and the p-value becomes exactly what Fisher intended: a reasonable way to weigh evidence against chance.


Giorgi Tchumburidze
August, 2026

Leave a Comment