The number that ate a science.
A convenience threshold from the 1920s became the gatekeeper of an entire discipline — and then the centerpiece of its crisis. This is the story of p < 0.05: what it means, how it broke psychology, and how the repair looks.
How a convenience became a gate#
Mid-century psychology wanted the authority of physics. It got a threshold instead.
After the hard sciences’ mid-century triumphs, psychology sought definitive, universally valid results. It reached for statistics: the p-value, descended from Fisher’s informal screening tool, became the field’s test of truth. A p below 0.05 came to mean “real”; journals began publishing little else. The gate created the incentive, and the incentive created the crisis: researchers bent analyses — p-hacking, in today’s vocabulary — until the number cleared the bar.
What p < 0.05 does and does not say#
Most misuse comes from wishing the number answered a different question than it does.
| Claim | Verdict |
|---|---|
| “p < 0.05 means the effect is real” | No — it means the data are unusual under the null. Rare things still happen. |
| “p > 0.05 means there is no effect” | No — absence of evidence is not evidence of absence; the study may simply be underpowered. |
| “p is the probability the hypothesis is true” | No — that quantity needs Bayesian machinery the p-value does not have. |
| “Significant means important” | No — with enough data, trivial effects clear 0.05 easily. Effect size is the question. |
In 2016 the American Statistical Association took the unprecedented step of issuing a formal statement on p-values — a scientific society warning the world about its own most popular tool.
The crisis the threshold built#
Publication filters plus a gameable number produced exactly what the incentives promised.
Editors like Geoffrey Loftus — who ran Memory & Cognition in the mid-1990s and urged researchers to report means and effect sizes instead of significance rituals — saw it early and changed little. The reckoning arrived with the replication crisis: when the Open Science Collaboration re-ran 100 published psychology studies in 2015, only about a third of the significant findings held. Gerd Gigerenzer of the Harding Center for Risk Literacy has argued the deeper problem is ritual itself — null-hypothesis testing as “mindless statistics,” a diagnostic stamp replacing the theory-building that pioneers like Pavlov and Piaget did with simple methods and strong predictions. Rejecting a null tells you almost nothing; it merely licenses speculation about what caused the blip.
What good practice looks like now#
The field did not abandon statistics. It abandoned the ritual.
How big, with what uncertainty — confidence intervals and effect sizes say what p never could.
Registered Reports accept papers on design, before results exist — removing the incentive to torture the data.
Theories earn credibility from specific forecasts that could fail, not from piles of rejected nulls.
A screening heuristic in Fisher’s hands; a decision rule only when paired with judgment, power, and context.
Statistical significance tells you the world is probably not exactly one specific way. Science needs to know how the world is.
Which repair do you need?
Sources#
Every number on this page traces to a primary source. Here they are.
- Fisher, R.A. (1925) — Statistical Methods for Research Workers. Where the informal 5% screening convention comes from; Fisher never intended it as a publication gate. Oliver & Boyd. Overview
- Wasserstein & Lazar (2016) — The ASA’s Statement on p-Values. The American Statistician, 70(2):129–133. The unprecedented formal warning referenced in Section 02. doi.org
- Wasserstein, Schirm & Lazar (2019) — Moving to a World Beyond “p < 0.05”. The American Statistician, 73(sup1):1–19. The follow-up: don’t say “statistically significant” at all. doi.org
- Open Science Collaboration (2015) — Estimating the reproducibility of psychological science. Science, 349(6251). 100 studies re-run; 97% of originals were significant, only 36% of replications were — the ~36% figure in Section 03. doi.org
- Simmons, Nelson & Simonsohn (2011) — False-Positive Psychology. Psychological Science, 22(11):1359–1366. The paper that showed how researcher degrees of freedom make p-hacking almost effortless. doi.org
- Amrhein, Greenland & McShane (2019) — Scientists rise up against statistical significance. Nature, 567:305–307, signed by 800+ researchers. The case for retiring the dichotomy. doi.org
- Center for Open Science — Registered Reports. The preregistration venue named in repair step 2: peer review on design, before results exist. cos.io




