Glossary

p-value

The probability of observing results at least as extreme as the data if the null hypothesis were true — widely misunderstood, widely misused, and insufficient on its own for evaluating scientific evidence.


What it means

A p-value is the probability of obtaining test results at least as extreme as the observed results, assuming the null hypothesis is true. A p-value of 0.03 means: if there were truly no effect (null hypothesis), there is a 3% chance of observing data this extreme or more extreme by chance alone.

What p-values do not tell you:

  • The probability that the null hypothesis is true
  • The probability that your result will replicate
  • The size or importance of the effect
  • Whether the result is scientifically meaningful

The p < 0.05 threshold for “statistical significance” is a convention, not a law of nature. It was adopted arbitrarily and has caused significant harm through selective reporting, multiple testing problems, and the conflation of statistical and practical significance.

The p-value and the replication crisis

Much of the replication crisis in psychology, medicine, and social science traces to p-value misuse:

  • p-hacking: Running multiple analyses and reporting only the one that crosses 0.05
  • HARKing (Hypothesizing After Results are Known): Presenting exploratory findings as confirmatory
  • Publication bias: Journals preferentially publishing p < 0.05 results, making the literature unrepresentative

AI tools that mine the literature inherit these biases. A meta-analysis trained on published studies with positive results will overestimate true effect sizes.

What to use instead (or alongside)

Effect size — how large is the effect, in practically meaningful units? A p < 0.001 result with a Cohen’s d of 0.05 is statistically significant but practically trivial.

Confidence intervals — the range of effect sizes consistent with the data. More informative than a binary significant/not-significant judgment.

Pre-registration — specifying hypotheses and analysis plans before collecting data eliminates most opportunities for p-hacking.

Bayesian methods — provide posterior probabilities that directly answer “how likely is this effect to be real, given what we know?”