Statistical Power
The probability that a study will detect a true effect if one exists — determined by sample size, effect size, and alpha level. Underpowered studies waste resources and produce unreliable results.
What it means
Statistical power is the probability that a statistical test will correctly reject a false null hypothesis — in other words, the probability of detecting a real effect if one actually exists. Power ranges from 0 to 1 (or 0–100%).
The standard target is 80% power: if a true effect exists, a well-designed study should detect it 80% of the time. Higher power costs more (requires more participants or observations); the 80% convention is a practical balance.
Power is determined by four interrelated quantities:
- Sample size (n) — more observations → more power
- Effect size (δ) — larger effects are easier to detect
- Significance threshold (α) — stricter thresholds (lower α) require more power to achieve
- Variability (σ) — less noise → more power
Given any three, you can solve for the fourth. Power analysis before a study calculates the sample size needed to achieve 80% power given an expected effect size, assumed variability, and chosen α.
Why underpowered studies are a serious problem
An underpowered study that finds p < 0.05 has likely detected an inflated effect size. This is the winner’s curse: among the studies that happen to achieve significance despite being underpowered, the ones that reach p < 0.05 are disproportionately those that overestimated the effect by chance. The published literature systematically overstates effect sizes for this reason.
Conversely, an underpowered study that finds p > 0.05 cannot conclude there is no effect — it only demonstrates the study was too small to detect an effect of the expected size.
AI tools and power analysis
AI assistants (Claude, ChatGPT) can help with power calculations — describe your design, outcome measure, and expected effect size, and they will generate the R or Python code using pwr, G*Power, or scipy.stats. The generated code is usually correct for standard designs; verify the effect size assumption carefully, since this drives the result more than anything else.
Where LLMs are useful: Explaining power concepts, generating boilerplate code for standard tests, helping interpret power analysis output.
Where LLMs are unreliable: Choosing an appropriate effect size for a specific clinical or scientific context — this requires domain knowledge about what effect size is plausible and what would be clinically meaningful.