Glossary

Confounding

A variable that causally affects both the exposure and outcome in a study, creating a spurious association — the central challenge in observational research and a key source of unreliable findings that AI tools can amplify rather than correct.


What it means

A confounder is a variable that is associated with both the exposure (treatment, risk factor) and the outcome, and that can create or distort the apparent relationship between them. Confounding is why observational studies cannot establish causation without careful design and analysis.

Classic example: Studies in the 1950s found that carrying a lighter was strongly associated with lung cancer. Lighters do not cause lung cancer — smoking does. Smoking caused people to carry lighters (to light cigarettes) and also caused lung cancer. The lighter-cancer association was entirely explained by the confounder: smoking.

In practice: Any analysis comparing groups that differ in ways other than the exposure of interest risks confounding. In clinical research: patients who receive a treatment may be systematically sicker (or healthier) than those who don’t, making treatment effects difficult to estimate. In genomics: population stratification (different ancestry distributions in cases vs. controls) creates spurious GWAS associations.

Confounding and AI/ML

Spurious correlations in training data. ML models trained on observational data learn confounded associations. A chest X-ray model trained in hospitals might learn to detect the portable X-ray markers that accompany critically ill patients as a proxy for mortality — not because the hardware causes death but because the same patients received both. The model performs well on training data and fails on deployment where this confound doesn’t exist.

Shortcut learning / clever Hans effects. A model that achieves high accuracy on a benchmark by exploiting a confounded feature (image watermarks, sentence length, hospital-specific artifacts) rather than the true signal is exhibiting confounding-driven failure. These failures often go undetected until external validation or adversarial probing.

LLMs and literature confounding. When AI tools mine observational literature to extract associations, they extract confounded associations as readily as causal ones. An AI that reads 10,000 observational studies and concludes “X causes Y” may be wrong if all those studies shared the same confounders.

Controlling for confounding

Randomization — the gold standard. Random treatment assignment ensures that confounders are balanced across groups, in expectation.

Adjustment methods — in observational data: regression adjustment, propensity score matching, inverse probability weighting. These work when all important confounders are measured; unmeasured confounders remain a threat.

Instrumental variables / Mendelian randomization — uses genetic variants as instruments to estimate causal effects in observational data. Widely used in epidemiology.