Glossary

Reproducibility

The ability of independent researchers to obtain the same results using the same data and methods — distinct from replicability (same results with new data) and a growing concern in AI research where model and training variability make reproduction difficult.


What it means

Reproducibility in science means that another researcher, following the same methods with the same data, should obtain the same results. It is the minimum standard for scientific validity — a finding that cannot be reproduced by a second group is suspect.

The term is often conflated with related concepts:

Reproducibility — same data + same methods → same results. A computational or analytical question: given the original data, can I recreate the analysis?

Replicability — different data + same methods → similar results. A scientific question: does the same effect appear in an independent sample?

Generalizability — does the finding hold in a broader population or context than the original study?

The replication crisis — a term coined around 2011 after large replication projects in psychology, social science, and medicine found that a substantial fraction of published findings (40–60% in some fields) did not replicate. Key causes include underpowered studies, p-hacking, publication bias, and HARKing.

Reproducibility in AI research

AI research faces particularly severe reproducibility problems:

Random seeds matter. Neural network training involves stochastic gradient descent with random weight initialization and mini-batch sampling. The same code with a different random seed can produce meaningfully different results, especially for small datasets or sensitive tasks.

Unreported hyperparameters. Papers frequently omit or underspecify hyperparameter choices, making exact reproduction difficult even with the original code and data.

Hardware and framework dependence. Floating-point operations on different GPU architectures or in different library versions (PyTorch 1.x vs 2.x, CUDA version) can produce different numerical results that accumulate through training.

Benchmark contamination. Models trained on data that inadvertently includes benchmark test set content will appear to outperform on those benchmarks. As internet-scraped training corpora have grown, contamination has become a serious concern for LLM evaluation.

Tools and practices for reproducibility

For data analysis (R/Python):

  • Renv / conda environment files to pin package versions
  • Quarto / R Markdown for computational notebooks that combine code and narrative
  • DVC (Data Version Control) for tracking dataset versions

For ML experiments:

  • Setting and reporting all random seeds
  • Weights & Biases or MLflow for experiment tracking
  • Full hyperparameter disclosure

For literature-based analysis:

  • Pre-registration of hypotheses and analysis plans (Pre-registration)
  • PRISMA reporting for systematic reviews