Glossary

Multi-omics

The integration of data from multiple 'omic' layers — genomics, transcriptomics, proteomics, metabolomics — to build a more complete picture of biological state than any single layer provides.


What it means

Multi-omics refers to the combined analysis of two or more “omic” data types — large-scale molecular measurements that characterize a biological system at a particular level:

Layer What is measured Key technology
Genomics DNA sequence, structural variants Whole-genome sequencing
Transcriptomics RNA expression levels (which genes are active) RNA-seq
Proteomics Protein abundance and modification state Mass spectrometry
Metabolomics Small molecule metabolite levels LC-MS, NMR
Epigenomics DNA methylation, histone modification ATAC-seq, ChIP-seq
Spatial omics Gene/protein expression with spatial location Visium, MERFISH

Each layer captures a different aspect of biology. Genomics tells you what is possible; transcriptomics tells you what is being expressed; proteomics tells you what is actually present and active; metabolomics tells you what the cell is doing. Integration across layers gives a more complete picture of disease, development, or treatment response.

Why integration is hard

Different scales and distributions. Genomic data is binary or categorical (variant present/absent); transcript data is count-based with heavy zero-inflation; proteomic data is continuous with heavy right skew. Combining them requires careful normalization.

Different sample requirements. Collecting transcriptomics, proteomics, and metabolomics from the same patient sample at the same time is technically challenging — different protocols require different sample preparation, and some require tissue amounts that conflict with clinical constraints.

Dimensionality. A single patient’s genome contributes ~4 million common variants; their transcriptome ~20,000 gene expression values; their proteome ~7,000 quantified proteins; their metabolome ~1,000 metabolites. The total feature space vastly exceeds the number of patients in most studies, requiring dimensionality reduction and regularization.

AI methods for multi-omics

Factor analysis and matrix factorization (MOFA, NMF) decompose multi-omics data into shared latent factors that represent biological axes of variation (disease progression, cell type composition, treatment response).

Graph neural networks on multi-omics data represent each sample as a node in a patient similarity network, with edges weighted by molecular similarity.

Deep learning integration — multi-input neural networks with separate encoders for each omic type, combined at a fusion layer. These require large sample sizes to avoid overfitting.

Practical relevance

Multi-omics data from large biobanks (UK Biobank, GTEx, TCGA) is publicly available for secondary analysis. For researchers entering this space, the Bioconductor ecosystem (R) and tools like MOFA+ provide well-documented starting points.