Multimodal Model
An AI model that processes and generates multiple data types — text, images, audio, or structured data — within a single model, rather than using separate specialized models for each modality.
What it means
A multimodal model is an AI system that can process and generate more than one type of data (modality) within a unified architecture. The most common combination is text and images, but scientific multimodal models also span sequences, structures, and spectra.
Examples in general AI tools researchers use:
- GPT-4o, Claude 3+, Gemini — accept text and images as input; useful for analyzing figures, microscopy images, gel images, and charts directly
- Whisper (OpenAI) — audio-to-text; useful for transcribing lectures, interviews, and conference talks
Examples in scientific AI:
- ESM3 (EvolutionaryScale) — protein model that jointly reasons over sequence, structure, and function in one model, enabling cross-modal prompting (e.g., “generate a sequence that folds to this structure and performs this function”)
- AlphaFold 3 — takes protein sequences alongside small molecule SMILES strings and DNA/RNA sequences as inputs; multimodal in the sense that multiple molecular types are jointly modeled
- BioViL / MedCLIP — vision-language models trained on radiology images and clinical reports
Why it matters for researchers
Directly analyzing figures: You can now paste a figure from a paper into Claude or GPT-4o and ask “what does this Western blot show?” or “does this calibration curve look linear?” The model can interpret visual scientific content — with the usual caveat that it can still be wrong and should be treated as a first-pass interpretation, not a definitive read.
Cross-modal scientific design: In structural biology, ESM3 allows you to specify design constraints in one modality (e.g., “this structural scaffold”) and receive outputs in another (e.g., a sequence that will fold to it). This is qualitatively different from tools that handle each modality separately.
Practical research use: For researchers not working in structural biology or molecular design, the most immediately useful multimodal capability is image understanding in general-purpose tools — extracting data from charts, interpreting microscopy or histology images, and reading figures from papers you’ve uploaded.