In-Context Learning
An LLM's ability to learn a new task from examples provided directly in the prompt, without any weight updates — the mechanism behind few-shot and zero-shot prompting.
What it means
In-context learning (ICL) is the ability of large language models to perform a new task from examples or instructions provided within the input prompt itself, without any modification to the model’s weights through training. The model “learns” the pattern from the examples in context and applies it to a new input, all within a single forward pass.
The spectrum of in-context learning:
- Zero-shot: The model is given only the task description, with no examples (“Classify this abstract as about drug discovery or not”)
- Few-shot: The model is given a small number of labeled examples alongside the task (“Here are 3 examples of abstracts classified as drug discovery / not drug discovery. Now classify this one”)
- Many-shot: Newer long-context models can accept hundreds of examples in the prompt, substantially improving performance on specialized tasks
Why it matters for researchers
In-context learning is how you customize LLM behavior without fine-tuning. If you want an LLM to extract a specific type of information from a paper, classify text according to a custom taxonomy, or generate outputs in a particular format, the practical way to do this is to include labeled examples of your desired input-output pattern in the prompt.
For research extraction tasks:
- Instead of “Extract the sample size from this paper,” try: “Here are examples of how I want you to extract sample sizes: [3 examples with the output format you want]. Now extract from this paper.”
- The quality of in-context examples matters significantly — clear, representative examples that edge cases are more predictive than many mediocre ones
Limitations:
- ICL performance degrades if examples are inconsistent or don’t represent the full variation in the task
- Beyond a certain number of examples, additional examples in context give diminishing returns compared to fine-tuning the model
- The “lost in the middle” problem in long contexts means examples placed early or late in a long prompt are recalled more reliably than those buried in the middle