Natural Language Processing (NLP)
The field of AI focused on enabling computers to understand, interpret, and generate human language — the foundation for LLMs, semantic search, text extraction, and most AI research tools.
What it means
Natural language processing (NLP) is the subfield of artificial intelligence concerned with the computational processing of human language — reading, understanding, generating, and translating text. It is the foundational discipline underlying large language models, semantic search, information extraction, machine translation, and automatic summarization.
NLP has been transformed by the transformer architecture and the large language model paradigm since 2017–2020. Most NLP tasks that were previously handled by separate specialized models (named entity recognition, sentiment analysis, question answering, summarization) are now handled by fine-tuning or prompting a single large model.
Core NLP tasks researchers encounter in AI tools:
| Task | What it does | Example in research tools |
|---|---|---|
| Named entity recognition (NER) | Identifies and classifies entities (genes, chemicals, diseases) in text | Extracting compound names from chemistry papers |
| Text classification | Assigns labels to text | Rayyan’s inclusion/exclusion classification |
| Semantic search | Retrieves semantically related documents | Elicit, Semantic Scholar |
| Summarization | Generates condensed summaries of text | NotebookLM, Elicit TLDR |
| Information extraction | Extracts structured data from unstructured text | Elicit’s custom extraction fields |
| Question answering | Answers questions given a context document | NotebookLM, Claude with PDFs |
Why it matters for researchers
All the tools on this site are built on NLP. Understanding what NLP can and cannot do reliably helps set appropriate expectations:
- NLP excels at tasks where the answer is contained explicitly in the source text
- NLP struggles with tasks requiring external knowledge, numerical reasoning over tables, and drawing valid inferences that require domain expertise
- Biomedical NLP is a distinct subfield: general NLP models perform worse on scientific text because the vocabulary, sentence structure, and reasoning patterns differ from the web text they were trained on. Domain-specific models (BioBERT, PubMedBERT, BioGPT) are often better for specialized medical/biological text extraction tasks