Glossary

Vector Database

A database that stores and searches over embedding vectors — the data infrastructure underlying retrieval-augmented generation and semantic search in AI research tools.


What it means

A vector database is a database system designed to store and retrieve high-dimensional numerical vectors (embeddings) efficiently, particularly for nearest-neighbor search. Where a traditional database retrieves rows by exact match on values (find all records where field = value), a vector database retrieves records by similarity (find the N records whose embedding vector is closest to this query vector).

Vector databases underlie most retrieval-augmented generation (RAG) systems: when you ask a tool like NotebookLM or Elicit a question, the system converts your question to an embedding vector, queries the vector database for the most similar document chunks, retrieves those chunks, and feeds them to the language model as context.

Common vector databases in production systems: Pinecone, Weaviate, Qdrant, Chroma, pgvector (PostgreSQL extension), and Milvus. These are primarily used by developers building research tools rather than by end-user researchers directly.

Why it matters for researchers

You use vector databases every time you use a semantic search tool. The underlying plumbing of Semantic Scholar, Elicit, and similar tools is a vector database storing embeddings for millions of papers.

For researchers building their own tools: If you want to build a custom literature assistant — for example, a system that answers questions over your group’s internal unpublished data — the standard approach involves:

  1. Chunking your documents into passages
  2. Embedding each passage into a vector (using a model like OpenAI’s text-embedding-3-small)
  3. Storing the vectors in a vector database
  4. At query time, embedding the question, finding the nearest passages, and feeding them to an LLM

This pipeline is called RAG (retrieval-augmented generation), and vector databases are the retrieval component.

The practical limitation: Vector databases return the most semantically similar chunks, not necessarily the most relevant ones for a given task. A passage that is about a similar topic but doesn’t contain the specific information you need may rank higher than a relevant passage that uses different vocabulary. Hybrid search (combining vector similarity with keyword matching) often outperforms pure semantic search in practice.