Tokenization
The process of splitting text into tokens — the basic units an LLM processes — which are roughly word pieces, not whole words, explaining why a 10,000-word document uses more than 10,000 tokens.
What it means
Tokenization is the process of converting text into the discrete units — called tokens — that a large language model actually processes. Tokens are not whole words: they are word pieces, chosen by an algorithm that balances vocabulary size against coverage of rare words and scientific terminology.
Rules of thumb for estimating tokens:
- Common English words: typically 1 token each
- Longer or less-common words: 2–3 tokens (e.g., “biosynthesis” → “bio” + “synthesis” = 2 tokens)
- Chemical names, SMILES strings, gene identifiers: can be many tokens per term
- A typical academic paper (~8,000 words): roughly 11,000–13,000 tokens
- The rough average: 1 token ≈ 0.75 words, or 100 tokens ≈ 75 words
Why it matters for researchers
Scientific text uses more tokens than general text. A methods section dense with chemical formulas, gene names, and species nomenclature may tokenize at closer to 1:1 word-to-token ratio than the 0.75 average. This means your document may hit a model’s context limit faster than expected if it contains a lot of specialized terminology.
Token limits affect pricing too. LLM API pricing is per token (both input and output). If you’re building a research pipeline that processes many documents, estimating token counts accurately matters for cost forecasting.
Why some terms are expensive to tokenize: Models learn their tokenizer vocabularies from training corpora. Words that appear rarely in the training data get broken into more subword pieces. Highly specialized scientific terminology — IUPAC chemical names, uncommon gene symbols, journal-specific abbreviations — often tokenizes inefficiently.
Practical tools: The Tiktokenizer website (by OpenAI) and similar tools let you paste text and see exactly how a specific model’s tokenizer splits it before sending it to the API.