When we interact with a Large Language Model (LLM), it feels natural: we write text, the model "understands" it, and generates meaningful responses.
But under the hood, LLMs never see raw text.
They operate entirely on tokens.
Understanding tokenization is crucial if you want to:
- Use LLM APIs efficiently
- Control costs and context length
- Debug unexpected model behavior
- Design better prompts and pipelines
Let's break it down step by step.
What Is Tokenization?
Tokenization is the process of converting raw text into a sequence of tokens, which are numerical units that the model can process.
A token is not necessarily:
- A word
- A character
- A syllable
Instead, a token is a statistical sub-unit of text, learned from massive corpora during model training.
Example:
Input text: "Tokenization is fascinating!"
Possible tokenization: ["Token", "ization", " is", " fascinating", "!"]
Each token is then mapped to an integer ID, which is what the model actually consumes.
Why LLMs Can't Use Words Directly
Neural networks work with numbers, not symbols.
To process language, an LLM needs:
- Discrete units (tokens)
- A numerical representation for each unit
- A consistent vocabulary shared across training and inference
Tokenization solves all three problems:
- It discretizes text
- It limits vocabulary size
- It balances flexibility and efficiency
Subword Tokenization: The Core Idea
Modern LLMs use subword tokenization, not word-level tokenization.
Why?
Word-level tokenization problems
- Massive vocabularies
- Poor handling of rare words
- No generalization to unseen words
Character-level tokenization problems
- Very long sequences
- High computational cost
- Harder long-range dependencies
Subword tokenization advantages
- Handles rare and unseen words
- Keeps vocabulary manageable
- Preserves semantic structure
- Efficient sequence lengths
Common Tokenization Algorithms
Byte Pair Encoding (BPE)
- Starts from characters
- Iteratively merges the most frequent character pairs
- Used in GPT-style models
Example: "lowest" → "low" + "est"
WordPiece
- Optimizes likelihood rather than frequency
- Common in BERT-based models
Unigram Language Model
- Starts with many candidates
- Removes tokens that reduce likelihood the least
- Used in SentencePiece
All of them aim to learn statistically meaningful subwords.
Tokens ≠ Words (Important!)
This is one of the most common misconceptions.
Example: "I love machine learning"
This sentence might be:
• 4 words
• 6–8 tokens (depending on tokenizer)
Even more surprising:
- Whitespace is often part of the token
- Punctuation usually gets its own token
- Emojis and symbols can be single or multiple tokens
- Different languages tokenize very differently
Token IDs → Embeddings
Once text is tokenized:
- Each token ID is mapped to an embedding vector
- Embeddings are dense numerical representations (e.g. 768–4096 dimensions)
- These vectors are the true "input" to the transformer
Conceptually:
Text → Tokens → Token IDs → Embeddings → Transformer Layers → Output Tokens
Context Length and Token Limits
LLMs have a maximum context window, measured in tokens (not characters or words).
For example:
- 4K tokens
- 8K tokens
- 32K+ tokens
This limit includes:
- System instructions
- User prompt
- Conversation history
- Generated output
That's why:
- Long prompts can silently truncate
- Token-efficient prompts matter
- Compression and summarization are critical in production systems
Why Tokenization Affects Cost and Performance
In most LLM APIs:
- You pay per token
- Latency scales with token count
- Memory usage grows with sequence length
Two prompts with the same meaning but different phrasing can have very different token costs.
Example:
"Summarize the following text"
vs
"Provide a concise summary of the text below"
Different tokens → different cost → different speed.
Generation Is Also Token-Based
LLMs don't generate words.
They generate one token at a time.
At each step:
- The model predicts a probability distribution over all possible next tokens
- A sampling strategy is applied (greedy, temperature, top-k, top-p)
- One token is selected and appended
- The process repeats
This continues until:
- An end-of-sequence token is generated
- A token limit is reached
Final Takeaway
Tokenization is not a preprocessing detail — it is foundational.
If you understand tokenization, you:
- Understand LLM limitations
- Design better prompts
- Optimize cost and latency
- Avoid context-related bugs
- Build more reliable AI systems
LLMs don't read text.
They read tokens — and tokens shape everything.
Practical Examples: Working with Tokenizers in Python
Below are ready-to-run Python examples demonstrating how to inspect tokens, token IDs, and counts using popular tokenizer libraries. These examples will help you understand how different tokenizers handle the same text.
Example 1: GPT-style BPE with tiktoken
The tiktoken library provides fast tokenization for OpenAI's GPT models. This example shows how to encode text into tokens and inspect individual token IDs.
pip install tiktoken
import tiktoken
text = "Tokenization is fascinating! "
# "cl100k_base" is a common modern encoding
enc = tiktoken.get_encoding("cl100k_base")
token_ids = enc.encode(text)
tokens = [enc.decode([tid]) for tid in token_ids]
print("TEXT:", text)
print("TOKENS COUNT:", len(token_ids))
print("TOKEN IDS:", token_ids)
print("TOKENS (decoded one-by-one):")
for i, t in enumerate(tokens):
print(f"{i:02d}: {repr(t)}")
print("\nDECODED BACK:", enc.decode(token_ids))
Example 2: Hugging Face Transformers Tokenizers
The transformers library from Hugging Face provides access to tokenizers from various pre-trained models. You can compare how different models tokenize the same text.
pip install transformers
GPT-2 BPE Tokenizer
GPT-2 uses Byte Pair Encoding (BPE), which merges frequent character pairs to create subword tokens.
from transformers import AutoTokenizer
text = "Tokenization is fascinating! "
tok = AutoTokenizer.from_pretrained("gpt2") # BPE
enc = tok(text, add_special_tokens=False)
print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKEN IDS:", enc["input_ids"])
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))
BERT WordPiece Tokenizer
BERT uses WordPiece tokenization, which represents continuation pieces with a ## prefix. Notice how the same text gets tokenized differently.
from transformers import AutoTokenizer
text = "Tokenization is fascinating! "
tok = AutoTokenizer.from_pretrained("bert-base-uncased") # WordPiece
enc = tok(text, add_special_tokens=False)
print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKEN IDS:", enc["input_ids"])
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))
Example 3: SentencePiece (T5 Model)
SentencePiece is language-agnostic and doesn't rely on whitespace pre-tokenization, making it ideal for multilingual models like T5.
from transformers import AutoTokenizer
text = "Tokenization is fascinating! "
tok = AutoTokenizer.from_pretrained("t5-small") # SentencePiece
enc = tok(text, add_special_tokens=False)
print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))
Example 4: Comparing Tokenizers Side-by-Side
This example compares how GPT-2, BERT, and T5 tokenize the same text. You'll see different token counts and different ways of splitting the same words.
from transformers import AutoTokenizer
text = "Tokenization is fascinating! "
models = ["gpt2", "bert-base-uncased", "t5-small"]
for m in models:
tok = AutoTokenizer.from_pretrained(m)
ids = tok(text, add_special_tokens=False)["input_ids"]
toks = tok.convert_ids_to_tokens(ids)
print("\n===", m, "===")
print("count:", len(ids))
print("tokens:", toks)
Example 5: Estimating Token Count for API Usage
Before sending prompts to LLM APIs, you can estimate token counts to optimize costs and stay within context limits.
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")
def count_tokens(s: str) -> int:
return len(enc.encode(s))
prompt = "Write a detailed explanation of tokenization in LLMs..."
print("Estimated tokens:", count_tokens(prompt))
Note: Exact token counts depend on the tokenizer and whether special tokens (chat templates, BOS/EOS, etc.) are added. For production use, always measure tokens with the same tokenizer that your target model uses.
Back to Home