HOWAI

Learn, Explore, and Master Artificial Intelligence

How Tokenization Works in Large Language Models

When we interact with a Large Language Model (LLM), it feels natural: we write text, the model "understands" it, and generates meaningful responses.

But under the hood, LLMs never see raw text.

They operate entirely on tokens.

Understanding tokenization is crucial if you want to:

Let's break it down step by step.

What Is Tokenization?

Tokenization is the process of converting raw text into a sequence of tokens, which are numerical units that the model can process.

A token is not necessarily:

Instead, a token is a statistical sub-unit of text, learned from massive corpora during model training.

Example:

Input text: "Tokenization is fascinating!"

Possible tokenization: ["Token", "ization", " is", " fascinating", "!"]

Each token is then mapped to an integer ID, which is what the model actually consumes.

Why LLMs Can't Use Words Directly

Neural networks work with numbers, not symbols.

To process language, an LLM needs:

Tokenization solves all three problems:

Subword Tokenization: The Core Idea

Modern LLMs use subword tokenization, not word-level tokenization.

Why?

Word-level tokenization problems

Character-level tokenization problems

Subword tokenization advantages

Common Tokenization Algorithms

Byte Pair Encoding (BPE)

Example: "lowest" → "low" + "est"

WordPiece

Unigram Language Model

All of them aim to learn statistically meaningful subwords.

Tokens ≠ Words (Important!)

This is one of the most common misconceptions.

Example: "I love machine learning"

This sentence might be:

• 4 words
• 6–8 tokens (depending on tokenizer)

Even more surprising:

Token IDs → Embeddings

Once text is tokenized:

Conceptually:

Text → Tokens → Token IDs → Embeddings → Transformer Layers → Output Tokens

Context Length and Token Limits

LLMs have a maximum context window, measured in tokens (not characters or words).

For example:

This limit includes:

That's why:

Why Tokenization Affects Cost and Performance

In most LLM APIs:

Two prompts with the same meaning but different phrasing can have very different token costs.

Example:

"Summarize the following text"
vs
"Provide a concise summary of the text below"


Different tokens → different cost → different speed.

Generation Is Also Token-Based

LLMs don't generate words.
They generate one token at a time.

At each step:

This continues until:

Final Takeaway

Tokenization is not a preprocessing detail — it is foundational.

If you understand tokenization, you:

LLMs don't read text.
They read tokens — and tokens shape everything.

Practical Examples: Working with Tokenizers in Python

Below are ready-to-run Python examples demonstrating how to inspect tokens, token IDs, and counts using popular tokenizer libraries. These examples will help you understand how different tokenizers handle the same text.

Example 1: GPT-style BPE with tiktoken

The tiktoken library provides fast tokenization for OpenAI's GPT models. This example shows how to encode text into tokens and inspect individual token IDs.

pip install tiktoken
import tiktoken

text = "Tokenization is fascinating! "

# "cl100k_base" is a common modern encoding
enc = tiktoken.get_encoding("cl100k_base")

token_ids = enc.encode(text)
tokens = [enc.decode([tid]) for tid in token_ids]

print("TEXT:", text)
print("TOKENS COUNT:", len(token_ids))
print("TOKEN IDS:", token_ids)
print("TOKENS (decoded one-by-one):")
for i, t in enumerate(tokens):
    print(f"{i:02d}: {repr(t)}")

print("\nDECODED BACK:", enc.decode(token_ids))

Example 2: Hugging Face Transformers Tokenizers

The transformers library from Hugging Face provides access to tokenizers from various pre-trained models. You can compare how different models tokenize the same text.

pip install transformers

GPT-2 BPE Tokenizer

GPT-2 uses Byte Pair Encoding (BPE), which merges frequent character pairs to create subword tokens.

from transformers import AutoTokenizer

text = "Tokenization is fascinating! "

tok = AutoTokenizer.from_pretrained("gpt2")  # BPE
enc = tok(text, add_special_tokens=False)

print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKEN IDS:", enc["input_ids"])
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))

BERT WordPiece Tokenizer

BERT uses WordPiece tokenization, which represents continuation pieces with a ## prefix. Notice how the same text gets tokenized differently.

from transformers import AutoTokenizer

text = "Tokenization is fascinating! "

tok = AutoTokenizer.from_pretrained("bert-base-uncased")  # WordPiece
enc = tok(text, add_special_tokens=False)

print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKEN IDS:", enc["input_ids"])
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))

Example 3: SentencePiece (T5 Model)

SentencePiece is language-agnostic and doesn't rely on whitespace pre-tokenization, making it ideal for multilingual models like T5.

from transformers import AutoTokenizer

text = "Tokenization is fascinating! "

tok = AutoTokenizer.from_pretrained("t5-small")  # SentencePiece
enc = tok(text, add_special_tokens=False)

print("TOKENS COUNT:", len(enc["input_ids"]))
print("TOKENS:", tok.convert_ids_to_tokens(enc["input_ids"]))
print("DECODED:", tok.decode(enc["input_ids"]))

Example 4: Comparing Tokenizers Side-by-Side

This example compares how GPT-2, BERT, and T5 tokenize the same text. You'll see different token counts and different ways of splitting the same words.

from transformers import AutoTokenizer

text = "Tokenization is fascinating! "

models = ["gpt2", "bert-base-uncased", "t5-small"]
for m in models:
    tok = AutoTokenizer.from_pretrained(m)
    ids = tok(text, add_special_tokens=False)["input_ids"]
    toks = tok.convert_ids_to_tokens(ids)
    print("\n===", m, "===")
    print("count:", len(ids))
    print("tokens:", toks)

Example 5: Estimating Token Count for API Usage

Before sending prompts to LLM APIs, you can estimate token counts to optimize costs and stay within context limits.

import tiktoken

enc = tiktoken.get_encoding("cl100k_base")

def count_tokens(s: str) -> int:
    return len(enc.encode(s))

prompt = "Write a detailed explanation of tokenization in LLMs..."
print("Estimated tokens:", count_tokens(prompt))

Note: Exact token counts depend on the tokenizer and whether special tokens (chat templates, BOS/EOS, etc.) are added. For production use, always measure tokens with the same tokenizer that your target model uses.

Back to Home