Training is where a language model learns, but inference is where all that learning is put into action. Inference is the moment when you type a question, press enter, and the model starts producing an answer. Nothing new is being learned here. No weights are updated. The model is simply using what it already knows to predict what should come next, one step at a time.
Although the experience feels instantaneous and conversational, under the hood inference is a precise and rigid pipeline, repeated over and over for every generated token.
From Prompt to Tokens to Embeddings
Everything starts with your prompt. The text you write is not processed as words or sentences, but as tokens. These tokens might correspond to whole words, parts of words, punctuation, or even whitespace. This step already shapes the model's behavior, because the model does not "see" text the way humans do, but as sequences of discrete symbols.
Once tokenized, each token is mapped to an embedding, which is a dense numerical vector. This is the point where raw text becomes something the neural network can work with. Embeddings capture latent properties of tokens: semantic similarity, syntactic role, stylistic cues. Words that are used in similar contexts tend to have embeddings that are close to each other in this space.
The transformation: At this stage, the prompt has been transformed from human-readable language into a sequence of vectors that represent meaning in a statistical sense.
The Forward Pass: Turning Context into Predictions
With embeddings ready, the model performs a forward pass through its layers. This is where self-attention, feed-forward networks, and residual connections all come into play. Each layer refines the representation of every token by mixing its information with that of the surrounding tokens.
Crucially, the model only attends to tokens that come before the current position. This causal structure ensures that the model is always predicting the next token based solely on past context, never future information.
By the time the data reaches the final layer, the model has distilled the entire prompt into a single representation that answers one question: given everything so far, what should the next token be? The output is not text, but a vector of scores—one score for every possible token in the vocabulary.
Sampling: From Scores to an Actual Choice
Those scores are converted into probabilities, forming a distribution over possible next tokens. This is where sampling strategies come into play, and where much of the model's "personality" seems to emerge.
Temperature controls how sharp or flat the probability distribution is. With a low temperature, the model becomes conservative, strongly favoring the most likely tokens and producing deterministic, often repetitive text. With a higher temperature, probabilities are flattened, allowing less likely tokens to be chosen and resulting in more creative but riskier outputs.
Top-k sampling restricts the model's choices to the k most probable tokens, cutting off the long tail of unlikely options. This prevents bizarre or incoherent continuations while still allowing some variability.
Top-p sampling, also known as nucleus sampling, works slightly differently. Instead of fixing the number of tokens, it selects the smallest set of tokens whose cumulative probability exceeds a given threshold. This adapts dynamically to the uncertainty of the model: confident situations lead to fewer choices, ambiguous ones to more.
Key insight: All of these mechanisms exist for one reason—the model does not produce text directly. It produces probabilities, and sampling is how those probabilities are turned into an actual decision.
Generation, One Token at a Time
Once a token is selected, it is appended to the sequence. Then everything happens again. The new token becomes part of the context, embeddings are updated, another forward pass is performed, and a new probability distribution is computed.
This autoregressive loop continues token by token until a stopping condition is reached, such as an end-of-sequence token or a maximum length limit. What feels like a coherent paragraph is actually the result of dozens or hundreds of independent next-token predictions, each conditioned on everything generated so far.
This also explains why early tokens matter so much. A small change at the beginning can cascade through the entire generation, pushing the model onto a completely different trajectory.
Cascade effect: Early tokens shape the entire generation. The first few predictions constrain all subsequent ones.
The Illusion of Interaction
From the outside, inference feels interactive and responsive. From the inside, it is mechanical and local. The model does not know it is answering a question. It does not plan an answer in advance. It does not revise earlier tokens. It simply extends text in the most statistically consistent way it can, step after step.
And yet, because language itself encodes structure, intention, and reasoning, this simple loop is enough to produce explanations, arguments, and dialogue that feel purposeful. Inference is where the gap between human intuition and machine reality is most visible: what looks like conversation is, at its core, controlled probability unfolding over time.
The fundamental reality: What feels like conversation is controlled probability unfolding token by token. Inference is where statistical prediction becomes perceived intelligence.