When a chatbot answers a question, it may seem as though the entire answer is already prepared. In reality, the process is simpler and more constrained. At each step, an autoregressive language model receives the text that already exists and computes a probability distribution for the next fragment, or token. Hugging Face documentation calls this mode causal language modeling: the model predicts the next token and sees only what is to its left.
First, the scope of this article. It concerns the model itself, not agents, internet search, or corporate systems built around it. The description primarily applies to a GPT-like model with causal attention. Not every neural network called a language model works this way: in the same documentation, the causal variant is distinguished from masked language modeling, in which the model reconstructs hidden fragments while seeing context on both sides.
The Full Route
Before examining the details, it is useful to see the entire path:
text → tokens → IDs → token vectors + positional information
→ repeated Transformer blocks
→ next-token scores and probabilities
→ token selection → append to context → repeat
This is an instructional diagram. It leaves out differences between implementations and later training stages, so it cannot be treated as a specification for every modern language model. The diagram is based on two works: the paper Attention Is All You Need, which describes the Transformer, and OpenAI’s paper on generative pre-training, which uses a single-decoder variant. An important qualification: the original Transformer was an encoder-decoder architecture for transforming one sequence into another. It did have an encoder; GPT-like models do without one, and from here on we follow their route.
Step 1. Text Becomes Numbers
A neural network does not work with letters directly. First, a tokenizer splits the text into fragments and assigns each one a numerical identifier (ID). The OpenAI guide How to count tokens with Tiktoken contains an example that is convenient to follow from start to finish. The recipe itself is marked as archived and may contain outdated information, but the tokenization shown in it remains a useful illustration. In the r50k_base encoding, the string “2 + 2 = 4” is split as follows:
| Byte fragment | ID |
|---|---|
| b'2' | 17 |
| b' +' | 1343 |
| b' 2' | 362 |
| b' =' | 796 |
| b' 4' | 604 |
The total is IDs [17, 1343, 362, 796, 604] and fragments [b'2', b' +', b' 2', b' =', b' 4']. The space is part of the token here: “ 2” with a space and “2” without one receive different numbers. This leads to the main point about tokens: a token does not necessarily correspond to a word, and its boundaries and number depend on the encoding. In the same guide, another encoding splits the same string differently. Therefore, numerical IDs mean nothing without the name of the encoding, and they cannot be transferred to another model.
For the rest of the example, we will take the first four tokens, that is, the input “2 + 2 =” with IDs [17, 1343, 362, 796]. We will consider the token “ 4” with ID 604 only as a possible continuation. This illustrates how the task works; it is not a measured prediction by any particular model.
Step 2. From IDs to Vectors: The Model Must Account for Position
An ID says nothing about meaning by itself. It is an address in a table in which each token corresponds to a trainable vector, a list of numbers whose values are selected during training. A card analogy is useful: the card number is not the card’s content; it merely helps locate it.
However, token vectors alone are not enough. Compare “2 + 2 =” and “= 2 + 2”. The symbols are the same, but the rearrangement also changes the tokens themselves: the equals sign at the beginning of the string has no leading space and is therefore encoded not as the token b' =' with ID 796, while the first “2”, after a space, becomes the token b' 2' (362). But even where the set of cards is the same, the vector table cannot distinguish order: the same ID always receives the same vector, regardless of the position where it appears. For computations to account for order, the architecture needs a way to encode or account for the position of a token. In the original Transformer, positional encodings were added to the token representations; in the OpenAI GPT model described, learned positional representations were used. These are specific options: other architectures may account for position differently, so a positional vector is not always added to a token vector.
We will return to the card analogy below, but with an advance qualification: it helps illustrate the flow of data and says nothing about thinking. The model does not “read” cards; it performs matrix calculations on vectors.
Step 3. Attention: How Positions Exchange Information
At the center of the Transformer is the attention mechanism. Using trainable projections, three matrices are computed from the current representations: queries Q, keys K, and values V. Then, according to the original paper, the following is calculated:
Attention(Q,K,V) = softmax(QKᵀ/√d_k)V
Notation:
- Q — the query matrix;
- K — the key matrix, Kᵀ — the transposed K;
- V — the value matrix;
- d_k — the key dimensionality;
- softmax — a transformation of scores into nonnegative weights whose sum is one.
The meaning of the formula is as follows. The product QKᵀ gives a score for how much each position “corresponds” to every other position; division by √d_k scales these scores; softmax turns them into weights; finally, the weights mix the values V. In terms of the cards, for the current position, attention temporarily weights information from other cards and assembles a new representation from it.
Now for the causal mask. It is needed so that a position cannot use future tokens. The mask is applied to the scores before softmax: scores for prohibited future positions are replaced, in the original paper, with −∞. Since exp(−∞) = 0, the weights of these positions become zero after softmax, while the weights of permitted positions remain positive. Thus, when processing “2 + 2 =”, the position of the “ +” token receives no information about tokens to its right.
The formula describes one attention mechanism. Multi-head attention repeats the same calculation in parallel with different trainable projections, and each “head” may weight positions in its own way.
One more important qualification. Attention is a calculation of weights for mixing representations, not human reasoning. It is tempting to regard large weights as an explanation of an answer, but the study Attention is not Explanation showed that attention weights alone do not provide a reliable explanation of why a model produced a particular result. This conclusion applies to the standard modules and tasks examined there; it is not a theorem about all architectures, but it is a good reason for caution.
Step 4. The Transformer Block and Its Repetition
Attention is only part of the block. In the original architecture, it is followed by a feed-forward network (MLP), which is applied separately to each position. Residual connections surround these components: the input to a sublayer is added to its output. Layer normalization is also used (LayerNorm): for each token, the mean and variance of its features are calculated, the values are normalized relative to these statistics, and they are then usually transformed using trainable scale and shift parameters. Normalization changes the distribution of activations, but does not keep them within a specified range.
Such blocks are placed one after another, with the output of one becoming the input of the next. The order of normalization and the specific function inside the MLP vary between implementations, so the diagram from the original paper should not be presented as the only possible one.
Step 5. From Probabilities to the Selected Token
After the final block, the representation of the last position is transformed into scores for every token in the vocabulary, and softmax turns them into probabilities. This is done both in the original Transformer and in GPT. During generation, only the scores for the last position are needed; during autoregressive pre-training, discussed below, they are computed for all positions at once and compared with the same sequence shifted by one token. In our example, the output would be a distribution over all tokens in the encoding, and “ 4” with ID 604 would be one of the candidates.
But a distribution is not yet an answer. Which token to take is determined by a separate decoding rule, described in the Hugging Face guide to generation strategies:
- greedy decoding selects the token with the highest probability;
- sampling selects a token randomly according to the distribution.
The practical consequence is that the most probable token is not always selected, and the same prompt can produce different continuations when sampling is used. The selected token is appended to the end of the sequence, and the entire calculation is repeated for the next position. The answer is built from this chain of steps, one token at a time.
Training and Inference: When the Weights Change
All trainable vectors, projections, and MLP weights are collectively called parameters. They are selected during training. During autoregressive pre-training, as described in the paper on GPT, the goal is to predict the next token. However, this is not the only stage: in the same paper, supervised fine-tuning follows pre-training, with a target task that depends on the specific assignment, and later fine-tuning stages in general may use other objectives. The parameter-update cycle looks like the one shown in the PyTorch tutorial Optimizing Model Parameters:
- the model makes predictions;
- they are compared with target values (during pre-training, with the next tokens), and a loss function is calculated;
- backpropagation produces gradients;
- the optimizer changes the parameters.
During ordinary inference, when the model answers a prompt, only a forward pass with the already trained weights is performed, without an update step. If you rewrite the prompt or change the generation settings, the answer will change, but the weights will remain the same. In this sense, the model does not “learn” from your conversation; this discussion concerns standard inference without fine-tuning during the request.
Practical Limits
Context window. The number of tokens that a model uses in one request is limited. In the OpenAI guide Conversation state, for the API described, the limit includes input and output tokens and, for the corresponding models, reasoning tokens as well. The specific limit depends on the model and how it is used, and the rule for handling an excess cannot be generalized to all implementations. Most importantly, the window is a limit on the current request, not evidence of long-term memory.
Parameters Are Not a Database. Parameters can implicitly contain information, but, as the authors of the paper on Retrieval-Augmented Generation emphasize, they do not provide a reliable mechanism for addressable retrieval, provenance checking, and targeted updating of facts, as an explicit external store does. At the same time, “not a database” does not mean “does not memorize anything”: the study Extracting Training Data from Large Language Models managed to extract individual verbatim memorized training fragments from GPT-2.
Plausibility Is Not Truth. A model can confidently produce a smooth but false statement. The OpenAI article Why language models hallucinate notes that it is especially difficult to infer arbitrary rare facts from textual regularities. The source examines possible causes of errors, not the inevitability of every specific error, but the practical conclusion for users is direct: a confident tone does not guarantee accuracy, and factual answers should be checked independently.
Architectures Differ. The route traced above is a simplification. The original Transformer was an encoder-decoder, GPT was decoder-only, and modern models differ in how they account for position, how their blocks are structured, and what additional training stages they use. The diagram helps explain what happens to text between input and answer, but it does not replace the documentation for a specific model.