Models · reviewed · reviewed Oct 5, 2026 · 4 min
How does a large language model work, end to end?
Follow text into a model, watch it choose the next token, and see what changes when it learns.
One token at a time
Ask a language model to finish “The cat sat on the …”. It produces scores for possible next tokens. A decoding rule picks one; that token becomes part of the input for the next prediction. Repeat, and a sentence takes shape.
This tutorial follows a decoder-only, autoregressive text model, the family behind GPT-style generation. Other language models can use different architectures and objectives. The key idea is a learned network that turns the supplied context into predictions. GPT-3 paper.

An artistic metaphor for the model and its feedback loop. The interactive diagrams below explain the actual relationships.
First, text becomes numbers
A tokenizer splits text into vocabulary pieces and assigns each an ID. A token might be a whole word, part of a word, punctuation, or another piece of encoded text. Different tokenizers split the same sentence differently.
An embedding maps each ID to a learned list of numbers: a vector. It gives the network a numerical starting point for processing that token. Position information supplies the order of the sequence. The numbers are learned features, rather than a dictionary of human-readable meanings. The Illustrated GPT-2.
The path is text → token IDs → vectors → transformer layers → next-token scores. Those scores become a probability distribution. A decoder chooses from it, then feeds the token back into the sequence. Explore the loop before looking inside a layer.
Watch a sentence grow
Keep your eye on the input: a selected token becomes part of the next prediction.
Current context · 5 toy tokens
Predict the next token
The model scores possible continuations from the current context. The answer is still a distribution.
Selected token: none yet
Next-token probabilities
mat69.8%sofa23.2%grass7.0%
A local simulation with a tiny word-level vocabulary and hand-authored scores. Temperature reshapes the same scores; it does not make facts more accurate. Samples A/B/C use fixed draws (0.2/0.6/0.9) for reproducibility. No LLM is called. Explore decoding →
Context changes the prediction
“The cat sat on the …” and “The camper sat on the …” have different plausible endings. The model’s weights can stay identical while a change in input changes its predictions. Instructions, earlier messages, and examples also become part of that input. This is how in-context learning can alter behavior without a training update. GPT-3, approach.
Inside a transformer, attention lets a position combine information from other permitted positions. Its scores control a weighted mix of vectors. In a causal text model, a position can use earlier tokens and itself, but cannot look ahead at the token it is supposed to predict. Attention Is All You Need, §3.2.
What can this position see?
Choose a token. Earlier positions can contribute information; later positions are hidden.
Information available to “it”
- The8.5%Earlier context
- robot56.9%Earlier context
- carried11.5%Earlier context
- it23.1%This position
Here, “it” receives the largest contribution from “robot”. The resulting vector combines information from several positions.
Invented scores for one attention head, normalized over visible positions. Real models have many heads and layers. Attention weights describe a weighted mix of vectors; they are not a complete explanation of reasoning. Explore transformers →
Attention is one part of the network
Each transformer layer also contains a feed-forward network that transforms the information at each position. Residual connections carry information between layers, and normalization helps keep the computations well behaved. Repeated layers build richer representations before the final projection produces vocabulary scores. Transformer architecture, §3.
You do not need to multiply matrices to follow the mechanism: start with numerical token representations, mix relevant context, transform those representations, then score continuations. For a closer look, visit embeddings or transformers and attention.
Where the weights come from
During pretraining, the next token is already present in the training text. The model predicts it; a loss measures prediction error; backpropagation calculates how parameters contributed to that error; an optimizer adjusts them. Many examples and updates teach useful patterns. Stanford’s Language Modeling from Scratch connects these stages, from data and tokenization to training and inference.
During ordinary inference, that update step is absent. The context grows while the learned weights stay fixed. Adding a fact to a conversation can influence the next reply, but does not by itself train the model. Try the difference with a predictor small enough to watch.
Change the weights, or use them
A tiny predictor knows two continuations of “The sky is”. Teach it an example, then use the same parameter to generate.
Learned parameter: 0.000
0 training updates · 0 generated samples
Prediction error for “blue”: 0.693Lower means more probability on the observed token.
Current prediction for “The sky is …”
blue50.0%grey50.0%
What this tiny predictor actually computes
Two scores, [w, 0], become probabilities through softmax. A training click applies gradient descent to next-token cross-entropy: w ← w − 0.8 × (P(blue) − target), where target is 1 for blue and 0 for grey. Generation alternates fixed draws 0.2 and 0.9. This is a real scalar learning update, not a transformer or a factual model of the sky.
This isolates the difference between learning and using a learned model. A real LLM updates many parameters across batches; ordinary inference keeps its learned parameters fixed.
A useful answer still needs evidence
Pretraining teaches broad patterns. Later post-training can teach instruction following and preferred responses using demonstrations, feedback, and other learning objectives. Those stages change parameters; writing a new prompt supplies a new input. See training versus inference and fine-tuning and preference training.
Next-token probability is not a fact-checking score. A fluent continuation can contain an invented detail. Important claims need appropriate evidence and verification; adding an authoritative source to context gives the model information to use, rather than a guarantee that it uses it correctly.
The next loop lives in the harness
Generating the text of a function call does not execute the function. The surrounding harness supplies context and tools, checks the request, runs the permitted operation, and returns its result for another model call.
That gives us two connected loops: inside a model call, tokens extend the sequence; around model calls, a harness connects proposals with actions and observations. Continue with How a harness works to see this boundary in action.
For a longer visual treatment, Hands-On Large Language Models by Jay Alammar and Maarten Grootendorst offers chapters on tokens, embeddings, and transformer internals. Its official example repository accompanies the book. This tutorial uses original illustrations and deliberately small teaching models.
Sources
Sources and further reading
- 01Language Models are Few-Shot LearnersBrown et al. · research · published May 28, 2020 · source checked Aug 30, 2026
Primary source for autoregressive GPT-style language modelling and in-context task specification.
- 02Attention Is All You NeedVaswani et al. · research · published Jun 12, 2017 · source checked Aug 30, 2026
Primary architecture source for transformer attention, feed-forward layers, residual connections, and positional information.
- 03Language Modeling from ScratchStanford University · guide · source checked Aug 30, 2026
A current engineering map from tokenizer and transformer construction through training, scaling, and inference.
- 04The Illustrated GPT-2Jay Alammar · guide · source checked Oct 5, 2026
Author-created visual explanation of decoder-only generation, token embeddings, causal attention, vocabulary scores, and feeding selected tokens back into context.
- 05Hands-On Large Language Models — official companionJay Alammar and Maarten Grootendorst · guide · source checked Oct 5, 2026
Official companion repository and public chapter map for the book, including tokens and embeddings and transformer internals. A further-reading path; the full book is not reproduced here.
