Watch it above, or on the episode page with chapters and every source: engineering.fm/episodes/02-llm-one-token
The 2-minute version
One token at a time. Before the first word of a reply appears, the model makes exactly one decision: the next token. Then it appends it and runs the whole machine again.
Tokens aren’t words. GPT-2’s list has 50,257 entries: 256 raw bytes, 50,000 learned merges and one end-of-text marker. “Tokenization” alone takes two.
Each token becomes a vector. Token 3797 (” cat”) picks out 768 learned numbers, which then climb a stack of 12 layers.
Attention is a lookup. A query meets every earlier key by dot product, softmax turns the scores into weights that add up to one, and the token takes that blend of values. It can only look back.
The output is 50,257 scores. For our prompt every one is below −80. Only the gaps matter, and softmax turns them into probabilities.
Temperature divides the scores first. Toy scores 2, 1 and 0 give 67%, 24% and 9%. At temperature 0.5: 87, 12 and under 2. At 2: 51, 31 and 19.
Then one weighted roll. GPT-2’s top bars for “The cat sat on the” are floor 7.6%, bed 6.5% and couch 5.4%. Trim the tail, roll the dice: “floor”.
The question
You type a prompt and hit Enter. Every word you then read is a separate trip through the whole model. So what happens on one trip?
The analogy: the suggestion bar
Your phone’s keyboard runs a tiny version of this machine: a few guesses for the next word. A language model plays the same game at a ridiculous scale, scoring every token it knows, every single time, then picking one, adding it to the end, and playing again.
We follow one pick through GPT-2, a small, fully open model from 2019, so every number below is one we measured ourselves on a laptop.
Zoom 1: the whole loop
Start with “The cat sat on the”. A tokenizer chops it into pieces from a fixed list: five tokens, one per word, each carrying its leading space. But “Tokenization isn’t magic.” comes out as six pieces, counting the full stop. GPT-2’s list has 50,257 entries; Llama 3’s has about 128,000.
To the model, each token is just an id (” cat” is 3797), and that id picks out a learned vector of 768 numbers: the token’s starting meaning. The vectors climb a stack of layers (12 in this GPT-2, 126 in the biggest Llama 3), and each layer refines each vector using the ones before it.
At the top, the last position’s vector becomes a score for every token in the list. Pick one, append it, run the stack again, until the model picks end-of-text or hits a length limit. On a laptop, this GPT-2 makes one trip in about a hundredth of a second.
Zoom 2: inside one layer
Each layer does two things: attention, where tokens look at each other, then a feed-forward block that works on each token on its own.
Attention is a lookup. Each token makes three vectors: a query (”what am I looking for?”), a key (”what do I contain?”) and a value (”what I’ll hand over”). The last token compares its query with every key so far using a dot product: multiply the numbers in pairs and add them up. Softmax squashes those scores into weights that add up to one, and the token takes that blend of values. Later positions are masked out, so no token can peek at the future.
You can measure this. In GPT-2’s fifth layer, one head sends 57% of the last token’s attention straight to “cat”. Across the model, 98 of 144 heads put more than half their attention on the very first token. Researchers call that an attention sink: the weights must add up to one, so spare attention has to go somewhere.
Nobody wrote these rules by hand. All of this GPT-2’s roughly 124 million weights were learned, from about 40 GB of web text. (The biggest Llama 3 has 405 billion.) And because earlier tokens’ keys and values never change, they’re cached: each new trip only computes the newest token.
Zoom 3: the pick
The top vector becomes 50,257 raw scores called logits. They aren’t probabilities: for our prompt, every one sits between about −111 and −81. That’s fine, because only the gaps matter.
Softmax turns gaps into probabilities: raise e (about 2.7) to the power of each score, then divide by the total. Toy scores 2, 1 and 0 become about 7.4, 2.7 and 1, which total 11.1: so 67%, 24% and 9%.
Temperature divides every score first. At 0.5 the gaps double (87, 12, under 2); at 2 they halve (51, 31, 19). Near zero, everything piles onto the top bar: that’s greedy decoding.
GPT-2’s real bars for our prompt: floor 7.6%, bed 6.5%, couch 5.4%, then ground, edge and bench. The top eight hold under 40%. So samplers cut the long tail first: keep the top k (the GPT-2 paper used 40), or the smallest set that adds up to, say, 90%. That’s top-p, and here it still leaves 728 tokens. Then one weighted dice roll. This time: “floor”.
And the top bar isn’t a fact. Ask this GPT-2 to finish “The capital of France is” and its favourite is “the”; Paris comes fifth. Try “…is the city of” and the first piece of “Marseille” edges out Paris, 12.9% to 11.1%. It isn’t looking anything up. It’s scoring what usually comes next.
Recap
Text becomes tokens, tokens become vectors, and the vectors climb the stack, looking back at each other in every layer. The last one becomes a score for every token. Softmax makes them probabilities, temperature shapes them, the tail gets trimmed, and one weighted roll picks the winner. Then it’s appended, and the trip starts again, for every single token you read.
Try it yourself
Our numbers come from GPT-2 small (openai-community/gpt2) on a laptop CPU (I have a base Apple M3 Max with 36GB of ram), with Hugging Face Transformers 5.18.0 and PyTorch 2.14.1, eager attention, no sampling. The core:
from transformers import GPT2TokenizerFast, GPT2LMHeadModel
tok = GPT2TokenizerFast.from_pretrained("gpt2")
m = GPT2LMHeadModel.from_pretrained("gpt2", attn_implementation="eager").eval()
ids = tok("The cat sat on the", return_tensors="pt").input_ids # [464, 3797, 3332, 319, 262]
out = m(ids, output_attentions=True)
p = out.logits[0, -1].softmax(-1) # one probability per token
for prob, i in zip(*p.topk(8)):
print(repr(tok.decode(int(i))), f"{prob.item():.2%}")
You should see “ floor” on top at about 7.6%. The same pass holds the raw logits (out.logits[0, -1]) and the attention weights (out.attentions, one tensor per layer), so you can find the head that looks at “cat”: layer index 4, head index 3, counting from zero.
Your turn
What’s the strangest top pick you’ve seen a model make? Tell us in the comments.
Sources
Radford et al. (2019): Language Models are Unsupervised Multitask Learners (GPT-2)
Hugging Face: GPT-2 small, the model behind every measured number in this episode
Xiao et al. (2023): Efficient Streaming Language Models with Attention Sinks
Hugging Face Transformers docs: How caching works (KV cache)
Holtzman et al. (2019): The Curious Case of Neural Text Degeneration (top-k, top-p, temperature)
Hugging Face Transformers docs: Logits processors (temperature)
Hugging Face Transformers docs: Generation strategies (greedy, sampling)
CNBC: OpenAI announces rollout of GPT-6 Astra model (2026-09-03)
That’s engineering.fm. One machine, taken apart, every episode.



