Modern Transformer models combine changes to architecture, training and inference. Attention determines how token representations interact; training shapes the model’s learned numerical values, called parameters; inference uses those parameters to generate an answer. DeepSeek provides selected examples of how these parts developed.
This article follows those examples from the original Transformer onwards. At each step, the question is: what changed, which problem did it address, and what does the evidence establish?
What you will learn
- Distinguish architecture, training and inference.
- Calculate a small attention output from a query, keys and values.
- Explain why attention compression and expert routing address different constraints.
- Read reasoning claims in relation to training and evaluation conditions.
The central idea
The Transformer evolved through changes to its structure, its training and its execution. Keeping those questions separate makes the development easier to follow, even where one design choice affects several of them.
| Question | What to inspect |
|---|---|
| Architecture: how does information move? | Layers, attention, connections and routing. |
| Training: how are parameters shaped? | Data, objectives, examples, rewards and optimisation. |
| Inference: how is an answer produced? | Active computation, cached representations and generation budget. |
2017: attention without sequence-aligned recurrence
Vaswani et al. introduced an encoder–decoder Transformer that removed recurrent and convolutional sequence processing. Attention relates positions directly, supporting greater parallelism during training. The decoder still generates output sequentially. Attention was already in use; the contribution was the architecture built around it. Vaswani et al. (2017), sections 1–3 (opens in a new tab).
From translation to causal language modelling
The original Transformer encodes an input sequence and decodes an output sequence. A decoder-only language model instead processes a growing token sequence with causal attention. At each position, causal self-attention can use the current and earlier token representations, while masking later positions. Its output contributes to predicting the next token. Pre-training teaches next-token prediction; generation repeatedly selects a token and appends it to the sequence.
GPT-3 is a documented example. Its few-shot evaluations supplied instructions and examples in the input without updating trained parameters. That is adaptation through context, distinct from further training. Brown et al. (2020), sections 1–2 (opens in a new tab).
2023: refining the architecture and training budget
LLaMA illustrates further changes within the family, including normalisation before sub-layers and a different way to represent token positions. It also considers training smaller models for longer to reduce the cost of serving them. The important development is the interaction between design, training budget and inference cost. Touvron et al. (2023), sections 1–2 (opens in a new tab).
A worked example
Scaled dot-product attention compares a query with keys, divides the scores by the square root of the key dimension, applies softmax, then combines values using those weights. Multiple heads use different learned projections; positional information supplies sequence order. Vaswani et al. (2017), sections 3.2–3.5 (opens in a new tab).
A query represents what the current position is matching against. Keys provide representations to compare with that query; values provide the information combined using the resulting weights. Real models learn the projections that produce these representations.
Use this invented query and three keys, labelled A, B and C in order:
query = [1, 0]
keys = [[1, 0], [0, 1], [1, 1]]
values = [10, 20, 40]
The dot products are 1, 0, 1. With two key dimensions, divide by sqrt(2) to obtain approximately 0.7071, 0, 0.7071.
Softmax exponentiates the scores and divides by their sum. The weights are approximately 0.4011, 0.1978, 0.4011. Using the rounded weights gives:
0.4011 × 10 + 0.1978 × 20 + 0.4011 × 40 = 24.011
Keeping the full precision until the final result gives 24.0111 to four decimal places. Here is the calculation in Python:
from math import exp, sqrt
query = [1.0, 0.0]
keys = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
values = [10.0, 20.0, 40.0]
scores = [
sum(q * k for q, k in zip(query, key)) / sqrt(len(query))
for key in keys
]
largest = max(scores)
exponentials = [exp(score - largest) for score in scores]
total = sum(exponentials)
weights = [value / total for value in exponentials]
output = sum(weight * value for weight, value in zip(weights, values))
print([round(weight, 4) for weight in weights])
print(round(output, 4))
Expected output:
[0.4011, 0.1978, 0.4011]
24.0111
Subtracting the largest score avoids unnecessarily large exponentials and leaves softmax unchanged. This example calculates one query's output with scalar values; real models use value vectors and learned projections. It omits masking and the surrounding network.
These weights describe how this attention operation combines values. They are not probabilities that the values are correct, and this calculation alone does not demonstrate reasoning.
Before running it, explain why the first and third positions receive equal weights even though their values differ.
2024–2025: from attention to DeepSeek’s efficiency choices
A cache stores information for reuse. In Transformer generation, a key–value cache retains earlier attention representations so they need not be calculated again for every new token. DeepSeek-V3 retains the Transformer framework and carries forward MLA and DeepSeekMoE from V2. Multi-head Latent Attention (MLA) compresses keys and values into a smaller learned internal representation, called a latent representation, with a separate positional component, reducing the cache retained during generation. Mixture-of-Experts (MoE) uses shared experts and selects routed experts for each token in its feed-forward layers, which transform each position’s representation after attention has combined information across positions. DeepSeek-AI (2025a), section 2 (opens in a new tab).
These choices answer different questions: how much attention information must be stored, and which parameters must perform computation. Sparse activation does not mean the remaining parameters disappear from the model or its storage requirements.
A common misconception
Expert routing, reasoning training and additional computation while answering are different mechanisms. Selecting an expert is a choice about which parameters execute. Reward-based training changes parameters. Generating a longer intermediate sequence spends more computation at inference; it does not itself retrain the model.
Neither a long explanation nor an architectural label establishes correctness. Check the answer against evidence appropriate to the task.
Pre-training: adjusting parameters through prediction
DeepSeek-V3's pre-training shapes its parameters using a large text corpus and prediction objectives, including multi-token prediction. It precedes the report's supervised fine-tuning and reinforcement-learning stages. DeepSeek-AI (2025a), sections 2.2, 4–5 (opens in a new tab).
2025: R1 changes the training story
R1-Zero starts from the pretrained DeepSeek-V3-Base and applies reinforcement learning without preliminary supervised fine-tuning. That distinction matters: it is not a model learning from an untrained starting point. The authors report improved reasoning-task performance alongside readability and language-mixing problems. R1 adds cold-start examples and a pipeline with supervised and reinforcement-learning stages. Its smaller distilled models use supervised fine-tuning on generated examples. DeepSeek-AI (2025b), sections 2.2–2.4 (opens in a new tab).
A small reward-training illustration
Consider the invented task: “Solve 3x + 2 = 11.”
| Candidate response | Final-answer reward |
|---|---|
Subtract 2, then divide by 3: x = 3. |
1 |
Divide 11 by 3: x ≈ 3.67. |
0 |
Guess x = 3, without valid working. |
1 |
A simplified training cycle generates candidates, checks their answers, assigns rewards and uses those rewards to update parameters. A final-answer check can reward a correct guess as well as valid working.
With parameters fixed, generating candidates and selecting by agreement is inference. It makes no parameter update.
This invented example illustrates the distinction; it does not reproduce GRPO or DeepSeek’s full reward scheme. DeepSeek-AI (2025b), sections 2.2.1–2.2.2 (opens in a new tab).
What counts as better reasoning?
Here, improved reasoning performance means better results on specified mathematical, coding or logical tasks under stated evaluation conditions. It does not establish correctness on every unfamiliar problem.
Table 2 reports these AIME 2024 results for R1-Zero:
| Procedure | Result | Interpretation |
|---|---|---|
| pass@1 | 71.0% | Estimated correctness of one sampled response. |
| consensus@64 | 86.7% | Correctness after majority voting over 64 responses. |
Repeated samples can estimate pass@1 without being combined into one answer. Figure 2 uses 16 responses per question to track training. Separately, section 3 specifies temperature 0.6, top-p 0.95, a maximum generation length of 32,768 tokens, and typically 4–64 samples per question. These locations should be distinguished when interpreting the figures. DeepSeek-AI (2025b), section 2.2.4, table 2, figure 2 and section 3 (opens in a new tab).
Does more training always make a model better?
No. Additional training changes parameter values; it need not increase the parameter count. Its value depends on the data, objective, procedure and available compute. Compute-optimal training research examines these tradeoffs rather than treating model size or training duration as sufficient alone. Hoffmann et al. (2022) (opens in a new tab).
A fixed parameter count does not establish that useful learning has stopped: LLaMA reports continued improvement for its 7B model beyond one trillion training tokens. Touvron et al. (2023), section 1 (opens in a new tab).
Measure the behaviour you want. Better syntax, stronger performance on unfamiliar problems and passing a particular set of tests are different outcomes.
2025–2026: three further changes to distinguish
| Release | Documented change | Question it addresses |
|---|---|---|
| V3.2 | A sparse attention mechanism selects a subset of context positions for attention; expanded reinforcement learning changes training. | Which context positions receive computation, and how is behaviour trained? |
| V4 | Hybrid attention combines compressed sparse attention with more heavily compressed representations attended to densely. | How can long-context processing use less memory and computation? |
| V4.1-Flash | Decoder global keys and values are projected from final encoder states, separating parts of input processing from output generation. | How can processing the input and generating output use different computation? |
Sources: DeepSeek-AI (2025c) (opens in a new tab), DeepSeek-AI (2026a) (opens in a new tab), and DeepSeek-AI (2026b), model card (opens in a new tab).
V4.1-Flash's terminology does not mean a return to the exact 2017 translation architecture. These descriptions come from the developer; this article does not independently reproduce efficiency or benchmark results. A later release needs its own evidence.
Try it yourself
Classify each change, then explain what resource or behaviour it affects:
- Replace an attention component with one that retains a compressed cache.
- Adjust parameters using rewards for checkable solutions.
- Keep the trained parameters fixed, generate several candidate answers, then select by agreement.
For the arithmetic example, change the query to [0, 1]. Predict which position loses weight before calculating the new output.
Answer
- Architecture with an inference consequence: the representation changes, affecting retained cache memory. That alone does not establish better reasoning.
- Training: the reward influences parameter updates. What behaviour improves depends on the objective, data and evaluation.
- Inference and evaluation procedure: additional candidates spend more computation without changing parameters. Agreement can still be wrong.
With query [0, 1], the dot products become 0, 1, 1. Position A loses weight; B gains it; C retains the same weight. The weights are approximately 0.1978, 0.4011, 0.4011, producing 26.0445.
The values did not change. The query changed how they were combined.
Use it in practice
When reading a model announcement, record four things: the exact release, the mechanism, the claimed benefit and the conditions under which it was measured.
For an efficiency claim, identify the resource: cache memory, active computation, training effort, latency or throughput. For a reasoning claim, identify the tasks, tools, attempt count and generation budget. A gain on one measure does not settle the others.
For learning, ask yourself to predict a changed example and explain the result before running the code. Being able to repeat the terminology is a weaker check than being able to use the distinction.
How to judge further training
Use separate data for separate decisions:
Data Purpose Training data Fit the model's parameters. Validation data Guide choices such as hyperparameters, checkpoint selection and when to stop training. Held-out test set Evaluate the selected model after those decisions are complete. Keep the final test set out of training and model selection. Repeatedly using its results to decide what to change compromises its role as a final evaluation. Google (n.d.), Datasets: Dividing the original dataset (opens in a new tab).
If training error keeps falling while validation error rises on a comparable measure, the model is improving on its training examples while becoming less reliable on the validation examples. That widening gap is a warning sign of overfitting; evaluate on held-out tasks before continuing. Google (n.d.), Overfitting (opens in a new tab).
Decide whether to continue using held-out validation problems, then check the final result on a separate test set.