---
title: "How Transformer models changed: architecture, training and inference — Erkan Malcok"
description: "Explore attention, expert routing and reasoning training through a worked calculation and selected examples from DeepSeek’s research."
date: "2026-10-01"
updated: "2026-10-01"
canonical: "https://erkanmalcok.com/articles/from-attention-to-reasoning/"
kind: "Article"
series: "education-and-ai"
tags:
  - "Computer science"
  - "AI models"
---

# How Transformer models changed: architecture, training and inference — Erkan Malcok

Architecture, training and inference answer different questions.

Modern Transformer models combine changes to architecture, training and inference. Attention determines how token representations interact; training shapes the model’s learned numerical values, called parameters; inference uses those parameters to generate an answer. DeepSeek provides selected examples of how these parts developed.

This article follows those examples from the original Transformer onwards. At each step, the question is: **what changed, which problem did it address, and what does the evidence establish?**

## What you will learn

- Distinguish architecture, training and inference.
- Calculate a small attention output from a query, keys and values.
- Explain why attention compression and expert routing address different constraints.
- Read reasoning claims in relation to training and evaluation conditions.

## The central idea

The Transformer evolved through changes to its structure, its training and its execution. Keeping those questions separate makes the development easier to follow, even where one design choice affects several of them.

| Question | What to inspect |
| --- | --- |
| **Architecture:** how does information move? | Layers, attention, connections and routing. |
| **Training:** how are parameters shaped? | Data, objectives, examples, rewards and optimisation. |
| **Inference:** how is an answer produced? | Active computation, cached representations and generation budget. |

### 2017: attention without sequence-aligned recurrence

Vaswani et al. introduced an encoder–decoder Transformer that removed recurrent and convolutional sequence processing. Attention relates positions directly, supporting greater parallelism during training. The decoder still generates output sequentially. Attention was already in use; the contribution was the architecture built around it. [Vaswani et al. (2017), sections 1–3](https://arxiv.org/pdf/1706.03762v7).

### From translation to causal language modelling

The original Transformer encodes an input sequence and decodes an output sequence. A decoder-only language model instead processes a growing token sequence with causal attention. At each position, causal self-attention can use the current and earlier token representations, while masking later positions. Its output contributes to predicting the next token. Pre-training teaches next-token prediction; generation repeatedly selects a token and appends it to the sequence.

GPT-3 is a documented example. Its few-shot evaluations supplied instructions and examples in the input without updating trained parameters. That is adaptation through context, distinct from further training. [Brown et al. (2020), sections 1–2](https://arxiv.org/html/2005.14165v4).

### 2023: refining the architecture and training budget

LLaMA illustrates further changes within the family, including normalisation before sub-layers and a different way to represent token positions. It also considers training smaller models for longer to reduce the cost of serving them. The important development is the interaction between design, training budget and inference cost. [Touvron et al. (2023), sections 1–2](https://arxiv.org/pdf/2302.13971v1).

## A worked example

Scaled dot-product attention compares a query with keys, divides the scores by the square root of the key dimension, applies softmax, then combines values using those weights. Multiple heads use different learned projections; positional information supplies sequence order. [Vaswani et al. (2017), sections 3.2–3.5](https://arxiv.org/pdf/1706.03762v7).

A query represents what the current position is matching against. Keys provide representations to compare with that query; values provide the information combined using the resulting weights. Real models learn the projections that produce these representations.

Use this invented query and three keys, labelled A, B and C in order:

```text
query = [1, 0]
keys  = [[1, 0], [0, 1], [1, 1]]
values = [10, 20, 40]
```

The dot products are `1, 0, 1`. With two key dimensions, divide by `sqrt(2)` to obtain approximately `0.7071, 0, 0.7071`.

Softmax exponentiates the scores and divides by their sum. The weights are approximately `0.4011, 0.1978, 0.4011`. Using the rounded weights gives:

```text
0.4011 × 10 + 0.1978 × 20 + 0.4011 × 40 = 24.011
```

Keeping the full precision until the final result gives `24.0111` to four decimal places. Here is the calculation in Python:

```python
from math import exp, sqrt

query = [1.0, 0.0]
keys = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
values = [10.0, 20.0, 40.0]

scores = [
    sum(q * k for q, k in zip(query, key)) / sqrt(len(query))
    for key in keys
]
largest = max(scores)
exponentials = [exp(score - largest) for score in scores]
total = sum(exponentials)
weights = [value / total for value in exponentials]
output = sum(weight * value for weight, value in zip(weights, values))

print([round(weight, 4) for weight in weights])
print(round(output, 4))
```

Expected output:

```text
[0.4011, 0.1978, 0.4011]
24.0111
```

Subtracting the largest score avoids unnecessarily large exponentials and leaves softmax unchanged. This example calculates one query's output with scalar values; real models use value vectors and learned projections. It omits masking and the surrounding network.

These weights describe how this attention operation combines values. They are not probabilities that the values are correct, and this calculation alone does not demonstrate reasoning.

Before running it, explain why the first and third positions receive equal weights even though their values differ.

### 2024–2025: from attention to DeepSeek’s efficiency choices

A cache stores information for reuse. In Transformer generation, a key–value cache retains earlier attention representations so they need not be calculated again for every new token. DeepSeek-V3 retains the Transformer framework and carries forward MLA and DeepSeekMoE from V2. **Multi-head Latent Attention (MLA)** compresses keys and values into a smaller learned internal representation, called a latent representation, with a separate positional component, reducing the cache retained during generation. **Mixture-of-Experts (MoE)** uses shared experts and selects routed experts for each token in its feed-forward layers, which transform each position’s representation after attention has combined information across positions. [DeepSeek-AI (2025a), section 2](https://arxiv.org/html/2412.19437v2).

These choices answer different questions: how much attention information must be stored, and which parameters must perform computation. Sparse activation does not mean the remaining parameters disappear from the model or its storage requirements.

## A common misconception

**Expert routing, reasoning training and additional computation while answering are different mechanisms.** Selecting an expert is a choice about which parameters execute. Reward-based training changes parameters. Generating a longer intermediate sequence spends more computation at inference; it does not itself retrain the model.

Neither a long explanation nor an architectural label establishes correctness. Check the answer against evidence appropriate to the task.

### Pre-training: adjusting parameters through prediction

DeepSeek-V3's pre-training shapes its parameters using a large text corpus and prediction objectives, including multi-token prediction. It precedes the report's supervised fine-tuning and reinforcement-learning stages. [DeepSeek-AI (2025a), sections 2.2, 4–5](https://arxiv.org/html/2412.19437v2).

### 2025: R1 changes the training story

R1-Zero starts from the pretrained DeepSeek-V3-Base and applies reinforcement learning without preliminary supervised fine-tuning. That distinction matters: it is not a model learning from an untrained starting point. The authors report improved reasoning-task performance alongside readability and language-mixing problems. R1 adds cold-start examples and a pipeline with supervised and reinforcement-learning stages. Its smaller distilled models use supervised fine-tuning on generated examples. [DeepSeek-AI (2025b), sections 2.2–2.4](https://arxiv.org/html/2501.12948v1).

### A small reward-training illustration

Consider the invented task: “Solve `3x + 2 = 11`.”

| Candidate response | Final-answer reward |
| --- | --- |
| Subtract 2, then divide by 3: `x = 3`. | 1 |
| Divide 11 by 3: `x ≈ 3.67`. | 0 |
| Guess `x = 3`, without valid working. | 1 |

A simplified training cycle generates candidates, checks their answers, assigns rewards and uses those rewards to update parameters. A final-answer check can reward a correct guess as well as valid working.

With parameters fixed, generating candidates and selecting by agreement is inference. It makes no parameter update.

This invented example illustrates the distinction; it does not reproduce GRPO or DeepSeek’s full reward scheme. [DeepSeek-AI (2025b), sections 2.2.1–2.2.2](https://arxiv.org/html/2501.12948v1).

### What counts as better reasoning?

Here, improved reasoning performance means better results on specified mathematical, coding or logical tasks under stated evaluation conditions. It does not establish correctness on every unfamiliar problem.

Table 2 reports these AIME 2024 results for R1-Zero:

| Procedure | Result | Interpretation |
| --- | --- | --- |
| pass@1 | 71.0% | Estimated correctness of one sampled response. |
| consensus@64 | 86.7% | Correctness after majority voting over 64 responses. |

Repeated samples can estimate pass@1 without being combined into one answer. Figure 2 uses 16 responses per question to track training. Separately, section 3 specifies temperature `0.6`, top-p `0.95`, a maximum generation length of `32,768` tokens, and typically 4–64 samples per question. These locations should be distinguished when interpreting the figures. [DeepSeek-AI (2025b), section 2.2.4, table 2, figure 2 and section 3](https://arxiv.org/html/2501.12948v1).

### Does more training always make a model better?

No. Additional training changes parameter values; it need not increase the parameter count. Its value depends on the data, objective, procedure and available compute. Compute-optimal training research examines these tradeoffs rather than treating model size or training duration as sufficient alone. [Hoffmann et al. (2022)](https://arxiv.org/abs/2203.15556).

A fixed parameter count does not establish that useful learning has stopped: LLaMA reports continued improvement for its 7B model beyond one trillion training tokens. [Touvron et al. (2023), section 1](https://arxiv.org/pdf/2302.13971v1).

Measure the behaviour you want. Better syntax, stronger performance on unfamiliar problems and passing a particular set of tests are different outcomes.

### 2025–2026: three further changes to distinguish

| Release | Documented change | Question it addresses |
| --- | --- | --- |
| **V3.2** | A sparse attention mechanism selects a subset of context positions for attention; expanded reinforcement learning changes training. | Which context positions receive computation, and how is behaviour trained? |
| **V4** | Hybrid attention combines compressed sparse attention with more heavily compressed representations attended to densely. | How can long-context processing use less memory and computation? |
| **V4.1-Flash** | Decoder global keys and values are projected from final encoder states, separating parts of input processing from output generation. | How can processing the input and generating output use different computation? |

Sources: [DeepSeek-AI (2025c)](https://arxiv.org/abs/2512.02556v1), [DeepSeek-AI (2026a)](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf), and [DeepSeek-AI (2026b), model card](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/2cba9e42aa026125f3ed06c6d98c1db82f7ca027/README.md).

V4.1-Flash's terminology does not mean a return to the exact 2017 translation architecture. These descriptions come from the developer; this article does not independently reproduce efficiency or benchmark results. A later release needs its own evidence.

## Try it yourself

Classify each change, then explain what resource or behaviour it affects:

1. Replace an attention component with one that retains a compressed cache.
2. Adjust parameters using rewards for checkable solutions.
3. Keep the trained parameters fixed, generate several candidate answers, then select by agreement.

For the arithmetic example, change the query to `[0, 1]`. Predict which position loses weight before calculating the new output.



### Answer

1. **Architecture with an inference consequence:** the representation changes, affecting retained cache memory. That alone does not establish better reasoning.
2. **Training:** the reward influences parameter updates. What behaviour improves depends on the objective, data and evaluation.
3. **Inference and evaluation procedure:** additional candidates spend more computation without changing parameters. Agreement can still be wrong.

With query `[0, 1]`, the dot products become `0, 1, 1`. Position A loses weight; B gains it; C retains the same weight. The weights are approximately `0.1978, 0.4011, 0.4011`, producing `26.0445`.

The values did not change. The query changed how they were combined.



## Use it in practice

When reading a model announcement, record four things: the exact release, the mechanism, the claimed benefit and the conditions under which it was measured.

For an efficiency claim, identify the resource: cache memory, active computation, training effort, latency or throughput. For a reasoning claim, identify the tasks, tools, attempt count and generation budget. A gain on one measure does not settle the others.

For learning, ask yourself to predict a changed example and explain the result before running the code. Being able to repeat the terminology is a weaker check than being able to use the distinction.

> **How to judge further training**
>
> Use separate data for separate decisions:
>
> | Data | Purpose |
> | --- | --- |
> | **Training data** | Fit the model's parameters. |
> | **Validation data** | Guide choices such as hyperparameters, checkpoint selection and when to stop training. |
> | **Held-out test set** | Evaluate the selected model after those decisions are complete. |
>
> Keep the final test set out of training and model selection. Repeatedly using its results to decide what to change compromises its role as a final evaluation. [Google (n.d.), Datasets: Dividing the original dataset](https://developers.google.com/machine-learning/crash-course/overfitting/dividing-datasets).
>
> If training error keeps falling while validation error rises on a comparable measure, the model is improving on its training examples while becoming less reliable on the validation examples. That widening gap is a warning sign of overfitting; evaluate on held-out tasks before continuing. [Google (n.d.), Overfitting](https://developers.google.com/machine-learning/crash-course/overfitting/overfitting).
>
> Decide whether to continue using held-out validation problems, then check the final result on a separate test set.

## Related reading

- [Inside an AI model: from tokens to reasoning](/articles/inside-an-ai-model/)
- [LLM fluency is not reasoning](/articles/llm-fluency-is-not-reasoning/)
- [Measure before you optimise](/articles/measure-before-you-optimise/)

## References

- [Brown et al. \(2020\) Language Models are Few-Shot Learners. arXiv, version 4. Available at: https://arxiv.org/html/2005.14165v4 \(Accessed: 1 October 2026\).](https://arxiv.org/html/2005.14165v4)
- [Hoffmann et al. \(2022\) Training Compute-Optimal Large Language Models. arXiv. Available at: https://arxiv.org/abs/2203.15556 \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2203.15556)
- [Google \(n.d.\) Datasets: Dividing the original dataset. Machine Learning Crash Course. Available at: https://developers.google.com/machine-learning/crash-course/overfitting/dividing-datasets \(Accessed: 1 October 2026\).](https://developers.google.com/machine-learning/crash-course/overfitting/dividing-datasets)
- [Google \(n.d.\) Overfitting. Machine Learning Crash Course. Available at: https://developers.google.com/machine-learning/crash-course/overfitting/overfitting \(Accessed: 1 October 2026\).](https://developers.google.com/machine-learning/crash-course/overfitting/overfitting)
- [Vaswani et al. \(2017\) Attention Is All You Need. arXiv, version 7, revised 2023. Available at: https://arxiv.org/pdf/1706.03762v7 \(Accessed: 1 October 2026\).](https://arxiv.org/pdf/1706.03762v7)
- [Touvron et al. \(2023\) LLaMA: Open and Efficient Foundation Language Models. arXiv, version 1. Available at: https://arxiv.org/pdf/2302.13971v1 \(Accessed: 1 October 2026\).](https://arxiv.org/pdf/2302.13971v1)
- [DeepSeek-AI \(2025a\) DeepSeek-V3 Technical Report. arXiv, version 2, 18 February. Originally submitted December 2024. Available at: https://arxiv.org/html/2412.19437v2 \(Accessed: 1 October 2026\).](https://arxiv.org/html/2412.19437v2)
- [DeepSeek-AI \(2025b\) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv, version 1. Available at: https://arxiv.org/html/2501.12948v1 \(Accessed: 1 October 2026\).](https://arxiv.org/html/2501.12948v1)
- [DeepSeek-AI \(2025c\) DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv, version 1. Available at: https://arxiv.org/abs/2512.02556v1 \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2512.02556v1)
- [DeepSeek-AI \(2026a\) DeepSeek V4 technical documentation. Available at: https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf \(Accessed: 1 October 2026\).](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf)
- [DeepSeek-AI \(2026b\) DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. Model card, revision 2cba9e42aa026125f3ed06c6d98c1db82f7ca027. Available at: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/2cba9e42aa026125f3ed06c6d98c1db82f7ca027/README.md \(Accessed: 1 October 2026\).](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/2cba9e42aa026125f3ed06c6d98c1db82f7ca027/README.md)
