---
title: "Inside an AI model: from tokens to reasoning — Erkan Malcok"
description: "Understand tokens, attention, expert routing and reasoning, then compare what OpenAI, Anthropic, Google and DeepSeek disclose."
date: "2026-10-01"
updated: "2026-10-01"
canonical: "https://erkanmalcok.com/articles/inside-an-ai-model/"
kind: "Article"
series: "education-and-ai"
tags:
  - "AI models"
  - "Computer science"
---

# Inside an AI model: from tokens to reasoning — Erkan Malcok

Five questions for understanding model architecture, training and disclosure.

When you send a sentence to an AI model, what happens between your words and its answer?

A useful starting point is:

```text
Text → token IDs → numerical representations
     → attention and neural computation
     → next-token scores → select a token
     → append it to the sequence → repeat
```

This is a teaching diagram, not a complete blueprint for every model.

Keep three layers distinct: **architecture describes the machinery; training adjusts its learned parameters; the application supplies interfaces, tools and memory.** ChatGPT is an application. Architectural comparisons concern the particular model running inside it.

## What you will learn

- Explain the roles of tokens, vectors and attention.
- Distinguish expert routing from additional reasoning computation.
- Separate parameter learning from adaptation within a conversation.
- Compare model disclosures without treating missing detail as evidence of absence.

## The central idea

Most of the introductory explanation applies broadly across Transformer-based models. Their differences become clearer when we ask five questions:

1. Which parameters perform the computation?
2. How is context represented, compressed or selected?
3. How does information move through layers?
4. How are capabilities trained?
5. How are different input types handled?

DeepSeek provides particularly concrete examples because its published reports disclose more architectural detail. That makes mechanisms easier to inspect; it does not establish better performance.

## A worked example

### From a sentence to representations

Consider:

> I deposited money at the bank.

A tokeniser converts the text into integer identifiers. Tokens may represent words, fragments or punctuation. Their numerical identifiers do not express degrees of meaning.

An embedding maps each identifier to a learned vector. Individual dimensions should not be assumed to represent neat human concepts.

The model also represents token position. This helps distinguish “The dog chased the cat” from “The cat chased the dog”. Position may be incorporated into representations or attention calculations; implementations vary. The initial token embedding is the lookup vector before any position-dependent operations.

Suppose `bank` receives the same token ID in these sentences:

> I deposited money at the bank.
>
> I sat beside the river bank.

Its initial token embedding is the same, but later representations can differ with context.

Attention combines information from accessible token positions. Queries and keys determine weights; those weights combine value vectors. In causal attention, a position cannot attend to later positions.

For one simplified attention head, suppose the following three positions receive these invented weights and value vectors:

```text
deposited: 0.5 × [2, 0]
money:     0.4 × [0, 3]
the:       0.1 × [1, 1]

Combined value:
0.5[2, 0] + 0.4[0, 3] + 0.1[1, 1] = [1.1, 1.3]
```

Attention can blend information from multiple positions. These numbers illustrate an operation, rather than measured behaviour inside a particular model. [Vaswani *et al.* (2017)](https://arxiv.org/abs/1706.03762v1).

### From representations to generated text

Across successive Transformer layers, attention mixes information between positions, while **feed-forward networks** transform each position’s representation. **Residual connections** carry representations forward by adding them to a sub-layer’s output; normalisation helps control their scale. Precise arrangements vary between architectures.

At the output, scores become probabilities over possible next tokens. A **decoding procedure** selects a token, adds it to the sequence and repeats until a stopping condition is reached. It may choose the highest-probability token or sample from a distribution. This is a conventional autoregressive teaching model, rather than a complete account of every serving implementation. [Vaswani *et al.* (2017), sections 3.1–3.4](https://arxiv.org/abs/1706.03762v1).

During decoding, a KV cache stores earlier tokens’ attention keys and values for reuse. These representations are model-specific. KV-Lingo trains translators between selected frozen model pairs, rather than reusing one model’s cache unchanged in another. [Castin *et al.* (2026)](https://arxiv.org/abs/2609.32610v1).

### Expert routing, context and connections

In a **mixture of experts (MoE)** model, a router selects expert networks for each token. Those experts perform part of the neural computation; the model also has shared operations. Active parameter counts therefore describe more than an expert count.

DeepSeek’s reports give concrete examples of three different design choices:

- **Multi-head Latent Attention (MLA)** in V3 compresses key/value representations into a smaller latent representation, reducing what must be cached. This is not the same as discarding whole token positions. [DeepSeek-AI (2025a), section 2.1.1](https://arxiv.org/abs/2412.19437v2).
- **Compressed Sparse Attention (CSA)** in V4 compresses the cache along the sequence dimension and applies sparse attention. **Heavily Compressed Attention (HCA)** uses heavier compression with dense attention over the compressed representations. [DeepSeek-AI (2026a)](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf).
- **Manifold-constrained hyper-connections (mHC)** in V4 constrain residual mixing to stabilise information flow between layers. They address a different problem from selecting experts or compressing context. [DeepSeek-AI (2026a)](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf).

Expert routing selects computation. Context compression changes stored representations. Connections between layers govern information flow.

### Selected public disclosures

This is a comparison of selected public disclosures, not a performance ranking. Each cell names the release behind an affirmative claim. The columns cover different numbers of releases; they do not describe one fixed architecture per provider. “Undisclosed” means not established by the listed sources reviewed on 1 October 2026.

The source set is bounded:

- **OpenAI:** the [GPT-4 report (2023)](https://cdn.openai.com/papers/gpt-4.pdf), [GPT-4o description (2024a)](https://openai.com/index/hello-gpt-4o/) and [o1 research (2024b)](https://openai.com/index/learning-to-reason-with-llms/).
- **Anthropic:** the [Opus 4.6 summary (2026)](https://www.anthropic.com/transparency).
- **Google DeepMind:** the [3 Pro card (2026a)](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf) and [3.1 Pro card (2026b)](https://deepmind.google/models/model-cards/gemini-3-1-pro/).
- **DeepSeek:** the [V3 report (2025a)](https://arxiv.org/abs/2412.19437v2), [R1 report (2025b)](https://arxiv.org/abs/2501.12948v1), [V4 documentation (2026a)](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf) and [V4.1-Flash announcement (2026b)](https://www.deepseek.com/en/news/deepseek-v4-1-flash/).

Author-date citations in the table identify the supporting sources; full references are linked below the article.

| Question | OpenAI examples | Claude Opus 4.6 | Gemini 3 Pro / 3.1 Pro | DeepSeek examples |
|---|---|---|---|---|
| **Computation** | GPT-4’s detailed architecture and size are withheld. OpenAI (2023). | Opus 4.6’s dense-versus-MoE design and active counts are undisclosed. Anthropic (2026). | 3 Pro uses sparse MoE; counts undisclosed. 3.1 Pro refers back to its architecture. Google DeepMind (2026a), (2026b). | V3: 671 billion total parameters, 37 billion active per token. DeepSeek-AI (2025a). |
| **Context** | GPT-4’s report does not specify detailed internal selection or compression. OpenAI (2023). | Opus 4.6’s detailed internal mechanisms are undisclosed. Anthropic (2026). | 3 Pro and 3.1 Pro document long context, without a complete internal blueprint. Google DeepMind (2026a), (2026b). | V3 documents MLA; V4 documents CSA and HCA. DeepSeek-AI (2025a), (2026a). |
| **Information flow** | GPT-4 is Transformer-style; complete layer arrangements are undisclosed. OpenAI (2023). | Opus 4.6’s summary does not support a precise layer comparison. Anthropic (2026). | 3 Pro combines Transformer processing with MoE; complete arrangements are undisclosed. Google DeepMind (2026a). | V4 documents mHC. DeepSeek-AI (2026a). |
| **Learning** | GPT-4: pre-training and human-feedback reinforcement learning. o1: reasoning reinforcement learning. OpenAI (2023), (2024b). | Opus 4.6: pre-training, human and AI feedback, and constitutional principles. Anthropic (2026). | 3 Pro: pre-training and post-training, including reinforcement learning. Google DeepMind (2026a). | V3: supervised fine-tuning and reinforcement learning. R1: reasoning training. DeepSeek-AI (2025a), (2025b). |
| **Modalities** | GPT-4o jointly processes text, vision and audio. OpenAI (2024a). | Opus 4.6 documents text and image understanding. Anthropic (2026). | 3 Pro and 3.1 Pro: text, image, audio and video inputs; text output. Google DeepMind (2026a), (2026b). | V3/R1 are text models; V4.1-Flash adds native visual understanding. DeepSeek-AI (2025a), (2025b), (2026b). |
| **Public detail** | Broad descriptions; limited internal blueprint. | Training and capability descriptions; limited internal blueprint. | Broad architectural disclosures; limited implementation detail. | Detailed reports and released weights for published open models. |

## A common misconception

### Routing is not reasoning

MoE routing selects parameters used to process a token. Additional reasoning computation allows more intermediate work before answering. Neither mechanism, by itself, establishes better reasoning or generalisation.

Gemini’s Deep Think, OpenAI’s reasoning models and Claude’s supported thinking modes should therefore be compared using their documented behaviour and matched evaluation conditions. Their implementations need not be identical. [OpenAI (2024b)](https://openai.com/index/learning-to-reason-with-llms/), [Anthropic (n.d.)](https://platform.claude.com/docs/en/build-with-claude/extended-thinking), [Google DeepMind (2026a)](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf).

### Learning is not only one training stage

Pre-training develops representations through prediction objectives. **Supervised fine-tuning** trains on example responses. **Reinforcement learning** adjusts parameters using a reward signal associated with generated responses. Reasoning training may form part of post-training, rather than a separate stage every model follows.

DeepSeek’s R1-Zero experiment began with an already pre-trained model and applied reinforcement learning without initial supervised fine-tuning. The report describes emerging reasoning behaviours, alongside readability problems and language mixing. Final R1 used a broader pipeline, including initial examples. [DeepSeek-AI (2025b), section 2](https://arxiv.org/abs/2501.12948v1).

OpenAI also documents reinforcement learning for reasoning. The defensible distinction is between particular training methods and their disclosure, rather than “one model learns while another merely predicts”.

During ordinary inference, the trained model generates a response using its current parameters. Additional thinking computation does not itself mean those parameters are being updated. [OpenAI (2024b)](https://openai.com/index/learning-to-reason-with-llms/).

### Conversation adaptation is not parameter learning

Your corrections can influence subsequent responses because they enter the conversation context. That does not, by itself, update the model’s trained parameters.

```text
Training:
Training material and feedback → parameter updates

Conversation:
Your examples and corrections → current context → adapted response
```

Application memory can retain information for later use. Retaining that information is different from retraining the underlying model.

## Try it yourself

Compare these new sentences:

> She swung the bat.
>
> The bat flew overhead.

Assuming `bat` has the same token ID, what remains unchanged and what can change? Would that contextual change persist into an unrelated conversation? Explain why.

A second system accepts spoken questions and reads its answers aloud. Does that establish that its language model directly processes audio? What evidence would you need?



### Answer

The initial token embedding remains the same under that assumption. Position-dependent operations and attention can produce different representations as the model processes each sentence.

Those contextual representations belong to processing these inputs. They do not, by themselves, persist into an unrelated conversation or update the embedding parameters. A separate memory or training mechanism would be a different claim.

A voice interface is insufficient evidence of native audio processing: it could use separate speech recognition and text-to-speech systems. Look for model-specific documentation. OpenAI’s GPT-4o description explicitly distinguishes an earlier speech pipeline from joint text, vision and audio processing. [OpenAI (2024a)](https://openai.com/index/hello-gpt-4o/).



## Use it in practice

For architectural claims, record the exact release, source and disclosure limit. **Undisclosed does not mean absent.**

For benchmark comparisons, match model versions, reasoning budgets, tool access and attempts. Use unfamiliar tasks with independently checkable answers.

For learning, evaluate the interaction:

> Ask me to predict first. Give one hint at a time. After my attempt, identify the misconception and ask a changed question that checks my understanding.

For programming, work through **prediction → explanation → execution → debugging → a changed example**.

A convincing explanation is not proof of correctness. Judge the lesson by whether you can explain and solve the changed example independently.

## Related reading

- [LLM fluency is not reasoning](/articles/llm-fluency-is-not-reasoning/)
- [Fluent answers are not understanding](/articles/fluent-answers-are-not-understanding/)

### Optional further reading

[Cheng *et al.* (2026), *Conditional Memory via Scalable Lookup*](https://arxiv.org/abs/2601.07372v2) introduces Engram: learned lookup for local token patterns alongside neural computation. This extends the architectural discussion beyond the introductory pipeline. Its inclusion must be verified separately for each released model; it is not application conversation memory.

## References

- [Vaswani et al. \(2017\) Attention Is All You Need. arXiv. \(Accessed: 1 October 2026\).](https://arxiv.org/abs/1706.03762v1)
- [OpenAI \(2023\) GPT-4 technical report. \(Accessed: 1 October 2026\).](https://cdn.openai.com/papers/gpt-4.pdf)
- [OpenAI \(2024a\) Hello GPT-4o. \(Accessed: 1 October 2026\).](https://openai.com/index/hello-gpt-4o/)
- [OpenAI \(2024b\) Learning to reason with LLMs. \(Accessed: 1 October 2026\).](https://openai.com/index/learning-to-reason-with-llms/)
- [Anthropic \(2026\) Transparency Hub: Claude Opus 4.6. \(Accessed: 1 October 2026\).](https://www.anthropic.com/transparency)
- [Anthropic \(n.d.\) Extended thinking. Claude Platform Docs. \(Accessed: 1 October 2026\).](https://platform.claude.com/docs/en/build-with-claude/extended-thinking)
- [Google DeepMind \(2026a\) Gemini 3 Pro model card. Updated May 2026. \(Accessed: 1 October 2026\).](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf)
- [Google DeepMind \(2026b\) Gemini 3.1 Pro model card. \(Accessed: 1 October 2026\).](https://deepmind.google/models/model-cards/gemini-3-1-pro/)
- [DeepSeek-AI \(2025a\) DeepSeek-V3 technical report. arXiv, version 2, 18 February. Originally submitted December 2024. \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2412.19437v2)
- [DeepSeek-AI \(2025b\) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv, version 1, 22 January. \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2501.12948v1)
- [DeepSeek-AI \(2026a\) DeepSeek V4 technical documentation. \(Accessed: 1 October 2026\).](https://fe-static.deepseek.com/chat/transparency/deepseek-V4-model-card-EN.pdf)
- [DeepSeek-AI \(2026b\) Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. \(Accessed: 1 October 2026\).](https://www.deepseek.com/en/news/deepseek-v4-1-flash/)
- [Cheng et al. \(2026\) Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models. arXiv. \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2601.07372v2)
- [Castin, V. et al. \(2026\) KV-Lingo: Learning KV-Cache Translators with Distillation. arXiv preprint, version 1, 26 September. \(Accessed: 1 October 2026\).](https://arxiv.org/abs/2609.32610v1)
