Skip to content
Back to articles

AI models · Computer science

Inside an AI model: from tokens to reasoning — Erkan Malcok

Five questions for understanding model architecture, training and disclosure.

On this page
  1. What you will learn
  2. The central idea
  3. A worked example
  4. A common misconception
  5. Try it yourself
  6. Use it in practice
  7. Related reading
  8. References

When you send a sentence to an AI model, what happens between your words and its answer?

A useful starting point is:

Text → token IDs → numerical representations
     → attention and neural computation
     → next-token scores → select a token
     → append it to the sequence → repeat

This is a teaching diagram, not a complete blueprint for every model.

Keep three layers distinct: architecture describes the machinery; training adjusts its learned parameters; the application supplies interfaces, tools and memory. ChatGPT is an application. Architectural comparisons concern the particular model running inside it.

What you will learn

  • Explain the roles of tokens, vectors and attention.
  • Distinguish expert routing from additional reasoning computation.
  • Separate parameter learning from adaptation within a conversation.
  • Compare model disclosures without treating missing detail as evidence of absence.

The central idea

Most of the introductory explanation applies broadly across Transformer-based models. Their differences become clearer when we ask five questions:

  1. Which parameters perform the computation?
  2. How is context represented, compressed or selected?
  3. How does information move through layers?
  4. How are capabilities trained?
  5. How are different input types handled?

DeepSeek provides particularly concrete examples because its published reports disclose more architectural detail. That makes mechanisms easier to inspect; it does not establish better performance.

A worked example

From a sentence to representations

Consider:

I deposited money at the bank.

A tokeniser converts the text into integer identifiers. Tokens may represent words, fragments or punctuation. Their numerical identifiers do not express degrees of meaning.

An embedding maps each identifier to a learned vector. Individual dimensions should not be assumed to represent neat human concepts.

The model also represents token position. This helps distinguish “The dog chased the cat” from “The cat chased the dog”. Position may be incorporated into representations or attention calculations; implementations vary. The initial token embedding is the lookup vector before any position-dependent operations.

Suppose bank receives the same token ID in these sentences:

I deposited money at the bank.

I sat beside the river bank.

Its initial token embedding is the same, but later representations can differ with context.

Attention combines information from accessible token positions. Queries and keys determine weights; those weights combine value vectors. In causal attention, a position cannot attend to later positions.

For one simplified attention head, suppose the following three positions receive these invented weights and value vectors:

deposited: 0.5 × [2, 0]
money:     0.4 × [0, 3]
the:       0.1 × [1, 1]

Combined value:
0.5[2, 0] + 0.4[0, 3] + 0.1[1, 1] = [1.1, 1.3]

Attention can blend information from multiple positions. These numbers illustrate an operation, rather than measured behaviour inside a particular model. Vaswani et al. (2017) (opens in a new tab).

From representations to generated text

Across successive Transformer layers, attention mixes information between positions, while feed-forward networks transform each position’s representation. Residual connections carry representations forward by adding them to a sub-layer’s output; normalisation helps control their scale. Precise arrangements vary between architectures.

At the output, scores become probabilities over possible next tokens. A decoding procedure selects a token, adds it to the sequence and repeats until a stopping condition is reached. It may choose the highest-probability token or sample from a distribution. This is a conventional autoregressive teaching model, rather than a complete account of every serving implementation. Vaswani et al. (2017), sections 3.1–3.4 (opens in a new tab).

Expert routing, context and connections

In a mixture of experts (MoE) model, a router selects expert networks for each token. Those experts perform part of the neural computation; the model also has shared operations. Active parameter counts therefore describe more than an expert count.

DeepSeek’s reports give concrete examples of three different design choices:

  • Multi-head Latent Attention (MLA) in V3 compresses key/value representations into a smaller latent representation, reducing what must be cached. This is not the same as discarding whole token positions. DeepSeek-AI (2025a), section 2.1.1 (opens in a new tab).
  • Compressed Sparse Attention (CSA) in V4 compresses the cache along the sequence dimension and applies sparse attention. Heavily Compressed Attention (HCA) uses heavier compression with dense attention over the compressed representations. DeepSeek-AI (2026a) (opens in a new tab).
  • Manifold-constrained hyper-connections (mHC) in V4 constrain residual mixing to stabilise information flow between layers. They address a different problem from selecting experts or compressing context. DeepSeek-AI (2026a) (opens in a new tab).

Expert routing selects computation. Context compression changes stored representations. Connections between layers govern information flow.

Selected public disclosures

This is a comparison of selected public disclosures, not a performance ranking. Each cell names the release behind an affirmative claim. The columns cover different numbers of releases; they do not describe one fixed architecture per provider. “Undisclosed” means not established by the listed sources reviewed on 1 October 2026.

The source set is bounded:

Author-date citations in the table identify the supporting sources; full references are linked below the article.

Question OpenAI examples Claude Opus 4.6 Gemini 3 Pro / 3.1 Pro DeepSeek examples
Computation GPT-4’s detailed architecture and size are withheld. OpenAI (2023). Opus 4.6’s dense-versus-MoE design and active counts are undisclosed. Anthropic (2026). 3 Pro uses sparse MoE; counts undisclosed. 3.1 Pro refers back to its architecture. Google DeepMind (2026a), (2026b). V3: 671 billion total parameters, 37 billion active per token. DeepSeek-AI (2025a).
Context GPT-4’s report does not specify detailed internal selection or compression. OpenAI (2023). Opus 4.6’s detailed internal mechanisms are undisclosed. Anthropic (2026). 3 Pro and 3.1 Pro document long context, without a complete internal blueprint. Google DeepMind (2026a), (2026b). V3 documents MLA; V4 documents CSA and HCA. DeepSeek-AI (2025a), (2026a).
Information flow GPT-4 is Transformer-style; complete layer arrangements are undisclosed. OpenAI (2023). Opus 4.6’s summary does not support a precise layer comparison. Anthropic (2026). 3 Pro combines Transformer processing with MoE; complete arrangements are undisclosed. Google DeepMind (2026a). V4 documents mHC. DeepSeek-AI (2026a).
Learning GPT-4: pre-training and human-feedback reinforcement learning. o1: reasoning reinforcement learning. OpenAI (2023), (2024b). Opus 4.6: pre-training, human and AI feedback, and constitutional principles. Anthropic (2026). 3 Pro: pre-training and post-training, including reinforcement learning. Google DeepMind (2026a). V3: supervised fine-tuning and reinforcement learning. R1: reasoning training. DeepSeek-AI (2025a), (2025b).
Modalities GPT-4o jointly processes text, vision and audio. OpenAI (2024a). Opus 4.6 documents text and image understanding. Anthropic (2026). 3 Pro and 3.1 Pro: text, image, audio and video inputs; text output. Google DeepMind (2026a), (2026b). V3/R1 are text models; V4.1-Flash adds native visual understanding. DeepSeek-AI (2025a), (2025b), (2026b).
Public detail Broad descriptions; limited internal blueprint. Training and capability descriptions; limited internal blueprint. Broad architectural disclosures; limited implementation detail. Detailed reports and released weights for published open models.

A common misconception

Routing is not reasoning

MoE routing selects parameters used to process a token. Additional reasoning computation allows more intermediate work before answering. Neither mechanism, by itself, establishes better reasoning or generalisation.

Gemini’s Deep Think, OpenAI’s reasoning models and Claude’s supported thinking modes should therefore be compared using their documented behaviour and matched evaluation conditions. Their implementations need not be identical. OpenAI (2024b) (opens in a new tab), Anthropic (n.d.) (opens in a new tab), Google DeepMind (2026a) (opens in a new tab).

Learning is not only one training stage

Pre-training develops representations through prediction objectives. Supervised fine-tuning trains on example responses. Reinforcement learning adjusts parameters using a reward signal associated with generated responses. Reasoning training may form part of post-training, rather than a separate stage every model follows.

DeepSeek’s R1-Zero experiment began with an already pre-trained model and applied reinforcement learning without initial supervised fine-tuning. The report describes emerging reasoning behaviours, alongside readability problems and language mixing. Final R1 used a broader pipeline, including initial examples. DeepSeek-AI (2025b), section 2 (opens in a new tab).

OpenAI also documents reinforcement learning for reasoning. The defensible distinction is between particular training methods and their disclosure, rather than “one model learns while another merely predicts”.

During ordinary inference, the trained model generates a response using its current parameters. Additional thinking computation does not itself mean those parameters are being updated. OpenAI (2024b) (opens in a new tab).

Conversation adaptation is not parameter learning

Your corrections can influence subsequent responses because they enter the conversation context. That does not, by itself, update the model’s trained parameters.

Training:
Training material and feedback → parameter updates

Conversation:
Your examples and corrections → current context → adapted response

Application memory can retain information for later use. Retaining that information is different from retraining the underlying model.

Try it yourself

Compare these new sentences:

She swung the bat.

The bat flew overhead.

Assuming bat has the same token ID, what remains unchanged and what can change? Would that contextual change persist into an unrelated conversation? Explain why.

A second system accepts spoken questions and reads its answers aloud. Does that establish that its language model directly processes audio? What evidence would you need?

Answer

The initial token embedding remains the same under that assumption. Position-dependent operations and attention can produce different representations as the model processes each sentence.

Those contextual representations belong to processing these inputs. They do not, by themselves, persist into an unrelated conversation or update the embedding parameters. A separate memory or training mechanism would be a different claim.

A voice interface is insufficient evidence of native audio processing: it could use separate speech recognition and text-to-speech systems. Look for model-specific documentation. OpenAI’s GPT-4o description explicitly distinguishes an earlier speech pipeline from joint text, vision and audio processing. OpenAI (2024a) (opens in a new tab).

Use it in practice

For architectural claims, record the exact release, source and disclosure limit. Undisclosed does not mean absent.

For benchmark comparisons, match model versions, reasoning budgets, tool access and attempts. Use unfamiliar tasks with independently checkable answers.

For learning, evaluate the interaction:

Ask me to predict first. Give one hint at a time. After my attempt, identify the misconception and ask a changed question that checks my understanding.

For programming, work through prediction → explanation → execution → debugging → a changed example.

A convincing explanation is not proof of correctness. Judge the lesson by whether you can explain and solve the changed example independently.

Optional further reading

Cheng et al. (2026), Conditional Memory via Scalable Lookup (opens in a new tab) introduces Engram: learned lookup for local token patterns alongside neural computation. This extends the architectural discussion beyond the introductory pipeline. Its inclusion must be verified separately for each released model; it is not application conversation memory.

References

  1. Vaswani et al. (2017) Attention Is All You Need. arXiv. (Accessed: 1 October 2026). (opens in a new tab)
  2. OpenAI (2023) GPT-4 technical report. (Accessed: 1 October 2026). (opens in a new tab)
  3. OpenAI (2024a) Hello GPT-4o. (Accessed: 1 October 2026). (opens in a new tab)
  4. OpenAI (2024b) Learning to reason with LLMs. (Accessed: 1 October 2026). (opens in a new tab)
  5. Anthropic (2026) Transparency Hub: Claude Opus 4.6. (Accessed: 1 October 2026). (opens in a new tab)
  6. Anthropic (n.d.) Extended thinking. Claude Platform Docs. (Accessed: 1 October 2026). (opens in a new tab)
  7. Google DeepMind (2026a) Gemini 3 Pro model card. Updated May 2026. (Accessed: 1 October 2026). (opens in a new tab)
  8. Google DeepMind (2026b) Gemini 3.1 Pro model card. (Accessed: 1 October 2026). (opens in a new tab)
  9. DeepSeek-AI (2025a) DeepSeek-V3 technical report. arXiv, version 2, 18 February. Originally submitted December 2024. (Accessed: 1 October 2026). (opens in a new tab)
  10. DeepSeek-AI (2025b) DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv, version 1, 22 January. (Accessed: 1 October 2026). (opens in a new tab)
  11. DeepSeek-AI (2026a) DeepSeek V4 technical documentation. (Accessed: 1 October 2026). (opens in a new tab)
  12. DeepSeek-AI (2026b) Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. (Accessed: 1 October 2026). (opens in a new tab)
  13. Cheng et al. (2026) Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models. arXiv. (Accessed: 1 October 2026). (opens in a new tab)
More writing