Home / Articles / Inside the Computational Brain of an LLM: How a Transformer Predicts the Next

This article is published in English.

Inside the Computational Brain of an LLM: How a Transformer Predicts the Next

Operable walkthrough of Inside the Computational Brain of an LLM: How a Transformer Predicts the Next: contracts, checks, and drop-in code slots for teams shipping this pattern.

3045 words

The following notes reconstruct a practical path around “ Inside the Computational Brain of an LLM: How a Transformer Predicts the Next Token”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The Big Picture

The The Big Picture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Text
 ↓
Tokenization
 ↓
Token IDs
 ↓
Token Embeddings
 +
Positional Embeddings
 ↓
Transformer Blocks
 ↓
Layer Normalization
 ↓
Multi-Head Causal Self-Attention
 ↓
Residual Connection
 ↓
Layer Normalization
 ↓
Feed-Forward Network
 ↓
Residual Connection
 ↓
(repeated many times)
 ↓
Layer Normalization
 ↓
Linear Layer
 ↓
Softmax
 ↓
Next-token probabilities

Everything Starts With Text

The Everything Starts With Text stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

"How"     → token
"to"      → token
"predict" → token
How      → 2437
to       → 284
predict  → 4331

Token IDs → Token Embeddings

The Token IDs Token Embeddings stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Token IDs Token Embeddings stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

2437
 ↓
[0.12, -0.34, 0.72, ...]
How
to
predict

The Model Needs to Know the Position of Each Token

For the The Model Needs to stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Token Embedding
       +
Positional Embedding
       ↓
Final input representation

Now We Enter the Transformer

For the Now We Enter the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Input
  ↓
Layer Normalization
  ↓
Multi-Head Causal Self-Attention
  ↓
Residual Connection
  ↓
Layer Normalization
  ↓
Feed-Forward Network
  ↓
Residual Connection
  ↓
Output
Input
  ↓
Transformer Block 1
  ↓
Transformer Block 2
  ↓
Transformer Block 3
  ↓
...
  ↓
Transformer Block N

Layer Normalization

For the Layer Normalization stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Layer Normalization stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Self-Attention — Where Tokens Look at Other Tokens

When working through the Self-Attention Where Tokens Look stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Self-Attention

When working through the Self-Attention stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

The   cat   sat   on   the   mat   because   it   was   tired
       ↑                                  ↑
       └──────── relationship ────────────┘

Query, Key and Value

When working through the Query Key and Value stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Query Key and Value stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Query (Q)

The Query Q stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Key (K)

The Key K stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Value (V)

The Value V stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Value V stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Q + K
 ↓
Attention Scores
 ↓
How much attention should each token receive?
 ↓
Use V
 ↓
Updated token representation

Why “Causal” Self-Attention?

For the Why Causal Self-Attention stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

I
love
eating
pizza

Why Multiple Attention Heads?

For the Why Multiple Attention Heads stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Residual / Skip Connections

For the Residual Skip Connections stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Residual Skip Connections stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

┌─────────────────────┐
             │                     ↓
Input ───────┼──→ Attention ───→ Add
             │                     ↑
             └─────────────────────┘
Output = Input + Transformation(Input)

Feed-Forward Network

When working through the Feed-Forward Network stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Attention
+
Feed-Forward Network
+
Normalization
+
Residual Connections

And Then We Repeat

When working through the And Then We Repeat stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Input
  ↓
Transformer Block 1
  ↓
Transformer Block 2
  ↓
Transformer Block 3
  ↓
   ...
  ↓
Transformer Block N

What Happens After the Last Transformer Block?

When working through the What Happens After the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the What Happens After the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Transformer output
       ↓
Layer Normalization
       ↓
Linear Layer
       ↓
Logits
       ↓
Softmax
       ↓
Probabilities

The Linear Layer

The The Linear Layer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

pizza
burger
food
eat
...
pizza    → 5.7
food     → 3.2
burger   → 1.8
eat      → 2.1

Softmax — Turning Scores Into Probabilities

The Softmax Turning Scores Into stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

pizza    → 0.68
food     → 0.15
burger   → 0.10
eat      → 0.07

The Model Generates One Token at a Time

The The Model Generates One stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The The Model Generates One stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

the       → 0.35
next      → 0.25
stock     → 0.15
weather   → 0.08
...
Context
   ↓
Transformer
   ↓
Next-token probabilities
   ↓
Select next token
   ↓
Add token to context
   ↓
Repeat

Putting the Entire Architecture Together

For the Putting the Entire Architecture stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

                INPUT TEXT
                     ↓
                Tokenization
                     ↓
                  Token IDs
                     ↓
               Token Embedding
                     +
              Positional Information
                     ↓
              ┌───────────────┐
              │ Transformer   │
              │    Block      │
              │               │
              │ LayerNorm     │
              │      ↓        │
              │ Multi-Head    │
              │ Causal        │
              │ Self-Attention│
              │      ↓        │
              │ Residual      │
              │      ↓        │
              │ LayerNorm     │
              │      ↓        │
              │ Feed Forward  │
              │      ↓        │
              │ Residual      │
              └───────────────┘
                     ↓
                  Repeat N
                     ↓
                LayerNorm
                     ↓
                  Linear
                     ↓
                  Logits
                     ↓
                 Softmax
                     ↓
          Next-token probabilities
                     ↓
              Select next token

So What Does GPT Actually Mean?

For the So What Does GPT stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

G → Generative

For the G Generative stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the G Generative stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

P → Pre-trained

When working through the P Pre-trained stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

T → Transformer

When working through the T Transformer stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

GPT-1, GPT-2, GPT-3… What’s Different?

the Biggest Takeaway

Operational checklist