This article is published in English.
Practical notes: From Next-Token Prediction to ChatGPT <> Build an LLM From
Operable walkthrough of Practical notes: From Next-Token Prediction to ChatGPT <> Build an LLM From: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “From Next-Token Prediction to ChatGPT <> Build an LLM From Scratch [6]”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Act 1: Pretraining Taught Completion, Not Obedience
For the Act 1 Pretraining Taught stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Explain why the sky appears blue.
Explain why the sky appears blue.
The following questions are commonly asked in introductory physics...
The sky appears blue because Earth's atmosphere scatters
shorter wavelengths of sunlight more strongly than longer
wavelengths...
What text probably comes next?
When the text looks like an instruction,
what kind of continuation should I produce?
Act 2: Turn Conversations Into Training Examples
For the Act 2 Turn Conversations stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
example = {
"instruction": "Convert the following sentence to passive voice.",
"input": "The developer fixed the bug.",
"output": "The bug was fixed by the developer."
}
Below is an instruction that describes a task.
### Instruction:
Convert the following sentence to passive voice.
### Input:
The developer fixed the bug.
### Response:
The bug was fixed by the developer.
instruction
↓
optional context
↓
response
What the model actually sees
For the What the model actually stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the What the model actually stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
[21106, 318, 281, 12064, ... 4435, 257, 3126, ...]
Tokens:
[A, B, C, D, E]
Input:
[A, B, C, D]
Labels:
[B, C, D, E]
instruction → useful response
Padding needs special treatment
When working through the Padding needs special treatment stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Example 1:
[10, 20, 30, 40]
Example 2:
[11, 21]
Example 1:
[10, 20, 30, 40]
Example 2:
[11, 21, PAD, PAD]
PAD → PAD → PAD → PAD
ignore_index = -100
loss = cross_entropy(
logits.reshape(-1, vocab_size),
targets.reshape(-1),
ignore_index=-100,
)
Act 3: Fine-Tune the Behavior
When working through the Act 3 Fine-Tune the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
for batch in train_loader:
optimizer.zero_grad()
input_ids = batch[:, :-1]
targets = batch[:, 1:]
logits = model(input_ids)
loss = cross_entropy(
logits.flatten(0, 1),
targets.flatten(),
ignore_index=-100,
)
loss.backward()
optimizer.step()
model.learn_to_be_helpful()
### Instruction:
Summarize this paragraph.
### Response:
<clear concise summary>
### Instruction:
Write Python code that reverses a string.
### Response:
def reverse_string(s):
return s[::-1]
The CPU cache is...
### Instruction:
Explain CPU caching to a beginner.
### Response:
A CPU cache is...
Act 4: Evaluation Becomes Uncomfortable
When working through the Act 4 Evaluation Becomes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Act 4 Evaluation Becomes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Prediction: SPAM
Label: SPAM
Correct.
Recursion is when a function solves a problem by calling
itself on a smaller version of the same problem.
Recursion is a technique where a problem is reduced into
smaller instances of itself until a stopping condition is reached.
Reference answers still help.
The Reference answers still help stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
{
"instruction": "Explain recursion simply.",
"reference": "Recursion solves a problem by repeatedly reducing it...",
"model_response": "A recursive function calls itself..."
}
Ask another LLM to evaluate the answer
The Ask another LLM to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Instruction:
Explain recursion simply.
Reference answer:
...
Model answer:
...
Rate the model answer from 0 to 100 based on correctness,
relevance, and clarity.
1. Inspect representative outputs manually
2. Compare against reference answers where useful
3. Use an LLM judge for aggregate comparison
Act 5: Then We Hit a Harder Problem <> Preferences
The Act 5 Then We stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Act 5 Then We stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Response A
For the Response A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Recursion is when a function calls itself.
Response B
For the Response B stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Recursion is a technique where a function solves a problem
by calling itself on a smaller version of that problem.
The process stops when it reaches a base case.
correct vs incorrect
chosen vs rejected
{
"prompt": "Explain recursion simply.",
"chosen": """
Recursion solves a problem by repeatedly reducing it
until reaching a base case.
""",
"rejected": """
Recursion is when recursion happens recursively.
"""
}
Prompt
↓
Chosen response ─────┐
├── preference objective
Rejected response ───┘
↓
model update
Pretraining
↓
learn language patterns
------------------------------
Instruction fine-tuning
↓
learn instruction → response behavior
------------------------------
Preference tuning
↓
learn which acceptable responses we prefer
What Actually Changed?
For the What Actually Changed stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the What Actually Changed stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
### Instruction:
Pretraining:
What token should come next?
Instruction tuning:
What does a good answer look like after an instruction?
Preference tuning:
Of several reasonable answers, which behavior should we prefer?
Final Thoughts
When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
If this guide helped you…
When working through the If this guide helped stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Keep render work cheap and push expensive derivation behind memoization only after measuring. Premature memo can hide stale props bugs.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 6748f099cda4: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.