This article is published in English.
Practical notes: Your AI Model(LLMs) Is Not Running on Python
Operable walkthrough of Practical notes: Your AI Model(LLMs) Is Not Running on Python: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Your AI Model(LLMs) Is Not Running on Python”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
output = model(input)
Let’s start with something simple
The Let s start with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
Input
↓
Matrix Multiplication
↓
ReLU
↓
Matrix Multiplication
↓
Output
y = relu(x @ W + b)
relu()
x @ W
The journey
The The journey stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
AI Model
↓
PyTorch / TensorFlow
↓
Computational Graph
↓
Intermediate Representation (IR)
↓
Optimisations
↓
Lowering
↓
Hardware-specific Code
↓
Machine Instructions
↓
CPU / GPU / NPU
Step 1: you write Python
The Step 1 you write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The Step 1 you write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
y = torch.relu(x @ W + b)
Step 2: The model becomes a graph
For the Step 2 The model stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
y = relu(x @ W + b)
Step 3: Optimisation
For the Step 3 Optimisation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
MatMul
↓
Add
↓
ReLU
Step 4: Intermediate Representation
For the Step 4 Intermediate Representation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Step 4 Intermediate Representation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Source Code
Machine Code
variables
functions
loops
tensors
shapes
matrix operations
convolutions
attention
memory layouts
Step 5: Lowering
When working through the Step 5 Lowering stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
High-level ML operation
↓
Tensor operation
↓
Lower-level operation
↓
Hardware-oriented operation
↓
Machine instructions
Step 6: The hardware finally gets involved
When working through the Step 6 The hardware stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
model(input)
y = relu(x @ W + b)
So where did Python go?
When working through the So where did Python stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the So where did Python stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
AI is not only about models
The AI is not only stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
Data → Model → Prediction
Model
↓
Compiler
↓
Hardware
Why do AI companies need compilers?
The Why do AI companies stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
WHAT
The ML model wants to compute
HOW
A specific piece of hardware should execute it efficiently
The part you found most fascinating
The The part you found stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The The part you found stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Source Code
↓
Lexer
↓
Parser
↓
AST
↓
IR
↓
Optimisation
↓
Machine Code
ML Model
↓
ML Representation
↓
ML IR
↓
Optimisation
↓
Lowering
↓
Hardware Code
And now you have a new question
For the And now you have stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
Simple ML Language
↓
Parser
↓
ML AST
↓
ML IR
↓
Simple Optimiser
↓
Lowering
↓
LLVM / Backend
Final thought
For the Final thought stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
model(input)
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 3443330c9254: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.