Home / Articles / Practical notes: Your AI Model(LLMs) Is Not Running on Python

This article is published in English.

Practical notes: Your AI Model(LLMs) Is Not Running on Python

Operable walkthrough of Practical notes: Your AI Model(LLMs) Is Not Running on Python: contracts, checks, and drop-in code slots for teams shipping this pattern.

2027 words

The following notes reconstruct a practical path around “Your AI Model(LLMs) Is Not Running on Python”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

output = model(input)

Let’s start with something simple

The Let s start with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Input
  ↓
Matrix Multiplication
  ↓
ReLU
  ↓
Matrix Multiplication
  ↓
Output
y = relu(x @ W + b)
relu()
x @ W

The journey

The The journey stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

AI Model
   ↓
PyTorch / TensorFlow
   ↓
Computational Graph
   ↓
Intermediate Representation (IR)
   ↓
Optimisations
   ↓
Lowering
   ↓
Hardware-specific Code
   ↓
Machine Instructions
   ↓
CPU / GPU / NPU

Step 1: you write Python

The Step 1 you write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The Step 1 you write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

y = torch.relu(x @ W + b)

Step 2: The model becomes a graph

For the Step 2 The model stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

y = relu(x @ W + b)

Step 3: Optimisation

For the Step 3 Optimisation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

MatMul
   ↓
Add
   ↓
ReLU

Step 4: Intermediate Representation

For the Step 4 Intermediate Representation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Step 4 Intermediate Representation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Source Code
Machine Code
variables
functions
loops
tensors
shapes
matrix operations
convolutions
attention
memory layouts

Step 5: Lowering

When working through the Step 5 Lowering stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

High-level ML operation
          ↓
Tensor operation
          ↓
Lower-level operation
          ↓
Hardware-oriented operation
          ↓
Machine instructions

Step 6: The hardware finally gets involved

When working through the Step 6 The hardware stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

model(input)
y = relu(x @ W + b)

So where did Python go?

When working through the So where did Python stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the So where did Python stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

AI is not only about models

The AI is not only stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Data → Model → Prediction
Model
  ↓
Compiler
  ↓
Hardware

Why do AI companies need compilers?

The Why do AI companies stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

WHAT
The ML model wants to compute
HOW
A specific piece of hardware should execute it efficiently

The part you found most fascinating

The The part you found stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The The part you found stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Source Code
    ↓
Lexer
    ↓
Parser
    ↓
AST
    ↓
IR
    ↓
Optimisation
    ↓
Machine Code
ML Model
    ↓
ML Representation
    ↓
ML IR
    ↓
Optimisation
    ↓
Lowering
    ↓
Hardware Code

And now you have a new question

For the And now you have stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Simple ML Language
        ↓
       Parser
        ↓
      ML AST
        ↓
       ML IR
        ↓
   Simple Optimiser
        ↓
    Lowering
        ↓
   LLVM / Backend

Final thought

For the Final thought stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

model(input)

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 3443330c9254: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.