This article is published in English.
Practical notes: I Built a Local AI Agent With Ollama — and the Hard Part
Operable walkthrough of Practical notes: I Built a Local AI Agent With Ollama — and the Hard Part: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: I Built a Local AI Agent With Ollama — and the Hard Part Wasn’t the Model. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Why you chose Ollama
When working through the Why you chose Ollama stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
ollama pull qwen3
pip install ollama
from ollama import chat
response = chat(
model="qwen3",
messages=[
{"role": "user", "content": "Explain what an overdue invoice is."}
],
)print(response.message.content)
A chatbot answers; an agent takes steps
When working through the A chatbot answers an stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Start with narrow tools
When working through the Start with narrow tools stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the Start with narrow tools stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
CUSTOMERS = {
"acme plumbing": {
"customer_id": "cus_1042",
"name": "Acme Plumbing",
"email": "billing@example.com",
}
}
INVOICES = [
{
"invoice_id": "INV-2048",
"customer_id": "cus_1042",
"amount": 1850.00,
"days_overdue": 18,
}
]
def find_customer(name: str) -> dict:
customer = CUSTOMERS.get(name.strip().lower())
return customer or {"error": "customer_not_found"}
def get_overdue_invoices(customer_id: str) -> dict:
matches = [
invoice
for invoice in INVOICES
if invoice["customer_id"] == customer_id
and invoice["days_overdue"] > 0
]
return {"invoices": matches, "count": len(matches)}
Give the model tools, not imaginary access
The Give the model tools stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
import json
from ollama import chat
def find_customer(name: str) -> dict:
"""Find a customer by business name and return its verified record."""
customer = CUSTOMERS.get(name.strip().lower())
return customer or {"error": "customer_not_found"}
def get_overdue_invoices(customer_id: str) -> dict:
"""Return overdue invoices for a verified customer ID."""
matches = [
invoice
for invoice in INVOICES
if invoice["customer_id"] == customer_id
and invoice["days_overdue"] > 0
]
return {"invoices": matches, "count": len(matches)}
TOOLS = {
"find_customer": find_customer,
"get_overdue_invoices": get_overdue_invoices,
}
Build the agent loop
The Build the agent loop stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
SYSTEM_PROMPT = """
You are an invoice assistant.
Rules:
- Never invent a customer, invoice, email address, balance, or date.
- Use find_customer before requesting invoices.
- Only use customer IDs returned by tools.
- If a tool returns an error or no records, explain that clearly.
- You may draft communication, but you cannot send it.
"""
def run_agent(user_request: str) -> str:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_request},
] for _ in range(6):
response = chat(
model="qwen3",
messages=messages,
tools=list(TOOLS.values()),
) messages.append(response.message) if not response.message.tool_calls:
return response.message.content for call in response.message.tool_calls:
name = call.function.name
arguments = call.function.arguments if name not in TOOLS:
result = {"error": "tool_not_allowed"}
else:
try:
result = TOOLS[name](**arguments)
except (TypeError, ValueError) as error:
result = {
"error": "invalid_tool_arguments",
"detail": str(error),
} messages.append(
{
"role": "tool",
"tool_name": name,
"content": json.dumps(result),
}
) return "I stopped because the task exceeded the maximum number of steps."
The real fix was not a better prompt
The The real fix was stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The The real fix was stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Add structured output at the boundary
For the Add structured output at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
from pydantic import BaseModel, Field
class ReminderReview(BaseModel):
customer_name: str
invoice_ids: list[str]
total_due: float = Field(ge=0)
draft_subject: str
draft_body: str
requires_approval: bool = True
review_response = chat(
model="qwen3",
messages=messages,
format=ReminderReview.model_json_schema(),
)
review = ReminderReview.model_validate_json(
review_response.message.content
)
State and memory are different things
For the State and memory are stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
task_state = {
"customer_id": "cus_1042",
"verified_invoice_ids": ["INV-2048"],
"approved_actions": [],
}
Local does not automatically mean safe
For the Local does not automatically stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Local does not automatically stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
How to tested the agent
When working through the How to tested the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
What the working version looked like
When working through the What the working version stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
User request
→ find_customer(name="Acme Plumbing")
→ verified customer_id: cus_1042
→ get_overdue_invoices(customer_id="cus_1042")
→ verified invoice: INV-2048, $1,850, 18 days overdue
→ generate draft
→ wait for human approval
Final lesson
When working through the Final lesson stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the Final lesson stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for a5f763eecd03: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.