This article is published in English.
Practical notes: Run a Useful Local LLM in 30 Minutes (Coding, RAG, Voice)
Operable walkthrough of Practical notes: Run a Useful Local LLM in 30 Minutes (Coding, RAG, Voice): contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Run a Useful Local LLM in 30 Minutes (Coding, RAG, Voice)”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through Overview, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
$ ollama run qwen3:8b
>>> rewrite this function to use async/await
The Five-Minute Shared Setup
The Five-Minute Shared Setup works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
brew install ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen3:8b
ollama run qwen3:8b "Write a python script to reverse a string."
Path A: The Local Coding Assistant
Path A: The Local Coding Assistant works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
# For 24GB+ VRAM or 32GB+ Mac Unified Memory
ollama pull qwen3-coder:30b
# For 16GB RAM laptops
ollama pull qwen2.5-coder:7b
A Cline-Tuned Ollama Model
A Cline-Tuned Ollama Model works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. A Cline-Tuned Ollama Model works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# Modelfile — cline-tuned qwen3-coder
# Save as: ./Modelfile
FROM qwen3-coder:30b
# Cline's system prompt is ~25-30K tokens before your code is added.
# Ollama's default num_ctx for Qwen 3 is 40K, but Cline ships an
# expanded system prompt in v3+ that comfortably exceeds it on any
# real task. 65536 is the safe floor for serious use; 131072 if
# you have RAM/VRAM to spare.
PARAMETER num_ctx 65536
# Code edits want low-variance output. The default 0.7 is too loose.
PARAMETER temperature 0.2
# Stop after the model's natural turn. Cline parses this.
PARAMETER stop "<|im_end|>"
# Build the Cline-tuned model from the Modelfile in the current dir
ollama create qwen3-coder-cline -f ./Modelfile
# Verify the context window actually took
ollama show qwen3-coder-cline --modelfile | grep num_ctx
# Expected output: PARAMETER num_ctx 65536
# Sanity-check it answers
ollama run qwen3-coder-cline "write a python one-liner to read /etc/hostname"
API Provider: Ollama
Base URL: http://localhost:11434
Model: qwen3-coder-cline
Context Window: 65536
Path B: RAG Over Your Own Documents
For Path B: RAG Over Your Own Documents, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
llama pull nomic-embed-text
pip install ollama numpy
"""Local RAG over a folder of notes — Ollama embeddings + numpy cosine.
No vector DB. For a few hundred docs, an in-memory list + numpy is
faster to set up and fast enough to run. Swap in sqlite-vec when
the corpus outgrows it (typically past ~5000 chunks).
Setup:
ollama pull qwen3:8b
ollama pull nomic-embed-text
pip install ollama numpy
Run:
python rag.py ./notes "what did the team decide about pricing?"
"""
from __future__ import annotations
import glob
import os
import sys
import numpy as np
import ollama
EMBED_MODEL = "nomic-embed-text"
CHAT_MODEL = "qwen3:8b"
CHUNK_WORDS = 400 # ~500 tokens; fits 4 chunks in an 8K context
TOP_K = 3
def load_chunks(folder: str) -> list[str]:
"""Read every .md / .txt under folder, split into ~CHUNK_WORDS chunks."""
chunks: list[str] = []
for path in glob.glob(os.path.join(folder, "**/*"), recursive=True):
if not path.endswith((".md", ".txt")):
continue
with open(path, encoding="utf-8") as f:
words = f.read().split()
for i in range(0, len(words), CHUNK_WORDS):
chunk = " ".join(words[i : i + CHUNK_WORDS])
if chunk.strip():
chunks.append(chunk)
return chunks
def embed(texts: list[str]) -> np.ndarray:
"""Embed a batch of texts via Ollama. Returns an (N, D) matrix."""
resp = ollama.embed(model=EMBED_MODEL, input=texts)
return np.array(resp["embeddings"], dtype=np.float32)
def top_k_indices(
query_vec: np.ndarray, doc_mat: np.ndarray, k: int
) -> list[int]:
"""Return indices of the k most cosine-similar rows in doc_mat.
The trick: if both vectors are unit-length (norm == 1), their
dot product equals their cosine similarity. So we normalize
once, then one matrix multiply gives a similarity score for
every chunk against the query — no Python loop needed.
"""
# Normalize query and every chunk to unit length.
# The 1e-8 prevents division by zero on a zero vector.
q = query_vec / (np.linalg.norm(query_vec) + 1e-8)
d = doc_mat / (np.linalg.norm(doc_mat, axis=1, keepdims=True) + 1e-8)
# One matmul across all chunks: scores[i] = cos(query, chunk_i).
scores = d @ q
# Sort descending, take the first k indices.
return np.argsort(scores)[::-1][:k].tolist()
def answer(query: str, chunks: list[str], doc_mat: np.ndarray) -> str:
"""Retrieve top-K chunks, stuff them into a prompt, generate."""
q_vec = embed([query])[0]
idx = top_k_indices(q_vec, doc_mat, k=TOP_K)
context = "\n\n---\n\n".join(chunks[i] for i in idx)
prompt = (
"Answer the question using ONLY the context below. "
"If the context does not contain the answer, say so plainly.\n\n"
f"CONTEXT:\n{context}\n\nQUESTION: {query}"
)
resp = ollama.chat(
model=CHAT_MODEL,
messages=[{"role": "user", "content": prompt}],
)
return resp["message"]["content"]
if __name__ == "__main__":
if len(sys.argv) != 3:
sys.exit('usage: python rag.py <folder> "<question>"')
folder, query = sys.argv[1], sys.argv[2]
chunks = load_chunks(folder)
if not chunks:
sys.exit(f"no .md or .txt files found under {folder}")
print(f"indexing {len(chunks)} chunks...")
doc_mat = embed(chunks)
print(answer(query, chunks, doc_mat))
python rag.py ~/Documents/meeting_notes "what did the team decide about pricing?"
Path C: The Voice Loop
For Path C: The Voice Loop, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
brew install whisper-cpp ffmpeg
pip install -U sounddevice ollama kokoro-onnx
The Voice Pipeline
For The Voice Pipeline, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For The Voice Pipeline, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
"""Local voice loop — mic → whisper.cpp → Ollama → Kokoro → speaker.
Everything runs offline. Nothing leaves the machine.
Setup:
brew install whisper-cpp ffmpeg # macOS
# (Linux: apt install ffmpeg; build whisper.cpp from source)
pip install -U sounddevice ollama kokoro-onnx
# Whisper GGML model — pick one:
# ggml-base.en.bin (~150MB, fast, English-only)
# ggml-large-v3.bin (~3GB, accurate, multilingual)
# Download from https://huggingface.co/ggerganov/whisper.cpp
# Kokoro model files auto-download to ~/.cache/kokoro-onnx/
# on first run, or grab them manually from
# https://github.com/thewh1teagle/kokoro-onnx/releases
Run:
python voice.py
"""
from __future__ import annotations
import subprocess
import tempfile
import wave
from pathlib import Path
import sounddevice as sd
import ollama
from kokoro_onnx import Kokoro
WHISPER_MODEL = Path.home() / "models" / "ggml-base.en.bin"
CHAT_MODEL = "qwen3:8b"
KOKORO_VOICE = "af_sarah" # see kokoro voices.json for all options
SAMPLE_RATE = 16000 # whisper.cpp requires 16kHz mono
RECORD_SECONDS = 5
# Kokoro auto-downloads the model files on first instantiation.
kokoro = Kokoro("kokoro-v1.0.onnx", "voices-v1.0.bin")
def record() -> Path:
"""Record from the default mic, write a 16kHz mono WAV, return its path."""
print(f"listening for {RECORD_SECONDS}s...")
audio = sd.rec(
int(RECORD_SECONDS * SAMPLE_RATE),
samplerate=SAMPLE_RATE,
channels=1, # mono
dtype="int16", # 16-bit PCM is what wave.open expects
)
sd.wait() # block until recording finishes
# Fixed path in the system tempdir — overwritten each invocation.
wav_path = Path(tempfile.gettempdir()) / "voice_in.wav"
with wave.open(str(wav_path), "wb") as w:
w.setnchannels(1) # mono
w.setsampwidth(2) # 2 bytes per sample == int16
w.setframerate(SAMPLE_RATE) # 16 kHz — whisper.cpp's required rate
w.writeframes(audio.tobytes())
return wav_path
def transcribe(wav_path: Path) -> str:
"""Run whisper.cpp on a WAV file, return the transcript text."""
result = subprocess.run(
[
"whisper-cli",
"-m", str(WHISPER_MODEL),
"-f", str(wav_path),
"-otxt", # write the transcript to <input>.txt
"-np", # suppress progress prints on stdout
],
capture_output=True,
text=True,
timeout=60,
)
if result.returncode != 0:
raise RuntimeError(f"whisper-cli failed: {result.stderr}")
# -otxt writes the transcript next to the input file, with .txt
# tacked onto the existing name — so /tmp/voice_in.wav becomes
# /tmp/voice_in.wav.txt (same directory, same stem, extra suffix).
txt_path = wav_path.with_suffix(wav_path.suffix + ".txt")
return txt_path.read_text(encoding="utf-8").strip()
def respond(text: str) -> str:
"""Send the transcript to the local LLM, return the reply."""
resp = ollama.chat(
model=CHAT_MODEL,
messages=[
{
"role": "system",
"content": (
"You are a voice assistant. Keep replies to one or "
"two short sentences. No markdown, no lists."
),
},
{"role": "user", "content": text},
],
)
return resp["message"]["content"]
def speak(text: str) -> None:
"""Synthesize the reply with Kokoro and play it back."""
samples, sample_rate = kokoro.create(
text, voice=KOKORO_VOICE, speed=1.0, lang="en-us"
)
sd.play(samples, sample_rate)
sd.wait()
if __name__ == "__main__":
wav = record()
heard = transcribe(wav)
if not heard:
print("nothing heard. exiting.")
raise SystemExit(0)
print(f"you said: {heard}")
reply = respond(heard)
print(f"model: {reply}")
speak(reply)
The Hardware Reality
When working through The Hardware Reality, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Continue Reading
When working through Continue Reading, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Operational checklist
For Operational checklist, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 9f628082e0d0: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
hardening note 0 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For hardening note 1, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through hardening note 2, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
hardening note 3 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For hardening note 4, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 4/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through hardening note 5, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 5/770: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.