《实用指南》:30分钟内运行实用的本地大语言模型(编程、RAG、语音功能)
《实用指南》操作流程详解:30分钟内运行实用的本地大语言模型(编程、RAG、语音功能):适用于采用该模式的团队的合同条款、检查清单及可直接使用的代码片段。
以下笔记为“30分钟内运行实用的本地大语言模型(编码、RAG、语音)”提供了实用的操作路径。重点在于契约定义、检查项以及可直接使用的代码占位符,而非激励性表述。 在阅读概述时,首先写下契约内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误处理都属于产品功能的一部分,而非后续需要完善的内容。
$ ollama run qwen3:8b
>>> rewrite this function to use async/await
五分钟快速配置
“五分钟共享设置”法在被视为可量化的框架时效果最佳。在扩大范围之前,先记录一份优秀的操作案例、一个故障实例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出错时,故障应能指向单一责任主体,而非复杂的流程链。 为每轮及每次会话设定token预算。智能工具会大量消耗上下文资源,设置上限可避免演示过程变成意外的费用账单。
brew install ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen3:8b
ollama run qwen3:8b "Write a python script to reverse a string."
路径A:本地编码助手
路径A:将本地编码助手视为可度量的对象使用效果最佳。在扩大范围之前,先记录一份优秀的处理结果、一个失败案例以及回滚说明。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许默默完成部分任务。 为每轮及每次会话设定token预算。智能工具会过度扩展上下文;设置上限可避免演示变成意外的费用账单。
# For 24GB+ VRAM or 32GB+ Mac Unified Memory
ollama pull qwen3-coder:30b
# For 16GB RAM laptops
ollama pull qwen2.5-coder:7b
经过Cline调优的Ollama模型
将经过Cline调优的Ollama模型视为可测量的对象时,其性能表现最佳。在扩大应用范围之前,需记录一份理想的处理结果、一个故障案例以及回滚说明。 在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在教授循环处理逻辑之前,需锁定解释器及依赖项。笔记本电脑与持续集成环境之间的差异是API演示中最常见的隐性故障来源。 将经过Cline调优的Ollama模型视为可测量的对象时,其性能表现最佳。在扩大应用范围之前,需记录一份理想的处理结果、一个故障案例以及回滚说明。 需同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。
# Modelfile — cline-tuned qwen3-coder
# Save as: ./Modelfile
FROM qwen3-coder:30b
# Cline's system prompt is ~25-30K tokens before your code is added.
# Ollama's default num_ctx for Qwen 3 is 40K, but Cline ships an
# expanded system prompt in v3+ that comfortably exceeds it on any
# real task. 65536 is the safe floor for serious use; 131072 if
# you have RAM/VRAM to spare.
PARAMETER num_ctx 65536
# Code edits want low-variance output. The default 0.7 is too loose.
PARAMETER temperature 0.2
# Stop after the model's natural turn. Cline parses this.
PARAMETER stop "<|im_end|>"
# Build the Cline-tuned model from the Modelfile in the current dir
ollama create qwen3-coder-cline -f ./Modelfile
# Verify the context window actually took
ollama show qwen3-coder-cline --modelfile | grep num_ctx
# Expected output: PARAMETER num_ctx 65536
# Sanity-check it answers
ollama run qwen3-coder-cline "write a python one-liner to read /etc/hostname"
API Provider: Ollama
Base URL: http://localhost:11434
Model: qwen3-coder-cline
Context Window: 65536
路径B:基于自有文档的RAG技术
对于路径B:基于自身文档的RAG方案,在修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。若后续步骤为代码或工具调用,应优先使用具有结构化格式且经过模式验证的输出,而非自由形式的文字描述。
llama pull nomic-embed-text
pip install ollama numpy
"""Local RAG over a folder of notes — Ollama embeddings + numpy cosine.
No vector DB. For a few hundred docs, an in-memory list + numpy is
faster to set up and fast enough to run. Swap in sqlite-vec when
the corpus outgrows it (typically past ~5000 chunks).
Setup:
ollama pull qwen3:8b
ollama pull nomic-embed-text
pip install ollama numpy
Run:
python rag.py ./notes "what did the team decide about pricing?"
"""
from __future__ import annotations
import glob
import os
import sys
import numpy as np
import ollama
EMBED_MODEL = "nomic-embed-text"
CHAT_MODEL = "qwen3:8b"
CHUNK_WORDS = 400 # ~500 tokens; fits 4 chunks in an 8K context
TOP_K = 3
def load_chunks(folder: str) -> list[str]:
"""Read every .md / .txt under folder, split into ~CHUNK_WORDS chunks."""
chunks: list[str] = []
for path in glob.glob(os.path.join(folder, "**/*"), recursive=True):
if not path.endswith((".md", ".txt")):
continue
with open(path, encoding="utf-8") as f:
words = f.read().split()
for i in range(0, len(words), CHUNK_WORDS):
chunk = " ".join(words[i : i + CHUNK_WORDS])
if chunk.strip():
chunks.append(chunk)
return chunks
def embed(texts: list[str]) -> np.ndarray:
"""Embed a batch of texts via Ollama. Returns an (N, D) matrix."""
resp = ollama.embed(model=EMBED_MODEL, input=texts)
return np.array(resp["embeddings"], dtype=np.float32)
def top_k_indices(
query_vec: np.ndarray, doc_mat: np.ndarray, k: int
) -> list[int]:
"""Return indices of the k most cosine-similar rows in doc_mat.
The trick: if both vectors are unit-length (norm == 1), their
dot product equals their cosine similarity. So we normalize
once, then one matrix multiply gives a similarity score for
every chunk against the query — no Python loop needed.
"""
# Normalize query and every chunk to unit length.
# The 1e-8 prevents division by zero on a zero vector.
q = query_vec / (np.linalg.norm(query_vec) + 1e-8)
d = doc_mat / (np.linalg.norm(doc_mat, axis=1, keepdims=True) + 1e-8)
# One matmul across all chunks: scores[i] = cos(query, chunk_i).
scores = d @ q
# Sort descending, take the first k indices.
return np.argsort(scores)[::-1][:k].tolist()
def answer(query: str, chunks: list[str], doc_mat: np.ndarray) -> str:
"""Retrieve top-K chunks, stuff them into a prompt, generate."""
q_vec = embed([query])[0]
idx = top_k_indices(q_vec, doc_mat, k=TOP_K)
context = "\n\n---\n\n".join(chunks[i] for i in idx)
prompt = (
"Answer the question using ONLY the context below. "
"If the context does not contain the answer, say so plainly.\n\n"
f"CONTEXT:\n{context}\n\nQUESTION: {query}"
)
resp = ollama.chat(
model=CHAT_MODEL,
messages=[{"role": "user", "content": prompt}],
)
return resp["message"]["content"]
if __name__ == "__main__":
if len(sys.argv) != 3:
sys.exit('usage: python rag.py <folder> "<question>"')
folder, query = sys.argv[1], sys.argv[2]
chunks = load_chunks(folder)
if not chunks:
sys.exit(f"no .md or .txt files found under {folder}")
print(f"indexing {len(chunks)} chunks...")
doc_mat = embed(chunks)
print(answer(query, chunks, doc_mat))
python rag.py ~/Documents/meeting_notes "what did the team decide about pricing?"
路径C:语音循环
对于路径C:语音循环,在修改代码之前需明确输入内容、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 当下一步操作为代码编写或工具调用时,应优先采用具有架构验证的结构化输出,而非自由形式的文本描述。
brew install whisper-cpp ffmpeg
pip install -U sounddevice ollama kokoro-onnx
语音处理流程
对于 The Voice Pipeline,在修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在流程从演示环境转向共享环境时出现意外费用。当下一步操作是编写代码或调用工具时,应优先使用具有架构验证的结构化输出,而非自由形式的文本。对于 The Voice Pipeline,在修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。需同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的一部分,而非后续需要补充的功能。
"""Local voice loop — mic → whisper.cpp → Ollama → Kokoro → speaker.
Everything runs offline. Nothing leaves the machine.
Setup:
brew install whisper-cpp ffmpeg # macOS
# (Linux: apt install ffmpeg; build whisper.cpp from source)
pip install -U sounddevice ollama kokoro-onnx
# Whisper GGML model — pick one:
# ggml-base.en.bin (~150MB, fast, English-only)
# ggml-large-v3.bin (~3GB, accurate, multilingual)
# Download from https://huggingface.co/ggerganov/whisper.cpp
# Kokoro model files auto-download to ~/.cache/kokoro-onnx/
# on first run, or grab them manually from
# https://github.com/thewh1teagle/kokoro-onnx/releases
Run:
python voice.py
"""
from __future__ import annotations
import subprocess
import tempfile
import wave
from pathlib import Path
import sounddevice as sd
import ollama
from kokoro_onnx import Kokoro
WHISPER_MODEL = Path.home() / "models" / "ggml-base.en.bin"
CHAT_MODEL = "qwen3:8b"
KOKORO_VOICE = "af_sarah" # see kokoro voices.json for all options
SAMPLE_RATE = 16000 # whisper.cpp requires 16kHz mono
RECORD_SECONDS = 5
# Kokoro auto-downloads the model files on first instantiation.
kokoro = Kokoro("kokoro-v1.0.onnx", "voices-v1.0.bin")
def record() -> Path:
"""Record from the default mic, write a 16kHz mono WAV, return its path."""
print(f"listening for {RECORD_SECONDS}s...")
audio = sd.rec(
int(RECORD_SECONDS * SAMPLE_RATE),
samplerate=SAMPLE_RATE,
channels=1, # mono
dtype="int16", # 16-bit PCM is what wave.open expects
)
sd.wait() # block until recording finishes
# Fixed path in the system tempdir — overwritten each invocation.
wav_path = Path(tempfile.gettempdir()) / "voice_in.wav"
with wave.open(str(wav_path), "wb") as w:
w.setnchannels(1) # mono
w.setsampwidth(2) # 2 bytes per sample == int16
w.setframerate(SAMPLE_RATE) # 16 kHz — whisper.cpp's required rate
w.writeframes(audio.tobytes())
return wav_path
def transcribe(wav_path: Path) -> str:
"""Run whisper.cpp on a WAV file, return the transcript text."""
result = subprocess.run(
[
"whisper-cli",
"-m", str(WHISPER_MODEL),
"-f", str(wav_path),
"-otxt", # write the transcript to <input>.txt
"-np", # suppress progress prints on stdout
],
capture_output=True,
text=True,
timeout=60,
)
if result.returncode != 0:
raise RuntimeError(f"whisper-cli failed: {result.stderr}")
# -otxt writes the transcript next to the input file, with .txt
# tacked onto the existing name — so /tmp/voice_in.wav becomes
# /tmp/voice_in.wav.txt (same directory, same stem, extra suffix).
txt_path = wav_path.with_suffix(wav_path.suffix + ".txt")
return txt_path.read_text(encoding="utf-8").strip()
def respond(text: str) -> str:
"""Send the transcript to the local LLM, return the reply."""
resp = ollama.chat(
model=CHAT_MODEL,
messages=[
{
"role": "system",
"content": (
"You are a voice assistant. Keep replies to one or "
"two short sentences. No markdown, no lists."
),
},
{"role": "user", "content": text},
],
)
return resp["message"]["content"]
def speak(text: str) -> None:
"""Synthesize the reply with Kokoro and play it back."""
samples, sample_rate = kokoro.create(
text, voice=KOKORO_VOICE, speed=1.0, lang="en-us"
)
sd.play(samples, sample_rate)
sd.wait()
if __name__ == "__main__":
wav = record()
heard = transcribe(wav)
if not heard:
print("nothing heard. exiting.")
raise SystemExit(0)
print(f"you said: {heard}")
reply = respond(heard)
print(f"model: {reply}")
speak(reply)
硬件现实
在编写《硬件现实》相关内容时,首先列出规范:所需的输入参数、成功信号以及部分故障时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非冗长的脚本。当某个步骤出错时,故障应能指向单一责任模块,而非复杂的流程链。 缓存稳定的系统指令和工具结构。重复发送相同的开头信息是导致资源浪费的常见原因。
继续阅读
在处理“继续阅读”功能时,首先需写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与验证后输出之间的契约。为相关产物命名,明确成功判定标准,杜绝无声的半完成状态。 缓存稳定的系统指令和工具结构。重复发送相同的开头信息是导致资源浪费的常见原因。
操作检查清单
对于操作检查清单,应在修改代码之前明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏的状态。
将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。
当下一步操作是编写代码或调用工具时,应优先选择经过模式验证的结构化输出,而非自由形式的文字描述。
在调整提示词之前,需先使用固定问题集来衡量召回率。频繁更换提示词往往无法解决检索效果不佳的问题。
锁定依赖版本,并记录用于演示的图像摘要。可重复性比经验知识更为重要。
将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,拒绝默许的不完整处理。
在推广整个技术栈之前,应冻结版本,为关键流程保存标准记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其追求华而不实的临时演示,不如注重扎实可靠的性能。
关于9f628082e0d0的批量处理说明:不要将提供者密钥放入代码仓库,为每个会话设置令牌使用上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。
强化措施0作为可量化指标来使用效果最佳。在扩大范围之前,先收集一份理想的转录样本、一个故障案例以及回滚说明。配置应置于应用程序代码之外,环境文件、密钥存储和功能开关应集中存放于一个位置,以便操作人员无需查看全部内容即可进行审计。
强化措施细节0/770:针对此措施需统计运行时间、错误类型及令牌消耗情况,然后依据固定的评估标准而非主观判断来决定是否保留该变更。
针对强化措施1,在修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某一步骤失败时,故障原因应能指向具体的责任主体,而非复杂的流程问题。
强化措施细节1/770:需记录该措施的耗时、错误类型以及令牌消耗情况,然后依据固定的评估标准而非主观判断来决定是否保留该更改。
在处理强化措施笔记2时,首先列出相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在代码从演示环境转向共享环境时出现意外费用。
强化措施细节2/770:为该笔记测量实际执行时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该修改。
将强化措施笔记3视为可测量的对象来处理效果最佳。在扩大范围之前,先记录一个理想运行案例、一个失败案例以及回滚说明。 同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的内容。
强化措施细节3/770:记录该步骤的运行时间、错误类型以及代币消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。
对于强化措施笔记4,在修改代码之前需明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。应将此阶段视为输入参数与验证后输出结果之间的契约,为相关成果命名、定义成功判定条件,并拒绝默许的不完整完成情况。
强化措施细节4/770:记录该步骤的运行时间、错误类型以及代币消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。
在处理强化措施第5条时,首先写下相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。
强化措施细节5/770:针对此条要求,需测量处理时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该修改。