首页 / 文章 / 《实用指南》:30分钟内运行实用的本地大语言模型(编程、RAG、语音功能)

《实用指南》:30分钟内运行实用的本地大语言模型(编程、RAG、语音功能)

《实用指南》操作流程详解:30分钟内运行实用的本地大语言模型(编程、RAG、语音功能):适用于采用该模式的团队的合同条款、检查清单及可直接使用的代码片段。

2857 词

以下笔记为“30分钟内运行实用的本地大语言模型(编码、RAG、语音)”提供了实用的操作路径。重点在于契约定义、检查项以及可直接使用的代码占位符,而非激励性表述。 在阅读概述时,首先写下契约内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误处理都属于产品功能的一部分,而非后续需要完善的内容。

$ ollama run qwen3:8b
>>> rewrite this function to use async/await

五分钟快速配置

“五分钟共享设置”法在被视为可量化的框架时效果最佳。在扩大范围之前,先记录一份优秀的操作案例、一个故障实例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出错时,故障应能指向单一责任主体,而非复杂的流程链。 为每轮及每次会话设定token预算。智能工具会大量消耗上下文资源,设置上限可避免演示过程变成意外的费用账单。

brew install ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen3:8b
ollama run qwen3:8b "Write a python script to reverse a string."

路径A:本地编码助手

路径A:将本地编码助手视为可度量的对象使用效果最佳。在扩大范围之前,先记录一份优秀的处理结果、一个失败案例以及回滚说明。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许默默完成部分任务。 为每轮及每次会话设定token预算。智能工具会过度扩展上下文;设置上限可避免演示变成意外的费用账单。

# For 24GB+ VRAM or 32GB+ Mac Unified Memory
ollama pull qwen3-coder:30b

# For 16GB RAM laptops
ollama pull qwen2.5-coder:7b

经过Cline调优的Ollama模型

将经过Cline调优的Ollama模型视为可测量的对象时,其性能表现最佳。在扩大应用范围之前,需记录一份理想的处理结果、一个故障案例以及回滚说明。 在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在教授循环处理逻辑之前,需锁定解释器及依赖项。笔记本电脑与持续集成环境之间的差异是API演示中最常见的隐性故障来源。 将经过Cline调优的Ollama模型视为可测量的对象时,其性能表现最佳。在扩大应用范围之前,需记录一份理想的处理结果、一个故障案例以及回滚说明。 需同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

# Modelfile — cline-tuned qwen3-coder
# Save as: ./Modelfile

FROM qwen3-coder:30b

# Cline's system prompt is ~25-30K tokens before your code is added.
# Ollama's default num_ctx for Qwen 3 is 40K, but Cline ships an
# expanded system prompt in v3+ that comfortably exceeds it on any
# real task. 65536 is the safe floor for serious use; 131072 if
# you have RAM/VRAM to spare.
PARAMETER num_ctx 65536

# Code edits want low-variance output. The default 0.7 is too loose.
PARAMETER temperature 0.2

# Stop after the model's natural turn. Cline parses this.
PARAMETER stop "<|im_end|>"
# Build the Cline-tuned model from the Modelfile in the current dir
ollama create qwen3-coder-cline -f ./Modelfile

# Verify the context window actually took
ollama show qwen3-coder-cline --modelfile | grep num_ctx
# Expected output:  PARAMETER num_ctx 65536

# Sanity-check it answers
ollama run qwen3-coder-cline "write a python one-liner to read /etc/hostname"
API Provider:      Ollama
Base URL:          http://localhost:11434
Model:             qwen3-coder-cline
Context Window:    65536

路径B:基于自有文档的RAG技术

对于路径B:基于自身文档的RAG方案,在修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。若后续步骤为代码或工具调用,应优先使用具有结构化格式且经过模式验证的输出,而非自由形式的文字描述。

llama pull nomic-embed-text
pip install ollama numpy
"""Local RAG over a folder of notes — Ollama embeddings + numpy cosine.

No vector DB. For a few hundred docs, an in-memory list + numpy is
faster to set up and fast enough to run. Swap in sqlite-vec when
the corpus outgrows it (typically past ~5000 chunks).

Setup:
    ollama pull qwen3:8b
    ollama pull nomic-embed-text
    pip install ollama numpy

Run:
    python rag.py ./notes "what did the team decide about pricing?"
"""
from __future__ import annotations
import glob
import os
import sys

import numpy as np
import ollama

EMBED_MODEL = "nomic-embed-text"
CHAT_MODEL = "qwen3:8b"
CHUNK_WORDS = 400          # ~500 tokens; fits 4 chunks in an 8K context
TOP_K = 3


def load_chunks(folder: str) -> list[str]:
    """Read every .md / .txt under folder, split into ~CHUNK_WORDS chunks."""
    chunks: list[str] = []
    for path in glob.glob(os.path.join(folder, "**/*"), recursive=True):
        if not path.endswith((".md", ".txt")):
            continue
        with open(path, encoding="utf-8") as f:
            words = f.read().split()
        for i in range(0, len(words), CHUNK_WORDS):
            chunk = " ".join(words[i : i + CHUNK_WORDS])
            if chunk.strip():
                chunks.append(chunk)
    return chunks


def embed(texts: list[str]) -> np.ndarray:
    """Embed a batch of texts via Ollama. Returns an (N, D) matrix."""
    resp = ollama.embed(model=EMBED_MODEL, input=texts)
    return np.array(resp["embeddings"], dtype=np.float32)


def top_k_indices(
    query_vec: np.ndarray, doc_mat: np.ndarray, k: int
) -> list[int]:
    """Return indices of the k most cosine-similar rows in doc_mat.

    The trick: if both vectors are unit-length (norm == 1), their
    dot product equals their cosine similarity. So we normalize
    once, then one matrix multiply gives a similarity score for
    every chunk against the query — no Python loop needed.
    """
    # Normalize query and every chunk to unit length.
    # The 1e-8 prevents division by zero on a zero vector.
    q = query_vec / (np.linalg.norm(query_vec) + 1e-8)
    d = doc_mat / (np.linalg.norm(doc_mat, axis=1, keepdims=True) + 1e-8)

    # One matmul across all chunks: scores[i] = cos(query, chunk_i).
    scores = d @ q

    # Sort descending, take the first k indices.
    return np.argsort(scores)[::-1][:k].tolist()


def answer(query: str, chunks: list[str], doc_mat: np.ndarray) -> str:
    """Retrieve top-K chunks, stuff them into a prompt, generate."""
    q_vec = embed([query])[0]
    idx = top_k_indices(q_vec, doc_mat, k=TOP_K)
    context = "\n\n---\n\n".join(chunks[i] for i in idx)

    prompt = (
        "Answer the question using ONLY the context below. "
        "If the context does not contain the answer, say so plainly.\n\n"
        f"CONTEXT:\n{context}\n\nQUESTION: {query}"
    )
    resp = ollama.chat(
        model=CHAT_MODEL,
        messages=[{"role": "user", "content": prompt}],
    )
    return resp["message"]["content"]


if __name__ == "__main__":
    if len(sys.argv) != 3:
        sys.exit('usage: python rag.py <folder> "<question>"')

    folder, query = sys.argv[1], sys.argv[2]

    chunks = load_chunks(folder)
    if not chunks:
        sys.exit(f"no .md or .txt files found under {folder}")

    print(f"indexing {len(chunks)} chunks...")
    doc_mat = embed(chunks)

    print(answer(query, chunks, doc_mat))
python rag.py ~/Documents/meeting_notes "what did the team decide about pricing?"

路径C:语音循环

对于路径C:语音循环,在修改代码之前需明确输入内容、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 当下一步操作为代码编写或工具调用时,应优先采用具有架构验证的结构化输出,而非自由形式的文本描述。

brew install whisper-cpp ffmpeg
pip install -U sounddevice ollama kokoro-onnx

语音处理流程

对于 The Voice Pipeline,在修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在流程从演示环境转向共享环境时出现意外费用。当下一步操作是编写代码或调用工具时,应优先使用具有架构验证的结构化输出,而非自由形式的文本。对于 The Voice Pipeline,在修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。需同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的一部分,而非后续需要补充的功能。

"""Local voice loop — mic → whisper.cpp → Ollama → Kokoro → speaker.

Everything runs offline. Nothing leaves the machine.

Setup:
    brew install whisper-cpp ffmpeg                  # macOS
    # (Linux: apt install ffmpeg; build whisper.cpp from source)

    pip install -U sounddevice ollama kokoro-onnx

    # Whisper GGML model — pick one:
    #   ggml-base.en.bin   (~150MB, fast, English-only)
    #   ggml-large-v3.bin  (~3GB, accurate, multilingual)
    # Download from https://huggingface.co/ggerganov/whisper.cpp

    # Kokoro model files auto-download to ~/.cache/kokoro-onnx/
    # on first run, or grab them manually from
    # https://github.com/thewh1teagle/kokoro-onnx/releases

Run:
    python voice.py
"""
from __future__ import annotations
import subprocess
import tempfile
import wave
from pathlib import Path

import sounddevice as sd
import ollama
from kokoro_onnx import Kokoro

WHISPER_MODEL = Path.home() / "models" / "ggml-base.en.bin"
CHAT_MODEL = "qwen3:8b"
KOKORO_VOICE = "af_sarah"      # see kokoro voices.json for all options
SAMPLE_RATE = 16000            # whisper.cpp requires 16kHz mono
RECORD_SECONDS = 5

# Kokoro auto-downloads the model files on first instantiation.
kokoro = Kokoro("kokoro-v1.0.onnx", "voices-v1.0.bin")


def record() -> Path:
    """Record from the default mic, write a 16kHz mono WAV, return its path."""
    print(f"listening for {RECORD_SECONDS}s...")
    audio = sd.rec(
        int(RECORD_SECONDS * SAMPLE_RATE),
        samplerate=SAMPLE_RATE,
        channels=1,                    # mono
        dtype="int16",                 # 16-bit PCM is what wave.open expects
    )
    sd.wait()                          # block until recording finishes

    # Fixed path in the system tempdir — overwritten each invocation.
    wav_path = Path(tempfile.gettempdir()) / "voice_in.wav"
    with wave.open(str(wav_path), "wb") as w:
        w.setnchannels(1)              # mono
        w.setsampwidth(2)              # 2 bytes per sample == int16
        w.setframerate(SAMPLE_RATE)    # 16 kHz — whisper.cpp's required rate
        w.writeframes(audio.tobytes())
    return wav_path


def transcribe(wav_path: Path) -> str:
    """Run whisper.cpp on a WAV file, return the transcript text."""
    result = subprocess.run(
        [
            "whisper-cli",
            "-m", str(WHISPER_MODEL),
            "-f", str(wav_path),
            "-otxt",                   # write the transcript to <input>.txt
            "-np",                     # suppress progress prints on stdout
        ],
        capture_output=True,
        text=True,
        timeout=60,
    )
    if result.returncode != 0:
        raise RuntimeError(f"whisper-cli failed: {result.stderr}")

    # -otxt writes the transcript next to the input file, with .txt
    # tacked onto the existing name — so /tmp/voice_in.wav becomes
    # /tmp/voice_in.wav.txt (same directory, same stem, extra suffix).
    txt_path = wav_path.with_suffix(wav_path.suffix + ".txt")
    return txt_path.read_text(encoding="utf-8").strip()


def respond(text: str) -> str:
    """Send the transcript to the local LLM, return the reply."""
    resp = ollama.chat(
        model=CHAT_MODEL,
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a voice assistant. Keep replies to one or "
                    "two short sentences. No markdown, no lists."
                ),
            },
            {"role": "user", "content": text},
        ],
    )
    return resp["message"]["content"]


def speak(text: str) -> None:
    """Synthesize the reply with Kokoro and play it back."""
    samples, sample_rate = kokoro.create(
        text, voice=KOKORO_VOICE, speed=1.0, lang="en-us"
    )
    sd.play(samples, sample_rate)
    sd.wait()


if __name__ == "__main__":
    wav = record()
    heard = transcribe(wav)

    if not heard:
        print("nothing heard. exiting.")
        raise SystemExit(0)

    print(f"you said: {heard}")
    reply = respond(heard)
    print(f"model:    {reply}")
    speak(reply)

硬件现实

在编写《硬件现实》相关内容时,首先列出规范:所需的输入参数、成功信号以及部分故障时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非冗长的脚本。当某个步骤出错时,故障应能指向单一责任模块,而非复杂的流程链。 缓存稳定的系统指令和工具结构。重复发送相同的开头信息是导致资源浪费的常见原因。

继续阅读

在处理“继续阅读”功能时,首先需写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与验证后输出之间的契约。为相关产物命名,明确成功判定标准,杜绝无声的半完成状态。 缓存稳定的系统指令和工具结构。重复发送相同的开头信息是导致资源浪费的常见原因。

操作检查清单

对于操作检查清单,应在修改代码之前明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏的状态。

将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。

当下一步操作是编写代码或调用工具时,应优先选择经过模式验证的结构化输出,而非自由形式的文字描述。

在调整提示词之前,需先使用固定问题集来衡量召回率。频繁更换提示词往往无法解决检索效果不佳的问题。

锁定依赖版本,并记录用于演示的图像摘要。可重复性比经验知识更为重要。

将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,拒绝默许的不完整处理。

在推广整个技术栈之前,应冻结版本,为关键流程保存标准记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其追求华而不实的临时演示,不如注重扎实可靠的性能。

关于9f628082e0d0的批量处理说明:不要将提供者密钥放入代码仓库,为每个会话设置令牌使用上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。

强化措施0作为可量化指标来使用效果最佳。在扩大范围之前,先收集一份理想的转录样本、一个故障案例以及回滚说明。配置应置于应用程序代码之外,环境文件、密钥存储和功能开关应集中存放于一个位置,以便操作人员无需查看全部内容即可进行审计。

强化措施细节0/770:针对此措施需统计运行时间、错误类型及令牌消耗情况,然后依据固定的评估标准而非主观判断来决定是否保留该变更。

针对强化措施1,在修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某一步骤失败时,故障原因应能指向具体的责任主体,而非复杂的流程问题。

强化措施细节1/770:需记录该措施的耗时、错误类型以及令牌消耗情况,然后依据固定的评估标准而非主观判断来决定是否保留该更改。

在处理强化措施笔记2时,首先列出相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在代码从演示环境转向共享环境时出现意外费用。

强化措施细节2/770:为该笔记测量实际执行时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该修改。

将强化措施笔记3视为可测量的对象来处理效果最佳。在扩大范围之前,先记录一个理想运行案例、一个失败案例以及回滚说明。 同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的内容。

强化措施细节3/770:记录该步骤的运行时间、错误类型以及代币消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。

对于强化措施笔记4,在修改代码之前需明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。应将此阶段视为输入参数与验证后输出结果之间的契约,为相关成果命名、定义成功判定条件,并拒绝默许的不完整完成情况。

强化措施细节4/770:记录该步骤的运行时间、错误类型以及代币消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。

在处理强化措施第5条时,首先写下相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。

强化措施细节5/770:针对此条要求,需测量处理时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该修改。