Home / Articles / 74ce6b13862c for production systems — contracts and checks

This article is published in English.

74ce6b13862c for production systems — contracts and checks

Operable walkthrough of 74ce6b13862c for production systems — contracts and checks: contracts, checks, and drop-in code slots for teams shipping this pattern.

2185 words

This walkthrough rebuilds the path from raw materials to a working system for: . The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent.

Run Qwen3.8-Flash-Next Locally: A Complete Guide to Building a Local Coding Agent

For the Run Qwen3 8-Flash-Next Locally stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Qwen3.8-Flash-Next
        ↓
UD-Q4_K_XL GGUF
        ↓
llama.cpp
        ↓
OpenAI-compatible API
        ↓
OpenCode
        ↓
Local Coding Agent

What Is Qwen3.8-Flash-Next?

For the What Is Qwen3 8-Flash-Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Colocate state with the component that owns the mutation. Lifting everything to a global store makes timing bugs harder to see.

Hardware Requirements

For the Hardware Requirements stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Step 1 — Check Your NVIDIA GPU

For the Step 1 Check Your stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

nvidia-smi

Step 2 — Install the Build Dependencies

For the Step 2 Install the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow. For the Step 2 Install the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

sudo apt update
sudo apt install -y \
  git \
  cmake \
  build-essential \
  curl \
  libcurl4-openssl-dev \
  python3-pip

Step 3 — Build llama.cpp

When working through the Step 3 Build llama stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

cd /workspace
git clone \
  --branch qwen4exp/qwen3.8-flash-next \
  https://github.com/unslothai/llama.cpp.git
cd llama.cpp
cmake -B build \
  -DGGML_CUDA=ON \
  -DCMAKE_BUILD_TYPE=Release
cmake --build build \
  --config Release \
  -j"$(nproc)"
./build/bin/llama-server --version

Step 4 — Download the GGUF Model

When working through the Step 4 Download the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

pip install -U huggingface_hub
export HF_HUB_DISABLE_XET=1
unset HF_XET_HIGH_PERFORMANCE
unset HF_XET_NUM_CONCURRENT_RANGE_GETS
unset HF_HUB_ENABLE_HF_TRANSFER
cd /workspace
mkdir -p Qwen3.8-Flash-Next-GGUF
for i in 1 2 3 4; do
  shard=$(printf "%05d" "$i")
  hf download unsloth/Qwen3.8-Flash-Next-GGUF \
    "UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-${shard}-of-00004.gguf" \
    --local-dir Qwen3.8-Flash-Next-GGUF &
donewait

Step 5 — Start the Local Model Server

When working through the Step 5 Start the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Step 5 Start the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

cd /workspace/llama.cpp
./build/bin/llama-server \
  -m /workspace/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
  --alias qwen3.8-flash-next \
  --host 0.0.0.0 \
  --port 8080 \
  --ctx-size 131072 \
  --parallel 1 \
  --flash-attn on \
  --fit on \
  --fit-target 4096 \
  --jinja \
  --batch-size 1024 \
  --ubatch-size 512 \
  --temp 1.0 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0

Step 6 — Test the API

The Step 6 Test the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

curl http://127.0.0.1:8080/v1/models
qwen3.8-flash-next
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-flash-next",
    "messages": [
      {
        "role": "user",
        "content": "Write a Python function that checks whether a number is prime."
      }
    ]
  }'

Step 7 — Use the llama.cpp WebUI

The Step 7 Use the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

http://localhost:8080

Step 8 — Install OpenCode

The Step 8 Install OpenCode stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge. The Step 8 Install OpenCode stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

curl -fsSL https://opencode.ai/install | bash
opencode --version

Step 9 — Connect OpenCode to llama.cpp

For the Step 9 Connect OpenCode stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

mkdir -p ~/.config/opencode
printf '%s\n' '{"$schema":"https://opencode.ai/config.json","model":"llama.cpp/qwen3.8-flash-next","provider":{"llama.cpp":{"npm":"@ai-sdk/openai-compatible","name":"Qwen3.8 Flash Next Local","options":{"baseURL":"http://127.0.0.1:8080/v1"},"models":{"qwen3.8-flash-next":{"name":"Qwen3.8 Flash Next","limit":{"context":65536,"output":32768}}}}}}' > ~/.config/opencode/opencode.json
http://127.0.0.1:8080/v1

Step 10 — Start Your Local Coding Agent

For the Step 10 Start Your stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

cd /workspace/my-project
opencode
Your Project
     ↓
OpenCode
     ↓
llama.cpp API
     ↓
Qwen3.8-Flash-Next
     ↓
Local AI Coding Agent
Build a modern system analytics and task-management dashboard.
Monitor CPU, RAM, VRAM, GPU usage, temperatures, disk usage,
running processes, and temporary files.

Performance

For the Performance stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow. For the Performance stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

What you Learned

When working through the What you Learned stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Final Architecture

When working through the Final Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

┌──────────────────────────┐
│ Qwen3.8-Flash-Next       │
│ 125B MoE / ~6B active   │
└────────────┬─────────────┘
             │
             ▼
┌──────────────────────────┐
│ UD-Q4_K_XL GGUF          │
│ ~111GB                   │
└────────────┬─────────────┘
             │
             ▼
┌──────────────────────────┐
│ llama.cpp                │
│ CUDA + 131K Context      │
└────────────┬─────────────┘
             │
             ▼
┌──────────────────────────┐
│ OpenAI-Compatible API    │
│ localhost:8080/v1        │
└────────────┬─────────────┘
             │
             ▼
┌──────────────────────────┐
│ OpenCode                 │
│ Agentic Coding           │
└────────────┬─────────────┘
             │
             ▼
┌──────────────────────────┐
│ Fully Local AI Developer │
│ Environment              │
└──────────────────────────┘

Conclusion

When working through the Conclusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Operational checklist