This article is published in English.
74ce6b13862c for production systems — contracts and checks
Operable walkthrough of 74ce6b13862c for production systems — contracts and checks: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: . The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent.
Run Qwen3.8-Flash-Next Locally: A Complete Guide to Building a Local Coding Agent
For the Run Qwen3 8-Flash-Next Locally stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Qwen3.8-Flash-Next
↓
UD-Q4_K_XL GGUF
↓
llama.cpp
↓
OpenAI-compatible API
↓
OpenCode
↓
Local Coding Agent
What Is Qwen3.8-Flash-Next?
For the What Is Qwen3 8-Flash-Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Colocate state with the component that owns the mutation. Lifting everything to a global store makes timing bugs harder to see.
Hardware Requirements
For the Hardware Requirements stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Step 1 — Check Your NVIDIA GPU
For the Step 1 Check Your stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
nvidia-smi
Step 2 — Install the Build Dependencies
For the Step 2 Install the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow. For the Step 2 Install the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
sudo apt update
sudo apt install -y \
git \
cmake \
build-essential \
curl \
libcurl4-openssl-dev \
python3-pip
Step 3 — Build llama.cpp
When working through the Step 3 Build llama stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
cd /workspace
git clone \
--branch qwen4exp/qwen3.8-flash-next \
https://github.com/unslothai/llama.cpp.git
cd llama.cpp
cmake -B build \
-DGGML_CUDA=ON \
-DCMAKE_BUILD_TYPE=Release
cmake --build build \
--config Release \
-j"$(nproc)"
./build/bin/llama-server --version
Step 4 — Download the GGUF Model
When working through the Step 4 Download the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
pip install -U huggingface_hub
export HF_HUB_DISABLE_XET=1
unset HF_XET_HIGH_PERFORMANCE
unset HF_XET_NUM_CONCURRENT_RANGE_GETS
unset HF_HUB_ENABLE_HF_TRANSFER
cd /workspace
mkdir -p Qwen3.8-Flash-Next-GGUF
for i in 1 2 3 4; do
shard=$(printf "%05d" "$i")
hf download unsloth/Qwen3.8-Flash-Next-GGUF \
"UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-${shard}-of-00004.gguf" \
--local-dir Qwen3.8-Flash-Next-GGUF &
donewait
Step 5 — Start the Local Model Server
When working through the Step 5 Start the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Step 5 Start the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
cd /workspace/llama.cpp
./build/bin/llama-server \
-m /workspace/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
--alias qwen3.8-flash-next \
--host 0.0.0.0 \
--port 8080 \
--ctx-size 131072 \
--parallel 1 \
--flash-attn on \
--fit on \
--fit-target 4096 \
--jinja \
--batch-size 1024 \
--ubatch-size 512 \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0
Step 6 — Test the API
The Step 6 Test the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
curl http://127.0.0.1:8080/v1/models
qwen3.8-flash-next
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-flash-next",
"messages": [
{
"role": "user",
"content": "Write a Python function that checks whether a number is prime."
}
]
}'
Step 7 — Use the llama.cpp WebUI
The Step 7 Use the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
http://localhost:8080
Step 8 — Install OpenCode
The Step 8 Install OpenCode stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge. The Step 8 Install OpenCode stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
curl -fsSL https://opencode.ai/install | bash
opencode --version
Step 9 — Connect OpenCode to llama.cpp
For the Step 9 Connect OpenCode stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
mkdir -p ~/.config/opencode
printf '%s\n' '{"$schema":"https://opencode.ai/config.json","model":"llama.cpp/qwen3.8-flash-next","provider":{"llama.cpp":{"npm":"@ai-sdk/openai-compatible","name":"Qwen3.8 Flash Next Local","options":{"baseURL":"http://127.0.0.1:8080/v1"},"models":{"qwen3.8-flash-next":{"name":"Qwen3.8 Flash Next","limit":{"context":65536,"output":32768}}}}}}' > ~/.config/opencode/opencode.json
http://127.0.0.1:8080/v1
Step 10 — Start Your Local Coding Agent
For the Step 10 Start Your stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
cd /workspace/my-project
opencode
Your Project
↓
OpenCode
↓
llama.cpp API
↓
Qwen3.8-Flash-Next
↓
Local AI Coding Agent
Build a modern system analytics and task-management dashboard.
Monitor CPU, RAM, VRAM, GPU usage, temperatures, disk usage,
running processes, and temporary files.
Performance
For the Performance stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow. For the Performance stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
What you Learned
When working through the What you Learned stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Final Architecture
When working through the Final Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
┌──────────────────────────┐
│ Qwen3.8-Flash-Next │
│ 125B MoE / ~6B active │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ UD-Q4_K_XL GGUF │
│ ~111GB │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ llama.cpp │
│ CUDA + 131K Context │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ OpenAI-Compatible API │
│ localhost:8080/v1 │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ OpenCode │
│ Agentic Coding │
└────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ Fully Local AI Developer │
│ Environment │
└──────────────────────────┘
Conclusion
When working through the Conclusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.