Articles for people who
ship the stack
Original rewrites on React, Node.js, TypeScript and AI — practical notes from the same engineering practice behind our operator software. Article bodies are in English.
Tagged: qwen
Practical notes: ToolCallingAgent vs. CodeAgent: Which Performs Better on a
Operable walkthrough of Practical notes: ToolCallingAgent vs. CodeAgent: Which Performs Better on a: contracts, checks, and drop-in code slots for teams shipping this pattern.
916 wordsRead articlePractical notes: Local LLMs for Graph RAG Extraction: The Mid-2026 Re-Benchmark
Operable walkthrough of Practical notes: Local LLMs for Graph RAG Extraction: The Mid-2026 Re-Benchmark: contracts, checks, and drop-in code slots for teams shipping this pattern.
1567 wordsRead article220 tok/s Qwen reports on dual 3090s — what to verify
Reproduce community throughput claims with careful batching, quantization, and measurement notes.
1409 wordsRead articleLangChain plus Ollama: Local Qwen Chains, Tools, and Structure
Point langchain-ollama at a local qwen tag for chains, streaming, tools, and structured output without cloud API keys.
890 wordsRead articleServing Qwen3.8-27B on One RTX 3090 With a Patched vLLM and DFlash2
How a pinned vLLM 0.28.0 fork with requantized embeddings and DFlash2 speculation serves a 27B hybrid model on a 24 GB card, and where long context slows it down.
3259 wordsRead articleServing Qwen3.8-Flash-Next on RTX 3090s with llama.cpp Tensor Offloading
Why an 88GB, 180B-parameter sparse model fits consumer GPUs, and how llama.cpp flags like -ot, -ncmoe and mmap split it across VRAM, RAM and NVMe.
2715 wordsRead articleBuilding a Local Angry Birds Clone With Qwen3.8-27B and Pi
Learn how to set up a fully local AI coding workflow using LM Studio and the Pi agent to build a playable Angry Birds level with Qwen3.8-27B.
2102 wordsRead article
About these articles
Request a 24h estimate
Need the same stack in a production operator layer? Send the brief — estimate within 24 hours.