Articles for people who
ship the stack
Original rewrites on React, Node.js, TypeScript and AI — practical notes from the same engineering practice behind our operator software. Article bodies are in English.
Tagged: ollama
LangGraph fleet triage: emergency short-circuit before RAG diagnostics
WhatsApp vehicle reports need triage first—emergency stop orders bypass retrieval, while ordinary symptoms run manuals, Llama generation, and faithfulness eval.
1225 wordsRead articleLLMOps for Small Language Models: Serving Frameworks and Production Playbooks
Why SLMs win on cost and privacy, how vLLM, SGLang, TGI, llama.cpp, Ollama, WebLLM, ONNX, and TensorRT-LLM compare, and how to quantize, evaluate, and route them in production.
4127 wordsRead articleModel Context Protocol for Beginners with FastMCP and Ollama
Learn MCP roles—host, client, server, transport—then wire a weather tool server to a local qwen3:8b model through FastMCP and STDIO.
1841 wordsRead articleFrom Scanned Lab Reports to Structured Data: OCR Architecture Trade-offs
How a budget-constrained health startup can weigh OCR libraries, Google Cloud APIs and a self-hosted vision LLM when turning medical scans into structured JSON.
1779 wordsRead articleSelf-Hosted Vision LLM OCR in Go: Rasterizing, Prompting, Parsing
How a Go worker turns medical PDFs into structured JSON with Ollama, pdftoppm, strict prompts, fallback parsing and Redis Streams consumer groups.
1684 wordsRead articleDriving Docling Pipelines Over HTTP: From Project Setup to Indexed Chunks
Walk through the Docling Pipelines REST API step by step: start the server, discover operators, validate and run an ingestion DAG, and read its execution telemetry.
3355 wordsRead articleLangChain 1.x in Practice: Chains, RAG, Tools, and Agents Locally
Learn to build chains, retrieval-augmented generation, tools, and agentic RAG with LangChain 1.x using a free local Ollama setup, no API keys required.
3655 wordsRead article
About these articles
Request a 24h estimate
Need the same stack in a production operator layer? Send the brief — estimate within 24 hours.