Articles for people who
ship the stack
Original rewrites on React, Node.js, TypeScript and AI — practical notes from the same engineering practice behind our operator software. Article bodies are in English.
Tagged: mlops
LLMOps for Small Language Models: Serving Frameworks and Production Playbooks
Why SLMs win on cost and privacy, how vLLM, SGLang, TGI, llama.cpp, Ollama, WebLLM, ONNX, and TensorRT-LLM compare, and how to quantize, evaluate, and route them in production.
4127 wordsRead articleEvaluation Harnesses for LLM Apps: Datasets, Scoring and Regression Gates
Learn what an LLM evaluation harness is, its four core components, and how it catches prompt regressions and compares models fairly before anything reaches users.
1985 wordsRead articleFloating Ground Truth: Dataset Version Pinning in LLM Evaluation Suites
A count of lm-evaluation-harness task configs shows that almost none pin a dataset revision. Learn what that means for score comparisons and how to audit your own evals.
2162 wordsRead articleA Tiered Map of AI Engineering Concepts and When They Matter
Learn which AI engineering concepts determine whether a system works at all, which matter once you build for production, and which can wait.
2409 wordsRead articleManaging LLM-as-a-Judge as a Living Production System
Learn how Netflix's four-stage lifecycle—ground-truth data, rubric-tuned training, safe rollout, and ongoing monitoring—keeps an LLM judge accurate at scale.
4515 wordsRead articleA Curated Roadmap of 50 Hands-On AI Projects to Build in 2026
A categorized list of 50 practical AI engineering and GenAI/agentic project ideas designed to help engineers build deployable, portfolio-ready skills for 2026.
1302 wordsRead article
About these articles
Request a 24h estimate
Need the same stack in a production operator layer? Send the brief — estimate within 24 hours.