Articles for people who
ship the stack
Original rewrites on React, Node.js, TypeScript and AI — practical notes from the same engineering practice behind our operator software. Article bodies are in English.
Tagged: llm-as-a-judge
Practical notes: How to Diagnose RAG Failures with a Groundedness Judge
Operable walkthrough of Practical notes: How to Diagnose RAG Failures with a Groundedness Judge: contracts, checks, and drop-in code slots for teams shipping this pattern.
1820 wordsRead articleTraining Your Own LLM Judge: From PandaLM and JudgeLM to Prometheus
How finetuned evaluator models are built, from data and rubrics to bias control and meta-evaluation, and a practical recipe for training a domain-specific LLM judge.
8925 wordsRead articleA Dependency-Free LLM Eval With a Judge You Can Actually Trust
Build a small LLM evaluation from real logs, code checks and a single-criterion judge, then calibrate that judge against human labels so its scores mean something.
2147 wordsRead articleManaging LLM-as-a-Judge as a Living Production System
Learn how Netflix's four-stage lifecycle—ground-truth data, rubric-tuned training, safe rollout, and ongoing monitoring—keeps an LLM judge accurate at scale.
4544 wordsRead article
About these articles
Request a 24h estimate
Need the same stack in a production operator layer? Send the brief — estimate within 24 hours.