Articles for people who
ship the stack
Original rewrites on React, Node.js, TypeScript and AI — practical notes from the same engineering practice behind our operator software. Article bodies are in English.
Tagged: evaluation
Six RAG evaluation metrics that matter in production
Recall@K, nDCG, MRR, faithfulness, latency, and cost—how to debug retrieval that looks better on paper while answers get worse.
2244 wordsRead articleTyped Decision Models vs LLM Calls: Latency and Stability at the Boundary
A 240-call experiment comparing Jev's native probabilities with an LLM's self-reported ones, and how to turn that signal into a LangChain model router.
2630 wordsRead articleDeciding on an Agentic Model Upgrade by Cost per Successful Task
A practical framework for judging whether a more autonomous model is worth adopting: six metrics, a reproducible coding-agent trial, and the controls it demands.
1756 wordsRead articlePricing a Typed Decision Router in Front of Frontier LLM Calls
How to model the expected cost of gating LLM traffic with a typed decision model like Jev, set confidence thresholds, and validate routing quality before you ship it.
2655 wordsRead articleTesting AI Agents by Outcome: Verifying State Instead of Trusting Replies
Learn to evaluate tool-using AI agents by checking final state, action traces and truthfulness, including timeouts, no-action cases and repeated-trial reliability.
2131 wordsRead articleTraining Your Own LLM Judge: From PandaLM and JudgeLM to Prometheus
How finetuned evaluator models are built, from data and rubrics to bias control and meta-evaluation, and a practical recipe for training a domain-specific LLM judge.
8925 wordsRead articleFifteen LLM Concepts Traced Through a Single Support-Bot Request
Follow one customer question through an AI assistant to learn what models, tokens, embeddings, context, RAG, agents and evaluation really do, and where each one stops.
3094 wordsRead articleImproving RAG Answers One Measured Change at a Time
A measure-first workflow for fixing weak RAG answers: tune chunking, top_k, reranking, hybrid search and query rewriting one at a time and track retrieval metrics.
868 wordsRead articleEvaluation Harnesses for LLM Apps: Datasets, Scoring and Regression Gates
Learn what an LLM evaluation harness is, its four core components, and how it catches prompt regressions and compares models fairly before anything reaches users.
1985 wordsRead article
About these articles
Request a 24h estimate
Need the same stack in a production operator layer? Send the brief — estimate within 24 hours.