Home / Articles / Practical notes: Eval-Driven Development: A Software Engineering Approach to

This article is published in English.

Practical notes: Eval-Driven Development: A Software Engineering Approach to

Operable walkthrough of Practical notes: Eval-Driven Development: A Software Engineering Approach to: contracts, checks, and drop-in code slots for teams shipping this pattern.

3101 words

This walkthrough rebuilds the path from raw materials to a working system for: Eval-Driven Development: A Software Engineering Approach to Production-Grade AI Agents. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Summary

When working through the Summary stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Introduction

When working through the Introduction stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

The Production Pipeline: The High Level Flow

When working through the The Production Pipeline The stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

1. The Contract: Decouple and Version the Prompt

When working through the 1 The Contract Decouple stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

[
  {
    "agent_id": "financial_market_headlines",
    "version": 1,
    "agent_model": "openai:gpt-4",
    "prompt": "What are today's major financial market headlines?",
    "eval": {
      "contains": ["market"],
      "max_model_requests": 3,
      "min_tool_calls": 1,
      "max_tool_calls": 5,
      "judge_rubric": "The answer should be a useful response to the user's financial markets question. It should summarize market-relevant information, avoid obviously unrelated content, avoid investment advice, and avoid claiming certainty beyond what the retrieved information supports."
    }
  }
]

2. The Runtime: The Execution Engine

When working through the 2 The Runtime The stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

3. The Observability Layer: No Blind Spots

When working through the 3 The Observability Layer stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

4. The Evaluation Layer: The CI/CD of AI

When working through the 4 The Evaluation Layer stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Evals: Forcing Determinism on a Non-Deterministic System

When working through the Evals Forcing Determinism on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Evals Forcing Determinism on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Layer 1: Live Deterministic tests

The Layer 1 Live Deterministic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

uv run python tests/unittest_eval_agent.py
evaluators=[
    *[Contains(value, case_sensitive=False) for value in case.eval.contains],
    MaxModelRequests(case.eval.max_model_requests),
    MinToolCalls(case.eval.min_tool_calls),
    MaxToolCalls(case.eval.max_tool_calls),
]
(financial-agent2) alex@pop-os:/ssd/ai_works/financial_agent2$ uv run python tests/unittest_eval_agent.py
Running each eval case 3 time(s)
Evaluating task ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% 0:00:00
                               Evaluation Summary: task
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Case ID                             ┃ Metrics               ┃ Assertions ┃ Duration ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━┩
│ financial_market_headlines_v1 [1/3] │ tool_calls: 1         │ ✔✔✔✔       │    18.8s │
│                                     │ requests: 2           │            │          │
│                                     │ input_tokens: 1,325   │            │          │
│                                     │ output_tokens: 278    │            │          │
│                                     │ cost: 0.0564          │            │          │
├─────────────────────────────────────┼───────────────────────┼────────────┼──────────┤
│ financial_market_headlines_v1 [2/3] │ tool_calls: 1         │ ✔✔✔✔       │     8.8s │
│                                     │ requests: 2           │            │          │
│                                     │ input_tokens: 1,195   │            │          │
│                                     │ output_tokens: 82     │            │          │
│                                     │ cost: 0.0408          │            │          │
├─────────────────────────────────────┼───────────────────────┼────────────┼──────────┤
│ financial_market_headlines_v1 [3/3] │ tool_calls: 1         │ ✔✔✔✔       │    13.5s │
│                                     │ requests: 2           │            │          │
│                                     │ input_tokens: 1,292   │            │          │
│                                     │ output_tokens: 388    │            │          │
│                                     │ cost: 0.0620          │            │          │
├─────────────────────────────────────┼───────────────────────┼────────────┼──────────┤
│ Averages                            │ requests: 2.00        │ 100.0% ✔   │    13.7s │
│                                     │ output_tokens: 249.3  │            │          │
│                                     │ tool_calls: 1.00      │            │          │
│                                     │ cost: 0.0531          │            │          │
│                                     │ input_tokens: 1,270.7 │            │          │
└─────────────────────────────────────┴───────────────────────┴────────────┴──────────┘

Layer 2: Deterministic regression from Traces

The Layer 2 Deterministic regression stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

uv run python tests/regression_eval_traces.py
Trace: ff5a53e3392dc26cd0a2890782be70ec
Eval case: financial_market_headlines v1
Status: PASS
Model requests: 2
Tool calls: 1
contains('market'): PASS
MaxModelRequests: PASS
MaxToolCalls: PASS

Layer 3: Non Deterministic Tests — LLM As A Judge

The Layer 3 Non Deterministic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Layer 3 Non Deterministic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

[
  {
    "agent_id": "financial_market_headlines",
    "version": 1,
    "agent_model": "openai:gpt-4",
    "prompt": "What are today's major financial market headlines?",
    "eval": {
      "contains": ["market"],
      "max_model_requests": 3,
      "min_tool_calls": 1,
      "max_tool_calls": 5,
      "judge_rubric": "The answer should be a useful response to the user's financial markets question. It should summarize market-relevant information, avoid obviously unrelated content, avoid investment advice, and avoid claiming certainty beyond what the retrieved information supports."
    }
  }
]
LLMJudge(
    rubric=case.eval_case.eval.judge_rubric,
    model=args.judge_model,
    include_input=True,
    score={"evaluation_name": "judge_score", "include_reason": True},
    assertion={"evaluation_name": "judge_pass", "include_reason": True},
)
uv run python tests/regression_llm_judge_traces.py   --sample-percent 50
Trace selection: fetched=1 sampled=1 judging=1 lookback_minutes=1440 sample_percent=50 max_traces=5
Trace case: trace_id=3891f0b8432c9bbcb00e5e8adce1600b agent_id=financial_market_headlines prompt_version=1 answer_chars=207
Judging 1 trace(s) with openai:gpt-5.4
Evaluating task ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% 0:00:00
                                                   Evaluation Summary: task
┏━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Case ID              ┃ Inputs               ┃ Outputs              ┃ Scores               ┃ Assertions          ┃ Duration ┃
┡━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ financial_market_he… │ {'trace_id':         │ Today's major        │ judge_score: 0.000   │ judge_pass: ✗       │    469µs │
│                      │ '3891f0b8432c9bbcb00 │ financial market     │   Reason: The        │   Reason: The       │          │
│                      │ e5e8adce1600b',      │ headlines can be     │ response does not    │ response does not   │          │
│                      │ 'agent_id':          │ found on major       │ summarize any actual │ summarize any       │          │
│                      │ 'financial_market_he │ business and finance │ market headlines or  │ actual market       │          │
│                      │ adlines',            │ news outlets         │ provide              │ headlines or        │          │
│                      │ 'prompt_version': 1, │ including CNBC,      │ market-relevant      │ provide             │          │
│                      │ 'agent_model':       │ Yahoo Finance,       │ information; it only │ market-relevant     │          │
│                      │ 'openai:gpt-4',      │ Reuters, and         │ redirects the user   │ information; it     │          │
│                      │ 'judge_model':       │ Bloomberg. For more  │ to news websites. It │ only redirects the  │          │
│                      │ 'openai:gpt-5.4',    │ specific stories,    │ avoids investment    │ user to news        │          │
│                      │ 'prompt': "What are  │ please visit their   │ advice, but it is    │ websites. It avoids │          │
│                      │ today's major        │ websites.            │ not a useful answer  │ investment advice,  │          │
│                      │ financial market     │                      │ to the user's        │ but it is not a     │          │
│                      │ headlines?"}         │                      │ question.            │ useful answer to    │          │
│                      │                      │                      │                      │ the user's          │          │
│                      │                      │                      │                      │ question.           │          │
│                      │                      │                      │                      │                     │          │
│                      │                      │                      │                      │                     │          │
├──────────────────────┼──────────────────────┼──────────────────────┼──────────────────────┼─────────────────────┼──────────┤
│ Averages             │                      │                      │ judge_score: 0.000   │ 0.0% ✔              │    469µs │
└──────────────────────┴──────────────────────┴──────────────────────┴──────────────────────┴─────────────────────┴──────────┘
uv run python tests/regression_llm_judge_traces.py  --sample-percent 10 --lookback-minutes 30
Trace selection: fetched=1 sampled=1 judging=1 lookback_minutes=30 sample_percent=10 max_traces=5
Trace case: trace_id=38fa3a54134d8818c7d7a5afbf50967e agent_id=financial_market_headlines prompt_version=1 answer_chars=654
Judging 1 trace(s) with openai:gpt-5.4
Evaluating task ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% 0:00:00
                                             Evaluation Summary: task
┏━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Case ID            ┃ Inputs             ┃ Outputs           ┃ Scores             ┃ Assertions        ┃ Duration ┃
┡━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ financial_market_… │ {'trace_id':       │ Here are today's  │ judge_score: 0.450 │ judge_pass: ✗     │    441µs │
│                    │ '38fa3a54134d8818c │ major financial   │   Reason: The      │   Reason: The     │          │
│                    │ 7d7a5afbf50967e',  │ market headlines: │ response is        │ response is       │          │
│                    │ 'agent_id':        │                   │ market-related and │ market-related    │          │
│                    │ 'financial_market_ │ 1. Wall Street's  │ avoids investment  │ and avoids        │          │
│                    │ headlines',        │ riskiest trades   │ advice, but it     │ investment        │          │
│                    │ 'prompt_version':  │ are suddenly back │ mostly lists       │ advice, but it    │          │
│                    │ 1, 'agent_model':  │ on top: Chart of  │ article headlines  │ mostly lists      │          │
│                    │ 'openai:gpt-4',    │ the Day - Yahoo   │ and links rather   │ article headlines │          │
│                    │ 'judge_model':     │ Finance           │ than providing a   │ and links rather  │          │
│                    │ 'openai:gpt-5.4',  │ [Link](https://fi │ useful summary of  │ than providing a  │          │
│                    │ 'prompt': "What    │ nance.yahoo.com/) │ the key financial  │ useful summary of │          │
│                    │ are today's major  │ 2. S&P 500        │ market             │ the key financial │          │
│                    │ financial market   │ notches           │ developments.      │ market            │          │
│                    │ headlines?"}       │ record-high close │                    │ developments.     │          │
│                    │                    │ as rate-hike      │                    │                   │          │
│                    │                    │ worries ease -    │                    │                   │          │
│                    │                    │ Reuters           │                    │                   │          │
│                    │                    │ [Link](https://ww │                    │                   │          │
│                    │                    │ w.reuters.com/mar │                    │                   │          │
│                    │                    │ kets/us/)         │                    │                   │          │
│                    │                    │ 3. Treasury       │                    │                   │          │
│                    │                    │ yields rise as    │                    │                   │          │
│                    │                    │ U.S. threatens    │                    │                   │          │
│                    │                    │ Iran with more    │                    │                   │          │
│                    │                    │ economic          │                    │                   │          │
│                    │                    │ sanctions - CNBC  │                    │                   │          │
│                    │                    │ [Link](https://ww │                    │                   │          │
│                    │                    │ w.cnbc.com/)      │                    │                   │          │
│                    │                    │ 4. Latest stock   │                    │                   │          │
│                    │                    │ market, financial │                    │                   │          │
│                    │                    │ and business news │                    │                   │          │
│                    │                    │ - MarketWatch     │                    │                   │          │
│                    │                    │ [Link](https://ww │                    │                   │          │
│                    │                    │ w.marketwatch.com │                    │                   │          │
│                    │                    │ /)                │                    │                   │          │
│                    │                    │ 5. Latest finance │                    │                   │          │
│                    │                    │ and stock market  │                    │                   │          │
│                    │                    │ news covering the │                    │                   │          │
│                    │                    │ Dow, S&P 500,     │                    │                   │          │
│                    │                    │ banking,          │                    │                   │          │
│                    │                    │ investing and     │                    │                   │          │
│                    │                    │ regulation - WSJ  │                    │                   │          │
│                    │                    │ [Link](https://ww │                    │                   │          │
│                    │                    │ w.wsj.com/finance │                    │                   │          │
│                    │                    │ )                 │                    │                   │          │
├────────────────────┼────────────────────┼───────────────────┼────────────────────┼───────────────────┼──────────┤
│ Averages           │                    │                   │ judge_score: 0.450 │ 0.0% ✔            │    441µs │
└────────────────────┴────────────────────┴───────────────────┴────────────────────┴───────────────────┴──────────┘mar
  {
    "id": "financial_market_headlines",
    "version": 2,
    "agent": "websearch",
    "prompt": "Use the web-search tool to find and verify today's major financial-market headlines. Report the 3 to 5 most consequential developments across equities, rates, currencies, commodities, or macroeconomic policy. For each item, state what happened, identify the affected market or region, explain briefly why it matters, and name the source with a link when available. Include the relevant date and units for numerical claims. Cross-check any surprising index level, percentage move, policy decision, or economic release against a second reliable source; if it cannot be verified, omit it or clearly label it as unconfirmed. State when the information was current, distinguish facts from developing reports or interpretation, and say when reliable current information is insufficient. Do not invent facts, present stale information as today's news, or give personalized investment advice.",
    "contains": ["market"],
    "max_model_requests": 4,
    "min_tool_calls": 1,
    "max_tool_calls": 7,
    "judge_rubric": "The response should provide 3 to 5 current, consequential financial-market developments based on web research. Each item should identify what happened, the affected market or region, why it matters, and its source, preferably with a link. Numerical claims should include meaningful dates and units; surprising figures should be corroborated by a second reliable source or explicitly marked unconfirmed. The answer should state when the information was current, distinguish verified facts from developing reports or interpretation, and acknowledge insufficient evidence rather than inventing details. It must stay relevant, avoid stale news presented as current, avoid unsupported certainty, and avoid personalized investment advice. A polished but uncited answer containing an implausible or unverifiable market figure should fail."
  },

Turning the pattern into CI/CD

For the Turning the pattern into stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

References

For the References stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Score single-turn answers and multi-turn trajectories separately. Aggregate chat scores bury tool-loop failures.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 4a86f3fd2d9a: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.