This article is published in English.
Practical notes: So Which Agent SDK Should You Build With?
Operable walkthrough of Practical notes: So Which Agent SDK Should You Build With?: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: So Which Agent SDK Should You Build With?. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What is an agent?
When working through the What is an agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The Nutrition Label Investigator Agent
When working through the The Nutrition Label Investigator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Keeping the comparison fair
When working through the Keeping the comparison fair stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Keeping the comparison fair stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Anthropic’s Claude Agent SDK
The Anthropic s Claude Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
@tool(
"check_sfa_additive",
"Check whether a food additive is permitted by the Singapore Food Agency, "
"using the full SFA permitted-additives list parsed from the official PDF. "
"Accepts an E-number ('E211') OR a name ('Sodium Benzoate', 'Soy Lecithin'). "
"If you pass a name and also know its E/INS number, set e_number_hint — an "
"unrecognised name that the hint resolves is REMEMBERED for next time.",
{"additive": str, "e_number_hint": str},
)
async def check_sfa_additive(args: dict) -> dict:
raw = args["additive"].strip()
hint = args.get("e_number_hint", "").strip()
entry = _resolve_additive(raw)
if not entry and hint:
entry = _resolve_additive(hint)
if entry: # learn: this label name -> this number, persisted to disk
num = entry["e_number"] or entry.get("ins") or ""
_MEMORY["learned_aliases"][_norm(raw)] = re.sub(r"\(.*\)", "", num).lower()
_save_memory(_MEMORY)
return _ok(_format_entry(entry, raw))
server = create_sdk_mcp_server(
name="sg-nutrition-tools",
version="1.0.0",
tools=[recall_product, check_sfa_additive, check_hcs, calculate_nutri_grade],
)
options = ClaudeAgentOptions(
system_prompt=SYSTEM_PROMPT,
mcp_servers={"sg": server},
allowed_tools=[
"mcp__sg__recall_product",
"mcp__sg__check_sfa_additive",
"mcp__sg__check_hcs",
"mcp__sg__calculate_nutri_grade",
],
model=CLAUDE_MODEL,
permission_mode="bypassPermissions",
max_turns=20,
# Point the bundled CLI at the LiteLLM proxy's Anthropic endpoint.
env={
**os.environ,
"ANTHROPIC_BASE_URL": BASE_URL,
"ANTHROPIC_AUTH_TOKEN": API_KEY,
},
)
async def message_stream():
yield {
"type": "user",
"message": {"role": "user", "content": [
{"type": "text", "text": "Investigate this nutrition label."},
{"type": "image", "source": {"type": "base64", "media_type": mime, "data": b64}},
]},
}
async with ClaudeSDKClient(options=options) as client:
await client.query(message_stream())
async for message in client.receive_response():
... # stream ToolUseBlock / TextBlock as they arrive
sausheong claude-demo % uv run python main.py ../milo.jpeg
Investigating image: ../milo.jpeg
────────────────────────────────────────────────────────────
I'll start by reading the label carefully, then run all checks in parallel.
**Reading from the label:**
- **Product:** Nestlé MILO 3in1
- **Net weight:** 27g powder sachet (add hot water to prepare — this is a **powder, not a ready-to-drink beverage**)
- **Ingredients:** Sugar, PROTOMAL® Malt Extract (Contains Barley), Skimmed Milk (Cow's Milk), Cocoa, Palm Oil, Minerals (Dicalcium Phosphate, Disodium Phosphate, Ferric Pyrophosphate), Resistant Dextrin, Flavourings, Vitamins (C, B3, D, B6, B2, B12)
- **Identifiable additives:** Dicalcium Phosphate (E341), Disodium Phosphate (E339), Ferric Pyrophosphate; Flavourings (generic — cannot identify specific ones)
Kicking off recall + HCS check + all additive checks simultaneously:
→ tool: recall_product({'product_name': 'Nestlé MILO 3in1'})
→ tool: check_hcs({'product_name': 'MILO 3in1'})
→ tool: check_sfa_additive({'additive': 'Dicalcium Phosphate', 'e_number_hint': 'E341'})
→ tool: check_sfa_additive({'additive': 'Disodium Phosphate', 'e_number_hint': 'E339'})
→ tool: check_sfa_additive({'additive': 'Ferric Pyrophosphate', 'e_number_hint': 'E450'})
All checks returned. Note: this is a **powder**, not a ready-to-drink beverage, so Nutri-Grade (which requires per-100ml liquid values) does **not apply** — I will skip `calculate_nutri_grade`.
---
## Reasoning
**Step 1 — Product recall.** This product has been seen before. The prior verdict is consistent with what I read now — I'll treat this fresh investigation as the authoritative one and use the prior as a cross-check.
**Step 2 — Label reading.**
- MILO 3in1 is a **27g powder sachet** — you add hot water to make a drink. It is classified as a "Tonic Food Drink" on the label. Because it is sold as a dry powder (not a pre-packaged ready-to-drink liquid), Nutri-Grade per-100ml calculation does not apply.
- Nutritional values per serving (27g): Energy 110 kcal, Total Fat ~2.4g, Saturated Fat ~1.2g, Total Sugars ~13.8g (of which Added Sugars ~10g), Sodium ~38mg.
- **Sugar is the #1 ingredient** — it appears first in the ingredient list, before the malt extract.
**Step 3 — Additives.**
- **Dicalcium Phosphate (E341):** SFA-permitted ✅ — used as a mineral/calcium source.
- **Disodium Phosphate (E339):** SFA-permitted ✅ — used as an acidity regulator/mineral.
- **Ferric Pyrophosphate:** SFA-permitted ✅ — used as an iron supplement/mineral source.
- **Resistant Dextrin:** A dietary fibre ingredient (prebiotic fibre), not a regulated additive — no E number, no SFA check needed.
- **Flavourings:** Generic label — cannot identify specific flavouring compounds; SFA check not possible.
- **Vitamins (C, B3, D, B6, B2, B12):** Added as nutritional fortification, not food additives in the regulatory sense. Riboflavin (B2) carries E101 when used as a colourant, but here it's explicitly declared as a vitamin supplement — standard and expected.
**Step 4 — Healthier Choice Symbol.** HCS certified ✅ (MILO 3in1 Activ-Go appears multiple times in the HPB dataset — `check_hcs`).
**Step 5 — Sugar content reality check.** Each 27g sachet contains ~13.8g total sugars (~10g added sugars). That's more than **half the sachet's weight is sugar**, and exceeds the WHO daily added-sugar guideline of 25g in less than one cup. This is label-read, not a tool finding.
---
## Verdict
### 🟢 GREEN — No Concern
- **All three declared additives are SFA-permitted** (Dicalcium Phosphate E341, Disodium Phosphate E339, Ferric Pyrophosphate) — `check_sfa_additive` ×3.
- **Carries the Healthier Choice Symbol** — HPB has certified MILO 3in1 Activ-Go — `check_hcs`.
- **Vitamins and minerals fortified** — B2, B3, B6, B12, C, D, plus calcium and iron — label.
- **Contains barley malt extract** — a whole-grain source of energy — label.
- **Resistant Dextrin** (prebiotic dietary fibre) is a functional ingredient, not a concern — label.
### 🟡 AMBER — Worth Knowing
- **Palm Oil present** — a saturated fat source. Saturated fat is 1.2g per sachet (~4.4g per 100g of powder). Not alarming at one serving, but worth knowing if you have multiple cups a day — label.
- **"Flavourings" are not specified** — cannot determine whether natural or artificial. No SFA check was possible — label.
- **Nutri-Grade does not apply here** — this is sold as a powder. If you reconstitute it and compare the drink to ready-to-drink alternatives, it would likely land in Grade C or D territory based on sugar content, but no regulatory grade is required on the packet.
### 🔴 RED — Genuinely Concerning
- **Sugar is the #1 ingredient** — it appears first in the list, meaning it is the dominant ingredient by weight. Each sachet delivers **~13.8g total sugars / ~10g added sugars** — label. That's 40% of the WHO's recommended 25g/day added-sugar limit in a single cup. If you're drinking 2–3 sachets a day (as Nestlé's own label suggests for daily energy), you are consuming **20–30g of added sugar from MILO alone**, likely exceeding the daily limit.
- **Marketing framing vs. nutritional reality:** The packet prominently says "Nourishing Energy" and shows vitamins, but the sugar load is the dominant nutritional feature. The fortification (vitamins, minerals) does not offset this — label.
**Bottom line for a Singapore shopper:** MILO 3in1 is safe from an additives and regulatory standpoint — everything checks out with SFA and HPB. But it is a **sugar-heavy product** dressed up in health language. Treat it as an occasional treat rather than a daily nutritious breakfast drink, especially for children.
OpenAI Agents SDK
The OpenAI Agents SDK stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
@function_tool
async def check_sfa_additive(additive: str, e_number_hint: str = "") -> str:
"""Check whether a food additive is permitted by the Singapore Food Agency.
Looks the additive up in the full SFA permitted-additives list (parsed from
the official SFA PDF). Accepts either an E-number or a plain-English name as
printed on a Singapore label.
Args:
additive: An E-number ('E211', 'e211', 'en:e211') OR an additive name
('Sodium Benzoate', 'Soy Lecithin', 'MSG').
e_number_hint: Optional. If you pass a name and also know its E/INS
number, supply it here. If the name isn't recognised but the number
is, the mapping is REMEMBERED so the name resolves directly next time. """
...
class NutritionVerdict(BaseModel):
reasoning: str # declared first: the model works out loud before committing
product_name: str
summary: str
green: list[Finding]
amber: list[Finding]
red: list[Finding]
recommendation: str
client = AsyncOpenAI(base_url=BASE_URL, api_key=API_KEY)
model = OpenAIChatCompletionsModel(model=OPENAI_MODEL, openai_client=client)
agent = Agent(name="SG Nutrition Investigator", model=model,
instructions=INSTRUCTIONS, output_type=NutritionVerdict,
tools=[recall_product, check_sfa_additive, check_hcs, calculate_nutri_grade])
user_input = [{"role": "user", "content": [
{"type": "input_text", "text": "Investigate this nutrition label."},
{"type": "input_image", "image_url": _data_url(image_path)}, # base64 data URL
]}]
result = await Runner.run(agent, input=user_input)
verdict: NutritionVerdict = result.final_output
sausheong openai-demo % uv run python main.py ../milo.jpeg
Investigating image: ../milo.jpeg
────────────────────────────────────────────────────────────
Product: Nestlé MILO 3 in 1 Activ-Go
Reasoning: I read the sachet as **Nestlé MILO 3 in 1 Activ-Go**. The front says “goodness of malt, milk & cocoa” and the back says “MILO Ionic Food Drink,” so this is a **powdered drink mix**, not a ready-to-drink beverage. Because it is not sold as a beverage with per-100ml nutrition for the prepared drink as consumed, I did **not** calculate Nutri-Grade; that tool is only for beverages.
I first checked whether this product had been investigated before. A prior record existed, and it broadly matched what I can read here: a sweetened powdered malt-cocoa-milk drink sachet with permitted phosphate additives, relatively high in sugar for a single sachet, and not found in the HCS database. I used that only as background and then re-read this label fresh.
From the nutrition panel, I can read:
- Serving size: **1 sachet (27g)**
- Servings per package: **1**
- Per 100g / per serving shown side by side
- **Total sugars: 50.1g per 100g; 13.8g per serving**
- **Saturated fat: 0.5g per 100g; 1.2g per serving**
This is a powder, so the relevant direct reading is per 100g, not per 100ml.
I can read the ingredient list as:
**Sugar, Protomalt™ Malt Extract (contains barley), skimmed milk/cow’s milk, cocoa, palm oil, minerals (dicalcium phosphate, disodium phosphate), fibre (pyrophosphate), resistant dextrin, flavourings, vitamins (C, B3, D, B6, B2, B12).**
For additives, I checked each identifiable additive individually where possible:
- **Dicalcium phosphate**: checked and found **permitted by SFA**. That lands in GREEN.
- **Disodium phosphate**: checked and found **permitted by SFA**. That also lands in GREEN.
- The label also contains the broader terms **“phosphates”** and **“pyrophosphate”/“fibre (pyrophosphate)”** in wording around ingredients, but the image does not clearly identify a specific additive code for the pyrophosphate entry, and the generic umbrella term “phosphates” is not reliable enough on its own. I therefore do not treat the generic term as confirmed permitted just from the broad word alone. I only rely on the specifically named phosphates I could check.
- **Flavourings** is too generic to identify a specific additive, so I could not check it individually.
- The vitamins and minerals are nutrients rather than red-flag food additives in the SFA additive sense.
I also checked for the **Healthier Choice Symbol** using the product name. It was **not found** in the HCS database.
Main nutritional takeaways from the label:
- Sugar is the **first ingredient**, which usually means it is the largest ingredient by weight.
- The product contains **50.1g sugar per 100g** powder, and **13.8g sugar per 27g sachet**. That is quite a lot for one small sachet drink.
- Saturated fat is **1.2g per sachet**, which is not extreme, but the bigger concern here is sugar.
So the overall picture is: the identifiable phosphate additives appear permitted, but nutritionally this is still a **sweetened convenience drink mix** rather than a lower-sugar everyday option.
Summary: This is a powdered MILO drink sachet, not a ready-to-drink beverage. The clearly identifiable phosphate additives on the label are permitted by SFA, but sugar is the first ingredient and the product contains 50.1g sugar per 100g, or 13.8g per 27g sachet. It was not found in the Healthier Choice Symbol database.
🟢 Dicalcium phosphate is permitted by SFA. [check_sfa_additive]
🟢 Disodium phosphate is permitted by SFA. [check_sfa_additive]
🟢 This is a powdered drink mix rather than a ready-to-drink beverage, so Nutri-Grade was not applicable here. [label]
🟡 Generic terms such as “flavourings” and the unclear “fibre (pyrophosphate)” wording are not specific enough to confirm exact additive identities from the image alone. [label]
🟡 The product was not found in the Healthier Choice Symbol database. [check_hcs]
🟡 Saturated fat is 1.2g per 27g sachet (0.5g per 100g listed on the label panel image). [label]
🔴 Sugar is the first ingredient on the label. [label]
🔴 Total sugars are 50.1g per 100g and 13.8g per 27g sachet, which is high for a small single-serve drink mix. [label]
Recommendation: Fine as an occasional sachet drink if convenience matters, but not the best everyday choice if you are trying to reduce sugar. If you drink MILO often, consider using plain MILO powder with your own milk and less sugar, or choose a lower-sugar version.
Google’s Agent Development Kit (ADK)
The Google s Agent Development stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Google s Agent Development stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
async def check_sfa_additive(additive: str, e_number_hint: str = "") -> dict:
"""Check whether a food additive is permitted under Singapore Food Agency rules.
Looks the additive up in the full SFA permitted-additives list (parsed from
the official SFA PDF). Accepts either an E-number or a plain-English name as
printed on a Singapore label.
Args:
additive: An E-number ('E211', 'e211', 'en:e211') OR an additive name
('Sodium Benzoate', 'Soy Lecithin', 'MSG').
e_number_hint: Optional E/INS number for when you pass a name. If the
name isn't recognised but the number is, the mapping is REMEMBERED
so the name resolves directly next time.
Returns:
dict with the input, resolved e_number/name, permitted status, schedule
and any consumer note. found=False if it is not in the SFA list.
"""
...
adk-demo/
sg_nutrition_investigator/
__init__.py # from . import agent
agent.py # tools + root_agent
main.py # standalone Runner + InMemorySessionService
root_agent = LlmAgent(
name="sg_nutrition_investigator",
model=LiteLlm(model="openai/" + GEMINI_MODEL, api_base=BASE_URL, api_key=API_KEY),
instruction=INSTRUCTION,
tools=[FunctionTool(recall_product), FunctionTool(check_sfa_additive),
FunctionTool(check_hcs), FunctionTool(calculate_nutri_grade)],
)
# image input — the cleanest of the three:
message = types.Content(role="user", parts=[
types.Part(text="Investigate this nutrition label."),
types.Part.from_bytes(data=Path(image_path).read_bytes(), mime_type=mime),
])
sausheong adk-demo % uv run python main.py ../milo.jpeg
Investigating image: ../milo.jpeg
────────────────────────────────────────────────────────────
→ tool: recall_product({'product_name': 'Nestle MILO 3 in 1'})
→ tool: check_sfa_additive({'e_number_hint': 'E341', 'additive': 'Dicalcium Phosphate'})
→ tool: check_sfa_additive({'additive': 'Disodium Phosphate', 'e_number_hint': 'E339'})
→ tool: check_sfa_additive({'additive': 'Ferric Pyrophosphate'})
→ tool: check_hcs({'product_name': 'Nestle MILO 3 in 1'})
→ tool: calculate_nutri_grade({'saturated_fat_per_100ml': 0.5, 'sugar_per_100ml': 6.1})
## Reasoning
We investigated the label of **Nestle MILO 3 in 1 (with Activ-Go)**. Here is our step-by-step analysis:
1. **Product Categorisation & Reconstitution**:
* This is a dry beverage powder sold in single-serving 27g sachets.
* The nutritional information panel provides two columns: "Per Serving 27g" (the dry sachet) and "Per 100ml*" (as reconstituted by adding 200ml of hot water).
* For the **Singapore Nutri-Grade** system, the grading is calculated based on the beverage *as reconstituted* according to the directions.
2. **Sugar & Saturated Fat Analysis (Per 100ml prepared)**:
* **Total Sugars**: **6.1g per 100ml** (of which 4.8g is Added Sugar).
* **Saturated Fat**: **0.5g per 100ml**.
* Feeding these values into the Nutri-Grade calculator yields **Nutri-Grade C** (due to the sugar content being > 5g/100ml). Grade C beverages are classified as "less healthy" and must display the Nutri-Grade label at retail.
3. **Ingredients & Additives Analysis**:
* **Sugar is the #1 Ingredient**: Refined sugar is listed first on the ingredients list, meaning it is the largest ingredient by weight in this packet—even ahead of malt extract and milk powder.
* **Additives**:
* **Dicalcium Phosphate** (E341(ii)) and **Disodium Phosphate** (E339(ii)) are SFA-approved food additives used as stabilizers and emulsifiers.
* **Ferric Pyrophosphate** is an SFA-approved mineral nutrient used to fortify the drink with Iron.
* **Healthier Choice Symbol (HCS)**: The packet displays the red HCS logo. This is because it is formulated to be "25% lower in sugar" *compared to regular malted beverages*. However, in absolute terms, it still carries enough sugar to land in Nutri-Grade C.
---
## Verdict
### 🟢 Green Findings
* **Fortified with Vitamins & Minerals** `[label]`: Contains B-vitamins (B2, B3, B6, B12), Vitamin C, Vitamin D, Calcium, and Iron (via *Ferric Pyrophosphate* `[check_sfa_additive]`).
* **Contains Soluble Fibre** `[label]`: Features *Resistant Dextrin*, a prebiotic soluble dietary fibre (1.1g per serving) that supports gut health.
* **SFA-Approved Additives** `[check_sfa_additive]`: Phosphates used for stabilizing and emulsifying (*Dicalcium Phosphate* and *Disodium Phosphate*) are fully permitted.
* **Healthier Choice Certified** `[check_hcs]`: Carries Singapore's HCS logo for having at least 25% less sugar than standard malted chocolate drinks.
### 🟡 Amber Findings
* **Nutri-Grade C** `[calculate_nutri_grade]`: When prepared as instructed, it gets a **Grade C** because of its **6.1g of sugar per 100ml**. Grade C drinks are in the "less healthy" bracket.
* **Contains Palm Oil** `[label]`: Palm oil is added to give a rich mouthfeel, contributing **1.2g of saturated fat per serving** (0.5g/100ml).
### 🔴 Red Findings
* **Sugar is the Main Ingredient** `[label]`: Sugar is listed as the very first ingredient on the label, which means there is more plain sugar by weight than malt extract (*PROTOMALT*) or milk powder.
* **High Sugar Load per Cup** `[label]`: A single small 27g sachet contains **13.8g of total sugar** (equivalent to about **3 teaspoons of sugar**), of which **10.8g** (approx. 2.5 teaspoons) is added refined sugar. Drinking multiple cups a day can quickly max out your recommended daily limit for added sugars.
Running it yourself
For the Running it yourself stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
cd openai-demo # or claude-demo, or adk-demo
uv run python main.py ../milo.jpeg
uv run python main.py ../hl.jpeg
Going beyond a single shot
For the Going beyond a single stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
What each SDK provides, and what it doesn’t
For the What each SDK provides stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the What each SDK provides stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What is missing from the SDK?
When working through the What is missing from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
So which SDK should you use?
When working through the So which SDK should stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Conclusion
When working through the Conclusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Conclusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 5df04c582f40: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 2 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 3 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 4 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 4/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 5 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 5/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 6 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 6/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 7 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 7/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 8 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 8/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 9 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 9/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 10 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 10/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 11 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 11/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 12 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 12/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 13 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 13/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 14 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 14/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 15 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 15/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 16 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 16/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 17 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 17/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 18 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 18/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 19 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 19/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 20 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 20/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 21 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 21/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 22 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 22/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 23 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 23/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 24 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 24/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 25 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 25/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 26 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 26/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 27 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 27/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 28 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 28/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 29 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 29/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 30 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 30/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 31 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 31/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 32 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 32/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 33 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 33/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 34 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 34/815: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.