This article is published in English.
Planner, Generator, Healer: How I Ran a Full Playwright Suite With Three AI
Operable walkthrough of Planner, Generator, Healer: How I Ran a Full Playwright Suite With Three AI: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Planner, Generator, Healer: How I Ran a Full Playwright Suite With Three AI Agents And Playwright MCP”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The setup that makes agents useful
For the The setup that makes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
{
"servers": {
"playwright-test": {
"type": "stdio",
"command": "npx",
"args": ["playwright", "run-test-mcp-server"]
}
}
}
Agent 1 — the Planner: “What is worth testing?”
For the Agent 1 the Planner stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Agent 2 — the Generator: “Make this one scenario real”
For the Agent 2 the Generator stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Agent 2 the Generator stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
// spec: specs/checkout-e2e-plan.md
// seed: tests/web/seed.spec.js
import { test, expect } from '../../../fixtures';
import { users } from '../../../testdata';test.describe.configure({ mode: 'parallel' });test.describe('E2E checkout journey', () => {
test('positive: login → add product → checkout → thank you → home', async ({ flow }) => {
await flow.completeHappyPathPurchase();
});
});
async completeHappyPathPurchase(user = users.standard, info = checkout.valid) {
await this.goToCheckoutInformation(user);
await this.fillValidInformation(info);
await this.continueToOverview();
await this.app.checkoutOverview.expectProductVisible(products.first.name);
await this.finishOrder();
await this.backHomeToProducts();
}
Agent 3 — the Healer: “It broke. Fix the right layer.”
When working through the Agent 3 the Healer stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
How to actually run it, start to finish
When working through the How to actually run stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
1. Copy .env.example → .env, set BASE_URL (+ credentials)
2. npm install && npx playwright install
3. Run the Planner
→ specs/checkout-e2e-plan.md (22 scenarios, reviewed by me)4. Run the Generator, scenario by scenario
→ tests/web/login/login.spec.js
→ tests/web/cart/cart.spec.js
→ tests/web/checkout/checkout.spec.js
→ tests/web/e2e/checkout-journey.spec.js5. npm test (4 workers, fully parallel)6. Any red? Run the Healer → re-run → green
So how much faster is it, really?
When working through the So how much faster stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the So how much faster stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Four lessons worth stealing
The Four lessons worth stealing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Who should try this
The Who should try this stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for ebfa232e6bf7: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 0/959: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.