Home / Articles / Practical notes: Engineering Enterprise AI Agent Harnesses

This article is published in English.

Practical notes: Engineering Enterprise AI Agent Harnesses

Operable walkthrough of Practical notes: Engineering Enterprise AI Agent Harnesses: contracts, checks, and drop-in code slots for teams shipping this pattern.

6175 words

The following notes reconstruct a practical path around “Engineering Enterprise AI Agent Harnesses”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

1. Model, agent, and harness are different things

The 1 Model agent and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Observe → Reason → Validate → Act → Observe

2. Why enterprise systems need a harness

The 2 Why enterprise systems stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

3. The enterprise agent harness architecture

The 3 The enterprise agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

User and application boundary

The User and application boundary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Planner and state manager

The Planner and state manager stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Model runtime

The Model runtime stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Tool registry

The Tool registry stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Policy engine

The Policy engine stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Policy engine stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Guardrails

For the Guardrails stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Human approval gate

For the Human approval gate stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Tool executor and sandbox

For the Tool executor and sandbox stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Tool executor and sandbox stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Memory and data

When working through the Memory and data stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Observability and traces

When working through the Observability and traces stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Evaluations and feedback

When working through the Evaluations and feedback stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Evaluations and feedback stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

4. Step-by-step execution of an agent run

The 4 Step-by-step execution of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Step 1 — Receive and classify the goal

The Step 1 Receive and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Step 2 — Build bounded context

The Step 2 Build bounded stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Step 2 Build bounded stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Step 3 — Ask the model for the next action

For the Step 3 Ask the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Step 4 — Validate outside the model

For the Step 4 Validate outside stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Step 5 — Obtain approval when required

For the Step 5 Obtain approval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 5 Obtain approval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Step 6 — Execute in a controlled environment

When working through the Step 6 Execute in stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Step 7 — Normalize the observation

When working through the Step 7 Normalize the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Step 8 — Continue or stop

When working through the Step 8 Continue or stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Step 8 Continue or stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Step 9 — Produce the final answer and trace

The Step 9 Produce the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

5. Example: an accounts-receivable agent

The 5 Example an accounts-receivable stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

6. Failure modes the harness must prevent

The 6 Failure modes the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Prompt-only security

The Prompt-only security stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Direct model-to-tool wiring

The Direct model-to-tool wiring stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Overpowered shared credentials

The Overpowered shared credentials stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Untrusted content becoming instructions

The Untrusted content becoming instructions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Unlimited loops

The Unlimited loops stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Unlimited loops stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Logging everything

For the Logging everything stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Measuring only fluent answers

For the Measuring only fluent answers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

7. Set up a local AI harness environment

For the 7 Set up a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the 7 Set up a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Prerequisites

When working through the Prerequisites stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

mkdir enterprise-agent-harness
cd enterprise-agent-harness
python -m venv .venv
source .venv/bin/activate
.venv\Scripts\Activate.ps1
python -m pip install openai pydantic fastapi uvicorn python-dotenv

Keep configuration separate from secrets

When working through the Keep configuration separate from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

AI_PROVIDER=openai
AI_MODEL=<approved-model-id>
OPENAI_API_KEY=<set-locally-never-commit>
HARNESS_ENV=development
HARNESS_MAX_STEPS=8
HARNESS_RUN_TIMEOUT_SECONDS=120
HARNESS_MAX_TOOL_RETRIES=2
HARNESS_REQUIRE_APPROVAL_FOR_WRITES=true

Use a layered project structure

When working through the Use a layered project stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

enterprise-agent-harness/
├── app/
│   ├── api.py                 # authenticated HTTP entry point
│   ├── harness.py             # observe/reason/validate/act loop
│   ├── model_adapter.py       # provider-specific model calls
│   ├── schemas.py             # typed proposals and observations
│   ├── tools/
│   │   ├── registry.py        # available capabilities
│   │   ├── invoices.py        # example read-only tool
│   │   └── email.py           # example external-write tool
│   ├── policy/
│   │   ├── engine.py          # allow/deny/require-approval
│   │   └── rules.yaml         # reviewed declarative policy
│   ├── approvals.py           # exact-action approval records
│   ├── sandbox.py             # isolated execution adapter
│   ├── state.py               # run and conversation state
│   ├── telemetry.py           # sanitized events and traces
│   └── redaction.py           # secret and sensitive-data filtering
├── evals/
│   ├── cases.jsonl            # representative tasks and attacks
│   └── run_evals.py
├── tests/
│   ├── test_policy.py
│   ├── test_tools.py
│   └── test_harness.py
├── Dockerfile
├── compose.yaml
├── .env
.example
└── pyproject.toml

When working through the Use a layered project stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Define the autonomy boundary first

The Define the autonomy boundary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

8. Build a runnable reference harness

The 8 Build a runnable stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Define typed contracts

The Define typed contracts stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

from enum import Enum
from typing import Any, Literal
from pydantic import BaseModel, Field
class Risk(str, Enum):
READ_ONLY = "read_only"
SENSITIVE_READ = "sensitive_read"
EXTERNAL_WRITE = "external_write"
DESTRUCTIVE = "destructive"
class ToolProposal(BaseModel):
tool: str
arguments: dict[str, Any]
reason: str
class PolicyDecision(BaseModel):
outcome: Literal["allow", "deny", "require_approval"]
reason_code: str
explanation: str
class Observation(BaseModel):
tool: str
ok: bool
data: dict[str, Any] = Field(default_factory=dict)
error_code: str | None = None
class RunState(BaseModel):
run_id: str
user_id: str
tenant_id: str
goal: str
step: int = 0
observations: list[Observation] = Field(default_factory=list)
finished: bool = False

The Define typed contracts stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Register tools with security metadata

For the Register tools with security stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

from dataclasses import dataclass
from typing import Callable
@dataclass(frozen=True)
class ToolDefinition:
name: str
risk: Risk
handler: Callable[..., dict]
timeout_seconds: int
idempotent: bool
TOOLS = {
"list_overdue_invoices": ToolDefinition(
name="list_overdue_invoices",
risk=Risk.SENSITIVE_READ,
handler=list_overdue_invoices,
timeout_seconds=10,
idempotent=True,
),
"send_email": ToolDefinition(
name="send_email",
risk=Risk.EXTERNAL_WRITE,
handler=send_email,
timeout_seconds=15,
idempotent=False,
),
}

Implement deterministic policy

For the Implement deterministic policy stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

def authorize(state: RunState, proposal: ToolProposal) -> PolicyDecision:
definition = TOOLS.get(proposal.tool)
if definition is None:
return PolicyDecision(
outcome="deny",
reason_code="UNKNOWN_TOOL",
explanation="The requested capability is not registered.",
)
if not identity_can_use_tool(
user_id=state.user_id,
tenant_id=state.tenant_id,
tool=definition.name,
arguments=proposal.arguments,
):
return PolicyDecision(
outcome="deny",
reason_code="NOT_AUTHORIZED",
explanation="The caller lacks permission for this resource.",
)if definition.risk in {Risk.EXTERNAL_WRITE, Risk.DESTRUCTIVE}:
return PolicyDecision(
outcome="require_approval",
reason_code="CONSEQUENTIAL_ACTION",
explanation="The exact action must be approved before execution.",
)return PolicyDecision(
outcome="allow",
reason_code="POLICY_ALLOWED",
explanation="The action is permitted within the caller's scope.",
)

Bind approval to the exact action

For the Bind approval to the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Bind approval to the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

import hashlib
import json
def action_digest(state: RunState, proposal: ToolProposal) -> str:
value = {
"user_id": state.user_id,
"tenant_id": state.tenant_id,
"tool": proposal.tool,
"arguments": proposal.arguments,
}
canonical = json.dumps(value, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canonical.encode("utf-8")).hexdigest()

Implement the controlled loop

When working through the Implement the controlled loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

async def run_harness(state: RunState, model, approvals, executor):
max_steps = 8
while not state.finished and state.step < max_steps:
state.step += 1context = build_bounded_context(state, TOOLS)
response = await model.next_action(context)if response.final_answer is not None:
state.finished = True
return complete_run(state, response.final_answer)proposal = ToolProposal.model_validate(response.tool_proposal)
decision = authorize(state, proposal)
trace_policy_decision(state, proposal, decision)if decision.outcome == "deny":
state.observations.append(Observation(
tool=proposal.tool,
ok=False,
error_code=decision.reason_code,
))
continueif decision.outcome == "require_approval":
digest = action_digest(state, proposal)
if not approvals.has_valid_approval(digest):
return pause_for_approval(state, proposal, digest)observation = await executor.execute(
definition=TOOLS[proposal.tool],
arguments=proposal.arguments,
identity={
"user_id": state.user_id,
"tenant_id": state.tenant_id,
},
)
state.observations.append(redact_observation(observation))return stop_run(state, reason="STEP_LIMIT_REACHED")

Connect the model through an adapter

When working through the Connect the model through stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

class ModelAdapter:
async def next_action(self, context) -> "ModelTurn":
raise NotImplementedError

Expose an authenticated API

When working through the Expose an authenticated API stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

from fastapi import Depends, FastAPI
app = FastAPI()
@app.post("/runs")
async def create_run(request: RunRequest, identity=Depends(require_identity)):
state = RunState(
run_id=new_run_id(),
user_id=identity.user_id,
tenant_id=identity.tenant_id,
goal=request.goal,
)
return await run_harness(state, model, approvals, executor)
@app.post("/runs/{run_id}/approvals")
async def approve_run_action(
run_id: str,
request: ApprovalRequest,
identity=Depends(require_identity),
):
return approve_exact_action(run_id, request.action_digest, identity)

When working through the Expose an authenticated API stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Run locally

The Run locally stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

uvicorn app.api:app --host 127.0.0.1 --port 8000 --reload

9. Test and evaluate the harness

The 9 Test and evaluate stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Deterministic unit and integration tests

The Deterministic unit and integration stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Deterministic unit and integration stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

def test_external_write_requires_approval():
proposal = ToolProposal(
tool="send_email",
arguments={"to": "customer@example.test", "body": "Reminder"},
reason="Send the approved reminder",
)
decision = authorize(make_finance_run(), proposal)assert decision.outcome == "require_approval"
assert decision.reason_code == "CONSEQUENTIAL_ACTION"

Model and workflow evals

For the Model and workflow evals stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

10. Deploy the harness at enterprise scale

For the 10 Deploy the harness stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Production topology

For the Production topology stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Client
│
▼
API Gateway / WAF
│  authentication, rate limits, request ID
▼
Harness API
├── Policy Service
├── Approval Service
├── Model Gateway ─────► Model Provider
├── State Store
├── Trace / Audit Pipeline
└── Tool Executor Queue
│
▼
Isolated Workers
├── MCP / SaaS APIs
├── Browser Sandbox
├── Code Sandbox
└── Enterprise Data

For the Production topology stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Containerize the control plane

When working through the Containerize the control plane stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

FROM python:3.12-slim
WORKDIR /app
COPY pyproject.toml ./
RUN pip install --no-cache-dir .COPY app ./appRUN useradd --create-home --uid 10001 harness
USER harnessENV PYTHONDONTWRITEBYTECODE=1
ENV PYTHONUNBUFFERED=1CMD ["uvicorn", "app.api:app", "--host", "0.0.0.0", "--port", "8000"]

Production controls

When working through the Production controls stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Development-to-production path

When working through the Development-to-production path stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Development-to-production path stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

11. Enterprise readiness checklist

The 11 Enterprise readiness checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Conclusion

The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

References

Operational checklist