Home / Articles / Practical notes: Keeping an Autonomous Agent Inside the Lines

This article is published in English.

Practical notes: Keeping an Autonomous Agent Inside the Lines

Operable walkthrough of Practical notes: Keeping an Autonomous Agent Inside the Lines: contracts, checks, and drop-in code slots for teams shipping this pattern.

1537 words

Use this as an operator-facing rebuild of the ideas in “Keeping an Autonomous Agent Inside the Lines”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Why a plain string check is not enough

For the Why a plain string stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

curl http://localhost/exec?cmd=ping%20192.168.2.1

Peeling the layers off before judging anything

For the Peeling the layers off stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

def _de_cloak_payloads(cmd: str, depth: int = 0, max_depth: int = 3) -> set[str]:
    """
    Recursively searches for base64, hex, and URL encodings inside cmd.
    Returns a set of all extracted/decoded plain strings.
    """
    extracted = {cmd}
    if depth >= max_depth:
        return extracted

    # 1. URL Decoding
    decoded_url = urllib.parse.unquote(cmd)
    if decoded_url != cmd:
        extracted.update(_de_cloak_payloads(decoded_url, depth + 1, max_depth))

    # 2. Hex escape and raw hex string decoding
    for match in re.findall(r"(?:\\x[0-9a-fA-F]{2})+", cmd):
        hex_bytes = bytes.fromhex(match.replace("\\x", ""))
        extracted.update(_de_cloak_payloads(hex_bytes.decode("utf-8", errors="ignore"), depth + 1, max_depth))

    # 3. Base64 decoding
    for match in re.findall(r"\b[A-Za-z0-9+/]{12,}={0,2}\b", cmd):
        padded = match + "=" * ((4 - len(match) % 4) % 4)
        decoded = base64.b64decode(padded.encode("ascii")).decode("utf-8", errors="ignore")
        extracted.update(_de_cloak_payloads(decoded, depth + 1, max_depth))

    return extracted

The same trick works on numbers, not just text

For the The same trick works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The same trick works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

def _normalize_ip_token(token: str) -> str | None:
    token = token.strip().lower()

    # A bare integer or hex integer standing in for a full IP
    if token.isdigit() or (token.startswith("0x") and all(c in "0123456789abcdef" for c in token[2:])):
        val = int(token, 16) if token.startswith("0x") else int(token)
        if 0 <= val <= 0xFFFFFFFFFFFFFFFFFFFFFFFFFFFFFFFF:
            if val > 1024 or val in (0,):
                return str(ipaddress.ip_address(val))

    # Dotted octal or dotted hex, one part at a time
    if "." in token:
        parts = token.split(".")
        if len(parts) == 4:
            normalized_parts = []
            for p in parts:
                val = int(p, 16) if p.startswith("0x") else int(p, 8) if p.startswith("0") and len(p) > 1 else int(p)
                if 0 <= val <= 255:
                    normalized_parts.append(str(val))
            if len(normalized_parts) == 4:
                return ".".join(normalized_parts)
    return None

Domains need a different kind of care, in the other direction

When working through the Domains need a different stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

tld = token.rsplit(".", 1)[-1]
if tld in _FILE_EXTENSIONS:
    continue
if not any(token == d or token.endswith("." + d) for d in self.domains):
    raise ScopeViolation(f"Domain {token!r} is outside engagement scope.")

Putting it together

When working through the Putting it together stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

def check(self, cmd: str) -> None:
    payloads = _de_cloak_payloads(cmd)

    for payload in payloads:
        if self.networks:
            for m in _IPV4_RE.finditer(payload):
                self._validate_ip(m.group(1))
            for token in re.split(r"[\s\"'$,;()|&<>`\\/]", payload):
                normalized = _normalize_ip_token(token)
                if normalized:
                    self._validate_ip(normalized)

        if self.domains:
            for m in _DOMAIN_RE.finditer(payload):
                token = m.group(0).lower()
                tld = token.rsplit(".", 1)[-1]
                if tld in _FILE_EXTENSIONS:
                    continue
                if not any(token == d or token.endswith("." + d) for d in self.domains):
                    raise ScopeViolation(f"Domain {token!r} is outside engagement scope.")

What’s next

When working through the What s next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What s next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 373547c620fe: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 0/771: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 1/771: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.