Home / Articles / Practical notes: I Built an Autonomous SRE Agent That Resolves Production

This article is published in English.

Practical notes: I Built an Autonomous SRE Agent That Resolves Production

Operable walkthrough of Practical notes: I Built an Autonomous SRE Agent That Resolves Production: contracts, checks, and drop-in code slots for teams shipping this pattern.

2339 words

This walkthrough rebuilds the path from raw materials to a working system for: I Built an Autonomous SRE Agent That Resolves Production Incidents While I Sleep. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

GitHub Repo

When working through the GitHub Repo stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Why ReAct, and Why LangGraph

When working through the Why ReAct and Why stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

The Four Memory Layers

When working through the The Four Memory Layers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Four Memory Layers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

How episodic memory actually works

The How episodic memory actually stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The Autonomy System: L0 Through L3

The The Autonomy System L0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The Full Tool Arsenal

The The Full Tool Arsenal stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The The Full Tool Arsenal stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

The ReAct Loop in Action

For the The ReAct Loop in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

[MEMORY] Loading episodic context for payment-service...
[MEMORY] Found 2 similar past incidents (similarity: 0.89, 0.76)
[MEMORY] Loaded 3 procedural policies for payment-service/critical

[THOUGHT] Incident: payment-service OOMKill, 3 pods CrashLooping.
Past incident INC-2801 had identical symptoms — root cause was
a memory leak in the gRPC connection pool after a deploy.
Starting with get_logs to confirm before acting.

[ACTION] Calling tool: get_logs with {'service_name': 'payment-service', 'log_level': 'error'}
→ top_errors: [{"pattern": "out of memory: kill process", "count": 47},
{"pattern": "grpc connection pool exhausted", "count": 23}]

[THOUGHT] OOM confirmed. gRPC pool exhaustion matches INC-2801 pattern.
Retrieve runbook before acting — want the fix procedure.

[ACTION] Calling tool: retrieve_runbook with {'query': 'payment service grpc oom memory leak fix'}
→ [runbooks / payment-service-oom.md] Recommended action: rolling restart
resets connection pool. If recurs within 1h, rollback to previous image.

[THOUGHT] Runbook confirms rolling restart. Checking policy for autonomy L2...

[ACTION] Calling tool: restart_deployment with {'service': 'payment-service',
'deployment': 'payment-worker', 'namespace': 'production',
'severity': 'critical', 'autonomy_level': 'L2'}
→ policy: ALLOWED (L2 ≥ min_level L2 for restart_deployment)
→ draining pods... restarting... polling rollout...
→ health_check: healthy=True, pods_ready=3/3, error_rate=0.008

[THOUGHT] Restart succeeded. Error rate 0.8% — below 5% threshold.
Verifying with check_service_health before closing.

[ACTION] Calling tool: check_service_health with {'service': 'payment-service',
'deployment': 'payment-worker', 'namespace': 'production'}
→ healthy=True, pods_total=3, pods_ready=3, error_rate=0.007

[THOUGHT] Service fully recovered. Writing audit log and resolving.

[ACTION] Calling tool: write_audit_log ...
→ audit_log written: restart_deployment for incident INC-2847

RESOLVED: payment-service recovered after rolling restart. Root cause: gRPC
connection pool OOM — matched pattern from INC-2801. All 3 pods healthy,
error rate 0.7%. MTTR: 187 seconds.

[MEMORY] Saving incident INC-2847 to episodic memory...
[NOTIFY] Sending resolution to Slack channel: #all-aegisops ✅
Invoke-RestMethod -Uri "http://localhost:8000/webhook/incident" `
  -Method Post `
  -ContentType "application/json" `
  -Body '{
    "source": "webhook",
    "severity": "critical",
    "service": "payment-service",
    "symptoms": ["High 500 errors", "DB connection timeouts"],
    "autonomy_level": "L2"
  }'

What the Slack Channel Looks Like

For the What the Slack Channel stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The Infrastructure Stack

For the The Infrastructure Stack stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The Infrastructure Stack stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

version: "3.9"

services:
  postgres:
    image: ankane/pgvector:latest
    ports:
      - "5445:5432"
    environment:
      POSTGRES_USER: postgres
      POSTGRES_PASSWORD: admin123
      POSTGRES_DB: aegisops
    volumes:
      - postgres_data:/var/lib/postgresql/data

  redis:
    image: redis:alpine
    ports:
      - "6380:6379"

  app:
    build: .
    ports:
      - "8000:8000"
    env_file: .env
    depends_on:
      - postgres
      - redis

  prometheus:
    image: prom/prometheus:latest
    ports:
      - "9090:9090"
    volumes:
      - ./infra/prometheus.yml:/etc/prometheus/prometheus.yml

  grafana:
    image: grafana/grafana:latest
    ports:
      - "3000:3000"

volumes:
  postgres_data:

The Audit Trail

When working through the The Audit Trail stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

What’s Next

When working through the What s Next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Final Thoughts

When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

About This Project

The About This Project stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 102175b0ac7c: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 0/800: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 1/800: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.