This article is published in English.
Practical notes: I Built an Autonomous SRE Agent That Resolves Production
Operable walkthrough of Practical notes: I Built an Autonomous SRE Agent That Resolves Production: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: I Built an Autonomous SRE Agent That Resolves Production Incidents While I Sleep. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
GitHub Repo
When working through the GitHub Repo stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Why ReAct, and Why LangGraph
When working through the Why ReAct and Why stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The Four Memory Layers
When working through the The Four Memory Layers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Four Memory Layers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
How episodic memory actually works
The How episodic memory actually stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The Autonomy System: L0 Through L3
The The Autonomy System L0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The Full Tool Arsenal
The The Full Tool Arsenal stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The The Full Tool Arsenal stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The ReAct Loop in Action
For the The ReAct Loop in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
[MEMORY] Loading episodic context for payment-service...
[MEMORY] Found 2 similar past incidents (similarity: 0.89, 0.76)
[MEMORY] Loaded 3 procedural policies for payment-service/critical
[THOUGHT] Incident: payment-service OOMKill, 3 pods CrashLooping.
Past incident INC-2801 had identical symptoms — root cause was
a memory leak in the gRPC connection pool after a deploy.
Starting with get_logs to confirm before acting.
[ACTION] Calling tool: get_logs with {'service_name': 'payment-service', 'log_level': 'error'}
→ top_errors: [{"pattern": "out of memory: kill process", "count": 47},
{"pattern": "grpc connection pool exhausted", "count": 23}]
[THOUGHT] OOM confirmed. gRPC pool exhaustion matches INC-2801 pattern.
Retrieve runbook before acting — want the fix procedure.
[ACTION] Calling tool: retrieve_runbook with {'query': 'payment service grpc oom memory leak fix'}
→ [runbooks / payment-service-oom.md] Recommended action: rolling restart
resets connection pool. If recurs within 1h, rollback to previous image.
[THOUGHT] Runbook confirms rolling restart. Checking policy for autonomy L2...
[ACTION] Calling tool: restart_deployment with {'service': 'payment-service',
'deployment': 'payment-worker', 'namespace': 'production',
'severity': 'critical', 'autonomy_level': 'L2'}
→ policy: ALLOWED (L2 ≥ min_level L2 for restart_deployment)
→ draining pods... restarting... polling rollout...
→ health_check: healthy=True, pods_ready=3/3, error_rate=0.008
[THOUGHT] Restart succeeded. Error rate 0.8% — below 5% threshold.
Verifying with check_service_health before closing.
[ACTION] Calling tool: check_service_health with {'service': 'payment-service',
'deployment': 'payment-worker', 'namespace': 'production'}
→ healthy=True, pods_total=3, pods_ready=3, error_rate=0.007
[THOUGHT] Service fully recovered. Writing audit log and resolving.
[ACTION] Calling tool: write_audit_log ...
→ audit_log written: restart_deployment for incident INC-2847
RESOLVED: payment-service recovered after rolling restart. Root cause: gRPC
connection pool OOM — matched pattern from INC-2801. All 3 pods healthy,
error rate 0.7%. MTTR: 187 seconds.
[MEMORY] Saving incident INC-2847 to episodic memory...
[NOTIFY] Sending resolution to Slack channel: #all-aegisops ✅
Invoke-RestMethod -Uri "http://localhost:8000/webhook/incident" `
-Method Post `
-ContentType "application/json" `
-Body '{
"source": "webhook",
"severity": "critical",
"service": "payment-service",
"symptoms": ["High 500 errors", "DB connection timeouts"],
"autonomy_level": "L2"
}'
What the Slack Channel Looks Like
For the What the Slack Channel stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The Infrastructure Stack
For the The Infrastructure Stack stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The Infrastructure Stack stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
version: "3.9"
services:
postgres:
image: ankane/pgvector:latest
ports:
- "5445:5432"
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: admin123
POSTGRES_DB: aegisops
volumes:
- postgres_data:/var/lib/postgresql/data
redis:
image: redis:alpine
ports:
- "6380:6379"
app:
build: .
ports:
- "8000:8000"
env_file: .env
depends_on:
- postgres
- redis
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./infra/prometheus.yml:/etc/prometheus/prometheus.yml
grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
volumes:
postgres_data:
The Audit Trail
When working through the The Audit Trail stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
What’s Next
When working through the What s Next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Final Thoughts
When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
About This Project
The About This Project stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 102175b0ac7c: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/800: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/800: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.