This article is published in English.
Self-Healing Kubernetes: Wiring OpenTelemetry, SigNoz, and a Groq-Powered
Operable walkthrough of Self-Healing Kubernetes: Wiring OpenTelemetry, SigNoz, and a Groq-Powered: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Self-Healing Kubernetes: Wiring OpenTelemetry, SigNoz, and a Groq-Powered Remediation Agent (PART 6)”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
cd ~/k8s-aiops-blog
mkdir -p dashboard/backend dashboard/frontend
cd ~/k8s-aiops-blog/dashboard/backend
cat > main.py <<'EOF'
"""
Approval dashboard backend.
Receives incident cards from the remediation agent, serves them to the
React frontend, and applies approved kubectl commands via the Kubernetes
Python client.
"""
import logging
import shlex
import subprocess
import time
from typing import Literal
from uuid import uuid4from fastapi import FastAPI, HTTPException
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModellogging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(name)s — %(message)s",
)
log = logging.getLogger("dashboard")app = FastAPI(title="Approval Dashboard")app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_methods=["*"],
allow_headers=["*"],
)# ---------------------------------------------------------------------------
# In-memory store. Fine for a demo; swap for SQLite or Redis in production.
# ---------------------------------------------------------------------------
INCIDENTS: dict[str, dict] = {}# ---------------------------------------------------------------------------
# kubectl command allowlist.
# The LLM proposes a fix. We only execute it if it matches one of these
# safe prefixes. Nothing with delete, drain, or --all gets through.
# ---------------------------------------------------------------------------
ALLOWED_PREFIXES = (
"kubectl set env",
"kubectl rollout restart",
"kubectl scale",
"kubectl set resources",
)BLOCKED_TERMS = ("delete", "drain", "--all", "cluster-admin", "exec", "cp")
# ---------------------------------------------------------------------------
# Models
# ---------------------------------------------------------------------------
class IncidentCard(BaseModel):
id: str
timestamp: int
service: str
signals: dict
diagnosis: dict
status: Literal["pending", "approved", "rejected"] = "pending"
class ActionRequest(BaseModel):
actor: str = "operator"
# ---------------------------------------------------------------------------
# Routes
# ---------------------------------------------------------------------------
@app.get("/healthz")
def healthz():
return {"ok": True}
@app.post("/api/incidents", status_code=201)
def receive_incident(card: IncidentCard):
INCIDENTS[card.id] = card.dict()
log.info("incident received: id=%s type=%s", card.id, card.diagnosis.get("incident_type"))
return {"id": card.id}
@app.get("/api/incidents")
def list_incidents():
# Most recent first
return sorted(INCIDENTS.values(), key=lambda c: c["timestamp"], reverse=True)
@app.get("/api/incidents/{incident_id}")
def get_incident(incident_id: str):
if incident_id not in INCIDENTS:
raise HTTPException(status_code=404, detail="not found")
return INCIDENTS[incident_id]
@app.post("/api/incidents/{incident_id}/approve")
def approve_incident(incident_id: str, req: ActionRequest):
card = INCIDENTS.get(incident_id)
if not card:
raise HTTPException(status_code=404, detail="not found")
if card["status"] != "pending":
raise HTTPException(status_code=409, detail=f"incident already {card['status']}") proposed = card["diagnosis"].get("proposed_fix", "")
log.info("approve requested: id=%s fix=%r actor=%s", incident_id, proposed, req.actor) # Validate against allowlist
allowed = any(proposed.startswith(p) for p in ALLOWED_PREFIXES)
blocked = any(term in proposed for term in BLOCKED_TERMS) if not allowed or blocked:
log.error("REJECTED by allowlist: %r", proposed)
raise HTTPException(
status_code=400,
detail=f"proposed command did not pass allowlist validation: {proposed!r}",
) # Apply the fix
try:
args = shlex.split(proposed)
result = subprocess.run(
args,
capture_output=True,
text=True,
timeout=30,
)
if result.returncode != 0:
log.error("kubectl failed: %s", result.stderr)
raise HTTPException(status_code=500, detail=f"kubectl error: {result.stderr}") log.info("kubectl succeeded: %s", result.stdout.strip())
card["status"] = "approved"
card["applied_at"] = int(time.time())
card["applied_by"] = req.actor
card["kubectl_output"] = result.stdout.strip()
INCIDENTS[incident_id] = card
return {"status": "approved", "kubectl_output": result.stdout.strip()} except subprocess.TimeoutExpired:
raise HTTPException(status_code=504, detail="kubectl timed out")
@app.post("/api/incidents/{incident_id}/reject")
def reject_incident(incident_id: str, req: ActionRequest):
card = INCIDENTS.get(incident_id)
if not card:
raise HTTPException(status_code=404, detail="not found")
if card["status"] != "pending":
raise HTTPException(status_code=409, detail=f"incident already {card['status']}") card["status"] = "rejected"
card["rejected_at"] = int(time.time())
card["rejected_by"] = req.actor
INCIDENTS[incident_id] = card
log.info("incident rejected: id=%s actor=%s", incident_id, req.actor)
return {"status": "rejected"}
EOF
cat > requirements.txt <<'EOF'
fastapi==0.115.0
uvicorn[standard]==0.32.0
pydantic==2.9.2
httpx==0.27.2
EOF
cat > Dockerfile <<'EOF'
FROM python:3.11-slim
# Install kubectl so the backend can apply fixes
RUN apt-get update && apt-get install -y curl && \
curl -LO "https://dl.k8s.io/release/$(curl -Ls https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl" && \
install -o root -g root -m 0755 kubectl /usr/local/bin/kubectl && \
rm kubectl && \
apt-get cleanWORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY main.py .EXPOSE 9000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "9000"]
EOF
cd ~/k8s-aiops-blog/dashboard/frontend
cat > index.html <<'EOF'
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0"/>
<title>SRE Approval Dashboard</title>
<script src="https://cdnjs.cloudflare.com/ajax/libs/react/18.2.0/umd/react.production.min.js"></script>
<script src="https://cdnjs.cloudflare.com/ajax/libs/react-dom/18.2.0/umd/react-dom.production.min.js"></script>
<script src="https://cdnjs.cloudflare.com/ajax/libs/babel-standalone/7.23.2/babel.min.js"></script>
<link href="https://cdnjs.cloudflare.com/ajax/libs/tailwindcss/2.2.19/tailwind.min.css" rel="stylesheet"/>
</head>
<body class="bg-gray-950 text-gray-100 min-h-screen">
<div id="root"></div><script type="text/babel">
const API = "http://localhost:9000";const SEVERITY_COLOR = {
critical: "bg-red-600",
high: "bg-orange-500",
medium: "bg-yellow-500",
low: "bg-blue-500",
};const TYPE_LABEL = {
crash_loop: "Crash Loop",
oom_kill: "OOM Kill",
latency_spike: "Latency Spike",
token_spike: "Token Spike",
unknown: "Unknown",
};function ConfidenceBadge({ value }) {
const pct = Math.round((value || 0) * 100);
const color = pct >= 80 ? "text-green-400" : pct >= 50 ? "text-yellow-400" : "text-red-400";
return (
<span className={`text-xs font-mono font-bold ${color}`}>
{pct}% confidence
</span>
);
}function IncidentCard({ card, onAction }) {
const d = card.diagnosis || {};
const sev = d.severity || "low";
const isPending = card.status === "pending"; return (
<div className="bg-gray-900 border border-gray-700 rounded-xl p-5 mb-4 shadow-lg">
{/* Header row */}
<div className="flex items-center justify-between mb-3">
<div className="flex items-center gap-2">
<span className={`text-xs font-bold px-2 py-0.5 rounded-full text-white ${SEVERITY_COLOR[sev] || "bg-gray-600"}`}>
{sev.toUpperCase()}
</span>
<span className="font-semibold text-white">
{TYPE_LABEL[d.incident_type] || d.incident_type}
</span>
<span className="text-gray-400 text-sm">— {card.service}</span>
</div>
<div className="flex items-center gap-3">
<ConfidenceBadge value={d.confidence} />
<span className="text-xs text-gray-500">
{new Date(card.timestamp * 1000).toLocaleTimeString()}
</span>
</div>
</div> {/* Root cause */}
<p className="text-sm text-gray-300 mb-3">
<span className="text-gray-500 mr-1">Root cause:</span>
{d.root_cause}
</p> {/* Proposed fix */}
<div className="bg-gray-800 rounded-lg px-4 py-2 mb-4 font-mono text-sm text-green-300 flex items-center gap-2">
<span className="text-gray-500 select-none">lt;/span>
{d.proposed_fix}
</div> {/* Low confidence warning */}
{(d.confidence || 0) < 0.6 && isPending && (
<div className="bg-yellow-900 border border-yellow-600 rounded-lg px-3 py-2 mb-3 text-yellow-300 text-xs">
⚠ Low confidence — review signals carefully before approving.
</div>
)} {/* Action buttons or status badge */}
{isPending ? (
<div className="flex gap-3 mt-2">
<button
onClick={() => onAction(card.id, "approve")}
className="bg-green-600 hover:bg-green-500 text-white text-sm font-semibold px-5 py-2 rounded-lg transition"
>
✓ Approve
</button>
<button
onClick={() => onAction(card.id, "reject")}
className="bg-red-700 hover:bg-red-600 text-white text-sm font-semibold px-5 py-2 rounded-lg transition"
>
✗ Reject
</button>
</div>
) : (
<div className={`inline-block text-xs font-bold px-3 py-1 rounded-full mt-1 ${
card.status === "approved" ? "bg-green-800 text-green-300" : "bg-red-900 text-red-300"
}`}>
{card.status === "approved" ? "✓ Approved" : "✗ Rejected"}
{card.applied_by && ` by ${card.applied_by}`}
</div>
)} {/* kubectl output on approved cards */}
{card.kubectl_output && (
<div className="mt-3 bg-gray-800 rounded-lg px-4 py-2 font-mono text-xs text-gray-400">
{card.kubectl_output}
</div>
)}
</div>
);
}function App() {
const [incidents, setIncidents] = React.useState([]);
const [loading, setLoading] = React.useState(true);
const [error, setError] = React.useState(null);
const [acting, setActing] = React.useState(null); const fetchIncidents = () => {
fetch(`${API}/api/incidents`)
.then(r => r.json())
.then(data => { setIncidents(data); setLoading(false); setError(null); })
.catch(e => { setError(e.message); setLoading(false); });
}; React.useEffect(() => {
fetchIncidents();
const id = setInterval(fetchIncidents, 10000); // poll every 10s
return () => clearInterval(id);
}, []); const handleAction = async (id, action) => {
setActing(id);
try {
const r = await fetch(`${API}/api/incidents/${id}/${action}`, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ actor: "operator" }),
});
if (!r.ok) {
const err = await r.json();
alert(`Action failed: ${err.detail}`);
}
fetchIncidents();
} catch (e) {
alert(`Network error: ${e.message}`);
} finally {
setActing(null);
}
}; const pending = incidents.filter(c => c.status === "pending");
const resolved = incidents.filter(c => c.status !== "pending"); return (
<div className="max-w-3xl mx-auto px-4 py-8">
{/* Header */}
<div className="mb-8">
<h1 className="text-2xl font-bold text-white">SRE Approval Dashboard</h1>
<p className="text-gray-400 text-sm mt-1">
Human-in-the-loop remediation — review every fix before it runs
</p>
</div> {loading && <p className="text-gray-500">Loading incidents…</p>}
{error && <p className="text-red-400">Could not reach backend: {error}</p>} {/* Pending */}
{pending.length > 0 && (
<section className="mb-8">
<h2 className="text-sm font-semibold text-gray-400 uppercase tracking-widest mb-3">
Awaiting Review ({pending.length})
</h2>
{pending.map(c => (
<IncidentCard key={c.id} card={c} onAction={handleAction} />
))}
</section>
)} {pending.length === 0 && !loading && (
<div className="bg-gray-900 border border-gray-700 rounded-xl p-8 text-center text-gray-500 mb-8">
No pending incidents. The cluster looks healthy.
</div>
)} {/* Resolved */}
{resolved.length > 0 && (
<section>
<h2 className="text-sm font-semibold text-gray-400 uppercase tracking-widest mb-3">
Resolved ({resolved.length})
</h2>
{resolved.map(c => (
<IncidentCard key={c.id} card={c} onAction={handleAction} />
))}
</section>
)}
</div>
);
}ReactDOM.createRoot(document.getElementById("root")).render(<App />);
</script>
</body>
</html>
EOF
cat > Dockerfile <<'EOF'
FROM nginx:alpine
COPY index.html /usr/share/nginx/html/index.html
EXPOSE 80
EOF
cd ~/k8s-aiops-blog/dashboard/backend
docker build -t approval-backend:v1 .
kind load docker-image approval-backend:v1 --name aiops
cd ~/k8s-aiops-blog/dashboard/frontend
docker build -t approval-frontend:v1 .
kind load docker-image approval-frontend:v1 --name aiops
cat > ~/k8s-aiops-blog/k8s/dashboard.yaml <<'EOF'
# ── Backend ────────────────────────────────────────────────────────────────
apiVersion: apps/v1
kind: Deployment
metadata:
name: approval-backend
namespace: apps
spec:
replicas: 1
selector:
matchLabels:
app: approval-backend
template:
metadata:
labels:
app: approval-backend
spec:
containers:
- name: backend
image: approval-backend:v1
imagePullPolicy: IfNotPresent
ports:
- containerPort: 9000
readinessProbe:
httpGet:
path: /healthz
port: 9000
initialDelaySeconds: 5
periodSeconds: 10
resources:
requests:
memory: "128Mi"
cpu: "100m"
limits:
memory: "256Mi"
cpu: "500m"
---
apiVersion: v1
kind: Service
metadata:
name: approval-dashboard
namespace: apps
spec:
selector:
app: approval-backend
ports:
- port: 9000
targetPort: 9000
---
# ── Frontend ───────────────────────────────────────────────────────────────
apiVersion: apps/v1
kind: Deployment
metadata:
name: approval-frontend
namespace: apps
spec:
replicas: 1
selector:
matchLabels:
app: approval-frontend
template:
metadata:
labels:
app: approval-frontend
spec:
containers:
- name: frontend
image: approval-frontend:v1
imagePullPolicy: IfNotPresent
ports:
- containerPort: 80
resources:
requests:
memory: "64Mi"
cpu: "50m"
limits:
memory: "128Mi"
cpu: "200m"
---
apiVersion: v1
kind: Service
metadata:
name: approval-frontend-svc
namespace: apps
spec:
selector:
app: approval-frontend
ports:
- port: 3000
targetPort: 80
EOF
kubectl apply -f ~/k8s-aiops-blog/k8s/dashboard.yaml
kubectl rollout status deployment/approval-backend -n apps
kubectl rollout status deployment/approval-frontend -n apps
kubectl port-forward -n apps svc/approval-dashboard 9000:9000
kubectl port-forward -n apps svc/approval-frontend-svc 3000:3000
kubectl logs -n apps deployment/remediation-agent --tail=5 -f
kubectl get pods -n apps -w
kubectl set env deployment/summarizer -n apps KILL_ME=true
curl -s http://localhost:8000/crash
... WARNING crash loop signal: error_rate=1.00
... INFO anomaly detected: crash_loop — calling Groq for diagnosis
... INFO diagnosis: type=crash_loop severity=critical confidence=0.91
fix=kubectl set env deployment/summarizer -n apps KILL_ME-
... INFO incident posted: id=<uuid> type=crash_loop
CRITICAL Crash Loop — summarizer 91% confidence
Root cause: The KILL_ME environment variable is set to true, causing the process to exit immediately on each start.
$ kubectl set env deployment/summarizer -n apps KILL_ME-
[ ✓ Approve ] [ ✗ Reject ]
✓ Approved by operator
deployment.apps/summarizer env updated
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 4a22289d433c: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.