This article is published in English.
Practical notes: kagent: The Kubernetes-Native AI Agent Framework You Should
Operable walkthrough of Practical notes: kagent: The Kubernetes-Native AI Agent Framework You Should: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: kagent: The Kubernetes-Native AI Agent Framework You Should Actually Know About. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
What Is kagent?
When working through the What Is kagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Architecture
When working through the Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
┌─────────────────────────────────────────────────────────┐
│ Kubernetes Cluster │
│ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Controller │────▶│ Agent CRDs │ │
│ │ (Go) │ │ ToolServers │ │
│ └──────────────┘ │ ModelConfig │ │
│ │ └──────────────┘ │
│ ▼ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Engine │────▶│ MCP Server │ │
│ │ (Python ADK │ │ (Built-in │ │
│ │ or Go ADK) │ │ tools) │ │
│ └──────────────┘ └──────────────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ Dashboard │ ← Web UI (React/TypeScript) │
│ │ + CLI │ ← kagent CLI (Go) │
│ └──────────────┘ │
└─────────────────────────────────────────────────────────┘
Controller
When working through the Controller stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Controller stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Engine (App)
The Engine App stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Tool Servers
The Tool Servers stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Human-in-the-Loop (HITL)
The Human-in-the-Loop HITL stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Human-in-the-Loop HITL stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
tools:
- type: McpServer
mcpServer:
name: kagent-tool-server
kind: RemoteMCPServer
apiGroup: kagent.dev
toolNames:
- k8s_get_resources # runs immediately
- k8s_apply_manifest # pauses for approval
- k8s_delete_resource # pauses for approval
requireApproval:
- k8s_apply_manifest
- k8s_delete_resource
Prompt Templates
For the Prompt Templates stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
systemMessage: |
You are a Kubernetes management agent named {{.AgentName}}.
{{include "builtin/safety-guardrails"}}
{{include "builtin/tool-usage-best-practices"}}
{{include "my-prompts/cluster-specific-rules"}}
Your tools: {{.ToolNames}}
Agent Memory and Context Compaction
For the Agent Memory and Context stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Agents as Tools (A2A)
For the Agents as Tools A2A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Agents as Tools A2A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Supported LLM Providers
When working through the Supported LLM Providers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
apiVersion: kagent.dev/v1alpha2
kind: ModelConfig
metadata:
name: default-model-config
namespace: kagent
spec:
provider: OpenAI
model: gpt-4o
apiKeySecret: kagent-openai
apiKeySecretKey: OPENAI_API_KEY
Pros
When working through the Pros stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Cons
When working through the Cons stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Cons stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Choosing a Model: Local vs API
The Choosing a Model Local stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
What works reliably with local models
The What works reliably with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
What needs an API model
The What needs an API stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The What needs an API stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Cost and privacy reality
For the Cost and privacy reality stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Competitor Landscape
For the Competitor Landscape stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
k8sgpt vs kagent
For the k8sgpt vs kagent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the k8sgpt vs kagent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
BotKube vs kagent
When working through the BotKube vs kagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
HolmesGPT vs kagent
When working through the HolmesGPT vs kagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
kubectl-ai vs kagent
When working through the kubectl-ai vs kagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the kubectl-ai vs kagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Demo: kagent on kind (Docker Desktop for Mac)
The Demo kagent on kind stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
git clone https://github.com/simonjday/kagent-demo
cd kagent-demo
Prerequisites
The Prerequisites stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
brew install kind kubectl helm kagent
Step 1: Create the kind Cluster
The Step 1 Create the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Step 1 Create the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
kind create cluster \
--name kagent-demo \
--config kind-cluster.yaml \
--wait 120s
export OPENAI_API_KEY="sk-..."
./scripts/install.sh
kubectl cluster-info --context kind-kagent-demo
kubectl get nodes
# NAME STATUS ROLES AGE
# kagent-demo-control-plane Ready control-plane 90s
Step 2: Install kagent with the Demo Profile
For the Step 2 Install kagent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# OpenAI
export OPENAI_API_KEY="sk-..."
# --- or Ollama (no API key needed) ---
# export KAGENT_DEFAULT_MODEL_PROVIDER=ollama
# (ensure: ollama serve && ollama pull qwen3:8b)
# Installs kagent + pre-built demo agents (k8s-agent, helm-agent, observability-agent, istio-agent)
kagent install --profile demo
# Verify the install
kubectl -n kagent get pods
# NAME READY STATUS RESTARTS
# kagent-controller-xxxx 1/1 Running 0
# kagent-xxxx 1/1 Running 0
Step 2b: Choose Your Model
For the Step 2b Choose Your stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
export ANTHROPIC_API_KEY="sk-ant-..."
kubectl -n kagent create secret generic kagent-anthropic \
--from-literal=ANTHROPIC_API_KEY="${ANTHROPIC_API_KEY}"
kubectl apply -f - <<'YAML'
apiVersion: kagent.dev/v1alpha2
kind: ModelConfig
metadata:
name: claude-sonnet
namespace: kagent
spec:
provider: Anthropic
model: claude-sonnet-4-6
apiKeySecret: kagent-anthropic
apiKeySecretKey: ANTHROPIC_API_KEY
anthropic:
apiVersion: "2023-06-01"
YAML
kubectl -n kagent get agents -o name | xargs -I {} kubectl -n kagent patch {} \
--type=merge -p '{"spec":{"declarative":{"modelConfig":"claude-sonnet"}}}'
ollama pull qwen2.5:14b
kubectl apply -f - <<'YAML'
apiVersion: kagent.dev/v1alpha2
kind: ModelConfig
metadata:
name: qwen25-14b
namespace: kagent
spec:
provider: Ollama
model: qwen2.5:14b
ollama:
host: host.docker.internal:11434
YAML
kubectl -n kagent get agents -o name | xargs -I {} kubectl -n kagent patch {} \
--type=merge -p '{"spec":{"declarative":{"modelConfig":"qwen25-14b"}}}'
Step 3: Open the Dashboard
For the Step 3 Open the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
kagent dashboard
# kagent dashboard is available at http://localhost:8082
# Press Enter to stop the port-forward...
For the Step 3 Open the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Step 4: Inspect What Was Created
When working through the Step 4 Inspect What stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# See the Agent CRDs
kubectl -n kagent get agents
# NAME AGE
# helm-agent 2m
# observability-agent 2m
# istio-agent 2m
# k8s-agent 2m
# Inspect the Kubernetes agent
kubectl -n kagent get agent k8s-agent -o yaml
Step 5: Demo Prompt 1 — Cluster Inventory
When working through the Step 5 Demo Prompt stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
List all namespaces in this cluster
How many pods are running in the kagent namespace?
Step 6: Demo Prompt 2 — Helm Awareness
When working through the Step 6 Demo Prompt stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Step 6 Demo Prompt stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
What Helm releases are installed in this cluster?
Step 7: Demo Prompt 3 — Human-in-the-Loop (Add a Failing Deployment)
The Step 7 Demo Prompt stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
name: broken-app
namespace: default
spec:
replicas: 2
selector:
matchLabels:
app: broken-app
template:
metadata:
labels:
app: broken-app
spec:
containers:
- name: app
image: nginx:nonexistent-tag
resources:
limits:
memory: "64Mi"
cpu: "250m"
EOF
There's a deployment called broken-app in the default namespace that's failing.
Diagnose what's wrong and propose a fix. Ask me before applying any changes.
kubectl delete deployment broken-app
Step 8: CLI Interaction
The Step 8 CLI Interaction stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# List agents
kagent get agent
# Invoke the helm agent directly
kagent invoke -t "List all Helm releases and flag any that haven't been updated in 30+ days" --agent helm-agent
# Run the k8s agent
kagent invoke -t "Are there any pods in a non-Running state across all namespaces?" --agent k8s-agent
Step 9: Build Your Own Agent (GitOps-Ready YAML)
The Step 9 Build Your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Step 9 Build Your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
apiVersion: kagent.dev/v1alpha2
kind: Agent
metadata:
name: platform-ops-agent
namespace: kagent
spec:
type: Declarative
declarative:
runtime: go # Fast startup; no Python deps
modelConfig: default-model-config
systemMessage: |
You are a platform operations agent for this Kubernetes cluster.
You help engineers diagnose issues, review Helm releases, and check
workload health. You MUST ask for approval before modifying any resource.
Be concise. Provide your reasoning before each tool call.
tools:
- type: McpServer
mcpServer:
name: kagent-tool-server
kind: RemoteMCPServer
apiGroup: kagent.dev
toolNames:
- k8s_get_resources
- k8s_describe_resource
- k8s_get_logs
- k8s_get_events
- k8s_apply_manifest
- k8s_delete_resource
- helm_list
- helm_status
requireApproval:
- k8s_apply_manifest
- k8s_delete_resource
context:
compaction:
compactionInterval: 5 # Summarise after every 5 exchanges
Step 10: Uninstall
For the Step 10 Uninstall stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
kagent uninstall
Should You Use It?
For the Should You Use It stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Conclusion
For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for d6fc3a8252f1: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.