This article is published in English.
Practical notes: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on
Operable walkthrough of Practical notes: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on Kubernetes. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Core Bottlenecks of Production Agentic AI
When working through the The Core Bottlenecks of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Meet the Stack: The Blueprint of the Agentic Mesh
When working through the Meet the Stack The stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
1. kagent: The Agent Runtime
When working through the 1 kagent The Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the 1 kagent The Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
2. agentgateway: The Traffic Policeman
The 2 agentgateway The Traffic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
3. llm-d: The Workload Splitter
The 3 llm-d The Workload stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
4. vLLM: The GPU Muscle
The 4 vLLM The GPU stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 4 vLLM The GPU stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
1. System Architecture
For the 1 System Architecture stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Under the Hood: Processing a Request in the Mesh
For the Under the Hood Processing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
2. Install Controllers & CRDs
For the 2 Install Controllers CRDs stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the 2 Install Controllers CRDs stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# 1. Install Kubernetes Gateway API & Inference Extension CRDs
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/standard-install.yaml
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v0.1.0/manifests.yaml
# 2. Add Helm Registries
helm repo add agentgateway oci://cr.agentgateway.dev/charts
helm repo add kagent https://charts.kagent.dev
helm repo update
# 3. Install agentgateway (Control & Data Plane)
helm upgrade -i agentgateway-crds agentgateway/agentgateway-crds -n agentgateway-system --create-namespace
helm upgrade -i agentgateway agentgateway/agentgateway -n agentgateway-system
# 4. Install kagent (Agent Runtime)
helm upgrade -i kagent kagent/kagent -n kagent-system --create-namespace
3. Deploy GPU Inference Infrastructure (vLLM & llm-d)
When working through the 3 Deploy GPU Inference stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Prefill & Decode Deployment (vllm-infrastructure.yaml)
When working through the Prefill Decode Deployment vllm-infrastructure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-prefill
namespace: llm-serving
labels:
app: vllm-prefill
spec:
replicas: 1
selector:
matchLabels:
app: vllm-prefill
template:
metadata:
labels:
app: vllm-prefill
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- "--model"
- "meta-llama/Meta-Llama-3-8B-Instruct"
- "--experimental-prefill-only" # Optimization: Dedicated Prefill role
- "--gpu-memory-utilization"
- "0.90"
- "--port"
- "8000"
ports:
- containerPort: 8000
name: http
resources:
limits:
nvidia.com/gpu: "1" # Schedule on premium compute node
cpu: "4"
memory: 16Gi
volumeMounts:
- mountPath: /root/.cache/huggingface
name: model-cache
- mountPath: /dev/shm
name: dshm
volumes:
- name: model-cache
persistentVolumeClaim:
claimName: vllm-model-cache-pvc
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 4Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-decode
namespace: llm-serving
labels:
app: vllm-decode
spec:
replicas: 3 # Scale out dynamically based on load
selector:
matchLabels:
app: vllm-decode
template:
metadata:
labels:
app: vllm-decode
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- "--model"
- "meta-llama/Meta-Llama-3-8B-Instruct"
- "--experimental-decode-only" # Optimization: Dedicated token generator role
- "--gpu-memory-utilization"
- "0.90"
- "--port"
- "8000"
ports:
- containerPort: 8000
name: http
resources:
limits:
nvidia.com/gpu: "1" # Schedule on low-cost L4/A10G nodes
cpu: "4"
memory: 16Gi
volumeMounts:
- mountPath: /root/.cache/huggingface
name: model-cache
- mountPath: /dev/shm
name: dshm
volumes:
- name: model-cache
persistentVolumeClaim:
claimName: vllm-model-cache-pvc
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 4Gi
llm-d Configuration (llmd-routing.yaml)
When working through the llm-d Configuration llmd-routing yaml stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the llm-d Configuration llmd-routing yaml stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
apiVersion: inference.networking.x-k8s.io/v1alpha1
kind: InferencePool
metadata:
name: llama3-decode-pool
namespace: llm-serving
spec:
selector:
matchLabels:
app: vllm-decode
targetPort: 8000
---
apiVersion: inference.networking.x-k8s.io/v1alpha1
kind: InferenceObjective
metadata:
name: llama3-serving-objective
namespace: llm-serving
spec:
modelName: "meta-llama/Meta-Llama-3-8B-Instruct"
prefillService:
name: vllm-prefill
port: 8000
decodePool:
name: llama3-decode-pool
4. Configure agentgateway
The 4 Configure agentgateway stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Gateway Definition (agent-gateway.yaml)
The Gateway Definition agent-gateway yaml stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: agent-gateway
namespace: agentgateway-system
spec:
gatewayClassName: agentgateway
listeners:
- name: http
protocol: HTTP
port: 8080
allowedRoutes:
namespaces:
from: All
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: llm-inference-route
namespace: llm-serving
spec:
parentRefs:
- name: agent-gateway
namespace: agentgateway-system
rules:
- matches:
- path:
type: PathPrefix
value: /v1/chat/completions
backendRefs:
- name: llama3-serving-objective
kind: InferenceObjective
group: inference.networking.x-k8s.io
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: mcp-tools-route
namespace: agent-tools
spec:
parentRefs:
- name: agent-gateway
namespace: agentgateway-system
rules:
- matches:
- path:
type: PathPrefix
value: /mcp/tools
backendRefs:
- name: mcp-tool-server-service
port: 50051
5. Deploy kagent AI Agent Pod
The 5 Deploy kagent AI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
apiVersion: apps/v1
kind: Deployment
metadata:
name: kagent-orchestrator
namespace: kagent-system
spec:
replicas: 2
selector:
matchLabels:
app: kagent-orchestrator
template:
metadata:
labels:
app: kagent-orchestrator
spec:
containers:
- name: agent-runtime
image: kagent/runtime:latest
env:
# Route all model and tool API calls through agentgateway
- name: OPENAI_API_BASE
value: "http://agent-gateway.agentgateway-system.svc.cluster.local:8080/v1"
- name: MCP_SERVER_URL
value: "http://agent-gateway.agentgateway-system.svc.cluster.local:8080/mcp/tools"
resources:
limits:
cpu: "2"
memory: 4Gi
requests:
cpu: "1"
memory: 2Gi