Home / Articles / Practical notes: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on

This article is published in English.

Practical notes: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on

Operable walkthrough of Practical notes: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on: contracts, checks, and drop-in code slots for teams shipping this pattern.

2055 words

This walkthrough rebuilds the path from raw materials to a working system for: Demystifying the Agentic Mesh: Scaling AI Agents & LLMs on Kubernetes. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The Core Bottlenecks of Production Agentic AI

When working through the The Core Bottlenecks of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Meet the Stack: The Blueprint of the Agentic Mesh

When working through the Meet the Stack The stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

1. kagent: The Agent Runtime

When working through the 1 kagent The Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the 1 kagent The Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

2. agentgateway: The Traffic Policeman

The 2 agentgateway The Traffic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

3. llm-d: The Workload Splitter

The 3 llm-d The Workload stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

4. vLLM: The GPU Muscle

The 4 vLLM The GPU stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 4 vLLM The GPU stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

1. System Architecture

For the 1 System Architecture stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Under the Hood: Processing a Request in the Mesh

For the Under the Hood Processing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

2. Install Controllers & CRDs

For the 2 Install Controllers CRDs stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the 2 Install Controllers CRDs stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# 1. Install Kubernetes Gateway API & Inference Extension CRDs
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/standard-install.yaml
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v0.1.0/manifests.yaml

# 2. Add Helm Registries
helm repo add agentgateway oci://cr.agentgateway.dev/charts
helm repo add kagent https://charts.kagent.dev
helm repo update

# 3. Install agentgateway (Control & Data Plane)
helm upgrade -i agentgateway-crds agentgateway/agentgateway-crds -n agentgateway-system --create-namespace
helm upgrade -i agentgateway agentgateway/agentgateway -n agentgateway-system

# 4. Install kagent (Agent Runtime)
helm upgrade -i kagent kagent/kagent -n kagent-system --create-namespace

3. Deploy GPU Inference Infrastructure (vLLM & llm-d)

When working through the 3 Deploy GPU Inference stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Prefill & Decode Deployment (vllm-infrastructure.yaml)

When working through the Prefill Decode Deployment vllm-infrastructure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-prefill
  namespace: llm-serving
  labels:
    app: vllm-prefill
spec:
  replicas: 1
  selector:
    matchLabels:
      app: vllm-prefill
  template:
    metadata:
      labels:
        app: vllm-prefill
    spec:
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args:
        - "--model"
        - "meta-llama/Meta-Llama-3-8B-Instruct"
        - "--experimental-prefill-only" # Optimization: Dedicated Prefill role
        - "--gpu-memory-utilization"
        - "0.90"
        - "--port"
        - "8000"
        ports:
        - containerPort: 8000
          name: http
        resources:
          limits:
            nvidia.com/gpu: "1" # Schedule on premium compute node
            cpu: "4"
            memory: 16Gi
        volumeMounts:
        - mountPath: /root/.cache/huggingface
          name: model-cache
        - mountPath: /dev/shm
          name: dshm
      volumes:
      - name: model-cache
        persistentVolumeClaim:
          claimName: vllm-model-cache-pvc
      - name: dshm
        emptyDir:
          medium: Memory
          sizeLimit: 4Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-decode
  namespace: llm-serving
  labels:
    app: vllm-decode
spec:
  replicas: 3 # Scale out dynamically based on load
  selector:
    matchLabels:
      app: vllm-decode
  template:
    metadata:
      labels:
        app: vllm-decode
    spec:
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args:
        - "--model"
        - "meta-llama/Meta-Llama-3-8B-Instruct"
        - "--experimental-decode-only" # Optimization: Dedicated token generator role
        - "--gpu-memory-utilization"
        - "0.90"
        - "--port"
        - "8000"
        ports:
        - containerPort: 8000
          name: http
        resources:
          limits:
            nvidia.com/gpu: "1" # Schedule on low-cost L4/A10G nodes
            cpu: "4"
            memory: 16Gi
        volumeMounts:
        - mountPath: /root/.cache/huggingface
          name: model-cache
        - mountPath: /dev/shm
          name: dshm
      volumes:
      - name: model-cache
        persistentVolumeClaim:
          claimName: vllm-model-cache-pvc
      - name: dshm
        emptyDir:
          medium: Memory
          sizeLimit: 4Gi

llm-d Configuration (llmd-routing.yaml)

When working through the llm-d Configuration llmd-routing yaml stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the llm-d Configuration llmd-routing yaml stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

apiVersion: inference.networking.x-k8s.io/v1alpha1
kind: InferencePool
metadata:
  name: llama3-decode-pool
  namespace: llm-serving
spec:
  selector:
    matchLabels:
      app: vllm-decode
  targetPort: 8000
---
apiVersion: inference.networking.x-k8s.io/v1alpha1
kind: InferenceObjective
metadata:
  name: llama3-serving-objective
  namespace: llm-serving
spec:
  modelName: "meta-llama/Meta-Llama-3-8B-Instruct"
  prefillService:
    name: vllm-prefill
    port: 8000
  decodePool:
    name: llama3-decode-pool

4. Configure agentgateway

The 4 Configure agentgateway stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Gateway Definition (agent-gateway.yaml)

The Gateway Definition agent-gateway yaml stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: agent-gateway
  namespace: agentgateway-system
spec:
  gatewayClassName: agentgateway
  listeners:
  - name: http
    protocol: HTTP
    port: 8080
    allowedRoutes:
      namespaces:
        from: All
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: llm-inference-route
  namespace: llm-serving
spec:
  parentRefs:
  - name: agent-gateway
    namespace: agentgateway-system
  rules:
  - matches:
    - path:
        type: PathPrefix
        value: /v1/chat/completions
    backendRefs:
    - name: llama3-serving-objective
      kind: InferenceObjective
      group: inference.networking.x-k8s.io
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: mcp-tools-route
  namespace: agent-tools
spec:
  parentRefs:
  - name: agent-gateway
    namespace: agentgateway-system
  rules:
  - matches:
    - path:
        type: PathPrefix
        value: /mcp/tools
    backendRefs:
    - name: mcp-tool-server-service
      port: 50051

5. Deploy kagent AI Agent Pod

The 5 Deploy kagent AI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: kagent-orchestrator
  namespace: kagent-system
spec:
  replicas: 2
  selector:
    matchLabels:
      app: kagent-orchestrator
  template:
    metadata:
      labels:
        app: kagent-orchestrator
    spec:
      containers:
      - name: agent-runtime
        image: kagent/runtime:latest
        env:
        # Route all model and tool API calls through agentgateway
        - name: OPENAI_API_BASE
          value: "http://agent-gateway.agentgateway-system.svc.cluster.local:8080/v1"
        - name: MCP_SERVER_URL
          value: "http://agent-gateway.agentgateway-system.svc.cluster.local:8080/mcp/tools"
        resources:
          limits:
            cpu: "2"
            memory: 4Gi
          requests:
            cpu: "1"
            memory: 2Gi

6. Optimization & Tuning Checklist

Operational checklist