Home / Articles / Practical notes: Building a Personal MCP Server With Your Team’s Runbooks and

This article is published in English.

Practical notes: Building a Personal MCP Server With Your Team’s Runbooks and

Operable walkthrough of Practical notes: Building a Personal MCP Server With Your Team’s Runbooks and: contracts, checks, and drop-in code slots for teams shipping this pattern.

3431 words

What We’re Building Specifically

You: we're getting high DB connection errors on payments-service,
     what's the runbook?

Claude: [searches your runbook MCP server]

Found your runbook: "Database Connection Pool Exhaustion"
Last updated March 14, 2026 by Priya.

Steps:
1. Check current pool size: kubectl exec -it <pod> -- env | grep DB_POOL
2. Temporarily increase: kubectl set env deploy/payments-service DB_POOL_SIZE=25
3. Monitor connections: watch -n2 'kubectl exec...'
4. If connections still growing, check for connection leaks in logs...

Project Structure

Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

runbook-mcp/
├── server.py           # The MCP server — main file
├── indexer.py          # Indexes docs into ChromaDB
├── sources/
│   ├── local.py        # Index local markdown files
│   ├── github.py       # Pull from GitHub repo
│   └── confluence.py   # Pull from Confluence (optional)
├── config.py           # Your settings
├── sync.py             # Keep index fresh
└── requirements.txt
pip install mcp chromadb sentence-transformers \
  httpx python-frontmatter watchdog \
  --break-system-packages

Step 1 — Config

the author the Step 1 Config stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. the author the Step 1 Config stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

import os
from dataclasses import dataclass, field

@dataclass
class Config:

    # ── Local docs ──────────────────────────────────────
    # Point these at wherever your docs actually are
    local_docs_paths: list = field(default_factory=lambda: [
        os.path.expanduser("~/docs/runbooks"),
        os.path.expanduser("~/docs/architecture"),
        os.path.expanduser("~/docs/postmortems"),
    ])

    # ── GitHub (optional) ───────────────────────────────
    github_token: str = os.getenv("GITHUB_TOKEN", "")
    github_repos: list = field(default_factory=lambda: [
        # format: "org/repo:path/to/docs"
        # "your-org/infrastructure:docs/runbooks",
        # "your-org/platform:docs/architecture",
    ])

    # ── Confluence (optional) ───────────────────────────
    confluence_url: str = os.getenv("CONFLUENCE_URL", "")
    confluence_token: str = os.getenv("CONFLUENCE_TOKEN", "")
    confluence_spaces: list = field(default_factory=lambda: [
        # "PLAT",   # Platform team space
        # "SRE",    # SRE space
    ])

    # ── Vector DB ───────────────────────────────────────
    vector_db_path: str = os.path.expanduser("~/.runbook-mcp/index")

    # ── Team context ────────────────────────────────────
    team_name: str = "Platform Team"
    company: str = "YourCompany"

config = Config()

Step 2 - The Document Sources

When working through the Step 2 The Document stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Local Files

When working through the Local Files stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

import os
import glob
import hashlib
import frontmatter
from datetime import datetime


def classify_doc(filepath: str, content: str) -> str:
    """Guess doc type from filename and content."""
    fp = filepath.lower()
    content_lower = content.lower()

    if any(kw in fp for kw in ["runbook", "procedure", "playbook"]):
        return "runbook"
    elif any(kw in fp for kw in ["postmortem", "incident", "retrospective"]):
        return "postmortem"
    elif any(kw in fp for kw in ["architecture", "design", "adr"]):
        return "architecture"
    elif any(kw in fp for kw in ["onboard", "setup", "getting-started"]):
        return "onboarding"
    elif any(kw in fp for kw in ["sop", "standard", "policy"]):
        return "sop"
    # Check content too
    elif "## steps" in content_lower or "## procedure" in content_lower:
        return "runbook"
    elif "root cause" in content_lower and "action items" in content_lower:
        return "postmortem"
    else:
        return "general"


def load_local_docs(paths: list[str]) -> list[dict]:
    """Load all markdown files from local directories."""
    docs = []

    for base_path in paths:
        if not os.path.exists(base_path):
            print(f"  ⚠️  Path not found: {base_path}")
            continue

        md_files = glob.glob(f"{base_path}/**/*.md", recursive=True)
        md_files += glob.glob(f"{base_path}/**/*.mdx", recursive=True)

        for filepath in md_files:
            try:
                with open(filepath, "r", errors="replace") as f:
                    raw = f.read()

                # Parse frontmatter if present
                try:
                    post = frontmatter.loads(raw)
                    content = post.content
                    metadata = dict(post.metadata)
                except:
                    content = raw
                    metadata = {}

                if not content.strip():
                    continue

                rel_path = os.path.relpath(filepath, base_path)
                doc_type = metadata.get("type") or classify_doc(rel_path, content)

                # Get title from frontmatter or first heading
                title = metadata.get("title") or _extract_title(content) or rel_path

                file_stat = os.stat(filepath)
                last_modified = datetime.fromtimestamp(
                    file_stat.st_mtime
                ).strftime("%Y-%m-%d")

                docs.append({
                    "content": content,
                    "source": rel_path,
                    "filepath": filepath,
                    "type": doc_type,
                    "title": title,
                    "last_modified": last_modified,
                    "tags": metadata.get("tags", []),
                    "hash": hashlib.md5(content.encode()).hexdigest()
                })

            except Exception as e:
                print(f"  ⚠️  Skipped {filepath}: {e}")

    return docs


def _extract_title(content: str) -> str | None:
    """Extract the first H1 heading from markdown."""
    for line in content.split("\n"):
        line = line.strip()
        if line.startswith("# "):
            return line[2:].strip()
    return None

GitHub Source

When working through the GitHub Source stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the GitHub Source stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

import httpx
import base64
from config import config


async def load_github_docs() -> list[dict]:
    """Pull markdown docs from GitHub repos."""
    if not config.github_token or not config.github_repos:
        return []

    headers = {
        "Authorization": f"token {config.github_token}",
        "Accept": "application/vnd.github.v3+json"
    }

    docs = []

    async with httpx.AsyncClient(headers=headers) as client:
        for repo_spec in config.github_repos:
            # Parse "org/repo:path"
            if ":" in repo_spec:
                repo, path = repo_spec.split(":", 1)
            else:
                repo, path = repo_spec, "docs"

            print(f"  📥 GitHub: {repo}/{path}")

            try:
                # Get all files in the path
                resp = await client.get(
                    f"https://api.github.com/repos/{repo}/contents/{path}",
                    timeout=15
                )

                if resp.status_code != 200:
                    print(f"     ⚠️  Failed: {resp.status_code}")
                    continue

                files = resp.json()
                if isinstance(files, dict):
                    files = [files]

                for file_info in files:
                    if not file_info.get("name", "").endswith(".md"):
                        continue

                    # Get file content
                    file_resp = await client.get(
                        file_info["url"], timeout=15
                    )
                    if file_resp.status_code != 200:
                        continue

                    file_data = file_resp.json()
                    content = base64.b64decode(
                        file_data["content"]
                    ).decode("utf-8", errors="replace")

                    docs.append({
                        "content": content,
                        "source": f"{repo}/{file_info['path']}",
                        "filepath": file_info["html_url"],
                        "type": "general",
                        "title": file_info["name"].replace(".md", ""),
                        "last_modified": "",
                        "tags": [],
                        "hash": file_data.get("sha", "")[:8]
                    })

            except Exception as e:
                print(f"  ⚠️  GitHub error for {repo}: {e}")

    print(f"  ✅ Loaded {len(docs)} docs from GitHub")
    return docs

Step 3 — The Indexer

The Step 3 The Indexer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

import os
import json
from pathlib import Path
import chromadb
from chromadb.utils import embedding_functions
from langchain.text_splitter import MarkdownTextSplitter
from sources.local import load_local_docs
from config import config


class RunbookIndexer:

    def __init__(self):
        os.makedirs(config.vector_db_path, exist_ok=True)

        self.db = chromadb.PersistentClient(path=config.vector_db_path)
        self.embedder = embedding_functions.SentenceTransformerEmbeddingFunction(
            model_name="all-MiniLM-L6-v2"
        )
        self.collection = self.db.get_or_create_collection(
            name="runbooks",
            embedding_function=self.embedder,
            metadata={"hnsw:space": "cosine"}
        )

        # Track ingested docs by hash
        self.index_file = Path(config.vector_db_path) / "doc_index.json"
        self.doc_index = self._load_doc_index()

        self.splitter = MarkdownTextSplitter(
            chunk_size=600,
            chunk_overlap=80
        )

    def _load_doc_index(self) -> dict:
        if self.index_file.exists():
            return json.loads(self.index_file.read_text())
        return {}

    def _save_doc_index(self):
        self.index_file.write_text(json.dumps(self.doc_index, indent=2))

    def build(self):
        """Build the full index from all sources."""
        print("📚 Loading documents from all sources...\n")

        all_docs = []

        # Local docs
        print("📁 Local files:")
        local_docs = load_local_docs(config.local_docs_paths)
        all_docs.extend(local_docs)
        print(f"   Loaded {len(local_docs)} local documents\n")

        # Index everything
        print(f"💾 Indexing {len(all_docs)} documents...")
        new_count = 0
        skip_count = 0

        for doc in all_docs:
            source = doc["source"]
            doc_hash = doc["hash"]

            # Skip unchanged docs
            if self.doc_index.get(source) == doc_hash:
                skip_count += 1
                continue

            # Remove old version
            try:
                self.collection.delete(where={"source": source})
            except:
                pass

            # Split into chunks
            chunks = self.splitter.split_text(doc["content"])
            if not chunks:
                continue

            self.collection.add(
                documents=chunks,
                metadatas=[{
                    "source": source,
                    "title": doc["title"],
                    "type": doc["type"],
                    "last_modified": doc["last_modified"],
                    "tags": ", ".join(doc.get("tags", [])),
                    "chunk_index": i,
                    "total_chunks": len(chunks)
                } for i, _ in enumerate(chunks)],
                ids=[f"{source}::chunk_{i}" for i in range(len(chunks))]
            )

            self.doc_index[source] = doc_hash
            new_count += 1
            print(f"  ✅ Indexed: {doc['title']} ({len(chunks)} chunks)")

        self._save_doc_index()

        print(f"\n✅ Done: {new_count} new, {skip_count} unchanged")
        print(f"   Total chunks in index: {self.collection.count()}")

    def search(self, query: str, n_results: int = 5, doc_type: str = None) -> list[dict]:
        """Search the index semantically."""
        where = {"type": doc_type} if doc_type else None

        try:
            results = self.collection.query(
                query_texts=[query],
                n_results=n_results,
                where=where,
                include=["documents", "metadatas", "distances"]
            )
        except Exception as e:
            return []

        matches = []
        for doc, meta, dist in zip(
            results["documents"][0],
            results["metadatas"][0],
            results["distances"][0]
        ):
            relevance = round((1 - dist) * 100, 1)
            if relevance < 30:
                continue

            matches.append({
                "content": doc,
                "source": meta["source"],
                "title": meta["title"],
                "type": meta["type"],
                "last_modified": meta.get("last_modified", ""),
                "relevance": relevance,
                "chunk_index": meta.get("chunk_index", 0),
                "total_chunks": meta.get("total_chunks", 1)
            })

        return matches

    def get_full_doc(self, source: str) -> str | None:
        """Get all chunks for a specific document and reconstruct it."""
        try:
            results = self.collection.get(
                where={"source": source},
                include=["documents", "metadatas"]
            )

            if not results["documents"]:
                return None

            # Sort chunks by index and join
            paired = list(zip(
                results["documents"],
                results["metadatas"]
            ))
            paired.sort(key=lambda x: x[1].get("chunk_index", 0))

            return "\n\n".join(doc for doc, _ in paired)

        except Exception as e:
            return None

    def list_docs(self, doc_type: str = None) -> list[dict]:
        """List all indexed documents."""
        try:
            where = {"type": doc_type} if doc_type else None
            results = self.collection.get(
                where=where,
                include=["metadatas"]
            )

            # Deduplicate by source
            seen = {}
            for meta in results["metadatas"]:
                source = meta["source"]
                if source not in seen:
                    seen[source] = {
                        "source": source,
                        "title": meta["title"],
                        "type": meta["type"],
                        "last_modified": meta.get("last_modified", "")
                    }

            return sorted(seen.values(), key=lambda x: x["title"])

        except Exception as e:
            return []

Step 4 — The MCP Server

The Step 4 The MCP stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

import asyncio
import json
from mcp.server import Server
from mcp.server.stdio import stdio_server
from mcp import types
from indexer import RunbookIndexer

app = Server("runbook-mcp-server")
indexer = RunbookIndexer()


@app.list_tools()
async def list_tools() -> list[types.Tool]:
    """Define all tools available to Claude."""
    return [

        types.Tool(
            name="search_runbooks",
            description=(
                "Search your team's runbooks and documentation semantically. "
                "Use this when someone asks about a procedure, error, "
                "or operational task. Returns the most relevant doc sections."
            ),
            inputSchema={
                "type": "object",
                "properties": {
                    "query": {
                        "type": "string",
                        "description": (
                            "What to search for. Be descriptive. "
                            "Examples: 'database connection pool exhaustion', "
                            "'nginx 502 errors', 'how to rotate API keys', "
                            "'deployment rollback procedure'"
                        )
                    },
                    "doc_type": {
                        "type": "string",
                        "enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
                        "description": "Filter by document type (optional)"
                    },
                    "n_results": {
                        "type": "integer",
                        "description": "Number of results to return (default: 5)"
                    }
                },
                "required": ["query"]
            }
        ),

        types.Tool(
            name="get_runbook",
            description=(
                "Get the full content of a specific runbook or doc by its source path. "
                "Use this after search_runbooks to get the complete document. "
                "The source path comes from search results."
            ),
            inputSchema={
                "type": "object",
                "properties": {
                    "source": {
                        "type": "string",
                        "description": "The source path from search results (e.g. 'runbooks/db-connection-pool.md')"
                    }
                },
                "required": ["source"]
            }
        ),

        types.Tool(
            name="list_runbooks",
            description=(
                "List all available runbooks and documents by type. "
                "Use when someone asks 'what runbooks do we have' or "
                "'list all incident procedures'."
            ),
            inputSchema={
                "type": "object",
                "properties": {
                    "doc_type": {
                        "type": "string",
                        "enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
                        "description": "Filter by type (optional — omit for all)"
                    }
                }
            }
        ),

        types.Tool(
            name="find_similar_incidents",
            description=(
                "Search past postmortems and incident reports for similar issues. "
                "Use when someone says 'have we seen this before' or "
                "'was there a similar incident'. Returns relevant past incidents."
            ),
            inputSchema={
                "type": "object",
                "properties": {
                    "description": {
                        "type": "string",
                        "description": "Description of the current issue or symptoms"
                    }
                },
                "required": ["description"]
            }
        ),

        types.Tool(
            name="get_onboarding_docs",
            description=(
                "Get onboarding and setup documentation. "
                "Use when someone is new or asks how to set something up."
            ),
            inputSchema={
                "type": "object",
                "properties": {
                    "topic": {
                        "type": "string",
                        "description": "What they need to set up or learn (e.g. 'kubectl access', 'AWS credentials', 'local development')"
                    }
                },
                "required": ["topic"]
            }
        ),

    ]


@app.call_tool()
async def call_tool(name: str, arguments: dict) -> list[types.TextContent]:
    """Handle tool calls from Claude."""

    # ── search_runbooks ────────────────────────────────────
    if name == "search_runbooks":
        query = arguments["query"]
        doc_type = arguments.get("doc_type")
        n = arguments.get("n_results", 5)

        results = indexer.search(query, n_results=n, doc_type=doc_type)

        if not results:
            return [types.TextContent(
                type="text",
                text=f"No documents found matching: '{query}'"
            )]

        output_parts = [f"Found {len(results)} relevant documents:\n"]

        for i, r in enumerate(results, 1):
            output_parts.append(
                f"\n--- Result {i} ---\n"
                f"Title: {r['title']}\n"
                f"Type: {r['type']}\n"
                f"Source: {r['source']}\n"
                f"Last modified: {r['last_modified']}\n"
                f"Relevance: {r['relevance']}%\n"
                f"Chunk {r['chunk_index']+1}/{r['total_chunks']}\n\n"
                f"{r['content']}"
            )

        return [types.TextContent(type="text", text="\n".join(output_parts))]

    # ── get_runbook ────────────────────────────────────────
    elif name == "get_runbook":
        source = arguments["source"]
        content = indexer.get_full_doc(source)

        if not content:
            return [types.TextContent(
                type="text",
                text=f"Document not found: {source}\n"
                     f"Try searching with search_runbooks first."
            )]

        return [types.TextContent(
            type="text",
            text=f"Full document: {source}\n\n{content}"
        )]

    # ── list_runbooks ──────────────────────────────────────
    elif name == "list_runbooks":
        doc_type = arguments.get("doc_type")
        docs = indexer.list_docs(doc_type=doc_type)

        if not docs:
            label = f" of type '{doc_type}'" if doc_type else ""
            return [types.TextContent(
                type="text",
                text=f"No documents found{label}. Run the indexer first."
            )]

        # Group by type
        by_type: dict[str, list] = {}
        for doc in docs:
            t = doc["type"]
            by_type.setdefault(t, []).append(doc)

        lines = [f"Available documents ({len(docs)} total):\n"]
        for dtype, dtype_docs in sorted(by_type.items()):
            lines.append(f"\n## {dtype.title()}s ({len(dtype_docs)})")
            for doc in dtype_docs:
                modified = f" — updated {doc['last_modified']}" if doc['last_modified'] else ""
                lines.append(f"  - {doc['title']}{modified}\n    [{doc['source']}]")

        return [types.TextContent(type="text", text="\n".join(lines))]

    # ── find_similar_incidents ─────────────────────────────
    elif name == "find_similar_incidents":
        description = arguments["description"]

        # Search specifically in postmortems
        results = indexer.search(
            query=description,
            n_results=5,
            doc_type="postmortem"
        )

        if not results:
            # Broaden to all docs if no postmortems found
            results = indexer.search(query=description, n_results=3)

        if not results:
            return [types.TextContent(
                type="text",
                text="No similar incidents found in the knowledge base."
            )]

        output = [f"Similar past incidents:\n"]
        for r in results:
            output.append(
                f"\n📄 {r['title']} [{r['source']}]\n"
                f"Relevance: {r['relevance']}%\n"
                f"Last modified: {r['last_modified']}\n\n"
                f"{r['content'][:600]}..."
            )

        return [types.TextContent(type="text", text="\n".join(output))]

    # ── get_onboarding_docs ────────────────────────────────
    elif name == "get_onboarding_docs":
        topic = arguments["topic"]

        results = indexer.search(
            query=f"setup onboarding {topic}",
            n_results=4,
            doc_type="onboarding"
        )

        if not results:
            # Fall back to all doc types
            results = indexer.search(
                query=f"how to setup {topic}", n_results=3
            )

        if not results:
            return [types.TextContent(
                type="text",
                text=f"No onboarding docs found for: {topic}"
            )]

        output = [f"Onboarding docs for '{topic}':\n"]
        for r in results:
            output.append(
                f"\n📄 {r['title']}\n"
                f"Source: {r['source']}\n\n"
                f"{r['content']}"
            )

        return [types.TextContent(type="text", text="\n".join(output))]

    return [types.TextContent(type="text", text=f"Unknown tool: {name}")]


async def main():
    async with stdio_server() as (read_stream, write_stream):
        await app.run(
            read_stream,
            write_stream,
            app.create_initialization_options()
        )


if __name__ == "__main__":
    asyncio.run(main())

Step 5 — Keep It Fresh

The Step 5 Keep It stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

"""
Run this to rebuild the index when docs change.
Can be triggered manually or via a cron job.
"""
from indexer import RunbookIndexer
from datetime import datetime

def sync():
    print(f"🔄 Syncing runbook index - {datetime.now().strftime('%Y-%m-%d %H:%M')}\n")
    indexer = RunbookIndexer()
    indexer.build()
    print("\n✅ Sync complete")

if __name__ == "__main__":
    sync()
# Rebuild every night at midnight
0 0 * * * cd ~/runbook-mcp && python sync.py >> ~/.runbook-mcp/sync.log 2>&1

# Or: rebuild every time someone pushes to the docs repo
# Add to your GitHub Actions workflow:
# - name: Sync MCP index
#   run: ssh your-server "cd ~/runbook-mcp && python sync.py"

Connect It to Claude Desktop

Mac:     ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
{
  "mcpServers": {
    "runbooks": {
      "command": "python",
      "args": ["/Users/yourname/runbook-mcp/server.py"],
      "env": {
        "GITHUB_TOKEN": "your-token-here"
      }
    }
  }
}

First-Time Setup — The Whole Thing in 5 Steps

# 1. Clone and install
git clone your-repo/runbook-mcp
cd runbook-mcp
pip install -r requirements.txt

# 2. Set up your docs paths in config.py

# Point local_docs_paths at wherever your runbooks actually are
nano config.py

# 3. Build the index
python sync.py

# 4. Test it works
python -c "
from indexer import RunbookIndexer
idx = RunbookIndexer()
results = idx.search('database connection pool')
for r in results:
    print(r['title'], r['relevance'])
"

# 5. Add to Claude Desktop config and restart

Real Conversations After Setup

You: we're seeing OOM kills on the API pods, what's the procedure?

Claude: [searches runbooks for "OOM killed API pods"]

Found your runbook: "Pod OOMKilled - Response Procedure"

Updated April 3, 2026.

Immediate steps:
1. Confirm: kubectl describe pod <pod-name> | grep -A5 "Last State"
2. Check current limits: kubectl get deploy api-service -o yaml | grep -A3 resources
3. Quick fix (buys time): kubectl patch deploy api-service -p '{"spec":{"template":{"spec":{"containers":[{"name":"api","resources":{"limits":{"memory":"1Gi"}}}]}}}}'
4. Restart: kubectl rollout restart deploy/api-service
5. Monitor: watch kubectl top pods -n production
Root fix: Update Helm values.yaml - memory limit should be at least 2x the p99 usage.
You: how do I get access to the production Kubernetes cluster?

Claude: [searches onboarding docs for "kubernetes cluster access production"]

Found your onboarding guide: "Getting Started - Cluster Access"

Updated Feb 10, 2026 by Arjun.

Steps for production cluster access:
1. Request access via this form: [link]
2. Wait for approval from your team lead (usually same day)
3. Install kubectl: brew install kubectl
4. Download our kubeconfig: aws eks update-kubeconfig --name prod-cluster --region us-east-1
5. Verify: kubectl get nodes

Note: Production is read-only by default. Write access requires a separate approval.
You: have we ever seen this S3 "SlowDown" error before?
     payments are failing because S3 is rate limiting us


Claude: [searches postmortems for "S3 SlowDown rate limiting payments"]

Found 2 similar past incidents:
📄 "S3 Rate Limiting - Payments Service" (March 2025)

Root cause: too many small PutObject calls in a tight loop.

Fix: batch uploads, added exponential backoff. See PR #4821.
📄 "S3 Throttling During Flash Sale" (November 2024)

Root cause: single-prefix S3 bucket hitting 3,500 req/s limit.
Fix: switched to key-prefix randomization. Docs in architecture/s3-patterns.md

Want me to pull the full postmortem for either of these?

Add More Sources Over Time

# sources/notion.py
# Use the Notion API to pull pages from a database
# Same pattern: fetch content → ingest into ChromaDB
# sources/confluence.py
# Use the Confluence REST API
# Fetch pages by space key, convert HTML to markdown
# sources/slack.py
# Pull important threads from your #incidents or #platform channel
# Save and index them as "informal runbooks"
# Already works with the GitHub source
# Just add "your-org/incidents:postmortems" to github_repos in config

Structuring Your Runbooks for Better Search

# Database Connection Pool Exhaustion — Response Runbook
# DB issue
## Symptoms
- PagerDuty: "connection pool exhausted"
- Error in logs: "too many clients already"
- API latency spike on database-heavy endpoints
---
title: Database Connection Pool Exhaustion
type: runbook
tags: [database, postgresql, connections, production]
last_updated: 2026-03-14
owner: platform-team
---

Honest Expectations

The Real Value

Operational checklist