Практические советы: создание личного сервера MCP с использованием руководств вашей команды и
Пошаговое руководство по практическим рекомендациям: создание личного сервера MCP с использованием руководств вашей команды, а также контрактов, проверок и слотов для вставки кода для команд, использующих эту схему.
Что именно мы разрабатываем
You: we're getting high DB connection errors on payments-service,
what's the runbook?
Claude: [searches your runbook MCP server]
Found your runbook: "Database Connection Pool Exhaustion"
Last updated March 14, 2026 by Priya.
Steps:
1. Check current pool size: kubectl exec -it <pod> -- env | grep DB_POOL
2. Temporarily increase: kubectl set env deploy/payments-service DB_POOL_SIZE=25
3. Monitor connections: watch -n2 'kubectl exec...'
4. If connections still growing, check for connection leaks in logs...
Структура проекта
Аутентификация осуществляется на шлюзе, а повторная авторизация — на уровне данных. Один только токен-носитель не является границей между тенантами.
runbook-mcp/
├── server.py # The MCP server — main file
├── indexer.py # Indexes docs into ChromaDB
├── sources/
│ ├── local.py # Index local markdown files
│ ├── github.py # Pull from GitHub repo
│ └── confluence.py # Pull from Confluence (optional)
├── config.py # Your settings
├── sync.py # Keep index fresh
└── requirements.txt
pip install mcp chromadb sentence-transformers \
httpx python-frontmatter watchdog \
--break-system-packages
Шаг 1 — настройка
Автор на этапе конфигурации Шаг 1 должен определить входные данные, ответственного за выполнение шага и критерии завершения перед изменением кода. Операторы должны иметь возможность перезапустить шаг с известной точки контроля, не догадываясь о скрытом состоянии. Записывайте время выполнения и стоимость токена или запроса рядом с функциональными результатами. Отображение стоимости заранее предотвращает неожиданные счета при переходе из демо-среды в общедоступные среды. Аутентифицируйтесь у шлюза и повторно авторизуйтесь на уровне данных. Один только токен-носитель не является границей аренды. Автор на этапе конфигурации Шаг 1 должен определить входные данные, ответственного за выполнение шага и критерии завершения перед изменением кода. Операторы должны иметь возможность перезапустить шаг с известной точки контроля, не догадываясь о скрытом состоянии. Документируйте одновременно успешный сценарий работы и сценарий восстановления. Повторные попытки, человеческое вмешательство и обработка некорректных сообщений являются частью продукта, а не элементами последующей доработки.
import os
from dataclasses import dataclass, field
@dataclass
class Config:
# ── Local docs ──────────────────────────────────────
# Point these at wherever your docs actually are
local_docs_paths: list = field(default_factory=lambda: [
os.path.expanduser("~/docs/runbooks"),
os.path.expanduser("~/docs/architecture"),
os.path.expanduser("~/docs/postmortems"),
])
# ── GitHub (optional) ───────────────────────────────
github_token: str = os.getenv("GITHUB_TOKEN", "")
github_repos: list = field(default_factory=lambda: [
# format: "org/repo:path/to/docs"
# "your-org/infrastructure:docs/runbooks",
# "your-org/platform:docs/architecture",
])
# ── Confluence (optional) ───────────────────────────
confluence_url: str = os.getenv("CONFLUENCE_URL", "")
confluence_token: str = os.getenv("CONFLUENCE_TOKEN", "")
confluence_spaces: list = field(default_factory=lambda: [
# "PLAT", # Platform team space
# "SRE", # SRE space
])
# ── Vector DB ───────────────────────────────────────
vector_db_path: str = os.path.expanduser("~/.runbook-mcp/index")
# ── Team context ────────────────────────────────────
team_name: str = "Platform Team"
company: str = "YourCompany"
config = Config()
Шаг 2 – Источники документов
При работе над этапом «Шаг 2: Документы» сначала запишите условия контракта: необходимые входные данные, сигнал о успешном выполнении и действия при частичной неудаче. Такой чек-лист поможет сохранять честность при последующих изменениях кода. Предпочитайте небольшие, тестируемые единицы кода вместо обширных скриптов. При сбое на каком-либо этапе причина неудачи должна указывать на конкретную ответственность, а не на запутанную цепочку операций. Фиксируйте название инструмента, хеш аргументов, время задержки и результат каждого вызова. Без такой информации отладка занимает часы.
Локальные файлы
При работе на этапе локальных файлов сначала запишите контракт: необходимые входные данные, сигнал о успешном выполнении и действия при частичной неудаче. Такой список помогает сохранять честность при последующих изменениях кода. Рассматривайте этот этап как контракт между входными данными и проверенными выходными результатами. Дайте названия создаваемым файлам, определите критерии успешности и не допускайте молчаливого частичного завершения работы. Записывайте название инструмента, хеш аргументов, время задержки и результат каждого вызова. Без такой записи отладка занимает гораздо больше времени.
import os
import glob
import hashlib
import frontmatter
from datetime import datetime
def classify_doc(filepath: str, content: str) -> str:
"""Guess doc type from filename and content."""
fp = filepath.lower()
content_lower = content.lower()
if any(kw in fp for kw in ["runbook", "procedure", "playbook"]):
return "runbook"
elif any(kw in fp for kw in ["postmortem", "incident", "retrospective"]):
return "postmortem"
elif any(kw in fp for kw in ["architecture", "design", "adr"]):
return "architecture"
elif any(kw in fp for kw in ["onboard", "setup", "getting-started"]):
return "onboarding"
elif any(kw in fp for kw in ["sop", "standard", "policy"]):
return "sop"
# Check content too
elif "## steps" in content_lower or "## procedure" in content_lower:
return "runbook"
elif "root cause" in content_lower and "action items" in content_lower:
return "postmortem"
else:
return "general"
def load_local_docs(paths: list[str]) -> list[dict]:
"""Load all markdown files from local directories."""
docs = []
for base_path in paths:
if not os.path.exists(base_path):
print(f" ⚠️ Path not found: {base_path}")
continue
md_files = glob.glob(f"{base_path}/**/*.md", recursive=True)
md_files += glob.glob(f"{base_path}/**/*.mdx", recursive=True)
for filepath in md_files:
try:
with open(filepath, "r", errors="replace") as f:
raw = f.read()
# Parse frontmatter if present
try:
post = frontmatter.loads(raw)
content = post.content
metadata = dict(post.metadata)
except:
content = raw
metadata = {}
if not content.strip():
continue
rel_path = os.path.relpath(filepath, base_path)
doc_type = metadata.get("type") or classify_doc(rel_path, content)
# Get title from frontmatter or first heading
title = metadata.get("title") or _extract_title(content) or rel_path
file_stat = os.stat(filepath)
last_modified = datetime.fromtimestamp(
file_stat.st_mtime
).strftime("%Y-%m-%d")
docs.append({
"content": content,
"source": rel_path,
"filepath": filepath,
"type": doc_type,
"title": title,
"last_modified": last_modified,
"tags": metadata.get("tags", []),
"hash": hashlib.md5(content.encode()).hexdigest()
})
except Exception as e:
print(f" ⚠️ Skipped {filepath}: {e}")
return docs
def _extract_title(content: str) -> str | None:
"""Extract the first H1 heading from markdown."""
for line in content.split("\n"):
line = line.strip()
if line.startswith("# "):
return line[2:].strip()
return None
Исходный код на GitHub
При работе на этапе GitHub Source сначала запишите условия использования: необходимые входные данные, сигнал о успешном выполнении и действия при частичной неудаче. Такой список помогает сохранять честность при последующих изменениях кода. Рядом с функциональными результатами записывайте время выполнения и стоимость токенов или запросов. Отслеживание затрат с самого начала предотвращает неожиданные счета при переходе от демо-среды к общедоступным средам. Фиксируйте название инструмента, хеш аргументов, время задержки и результат каждого вызова. Без такой записи отладка занимает часы. При работе на этапе GitHub Source сначала запишите условия использования: необходимые входные данные, сигнал о успешном выполнении и действия при частичной неудаче. Такой список помогает сохранять честность при последующих изменениях кода. Одновременно документируйте успешный сценарий работы и сценарий восстановления. Повторные попытки, проверки человеком и обработка неработоспособных сообщений являются частью продукта, а не элементами последующей доработки.
import httpx
import base64
from config import config
async def load_github_docs() -> list[dict]:
"""Pull markdown docs from GitHub repos."""
if not config.github_token or not config.github_repos:
return []
headers = {
"Authorization": f"token {config.github_token}",
"Accept": "application/vnd.github.v3+json"
}
docs = []
async with httpx.AsyncClient(headers=headers) as client:
for repo_spec in config.github_repos:
# Parse "org/repo:path"
if ":" in repo_spec:
repo, path = repo_spec.split(":", 1)
else:
repo, path = repo_spec, "docs"
print(f" 📥 GitHub: {repo}/{path}")
try:
# Get all files in the path
resp = await client.get(
f"https://api.github.com/repos/{repo}/contents/{path}",
timeout=15
)
if resp.status_code != 200:
print(f" ⚠️ Failed: {resp.status_code}")
continue
files = resp.json()
if isinstance(files, dict):
files = [files]
for file_info in files:
if not file_info.get("name", "").endswith(".md"):
continue
# Get file content
file_resp = await client.get(
file_info["url"], timeout=15
)
if file_resp.status_code != 200:
continue
file_data = file_resp.json()
content = base64.b64decode(
file_data["content"]
).decode("utf-8", errors="replace")
docs.append({
"content": content,
"source": f"{repo}/{file_info['path']}",
"filepath": file_info["html_url"],
"type": "general",
"title": file_info["name"].replace(".md", ""),
"last_modified": "",
"tags": [],
"hash": file_data.get("sha", "")[:8]
})
except Exception as e:
print(f" ⚠️ GitHub error for {repo}: {e}")
print(f" ✅ Loaded {len(docs)} docs from GitHub")
return docs
Шаг 3 — Индексатор
На третьем этапе, на стадии индексации, лучший результат достигается при рассмотрении её как измеримой величины. Соберите один идеальный пример работы, один случай сбоя и записку о возврате к предыдущему состоянию перед расширением объёма работ. Предпочитайте небольшие, тестируемые единицы кода вместо обширных скриптов. При сбое какого-либо шага причина должна быть связана с конкретной функцией, а не с запутанной цепочкой операций. Используйте инструменты с узкими схемами и чёткими метками о побочных эффектах. У операторов должна быть возможность узнать, какие вызовы изменяют состояние, прежде чем они автоматически одобрят действие.
import os
import json
from pathlib import Path
import chromadb
from chromadb.utils import embedding_functions
from langchain.text_splitter import MarkdownTextSplitter
from sources.local import load_local_docs
from config import config
class RunbookIndexer:
def __init__(self):
os.makedirs(config.vector_db_path, exist_ok=True)
self.db = chromadb.PersistentClient(path=config.vector_db_path)
self.embedder = embedding_functions.SentenceTransformerEmbeddingFunction(
model_name="all-MiniLM-L6-v2"
)
self.collection = self.db.get_or_create_collection(
name="runbooks",
embedding_function=self.embedder,
metadata={"hnsw:space": "cosine"}
)
# Track ingested docs by hash
self.index_file = Path(config.vector_db_path) / "doc_index.json"
self.doc_index = self._load_doc_index()
self.splitter = MarkdownTextSplitter(
chunk_size=600,
chunk_overlap=80
)
def _load_doc_index(self) -> dict:
if self.index_file.exists():
return json.loads(self.index_file.read_text())
return {}
def _save_doc_index(self):
self.index_file.write_text(json.dumps(self.doc_index, indent=2))
def build(self):
"""Build the full index from all sources."""
print("📚 Loading documents from all sources...\n")
all_docs = []
# Local docs
print("📁 Local files:")
local_docs = load_local_docs(config.local_docs_paths)
all_docs.extend(local_docs)
print(f" Loaded {len(local_docs)} local documents\n")
# Index everything
print(f"💾 Indexing {len(all_docs)} documents...")
new_count = 0
skip_count = 0
for doc in all_docs:
source = doc["source"]
doc_hash = doc["hash"]
# Skip unchanged docs
if self.doc_index.get(source) == doc_hash:
skip_count += 1
continue
# Remove old version
try:
self.collection.delete(where={"source": source})
except:
pass
# Split into chunks
chunks = self.splitter.split_text(doc["content"])
if not chunks:
continue
self.collection.add(
documents=chunks,
metadatas=[{
"source": source,
"title": doc["title"],
"type": doc["type"],
"last_modified": doc["last_modified"],
"tags": ", ".join(doc.get("tags", [])),
"chunk_index": i,
"total_chunks": len(chunks)
} for i, _ in enumerate(chunks)],
ids=[f"{source}::chunk_{i}" for i in range(len(chunks))]
)
self.doc_index[source] = doc_hash
new_count += 1
print(f" ✅ Indexed: {doc['title']} ({len(chunks)} chunks)")
self._save_doc_index()
print(f"\n✅ Done: {new_count} new, {skip_count} unchanged")
print(f" Total chunks in index: {self.collection.count()}")
def search(self, query: str, n_results: int = 5, doc_type: str = None) -> list[dict]:
"""Search the index semantically."""
where = {"type": doc_type} if doc_type else None
try:
results = self.collection.query(
query_texts=[query],
n_results=n_results,
where=where,
include=["documents", "metadatas", "distances"]
)
except Exception as e:
return []
matches = []
for doc, meta, dist in zip(
results["documents"][0],
results["metadatas"][0],
results["distances"][0]
):
relevance = round((1 - dist) * 100, 1)
if relevance < 30:
continue
matches.append({
"content": doc,
"source": meta["source"],
"title": meta["title"],
"type": meta["type"],
"last_modified": meta.get("last_modified", ""),
"relevance": relevance,
"chunk_index": meta.get("chunk_index", 0),
"total_chunks": meta.get("total_chunks", 1)
})
return matches
def get_full_doc(self, source: str) -> str | None:
"""Get all chunks for a specific document and reconstruct it."""
try:
results = self.collection.get(
where={"source": source},
include=["documents", "metadatas"]
)
if not results["documents"]:
return None
# Sort chunks by index and join
paired = list(zip(
results["documents"],
results["metadatas"]
))
paired.sort(key=lambda x: x[1].get("chunk_index", 0))
return "\n\n".join(doc for doc, _ in paired)
except Exception as e:
return None
def list_docs(self, doc_type: str = None) -> list[dict]:
"""List all indexed documents."""
try:
where = {"type": doc_type} if doc_type else None
results = self.collection.get(
where=where,
include=["metadatas"]
)
# Deduplicate by source
seen = {}
for meta in results["metadatas"]:
source = meta["source"]
if source not in seen:
seen[source] = {
"source": source,
"title": meta["title"],
"type": meta["type"],
"last_modified": meta.get("last_modified", "")
}
return sorted(seen.values(), key=lambda x: x["title"])
except Exception as e:
return []
Шаг 4 — Сервер MCP
На четвертом этапе работа MCP наиболее эффективна, если рассматривать её как измеримую среду. Соберите один эталонный пример работы, один случай сбоя и записку о возврате к предыдущему состоянию перед расширением объёма работ. Рассматривайте этот этап как контракт между входными данными и проверенными выходными результатами. Дайте названия создаваемым элементам, определите критерии успешного выполнения и не соглашайтесь на молчаливое частичное завершение работ. Используйте инструменты с узкими схемами и чёткими метками о побочных эффектах. У операторов должна быть возможность узнать, какие вызовы изменяют состояние, прежде чем они автоматически одобрят их.
import asyncio
import json
from mcp.server import Server
from mcp.server.stdio import stdio_server
from mcp import types
from indexer import RunbookIndexer
app = Server("runbook-mcp-server")
indexer = RunbookIndexer()
@app.list_tools()
async def list_tools() -> list[types.Tool]:
"""Define all tools available to Claude."""
return [
types.Tool(
name="search_runbooks",
description=(
"Search your team's runbooks and documentation semantically. "
"Use this when someone asks about a procedure, error, "
"or operational task. Returns the most relevant doc sections."
),
inputSchema={
"type": "object",
"properties": {
"query": {
"type": "string",
"description": (
"What to search for. Be descriptive. "
"Examples: 'database connection pool exhaustion', "
"'nginx 502 errors', 'how to rotate API keys', "
"'deployment rollback procedure'"
)
},
"doc_type": {
"type": "string",
"enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
"description": "Filter by document type (optional)"
},
"n_results": {
"type": "integer",
"description": "Number of results to return (default: 5)"
}
},
"required": ["query"]
}
),
types.Tool(
name="get_runbook",
description=(
"Get the full content of a specific runbook or doc by its source path. "
"Use this after search_runbooks to get the complete document. "
"The source path comes from search results."
),
inputSchema={
"type": "object",
"properties": {
"source": {
"type": "string",
"description": "The source path from search results (e.g. 'runbooks/db-connection-pool.md')"
}
},
"required": ["source"]
}
),
types.Tool(
name="list_runbooks",
description=(
"List all available runbooks and documents by type. "
"Use when someone asks 'what runbooks do we have' or "
"'list all incident procedures'."
),
inputSchema={
"type": "object",
"properties": {
"doc_type": {
"type": "string",
"enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
"description": "Filter by type (optional — omit for all)"
}
}
}
),
types.Tool(
name="find_similar_incidents",
description=(
"Search past postmortems and incident reports for similar issues. "
"Use when someone says 'have we seen this before' or "
"'was there a similar incident'. Returns relevant past incidents."
),
inputSchema={
"type": "object",
"properties": {
"description": {
"type": "string",
"description": "Description of the current issue or symptoms"
}
},
"required": ["description"]
}
),
types.Tool(
name="get_onboarding_docs",
description=(
"Get onboarding and setup documentation. "
"Use when someone is new or asks how to set something up."
),
inputSchema={
"type": "object",
"properties": {
"topic": {
"type": "string",
"description": "What they need to set up or learn (e.g. 'kubectl access', 'AWS credentials', 'local development')"
}
},
"required": ["topic"]
}
),
]
@app.call_tool()
async def call_tool(name: str, arguments: dict) -> list[types.TextContent]:
"""Handle tool calls from Claude."""
# ── search_runbooks ────────────────────────────────────
if name == "search_runbooks":
query = arguments["query"]
doc_type = arguments.get("doc_type")
n = arguments.get("n_results", 5)
results = indexer.search(query, n_results=n, doc_type=doc_type)
if not results:
return [types.TextContent(
type="text",
text=f"No documents found matching: '{query}'"
)]
output_parts = [f"Found {len(results)} relevant documents:\n"]
for i, r in enumerate(results, 1):
output_parts.append(
f"\n--- Result {i} ---\n"
f"Title: {r['title']}\n"
f"Type: {r['type']}\n"
f"Source: {r['source']}\n"
f"Last modified: {r['last_modified']}\n"
f"Relevance: {r['relevance']}%\n"
f"Chunk {r['chunk_index']+1}/{r['total_chunks']}\n\n"
f"{r['content']}"
)
return [types.TextContent(type="text", text="\n".join(output_parts))]
# ── get_runbook ────────────────────────────────────────
elif name == "get_runbook":
source = arguments["source"]
content = indexer.get_full_doc(source)
if not content:
return [types.TextContent(
type="text",
text=f"Document not found: {source}\n"
f"Try searching with search_runbooks first."
)]
return [types.TextContent(
type="text",
text=f"Full document: {source}\n\n{content}"
)]
# ── list_runbooks ──────────────────────────────────────
elif name == "list_runbooks":
doc_type = arguments.get("doc_type")
docs = indexer.list_docs(doc_type=doc_type)
if not docs:
label = f" of type '{doc_type}'" if doc_type else ""
return [types.TextContent(
type="text",
text=f"No documents found{label}. Run the indexer first."
)]
# Group by type
by_type: dict[str, list] = {}
for doc in docs:
t = doc["type"]
by_type.setdefault(t, []).append(doc)
lines = [f"Available documents ({len(docs)} total):\n"]
for dtype, dtype_docs in sorted(by_type.items()):
lines.append(f"\n## {dtype.title()}s ({len(dtype_docs)})")
for doc in dtype_docs:
modified = f" — updated {doc['last_modified']}" if doc['last_modified'] else ""
lines.append(f" - {doc['title']}{modified}\n [{doc['source']}]")
return [types.TextContent(type="text", text="\n".join(lines))]
# ── find_similar_incidents ─────────────────────────────
elif name == "find_similar_incidents":
description = arguments["description"]
# Search specifically in postmortems
results = indexer.search(
query=description,
n_results=5,
doc_type="postmortem"
)
if not results:
# Broaden to all docs if no postmortems found
results = indexer.search(query=description, n_results=3)
if not results:
return [types.TextContent(
type="text",
text="No similar incidents found in the knowledge base."
)]
output = [f"Similar past incidents:\n"]
for r in results:
output.append(
f"\n📄 {r['title']} [{r['source']}]\n"
f"Relevance: {r['relevance']}%\n"
f"Last modified: {r['last_modified']}\n\n"
f"{r['content'][:600]}..."
)
return [types.TextContent(type="text", text="\n".join(output))]
# ── get_onboarding_docs ────────────────────────────────
elif name == "get_onboarding_docs":
topic = arguments["topic"]
results = indexer.search(
query=f"setup onboarding {topic}",
n_results=4,
doc_type="onboarding"
)
if not results:
# Fall back to all doc types
results = indexer.search(
query=f"how to setup {topic}", n_results=3
)
if not results:
return [types.TextContent(
type="text",
text=f"No onboarding docs found for: {topic}"
)]
output = [f"Onboarding docs for '{topic}':\n"]
for r in results:
output.append(
f"\n📄 {r['title']}\n"
f"Source: {r['source']}\n\n"
f"{r['content']}"
)
return [types.TextContent(type="text", text="\n".join(output))]
return [types.TextContent(type="text", text=f"Unknown tool: {name}")]
async def main():
async with stdio_server() as (read_stream, write_stream):
await app.run(
read_stream,
write_stream,
app.create_initialization_options()
)
if __name__ == "__main__":
asyncio.run(main())
Этап 5 — Поддерживайте актуальность
На пятом этапе поддержания актуальности работа наиболее эффективна, если рассматривать его как измеримую среду. Соберите один эталонный пример работы, один случай сбоя и записку о возврате к предыдущему состоянию перед расширением объёма работ. Записывайте временные показатели, а также стоимость токенов или запросов вместе с функциональными результатами. Отображение стоимости на раннем этапе предотвращает неожиданные счета при переходе от демо-среды к общедоступным средам.
"""
Run this to rebuild the index when docs change.
Can be triggered manually or via a cron job.
"""
from indexer import RunbookIndexer
from datetime import datetime
def sync():
print(f"🔄 Syncing runbook index - {datetime.now().strftime('%Y-%m-%d %H:%M')}\n")
indexer = RunbookIndexer()
indexer.build()
print("\n✅ Sync complete")
if __name__ == "__main__":
sync()
# Rebuild every night at midnight
0 0 * * * cd ~/runbook-mcp && python sync.py >> ~/.runbook-mcp/sync.log 2>&1
# Or: rebuild every time someone pushes to the docs repo
# Add to your GitHub Actions workflow:
# - name: Sync MCP index
# run: ssh your-server "cd ~/runbook-mcp && python sync.py"
Подключите его к Claude Desktop
Mac: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"runbooks": {
"command": "python",
"args": ["/Users/yourname/runbook-mcp/server.py"],
"env": {
"GITHUB_TOKEN": "your-token-here"
}
}
}
}
Настройка впервые — всё за 5 шагов
# 1. Clone and install
git clone your-repo/runbook-mcp
cd runbook-mcp
pip install -r requirements.txt
# 2. Set up your docs paths in config.py
# Point local_docs_paths at wherever your runbooks actually are
nano config.py
# 3. Build the index
python sync.py
# 4. Test it works
python -c "
from indexer import RunbookIndexer
idx = RunbookIndexer()
results = idx.search('database connection pool')
for r in results:
print(r['title'], r['relevance'])
"
# 5. Add to Claude Desktop config and restart
Реальные разговоры после настройки
You: we're seeing OOM kills on the API pods, what's the procedure?
Claude: [searches runbooks for "OOM killed API pods"]
Found your runbook: "Pod OOMKilled - Response Procedure"
Updated April 3, 2026.
Immediate steps:
1. Confirm: kubectl describe pod <pod-name> | grep -A5 "Last State"
2. Check current limits: kubectl get deploy api-service -o yaml | grep -A3 resources
3. Quick fix (buys time): kubectl patch deploy api-service -p '{"spec":{"template":{"spec":{"containers":[{"name":"api","resources":{"limits":{"memory":"1Gi"}}}]}}}}'
4. Restart: kubectl rollout restart deploy/api-service
5. Monitor: watch kubectl top pods -n production
Root fix: Update Helm values.yaml - memory limit should be at least 2x the p99 usage.
You: how do I get access to the production Kubernetes cluster?
Claude: [searches onboarding docs for "kubernetes cluster access production"]
Found your onboarding guide: "Getting Started - Cluster Access"
Updated Feb 10, 2026 by Arjun.
Steps for production cluster access:
1. Request access via this form: [link]
2. Wait for approval from your team lead (usually same day)
3. Install kubectl: brew install kubectl
4. Download our kubeconfig: aws eks update-kubeconfig --name prod-cluster --region us-east-1
5. Verify: kubectl get nodes
Note: Production is read-only by default. Write access requires a separate approval.
You: have we ever seen this S3 "SlowDown" error before?
payments are failing because S3 is rate limiting us
Claude: [searches postmortems for "S3 SlowDown rate limiting payments"]
Found 2 similar past incidents:
📄 "S3 Rate Limiting - Payments Service" (March 2025)
Root cause: too many small PutObject calls in a tight loop.
Fix: batch uploads, added exponential backoff. See PR #4821.
📄 "S3 Throttling During Flash Sale" (November 2024)
Root cause: single-prefix S3 bucket hitting 3,500 req/s limit.
Fix: switched to key-prefix randomization. Docs in architecture/s3-patterns.md
Want me to pull the full postmortem for either of these?
Добавление дополнительных источников со временем
# sources/notion.py
# Use the Notion API to pull pages from a database
# Same pattern: fetch content → ingest into ChromaDB
# sources/confluence.py
# Use the Confluence REST API
# Fetch pages by space key, convert HTML to markdown
# sources/slack.py
# Pull important threads from your #incidents or #platform channel
# Save and index them as "informal runbooks"
# Already works with the GitHub source
# Just add "your-org/incidents:postmortems" to github_repos in config
Структурирование руководств для улучшения поиска
# Database Connection Pool Exhaustion — Response Runbook
# DB issue
## Symptoms
- PagerDuty: "connection pool exhausted"
- Error in logs: "too many clients already"
- API latency spike on database-heavy endpoints
---
title: Database Connection Pool Exhaustion
type: runbook
tags: [database, postgresql, connections, production]
last_updated: 2026-03-14
owner: platform-team
---