实用提示:利用团队的操作手册构建个人MCP服务器
《实用笔记》操作指南:如何利用团队的运行手册构建个人 MCP 服务器,以及为采用该模式的团队提供的合同、检查项和即插即用代码模块。
我们正在具体构建的内容
You: we're getting high DB connection errors on payments-service,
what's the runbook?
Claude: [searches your runbook MCP server]
Found your runbook: "Database Connection Pool Exhaustion"
Last updated March 14, 2026 by Priya.
Steps:
1. Check current pool size: kubectl exec -it <pod> -- env | grep DB_POOL
2. Temporarily increase: kubectl set env deploy/payments-service DB_POOL_SIZE=25
3. Monitor connections: watch -n2 'kubectl exec...'
4. If connections still growing, check for connection leaks in logs...
项目结构
在网关处进行身份验证,在数据层进行重新授权。仅凭承载令牌无法界定租户边界。
runbook-mcp/
├── server.py # The MCP server — main file
├── indexer.py # Indexes docs into ChromaDB
├── sources/
│ ├── local.py # Index local markdown files
│ ├── github.py # Pull from GitHub repo
│ └── confluence.py # Pull from Confluence (optional)
├── config.py # Your settings
├── sync.py # Keep index fresh
└── requirements.txt
pip install mcp chromadb sentence-transformers \
httpx python-frontmatter watchdog \
--break-system-packages
第一步 — 配置
在第一步配置阶段,作者应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前了解这些成本可以避免在流程从演示环境转向共享环境时出现意外费用。应在网关处进行身份验证,并在数据层面重新授权。仅凭承载令牌并不足以界定租户边界。在第一步配置阶段,作者应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。
import os
from dataclasses import dataclass, field
@dataclass
class Config:
# ── Local docs ──────────────────────────────────────
# Point these at wherever your docs actually are
local_docs_paths: list = field(default_factory=lambda: [
os.path.expanduser("~/docs/runbooks"),
os.path.expanduser("~/docs/architecture"),
os.path.expanduser("~/docs/postmortems"),
])
# ── GitHub (optional) ───────────────────────────────
github_token: str = os.getenv("GITHUB_TOKEN", "")
github_repos: list = field(default_factory=lambda: [
# format: "org/repo:path/to/docs"
# "your-org/infrastructure:docs/runbooks",
# "your-org/platform:docs/architecture",
])
# ── Confluence (optional) ───────────────────────────
confluence_url: str = os.getenv("CONFLUENCE_URL", "")
confluence_token: str = os.getenv("CONFLUENCE_TOKEN", "")
confluence_spaces: list = field(default_factory=lambda: [
# "PLAT", # Platform team space
# "SRE", # SRE space
])
# ── Vector DB ───────────────────────────────────────
vector_db_path: str = os.path.expanduser("~/.runbook-mcp/index")
# ── Team context ────────────────────────────────────
team_name: str = "Platform Team"
company: str = "YourCompany"
config = Config()
第2步——文档来源
在处理“第2步:文档”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议使用小型、可测试的单元而非庞大的脚本。当某一步骤失败时,故障应能指向单一的责任模块,而非复杂的流程链。 需为每次调用记录工具名称、参数哈希值、延迟时间以及执行结果。没有这些记录的话,调试过程将会浪费大量时间。
本地文件
在处理“本地文件”阶段时,首先写下契约内容:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功判定标准,杜绝无声的半完成状态。 需记录每次调用的工具名称、参数哈希值、延迟时间以及最终结果。没有这些记录,调试过程将会浪费大量时间。
import os
import glob
import hashlib
import frontmatter
from datetime import datetime
def classify_doc(filepath: str, content: str) -> str:
"""Guess doc type from filename and content."""
fp = filepath.lower()
content_lower = content.lower()
if any(kw in fp for kw in ["runbook", "procedure", "playbook"]):
return "runbook"
elif any(kw in fp for kw in ["postmortem", "incident", "retrospective"]):
return "postmortem"
elif any(kw in fp for kw in ["architecture", "design", "adr"]):
return "architecture"
elif any(kw in fp for kw in ["onboard", "setup", "getting-started"]):
return "onboarding"
elif any(kw in fp for kw in ["sop", "standard", "policy"]):
return "sop"
# Check content too
elif "## steps" in content_lower or "## procedure" in content_lower:
return "runbook"
elif "root cause" in content_lower and "action items" in content_lower:
return "postmortem"
else:
return "general"
def load_local_docs(paths: list[str]) -> list[dict]:
"""Load all markdown files from local directories."""
docs = []
for base_path in paths:
if not os.path.exists(base_path):
print(f" ⚠️ Path not found: {base_path}")
continue
md_files = glob.glob(f"{base_path}/**/*.md", recursive=True)
md_files += glob.glob(f"{base_path}/**/*.mdx", recursive=True)
for filepath in md_files:
try:
with open(filepath, "r", errors="replace") as f:
raw = f.read()
# Parse frontmatter if present
try:
post = frontmatter.loads(raw)
content = post.content
metadata = dict(post.metadata)
except:
content = raw
metadata = {}
if not content.strip():
continue
rel_path = os.path.relpath(filepath, base_path)
doc_type = metadata.get("type") or classify_doc(rel_path, content)
# Get title from frontmatter or first heading
title = metadata.get("title") or _extract_title(content) or rel_path
file_stat = os.stat(filepath)
last_modified = datetime.fromtimestamp(
file_stat.st_mtime
).strftime("%Y-%m-%d")
docs.append({
"content": content,
"source": rel_path,
"filepath": filepath,
"type": doc_type,
"title": title,
"last_modified": last_modified,
"tags": metadata.get("tags", []),
"hash": hashlib.md5(content.encode()).hexdigest()
})
except Exception as e:
print(f" ⚠️ Skipped {filepath}: {e}")
return docs
def _extract_title(content: str) -> str | None:
"""Extract the first H1 heading from markdown."""
for line in content.split("\n"):
line = line.strip()
if line.startswith("# "):
return line[2:].strip()
return None
GitHub 源码
在处理 GitHub Source 阶段时,首先需明确合同规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合约定。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 需为每次调用记录工具名称、参数哈希值、延迟时间以及最终结果。没有这些记录,调试过程将会浪费大量时间。 在处理 GitHub Source 阶段时,首先需明确合同规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合约定。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。
import httpx
import base64
from config import config
async def load_github_docs() -> list[dict]:
"""Pull markdown docs from GitHub repos."""
if not config.github_token or not config.github_repos:
return []
headers = {
"Authorization": f"token {config.github_token}",
"Accept": "application/vnd.github.v3+json"
}
docs = []
async with httpx.AsyncClient(headers=headers) as client:
for repo_spec in config.github_repos:
# Parse "org/repo:path"
if ":" in repo_spec:
repo, path = repo_spec.split(":", 1)
else:
repo, path = repo_spec, "docs"
print(f" 📥 GitHub: {repo}/{path}")
try:
# Get all files in the path
resp = await client.get(
f"https://api.github.com/repos/{repo}/contents/{path}",
timeout=15
)
if resp.status_code != 200:
print(f" ⚠️ Failed: {resp.status_code}")
continue
files = resp.json()
if isinstance(files, dict):
files = [files]
for file_info in files:
if not file_info.get("name", "").endswith(".md"):
continue
# Get file content
file_resp = await client.get(
file_info["url"], timeout=15
)
if file_resp.status_code != 200:
continue
file_data = file_resp.json()
content = base64.b64decode(
file_data["content"]
).decode("utf-8", errors="replace")
docs.append({
"content": content,
"source": f"{repo}/{file_info['path']}",
"filepath": file_info["html_url"],
"type": "general",
"title": file_info["name"].replace(".md", ""),
"last_modified": "",
"tags": [],
"hash": file_data.get("sha", "")[:8]
})
except Exception as e:
print(f" ⚠️ GitHub error for {repo}: {e}")
print(f" ✅ Loaded {len(docs)} docs from GitHub")
return docs
第 3 步 — 索引器
在第三步的索引阶段中,若能将其视为可度量的对象,则效果最佳。在扩大范围之前,先记录一个成功的示例、一个失败案例以及回滚说明。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某一步骤失败时,故障应能指向单一的责任主体,而非复杂的流程链。 应提供具有明确结构规范和清晰副作用标识的工具。主机需要在自动批准之前知道哪些调用会修改状态。
import os
import json
from pathlib import Path
import chromadb
from chromadb.utils import embedding_functions
from langchain.text_splitter import MarkdownTextSplitter
from sources.local import load_local_docs
from config import config
class RunbookIndexer:
def __init__(self):
os.makedirs(config.vector_db_path, exist_ok=True)
self.db = chromadb.PersistentClient(path=config.vector_db_path)
self.embedder = embedding_functions.SentenceTransformerEmbeddingFunction(
model_name="all-MiniLM-L6-v2"
)
self.collection = self.db.get_or_create_collection(
name="runbooks",
embedding_function=self.embedder,
metadata={"hnsw:space": "cosine"}
)
# Track ingested docs by hash
self.index_file = Path(config.vector_db_path) / "doc_index.json"
self.doc_index = self._load_doc_index()
self.splitter = MarkdownTextSplitter(
chunk_size=600,
chunk_overlap=80
)
def _load_doc_index(self) -> dict:
if self.index_file.exists():
return json.loads(self.index_file.read_text())
return {}
def _save_doc_index(self):
self.index_file.write_text(json.dumps(self.doc_index, indent=2))
def build(self):
"""Build the full index from all sources."""
print("📚 Loading documents from all sources...\n")
all_docs = []
# Local docs
print("📁 Local files:")
local_docs = load_local_docs(config.local_docs_paths)
all_docs.extend(local_docs)
print(f" Loaded {len(local_docs)} local documents\n")
# Index everything
print(f"💾 Indexing {len(all_docs)} documents...")
new_count = 0
skip_count = 0
for doc in all_docs:
source = doc["source"]
doc_hash = doc["hash"]
# Skip unchanged docs
if self.doc_index.get(source) == doc_hash:
skip_count += 1
continue
# Remove old version
try:
self.collection.delete(where={"source": source})
except:
pass
# Split into chunks
chunks = self.splitter.split_text(doc["content"])
if not chunks:
continue
self.collection.add(
documents=chunks,
metadatas=[{
"source": source,
"title": doc["title"],
"type": doc["type"],
"last_modified": doc["last_modified"],
"tags": ", ".join(doc.get("tags", [])),
"chunk_index": i,
"total_chunks": len(chunks)
} for i, _ in enumerate(chunks)],
ids=[f"{source}::chunk_{i}" for i in range(len(chunks))]
)
self.doc_index[source] = doc_hash
new_count += 1
print(f" ✅ Indexed: {doc['title']} ({len(chunks)} chunks)")
self._save_doc_index()
print(f"\n✅ Done: {new_count} new, {skip_count} unchanged")
print(f" Total chunks in index: {self.collection.count()}")
def search(self, query: str, n_results: int = 5, doc_type: str = None) -> list[dict]:
"""Search the index semantically."""
where = {"type": doc_type} if doc_type else None
try:
results = self.collection.query(
query_texts=[query],
n_results=n_results,
where=where,
include=["documents", "metadatas", "distances"]
)
except Exception as e:
return []
matches = []
for doc, meta, dist in zip(
results["documents"][0],
results["metadatas"][0],
results["distances"][0]
):
relevance = round((1 - dist) * 100, 1)
if relevance < 30:
continue
matches.append({
"content": doc,
"source": meta["source"],
"title": meta["title"],
"type": meta["type"],
"last_modified": meta.get("last_modified", ""),
"relevance": relevance,
"chunk_index": meta.get("chunk_index", 0),
"total_chunks": meta.get("total_chunks", 1)
})
return matches
def get_full_doc(self, source: str) -> str | None:
"""Get all chunks for a specific document and reconstruct it."""
try:
results = self.collection.get(
where={"source": source},
include=["documents", "metadatas"]
)
if not results["documents"]:
return None
# Sort chunks by index and join
paired = list(zip(
results["documents"],
results["metadatas"]
))
paired.sort(key=lambda x: x[1].get("chunk_index", 0))
return "\n\n".join(doc for doc, _ in paired)
except Exception as e:
return None
def list_docs(self, doc_type: str = None) -> list[dict]:
"""List all indexed documents."""
try:
where = {"type": doc_type} if doc_type else None
results = self.collection.get(
where=where,
include=["metadatas"]
)
# Deduplicate by source
seen = {}
for meta in results["metadatas"]:
source = meta["source"]
if source not in seen:
seen[source] = {
"source": source,
"title": meta["title"],
"type": meta["type"],
"last_modified": meta.get("last_modified", "")
}
return sorted(seen.values(), key=lambda x: x["title"])
except Exception as e:
return []
第四步 — MCP服务器
第4步:将MCP阶段视为可度量的对象来处理效果最佳。在扩大范围之前,需记录一份理想状态示例、一个故障案例以及回滚说明。 将该阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 使用具有严格结构定义且带有明确副作用标签的工具。主机需要在自动批准之前知晓哪些调用会改变系统状态。
import asyncio
import json
from mcp.server import Server
from mcp.server.stdio import stdio_server
from mcp import types
from indexer import RunbookIndexer
app = Server("runbook-mcp-server")
indexer = RunbookIndexer()
@app.list_tools()
async def list_tools() -> list[types.Tool]:
"""Define all tools available to Claude."""
return [
types.Tool(
name="search_runbooks",
description=(
"Search your team's runbooks and documentation semantically. "
"Use this when someone asks about a procedure, error, "
"or operational task. Returns the most relevant doc sections."
),
inputSchema={
"type": "object",
"properties": {
"query": {
"type": "string",
"description": (
"What to search for. Be descriptive. "
"Examples: 'database connection pool exhaustion', "
"'nginx 502 errors', 'how to rotate API keys', "
"'deployment rollback procedure'"
)
},
"doc_type": {
"type": "string",
"enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
"description": "Filter by document type (optional)"
},
"n_results": {
"type": "integer",
"description": "Number of results to return (default: 5)"
}
},
"required": ["query"]
}
),
types.Tool(
name="get_runbook",
description=(
"Get the full content of a specific runbook or doc by its source path. "
"Use this after search_runbooks to get the complete document. "
"The source path comes from search results."
),
inputSchema={
"type": "object",
"properties": {
"source": {
"type": "string",
"description": "The source path from search results (e.g. 'runbooks/db-connection-pool.md')"
}
},
"required": ["source"]
}
),
types.Tool(
name="list_runbooks",
description=(
"List all available runbooks and documents by type. "
"Use when someone asks 'what runbooks do we have' or "
"'list all incident procedures'."
),
inputSchema={
"type": "object",
"properties": {
"doc_type": {
"type": "string",
"enum": ["runbook", "postmortem", "architecture", "onboarding", "sop", "general"],
"description": "Filter by type (optional — omit for all)"
}
}
}
),
types.Tool(
name="find_similar_incidents",
description=(
"Search past postmortems and incident reports for similar issues. "
"Use when someone says 'have we seen this before' or "
"'was there a similar incident'. Returns relevant past incidents."
),
inputSchema={
"type": "object",
"properties": {
"description": {
"type": "string",
"description": "Description of the current issue or symptoms"
}
},
"required": ["description"]
}
),
types.Tool(
name="get_onboarding_docs",
description=(
"Get onboarding and setup documentation. "
"Use when someone is new or asks how to set something up."
),
inputSchema={
"type": "object",
"properties": {
"topic": {
"type": "string",
"description": "What they need to set up or learn (e.g. 'kubectl access', 'AWS credentials', 'local development')"
}
},
"required": ["topic"]
}
),
]
@app.call_tool()
async def call_tool(name: str, arguments: dict) -> list[types.TextContent]:
"""Handle tool calls from Claude."""
# ── search_runbooks ────────────────────────────────────
if name == "search_runbooks":
query = arguments["query"]
doc_type = arguments.get("doc_type")
n = arguments.get("n_results", 5)
results = indexer.search(query, n_results=n, doc_type=doc_type)
if not results:
return [types.TextContent(
type="text",
text=f"No documents found matching: '{query}'"
)]
output_parts = [f"Found {len(results)} relevant documents:\n"]
for i, r in enumerate(results, 1):
output_parts.append(
f"\n--- Result {i} ---\n"
f"Title: {r['title']}\n"
f"Type: {r['type']}\n"
f"Source: {r['source']}\n"
f"Last modified: {r['last_modified']}\n"
f"Relevance: {r['relevance']}%\n"
f"Chunk {r['chunk_index']+1}/{r['total_chunks']}\n\n"
f"{r['content']}"
)
return [types.TextContent(type="text", text="\n".join(output_parts))]
# ── get_runbook ────────────────────────────────────────
elif name == "get_runbook":
source = arguments["source"]
content = indexer.get_full_doc(source)
if not content:
return [types.TextContent(
type="text",
text=f"Document not found: {source}\n"
f"Try searching with search_runbooks first."
)]
return [types.TextContent(
type="text",
text=f"Full document: {source}\n\n{content}"
)]
# ── list_runbooks ──────────────────────────────────────
elif name == "list_runbooks":
doc_type = arguments.get("doc_type")
docs = indexer.list_docs(doc_type=doc_type)
if not docs:
label = f" of type '{doc_type}'" if doc_type else ""
return [types.TextContent(
type="text",
text=f"No documents found{label}. Run the indexer first."
)]
# Group by type
by_type: dict[str, list] = {}
for doc in docs:
t = doc["type"]
by_type.setdefault(t, []).append(doc)
lines = [f"Available documents ({len(docs)} total):\n"]
for dtype, dtype_docs in sorted(by_type.items()):
lines.append(f"\n## {dtype.title()}s ({len(dtype_docs)})")
for doc in dtype_docs:
modified = f" — updated {doc['last_modified']}" if doc['last_modified'] else ""
lines.append(f" - {doc['title']}{modified}\n [{doc['source']}]")
return [types.TextContent(type="text", text="\n".join(lines))]
# ── find_similar_incidents ─────────────────────────────
elif name == "find_similar_incidents":
description = arguments["description"]
# Search specifically in postmortems
results = indexer.search(
query=description,
n_results=5,
doc_type="postmortem"
)
if not results:
# Broaden to all docs if no postmortems found
results = indexer.search(query=description, n_results=3)
if not results:
return [types.TextContent(
type="text",
text="No similar incidents found in the knowledge base."
)]
output = [f"Similar past incidents:\n"]
for r in results:
output.append(
f"\n📄 {r['title']} [{r['source']}]\n"
f"Relevance: {r['relevance']}%\n"
f"Last modified: {r['last_modified']}\n\n"
f"{r['content'][:600]}..."
)
return [types.TextContent(type="text", text="\n".join(output))]
# ── get_onboarding_docs ────────────────────────────────
elif name == "get_onboarding_docs":
topic = arguments["topic"]
results = indexer.search(
query=f"setup onboarding {topic}",
n_results=4,
doc_type="onboarding"
)
if not results:
# Fall back to all doc types
results = indexer.search(
query=f"how to setup {topic}", n_results=3
)
if not results:
return [types.TextContent(
type="text",
text=f"No onboarding docs found for: {topic}"
)]
output = [f"Onboarding docs for '{topic}':\n"]
for r in results:
output.append(
f"\n📄 {r['title']}\n"
f"Source: {r['source']}\n\n"
f"{r['content']}"
)
return [types.TextContent(type="text", text="\n".join(output))]
return [types.TextContent(type="text", text=f"Unknown tool: {name}")]
async def main():
async with stdio_server() as (read_stream, write_stream):
await app.run(
read_stream,
write_stream,
app.create_initialization_options()
)
if __name__ == "__main__":
asyncio.run(main())
第5步 — 保持新鲜度
第5步的“保持新鲜度”阶段同样适合以可度量的对象来处理。在扩大范围之前,需记录一份理想状态示例、一个故障案例以及回滚说明。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外费用。
"""
Run this to rebuild the index when docs change.
Can be triggered manually or via a cron job.
"""
from indexer import RunbookIndexer
from datetime import datetime
def sync():
print(f"🔄 Syncing runbook index - {datetime.now().strftime('%Y-%m-%d %H:%M')}\n")
indexer = RunbookIndexer()
indexer.build()
print("\n✅ Sync complete")
if __name__ == "__main__":
sync()
# Rebuild every night at midnight
0 0 * * * cd ~/runbook-mcp && python sync.py >> ~/.runbook-mcp/sync.log 2>&1
# Or: rebuild every time someone pushes to the docs repo
# Add to your GitHub Actions workflow:
# - name: Sync MCP index
# run: ssh your-server "cd ~/runbook-mcp && python sync.py"
将其连接到 Claude Desktop
Mac: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"runbooks": {
"command": "python",
"args": ["/Users/yourname/runbook-mcp/server.py"],
"env": {
"GITHUB_TOKEN": "your-token-here"
}
}
}
}
首次设置——5步完成全部流程
# 1. Clone and install
git clone your-repo/runbook-mcp
cd runbook-mcp
pip install -r requirements.txt
# 2. Set up your docs paths in config.py
# Point local_docs_paths at wherever your runbooks actually are
nano config.py
# 3. Build the index
python sync.py
# 4. Test it works
python -c "
from indexer import RunbookIndexer
idx = RunbookIndexer()
results = idx.search('database connection pool')
for r in results:
print(r['title'], r['relevance'])
"
# 5. Add to Claude Desktop config and restart
设置完成后进行真实对话
You: we're seeing OOM kills on the API pods, what's the procedure?
Claude: [searches runbooks for "OOM killed API pods"]
Found your runbook: "Pod OOMKilled - Response Procedure"
Updated April 3, 2026.
Immediate steps:
1. Confirm: kubectl describe pod <pod-name> | grep -A5 "Last State"
2. Check current limits: kubectl get deploy api-service -o yaml | grep -A3 resources
3. Quick fix (buys time): kubectl patch deploy api-service -p '{"spec":{"template":{"spec":{"containers":[{"name":"api","resources":{"limits":{"memory":"1Gi"}}}]}}}}'
4. Restart: kubectl rollout restart deploy/api-service
5. Monitor: watch kubectl top pods -n production
Root fix: Update Helm values.yaml - memory limit should be at least 2x the p99 usage.
You: how do I get access to the production Kubernetes cluster?
Claude: [searches onboarding docs for "kubernetes cluster access production"]
Found your onboarding guide: "Getting Started - Cluster Access"
Updated Feb 10, 2026 by Arjun.
Steps for production cluster access:
1. Request access via this form: [link]
2. Wait for approval from your team lead (usually same day)
3. Install kubectl: brew install kubectl
4. Download our kubeconfig: aws eks update-kubeconfig --name prod-cluster --region us-east-1
5. Verify: kubectl get nodes
Note: Production is read-only by default. Write access requires a separate approval.
You: have we ever seen this S3 "SlowDown" error before?
payments are failing because S3 is rate limiting us
Claude: [searches postmortems for "S3 SlowDown rate limiting payments"]
Found 2 similar past incidents:
📄 "S3 Rate Limiting - Payments Service" (March 2025)
Root cause: too many small PutObject calls in a tight loop.
Fix: batch uploads, added exponential backoff. See PR #4821.
📄 "S3 Throttling During Flash Sale" (November 2024)
Root cause: single-prefix S3 bucket hitting 3,500 req/s limit.
Fix: switched to key-prefix randomization. Docs in architecture/s3-patterns.md
Want me to pull the full postmortem for either of these?
逐步添加更多数据源
# sources/notion.py
# Use the Notion API to pull pages from a database
# Same pattern: fetch content → ingest into ChromaDB
# sources/confluence.py
# Use the Confluence REST API
# Fetch pages by space key, convert HTML to markdown
# sources/slack.py
# Pull important threads from your #incidents or #platform channel
# Save and index them as "informal runbooks"
# Already works with the GitHub source
# Just add "your-org/incidents:postmortems" to github_repos in config
优化运行手册结构以便更好搜索
# Database Connection Pool Exhaustion — Response Runbook
# DB issue
## Symptoms
- PagerDuty: "connection pool exhausted"
- Error in logs: "too many clients already"
- API latency spike on database-heavy endpoints
---
title: Database Connection Pool Exhaustion
type: runbook
tags: [database, postgresql, connections, production]
last_updated: 2026-03-14
owner: platform-team
---