This article is published in English.
Practical notes: How MCP Works: A Deep Dive with Code
Operable walkthrough of Practical notes: How MCP Works: A Deep Dive with Code: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: How MCP Works: A Deep Dive with Code. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list",
"params": {}
}
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"tools": [
{
"name": "search_web",
"description": "Search the web for a given query",
"inputSchema": {
"type": "object",
"properties": {
"query": { "type": "string" }
},
"required": ["query"]
}
}
]
}
}
{
"jsonrpc": "2.0",
"id": 2,
"method": "tools/call",
"params": {
"name": "search_web",
"arguments": {
"query": "latest news on MCP protocol"
}
}
}
{
"jsonrpc": "2.0",
"id": 2,
"result": {
"content": [
{
"type": "text",
"text": "Anthropic released MCP in Nov 2024 as an open standard..."
}
]
}
}
What is FastMCP?
When working through the What is FastMCP stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
from fastmcp import FastMCP
mcp = FastMCP("My Server") # creates the server
@mcp.tool() # registers the function as an MCP tool
def add(a: float, b: float) -> float:
"""Add two numbers.""" # docstring → tool description sent to the LLM
return a + b # type hints → JSON Schema sent to the LLM
mcp.run() # starts the stdio message loop
The three MCP primitives
When working through the The three MCP primitives stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
# servers/math_server.py
@mcp.tool()
def divide(a: float, b: float) -> float:
"""Divide a by b. Raises an error if b is zero."""
if b == 0:
raise ValueError("Cannot divide by zero")
return a / b
# servers/math_server.py
@mcp.resource("math://constants")
def get_math_constants() -> str:
"""Common mathematical constants."""
return f"π = {math.pi}\n e = {math.e}\n ..."
@mcp.resource("math://formulas/{category}")
def get_formulas(category: str) -> str:
"""Retrieve mathematical formulas by category (geometry | algebra | statistics)."""
catalog = {
"geometry": (
"Geometry Formulas:\n"
" Circle area: A = π × r²\n"
" Circle circumference: C = 2π × r\n"
" Rectangle area: A = length × width\n"
" Triangle area: A = (base × height) / 2\n"
" Sphere volume: V = (4/3) × π × r³\n"
),
"algebra": (
"Algebra Formulas:\n"
" Quadratic formula: x = (−b ± √(b²−4ac)) / 2a\n"
" Difference of squares: a²−b² = (a+b)(a−b)\n"
" Perfect square: (a+b)² = a²+2ab+b²\n"
" Sum of arithmetic seq: S = n(a₁+aₙ)/2\n"
),
"statistics": (
"Statistics Formulas:\n"
" Mean: μ = Σx / n\n"
" Variance: σ² = Σ(x−μ)² / n\n"
" Std Dev: σ = √(Σ(x−μ)² / n)\n"
" Z-score: z = (x−μ) / σ\n"
),
}
return catalog.get(
category,
f"Unknown category '{category}'. Available: geometry, algebra, statistics",
)
# servers/math_server.py
@mcp.prompt()
def math_tutor(difficulty: str = "intermediate") -> str:
return (
f"You are an expert math tutor for {difficulty}-level students. "
"Break every problem into numbered steps..."
)
Server summary
When working through the Server summary stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
From the Client’s Perspective
When working through the From the Client s stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
You type a query
│
▼
main.py ← entry point, parses args, kicks off async loop
│
▼
client/agent.py ← spawns 3 MCP servers, builds the agent, invokes it
│ │
│ ▼
│ utils/tracker.py ← fires on every LLM call, tool call, and result
│
▼
LangGraph ReAct loop ← think → call tool → observe → repeat
│
▼
servers/{math,text,data}_server.py ← each runs as an isolated subprocess
Step 1 — the entry point
When working through the Step 1 the entry stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
async def _run(queries: list) -> None:
from client.agent import run_query # imported here (late) to keep startup fast
for i, q in enumerate(queries):
await run_query(q)
def main() -> None:
...
asyncio.run(_run(queries))
Generate 8 random numbers between 5 and 50 using seed=42,
then calculate their statistics, and tell me if there are any outliers.
python3 main.py --query "Generate 8 random numbers between 5 and 50 using seed=42, \
then calculate their statistics, and tell me if there are any outliers."
╭────────────────────────────── 🔌 MCP Example ───────────────────────────────╮
│ Multi-Server MCP Demo │
│ │
│ Three FastMCP servers, each exposing tools + resources + prompts: │
│ ● Math Server — add, subtract, multiply, divide, power, sqrt, │
│ percentage │
│ ● Text Server — count_words, word_frequency, reverse, transform, │
│ extract_emails │
│ ● Data Server — generate_numbers, calculate_statistics, find_outliers, │
│ normalise │
│ │
│ Stack : FastMCP · LangChain · LangGraph · OpenAI · Rich │
│ Track : live Rich panels + JSONL log files under logs/ │
╰──────────────────────────────────────────────────────────────────────────────╯
Step 2 — Building the agent
When working through the Step 2 Building the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
log_file = str(LOGS_DIR / f"run_{int(time.time())}.jsonl")
tracker = MCPTracker(log_file=log_file)
╭──── 💬 USER QUERY ────╮
│ Generate 8 random... │
╰────────────────────────╯
def _server_config() -> Dict[str, Any]:
silent_env = {**os.environ, "FASTMCP_LOG_LEVEL": "ERROR"}
return {
"math_server": {
"command": sys.executable,
"args": ["servers/math_server.py"],
"transport": "stdio",
"env": silent_env,
},
"text_server": { ... },
"data_server": { ... },
}
client = MultiServerMCPClient(_server_config())
tools = await client.get_tools()
Connected to 3 servers (math_server, text_server, data_server) with 19 tools: add, subtract, multiply,
divide, power, square_root, calculate_percentage, count_words, word_frequency, reverse_text,
transform_case, find_and_replace, extract_emails, count_vowels_consonants, generate_numbers,
calculate_statistics, find_outliers, sort_values, normalize_values
model = ChatOpenAI(model=model_name, temperature=0)
agent = create_react_agent(model, tools)
START
│
▼
[call_model] ─── no tool call ──▶ END
│
tool call requested
│
▼
[call_tools]
│
▼
[call_model] (loop again with tool result in context)
config = {"callbacks": [tracker]}
result = await agent.ainvoke({"messages": [("human", query)]}, config=config)
Step 3 — Watching everything in real time
When working through the Step 3 Watching everything stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
agent.ainvoke() called
│
├─▶ on_chain_start() "▶ AGENT STARTED" panel
│
├─▶ on_chat_model_start() "🤖 LLM CALL #1" panel (timer starts)
├─▶ on_llm_end() "LLM responded ⏱ 1.23s → will call: add, multiply"
│
├─▶ on_tool_start() "🔧 TOOL CALL #1" panel (timer starts)
├─▶ on_tool_end() "✓ TOOL RESULT ⏱ 0.01s" panel
│
├─▶ on_tool_start() (second tool, if any)
├─▶ on_tool_end()
│
├─▶ on_chat_model_start() "🤖 LLM CALL #2" (LLM synthesises final answer)
├─▶ on_llm_end() no tools this time → loop ends
│
└─▶ agent.ainvoke() returns
TOOL_SERVER_MAP = {
"add": "math_server",
"count_words": "text_server",
"generate_numbers": "data_server",
...
}
SERVER_COLORS = {
"math_server": "cyan",
"text_server": "green",
"data_server": "yellow",
}
── 🤖 LLM CALL #1 model=gpt-5-mini messages=1 ──╮
│ Role Content preview │
│ Human Generate 8 random numbers between 5 and 50… │
╰──────────────────────────────────────────────────────╯
LLM responded ⏱ 5.35s → will call: generate_numbers
{
"name": "generate_numbers",
"arguments": {"count": 8, "min_val": 5, "max_val": 50, "seed": 42}
}
╭── 🔧 TOOL CALL #1 ──────────────────╮
│ Tool : generate_numbers │
│ Server: data_server │
│ Args : │
│ {'count': 8, 'min_val': 5, │
│ 'max_val': 50, 'seed': 42} │
╰───────────────────────────────────────╯
{"method": "tools/call", "params": {"name": "generate_numbers",
"arguments": {"count": 8, "min_val": 5, "max_val": 50, "seed": 42}}}
╭── ✓ TOOL RESULT ⏱ 0.63s ────────────────────────────╮
│ [33.77, 6.13, 17.38, 15.04, 38.14, 35.45, 45.15, 8.91] │
╰──────────────────────────────────────────────────────────╯
random.seed(42)
return [round(random.uniform(5, 50), 2) for _ in range(8)]
# → [33.77, 6.13, 17.38, 15.04, 38.14, 35.45, 45.15, 8.91]
╭── 🤖 LLM CALL #2 model=gpt-5-mini messages=3 ──╮
│ Role Content preview │
│ Human Generate 8 random numbers… │
│ AI (the tool-call decision) │
│ Tool [33.77, 6.13, 17.38, ...] │
╰──────────────────────────────────────────────────────╯
Tool : calculate_statistics
Server: data_server
Args : {'numbers': [33.77, 6.13, 17.38, 15.04, 38.14, 35.45, 45.15, 8.91]}
Result: {"count":8, "min":6.13, "max":45.15, "mean":24.9962,
"median":25.575, "std_dev":14.8181, "variance":219.5755,
"range":39.02, "q1":15.04, "q3":38.14}
Tool : find_outliers
Server: data_server
Args : {'numbers': [...], 'threshold': 2}
Result: {"outliers": [], "outlier_count": 0,
"total_checked": 8, "mean": 24.9962, "std_dev": 14.8181}
╭── 🤖 LLM CALL #4 messages=7 ──╮
│ Human / AI / Tool │
│ AI / Tool │
│ AI / Tool │
╰───────────────────────────────────╯
LLM responded ⏱ 30.25s
╭── ✅ FINAL ANSWER ───────────────────────────────╮
│ Random numbers: [33.77, 6.13, 17.38, ...] │
│ Mean: 24.9962 / Median: 25.575 / Std Dev: 14.8181│
│ Outliers: none (all |z| < 2) │
╰───────────────────────────────────────────────────╯
╭─────────────────────────────────────────────────────╮
│ LLM Calls 4 │
│ Tool Calls 3 │
│ Total Time 42.25s │
│ Log File logs/run_1776864980.jsonl │
╰─────────────────────────────────────────────────────╯
The full message sequence
When working through the The full message sequence stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the The full message sequence stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
1 HumanMessage "Generate 8 random numbers…"
2 AIMessage [tool_call: generate_numbers({count:8, min:5, max:50, seed:42})]
3 ToolMessage [33.77, 6.13, 17.38, 15.04, 38.14, 35.45, 45.15, 8.91]
4 AIMessage [tool_call: calculate_statistics({numbers:[...]})]
5 ToolMessage {count:8, mean:24.9962, std_dev:14.8181, ...}
6 AIMessage [tool_call: find_outliers({numbers:[...], threshold:2})]
7 ToolMessage {outliers:[], outlier_count:0, ...}
8 AIMessage "Here are the results. Random numbers: …" ← final answer
The JSONL log
The The JSONL log stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
{"timestamp": "...", "event": "agent_start", "data": {"runnable": "LangGraph", "run_id": "..."}}
{"timestamp": "...", "event": "llm_call", "data": {"call_number": 1, "model": "gpt-5-mini", "messages": 1}}
{"timestamp": "...", "event": "llm_end", "data": {"elapsed": "5.35s", "tool_calls_requested": ["generate_numbers"]}}
{"timestamp": "...", "event": "tool_start", "data": {"call_number": 1, "tool": "generate_numbers", "server": "data_server", "args": "..."}}
{"timestamp": "...", "event": "tool_end", "data": {"elapsed": "0.63s", "output_preview": "[33.77, 6.13, ...]"}}
{"timestamp": "...", "event": "llm_call", "data": {"call_number": 2, "model": "gpt-5-mini", "messages": 3}}
...
{"timestamp": "...", "event": "run_summary", "data": {"llm_calls": 4, "tool_calls": 3, "total_elapsed": "42.25s"}}
Sources
The Sources stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
A message from our Founder
The A message from our stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The A message from our stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for c7efc4f69698: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 0/888: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 1/888: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 2/888: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 3/888: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.