Home / Articles / Practical notes: MCP Servers Fail in Two Ways, and Both Are Preventable. Here’s

This article is published in English.

Practical notes: MCP Servers Fail in Two Ways, and Both Are Preventable. Here’s

Operable walkthrough of Practical notes: MCP Servers Fail in Two Ways, and Both Are Preventable. Here’s: contracts, checks, and drop-in code slots for teams shipping this pattern.

1878 words

The following notes reconstruct a practical path around “MCP Servers Fail in Two Ways, and Both Are Preventable. Here’s the Guardrail Layer.”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

An MCP server is a trust boundary, not just an integration

The An MCP server is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Failure one: permission failures, the agent inherits more trust than the task needs

The Failure one permission failures stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Failure two: interface failures, the agent can’t tell what a tool does or trust what it gets back

The Failure two interface failures stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Failure two interface failures stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

The pattern everyone misses: you harden the model’s output and trust the tool layer’s input

For the The pattern everyone misses stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Building an MCP guardrail layer

For the Building an MCP guardrail stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

from dataclasses import dataclass
from datetime import datetime, timedelta
from typing import Optional

@dataclass
class ScopedToken:
 audience: str # which MCP server this token is valid for
 permissions: list[str] # e.g. ["read", "create"] - never assume "all"
 issued_at: datetime
 expires_at: datetime
 source_user: str # who originally triggered this, for audit
def issue_scoped_token(user_token: ScopedToken, tool_name: str,
 required_permissions: list[str]) -> ScopedToken:
 # Never grant more than the tool declares it needs
 granted = [p for p in required_permissions if p in user_token.permissions]
 if set(required_permissions) - set(granted):
 raise PermissionError(
 f"{tool_name} requires {required_permissions}, "
 f"caller only has {user_token.permissions}"
 )
 return ScopedToken(
 audience=tool_name,
 permissions=granted,
 issued_at=datetime.utcnow(),
 expires_at=datetime.utcnow() + timedelta(minutes=5),
 source_user=user_token.source_user,
 )
def enforce_audience(token: ScopedToken, expected_tool: str) -> None:
 if token.audience != expected_tool:
 raise PermissionError(
 f"Token issued for '{token.audience}' cannot be used on '{expected_tool}'"
 )
def file_support_request(customer_email: str, issue_type: str, description: str) -> dict:
 ticket = create_ticket(issue_type, description)
 add_comment(ticket.id, f"Filed by {customer_email}")
 assign_ticket(ticket.id, team=route_by_type(issue_type))
 notify_user(customer_email, ticket.id)
 return {
 "ticket_id": ticket.id,
 "status": "open",
 "assigned_team": ticket.team,
 }
def safe_error(internal_message: str, request_id: str) -> dict:
 # internal_message goes to your logs, never to the model
 log.error(internal_message, extra={"request_id": request_id})
 return {
 "content": [{
 "type": "text",
 "text": f"Unable to complete the request. Request ID: {request_id}. "
 f"Try again or contact support."
 }],
 "isError": True,
 }
MAX_TOOLS_PER_SERVER = 15

def register_tool(server, tool):
 if len(server.tools) >= MAX_TOOLS_PER_SERVER:
 raise ValueError(
 f"{server.name} already has {len(server.tools)} tools. "
 f"Split into a domain-specific server instead of adding more."
 )
 server.tools.append(tool)

How to tell if your MCP server already has this problem

For the How to tell if stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the How to tell if stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Where this fits in a production stack

When working through the Where this fits in stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

The tool call is where the trust decision happens

When working through the The tool call is stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Frequently asked questions

When working through the Frequently asked questions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Frequently asked questions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 964498802023: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 0/631: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 0/650: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 1/650: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.