Home / Articles / Stop Writing Custom APIs for Your AI Agents

This article is published in English.

Stop Writing Custom APIs for Your AI Agents

Operable walkthrough of Stop Writing Custom APIs for Your AI Agents: contracts, checks, and drop-in code slots for teams shipping this pattern.

1254 words

Use this as an operator-facing rebuild of the ideas in “Stop Writing Custom APIs for Your AI Agents: Build an MCP Server in 5 Minutes”: clear stages, ordered code slots, and recovery notes that survive a handoff. Overview works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Step 1: The Architecture & Prerequisites

For Step 1: The Architecture & Prerequisites, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

pip install mcp

Step 2: Building the MCP Server

For Step 2: Building the MCP Server, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

import sqlite3
import json
import os
import sys
from mcp.server.mcpserver import MCPServer

# Initialize the MCP server
mcp = MCPServer(name="Enterprise_SQL_Agent")

# Force the database to be created in the exact same folder as this script
BASE_DIR = os.path.dirname(os.path.abspath(__file__))
DB_PATH = os.path.join(BASE_DIR, "enterprise.db")
def setup_dummy_db():
    """Create a sample employee database for the demo"""
    try:
        conn = sqlite3.connect(DB_PATH)
        cursor = conn.cursor()

        cursor.execute('''CREATE TABLE IF NOT EXISTS employees
                          (id INTEGER PRIMARY KEY, name TEXT, role TEXT, salary INTEGER)''')
        cursor.execute("DELETE FROM employees")

        employees = [
            ("Alice", "Data Scientist", 120000),
            ("Bob", "DevOps Engineer", 115000),
            ("Charlie", "AI Researcher", 135000)
        ]

        cursor.executemany("INSERT INTO employees (name, role, salary) VALUES (?, ?, ?)", employees)
        conn.commit()
        conn.close()
        print("Database initialized successfully.", file=sys.stderr)
    except Exception as e:
        print(f"Database setup error: {e}", file=sys.stderr)

Step 3: Exposing the Database to the AI

For Step 3: Exposing the Database to the AI, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For Step 3: Exposing the Database to the AI, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

@mcp.tool()
def query_employee_database(sql_query: str) -> str:
    """
    Executes a SQL SELECT query against the enterprise.db database.

    The database contains an 'employees' table with columns:
    - id (INTEGER PRIMARY KEY)
    - name (TEXT)
    - role (TEXT)
    - salary (INTEGER)

    SECURITY: Only READ operations (SELECT) are permitted.
    """

    # Safety Check: Block destructive SQL commands
    dangerous_keywords = ["DROP", "DELETE", "UPDATE", "INSERT", "ALTER"]
    if any(keyword in sql_query.upper() for keyword in dangerous_keywords):
        return "Error: Only SELECT queries are authorized for this tool."

    try:
        conn = sqlite3.connect(DB_PATH)
        cursor = conn.cursor()
        cursor.execute(sql_query)
        results = cursor.fetchall()

        # Format the output as JSON so the LLM can read it cleanly
        column_names = [description[0] for description in cursor.description]
        formatted_results = [dict(zip(column_names, row)) for row in results]

        conn.close()
        return json.dumps(formatted_results, indent=2)

    except Exception as e:
        return f"Database error: {str(e)}"
if __name__ == "__main__":
    setup_dummy_db()
    mcp.run()

Step 4: Connecting Claude Desktop

When working through Step 4: Connecting Claude Desktop, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

{
  "mcpServers": {
    "enterprise-sql": {
      "command": "C:\\Users\\YourName\\.conda\\envs\\your_env\\python.exe",
      "args": [
        "D:\\Your\\Project\\Path\\mcp_server.py"
      ]
    }
  }
}

Step 5: The Payoff

When working through Step 5: The Payoff, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

What’s Next?

When working through What’s Next?, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through What’s Next?, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Operational checklist

When working through Operational checklist, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for aed9f8a61db3: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.