This article is published in English.
Practical notes: I Built an MCP Server That Keeps My Work Journal — Here’s
Operable walkthrough of Practical notes: I Built an MCP Server That Keeps My Work Journal — Here’s: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: I Built an MCP Server That Keeps My Work Journal — Here’s Everything I Learned. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
What it does
When working through the What it does stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
## 14:32 #bugfix #websocket
Fixed the race condition in the WebSocket broadcast queue
## 16:10 #testing
Wrote E2E test covering two-client sync
How to build one (the whole recipe)
When working through the How to build one stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
1. The skeleton is genuinely small
When working through the 1 The skeleton is stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the 1 The skeleton is stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
npm install @modelcontextprotocol/server zod
import { McpServer } from '@modelcontextprotocol/server';
import { StdioServerTransport } from '@modelcontextprotocol/server/stdio';
import * as z from 'zod/v4';
const server = new McpServer({ name: 'dev-diary', version: '1.0.0' });
server.registerTool(
'log_work',
{
description: 'Append a timestamped entry to the developer diary...',
inputSchema: z.object({
text: z.string().min(1),
tags: z.array(z.string()).optional(),
date: z.string().regex(/^\d{4}-\d{2}-\d{2}$/).optional(),
}),
},
async ({ text, tags = [], date }) => {
// ...append to diary/YYYY-MM-DD.md...
return { content: [{ type: 'text', text: 'Logged.' }] };
},
);
await server.connect(new StdioServerTransport());
2. Descriptions are prompts, not documentation
The 2 Descriptions are prompts stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
// ❌ documentation-style
description: 'Appends an entry to the diary.'
// ✅ prompt-style
description: 'Append a timestamped entry to the developer diary for today.
Use this whenever the user says they finished/did/fixed something and
wants it recorded.'
3. Design tools around questions, not tables
The 3 Design tools around stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
The demo problem (and its elegant solution)
The The demo problem and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The The demo problem and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
import { Client } from '@modelcontextprotocol/client';
import { StdioClientTransport } from '@modelcontextprotocol/client/stdio';
const transport = new StdioClientTransport({
command: 'node',
args: ['dist/server.js'], // spawns the server as a child process
});
const client = new Client({ name: 'demo-client', version: '1.0.0' });
await client.connect(transport);
// Exactly what Claude Desktop does under the hood:
const { tools } = await client.listTools();
await client.callTool({ name: 'log_work', arguments: {
text: 'Fixed the race condition in the broadcast queue',
tags: ['bugfix', 'websocket'],
}});
=== 1. listTools ===
• log_work — Append a timestamped entry to the developer diary...
• search_diary — Full-text search across every entry...
• daily_summary — Everything logged on a given date...
• stats — Totals, active days, streaks, top tags...
=== 2. log_work x3 ===
Logged to 2026-08-22.md at 19:05 (tags: bugfix, websocket)
...
=== 5. stats ===
📊 1 entries across 2 day(s)
🔥 Streak: 2 consecutive day(s)
🏷️ Top tags: #bugfix (1), #websocket (1)
Things the tutorials don’t tell you
For the Things the tutorials don stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Connecting it for real
For the Connecting it for real stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
{
"mcpServers": {
"dev-diary": {
"command": "node",
"args": ["/absolute/path/to/dev-diary-mcp/dist/server.js"],
"env": { "DIARY_DIR": "/home/you/journal" }
}
}
}
Why Markdown-as-database won
For the Why Markdown-as-database won stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Why Markdown-as-database won stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Try it
When working through the Try it stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
git clone https://github.com/rogeriolaa/dev-diary-mcp
cd dev-diary-mcp && npm install && npm run build && npm run demo
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 0f6d55786c75: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.