Home / Articles / I Audited 300 MCP-Related npm Packages. My Scanner Couldn’t Inspect Runtime

This article is published in English.

I Audited 300 MCP-Related npm Packages. My Scanner Couldn’t Inspect Runtime

Operable walkthrough of I Audited 300 MCP-Related npm Packages. My Scanner Couldn’t Inspect Runtime: contracts, checks, and drop-in code slots for teams shipping this pattern.

1540 words

Use this as an operator-facing rebuild of the ideas in “I Audited 300 MCP-Related npm Packages. My Scanner Couldn’t Inspect Runtime Code in 217. So I created Driftward.”: clear stages, ordered code slots, and recovery notes that survive a handoff.

The short version

The The short version stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Why MCP changes the risk calculation

The Why MCP changes the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

What you measured

The What you measured stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Results

The Results stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

The two prompt-injection hits were false positives

The The two prompt-injection hits stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The The two prompt-injection hits stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

A deliberately malicious test package returned zero risk

For the A deliberately malicious test stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Static analysis did not fail

For the Static analysis did not stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

What a useful static report should disclose

For the What a useful static stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the What a useful static stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Why runtime behavior still matters

When working through the Why runtime behavior still stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Limitations

When working through the Limitations stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Reproduce the audit

When working through the Reproduce the audit stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Reproduce the audit stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# self-bootstrapping — no install needed
git clone https://github.com/abdalhafeezbushara/driftward.git
cd driftward
python3 research/mcp-npm-audit-2026-09-01/audit.py \
  research/mcp-npm-audit-2026-09-01/results.json \
  /tmp/mcp-audit-rerun.json
N=300 ./detonate/fetch-corpus.sh /tmp/mcp-packages.txt
python3 research/mcp-npm-audit-2026-09-01/audit.py \
  /tmp/mcp-packages.txt /tmp/mcp-audit-current.json

A practical checklist for MCP users

The A practical checklist for stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

The practical conclusion

The The practical conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 0f0aff9c909a: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.