This article is published in English.
Production Scraping Behind Anti-Bot Defenses: Operable Patterns
Rate limits, identity hygiene, and resilient fetch paths for apps that must collect public data responsibly.
This walkthrough rebuilds an operable path for: How to Bypass Anti-Bot Walls for Production-Ready Apps. Focus on contracts, checks, and code you can drop into a repo without guessing intent. For Overview, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
you — The Illusion of the Clean Browser
For you — The Illusion of the Clean Browser, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility. Prefer boring reliability over clever one-off demos.
II — The Anatomy of an Immediate Rejection
For II — The Anatomy of an Immediate Rejection, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer boring reliability over clever one-off demos.
III — The “Fake” 200 Response Trap
For III — The “Fake” 200 Response Trap, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments. Prefer boring reliability over clever one-off demos. For III — The “Fake” 200 Response Trap, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
// The Lazy Scraping Pattern (Guaranteed to fail at scale)
const axios = require('axios');
async function naiveScrape() {
const targetUrl = 'https://www.amazon.com/s?k=ddr5+ram+128gb';
try {
const response = await axios.get(targetUrl, {
headers: { 'User-Agent': 'Mozilla/5.0…' }
});
console.log("Success! Status:", response.status, "Data Length:",
response.data.length);
} catch (error) {
console.error("Blocked!");
}
}
naiveScrape();
console.log("Does it contain Amazon's soft-block tag?", response.data.includes("a-no-js"));
console.log("First 300 characters of the page code:\n", response.data.substring(0, 300));
const axios = require('axios');
async function naiveScrape() {
// Switching to a Cloudflare-protected target for a clean firewall error
const targetUrl = 'https://www.cnn.com/';
try {
console.log("Initiating naive data extraction request…");
const response = await axios.get(targetUrl, {
headers: {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64)
AppleWebKit/537.36'
}
});
console.log("Success! Data Length:", response.data.length);
} catch (error) {
console.error("\n❌ Request Crushed by Anti-Bot Wall.");
if (error.response) {
// This will capture and print the hard 403 error code
console.error(`Status Code: ${error.response.status}`);
console.error(`Status Text: ${error.response.statusText}`);
console.error("Reason: Cloudflare WAF rejected the vanilla Axios
handshake.");
} else {
console.error("Error Message:", error.message);
}
}
}
naiveScrape();
IV — Moving to an Automated Gateway
For IV — Moving to an Automated Gateway, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility. Add a smoke test for the critical path in CI with fixtures when budgets allow.
Step 1: Account Activation and Token Generation
For Step 1: Account Activation and Token Generation, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Add a smoke test for the critical path in CI with fixtures when budgets allow.
Step 2: Testing Inside the Native Playground
For Step 2: Testing Inside the Native Playground, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments. Add a smoke test for the critical path in CI with fixtures when budgets allow. For Step 2: Testing Inside the Native Playground, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
V — Building LaunchRadar Agent
For V — Building LaunchRadar Agent, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility. Checkpoint after expensive steps so retries do not re-bill the same model call.
Setup LaunchRadar Project from Scratch
For Setup LaunchRadar Project from Scratch, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Write a short runbook: rotate keys, drain queues, roll back the last change.
# 1. Move directly to your local User Home root path to avoid file syncing errors
cd ~
# 2. Provision a dedicated application folder
mkdir LaunchRadar
# 3. Enter your newly initialized project workspace
cd LaunchRadar
# 4. Initialize an isolated local package architecture
npm init -y
# 5. Tell the Node.js runtime to process modern ES Module imports natively
npm pkg set type="module"
npm install @brightdata/sdk
/**
* Project: LaunchRadar
* Core Stack: Node.js + Official Bright Data JavaScript SDK
*/
import { bdclient } from '@brightdata/sdk';
const BRIGHTDATA_KEY = process.env.BRIGHTDATA_API_KEY || '';
async function deployLaunchRadar() {
// Targeting Product Hunt's active daily launch asset path
const targetUrl = 'https://www.producthunt.com/';
console.log("🤖 [LaunchRadar]: Initializing autonomous target ingestion loop...");
try {
// Initialize the client gateway via the native constructor
const client = new bdclient({
apiKey: BRIGHTDATA_KEY
});
console.log("🤖 [LaunchRadar]: Tunneling through Web Unlocker cluster...");
// Execute the single-step browser-emulated extraction natively
const htmlResponse = await client.scrapeUrl(targetUrl, {
zone: 'web_unlocker1',
format: 'raw' // Streams the target webpage as a raw HTML text string
});
console.log("\n=================== PIPELINE SUCCESS ===================");
console.log(`🤖 [LaunchRadar]: Ingestion complete. Streamed: ${htmlResponse.length} bytes.`);
console.log("========================================================\n");
// This will print true-proving your paid credit bypassed the security loop
const dataIsValid = htmlResponse.includes("Product Hunt");
console.log(`🤖 [LaunchRadar]: Target Data Authenticated: ${dataIsValid}`);
// Gracefully disconnect active execution sockets
await client.close();
} catch (error) {
console.error("❌ [LaunchRadar Fatal]: Pipeline exception encountered.");
console.error(`Reason: ${error.message}`);
}
}
deployLaunchRadar();
VI — Building Our Final Vanguard Agent Application
For VI — Building Our Final Vanguard Agent Application, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments. Checkpoint after expensive steps so retries do not re-bill the same model call. For VI — Building Our Final Vanguard Agent Application, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
npm install cheerio
const htmlResponse = await client.scrapeUrl(targetUrl, {
zone: "web_unlocker1",
format: "raw",
country: "us",
dataFormat: "html"
});
/**
* Project: Vanguard Agent
* Function: Autonomous CNN Trend & Headline Generator
* Core Stack: Node.js (ESM) + Official Bright Data JavaScript SDK + Cheerio
*
* IMPORTANT: CNN's own robots.txt disallows /search (Disallow: /search),
* so Bright Data's Web Unlocker refuses it in immediate-access mode —
* that's not a bug, it's the site opting out of that specific path for
* compliant bots. Instead, this hits CNN's own News Sitemap
* (https://www.cnn.com/sitemap/news.xml), which robots.txt explicitly
* lists as a Sitemap: entry — meaning it's meant to be fetched by bots.
* It contains real headline titles and real publish timestamps for the
* last ~48 hours of coverage across every section, which is arguably a
* better fit for "trend monitoring" than a search page anyway.
*/
import fs from 'fs';
import readline from 'readline';
import * as cheerio from 'cheerio';
import { bdclient } from '@brightdata/sdk';
const REPORT_FILE = 'vanguard_trend_report.md';
const CNN_NEWS_SITEMAP = 'https://www.cnn.com/sitemap/news.xml';
// Key comes ONLY from the environment - no hardcoded fallback, ever.
// If a real key has been pasted into a file, a chat, or a log at any
// point, treat it as compromised and rotate it in your Bright Data
// dashboard immediately, even after removing it from the code.
const BRIGHTDATA_KEY = process.env.BRIGHTDATA_API_KEY || '';
if (!BRIGHTDATA_KEY) {
console.error(
"\n❌ [Vanguard Config Error] BRIGHTDATA_API_KEY is not set.\n" +
" Run it like this instead:\n\n" +
" PowerShell: $env:BRIGHTDATA_API_KEY=\"your_actual_key\"; node vanguard.mjs\n" +
" cmd.exe: set BRIGHTDATA_API_KEY=your_actual_key && node vanguard.mjs\n" +
" macOS/Linux: BRIGHTDATA_API_KEY=your_actual_key node vanguard.mjs\n"
);
process.exit(1);
}
// CNN article URLs follow /YYYY/MM/DD/category/slug - used here purely to
// derive a topic category for the breakdown, not for date (the sitemap
// gives us a real publish timestamp directly).
const CNN_ARTICLE_PATTERN = /^\/(\d{4})\/(\d{2})\/(\d{2})\/([a-z0-9-]+)\//i;
const printDashboardHeader = () => {
console.clear();
console.log("\x1b[38;5;39m%s\x1b[0m", " ┌────────────────────────────────────────────────────────┐");
console.log("\x1b[38;5;39m%s\x1b[0m", " │ VANGUARD INTELLIGENCE SERVICE // CNN TREND ENGINE │");
console.log("\x1b[38;5;39m%s\x1b[0m", " └────────────────────────────────────────────────────────┘");
};
const promptUser = (queryMessage) => {
const interfaceInstance = readline.createInterface({
input: process.stdin,
output: process.stdout
});
return new Promise((resolve) => {
interfaceInstance.question(queryMessage, (inputAnswer) => {
interfaceInstance.close();
resolve(inputAnswer.trim());
});
});
};
function escapeRegExp(string) {
return string.replace(/[.*+?^${}()|[\]\\]/g, '\\amp;');
}
function categoryFromUrl(loc) {
try {
const path = new URL(loc).pathname;
const dated = path.match(CNN_ARTICLE_PATTERN);
if (dated) return dated[4].replace(/-/g, ' ');
// Some CNN URLs skip the date prefix entirely (e.g. video pages
// like /politics/video/slug). Fall back to the first path segment
// instead of dumping everything into "uncategorized".
const segment = path.split('/').filter(Boolean)[0];
return segment ? segment.replace(/-/g, ' ') : 'uncategorized';
} catch {
return 'uncategorized';
}
}
async function deployVanguardService() {
printDashboardHeader();
console.log("\x1b[38;5;244m%s\x1b[0m", "\n [CONFIGURE INTELLIGENCE FILTER]");
const searchQuery = await promptUser('🔍 Enter a keyword to filter recent CNN headlines (e.g., startup, election, climate): ');
if (!searchQuery) {
console.error("\n❌ [Vanguard Error] Input token context empty. Pipeline aborted.");
return;
}
printDashboardHeader();
console.log(`\n\x1b[38;5;214m🤖 [Vanguard]: Connecting to proxy fabric for target: ${CNN_NEWS_SITEMAP}\x1b[0m`);
console.log("🤖 [Vanguard]: Initializing automated browser emulation...");
const client = new bdclient({ apiKey: BRIGHTDATA_KEY });
try {
// dataFormat: 'html' returns the raw response body as-is, which
// works fine for XML too - we just parse it with cheerio's XML
// mode below instead of treating it as an HTML page.
const xmlResponse = await client.scrapeUrl(CNN_NEWS_SITEMAP, {
zone: 'web_unlocker1',
dataFormat: 'html',
country: 'us'
});
console.log(`\x1b[38;5;46m🤖 [Vanguard]: Data stream ingestion complete. Length: ${xmlResponse.length} bytes.\x1b[0m`); if (!xmlResponse || xmlResponse.length < 500) { const preview = xmlResponse ? xmlResponse.slice(0, 300) : '(empty response)'; throw new Error( `Target returned only ${xmlResponse ? xmlResponse.length : 0} bytes - this is too small to be ` + `a real sitemap. Raw response below:\n\n${preview}\n\n` + `Checklist:\n` + ` 1. Confirm BRIGHTDATA_API_KEY is a valid, active key.\n` + ` 2. Confirm 'web_unlocker1' matches an active Web Unlocker zone name in your ` + `Bright Data dashboard - zone names are account-specific.\n` + ` 3. If the text above mentions a block/captcha, retry in a moment or try a ` + `different 'country' value.` ); } console.log("🤖 [Vanguard]: Processing resilient extraction analytics..."); // xmlMode: true keeps namespaced tags (e.g. news:title) intact // instead of cheerio's default HTML-parsing/lowercasing behavior.
const $ = cheerio.load(xmlResponse, { xmlMode: true });
const insights = [];
$('url').each((_, el) => {
const loc = $(el).find('loc').first().text().trim();
const title = $(el).find('news\\:title').first().text().trim();
const pubDateRaw = $(el).find('news\\:publication_date').first().text().trim()
|| $(el).find('lastmod').first().text().trim();
if (!loc || !title) return; // skip entries without a real headline (e.g. galleries/video-only)
insights.push({
title: title.replace(/\s+/g, ' '),
date: pubDateRaw ? pubDateRaw.slice(0, 10) : 'unknown',
category: categoryFromUrl(loc),
postUrl: loc
});
});
const uniqueRawInsights = Array.from(new Map(insights.map(item => [item.postUrl, item])).values());
// Sort newest first using the sitemap's own publish timestamp. uniqueRawInsights.sort((a, b) => (a.date < b.date ? 1 : a.date > b.date ? -1 : 0)); // Whole-word match, not substring - plain .includes() would match // "war" inside "Warner" or "warrant", which are false positives // for a search on "war". \b enforces a real word boundary. const keywordPattern = new RegExp(`\\b${escapeRegExp(searchQuery)}\\b`, 'i'); let filteredInsights = uniqueRawInsights.filter(item => keywordPattern.test(item.title)); // Phase 2: fallback to the newest overall headlines if no direct match let isFallbackMode = false; if (filteredInsights.length === 0) { filteredInsights = uniqueRawInsights; isFallbackMode = true; } const uniqueInsights = filteredInsights.slice(0, 10); if (uniqueInsights.length === 0) { throw new Error('Zero real headline entries isolated inside the sitemap buffer.'); } const categoryCounts = uniqueRawInsights.reduce((acc, item) => { acc[item.category] = (acc[item.category] || 0) + 1;
return acc;
}, {});
const topCategories = Object.entries(categoryCounts)
.sort((a, b) => b[1] - a[1])
.slice(0, 5);
console.log("\n\x1b[38;5;51m┌────────────────────────────────────────────────────────────────────────┐");
console.log(`│ ► VANGUARD INTELLIGENCE STREAM // CNN // '${searchQuery}'`.padEnd(76, ' ') + '│');
console.log("└────────────────────────────────────────────────────────────────────────┘\x1b[0m");
if (isFallbackMode) {
console.log(`\x1b[33m ⚠️ [Notice] No direct matches for '${searchQuery}'. Showing newest CNN headlines instead.\x1b[0m\n`);
}
uniqueInsights.forEach((item, index) => {
console.log(` \x1b[38;5;46m[${String(index + 1).padStart(2, '0')}]\x1b[0m \x1b[1m${item.title.substring(0, 75)}\x1b[0m`);
console.log(` ├─ Category : ${item.category}`);
console.log(` ├─ Date : ${item.date}`);
console.log(` └─ Link : \x1b[38;5;39m\x1b[4m${item.postUrl}\x1b[0m\n`);
});
console.log("\x1b[38;5;51m└────────────────────────────────────────────────────────────────────────┘\x1b[0m");
console.log('\x1b[38;5;244m📊 Top categories in this result set:\x1b[0m');
topCategories.forEach(([cat, count]) => console.log(` • ${cat} (${count})`));
const markdownReport = [
`# Vanguard CNN Intelligence Digest`,
`* **Search Keyword:** \`${searchQuery}\``,
`* **Source:** CNN News Sitemap (robots.txt-approved)`,
`* **Data Capture Status:** ${isFallbackMode ? 'Newest-headlines fallback' : 'Direct keyword match'}`,
`* **Timestamp:** ${new Date().toISOString()}`,
`\n## Top Categories`,
...topCategories.map(([cat, count]) => `* **${cat}** - ${count} article(s)`),
`\n## Headlines`,
...uniqueInsights.map((p, idx) =>
`### ${idx + 1}. ${p.title}\n* **Date:** ${p.date}\n* **Category:** ${p.category}\n* **Link:** [${p.postUrl}](${p.postUrl})`
)
].join('\n');
fs.writeFileSync(REPORT_FILE, markdownReport, 'utf8');
console.log(`\x1b[38;5;46m💾 [Vanguard] System Report written successfully -> ${REPORT_FILE}\x1b[0m\n`);
} catch (pipelineException) {
console.error("\n❌ [Vanguard Fatal]: Operational routine aborted.");
console.error(`Reason: ${pipelineException.message}`);
} finally {
await client.close();
}
}
deployVanguardService();
VII — Moving Past The Anti-Bot Maintenance
For VII — Moving Past The Anti-Bot Maintenance, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility. Prefer boring reliability over clever one-off demos.
Frequently Asked Questions
For Frequently Asked Questions, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer boring reliability over clever one-off demos.
Can Web Unlocker scrape JavaScript-heavy websites?
For Can Web Unlocker scrape JavaScript-heavy websites?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments. Pin runtime versions and record the digest that ran the demo. For Can Web Unlocker scrape JavaScript-heavy websites?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
Do you still need to rotate proxies manually?
For Do you still need to rotate proxies manually?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility. Add a smoke test for the critical path in CI with fixtures when budgets allow.
Can you use this architecture for AI agents?
For Can you use this architecture for AI agents?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Can you extend Vanguard to other websites?
For Can you extend Vanguard to other websites?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments. Add a smoke test for the critical path in CI with fixtures when budgets allow. For Can you extend Vanguard to other websites?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
Operational checklist
For Operational checklist, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit.
Add a smoke test for the critical path in CI with fixtures when budgets allow.
Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments.
Prefer boring reliability over clever one-off demos.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation.
For hardening note 0, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility.
Add a smoke test for the critical path in CI with fixtures when budgets allow.
For hardening note 1, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments.
Write a short runbook: rotate keys, drain queues, roll back the last change.
For hardening note 2, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
Prefer boring reliability over clever one-off demos.
For hardening note 3, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Add a smoke test for the critical path in CI with fixtures when budgets allow.
For hardening note 4, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit.
Write a short runbook: rotate keys, drain queues, roll back the last change.
For hardening note 5, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility.
Prefer boring reliability over clever one-off demos.
For hardening note 6, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Record timings and cost next to functional results. Visibility early prevents surprise bills in shared environments.
Add a smoke test for the critical path in CI with fixtures when budgets allow.
For hardening note 7, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
Write a short runbook: rotate keys, drain queues, roll back the last change.
For hardening note 8, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Prefer boring reliability over clever one-off demos.
For hardening note 0, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Document the happy path and the recovery path together. Retries and dead-letter handling are part of the product.
Prefer boring reliability over clever one-off demos.
For hardening note 1, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Add a smoke test for the critical path in CI with fixtures when budgets allow.