This article is published in English.
Why raising concurrency causes HTTP 429s: start rate and per-host limits
Raising worker count without shaping start rate and per-host caps creates bursts that trip 429s. Throughput comes from paced concurrency, backoff, and honest queue design.
TL;DR: Raising concurrency is sold as free parallelism: more workers, more requests, more data. Ten feels good, so twenty must be better, and fifty irresistible. The catch is that raising the pool does not only raise concurrency. It also decides how hard the client hits a host at the first instant, and—on multi-host workloads—what per-host concurrency emerges without anyone setting it. When those hidden dials fail, the status code is often HTTP 429, which looks exactly like “too many concurrent requests.” Teams then shrink the pool or bolt on proxies and move on. Without isolating the real cause, they leave useful throughput on the table.
Part I — Concurrency vs start rate: one pool setting controls two limits
A worker pool size (p-limit(n), a semaphore, a thread count) bounds how many operations may be in flight. It does not bound how quickly new work may begin. At t=0, a cold pool with fifty waiting tasks can start fifty requests in one tick—the start-rate spike—even though steady-state concurrency will later look “only fifty.” Rate limiters care about that opening burst as much as about sustained parallelism.
Measuring start-rate burst
A reproducible harness helps. Take a small set of Arch Linux Wiki article URLs:
https://wiki.archlinux.org/title/Arch_Linux
https://wiki.archlinux.org/title/Installation_guide
https://wiki.archlinux.org/title/Pacman
https://wiki.archlinux.org/title/Systemd
...and so on
Build a 100-request workload by cycling the list:
function buildWorkload(urls, n) {
return Array.from({ length: n }, (_, i) => urls[i % urls.length]);
}
async function runPooledFetch(urls, concurrencyLimit, agent, options = {}) {
const limit = pLimit(concurrencyLimit);
const results = await Promise.all(
urls.map((url) => limit(() => fetchOne(url, agent)))
);
// ...
}
Run with different pool sizes, classify each outcome as ok versus 429 (and other failures), and record completed requests per second alongside ok counts. The distinction matters: a fast 429 still contributes to “completed req/s” vanity metrics while delivering no page.
Useful throughput falls as global concurrency rises
With a fixed 100-request workload and only pool size changing, a direct connection behaves harshly. At concurrency 1, essentially all 100 succeed and completed req/s sits near 2.4/s—you wait on full renders without parallel starts. At concurrency 10, useful data collapses: about 42 pages succeed and 58% of the run returns 429, while completed req/s jumps to roughly 23.5/s. At concurrency 25 and 50, ok counts keep falling (about 24/100, then 16/100) while completed req/s remains in the high teens or twenties (~18.2/s and 21.8/s).
The vanity peak is concurrency 10 at 23.5 completed req/s—highest in the table—paired with one of the worst ok counts. Optimizing for completed requests per second would recommend the setting that burns most of the budget on bans. Useful throughput is ok-per-second, not raw completions.
Adding a start-rate cap at the same concurrency fixes it
Keep the pool size unchanged and gate how soon the next request may start. Derive a minimum gap from a max requests-per-second setting:
// minGapMs derived from --max-rps; pool concurrency is untouched
if (minGapMs > 0) {
const now = Date.now();
const waitMs = Math.max(0, nextStartAt - now);
if (waitMs > 0) await sleep(waitMs);
nextStartAt = Math.max(Date.now(), nextStartAt) + minGapMs;
}
With pacing, the same concurrency that previously flooded the host can finish nearly all pages because the opening burst disappears. The correct mental model is two dials: concurrency (in flight) and start rate (admissions per second). Collapsing them into one number hides the failure mode.
Why a pooled HTTP client sends a burst of requests at t=0
Pools do not know about polite pacing. They know about free slots. At startup every slot is free, so every queued task that can run does run. Connection reuse and HTTP/2 multiplexing can make that burst even denser on the wire. Only an admission controller—token bucket, min gap, leaky bucket—shapes starts independently of in-flight caps.
Do residential proxies fix start-rate burst?
Pacing is the correctness fix: gather data without spiking the target. Engineering reality sometimes demands speed that a 2.4 req/s throttle cannot meet. The same naive sweep—no start-rate cap, pool bursting at t=0—routed through residential proxies can hold roughly 98–99/100 successes with essentially no bans across concurrency levels, while the direct path still craters.
That does not mean proxies erase the need to understand start rate. They change IP reputation and how aggressively a site attributes traffic. They can mask a burst that would ban a single egress IP. If the product requirement is “collect without ever looking like a stampede,” pacing remains the primary control; proxies are a capacity and reputation layer, not a substitute for measuring admissions.
Proxy agent construction typically pulls credentials from the environment:
function buildProxyAgent() {
const user = process.env.BRIGHT_DATA_PROXY_RESIDENTIAL_USERNAME;
const pass = process.env.BRIGHT_DATA_PROXY_RESIDENTIAL_PASSWORD;
if (!user || !pass) return null;
const host = process.env.BRIGHT_DATA_PROXY_HOST || 'brd.superproxy.io';
const port = process.env.BRIGHT_DATA_PROXY_PORT || '33335';
const proxyUrl = `http://${encodeURIComponent(user)}:${encodeURIComponent(pass)}@${host}:${port}`;
return new HttpsProxyAgent(proxyUrl);
}
const residentialAgent = buildProxyAgent();
await runPooledFetch(tasks, 50, residentialAgent);
Use them deliberately, log whether a run was direct or proxied, and still record observed start RPS so you know which dial moved the outcome.
Part II — Per-host concurrency
Global pools also hide a second emergent limit. Consider:
const limit = pLimit(50);
await Promise.all(urls.map((url) => limit(() => fetchOne(url))));
Fifty global slots shared across many hosts do not mean each host sees at most fifty in flight. Queue order and response-time skew decide how many requests one hostname actually absorbs. A mixed list can look balanced on paper:
1. arch
2. github
3. arch
4. mdn
5. npm
6. arch
7. cloudflare
yet still deliver bursts to the strict host when its URLs cluster in the runnable set.
Benchmarking per-host concurrency
Reuse the harness with two hosts:
- A strict host — Arch Linux Wiki (10 article URLs) that blocks under pressure, as in Part I.
- A lenient host —
books.toscrape.comcatalogue pages (10 URLs) that rarely block, acting as a control. If the sandbox fails, the client is broken.
An alternating list yields 20 URLs × 5 repeats = 100 requests at global concurrency 50. Fixture lists and builders live in the open concurrency-trap-bench repository (for example urls-mixed-arch-books.txt).
https://wiki.archlinux.org/title/Arch_Linux
https://books.toscrape.com/catalogue/page-1.html
https://wiki.archlinux.org/title/Installation_guide
https://books.toscrape.com/catalogue/page-2.html
https://wiki.archlinux.org/title/Pacman
https://books.toscrape.com/catalogue/page-3.html
...and so on
Ordered workloads can round-robin, block by host, or shuffle with a seed:
function buildOrderedWorkload(urls, pattern, { repeatsPerUrl = 5, seed = null } = {}) {
const n = urls.length * repeatsPerUrl;
let list = Array.from({ length: n }, (_, i) => urls[i % urls.length]);
if (pattern === 'block') {
// AxN, BxN, and so on. Repeat each URL before advancing.
const out = [];
for (const url of urls) {
for (let i = 0; i < repeatsPerUrl; i++) out.push(url);
}
return out;
}
if (pattern === 'shuffle') {
// Fisher–Yates with a fixed seed so the run is reproducible
for (let i = list.length - 1; i > 0; i--) {
seed = (Math.imul(1664525, seed) + 1013904223) >>> 0;
const j = seed % (i + 1);
[list[i], list[j]] = [list[j], list[i]];
}
}
// 'round-robin' leaves the alternating list as-is
return list;
}
Measure peak in-flight per host from start/finish events:
function maxConcurrentPerHost(results) {
const eventsByHost = new Map();
for (const r of results) {
const host = new URL(r.url).hostname;
if (!eventsByHost.has(host)) eventsByHost.set(host, []);
const end = r.startedAt + r.ms;
eventsByHost.get(host).push({ t: r.startedAt, delta: 1 }, { t: end, delta: -1 });
}
// sort events by time, sweep: +1 on start, -1 on finish, track max
}
Global concurrency stays fixed; only ordering changes. That isolates per-host pressure from pool size.
How queue order changes per-host concurrency
Real crawlers never keep perfect round-robin forever. Sitemaps, dependency graphs, and retry queues reshuffle work. Identical global settings can therefore produce different per-host peaks.
1. Round-robin: alternating hosts evenly
The alternating list (A B A B …) spreads work. Peak in-flight on the strict host stays comparatively moderate because the other host keeps claiming slots.
2. Block: clustering same-host requests
Grouping all Arch URLs then all books URLs hands the strict host a long runnable streak. Peak in-flight on that host climbs toward the global pool size even though “concurrency is still 50.”
3. Shuffle (seed 42): randomized order
A seeded shuffle sits between the extremes and mirrors accidental production ordering. Peaks move with the seed; the lesson is sensitivity, not a magic permutation.
Same global concurrency, different per-host pressure
Across those patterns, ok rates on the strict host track its measured peak in-flight more closely than the constant global c=50. The lenient host stays healthy. Lowering the global pool to “fix” the strict host would also punish the lenient one and still would not pin the strict host’s peak under unlucky ordering.
Fix: limit per-host concurrency separately
Part I added a start-rate cap beside the pool. Part II adds a per-host cap beside the pool: a nested limit on how many in-flight requests any single hostname may hold.
Recall the three terms:
- Global concurrency — shared pool size (for example
p-limit(50)). You set it. - Per-host concurrency — what one host actually saw (
peak_inflight). You measure it. - Per-host cap — a separate maximum you choose for how many concurrent calls any single hostname may hold, nested under the global pool.
If that host-level brake is missing, every free global slot can be claimed by whichever URLs happen to be ready—including a stampede toward one strict origin. Adding the nested limit gives each hostname its own queue gate:
async function runPooledFetch(urls, concurrencyLimit, agent, { perHostLimit } = {}) {
const globalLimit = pLimit(concurrencyLimit);
const hostLimiters = new Map();
async function fetchWithLimits(url, execute) {
return globalLimit(async () => {
if (perHostLimit && perHostLimit >= 1) {
const host = new URL(url).hostname;
if (!hostLimiters.has(host)) hostLimiters.set(host, pLimit(perHostLimit));
return hostLimiters.get(host)(execute);
}
return execute();
});
}
// ...
}
Now a cluster of strict-host URLs cannot occupy every global slot at once. The lenient host can still use remaining capacity. Bans correlate with the capped variable you intended to control.
Why a shared worker pool concentrates load on one host
The pool fills slots with whatever is runnable. Fast hosts free slots quickly and pull more of their own work—until a cluster of strict-host URLs becomes runnable together and inherits a burst. Burst size depends on ordering, latency, and mix—none of which appear in the single concurrency integer. That is why a per-host limiter beats blindly shrinking the global pool.
A practical checklist
- Treat start rate as its own setting. A semaphore is not an RPS limiter. Add an explicit start-rate cap beside concurrency.
- Treat per-host concurrency as its own setting. Multi-host workloads need a per-host limiter inside the global pool.
- Log
observed_start_rpsand per-hostpeak_inflight. You cannot reason about limits you never measured. - Optimize useful throughput. Score ok-per-second, not completed-per-second; fast 429s are still failures.
Exact thresholds in any one blog run are not portable: limiters are stateful and depend on time, traffic history, and IP reputation. The portable pattern is isolation. A failure that looks like “too much concurrency” may be a start-rate problem, a per-host scheduling problem, or both. Until those variables are separated, shrinking the pool only treats a symptom—and often the wrong one.
Reading the metrics without fooling yourself
Completed requests per second rises whenever the client can open sockets quickly—including when most responses are refusals. Dashboards that celebrate concurrency sweeps without an ok filter will recommend hostile settings. Pair every throughput chart with an ok ratio and, when possible, with bytes of useful content retained. If your pipeline retries 429s, count retries separately so a storm of retries does not look like productive parallelism.
Start-rate logs should capture the timestamp of each admission, not only of each completion. From admissions you can reconstruct bursts in the first 100–500 ms, which is where many pooled clients look identical regardless of the steady-state pool size you configured. If two configurations share the same opening burst, do not be surprised when they share the same ban pattern.
Proxies, reputation, and honesty about trade-offs
Residential or datacenter proxies redistribute identity. They can convert a single-IP stampede into many quieter streams. That helps collection projects and can hide poor admission control from a destination that keys limits on IP. It does not remove ethical and contractual duties toward the sites you fetch, nor does it remove the engineering need to understand your own client. Prefer pacing first when you control the client; use proxies when the product truly requires higher aggregate throughput across many identities—and keep measuring per-identity and per-host pressure so you are not flying blind behind the proxy layer.
Bringing the two parts together
Production crawlers usually need both dials: a global pool for resource safety on your side, a start-rate cap for politeness at admission time, and per-host caps when multiple destinations share the pool. Skipping any one of them recreates a failure mode that “looks like concurrency” in HTTP status codes. The repair is not mystical; it is instrumentation plus a second limiter aimed at the variable you actually observed.
Harness design notes that keep comparisons fair
Keep success taxonomy identical across sweeps: ok HTML, HTTP 429, other 4xx/5xx, timeouts, and parse failures should be labeled the same way every run. Change only the variable under test—pool size, min start gap, proxy on/off, or queue order. Warm DNS and TLS where possible so the first points of a curve are not dominated by cold-handshake noise unless cold start is explicitly part of the story.
Repeat workloads enough times to damp noise, but not so many that a site’s adaptive limiter permanently shifts mid-experiment. When limiters are stateful, note the time of day and whether prior bans might still be cooling down. Publish seed values for shuffles so others can reproduce ordering effects.
URL fixtures should be stable. Wiki titles and sandbox catalogue pages that disappear mid-study corrupt ok counts. Pin lists in version control beside the harness, as the concurrency-trap-bench project does, so charts refer to known inputs.
What “useful throughput” looks like in a pipeline
Downstream systems care about accepted documents, not socket churn. If a scraper feeds an indexer, count indexed documents per minute. If it feeds a price database, count validated rows. Align the optimization target with that business unit. Otherwise engineering will maximize a proxy metric—completed HTTP exchanges—that includes mountains of 429 bodies.
Retries interact badly with unpaced pools. A burst that earns bans, followed by immediate retries, can amplify start rate further. Back off on 429 with jitter, honor Retry-After when present, and never let retry storms bypass the start-rate gate. The admission controller must see retries as new starts.
Multi-tenant and multi-host schedulers
Services that crawl on behalf of many customers often already have global concurrency ceilings for process safety. They still need per-destination budgets so one customer’s URL list cannot monopolize a fragile origin. Nested limits—global, per-tenant, per-host—compose. Implement them as explicit layers rather than hoping fair queuing emerges from a single semaphore.
When hosts differ by orders of magnitude in latency, work-stealing pools bias toward the fast host. That bias is good for utilization and bad for the slow, strict host when its URLs finally become runnable in a clump. Per-host caps bound the damage; optional per-host weights can further express politeness policies.
Interpreting proxy results without magical thinking
If direct egress fails and proxied egress succeeds under identical pool settings, the destination is likely keying enforcement on network identity. That is useful operational knowledge. It is not proof that start-rate physics changed. Behind proxies you should still log admissions per exit identity and per target host. Otherwise you merely relocate the blind spot.
Compliance and robots policies remain yours to respect. Higher aggregate throughput through many exits increases the blast radius of a logic bug. Feature-flag pacing and per-host caps even when proxies are enabled so you can dial aggression down without redeploying identity infrastructure.
Closing the loop from experiment to defaults
Once measurements show that start-rate and per-host peaks predict bans better than global pool size alone, encode those findings as defaults in the client library: require an RPS (or min-gap) parameter, require a per-host ceiling for multi-host modes, and export metrics for both. Documentation should show the bad dashboard (completed req/s rising while ok falls) beside the good one. Teaching the failure mode prevents the next team from “fixing” concurrency into a quiet outage of useful data.
Reproducing the start-rate experiment cleanly
Pin Node and undici (or your HTTP stack) versions so connection pooling behavior stays comparable. Disable unrelated browser-like retry middleware during sweeps. Clear DNS caches between direct and proxied runs if your harness resolves hosts differently. Record wall-clock for the whole batch plus per-request start and end timestamps so you can chart admissions in the first half-second—the window where pooled clients look identical regardless of the steady-state concurrency you believe you set.
When plotting, always show ok-count beside completed req/s. A dual-axis chart that hides ok-count is how vanity metrics win arguments. Export CSV from the harness so others can recompute without trusting screenshots.
Interpreting Arch Wiki style strictness
Public documentation hosts vary: some throttle by IP and path, some by User-Agent, some by concurrent connections, some by request rate over sliding windows. A 429 today may become a quieter slowdown tomorrow as operator policy shifts. That is why the article’s numbers are illustrative patterns, not eternal constants. What transfers is the methodology: separate pool size, admission rate, and per-host peaks, then change one variable at a time.
If you substitute your own strict host, keep a lenient control host in the same run. When the control fails, your client is broken. When only the strict host fails, you are studying their policy interacting with your schedule.
Start-rate cap implementation details
A minimum gap derived from max RPS is simple and effective for single-process clients. Token buckets allow short bursts while enforcing long-run averages—useful when you want snappy interactive fetches but polite bulk crawls. Leaky buckets smooth harder. Whatever algorithm you pick, apply it to starts, including retries. A retry storm that bypasses the gate recreates the t=0 stampede after the first ban.
Multi-process crawlers need a distributed admission lock (Redis, etcd, or a central scheduler). Local gaps alone will not coordinate five workers each believing they may start two requests per second.
Per-host limit implementation details
Nested semaphores keyed by hostname are the usual pattern: acquire global, then acquire host, then fetch, then release in reverse. Decide whether www and apex share a key. Decide how redirects that change host count against caps. Path-based shards of the same site usually still share one host budget unless you have explicit CDN evidence otherwise.
Emit metrics: inflight_global, inflight_per_host{host}, admissions_per_second, http_429_total{host}. Alert when 429 rates rise rather than only when error budgets burn on 5xx.
Queue ordering in production schedulers
Sitemap order, BFS from a seed, priority queues for “important” URLs, and retry queues all reshape per-host peaks. A retry queue that prepends failed Arch URLs can accidentally recreate block ordering after a partial outage. Fair queuing across hosts—round-robin ready queues per host—reduces accidental concentration even before hard caps. Hard caps remain necessary when fairness alone cannot bound peaks under uneven latency.
Proxies without self-deception
Residential networks change identity distribution. They do not repeal physics: if each identity still bursts fifty starts at once, destinations that key on behavior rather than IP may still ban. Log admissions per exit identity. Rotate politely. Honor robots and contractual terms. Prefer pacing even when proxies are enabled so a bug cannot amplify into a wide blast radius.
Datacenter proxies are cheaper and easier to fingerprint; residential are costlier and ethically sensitive. Choose deliberately; do not treat “proxy” as a synonym for “fix.”
From experiment to library defaults
Bake two required knobs into shared HTTP client wrappers used by crawlers: maxInFlight and maxStartsPerSecond, plus maxInFlightPerHost when more than one host appears. Refuse to construct a client for multi-host mode without the per-host cap. Provide dashboard templates that chart ok-per-second. Teach onboarding with the dual chart: concurrency 10 winning vanity throughput while losing useful pages.
Checklist expanded
- Cap starts, not only in-flight slots.
- Cap per host inside the global pool.
- Measure admissions and per-host peaks every run.
- Score success by useful pages, not by socket completions.
- Apply admission control to retries.
- Keep a lenient control host in mixed tests.
- Treat proxy success as reputation shift, not proof that scheduling is healthy.
- Revisit numbers as destinations change policy; keep the methodology.
Until those habits stick, teams will keep “fixing concurrency” into quieter failure modes that still return HTTP 429 and still burn crawl budgets.