Skip to content

Latest commit

 

History

History
155 lines (121 loc) · 8.09 KB

File metadata and controls

155 lines (121 loc) · 8.09 KB

Workstream A — Dry run findings

Measured 2026-08-18, 10:09–10:22 UTC. 5 targets × 4 arms × 10 rounds = 200 requests, interleaved by round. Benchmark host: DigitalOcean droplet botproxy-bench, nyc3, the benchmark host.

Exit addresses seen by targets are pseudonymised in the published data (benchmark-host, dc-exit-N, res-exit-N) — see the note in README.md. The residential pool presented a different address on each run.

Raw data: results/dryrun/raw.csv, summary.json, tiers.json, excluded.json. Data integrity verified: 200 rows, all 10 rounds present, no duplicate (target, arm, round) keys.

Purpose

To validate that the metrics work and that the arms are genuinely isolated — not to produce publishable target results. It succeeded at that, and produced four findings that change the full run.

Results

Target direct dc dc+anti-detect residential Minimum viable tier
uk_companies_house 100% 100% 100% 100% none (no proxy needed)
uk_fsa_hygiene 100% 100% 100% 100% none (no proxy needed)
us_az_real_estate 100% 100% 100% 80% none (no proxy needed)
us_sec_edgar 100% 100% 100% 90% none (no proxy needed)
us_nc_real_estate 0% 0% 0% 0% none sufficient

Median latency (ms), successful requests:

Target direct dc dc+anti-detect residential
us_sec_edgar 13 15 116 1,682
us_az_real_estate 19 27 168 1,832
uk_fsa_hygiene 168 169 440 1,882
uk_companies_house 535 584 584 2,367

Excluded by the compliance gate: us_oh_sos — ohiosos.gov answered 403 to both the benchmark User-Agent and a browser User-Agent, from the US datacenter host, so its robots.txt could not be read and it was not benchmarked.

Finding 1 — the dry run measured the wrong URLs, and this is the main lesson

Four of five targets returned "no proxy needed". That result is true and nearly meaningless, because these URLs are mostly homepages and landing pages: static, CDN-cached, and unprotected. Registry protection sits on the search endpoint and the record-detail page — which is exactly what a customer fetches and exactly what was not tested.

Before the full run every URL must point at a real data page — a company record, a licence lookup result, a parcel detail page. Where the search requires POST, the harness needs a request-body field added. Run as it stands, the full benchmark would measure CDN cache hits across 36 targets and conclude the product is unnecessary. This is recorded in targets-full.draft.json under _deep_urls.

Finding 2 — residential measured worse, but the tier is one day old

Read this one with its caveat attached. Residential routing on the rotating proxy shipped 2026-08-17, the day before this run: ClickHouse holds 65 residential rotating-proxy rows in total, all from 2026-08-17 onward, and some of them are this dry run. Exit-pool selection, region routing and retry behaviour are not at steady state one day in, so what follows measures a brand-new integration, not a mature tier. It is a baseline to improve against, not a verdict.

With that said, residential was the worst arm on every target that worked:

  • 10–100× the latency (1.7–2.4 s median against 13–584 ms).
  • Two read timeouts against azre.gov and one 502 against SEC EDGAR — the only transport failures in the entire run, all on the residential arm.
  • On ncrec.gov it was the only arm served a reCAPTCHA; the datacenter arms got a different (also failing) response.

The directional result is plausible and worth testing properly: the intuition that residential is a strict upgrade is often wrong for US government targets, because clean datacenter ranges can be scored better than residential pools, which are heavily abused. Residential also costs 6× more per GB.

But this run cannot support that as a published claim. One pool, one country selector (RS-US), one short window, on a tier that went live the previous day. Re-measure after the residential path has had time to settle, and only then decide what the report says about it. Publishing "residential is worse" off a day-old integration would be both unfair to the product and easy for a reader to dismiss once it improves.

Finding 3 — ncrec.gov blocks everything, in two different ways

North Carolina Real Estate Commission returned HTTP 202 with a zero-length body to all three datacenter arms (direct included) on all 10 rounds, and a reCAPTCHA page to residential on all 10. No tier succeeded.

Two consequences:

  1. A status-code-only benchmark would have scored the 202s as non-errors and the reCAPTCHA 200s as successes. The content assertion is what caught both. The harness now classifies any 2xx with an empty body as soft_block (it was previously filed as other_status, which reads like a curiosity rather than the block it is).
  2. ncrec.gov carried 26,932 production requests over 254 days, so a real customer is retrieving data from it — presumably via a deeper URL or a POST search rather than the homepage. Another argument for Finding 1.

Finding 4 — geography gates access before anti-bot does

The same compliance check gave different answers from different hosts:

Target from EU workstation from US datacenter host
azre.gov 403 to both UAs robots.txt served normally
ncrec.gov 403 to both UAs robots.txt served normally
ohiosos.gov 403 to both UAs 403 to both UAs

Two of three US state sites refused a European residential IP outright and served a US datacenter IP without complaint. Exit geography, not fingerprint, was the binding constraint — which is a genuinely useful result for the "which tier do I need" page, and a reminder that the compliance gate must be run from the same host as the benchmark.

Harness defects found and fixed

  • RobotFileParser.read() fetches with Python's default User-Agent. Targets that 403 that UA made it set disallow_all, so the harness silently skipped SEC EDGAR and azre.gov — targets the compliance gate had already cleared. The skip message read like a policy decision rather than the fingerprint block it was. Both scripts now share compliance.fetch_robots().
  • The compliance gate aborted the whole run on any ungated target. It now excludes them, reports them, and writes excluded.json — §2 requires exclusions to be published, not to stop the study.
  • 2xx with an empty body was classified as other_status. Now soft_block.
  • Crawl-delay from robots.txt is now honoured per target, on top of the politeness floor.

Cost and pacing for the full run

200 requests took 12.1 minutes at a 2 s floor plus jitter — 3.64 s per request wall-clock, dominated by the residential arm. The full run at 36 targets × 4 arms × 10 rounds = 1,440 requests implies ~87 minutes, which is comfortable and stays far below any target's rate tolerance.

Traffic was 9.36 MB total (direct 2.23, dc_plain 2.23, dc_antidetect 2.24, residential 2.65 MB). Scaling to the full run gives roughly 19 MB per arm — so residential consumption lands near 20 MB against the plan's 1 GB included allowance. The estimated $5 of residential spend will not be approached; actual cost is a few cents. Deep data pages will weigh more than these landing pages, but not by three orders of magnitude.

Open item

bench/config.json credentials bench-off / bench-on were created on account id 1 and pinned to us-ny. The benchmark droplet is still running (DigitalOcean id 593228013) so the full run can start without re-provisioning. Destroy it when the study is finished.