Skip to content

--speedtest-url makes the run nearly stop committing results (0/338707 for 35 min; ~130x slower on a small list) #33

Description

@alberttt46-spec

Summary

Enabling --speedtest-url makes the check pipeline almost never commit a result: throughput drops from ~1.5 items/s to ~0.011 items/s (~130×) on a small list, and on a large run it sits at Checked 0/338707 for 35 minutes with zero observations written. Removing --speedtest-url and changing nothing else makes the same database check at 6–11 items/s.

Environment

  • Proxy Workbench 3.0.3, installed from the release wheel proxy_workbench-3.0.3-py3-none-any.whl (sha256 591c14592521b47d86800c035eca663201d3e69857e24881b8947c84c0eb9146, matches the published SHA256SUMS), into a clean venv.
  • Windows 10 22H2 (19045), Python 3.11.3, 16 cores.
  • Network note (possibly relevant): this uplink is behind a TUN that accepts every outbound TCP connect in ~5 ms and then blackholes the traffic. Proof: TCP-connect to 80 random candidates from the collected database: 80/80 "open", median 4.6 ms. So "connected" here means nothing and anything without a hard timeout will hang instead of failing fast. This makes the per-stage budgets matter a lot — but the contrast below is on the same machine and network, with the speedtest flag as the only variable.

Steps to reproduce

Collect first (this part is fine, 106 feeds):

proxy-workbench collect --data <DIR> --source-timeout 20
# -> 106 sources: 103 ok, 2 TimeoutError, 1 RemoteProtocolError
# -> 694500 raw rows -> 338707 unique candidates

A. Large run with a speedtest — stalls at zero

proxy-workbench scan --data <BIG_DIR> --url https://example.com/ --attempts 1 \
  --timeout 8 --connect-timeout 3 --workers 256 --max-requests 20000 \
  --speedtest-url "https://speed.cloudflare.com/__down?bytes=2000000" \
  --speedtest-bytes 2000000

Observed (twice, two independent jobs, one on each of two consecutive nights of data):

job wall time result line observations results job_item moved
job-3799c320d82e8a84 2090 s (then terminated) Checked 0/338707; matching 0; saved 0 0 0 pending 338707
job-38251cf1c365efaa ≥120 s Checked 0/338707; ... 0 0 pending 338707

During the 2090 s run the worker held 34 established sockets and had accumulated ~22 s of user CPU, i.e. it was alive and connected but never committed a single item.

B. Same database, same flags, speedtest removed — works

proxy-workbench scan --data <SAME_BIG_DIR> --url https://example.com/ --attempts 1 \
  --timeout 8 --connect-timeout 3 --workers 256 --max-requests 20000 --deadline 120

Observed: Checked 186/338707; 7.1 proxies/s after 50 s, 130 observations / 130 results written. A separate 5 000-candidate run behaved the same way (479 done in 45 s, ~11/s, worker RSS 59 MB).

C. Small list, no prefilter, speedtest on — reproduces in miniature

proxy-workbench run --no-sources --input 30-proxies.txt --data <SMALL_DIR> \
  --url https://example.com/ --attempts 1 --timeout 8 --workers 30 \
  --prefilter 0 --deadline 90 \
  --speedtest-url "https://speed.cloudflare.com/__down?bytes=2000000" \
  --speedtest-bytes 2000000

Observed: Checked 1/30; matching 1; saved 1 — 1 item committed in 90 s, the other 29 never did, although 30 workers were idle-ish and --timeout was 8 s. The single committed observation took 16.2 s (started_at → finished_at), and its stored speed payload is:

{"state": "error", "mbps": null, "bytes": 0, "chunks": 0, "ttfb_ms": null,
 "transfer_ms": null, "total_ms": null, "code": "WHOLE_PROBE_TIMEOUT",
 "detail": "весь срок 16 с исчерпан", "connection": "cold", "error": "WHOLE_PROBE_TIMEOUT"}

The identical command without --speedtest-url completed 30/30 in ~20 s (1.5–1.7 items/s) on the same machine.

Expected

Enabling the optional speed measurement should slow a run down and leave speed.mbps empty for proxies that cannot transfer bytes — not stop results from being committed. Running the three variants above on the same host, I'd expect all of them to make progress, with the speedtest one only slower proportionally to --speedtest-bytes.

Hypothesis

  • The speed probe has its own budget (WHOLE_PROBE_TIMEOUT, ~16 s per proxy here regardless of --timeout 8), which is fine by itself, but an item appears not to be committed until the speed stage finishes — so on a network where the transfer never completes, items pile up instead of being recorded as "checked, speed unavailable".
  • At 338 707 candidates this shows up as a total stall (0/338707), which is the scariest symptom: the progress line and job state say the job is running while the database receives nothing.

Workaround

Omit --speedtest-url / --speedtest-bytes; the rest of the pipeline then behaves as documented. The speed/bandwidth sort keys and the Mbit/s column in the UI are unavailable as a result.

Notes

  • Both large jobs ended as partial in the job table (the second one was terminated by me after ~35 min of zero progress; the first one likewise), but the measured fact is that in 2090 s not one of 338 707 items left pending.
  • I'm happy to re-run any variation, or attach a diagnose bundle, if that helps localise it. I did not include the database (≈500 MB) or personal paths here.

RU, кратко: при включённом --speedtest-url проверка практически перестаёт записывать результаты: на 30 прокси в базу попал 1 результат за 90 с против 30 за 20 с без флага, а на 338 707 кандидатах вывод 35 минут стоял на Checked 0/338707 при нуле наблюдений. Без --speedtest-url та же база проверяется на 6–11 адресов/с. Похоже, элемент не коммитится, пока не отработает этап замера скорости, а он в такой сети висит до WHOLE_PROBE_TIMEOUT (16 с).

Activity

  1. alberttt46-spec commented on Sep 28, 2026

    @alberttt46-spec
    Author

    Small correction for transparency, so the states in the report are not misread:

    • job-38251cf1c365efaa was terminated by me after it had shown Checked 0/338707 for ~110 s, so that row is left as running in the job table (hard kill, no graceful shutdown) and EXIT=1. The 417 rows of observations/results present in that database copy predate this job — they came from the control run without --speedtest-url that I ran in the same copy before starting it. The speedtest job itself wrote nothing.
    • job-3799c320d82e8a84 (the 2090 s one) also shows partial because it was likewise stopped by hand after ~35 min of zero progress.

    Final lines of the second run for the record:

    Workers: 32; full pass; profile 38611189595d6c2ca5ee
    job: job-38251cf1c365efaa
    Checked 0/338707; 0.0 proxies/s; ~0.0 min left
    ...
    

    I can re-run it with a proper --deadline only (no external kill) if you want a clean partial/failed state in the job row instead.

  2. alberttt46-spec commented on Sep 28, 2026

    @alberttt46-spec
    Author

    Final numbers for the control run (variant B, no --speedtest-url), for completeness:

    • job-d92349f6e056d159 reached Checked 434/338707; 7.2 proxies/s before I stopped it (it was on track for ~13 h, which is why I capped it) — its row is also left as running in the job table because of the hard kill.
    • 417 items done, 417 observations, 417 results written; all 338 290 remaining items still pending.

    So on the same 338 707-candidate database: 434 checked in ~60 s without the flag vs 0 checked in 2090 s with it. No further action needed from me on this thread — I just wanted the states and numbers to be reproducible rather than approximate.

  3. DavidVoitenko commented on Sep 28, 2026

    @DavidVoitenko
    Owner

    Thanks for the detailed report. I reproduced a real persistence bug behind the Checked 0/... symptom: when the basic target check failed, the pipeline skipped the optional speed stage, but saving the result had been deferred to that stage. Those failed checks were therefore left out of results and observations, and their job items stayed pending.

    I fixed this in draft PR #34: a failed basic check is now saved immediately; proxies that pass the basic check still proceed to the speed test. A regression test showed 0/8 saved before the change and 8/8 afterward, including observations and job item states. The related local tests pass; GitHub CI is still running.

    This fixes the lost results and stalled progress. It does not shorten the speed probe's timeout for a proxy that passes the basic check but cannot transfer from the speed URL. Your Windows/TUN setup would be useful for final confirmation once the fix is available. Thanks again for the careful measurements.

  4. DavidVoitenko commented on Sep 28, 2026

    @DavidVoitenko
    Owner

    Update: the fix is now in the published v3.0.4 release, including the proxy_workbench-3.0.4-py3-none-any.whl you used for 3.0.3. PR #34 is merged. CI passed on Windows, macOS, and Linux, and the release builds completed.

    I also added a socket-level reproduction: eight local proxies accept TCP, seven never answer the target request, and one answers the target but stalls on the speed URL. With --speedtest-url enabled, the new version records all eight results and observations, leaves no job items pending, and reports the speed failure for the one proxy that passed the basic check.

    If convenient, could you rerun only your 30-proxy variant on Windows/TUN with 3.0.4? There is no need to rerun the 338,707-candidate set. The speed probe can still take its timeout for a proxy that passes the basic check but cannot transfer bytes; the change fixes the missing results and frozen progress for failed basic checks. Your result will confirm the behavior in your specific network.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions