crawl4ai version
0.9.4
Expected Behavior
The other seeds are crawled, and the refused one comes back as a failed result carrying the same opaque message, the way a robots.txt refusal already does inside a batch.
Current Behavior
_normalize_and_validate_seeds checks every seed before crawling and raises on the first one refused (api.py#L650-L657), so the whole request fails with 400 {"detail": "URL blocked (SSRF protection): URL blocked"} and no results. resolve_and_pin also raises EgressBlocked when getaddrinfo fails (egress_broker.py#L92-L96), so a host that simply doesn't resolve (a dead domain from a stale sitemap) takes the batch down the same way. The caller can't tell which seed it was, so its only recovery is to split the batch and retry.
Keeping the message opaque, with no resolution oracle, is right, and a per-URL failure keeps it: it says no more than today's 400, because a caller can already find the refused seed by sending the URLs one at a time.
If per-URL failures are the direction you want, I can send a PR. cc @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- urls: ["https://example.com/", "https://no-such-host-12345.example/"]
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop), without CRAWL4AI_ALLOW_INTERNAL_URLS
2. POST /crawl with the two URLs above
3. 400, no results; the first URL alone crawls fine
Code snippets
import httpx
r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
json={"urls": ["https://example.com/", "https://no-such-host-12345.example/"]}, timeout=60)
print(r.status_code, r.text) # 400 {"detail":"URL blocked (SSRF protection): URL blocked"}
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
[seedbatch] one resolvable + one unresolvable seed -> HTTP 400: {"detail":"URL blocked (SSRF protection): URL blocked"}
[seedbatch] the resolvable seed alone -> HTTP 200, success=True
crawl4ai version
0.9.4
Expected Behavior
The other seeds are crawled, and the refused one comes back as a failed result carrying the same opaque message, the way a robots.txt refusal already does inside a batch.
Current Behavior
_normalize_and_validate_seedschecks every seed before crawling and raises on the first one refused (api.py#L650-L657), so the whole request fails with400 {"detail": "URL blocked (SSRF protection): URL blocked"}and no results.resolve_and_pinalso raisesEgressBlockedwhengetaddrinfofails (egress_broker.py#L92-L96), so a host that simply doesn't resolve (a dead domain from a stale sitemap) takes the batch down the same way. The caller can't tell which seed it was, so its only recovery is to split the batch and retry.Keeping the message opaque, with no resolution oracle, is right, and a per-URL failure keeps it: it says no more than today's 400, because a caller can already find the refused seed by sending the URLs one at a time.
If per-URL failures are the direction you want, I can send a PR. cc @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop), without CRAWL4AI_ALLOW_INTERNAL_URLS 2. POST /crawl with the two URLs above 3. 400, no results; the first URL alone crawls fineCode snippets
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
[seedbatch] one resolvable + one unresolvable seed -> HTTP 400: {"detail":"URL blocked (SSRF protection): URL blocked"}
[seedbatch] the resolvable seed alone -> HTTP 200, success=True