crawl4ai version
0.9.4
Expected Behavior
crawler_configs applies to every request that carries it: a /crawl with one URL, and /crawl/stream (or /crawl with stream: true). If some case can't support it, a 400 rather than a 200 that ran with a different config.
Current Behavior
Two gaps in the config-list support from #1852:
handle_crawl_request uses the list only when len(urls) > 1 (api.py#L726); with one URL it calls arun() with crawler_config (#L749).
stream_process doesn't pass crawler_configs down, and handle_stream_crawl_request has no parameter for it (server.py#L1030-L1036, api.py#L889-L895).
Both answer 200, so the caller can't tell its per-URL settings were dropped. The single-URL case is pinned by tests/test_issue_1837_config_list.py::test_single_url_ignores_crawler_configs ("arun only takes one config"), but arun_many handles a one-URL list fine.
cc @hafezparast @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- /crawl, one URL, crawler_configs: [{url_matcher: "*", css_selector: ".only-a"}]
- /crawl/stream, two URLs, the same list
Steps to Reproduce
1. Page with two sections, .only-a and .only-b
2. POST /crawl with that one URL and the crawler_configs above
3. The markdown still contains the .only-b section
4. Same list on /crawl/stream: same result
Code snippets
import httpx
# URL: a page with
# <div class="only-a"><p>SECTION-A is the part a per-URL css_selector keeps.</p></div>
# <div class="only-b"><p>SECTION-B is the part it drops.</p></div>
per_url = [{"type": "CrawlerRunConfig",
"params": {"url_matcher": "*", "css_selector": ".only-a", "cache_mode": "bypass"}}]
r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
json={"urls": [URL], "crawler_configs": per_url}, timeout=60)
md = r.json()["results"][0]["markdown"]["raw_markdown"]
print("SECTION-B" in md) # True: the per-URL css_selector was not applied
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
Local test page with .only-a / .only-b sections (CRAWL4AI_ALLOW_INTERNAL_URLS=true); A / B = that section is in the markdown:
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=True
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=True; A=True B=True
crawl4ai version
0.9.4
Expected Behavior
crawler_configsapplies to every request that carries it: a/crawlwith one URL, and/crawl/stream(or/crawlwithstream: true). If some case can't support it, a 400 rather than a 200 that ran with a different config.Current Behavior
Two gaps in the config-list support from #1852:
handle_crawl_requestuses the list only whenlen(urls) > 1(api.py#L726); with one URL it callsarun()withcrawler_config(#L749).stream_processdoesn't passcrawler_configsdown, andhandle_stream_crawl_requesthas no parameter for it (server.py#L1030-L1036, api.py#L889-L895).Both answer 200, so the caller can't tell its per-URL settings were dropped. The single-URL case is pinned by
tests/test_issue_1837_config_list.py::test_single_url_ignores_crawler_configs("arun only takes one config"), butarun_manyhandles a one-URL list fine.cc @hafezparast @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- /crawl, one URL, crawler_configs: [{url_matcher: "*", css_selector: ".only-a"}] - /crawl/stream, two URLs, the same listSteps to Reproduce
Code snippets
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
Local test page with .only-a / .only-b sections (CRAWL4AI_ALLOW_INTERNAL_URLS=true); A / B = that section is in the markdown:
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=True
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=True; A=True B=True