Skip to content

[Bug]: crawler_configs is silently ignored for a single URL and on /crawl/stream #2287

Description

@talelboussetta

crawl4ai version

0.9.4

Expected Behavior

crawler_configs applies to every request that carries it: a /crawl with one URL, and /crawl/stream (or /crawl with stream: true). If some case can't support it, a 400 rather than a 200 that ran with a different config.

Current Behavior

Two gaps in the config-list support from #1852:

  1. handle_crawl_request uses the list only when len(urls) > 1 (api.py#L726); with one URL it calls arun() with crawler_config (#L749).
  2. stream_process doesn't pass crawler_configs down, and handle_stream_crawl_request has no parameter for it (server.py#L1030-L1036, api.py#L889-L895).

Both answer 200, so the caller can't tell its per-URL settings were dropped. The single-URL case is pinned by tests/test_issue_1837_config_list.py::test_single_url_ignores_crawler_configs ("arun only takes one config"), but arun_many handles a one-URL list fine.

cc @hafezparast @ntohidi

Is this reproducible?

Yes

Inputs Causing the Bug

- /crawl, one URL, crawler_configs: [{url_matcher: "*", css_selector: ".only-a"}]
- /crawl/stream, two URLs, the same list

Steps to Reproduce

1. Page with two sections, .only-a and .only-b
2. POST /crawl with that one URL and the crawler_configs above
3. The markdown still contains the .only-b section
4. Same list on /crawl/stream: same result

Code snippets

import httpx

# URL: a page with
#   <div class="only-a"><p>SECTION-A is the part a per-URL css_selector keeps.</p></div>
#   <div class="only-b"><p>SECTION-B is the part it drops.</p></div>
per_url = [{"type": "CrawlerRunConfig",
            "params": {"url_matcher": "*", "css_selector": ".only-a", "cache_mode": "bypass"}}]
r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
               json={"urls": [URL], "crawler_configs": per_url}, timeout=60)
md = r.json()["results"][0]["markdown"]["raw_markdown"]
print("SECTION-B" in md)   # True: the per-URL css_selector was not applied

OS

Linux (Docker image built from develop @ 1f68e5b)

Python version

3.12 (the image's)

Browser

Chromium (Playwright, bundled)

Browser version

No response

Error logs & Screenshots (if applicable)

Local test page with .only-a / .only-b sections (CRAWL4AI_ALLOW_INTERNAL_URLS=true); A / B = that section is in the markdown:
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=True
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=True; A=True B=True

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    🐞 BugSomething isn't working🩺 Needs TriageNeeds attention of maintainers

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions