Problem / use case
Crawl4AI is positioned as a web-to-Markdown input for RAG and agent data pipelines. A typical ingestion job needs to answer three questions for every stored document: which source and crawl produced it, did the usable text actually change, and can the same processed artifact be traced later? Those answers drive re-embedding cost, stale-document replacement, and source attribution in downstream answers.
Today CrawlResult already exposes useful pieces such as url, redirected_url, response_headers, cached_at, cache_status, and Markdown (current model). The cache also has internal content hashes. But the result does not provide one documented, serialization-friendly record tying the selected output text to a stable digest, its actual fetch time, and its source. Every RAG integrator must assemble this independently, and a cache hit can otherwise be mistaken for a fresh fetch.
Proposed module: optional crawl source receipt
Add an opt-in, small source_receipt attached to each successful CrawlResult (and available through the Docker API serialization). Suggested fields:
- Requested URL and final URL, reusing existing redirect information.
fetched_at: time the underlying page was fetched; on a cache hit preserve the original fetch time. Optionally include served_at separately.
- Output kind (
raw_markdown or fit_markdown) and a versioned SHA-256 digest of the exact UTF-8 text returned for that kind. Make the normalization/encoding rule explicit. An unchanged digest lets an ingestion job skip embedding; a changed digest triggers replacement.
- Crawl4AI version and a fingerprint of content-affecting, non-secret processing options, so a changed extraction configuration does not look like unchanged content.
- Existing cache status, and a clear distinction between successful extraction and fallback/partial output.
The default path can remain unchanged, with receipt generation behind a config flag. Please exclude cookies, Authorization headers, tokens, proxy credentials, and arbitrary response headers from the receipt/fingerprint. Consumers can retain the receipt next to their vector-store document ID; this does not require Crawl4AI to own a vector store or scheduler.
Acceptance criteria
- Fresh crawl, redirect, and cache-hit examples show accurate URL and fetch-time semantics.
- Identical selected Markdown + processing configuration yields the same digest; a text or relevant configuration change changes it.
- Raw and fit Markdown cannot accidentally share a digest namespace.
- Library and Docker API outputs serialize the receipt consistently; existing responses remain compatible when the option is off.
- Tests cover secret exclusion and a failed/partial crawl.
I checked main and develop result fields, the roadmap, and issue/PR searches for provenance, source lineage, content fingerprints, and crawl manifests. I found a planned incremental embedding index in the roadmap, but not this per-crawl receipt. The receipt could serve that later index while remaining useful to external ingestion pipelines.
Problem / use case
Crawl4AI is positioned as a web-to-Markdown input for RAG and agent data pipelines. A typical ingestion job needs to answer three questions for every stored document: which source and crawl produced it, did the usable text actually change, and can the same processed artifact be traced later? Those answers drive re-embedding cost, stale-document replacement, and source attribution in downstream answers.
Today
CrawlResultalready exposes useful pieces such asurl,redirected_url,response_headers,cached_at,cache_status, and Markdown (current model). The cache also has internal content hashes. But the result does not provide one documented, serialization-friendly record tying the selected output text to a stable digest, its actual fetch time, and its source. Every RAG integrator must assemble this independently, and a cache hit can otherwise be mistaken for a fresh fetch.Proposed module: optional crawl source receipt
Add an opt-in, small
source_receiptattached to each successfulCrawlResult(and available through the Docker API serialization). Suggested fields:fetched_at: time the underlying page was fetched; on a cache hit preserve the original fetch time. Optionally includeserved_atseparately.raw_markdownorfit_markdown) and a versioned SHA-256 digest of the exact UTF-8 text returned for that kind. Make the normalization/encoding rule explicit. An unchanged digest lets an ingestion job skip embedding; a changed digest triggers replacement.The default path can remain unchanged, with receipt generation behind a config flag. Please exclude cookies, Authorization headers, tokens, proxy credentials, and arbitrary response headers from the receipt/fingerprint. Consumers can retain the receipt next to their vector-store document ID; this does not require Crawl4AI to own a vector store or scheduler.
Acceptance criteria
I checked
mainanddevelopresult fields, the roadmap, and issue/PR searches for provenance, source lineage, content fingerprints, and crawl manifests. I found a planned incremental embedding index in the roadmap, but not this per-crawl receipt. The receipt could serve that later index while remaining useful to external ingestion pipelines.