Skip to content

Feature request: opt-in crawl source receipt for RAG ingestion #2307

Description

@shuyan-code

Problem / use case

Crawl4AI is positioned as a web-to-Markdown input for RAG and agent data pipelines. A typical ingestion job needs to answer three questions for every stored document: which source and crawl produced it, did the usable text actually change, and can the same processed artifact be traced later? Those answers drive re-embedding cost, stale-document replacement, and source attribution in downstream answers.

Today CrawlResult already exposes useful pieces such as url, redirected_url, response_headers, cached_at, cache_status, and Markdown (current model). The cache also has internal content hashes. But the result does not provide one documented, serialization-friendly record tying the selected output text to a stable digest, its actual fetch time, and its source. Every RAG integrator must assemble this independently, and a cache hit can otherwise be mistaken for a fresh fetch.

Proposed module: optional crawl source receipt

Add an opt-in, small source_receipt attached to each successful CrawlResult (and available through the Docker API serialization). Suggested fields:

  • Requested URL and final URL, reusing existing redirect information.
  • fetched_at: time the underlying page was fetched; on a cache hit preserve the original fetch time. Optionally include served_at separately.
  • Output kind (raw_markdown or fit_markdown) and a versioned SHA-256 digest of the exact UTF-8 text returned for that kind. Make the normalization/encoding rule explicit. An unchanged digest lets an ingestion job skip embedding; a changed digest triggers replacement.
  • Crawl4AI version and a fingerprint of content-affecting, non-secret processing options, so a changed extraction configuration does not look like unchanged content.
  • Existing cache status, and a clear distinction between successful extraction and fallback/partial output.

The default path can remain unchanged, with receipt generation behind a config flag. Please exclude cookies, Authorization headers, tokens, proxy credentials, and arbitrary response headers from the receipt/fingerprint. Consumers can retain the receipt next to their vector-store document ID; this does not require Crawl4AI to own a vector store or scheduler.

Acceptance criteria

  1. Fresh crawl, redirect, and cache-hit examples show accurate URL and fetch-time semantics.
  2. Identical selected Markdown + processing configuration yields the same digest; a text or relevant configuration change changes it.
  3. Raw and fit Markdown cannot accidentally share a digest namespace.
  4. Library and Docker API outputs serialize the receipt consistently; existing responses remain compatible when the option is off.
  5. Tests cover secret exclusion and a failed/partial crawl.

I checked main and develop result fields, the roadmap, and issue/PR searches for provenance, source lineage, content fingerprints, and crawl manifests. I found a planned incremental embedding index in the roadmap, but not this per-crawl receipt. The receipt could serve that later index while remaining useful to external ingestion pipelines.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions