Skip to content

Harden revision discovery against incomplete or misleading xref chains - #2

Merged
overjoyde merged 3 commits into
overjoyde:mainfrom
pepo72:fix/revision-discovery
Sep 26, 2026
Merged

overjoyde merged 3 commits into
overjoyde:mainfrom
pepo72:fix/revision-discovery

Conversation

@pepo72

@pepo72 pepo72 commented Sep 26, 2026

Copy link
Copy Markdown

Summary

Revision boundaries were derived entirely from the chain of startxref and /Prev links, and a file was treated as linearized if the /Linearized keyword appeared anywhere in its first 2 KB. Both come from whoever wrote the file. In some structurally unusual files this meant an earlier version with different page content was never compared, and the summary could still state that no page content had been changed.

This PR makes revision discovery independent of the chain being complete and correct.

Changes

  • raw.py: linearization is recognised only from a linearization dictionary that is the file's first object. The first-page section must be the lowest section in the file and link forward with /Prev; an incremental update always links backwards, so it can no longer be merged into revision 1.
  • raw.py: a startxref ... %%EOF outside any stream that the chain does not reach, and that points at a real cross-reference section, now closes a recovered revision. Streams are skipped by their direct /Length, so a PDF carried inside a stream is not mistaken for a revision.
  • revisions.py: new finding structure.unlinked-revision (MEDIUM). Recovered revisions are compared like any other, so a content change in them produces the usual revisions.content-changed.
  • revisions.py: when a file has more revisions than --max-revisions, the original is kept as the baseline, so a change inside the skipped range is still compared. The limit is reported as revisions.not-all-compared (LOW).
  • summary.py: "No page content was changed through appended edits" is no longer stated when the chain is broken, a revision was recovered, or there are %%EOF markers the chain does not explain.
  • README and docs/ANALYSIS.md updated.

Testing

  • pytest: all tests pass, including new regression tests for linearized files with updates, broken and missing /Prev links, misleading linearization hints, and a PDF embedded in a stream.
  • The bank-statement examples give identical verdicts and findings before and after.
  • The raw parser was run over about 330 mutated PDFs without exceptions. A timing test checks that it stays linear on inputs with very many stream headers or startxref markers.
  • Checked against genuine files for false positives: plain, object streams, linearized (with and without object streams, with and without updates), signed variants of each, and PDFs attached to PDFs with direct and indirect /Length.

Happy to share more detail on the constructions privately if useful.

Revision boundaries were taken entirely from the startxref -> /Prev
chain, and linearization was inferred from the /Linearized keyword
anywhere in the first 2 KB. Both are written by whoever produced the
file, so an earlier version with different page content could go
unreported.

- Linearization is recognised only from a linearization dictionary that
  is the file's first object, and only the lowest section that links
  forward with /Prev can be treated as the first-page section. An
  incremental update always links backwards and can no longer be merged
  into revision 1.
- A startxref ... %%EOF outside any stream, not reached by the chain and
  pointing at a real cross-reference section, now closes a recovered
  revision. It is compared like any other revision and reported as
  structure.unlinked-revision (MEDIUM). Streams are skipped by their
  direct /Length so a PDF carried inside a stream is not mistaken for
  a revision.
- The summary no longer states that no page content was changed when
  the chain is broken, a revision was recovered, or there are %%EOF
  markers the chain does not explain.

Results for the bank-statement examples are unchanged. Adds regression
tests for linearized files with updates, and for chains that are
broken, missing a link, or carry a misleading linearization hint.
… are skipped

- Bound the look-back for a stream's dictionary and the forward scan of
  a candidate xref stream to fixed windows, cache the check per target,
  and cap the number of startxref markers examined outside the chain.
  Hitting the cap is recorded as a chain error. Repeated stream headers
  or startxref markers no longer make parsing quadratic.
- A startxref whose target is already a section in the chain is not a
  separate version. This avoids a false recovered revision when a file
  carries an uncompressed copy of itself with an indirect /Length.
- The direct-/Length pattern no longer matches a prefix of an indirect
  reference ("/Length 12 0 R").
- Resolve /Prev the same way as the chain before checking that the
  first-page section links forward.
- When there are more revisions than --max-revisions, the original is
  kept as the baseline so changes inside the skipped range are still
  compared, and revisions.not-all-compared (LOW) reports the limit.
@overjoyde
overjoyde merged commit df1c1ae into overjoyde:main Sep 26, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants