Repository navigation
Harden revision discovery against incomplete or misleading xref chains - #2
Merged
Merged
Conversation
Revision boundaries were taken entirely from the startxref -> /Prev chain, and linearization was inferred from the /Linearized keyword anywhere in the first 2 KB. Both are written by whoever produced the file, so an earlier version with different page content could go unreported. - Linearization is recognised only from a linearization dictionary that is the file's first object, and only the lowest section that links forward with /Prev can be treated as the first-page section. An incremental update always links backwards and can no longer be merged into revision 1. - A startxref ... %%EOF outside any stream, not reached by the chain and pointing at a real cross-reference section, now closes a recovered revision. It is compared like any other revision and reported as structure.unlinked-revision (MEDIUM). Streams are skipped by their direct /Length so a PDF carried inside a stream is not mistaken for a revision. - The summary no longer states that no page content was changed when the chain is broken, a revision was recovered, or there are %%EOF markers the chain does not explain. Results for the bank-statement examples are unchanged. Adds regression tests for linearized files with updates, and for chains that are broken, missing a link, or carry a misleading linearization hint.
… are skipped
- Bound the look-back for a stream's dictionary and the forward scan of
a candidate xref stream to fixed windows, cache the check per target,
and cap the number of startxref markers examined outside the chain.
Hitting the cap is recorded as a chain error. Repeated stream headers
or startxref markers no longer make parsing quadratic.
- A startxref whose target is already a section in the chain is not a
separate version. This avoids a false recovered revision when a file
carries an uncompressed copy of itself with an indirect /Length.
- The direct-/Length pattern no longer matches a prefix of an indirect
reference ("/Length 12 0 R").
- Resolve /Prev the same way as the chain before checking that the
first-page section links forward.
- When there are more revisions than --max-revisions, the original is
kept as the baseline so changes inside the skipped range are still
compared, and revisions.not-all-compared (LOW) reports the limit.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Revision boundaries were derived entirely from the chain of
startxrefand/Prevlinks, and a file was treated as linearized if the/Linearizedkeyword appeared anywhere in its first 2 KB. Both come from whoever wrote the file. In some structurally unusual files this meant an earlier version with different page content was never compared, and the summary could still state that no page content had been changed.This PR makes revision discovery independent of the chain being complete and correct.
Changes
raw.py: linearization is recognised only from a linearization dictionary that is the file's first object. The first-page section must be the lowest section in the file and link forward with/Prev; an incremental update always links backwards, so it can no longer be merged into revision 1.raw.py: astartxref ... %%EOFoutside any stream that the chain does not reach, and that points at a real cross-reference section, now closes a recovered revision. Streams are skipped by their direct/Length, so a PDF carried inside a stream is not mistaken for a revision.revisions.py: new findingstructure.unlinked-revision(MEDIUM). Recovered revisions are compared like any other, so a content change in them produces the usualrevisions.content-changed.revisions.py: when a file has more revisions than--max-revisions, the original is kept as the baseline, so a change inside the skipped range is still compared. The limit is reported asrevisions.not-all-compared(LOW).summary.py: "No page content was changed through appended edits" is no longer stated when the chain is broken, a revision was recovered, or there are%%EOFmarkers the chain does not explain.docs/ANALYSIS.mdupdated.Testing
pytest: all tests pass, including new regression tests for linearized files with updates, broken and missing/Prevlinks, misleading linearization hints, and a PDF embedded in a stream.startxrefmarkers./Length.Happy to share more detail on the constructions privately if useful.