Repository navigation
Add an edit timeline and a summarised PDF report - #5
Merged
Merged
Conversation
Signatures - Load each signature separately. pyHanko's embedded_signatures raised on the first unreadable signature container, which left every signature in the file unvalidated and only produced a LOW note that blamed a legacy format. - An unreadable signature container is now signature.unparseable (HIGH). A pyHanko failure while reading the signatures marks the analysis incomplete instead of only being recorded in the facts. - signature.validation-error keeps MEDIUM effective severity (confidence MEDIUM instead of LOW): integrity is unknown, which needs review. - pdfsig reporting that it did not verify a signature is now pdfsig.integrity-unknown instead of producing no finding. Summary, batch report and exit code - "Checked and found in order" lines are only written for analysers that ran and reported no error, for PDF and Office alike, and the signature line is withheld when any signature finding leaves integrity in doubt. - The batch report marks incomplete verdicts and counts them apart from documents without significant findings. - --fail-on also exits 1 when an analysis is incomplete or a file could not be analysed.
…s, and make --fail-on-incomplete opt-in Follow-up after review: - Document timestamps (/DocTimeStamp) are validated with validate_pdf_timestamp instead of failing validate_pdf_signature, so PAdES-LTA and other timestamped files are no longer flagged. - pdfsig: an unsigned signature field is skipped, and "integrity unknown" is only reported when pyHanko did not settle the integrity of that signature either (Poppler does not verify document timestamps). - pyHanko decrypts files that have only an owner password (empty user password), or uses --password, before reading the signatures. - signature.unparseable is only used when the CMS container itself cannot be read; other loader failures are validation errors. signature.not-validated is not added when pyHanko could not read the file at all, since the analysis is then already incomplete. - Field names are stored as plain strings (pyHanko returns proxy objects for encrypted files, which broke report serialisation). - --fail-on keeps its previous behaviour. The new --fail-on-incomplete exits 1 when an analysis is incomplete or a file could not be analysed. The batch summary lists incomplete files in its JSON.
…n, times are claims
…rouping, layout - A revision that did not write /Info or XMP gets no save time instead of inheriting the previous revision's, and the consistency check only compares times revisions wrote themselves. - Times from a signature or timestamp that does not match the file are only "claimed". Only pyHanko's signed signingTime attribute counts as signed; a /M fallback stays claimed. RFC 3161 tokens are labelled "issuer not verified by this tool", since trust roots are not configured. - Consecutive Word tracked changes by the same author at the same time are one change, so text split over formatting runs reads correctly. - Timezone-naive dates are compared as UTC with a 14-hour tolerance instead of being skipped, so a backdate without an offset is flagged. - Spreadsheets and presentations are marked as not covered, and the PDF report says so instead of "no edits found". - Long changes are split over several table rows by text length, so a fully rewritten page no longer makes report generation fail. - --pdf-report treats the path as a directory whenever several files were given, even if only one could be analysed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This adds an edit timeline and a summarised PDF report. For a PDF or Word document it shows what was edited, when, and what it was changed to, and
--pdf-reportwrites the result as a PDF.This branch builds on #4 (document timestamps are validated there). Until #4 is merged, the diff below also shows its two commits; the commits for this feature start at "Record signing and timestamp times from pyHanko".
What the timeline shows
Each row has a time, the source of that time, what changed, and before → after. For
examples/bank-statement/03_edited_in_acrobat.pdf:…Lön Exempelföretaget AB 32 450,00 36 602,52→…52 450,00 56 602,52, and five more linesSources: the text diff between revisions (already computed), the save time each revision wrote into Info or XMP, signature and document-timestamp times from pyHanko, XMP history, and Word tracked changes (author,
w:date, text).How much a time proves
Every time is labelled with what backs it, because a time inside a file is only as good as its source:
w:date). Anyone who edits the file can change it.A revision that did not write its own save time gets none, rather than inheriting the previous one. Times from a signature that does not match the file drop to claimed.
Ordering
PDF rows follow the order of the revisions, which the file's bytes fix, and are sorted by time within a revision. Sorting purely by claimed time would let a backdated edit appear before the signature it follows. When the claimed times contradict the revision order, a new finding
timeline.inconsistent-times(MEDIUM) is raised.05_signed_then_edited.pdfnow gets this finding: revision 3 claims 2026-09-18 but comes after a signature made on 2026-09-26. Its verdict is unchanged. If you prefer pure time order, it is a small change intimeline._sort.Changes
timeline.py(new): builds the events from other analysers' facts and runs throughcollect(), so a failure marks the analysis incomplete.pdfreport.py(new): the PDF report with reportlab. It contains the verdict, the key findings, the timeline table, what was checked, limitations and method. reportlab is imported only when a report is written.revisions.py: records the save time each revision claims.signatures.py: records signing and timestamp times.office.py: keeps each Word tracked change with author, date and text.cli.py:--pdf-report PATH, which takes a directory when several files are analysed. Without reportlab it exits 2 with an install hint.pyproject.toml: new extrareport = ["reportlab>=4.0"]. README: options, extras, licenses, findings.Excel and PowerPoint are not covered; the report says so.
Testing
pytest: all tests pass, also with-W error::DeprecationWarning. There are new tests for each source of time, the evidence labels (including broken signatures and a missing signingTime attribute), revision ordering with a backdated edit, timezone-naive dates, Word changes split over formatting runs, a revision that could not be reconstructed, and the PDF report. The report tests cover hostile markup, characters outside Windows-1252, a fully rewritten page, very long unbroken lines, a missing timeline, spreadsheets, batch output and the case without reportlab.05additionally getstimeline.inconsistent-times.03and05were rendered and checked visually. The tables wrap inside the page, å, ä and ö render, and the header repeats on each page.