Where blocky-writer is
Rust/WASM core for fixed-layout PDF and application-form filling, green on main
at e1b6225. Two exports: detect_blocks (find the fillable widgets in a PDF)
and fill_blocks (write values into the AcroForm by field name).
The gap you could fill
detect_blocks labels each widget from the AcroForm field's /T, falling back
to the parent field's /T, falling back to a synthesised field_<page>_<index>.
That last fallback is where the weakness is: a scanned or badly-authored form
often has no /T at all, and we currently emit a positional label that means
nothing to a human.
Your pipeline — OCR, NER, metadata extraction and classification — is the obvious
source of a better label. A block sitting at a given rectangle, on a page whose
text layer has been OCR'd, usually has a printed caption immediately to its left
or above it. We do not do that and are not planning to.
What you might use from here
The inverse direction. detect_blocks returns exact PDF user-space rectangles
per widget:
{ "label": "Name", "x": 50.0, "y": 700.0, "width": 200.0, "height": 24.0 }
If your extraction wants to know which regions of a page are interactive rather
than merely textual, that is a ready-made answer, and it comes with the AcroForm
field name when one exists.
What we are not proposing
An integration. No shared code, no dependency, no contract. This is a note that
the two shapes fit together, so that if either of us later needs the other we
start from a known seam rather than a guess.
Ref: hyperpolymath/blocky-writer#73
Where blocky-writer is
Rust/WASM core for fixed-layout PDF and application-form filling, green on
mainat
e1b6225. Two exports:detect_blocks(find the fillable widgets in a PDF)and
fill_blocks(write values into the AcroForm by field name).The gap you could fill
detect_blockslabels each widget from the AcroForm field's/T, falling backto the parent field's
/T, falling back to a synthesisedfield_<page>_<index>.That last fallback is where the weakness is: a scanned or badly-authored form
often has no
/Tat all, and we currently emit a positional label that meansnothing to a human.
Your pipeline — OCR, NER, metadata extraction and classification — is the obvious
source of a better label. A block sitting at a given rectangle, on a page whose
text layer has been OCR'd, usually has a printed caption immediately to its left
or above it. We do not do that and are not planning to.
What you might use from here
The inverse direction.
detect_blocksreturns exact PDF user-space rectanglesper widget:
{ "label": "Name", "x": 50.0, "y": 700.0, "width": 200.0, "height": 24.0 }If your extraction wants to know which regions of a page are interactive rather
than merely textual, that is a ready-made answer, and it comes with the AcroForm
field name when one exists.
What we are not proposing
An integration. No shared code, no dependency, no contract. This is a note that
the two shapes fit together, so that if either of us later needs the other we
start from a known seam rather than a guess.
Ref: hyperpolymath/blocky-writer#73