Skip to content

PDF: one form-like page makes the rest of the document interleave two-column text #2580

Description

@nicoacn

Summary

When PdfConverter finds at least one page that looks like a form or table, it extracts every other page with pdfplumber's page.extract_text(). That call reads straight across both columns of a two-column layout, so the prose pages come out interleaved. If no page looks like a form, the whole document goes through pdfminer, which reads each column top to bottom.

So the reading order of page 1 depends on whether page 11 has a table.

Reproduction

A synthetic PDF, generated with reportlab: two pages of two-column prose ("Article 1." … "Article 24.") plus, in the second file, one extra page with a 4-column grid.

File Article order from markitdown Order inversions
two-column pages only 1, 2, 3, 4, 5, 6, 7, 8, … 24 0
same pages + one form-like page 1, 7, 2, 8, 3, 9, 4, 10, … 10

pdftotext reads both files in order.

On a real document (an 11-page, two-column law in Spanish, 40 articles; only its last page is detected as form-like), markitdown 0.1.8 gives 8 order inversions; pdfminer on the same pages gives 1.

Where

packages/markitdown/src/markitdown/converters/_pdf_converter.py, around lines 552-574 (commit 1f9530a):

  • if form_page_count == 0: pdfminer.high_level.extract_text(pdf_bytes) for the whole document;
  • otherwise: plain pages use page.extract_text() (pdfplumber), joined with "\n\n".

plain_page_indices is filled in that loop but never used, which suggests the plain pages were meant to be handled separately.

A fix, and its trade-off

I have a small patch ready: keep the form pages as they are, and re-extract the plain pages with one pdfminer pass (extract_text(..., page_numbers=plain_page_indices), split on \f), keeping the pdfplumber text if that pass fails. It fixes the reproduction (10 → 0 inversions) and the real document (8 → 1), passes the existing PDF tests, and adds a regression test.

It is not free, which is why I'm asking before opening a PR:

Two options I can see:

  1. Plain pages always through pdfminer (the patch above), i.e. the same policy the prose path already has.
  2. Keep pdfplumber for plain pages, but pass layout-aware options or detect multi-column pages before choosing.

Happy to open the PR for whichever you prefer. The reproduction script (reportlab) is below.

Related but separate: page markers (#1304, #2272).

Environment

markitdown 0.1.8 (main @ 1f9530a), pdfminer.six 20251230+, pdfplumber 0.11.10, Python 3.14, macOS.

Reproduction script
"""Minimal repro: a two-column PDF reads in order, until one form-like page is added."""
import re, sys
from reportlab.lib.pagesizes import A4
from reportlab.pdfgen import canvas
from markitdown import MarkItDown

W, H = A4
FILLER = "This paragraph is long enough to wrap across several lines of its column so that the two columns run side by side on the page."

def column(c, x, y, first, last):
    for n in range(first, last + 1):
        c.setFont("Helvetica-Bold", 10); c.drawString(x, y, f"Article {n}."); y -= 14
        c.setFont("Helvetica", 9)
        words, line = FILLER.split(), ""
        for w in words:
            if c.stringWidth(line + " " + w, "Helvetica", 9) > 200:
                c.drawString(x, y, line.strip()); y -= 12; line = ""
            line += " " + w
        c.drawString(x, y, line.strip()); y -= 22

def build(path, with_form_page):
    c = canvas.Canvas(path, pagesize=A4)
    art = 1
    for _ in range(2):  # two pages, two columns, 6 articles per column
        column(c, 50, H - 60, art, art + 5); column(c, 320, H - 67, art + 6, art + 11)
        art += 12; c.showPage()
    if with_form_page:  # one page that looks like a form/table to the heuristic
        c.setFont("Helvetica", 10)
        for r in range(12):
            for k, x in enumerate((50, 200, 350, 470)):
                c.drawString(x, H - 80 - r * 18, f"{'Field' if k == 0 else 'Value'}{r}-{k}")
        c.showPage()
    c.save()

def order(md):
    seen = []
    for n in map(int, re.findall(r"Article (\d+)\.", md)):
        if n not in seen: seen.append(n)
    return seen

for flag in (False, True):
    p = f"repro_{'with' if flag else 'without'}_form_page.pdf"
    build(p, flag)
    s = order(MarkItDown().convert(p).markdown)
    print(f"{p}: {s}  -> {sum(b < a for a, b in zip(s, s[1:]))} order inversions")

Activity

  1. added a commit that references this issue on Oct 4, 2026
    b8bf2a7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions