Summary
When PdfConverter finds at least one page that looks like a form or table, it extracts every other page with pdfplumber's page.extract_text(). That call reads straight across both columns of a two-column layout, so the prose pages come out interleaved. If no page looks like a form, the whole document goes through pdfminer, which reads each column top to bottom.
So the reading order of page 1 depends on whether page 11 has a table.
Reproduction
A synthetic PDF, generated with reportlab: two pages of two-column prose ("Article 1." … "Article 24.") plus, in the second file, one extra page with a 4-column grid.
| File |
Article order from markitdown |
Order inversions |
| two-column pages only |
1, 2, 3, 4, 5, 6, 7, 8, … 24 |
0 |
| same pages + one form-like page |
1, 7, 2, 8, 3, 9, 4, 10, … |
10 |
pdftotext reads both files in order.
On a real document (an 11-page, two-column law in Spanish, 40 articles; only its last page is detected as form-like), markitdown 0.1.8 gives 8 order inversions; pdfminer on the same pages gives 1.
Where
packages/markitdown/src/markitdown/converters/_pdf_converter.py, around lines 552-574 (commit 1f9530a):
- if
form_page_count == 0: pdfminer.high_level.extract_text(pdf_bytes) for the whole document;
- otherwise: plain pages use
page.extract_text() (pdfplumber), joined with "\n\n".
plain_page_indices is filled in that loop but never used, which suggests the plain pages were meant to be handled separately.
A fix, and its trade-off
I have a small patch ready: keep the form pages as they are, and re-extract the plain pages with one pdfminer pass (extract_text(..., page_numbers=plain_page_indices), split on \f), keeping the pdfplumber text if that pass fails. It fixes the reproduction (10 → 0 inversions) and the real document (8 → 1), passes the existing PDF tests, and adds a regression test.
It is not free, which is why I'm asking before opening a PR:
Two options I can see:
- Plain pages always through pdfminer (the patch above), i.e. the same policy the prose path already has.
- Keep pdfplumber for plain pages, but pass layout-aware options or detect multi-column pages before choosing.
Happy to open the PR for whichever you prefer. The reproduction script (reportlab) is below.
Related but separate: page markers (#1304, #2272).
Environment
markitdown 0.1.8 (main @ 1f9530a), pdfminer.six 20251230+, pdfplumber 0.11.10, Python 3.14, macOS.
Reproduction script
"""Minimal repro: a two-column PDF reads in order, until one form-like page is added."""
import re, sys
from reportlab.lib.pagesizes import A4
from reportlab.pdfgen import canvas
from markitdown import MarkItDown
W, H = A4
FILLER = "This paragraph is long enough to wrap across several lines of its column so that the two columns run side by side on the page."
def column(c, x, y, first, last):
for n in range(first, last + 1):
c.setFont("Helvetica-Bold", 10); c.drawString(x, y, f"Article {n}."); y -= 14
c.setFont("Helvetica", 9)
words, line = FILLER.split(), ""
for w in words:
if c.stringWidth(line + " " + w, "Helvetica", 9) > 200:
c.drawString(x, y, line.strip()); y -= 12; line = ""
line += " " + w
c.drawString(x, y, line.strip()); y -= 22
def build(path, with_form_page):
c = canvas.Canvas(path, pagesize=A4)
art = 1
for _ in range(2): # two pages, two columns, 6 articles per column
column(c, 50, H - 60, art, art + 5); column(c, 320, H - 67, art + 6, art + 11)
art += 12; c.showPage()
if with_form_page: # one page that looks like a form/table to the heuristic
c.setFont("Helvetica", 10)
for r in range(12):
for k, x in enumerate((50, 200, 350, 470)):
c.drawString(x, H - 80 - r * 18, f"{'Field' if k == 0 else 'Value'}{r}-{k}")
c.showPage()
c.save()
def order(md):
seen = []
for n in map(int, re.findall(r"Article (\d+)\.", md)):
if n not in seen: seen.append(n)
return seen
for flag in (False, True):
p = f"repro_{'with' if flag else 'without'}_form_page.pdf"
build(p, flag)
s = order(MarkItDown().convert(p).markdown)
print(f"{p}: {s} -> {sum(b < a for a, b in zip(s, s[1:]))} order inversions")
Summary
When
PdfConverterfinds at least one page that looks like a form or table, it extracts every other page with pdfplumber'spage.extract_text(). That call reads straight across both columns of a two-column layout, so the prose pages come out interleaved. If no page looks like a form, the whole document goes through pdfminer, which reads each column top to bottom.So the reading order of page 1 depends on whether page 11 has a table.
Reproduction
A synthetic PDF, generated with reportlab: two pages of two-column prose ("Article 1." … "Article 24.") plus, in the second file, one extra page with a 4-column grid.
markitdownpdftotextreads both files in order.On a real document (an 11-page, two-column law in Spanish, 40 articles; only its last page is detected as form-like),
markitdown0.1.8 gives 8 order inversions; pdfminer on the same pages gives 1.Where
packages/markitdown/src/markitdown/converters/_pdf_converter.py, around lines 552-574 (commit1f9530a):form_page_count == 0:pdfminer.high_level.extract_text(pdf_bytes)for the whole document;page.extract_text()(pdfplumber), joined with"\n\n".plain_page_indicesis filled in that loop but never used, which suggests the plain pages were meant to be handled separately.A fix, and its trade-off
I have a small patch ready: keep the form pages as they are, and re-extract the plain pages with one pdfminer pass (
extract_text(..., page_numbers=plain_page_indices), split on\f), keeping the pdfplumber text if that pass fails. It fixes the reproduction (10 → 0 inversions) and the real document (8 → 1), passes the existing PDF tests, and adds a regression test.It is not free, which is why I'm asking before opening a PR:
fi), rotated watermarks one letter per line, and some table-like pages collapsing into long lines.Two options I can see:
Happy to open the PR for whichever you prefer. The reproduction script (reportlab) is below.
Related but separate: page markers (#1304, #2272).
Environment
markitdown 0.1.8 (main @
1f9530a), pdfminer.six 20251230+, pdfplumber 0.11.10, Python 3.14, macOS.Reproduction script