Skip to content

perf(html): render Markdown with turbohtml - #2557

Draft
Bernát Gábor (gaborbernat) wants to merge 10 commits into
microsoft:mainfrom
gaborbernat:turbohtml-markdown
Draft

Bernát Gábor (gaborbernat) wants to merge 10 commits into
microsoft:mainfrom
gaborbernat:turbohtml-markdown

Conversation

@gaborbernat

@gaborbernat Bernát Gábor (gaborbernat) commented Sep 27, 2026 •

Copy link
Copy Markdown

Use turbohtml 1.14.1 for HTML, Wikipedia, Bing SERP and RSS conversion on Python 3.11 and later. HTML parsing and Markdown rendering run in C. Python 3.10 retains BeautifulSoup and markdownify, including its plain-text fallback and strict=True recursion errors. Conditional dependencies keep Python 3.10 installable without turbohtml.

Conversion paths

%%{init: {"look": "classic", "htmlLabels": false, "markdownAutoWrap": false, "themeVariables": {"fontFamily": "Arial"}}}%%
flowchart TB
    Pages["HTML / Wikipedia / Bing"] --> Backend{"Python version"}
    Backend -->|3.11+| Native["turbohtml HTML tree"]
    Backend -->|3.10| Legacy["BeautifulSoup + markdownify"] --> Output
    Feed["RSS / Atom feed"] --> XML["XML feed parser"] --> Kind{"Entry content"}
    Kind -->|HTML / XHTML| Backend
    Kind -->|plain text| Plain["Keep text verbatim"]
    Kind -->|binary| Skip["Omit payload"]
    Native --> Target["Select converter content"]
    Target --> Render["turbohtml<br/>Markdown renderer"] --> Output["Markdown body"]
    Plain --> Output
    classDef input fill:#d8e9f8,stroke:#4979a5,color:#17324d;
    classDef render fill:#dcecdf,stroke:#4d8058,color:#203f28;
    classDef route fill:#f5e5c6,stroke:#ab8040,color:#533f20;
    class Pages,Feed,Native,XML input;
    class Target,Render,Output,Plain,Legacy render;
    class Kind,Skip,Backend route;
    linkStyle default stroke:#718096,stroke-width:1.5px;
Loading

Behavior differences

These examples compare Microsoft MarkItDown at 4cc9fa1, using markdownify 1.2.3 and BeautifulSoup 4.15.0, with this integration on Python 3.11+. Python 3.10 retains the baseline formatting. Each diff compares Markdown source: removed lines show MarkItDown 4cc9fa1; added lines show this PR. In HTML examples, \n denotes a newline.

  1. Deep HTML. 500 <div> wrappers around <p>Deep content with <b>bold text</b></p>, with recursion limit 200.

    @@ -1,2 +1 @@
    -Deep content with
    -bold text
    +Deep content with **bold text**

    Retain content and formatting beyond Python’s recursion limit. turbohtml’s iterative renderer uses an explicit stack instead of recursive calls. strict=True also succeeds instead of raising the recursion error.

  2. Paragraph whitespace. <p>one\ntwo</p>.

    @@ -1,2 +1 @@
    -one
    -two
    +one two

    Chromium renders normal-flow text as one two, and CommonMark soft breaks permit the same result. GitHub uses different soft-newline rules for Markdown files and issue/PR text; emitting a space avoids turning source indentation into a visible break. Separate paragraphs, <br> breaks and code-block newlines remain distinct. CSS-preserved whitespace needs the separate fix proposed below.

  3. Paragraphs inside a table cell. <td><p>one</p><p>two</p></td> inside a table.

    @@ -1,3 +1,3 @@
     |  |
     | --- |
    -| one  two |
    +| one two |

    GFM table cells hold inline content. Paragraphs inside a cell use one separator space. Both converters flatten those paragraphs into one line; this change removes one separator space.

  4. List children inside inline code. <code><ul><li>one</li><li>two</li></ul></code>.

    @@ -1,2 +1 @@
    -`* one
    -* two`
    +`one two`

    CommonMark and GitHub parse markdownify’s output as broken inline code followed by a list. The new output is one code span with both words and no invented list punctuation. It flattens the source’s block layout.

  5. Paragraph children inside inline code. <code><p>one</p><p>two</p></code>.

    @@ -1,3 +1 @@
    -`one
    -
    -two`
    +`one two`

    The blank line in markdownify’s output ends the paragraph and breaks the code span. The new output is one valid code span. As with lists inside inline code, it flattens the source’s block layout.

  6. Language classes. <pre><code class="language-python">one</code></pre>.

    @@ -1,3 +1,3 @@
    -```
    +```python
     one
     ```

    Preserve the declared language: HTML recommends language-* classes, and GitHub uses fence language identifiers for highlighting. This reads author-supplied metadata rather than guessing from the code. An explicit code_language overrides the class, including an empty string to disable the label. The callback retains priority over that option.

  7. Malformed table content. <table><p>Before</p><tr><td>Cell</td></tr></table>.

    @@ -1,3 +1,5 @@
     Before
     
    +|  |
    +| --- |
     | Cell |

    HTML5 parsing moves the paragraph outside the table. The renderer adds the delimiter row required for a GFM table without a header.

  8. First newline after an opening pre tag. <pre>\none</pre>, with strip_pre=None.

    @@ -1,4 +1,3 @@
     ```
    -
     one
     ```

    HTML5 parsing removes the first newline adjacent to the opening <pre> tag. strip_pre controls the text remaining after parsing.

  9. Backticks inside fenced code. <pre><code>```</code></pre>.

    @@ -1,3 +1,3 @@
    -```
    -```
    -```
    +````
    +```
    +````

    A four-backtick fence prevents the content from closing its own block. markdownify 1.2.3 uses three backticks for this input.

  10. Feed callback errors. For <pre>Garden</pre> in an RSS or Atom HTML entry, code_language_callback raises ValueError("Cannot classify Garden").

    -<pre>Garden</pre>
    +ValueError: Cannot classify Garden

    RssConverter.convert() propagates the callback error instead of returning the raw HTML as Markdown. This correction applies to both backends, matching the HTML converter’s error handling.

Ordinary <pre><code>one\ntwo</code></pre> keeps the newline, and <p>one</p><p>two</p> remains one\n\ntwo.

Compatibility fixes

Restore code-block options. code_language_callback receives a BeautifulSoup Tag, including its children and ancestors, and a falsey callback result falls back to code_language. Build that compatibility tree when callers supply a callback. strip_pre preserves markdownify 1.2.3’s trimming modes and raises ValueError for an invalid mode.

The following outputs match MarkItDown 4cc9fa1 and this PR:

  1. Language callback. For <pre><code class="python">one</code></pre>, a callback that reads the child’s class produces:

    ```python
    one
    ```
  2. Code trimming. For <pre><code>\n\none\n\n</code></pre>, the default strip_pre="strip" produces "```\none\n```"; strip_pre="strip_one" produces "```\n\none\n```"; strip_pre=None produces "```\n\n\none\n\n\n```". These quoted strings use \n for each newline.

  3. Empty code blocks. <pre></pre> produces empty output.

  4. Title and template fragments. Both <title>Garden</title><p>Beans</p> and <template>Garden</template><p>Beans</p> produce:

    Garden
    
    Beans

The fragment path retains title and template text that the native renderer would otherwise omit. RSS/Atom text entries remain text: type="text" content containing &lt;job_id&gt; keeps the literal <job_id>, while binary entries remain omitted. HTML/XHTML entry content uses the Markdown renderer.

Corrections and remaining work

turbohtml 1.14.1 includes the native language-class correction. This PR remains draft pending the CSS whitespace fix.

  1. Language class scanning. turbohtml 1.14.1 includes the C fix for HTML class tokens. Both <code class="language-python highlight">print(1)</code> and <code class="highlight language-python">print(1)</code> select python. The scanner accepts leading whitespace and whitespace around <code>, and recognizes language-* on <pre> as supported by Prism.

  2. Explicit language options. The integration preserves caller precedence. For <pre><code class="language-python">print(1)</code></pre> with code_language="ruby", the output is:

    ```ruby
    print(1)
    ```

    A callback uses its nonempty result or falls back to code_language, matching markdownify. Without a callback, precedence is explicit code_language, declared class, then no language. An explicit empty string disables the label. Use the empty-string opt-out for code that must display as text: GitHub renders mermaid, geojson, topojson and stl fences as diagrams or models, beyond syntax highlighting. This correction lives in MarkItDown’s option adapter; turbohtml’s language setting remains a fallback.

  3. Handle CSS-preserved whitespace in turbohtml. <p style="white-space: pre-wrap">one\ntwo</p> displays on two lines in Chromium. The current conversion removes that break:

    -one
    -two
    +one two

    CSS distinguishes normal whitespace from pre, pre-wrap and pre-line. turbohtml #1283 preserves inline and inherited whitespace modes in preformatted blocks. Exact spacing takes precedence over bold and link formatting inside those regions. A generic soft-newline option would still render as a space in a CommonMark document. The proposed fix considers inline declarations and inheritance; it does not load stylesheets.

Release validation

Benchmarks use the released turbohtml 1.14.1 wheel on macOS arm64, Python 3.14.7, comparing MarkItDown 4cc9fa1 with this integration at 7b04fd7. Five baseline/PR/PR/baseline cycles use fresh processes. Times below are the median of ten per-process medians for MarkItDown.convert_stream(), including file-type detection. Each process warms the input before measuring; tiny inputs use 1,000 calls and larger inputs use five. Negative changes mean less time.

Input Baseline, ms This PR, ms Time change
Blog 8.003 3.613 -54.9%
Wikipedia 112.723 13.105 -88.4%
Bing results 26.702 5.146 -80.7%
RSS 44.554 13.319 -70.1%
Title fragment 1.894 1.935 +2.1%
Template fragment 1.963 2.003 +2.1%
Table with a cell list 2.020 1.928 -4.6%
One-paragraph document 1.888 1.828 -3.2%
2,000 repeated paragraphs 86.394 5.038 -94.2%

Tiny-input results vary across cycles: the title fragment ranges from −25.3% to +19.4% in paired comparisons. Profiling 1,000 title-fragment conversions attributes 94% of the native pipeline time to stream-info detection, dominated by Magika’s ONNX inference. The HTML converter takes 54 ms in that profile versus 357 ms on the baseline. The profile supports a conversion improvement; it does not establish an end-to-end gain for tiny documents.

The four small inputs are <title>Garden notes</title><p>Planting beans.</p>, <template>Garden notes</template><p>Beans.</p>, <table><tr><td><ul><li>Beans</li><li>Peas</li></ul></td><td><p>Garden plan</p></td></tr></table>, and <html><body><p>Garden notes and planting beans.</p></body></html>. The repeated input contains 2,000 copies of <p>Garden notes, <strong>beans</strong> and <em>peas</em>.</p>.

Disclosure: I maintain turbohtml.

@gaborbernat

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

Markdownify's Python traversal slows large pages and fails on deep
nesting. Use turbohtml's C parser and renderer for HTML-derived
conversion while retaining converter-specific behavior.

Require turbohtml 1.13.0 for list and code-boundary corrections found
during migration.
Keep code_language_callback's BeautifulSoup Tag interface and its
code_language fallback. Honor strip_pre modes while retaining native
fence escaping and code layout inside lists and tables.
Keep converter keyword arguments open while typing the native Markdown
options. Propagate feed callback errors instead of returning raw HTML,
and verify links, parser failures, and formatting through public APIs.
Keep BeautifulSoup and markdownify on Python 3.10 while using turbohtml
on Python 3.11 and later. Preserve the existing plain-text fallback and
strict recursion errors on Python 3.10. Select dependencies by Python
version so the package remains installable without turbohtml there.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant