Repository navigation
perf(html): render Markdown with turbohtml - #2557
Draft
Bernát Gábor (gaborbernat) wants to merge 10 commits into
Draft
Bernát Gábor (gaborbernat) wants to merge 10 commits into
Bernát Gábor (gaborbernat) wants to merge 10 commits into
Conversation
Author
|
@microsoft-github-policy-service agree |
Bernát Gábor (gaborbernat)
marked this pull request as ready for review
September 30, 2026 00:19
Bernát Gábor (gaborbernat)
marked this pull request as draft
September 30, 2026 03:41
Bernát Gábor (gaborbernat)
force-pushed
the
turbohtml-markdown
branch
2 times, most recently
from
September 30, 2026 14:03
dca585a to
5abdcf3
Compare
Markdownify's Python traversal slows large pages and fails on deep nesting. Use turbohtml's C parser and renderer for HTML-derived conversion while retaining converter-specific behavior. Require turbohtml 1.13.0 for list and code-boundary corrections found during migration.
Bernát Gábor (gaborbernat)
force-pushed
the
turbohtml-markdown
branch
from
October 9, 2026 22:16
5abdcf3 to
b037c64
Compare
Keep code_language_callback's BeautifulSoup Tag interface and its code_language fallback. Honor strip_pre modes while retaining native fence escaping and code layout inside lists and tables.
Keep converter keyword arguments open while typing the native Markdown options. Propagate feed callback errors instead of returning raw HTML, and verify links, parser failures, and formatting through public APIs.
Bernát Gábor (gaborbernat)
force-pushed
the
turbohtml-markdown
branch
from
October 10, 2026 05:55
6bc35fd to
1915fa9
Compare
Keep BeautifulSoup and markdownify on Python 3.10 while using turbohtml on Python 3.11 and later. Preserve the existing plain-text fallback and strict recursion errors on Python 3.10. Select dependencies by Python version so the package remains installable without turbohtml there.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Use turbohtml 1.14.1 for HTML, Wikipedia, Bing SERP and RSS conversion on Python 3.11 and later. HTML parsing and Markdown rendering run in C. Python 3.10 retains BeautifulSoup and markdownify, including its plain-text fallback and
strict=Truerecursion errors. Conditional dependencies keep Python 3.10 installable without turbohtml.Conversion paths
%%{init: {"look": "classic", "htmlLabels": false, "markdownAutoWrap": false, "themeVariables": {"fontFamily": "Arial"}}}%% flowchart TB Pages["HTML / Wikipedia / Bing"] --> Backend{"Python version"} Backend -->|3.11+| Native["turbohtml HTML tree"] Backend -->|3.10| Legacy["BeautifulSoup + markdownify"] --> Output Feed["RSS / Atom feed"] --> XML["XML feed parser"] --> Kind{"Entry content"} Kind -->|HTML / XHTML| Backend Kind -->|plain text| Plain["Keep text verbatim"] Kind -->|binary| Skip["Omit payload"] Native --> Target["Select converter content"] Target --> Render["turbohtml<br/>Markdown renderer"] --> Output["Markdown body"] Plain --> Output classDef input fill:#d8e9f8,stroke:#4979a5,color:#17324d; classDef render fill:#dcecdf,stroke:#4d8058,color:#203f28; classDef route fill:#f5e5c6,stroke:#ab8040,color:#533f20; class Pages,Feed,Native,XML input; class Target,Render,Output,Plain,Legacy render; class Kind,Skip,Backend route; linkStyle default stroke:#718096,stroke-width:1.5px;Behavior differences
These examples compare Microsoft MarkItDown at
4cc9fa1, using markdownify 1.2.3 and BeautifulSoup 4.15.0, with this integration on Python 3.11+. Python 3.10 retains the baseline formatting. Each diff compares Markdown source: removed lines show MarkItDown4cc9fa1; added lines show this PR. In HTML examples,\ndenotes a newline.Deep HTML. 500
<div>wrappers around<p>Deep content with <b>bold text</b></p>, with recursion limit 200.Retain content and formatting beyond Python’s recursion limit. turbohtml’s iterative renderer uses an explicit stack instead of recursive calls.
strict=Truealso succeeds instead of raising the recursion error.Paragraph whitespace.
<p>one\ntwo</p>.Chromium renders normal-flow text as
one two, and CommonMark soft breaks permit the same result. GitHub uses different soft-newline rules for Markdown files and issue/PR text; emitting a space avoids turning source indentation into a visible break. Separate paragraphs,<br>breaks and code-block newlines remain distinct. CSS-preserved whitespace needs the separate fix proposed below.Paragraphs inside a table cell.
<td><p>one</p><p>two</p></td>inside a table.GFM table cells hold inline content. Paragraphs inside a cell use one separator space. Both converters flatten those paragraphs into one line; this change removes one separator space.
List children inside inline code.
<code><ul><li>one</li><li>two</li></ul></code>.CommonMark and GitHub parse markdownify’s output as broken inline code followed by a list. The new output is one code span with both words and no invented list punctuation. It flattens the source’s block layout.
Paragraph children inside inline code.
<code><p>one</p><p>two</p></code>.The blank line in markdownify’s output ends the paragraph and breaks the code span. The new output is one valid code span. As with lists inside inline code, it flattens the source’s block layout.
Language classes.
<pre><code class="language-python">one</code></pre>.Preserve the declared language: HTML recommends
language-*classes, and GitHub uses fence language identifiers for highlighting. This reads author-supplied metadata rather than guessing from the code. An explicitcode_languageoverrides the class, including an empty string to disable the label. The callback retains priority over that option.Malformed table content.
<table><p>Before</p><tr><td>Cell</td></tr></table>.HTML5 parsing moves the paragraph outside the table. The renderer adds the delimiter row required for a GFM table without a header.
First newline after an opening pre tag.
<pre>\none</pre>, withstrip_pre=None.HTML5 parsing removes the first newline adjacent to the opening
<pre>tag.strip_precontrols the text remaining after parsing.Backticks inside fenced code.
<pre><code>```</code></pre>.A four-backtick fence prevents the content from closing its own block. markdownify 1.2.3 uses three backticks for this input.
Feed callback errors. For
<pre>Garden</pre>in an RSS or Atom HTML entry,code_language_callbackraisesValueError("Cannot classify Garden").RssConverter.convert()propagates the callback error instead of returning the raw HTML as Markdown. This correction applies to both backends, matching the HTML converter’s error handling.Ordinary
<pre><code>one\ntwo</code></pre>keeps the newline, and<p>one</p><p>two</p>remainsone\n\ntwo.Compatibility fixes
Restore code-block options.
code_language_callbackreceives a BeautifulSoupTag, including its children and ancestors, and a falsey callback result falls back tocode_language. Build that compatibility tree when callers supply a callback.strip_prepreserves markdownify 1.2.3’s trimming modes and raisesValueErrorfor an invalid mode.The following outputs match MarkItDown
4cc9fa1and this PR:Language callback. For
<pre><code class="python">one</code></pre>, a callback that reads the child’s class produces:Code trimming. For
<pre><code>\n\none\n\n</code></pre>, the defaultstrip_pre="strip"produces"```\none\n```";strip_pre="strip_one"produces"```\n\none\n```";strip_pre=Noneproduces"```\n\n\none\n\n\n```". These quoted strings use\nfor each newline.Empty code blocks.
<pre></pre>produces empty output.Title and template fragments. Both
<title>Garden</title><p>Beans</p>and<template>Garden</template><p>Beans</p>produce:The fragment path retains title and template text that the native renderer would otherwise omit. RSS/Atom text entries remain text:
type="text"content containing<job_id>keeps the literal<job_id>, while binary entries remain omitted. HTML/XHTML entry content uses the Markdown renderer.Corrections and remaining work
turbohtml 1.14.1 includes the native language-class correction. This PR remains draft pending the CSS whitespace fix.
Language class scanning. turbohtml 1.14.1 includes the C fix for HTML class tokens. Both
<code class="language-python highlight">print(1)</code>and<code class="highlight language-python">print(1)</code>selectpython. The scanner accepts leading whitespace and whitespace around<code>, and recognizeslanguage-*on<pre>as supported by Prism.Explicit language options. The integration preserves caller precedence. For
<pre><code class="language-python">print(1)</code></pre>withcode_language="ruby", the output is:A callback uses its nonempty result or falls back to
code_language, matching markdownify. Without a callback, precedence is explicitcode_language, declared class, then no language. An explicit empty string disables the label. Use the empty-string opt-out for code that must display as text: GitHub rendersmermaid,geojson,topojsonandstlfences as diagrams or models, beyond syntax highlighting. This correction lives in MarkItDown’s option adapter; turbohtml’s language setting remains a fallback.Handle CSS-preserved whitespace in turbohtml.
<p style="white-space: pre-wrap">one\ntwo</p>displays on two lines in Chromium. The current conversion removes that break:CSS distinguishes normal whitespace from
pre,pre-wrapandpre-line. turbohtml #1283 preserves inline and inherited whitespace modes in preformatted blocks. Exact spacing takes precedence over bold and link formatting inside those regions. A generic soft-newline option would still render as a space in a CommonMark document. The proposed fix considers inline declarations and inheritance; it does not load stylesheets.Release validation
Benchmarks use the released turbohtml 1.14.1 wheel on macOS arm64, Python 3.14.7, comparing MarkItDown
4cc9fa1with this integration at7b04fd7. Five baseline/PR/PR/baseline cycles use fresh processes. Times below are the median of ten per-process medians forMarkItDown.convert_stream(), including file-type detection. Each process warms the input before measuring; tiny inputs use 1,000 calls and larger inputs use five. Negative changes mean less time.Tiny-input results vary across cycles: the title fragment ranges from −25.3% to +19.4% in paired comparisons. Profiling 1,000 title-fragment conversions attributes 94% of the native pipeline time to stream-info detection, dominated by Magika’s ONNX inference. The HTML converter takes 54 ms in that profile versus 357 ms on the baseline. The profile supports a conversion improvement; it does not establish an end-to-end gain for tiny documents.
The four small inputs are
<title>Garden notes</title><p>Planting beans.</p>,<template>Garden notes</template><p>Beans.</p>,<table><tr><td><ul><li>Beans</li><li>Peas</li></ul></td><td><p>Garden plan</p></td></tr></table>, and<html><body><p>Garden notes and planting beans.</p></body></html>. The repeated input contains 2,000 copies of<p>Garden notes, <strong>beans</strong> and <em>peas</em>.</p>.Disclosure: I maintain turbohtml.