Skip to content

Parse whitespace-only input in linear time - #103

Open
BrianWillows wants to merge 3 commits into
cdown:masterfrom
BrianWillows:linear-leading-whitespace
Open

BrianWillows wants to merge 3 commits into
cdown:masterfrom
BrianWillows:linear-leading-whitespace

Conversation

@BrianWillows

Copy link
Copy Markdown

srt.parse() takes quadratic time on a long run of whitespace that contains no subtitle. On current master:

input before after
32,000 spaces 22.8 s 1.8 ms
32,000 newlines 22.6 s 1.8 ms
"x" + 32,000 spaces 22.7 s 1.8 ms
valid cues + 32 KB whitespace tail 1.2 ms 1.3 ms

This matters for anything that parses user-uploaded subtitle files (media servers, subtitle converters, web upload handlers): a ~32 KB file costs one CPU core about 20 seconds.

Cause. SRT_REGEX starts with \s* (added in 0d7e5fd for #50, to accept leading whitespace). When there is no match, finditer retries at every position in the run, and each retry rescans the rest of the run.

Fix. SRT_REGEX itself is unchanged. parse() now iterates over _iter_srt_matches(), which returns exactly the same matches as finditer:

  • If a match starts at p and srt[p - 1] is whitespace, a match also starts at p - 1, because the leading \s* can absorb that character. So the leftmost match at or after pos starts either at pos or at a position not preceded by whitespace.
  • It tries SRT_REGEX.match(srt, pos) first, which is the normal path for contiguous subtitles. If that fails, it searches with (?<!\s) prepended, which skips the retries inside a whitespace run in constant time.

Verification.

  • A differential over 20,000 random inputs (cues, CRLF cues, cues without an index, garbage, whitespace runs, BOM) found 0 differences in match spans or groups against finditer.
  • Two new tests: a hypothesis test that _iter_srt_matches equals finditer, and a timing test. The full suite passes (46 tests). Without the change to srt.py, both new tests fail.
  • The change to srt.py uses only Python 2.7-compatible syntax, in line with python_requires.

Found with AI assistance (Claude) and verified by hand.

Igor Kozyrenko and others added 3 commits November 2, 2023 00:35
This is no longer supported or operational.
SRT_REGEX starts with \s* (added for cdown#50), so finditer retries every
position of a long whitespace run and each retry rescans the rest of the
run: 32000 spaces or newlines took about 22 seconds to parse.

A match starting at p with whitespace at p - 1 implies a match at p - 1,
so the leftmost match is either at the expected position or at one not
preceded by whitespace. _iter_srt_matches tries the expected position
first and otherwise searches with a (?<!\s) guard, returning exactly the
matches finditer did.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants