Repository navigation
Parse whitespace-only input in linear time - #103
Open
BrianWillows wants to merge 3 commits into
Open
BrianWillows wants to merge 3 commits into
BrianWillows wants to merge 3 commits into
Conversation
This is no longer supported or operational.
SRT_REGEX starts with \s* (added for cdown#50), so finditer retries every position of a long whitespace run and each retry rescans the rest of the run: 32000 spaces or newlines took about 22 seconds to parse. A match starting at p with whitespace at p - 1 implies a match at p - 1, so the leftmost match is either at the expected position or at one not preceded by whitespace. _iter_srt_matches tries the expected position first and otherwise searches with a (?<!\s) guard, returning exactly the matches finditer did.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
srt.parse()takes quadratic time on a long run of whitespace that contains no subtitle. On current master:"x"+ 32,000 spacesThis matters for anything that parses user-uploaded subtitle files (media servers, subtitle converters, web upload handlers): a ~32 KB file costs one CPU core about 20 seconds.
Cause.
SRT_REGEXstarts with\s*(added in 0d7e5fd for #50, to accept leading whitespace). When there is no match,finditerretries at every position in the run, and each retry rescans the rest of the run.Fix.
SRT_REGEXitself is unchanged.parse()now iterates over_iter_srt_matches(), which returns exactly the same matches asfinditer:pandsrt[p - 1]is whitespace, a match also starts atp - 1, because the leading\s*can absorb that character. So the leftmost match at or afterposstarts either atposor at a position not preceded by whitespace.SRT_REGEX.match(srt, pos)first, which is the normal path for contiguous subtitles. If that fails, it searches with(?<!\s)prepended, which skips the retries inside a whitespace run in constant time.Verification.
finditer._iter_srt_matchesequalsfinditer, and a timing test. The full suite passes (46 tests). Without the change tosrt.py, both new tests fail.srt.pyuses only Python 2.7-compatible syntax, in line withpython_requires.Found with AI assistance (Claude) and verified by hand.