Parallel typed CSV reads, text-based schema inference and in-memory encryption (v5.1.0) - #131
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…handle's row cursor Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…overlapped handle on Windows A synchronous handle serializes every read on its file object, so 16 chunk readers sharing the FileStream's handle queued behind each other (cell scan 52 -> 25 ms on a 160 MB file). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tion's columns The partitions were merged into the first one's builders on one thread (53 ms of a 160 MB parse) before BuildTable copied the result again. BuildTable now concatenates every partition's buffers into the native block once, rebasing string offsets and splicing validity bits. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e tight loops The single loop checked for the dot and for leading zeros on every digit. Significant digits are now only recounted when the raw digit count exceeds the limit, and Pow10 is a table. Identical results to the old parser over 20M random inputs; ~3-4% on sequential CSV and XLSX typed parses through the AOT library.
…on one thread xl_parse_typed and xl_parse_arrow now route through ParseTypedTable with a degree of parallelism of 1, but the chunk source was resolved before the partition check, so every file-backed CSV parse on Windows reopened the file for overlapped reads it never used. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
FastDouble accepts an exponent, so parseText typed a column of part codes such as 12E4 as Float64 and parse_typed then silently read it as 120000. Text with an e or E now stays a string, like leading-zero codes do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…and its _ex twin xl_parse_typed/xl_parse_typed_ex and xl_parse_arrow/xl_parse_arrow_ex were copies apart from the degree of parallelism. Each pair now forwards to one private helper, and the arrow path takes its cause from the caller, so xl_parse_arrow_ex names itself in a live-session error like xl_parse_typed_ex. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
After the partitions are parsed, copying every column into its native block ran on one thread and took about a quarter of a parallel parse. Columns now build under Parallel.For when the parse itself ran in parallel; a sequential parse keeps the single-threaded loop. On a 521 MB CSV with 16 threads the parse drops from ~370 ms to ~330 ms read from a file and to ~300 ms read from memory. The columns block is zeroed first so a failed build frees whichever columns finished, and the allocation hooks move from ThreadStatic to AsyncLocal so fault-injection tests still reach the pool threads. Merging a byte-aligned validity bitmap also copies straight into the destination instead of through a temporary array. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The double parser was inlined into AppendFrom, so every cell of every column type paid its eight-register prologue and stack zeroing just to be dispatched. Int64, Float64 and Bool now append through their own methods, and RecordValidity's no-null fast path is inlined. With the static PGO profile regenerated for the new methods, an AOT typed parse gets ~2.5% faster on CSV and ~2% on XLSB; XLSX is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t column build Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
xl_encrypt_package only takes file paths, so encrypting a workbook meant writing its plaintext package to disk first. The new export takes the package bytes and returns the encrypted workbook in an xl_buffer released with xl_free_buffer, so a package from xl_write_typed_to_memory never reaches disk. Both exports share the password length check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
encrypt_package_bytes() encrypts a package's bytes through xl_encrypt_package_to_memory, so write_workbook_to_bytes() plus it produces an encrypted workbook without the plaintext touching disk. write_arrow_to_bytes(), write_pandas_to_bytes() and write_polars_to_bytes() are the in-memory twins of the frame writers, sharing their Arrow-to-table conversion with the file versions. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #131 +/- ##
==========================================
+ Coverage 89.32% 89.47% +0.15%
==========================================
Files 180 181 +1
Lines 13416 13611 +195
Branches 2462 2514 +52
==========================================
+ Hits 11984 12179 +195
Misses 954 954
Partials 478 478 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Benchmark ResultsBaseline: master Moved by at least 10%
57 within ±10%. Full results per groupExcelReader.Benchmarks.ArrowConversionBenchmark
ExcelReader.Benchmarks.ChunkedParseBenchmark
ExcelReader.Benchmarks.ColdStartBenchmark
ExcelReader.Benchmarks.CsvParallelParseBenchmark
ExcelReader.Benchmarks.CsvParallelVsSepBenchmark
ExcelReader.Benchmarks.CsvParseBenchmark
ExcelReader.Benchmarks.CsvReadBenchmark
ExcelReader.Benchmarks.CsvWriteBenchmark
ExcelReader.Benchmarks.DataReaderBenchmark
ExcelReader.Benchmarks.EncryptedWorkbookBenchmark
ExcelReader.Benchmarks.NativeRowReadBenchmark
ExcelReader.Benchmarks.NativeTypedParseBenchmark
ExcelReader.Benchmarks.ParseBenchmark
ExcelReader.Benchmarks.ReadBenchmark
ExcelReader.Benchmarks.RealDataReadBenchmark
ExcelReader.Benchmarks.RealDataTypedParseBenchmark
ExcelReader.Benchmarks.RecordWriteBenchmark
ExcelReader.Benchmarks.StringHeavyReadBenchmark
ExcelReader.Benchmarks.WriteBenchmark
ExcelReader.Benchmarks.WritePathBenchmark
ExcelReader.Benchmarks.XlsReadBenchmark
ExcelReader.Benchmarks.XlsWriteBenchmark
ExcelReader.Benchmarks.XlsxSharedStringHotPathBenchmark
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
xl_parse_typed_ex/xl_parse_arrow_extake a degree ofparallelism; Python's
parse_typed,to_arrow,to_record_batchandto_polarsexpose it asparallelism(1 = one thread, the default; 0 = every core). ~3.6x faster on a 160 MB CSV with16 threads. Other formats, and CSVs under 256 KiB, still read on one thread with the same result.
Excel.InferSchema(..., parseText: true),xl_infer_schema_exwith
XL_INFER_PARSE_TEXT, and Python'sinfer_schema(parse_text=True)type text cells (everyCSV field) as integers, decimals, booleans and ISO dates/date-times when the text has exactly that
shape. Leading-zero codes, scientific notation, integers past
long, padded or comma-decimalnumbers and non-ISO dates stay strings.
xl_encrypt_package_to_memoryand Python'sencrypt_package_bytes()encrypt a package without writing its plaintext to disk.write_arrow_to_bytes,write_pandas_to_bytesandwrite_polars_to_bytescomplete thein-memory writers.
table's columns are built concurrently;
ColumnBuilder.AppendFromis a thin per-cell dispatcher(-2.5% CSV, -2% XLSB sequential in AOT). The static PGO profile is regenerated.
Fixes
parallel CSV aggregation).
(regression from this branch's parallel work).
Compatibility
XL_ABI_VERSIONstays 5; every new export is additive.InferSchema(reader, headerRow, sampleSize)keeps its binary signature; the new overloadadds
parseText = false._exexports andxl_encrypt_package_to_memory(the Rust
.defis updated so the import library links).Test plan
dotnet test tests/ExcelReader.Tests— 2444 passedpytest python/testsagainst a freshly published native library — 123 passed, 1 xfailed (pre-existing)