text: Avoid remeasuring scroll-table column widths - #3318
lurenjia534 wants to merge 1 commit into
Conversation
|
@huacnlee I'm still quite unsure whether this is the right tradeoff for the project. Our measurements show some benefit from avoiding repeated table-column measurement—about 9.8% higher present FPS in one Linux native scrolling fixture—but the improvement is not consistent across all scenarios. It also adds a modest memory cost, around 528 bytes per ordinary four-column table in our tests, plus cache invalidation complexity. Same-window dynamic-font invalidation remains a known limitation outside the scope of this PR, and is documented in the report. The full measurements and limitations are linked in the PR. Please point out any problems or concerns directly. I'm completely comfortable with this PR being closed if this isn't a direction you'd like the project to take; I'd appreciate your feedback on whether this approach is worth pursuing. |
Description
Scroll-layout Markdown/HTML tables currently remeasure ordinary-cell column widths on repeated draws. Cache their per-column text maxima on the parsed table and reuse them while the window text system and typography key remain unchanged. Enter the scroll-layout path before the wrap-layout text-length scan.
Clones share the current immutable cache. Newly parsed tables start empty. Custom/resource-backed cells are measured on every call, outside the cache lock, and their widths are not retained so they can shrink as well as grow. Initial measurement still scans all cells; cell layout and painting still run.
This PR addresses repeated column measurement only. The diff is confined to
crates/base/src/text/node.rs, including regression tests. It adds no public API or first-layout budget.Performance and memory
Native Base/TextView fixture, Linux Wayland / Ryzen 9 7950X / RX 7900 XTX (RADV Mesa 26.2.1), 800×600 logical window, scale 1.75, fixed 170 Hz display. Release builds, normal multicore scheduling, 120 warmup presents followed by 600 measured presents per process, alternating A/B and B/A. There are 68 valid processes and 40,800 measured presents.
Large scrolling improves by about 9.8% present FPS. Its median CPU draw cost falls from 9.754 to 8.836 ms/frame, saving 0.918 ms (9.4%; 95% interval [0.533, 4.149] ms). Present-interval P95/P99 falls from 16.105/18.582 to 14.662/17.176 ms. Native static-redraw results remain inconclusive, and its P95 interval point estimate worsens; the report retains those results. Medium/control FPS is refresh-capped.
FPS is calculated from actual first/last GPUI platform-submission timestamps, not inverse CPU draw duration. It is not compositor-confirmed physical scanout or full Story Gallery performance. Scrolling is driven through frame callbacks, not real wheel input. Estimates summarize independent process runs with paired bootstrap intervals; the report documents the small-sample and environment limitations.
A separate glibc allocation probe measures about 528 B extra per ordinary four-column table, including the
Tablefield growth and allocator rounding: approximately 0.50 MiB for 1,000 tables and 10.07 MiB for 20,000. This is not the exact memory cost of the eight-column FPS fixture. Repeated measurement and 1,000 invalidation replacements support release of old caches; memory has no global quota, custom cells require extra coordinates, and the allocator can retain RSS after drop. No first-draw improvement is established.English reports, frozen sources, raw samples, execution history, reproduction scripts, validation logs and checksums are preserved at an immutable commit on the fork's
evidence/table-column-width-cache-20260929branch, following #3273 and #3247. No evidence files are included in the code diff.Known limitation and scope
Same-window dynamic-font invalidation is a known limitation outside the scope of this PR. The key compares window text-system identity, text style, rem size, mono family and inline-code style, but does not observe the generation changed by
TextSystem::add_fonts. Font installation within an existing window may therefore leave widths stale. This PR does not address that case; any font-installation invalidation work will be handled separately. It is documented here so reviewers can assess the tradeoff.The measurements support this optimization in some repeated-render workloads, but do not establish that the memory cost and invalidation complexity are the right tradeoff for the project. Long-running native memory stress and macOS/Windows measurements have not been performed.
How to Test
The branch is based on synchronized upstream
4f861e8a0654064fa53255cf8d9597bb9d5ef826. Performance remains measured on frozen baseline11b04d9bf09cec68c98ce97d759afc6776b1cda0; the three intervening upstream commits did not changenode.rsor the lockfile. The submittednode.rsis byte-identical to the measured patched snapshot. Performance was not rerun after synchronization; these functional checks were:The five focused tests and all 1,247 Base library tests pass. Formatting, Base lib/tests Clippy and the code whitespace check pass. Tests cover plain widths and clone sharing, window/typography invalidation, reparsing, inline code and dynamic custom-cell growth/shrinkage.
Clippy --all-targetsis not claimed; an unrelated existing Criterion benchmark dependency conflict is documented in the evidence.AI Assistance
AI assisted with implementation, code review, regression validation, performance and memory measurements, analysis, documentation and PR preparation.
Checklist
cargo runfor related Story tests (native TextView fixture measured; full Story Gallery not run).