Skip to content

text: Avoid remeasuring scroll-table column widths - #3318

Open
lurenjia534 wants to merge 1 commit into
longbridge:mainfrom
lurenjia534:perf/table-column-width-cache
Open

lurenjia534 wants to merge 1 commit into
longbridge:mainfrom
lurenjia534:perf/table-column-width-cache

Conversation

@lurenjia534

@lurenjia534 lurenjia534 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Description

Scroll-layout Markdown/HTML tables currently remeasure ordinary-cell column widths on repeated draws. Cache their per-column text maxima on the parsed table and reuse them while the window text system and typography key remain unchanged. Enter the scroll-layout path before the wrap-layout text-length scan.

Clones share the current immutable cache. Newly parsed tables start empty. Custom/resource-backed cells are measured on every call, outside the cache lock, and their widths are not retained so they can shrink as well as grow. Initial measurement still scans all cells; cell layout and painting still run.

This PR addresses repeated column measurement only. The diff is confined to crates/base/src/text/node.rs, including regression tests. It adds no public API or first-layout budget.

Performance and memory

Native Base/TextView fixture, Linux Wayland / Ryzen 9 7950X / RX 7900 XTX (RADV Mesa 26.2.1), 800×600 logical window, scale 1.75, fixed 170 Hz display. Release builds, normal multicore scheduling, 120 warmup presents followed by 600 measured presents per process, alternating A/B and B/A. There are 68 valid processes and 40,800 measured presents.

Scenario Baseline present FPS Patched present FPS 95% interval for FPS change Higher-FPS pairs
48×6 vertical page scroll 169.986 169.993 [-0.049, +0.045] 3/8
200×8 static redraw 87.146 88.058 [-10.962, +19.994] 7/10
200×8 vertical page scroll 84.525 92.797 [+3.263, +18.331] 7/8
48×6 wrap-layout control 169.977 170.000 [-0.016, +0.233] 5/8

Large scrolling improves by about 9.8% present FPS. Its median CPU draw cost falls from 9.754 to 8.836 ms/frame, saving 0.918 ms (9.4%; 95% interval [0.533, 4.149] ms). Present-interval P95/P99 falls from 16.105/18.582 to 14.662/17.176 ms. Native static-redraw results remain inconclusive, and its P95 interval point estimate worsens; the report retains those results. Medium/control FPS is refresh-capped.

FPS is calculated from actual first/last GPUI platform-submission timestamps, not inverse CPU draw duration. It is not compositor-confirmed physical scanout or full Story Gallery performance. Scrolling is driven through frame callbacks, not real wheel input. Estimates summarize independent process runs with paired bootstrap intervals; the report documents the small-sample and environment limitations.

A separate glibc allocation probe measures about 528 B extra per ordinary four-column table, including the Table field growth and allocator rounding: approximately 0.50 MiB for 1,000 tables and 10.07 MiB for 20,000. This is not the exact memory cost of the eight-column FPS fixture. Repeated measurement and 1,000 invalidation replacements support release of old caches; memory has no global quota, custom cells require extra coordinates, and the allocator can retain RSS after drop. No first-draw improvement is established.

English reports, frozen sources, raw samples, execution history, reproduction scripts, validation logs and checksums are preserved at an immutable commit on the fork's evidence/table-column-width-cache-20260929 branch, following #3273 and #3247. No evidence files are included in the code diff.

Known limitation and scope

Same-window dynamic-font invalidation is a known limitation outside the scope of this PR. The key compares window text-system identity, text style, rem size, mono family and inline-code style, but does not observe the generation changed by TextSystem::add_fonts. Font installation within an existing window may therefore leave widths stale. This PR does not address that case; any font-installation invalidation work will be handled separately. It is documented here so reviewers can assess the tradeoff.

The measurements support this optimization in some repeated-render workloads, but do not establish that the memory cost and invalidation complexity are the right tradeoff for the project. Long-running native memory stress and macOS/Windows measurements have not been performed.

How to Test

The branch is based on synchronized upstream 4f861e8a0654064fa53255cf8d9597bb9d5ef826. Performance remains measured on frozen baseline 11b04d9bf09cec68c98ce97d759afc6776b1cda0; the three intervening upstream commits did not change node.rs or the lockfile. The submitted node.rs is byte-identical to the measured patched snapshot. Performance was not rerun after synchronization; these functional checks were:

cargo test --offline --locked -p gpui-base --lib table_column -- --test-threads=1
cargo test --offline --locked -p gpui-base --lib
cargo fmt -p gpui-base -- --check
cargo clippy --offline --locked -p gpui-base --lib --tests -- -D warnings
git diff --check

The five focused tests and all 1,247 Base library tests pass. Formatting, Base lib/tests Clippy and the code whitespace check pass. Tests cover plain widths and clone sharing, window/typography invalidation, reparsing, inline code and dynamic custom-cell growth/shrinkage. Clippy --all-targets is not claimed; an unrelated existing Criterion benchmark dependency conflict is documented in the evidence.

AI Assistance

AI assisted with implementation, code review, regression validation, performance and memory measurements, analysis, documentation and PR preparation.

Checklist

  • Read the contributing guide and kept the code PR focused on one problem.
  • Reviewed the AI-assisted code and recorded evidence and unresolved limitations.
  • Passed cargo run for related Story tests (native TextView fixture measured; full Story Gallery not run).
  • Tested macOS, Windows and Linux performance (Linux measurements completed; macOS/Windows not measured).

@lurenjia534

lurenjia534 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

@huacnlee I'm still quite unsure whether this is the right tradeoff for the project. Our measurements show some benefit from avoiding repeated table-column measurement—about 9.8% higher present FPS in one Linux native scrolling fixture—but the improvement is not consistent across all scenarios. It also adds a modest memory cost, around 528 bytes per ordinary four-column table in our tests, plus cache invalidation complexity. Same-window dynamic-font invalidation remains a known limitation outside the scope of this PR, and is documented in the report. The full measurements and limitations are linked in the PR. Please point out any problems or concerns directly. I'm completely comfortable with this PR being closed if this isn't a direction you'd like the project to take; I'd appreciate your feedback on whether this approach is worth pursuing.

@lurenjia534
lurenjia534 marked this pull request as ready for review September 30, 2026 13:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant