diff --git a/.claude/board/entries/2026-10-07-aperture16-u64-word-schedule.md b/.claude/board/entries/2026-10-07-aperture16-u64-word-schedule.md new file mode 100644 index 000000000..9a7d7075f --- /dev/null +++ b/.claude/board/entries/2026-10-07-aperture16-u64-word-schedule.md @@ -0,0 +1,33 @@ +# 2026-10-07 — The subtraction law holds; its unit is the u64 word, not the u16 cell + +## MEASURED + +`crates/lance-graph-benches/examples/aperture16_probe.rs`, over a 65 536-self universe: +- 11 densities × 8 layouts × 8 payload widths × 10 routes, plus VIA, K propagation over CSR, extent × mask, and target-side scheduling; +- every route asserted equal before timing. + +- The `[u16; 4096]` view of the 64K mask is zero-copy: it is a shift of the same 8 KiB, endianness-proof, and checked against `align_to`. As a schedule, though, it is **1.3–2.3× slower per layout** (2.7× over all ladder cases) than u64 set-bit visitation, and 7× slower on an empty mask (4096 tests against 1024). +- Adding a full-word dense path to the u64 walk (`adapt64`): geomean 0.80, 2.4–2.9× on runs, at most 17 % worse on any layout. +- Full 16-run detection (`adapt64>16`) wins on 16-aligned islands (0.40) but is 1.25–1.64× slower on scattered layouts. A branch-free form does not remove that cost. +- Extent and mask compose multiplicatively: 2.9 µs, against 29.1 for extent only and 11.1 for mask only. The extent alone suffices only when the mask is dense. +- The mask schedules VIA and K propagation with no selection vector. The ordinal list shows no reproducible win; every apparent win reversed on re-run. +- On the target side, writing the exact next-frontier mask in the same pass is up to ~20× faster than clearing and scanning `K_next`, and the fastest route in 20 of 30 cases with run order rotated. The `[u16; 4096]` histogram pre-pass is never the fastest route (0 of 30). + +## CONSEQUENCE + +- No new carrier and no new V4 opcode. The mask is authoritative; u64 words are the unit. +- BUILD candidates, unbuilt: + - gated `mask_gather_u32` (the D-GATED-GATHER-0 item); + - a full-word dense path in the fold kernels' u64 walk; + - next-frontier-mask-in-pass for multi-hop K. +- KILL: u16 cells as the scheduling unit; popcount-threshold switching; `MultiplicityAperture`. + +## OPEN + +- Single machine, single thread, hand-rolled routes. Only A1's count goes through mask-risc. +- Host timing drifted by up to 2.4× on an identical pass, with no CPU steal. A4 uses a per-row minimum over 5 runs and A6 a median over 4 rotated runs; the other modes are single runs. +- Review corrections (#1377): cache-line counts use the real allocation offset; the 100 % extent case covers the whole universe; A6 rotates route order. +- Assembly not inspected; whether the dense paths auto-vectorise is unverified. +- `quack::SemanticAperture` already uses the word "aperture" for a different concept (a facet-key care mask). + +Report: `.claude/research/D-APERTURE-16-0.md`, `D-APERTURE-16-MATRIX.md`, `D-APERTURE-16-tables.md`. diff --git a/.claude/board/entries/README.md b/.claude/board/entries/README.md index 6efeaaeca..bff431359 100644 --- a/.claude/board/entries/README.md +++ b/.claude/board/entries/README.md @@ -25,13 +25,14 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately opposite directions; the stranding this convention prevents shows up in exactly one of them, never both. -240 entries, 2026-08-06 .. 2026-10-07. +241 entries, 2026-08-06 .. 2026-10-07. | date | entry id | finding | file | |---|---|---|---| | 2026-10-07 | `tinker-janus-fold-harvest` | TinkerPop bulk = K (GroupReduce), ONE_BULK = S, GValue pinning = bundle invalidation; JanusGraph slice = OrderedLaneWitness→Range (30–38×); no new V4 op | [2026-10-07-tinker-janus-fold-harvest.md](2026-10-07-tinker-janus-fold-harvest.md) | | 2026-10-07 | `gated-gather-bit-schedule` | Production gather is 3.8× slow from a per-row branch; bit-gated visitation wins or ties at every density; ordinal vectors never win (R1 holds) | [2026-10-07-gated-gather-bit-schedule.md](2026-10-07-gated-gather-bit-schedule.md) | | 2026-10-07 | `bind-bundle-gated-gather` | One bound query, five routes, identical answers; the semijoin is never gated and dominates; a gated mask gather beats a selection vector (R1 survives) | [2026-10-07-bind-bundle-gated-gather.md](2026-10-07-bind-bundle-gated-gather.md) | +| 2026-10-07 | `aperture16-u64-word-schedule` | The u16 view of the 64K mask is free but 1.3–2.3× slower per layout as a schedule than u64 words; extent × mask compose; the exact next-frontier mask beats a u16 target histogram; no new carrier or opcode | [2026-10-07-aperture16-u64-word-schedule.md](2026-10-07-aperture16-u64-word-schedule.md) | | 2026-10-06 | `v3-v4-dual-reading-round3` | | [2026-10-06-v3-v4-dual-reading-round3.md](2026-10-06-v3-v4-dual-reading-round3.md) | | 2026-10-06 | `D-ANCESTRY-K1-0` | | [2026-10-06-streamdto-circuit-ancestry.md](2026-10-06-streamdto-circuit-ancestry.md) | | 2026-10-06 | `D-SPOG-W-0` | | [2026-10-06-spog-witness-sub-context.md](2026-10-06-spog-witness-sub-context.md) | diff --git a/.claude/knowledge/fold-execution-laws.md b/.claude/knowledge/fold-execution-laws.md index 6825ecfca..bae459c43 100644 --- a/.claude/knowledge/fold-execution-laws.md +++ b/.claude/knowledge/fold-execution-laws.md @@ -51,6 +51,19 @@ to compute it now · EXECUTE touches the bytes. No fifth verb was forced. 9. **Anti-lasagne.** A new IR must hold information that cannot live in the semantic plan, the binding, V4, the bundle choice or the backend program. None proposed so far does. +10. **Subtract from the known universe, in u64 words.** Scheduling is + extent → skip zero words → dense full words → walk set bits → touch + payload last. The unit is the **u64 mask word**, not a 16-self cell. The + u16 view of the same 8 KiB is a free shift, but as a schedule it is 1.3–2.3× + slower per layout (2.7× overall geomean), and 7× slower on an empty mask (4096 tests vs 1024). 16-cells earn + a role only in detecting full 16-runs inside a partial word, opt-in. Cost + tracks live rows for lanes up to 16 B and live cache lines from 32 B up, never N. + [MEASURED, `aperture16_probe`, D-APERTURE-16-0] +11. **The target side subtracts too: the next frontier IS the schedule.** For + K propagation, writing the exact next-frontier mask in the same pass and + clearing `K_next` by walking it beats clearing and scanning the dense array + by up to ~20× at sparse frontiers. A coarse `[u16; 4096]` target histogram + pre-pass is never the fastest route. [MEASURED, `aperture16_probe a6`] ## Measured state of the candidate V4 (R2IL) diff --git a/.claude/research/D-APERTURE-16-0.md b/.claude/research/D-APERTURE-16-0.md new file mode 100644 index 000000000..c4e51d655 --- /dev/null +++ b/.claude/research/D-APERTURE-16-0.md @@ -0,0 +1,355 @@ +# D-APERTURE-16-0 — 4096 × u16 aperture scheduling over a known 64K self-space + +**Status:** MEASURED, 2026-10-07. Probe: `crates/lance-graph-benches/examples/aperture16_probe.rs` (modes `a1 a2 occ a3 a4 a5 a6`). +**Machine:** 4-vCPU Xeon @ 2.80 GHz, AVX-512F. Per core: L1d 32 KiB, L2 1 MiB; L3 33 MiB shared. Load average 0.06 at start; release build with `debug = 0`; median of 101 runs (31 for `a2`). +**Correctness:** every route's answer is asserted equal to a reference before its time is printed. Only answers that passed were recorded. + +## 0. Answer first + +**The law survives; the representation does not.** + +``` +known self-space +→ remove impossible regions (ordered extent → [lo, hi)) +→ remove closed mask regions (skip zero mask words) +→ visit only live ordinals (set-bit walk) +→ touch payload bytes last +``` + +Every part of that rule is measured and holds. Its best execution unit is the **u64 mask word, not the u16 cell**. As a scheduling unit, the u16 cell is 1.3–2.3× slower per layout (2.7× over all 464 ladder cases), and slower almost everywhere. It earns its place in exactly one narrow role: finding full 16-runs inside a partial word, on data that contains them. + +## 1. Current mask representation (read before building) + +| what | where | shape | +|---|---|---| +| support mask | `lance-graph-mask-risc` `Planes::masks`, `Scratch` slots | `&[u64]`, row `i` = bit `i & 63` of word `i >> 6`, little-endian bit order, tail bits zero | +| mask iteration in fold terminals | ndarray `simd_masking_ops.rs`: `masked_sum_i32` (:813), `group_walk` (:1616), `mask_scatter_or_u32` (:1205) | **already** skip a zero word in one test, then walk set bits with `trailing_zeros`: the `u64-visit` route here, transcribed | +| `mask_gather_u32` | ndarray `simd_masking_ops.rs:1099` | full scan of `index`, a data-dependent branch per row, no gate (unchanged since D-GATED-GATHER-0) | +| `MaskOp::Gather` | mask-risc `ir.rs:269`, `exec.rs:1753` | `{ lane, foreign, dst }`; **no `under`**, unlike `MaskOp::Pred` (`ir.rs:227`) | +| `Pred::Range` / `execute_extent` | `ir.rs:180`, `exec.rs:1393`; tiles from `extent_tiles` (`exec.rs:1246`) | work proportional to the extent, edge words clipped by `clip` | +| `GroupReduce` / Via | `exec.rs:1920`: `masked_group_*_u32{,_via}` | all route through `group_walk`, so they are u64 set-bit walks | + +So production already runs the u64-word schedule everywhere except `Gather`. This lab asks whether 16-self granularity would improve on it. + +**Name collision.** `lance_graph_quack::SemanticAperture` (`quack/src/lib.rs:925`) already means a facet-key care mask, lowered to a bound. The execution idea measured here needs a different name, or the word comes to mean two things. This document uses "aperture" only for the brief's 16-self cell. + +## 2. The zero-copy u16 view + +`[u64; 1024]` and `[u16; 4096]` are the same 8 KiB. The probe reads cell `c` as `(words[c >> 2] >> ((c & 3) * 16)) as u16`. That is a shift of an existing word, with no allocation and no cast, and it gives the same bits on any endianness. On little-endian it is also exactly `align_to::()`; `views_agree()` checks all 4096 cells at start-up. **So the view is free, and its memory footprint is identical.** Any win therefore has to come from execution geometry (§7 of the brief), and no new type is needed to express it. + +## 3. Configurations + +- Universe: N = 65 536 selves. Exactly `round(d·N)` live selves per pattern. +- Densities: 100 / 75 / 50 / 25 / 10 / 5 / 1 / 0.1 / 0.01 / 0.001 / 0.0001 %. At N = 65 536 the last three are 7, 1 and 0 survivors, so they measure the empty-mask floor, not a density. +- Layouts, never averaged together: + - uniform; + - clustered (runs of 128–640); + - one contiguous range; + - tiny runs (2–4); + - fully dense 16-islands in a sparse universe; + - alternating bits; + - one survivor per 16-cell; + - one survivor per 64-bit word. +- Payload lanes: u8, u16, u32, u64, plus records of 16/32/64/128 bytes (every byte folded). Working sets run from 64 KiB to 8 MiB, which crosses L1, L2 and L3. +- Routes: + + | route | what it does | + |---|---| + | `scan-branch` | per row, branch on its bit | + | `full-branchless` | load every row, AND with its bit | + | `u64-visit` | production shape | + | `u16-visit` | test each u16 cell | + | `u64→u16` | u64 test, then u16 quarters | + | `adapt16` | per cell: skip / dense 16 / set-bit | + | `adapt64` | per word: skip / dense 64 / set-bit | + | `adapt64>16` | u64 skip, dense full words, full u16 quarters dense, otherwise set-bit | + | `adapt64q` | branch-free full-quarter detection, then one set-bit loop | + | `ordinal-list` | materialise `Vec`, then fold; the build is timed and its bytes reported | + +## 4. Correctness parity + +Every route's answer is asserted against a reference before timing, for every pattern × payload × route in every mode, and the run aborts on a mismatch. All runs completed: 348 + 4 640 + 85 + 288 + 125 + 192 measured rows, plus 4 × 120 for A6. The u16 view also agrees with `align_to::` on all 4096 cells (§2). + +The A6 checksum covers only the next-frontier mask, so it cannot see whether `K_next` was cleared. A route that skipped clearing would pass and time faster. The three schedule-driven A6 routes therefore also assert that `K_next` is all zero afterwards, and all three pass. + +Disable runs, each verified red after committing, with the patch anchor asserted: + +| disable | fails at | +|---|---| +| D1: stop clearing `K_next` in the target-mask route | the clean-`K_next` assert | +| D2: drop bit 15 of every u16 cell in `visit16` | the A2 parity assert | +| D3: run the `adapt64` dense path over 63 of 64 rows | the A2 parity assert | + +## 5. Density crossover — visitation only (A1) + +Times in µs. Full table: `.claude/research/D-APERTURE-16-tables.md` §A1. + +| layout | density | u64 visit | u16 visit | u64→u16 | per-row scan | ordinal list | +|---|---|---|---|---|---|---| +| uniform | 100 % | 46.3 | 47.2 | 54.5 | 53.2 | 56.2 | +| uniform | 50 % | 25.1 | 49.9 | 35.6 | 53.2 | 29.8 | +| uniform | 10 % | 4.7 | 21.9 | 7.5 | 53.3 | 6.3 | +| uniform | 1 % | 0.9 | 3.9 | 2.2 | 53.2 | 1.5 | +| uniform | 0 live | **0.5** | **3.6** | 0.7 | 53.2 | 0.5 | +| clustered | 50 % | 23.0 | 26.3 | 28.2 | 53.2 | 28.3 | +| islands16 | 50 % | 23.4 | 37.4 | 30.3 | 53.2 | 28.6 | + +- **The empty-mask floor is 4096 vs 1024 tests: 3.6 vs 0.5 µs, a factor of 7.** The u16 cell visit has to test four times as many units, and nothing it saves offsets that. +- At mid density the u16 visit is 2–5× slower on scattered data: each cell's set-bit loop exits four times as often, and those exits are mispredicted. +- At 100 % the two routes converge, because the work is all set-bit visiting. +- The u64 → u16 hierarchy nearly recovers the floor (0.7 µs) but never beats the plain u64 walk. + +## 6. Local occupancy crossover (`occ`) + +Every one of the 4096 cells holds exactly k live bits. Times in µs (full table §OCC). + +| payload | k = 1 | 4 | 8 | 10 | 12 | 15 | 16 | +|---|---|---|---|---|---|---|---| +| u32, u16 set-bit | 10.8 | 15.0 | 28.6 | 36.0 | 44.1 | 87.1 | 61.7 | +| u32, dense masked 16 | 46.4 | 34.1 | 34.1 | 34.1 | 34.1 | 43.0 | 34.5 | +| u32, u64 set-bit | 4.9 | 11.0 | 21.5 | 26.9 | 32.1 | 68.2 | 43.5 | +| u32, adapt16 (dense only at k = 16) | 13.4 | 17.4 | 31.5 | 39.1 | 81.1 | 61.5 | **10.8** | + +- A dense masked 16-lane loop is flat in k, at about 33–34 µs for u8/u32/u64. It overtakes a u16 set-bit walk at k ≥ 10, and a u64 set-bit walk only at k ≥ 13. +- For 32-byte records it never wins. For 128-byte records the loop is memory-bound and the rows are noise. +- **The one large local win is the unmasked dense path at k = 16:** 10.5–11.6 µs against 42–67 µs for set-bit visiting, about 4–6×. +- So local density can choose an execution mode without changing the carrier. But only `cell == 0xFFFF` is a reliable switch; a popcount threshold is a narrow, payload-dependent band and not worth a branch. + +## 7. Payload-width ladder (A2) + +Times in µs. Geometric mean over all 464 pattern × payload cases, relative to `u64-visit`: + +| route | geomean vs u64 visit | best | worst | +|---|---|---|---| +| u16-visit | 2.70 | 0.88 | 14.4 | +| u64→u16-visit | 1.34 | 0.56 | 2.80 | +| adapt16 | 2.01 | 0.17 | 13.6 | +| **adapt64** | **0.80** | 0.05 | 2.57 | +| adapt64>16 | **0.71** | 0.04 | 2.71 | +| adapt64q | 0.81 | 0.13 | 3.01 | +| ordinal-list | 1.30 | 0.34 | 4.12 | +| scan-branch | 14.3 | 0.80 | 218 | +| full-branchless | 17.5 | 0.77 | 711 | + +By layout, for cases where `u64-visit` takes ≥ 2 µs (below that the numbers sit at the timer floor): + +| layout | adapt64 | adapt64>16 | adapt64q | u16 visit | +|---|---|---|---|---| +| uniform | **0.87** | 0.97 | 1.04 | 1.60 | +| tiny runs | **0.86** | 0.99 | 1.08 | 1.67 | +| clustered | 0.41 | **0.35** | 0.41 | 1.55 | +| single range | 0.34 | **0.32** | 0.40 | 1.53 | +| islands16 | 0.82 | **0.40** | 0.47 | 1.56 | +| alternating | **1.11** | 1.25 | 1.18 | 1.30 | +| one-per-16 | **0.98** | 1.59 | 1.85 | 1.53 | +| one-per-64 | **1.17** | 1.64 | 2.08 | 2.29 | + +What the ladder says: + +- **A u16 granule helps only when there are full 16-runs that do not fill 64.** That is islands16 (0.40 vs 0.82) and partly clustered data. Everywhere else it costs. +- **The branch-free variant (`adapt64q`) does not remove that cost.** Searching for runs is O(live words) of comparisons, and on scattered data, which has no runs, the search is pure overhead: up to ~2× per layout, 3.0× in the worst single case. +- `adapt64`, which only adds a dense path for a full word, has the best risk profile. It is never much worse than `u64-visit` (1.17 at worst per layout, on one survivor per word) and is 2.4–2.9× better on runs. + +**Does the benefit grow with payload width?** It depends on density. Ratio `full-branchless / u64-visit`, uniform layout: + +| density | u8 | u32 | u64 | rec32 | rec128 | +|---|---|---|---|---|---| +| 100 % | 1.1 | 1.5 | 1.2 | 1.0 | 1.1 | +| 10 % | 10.7 | 8.8 | 15.4 | 8.6 | 6.3 | +| 1 % | 45 | 45 | 49 | 53 | **115** | +| 0.1 % | 75 | 74 | 44 | 154 | **469** | + +- **At ≤ 1 % the benefit grows with width:** saved bytes scale with W, while the skip costs a fixed 1024 tests. +- **At ≥ 10 % it does not.** Uniform survivors leave almost every 64-byte line live for lanes under 64 B (at 10 % density a 16-row u32 line is live with probability 1 − 0.9¹⁶ ≈ 81 %). Few bytes are saved, and the remaining win is compute. +- **Runtime tracks work actually left, not N** (correlation of `u64-visit` time with what remains): + + | payload | r with live rows | r with live payload lines | + |---|---|---| + | u8 | 0.96 | 0.74 | + | u16 | 0.95 | 0.81 | + | u32 | 0.95 | 0.82 | + | u64 | 0.96 | 0.95 | + | rec16 | 0.97 | 0.91 | + | rec32 | 0.97 | **0.98** | + | rec64 | 0.95 | **0.97** | + | rec128 | 0.93 | **0.96** | + + Line counts use the payload's real address modulo 64: a `Vec

` is only aligned to `align_of::

()`, so a record can straddle one more line than an aligned layout would. For narrow lanes, rows predict runtime better. From 32-byte records upward, live lines predict it better. That is the refinement "self first, bytes later" needs: **the bytes that matter are the touched lines, and a lane narrower than a line saves rows, not bytes.** + +## 8. VIA (A3) + +Routes: full scan, u64 gate, u16 gate, ordinal. `via[self] → target: u16`. Support (`target_mask |= 1`) and multiplicity (`K_next[target] += K[self]`) are timed separately. Each run includes clearing the target (8 KiB for support, 256 KiB for `K_next`). + +| target | src density | support: full / u64 / u16 / ord (µs) | multiplicity: full / u64 / u16 / ord (µs) | +|---|---|---|---| +| uniform | 100 % | 109 / **80** / 85 / 138 | 180 / **131** / 136 / 234 | +| uniform | 50 % | 106 / **45** / 66 / 74 | 431 / **76** / 88 / 113 | +| uniform | 1 % | 194 / **3.8** / 6.7 / 3.5 | 154 / **19.8** / 23.7 / 20.7 | +| uniform | 0.01 % | 106 / 1.8 / 5.7 / 1.6 | 147 / 18.5 / 22.8 / 18.4 | +| hotspot | 100 % | **112** / 137 / 137 / 165 | 219 / **111** / 117 / 180 | +| many-to-one | 50 % | 136 / **76** / 111 / 154 | 433 / **114** / 136 / 165 | + +- **The mask can schedule VIA with no selection vector.** The u64 gate wins or ties every case except one family: support at 100 % source density into concentrated targets. There a branch-free full scan wins: hotspot 112 vs 137 µs, zipf 123 vs 152, many-to-one 136 vs 158. Concentrated scattered writes defeat the visit loop, and at 100 % there is nothing to skip. +- The u16 gate is never better than the u64 gate. +- **Support and multiplicity behave differently at the floor.** At low density, support costs ~1.8 µs (clearing 8 KiB), but multiplicity costs ~18.5 µs regardless of the route, because clearing the dense 256 KiB `K_next` dominates. §10 removes that cost. + +## 9. K propagation over CSR (A5) + +`K_next(v) = Σ K(u)` for edges `u → v`, with fan-out 1, 4 or 16 (up to 1 M edges). Times in µs; full table §A5. + +| target | fanout | src density | edge scan | src u64 gate | src u16 gate | ordinal | +|---|---|---|---|---|---|---| +| uniform | 1 | 100 % | 186 | **182** | 187 | 242 | +| uniform | 4 | 100 % | 767 | **375** | 382 | 477 | +| uniform | 16 | 100 % | 2984 | 1288 | **1271** | 1376 | +| uniform | 16 | 10 % | 2732 | **171** | 186 | 183 | +| uniform | 16 | 0.1 % | 2591 | **20.3** | 24.8 | 21.8 | +| hotspot | 16 | 100 % | 2857 | 1890 | 1939 | **980** | + +- **Source gating beats the edge scan even at 100 % density** once fan-out ≥ 4: 2.0–2.8× for uniform, clustered and hotspot targets, but only 1.1–1.15× for many-to-one. It reads each source's K once and walks its CSR row contiguously, instead of testing a bit and re-reading `src_of_edge` per edge. +- **Not reproducible: the dense hotspot fan-out-16 row.** Its ordinal "win" (980 vs 1890 µs) reversed between runs: the first run gave 1711 vs 958. Concentrated writes into 64 hot targets make it unstable. Recorded as noise, not as a crossover. +- A clustered fan-out-4 "ordinal win" in the first run (372 vs 574 µs) also failed to reproduce: the second run gave 347 vs 300. +- At ≤ 10 % source density the u16 gate trails the u64 gate in every row. At 100 % the two are within run-to-run noise, and either can lead: uniform fan-out 16 gives 1271 vs 1288 µs, hotspot fan-out 1 gives 183 vs 213. + +## 10. Coarse target histogram — the multiplicity aperture (A6) + +The question: is a coarse target-side object worth an extra pass? Each route propagates K, clears `K_next` for the next round, and builds the next frontier mask (`K_next > 0`). + +**Single runs are not stable here.** The same route's time varies by up to 2.4× between runs. Both run position and the host timing drift described in §11 contribute, and this lab did not separate them. So A6 was run four times, rotating the route order (`A6_ROT=0..3`) so that every route took every position for every pattern. The figures below are the median over those four runs (each itself a median of 101). Times in µs; full table §A6. + +| target | src density | touched cells | exact + full consume | exact + cell-seen bitmap | u16 histogram pre-pass | **exact + target mask** | +|---|---|---|---|---|---|---| +| uniform | 100 % | 4096 | 244 | **209** | 244 | 216 | +| uniform | 10 % | 3279 | 138 | 71 | 100 | **36** | +| uniform | 1 % | 610 | 123 | 15.7 | 22.3 | **8.2** | +| uniform | 0.1 % | 66 | 121 | 7.0 | 14.7 | **6.3** | +| clustered | 10 % | 3310 | 162 | 70 | 132 | **37** | +| hotspot | 100 % | 3265 | 223 | 191 | 242 | **152** | +| hotspot | 0.1 % | 10 | 137 | **6.8** | 19.5 | 10.1 | +| zipf | 10 % | 36 | 153 | 31 | 67 | **25** | +| one-to-one | 100 % | 4096 | 276 | 288 | 275 | **266** | +| many-to-one | 100 % | 16 | 245 | 154 | 292 | **150** | + +- **The u16 histogram pre-pass is KILLED (F3).** It is the fastest route in 0 of 30 cases, and slower than a cell bitmap built during the same pass in 29 of 30. The exception is one-to-one at 100 % (275 vs 288 µs), where the exact target mask beat both (266 µs). A second pass to count contributions cannot be repaid by any work it lets the consumer skip. +- The coarse target-cell bitmap (`target >> 4`, one OR per contribution) does pay, up to ~20× over "clear and scan everything" at sparse frontiers. +- **But the smallest object that delivers the saving is the EXACT next-frontier mask, written in the same pass.** Here K(u) ≥ 1 for every live u, so `K_next > 0` iff `v` was touched. That mask is the result the next round needs anyway, and walking its set bits is also the schedule for clearing `K_next`. + - It is fastest in 20 of 30 cases, and about 2× ahead of the cell bitmap at 10 % density. + - The cell bitmap's 10 wins are almost all at ≤ 1 % density, by ≤ 1 µs. + - The exceptions are uniform 100 % (209 vs 216 µs) and hotspot 0.1 % (6.8 vs 10.1 µs). +- **Important:** the saving is not a coarse-granularity effect. It comes from not touching dead target ordinals when clearing and consuming, which is the same subtraction law applied on the target side. + +## 11. Bounded extent × mask (A4) + +An ordered i32 lane `ts = self / 4`, with a predicate `ts ∈ [a, b)`, `m(self)`, and a u32 payload. The window starts at `min(span/5, span − width)`, so "100 %" is the whole universe and every narrower window sits where it always did. "Mask only" walks u16 cells and tests `ts` per live self. + +Times in µs, each the per-row **minimum over 5 runs** (each run a median of 101). During this re-run the host's timing drifted: the identical full-universe pass ranged from 101.6 to 241 µs, with no CPU steal recorded. Drift only ever adds time, so the minimum is the right estimate. With it, the full-universe baseline is flat at 101.6–101.8 µs in every row. Full table §A4. + +| mask density | extent | extent selves | live in extent | full universe | extent only | mask only | extent + u16 | **extent + u64** | +|---|---|---|---|---|---|---|---|---| +| 100 % | 100 % | 65536 | 65536 | 101.6 | 58.0 | 83.5 | 50.6 | **42.9** | +| 100 % | 50 % | 32768 | 32768 | 101.7 | 29.1 | 70.2 | 25.4 | **21.9** | +| 10 % | 100 % | 65536 | 6554 | 101.6 | 58.0 | 12.2 | 10.0 | **5.5** | +| 10 % | 50 % | 32768 | 3321 | 101.6 | 29.1 | 11.1 | 5.1 | **2.9** | +| 1 % | 100 % | 65536 | 655 | 101.7 | 58.0 | 4.9 | 6.7 | **1.6** | +| 1 % | 10 % | 6552 | 65 | 101.8 | 5.9 | 4.7 | 0.7 | **0.2** | +| 0.1 % | 100 % | 65536 | 66 | 101.7 | 58.0 | 4.5 | 6.5 | **1.1** | + +**Extent and mask are complementary, and they compose multiplicatively.** "Extent + u64 words" beats both single schedulers whenever both remove work, for example 2.9 µs vs 29.1 (extent only) and 11.1 (mask only). + +- **F6 holds where the mask is dense.** At 100 % mask density the extent does almost all the work, and the mask adds little (21.9 vs 29.1 µs). +- At a whole-universe extent, the extent removes nothing (58 µs, a branch-free pass over every row), and the mask does all the work (1.1–5.5 µs). +- Extent + u16 cells is slower than extent + u64 words wherever the mask is sparse (6.7 vs 1.6 µs): the same 4× test count as §5. + +## 12. "Known work subtraction" accounting + +Every row in `a1`/`a2`/`a4` reports: + +- **A1:** `live`, zero/full cells and words, and mask tests issued. +- **A2:** payload loads and touched lines. +- **A4:** universe, extent selves, live-in-extent and closed cells in the extent. + +Removed work = universe − remaining, attributed in this order: + +1. outside the extent; +2. closed words (A1's `zero_words`); +3. dead bits inside open words. + +§7's correlations show runtime following *remaining* work (r ≈ 0.95–0.99), never N, which is constant. One correction to the model: **"closed regions" should be counted in u64 words, not u16 cells.** At 1 % uniform density, 3488 of 4096 cells are zero but only 542 of 1024 words are. The u16 view finds more closed regions, yet still loses (§5), because a closed cell is cheaper to skip as part of a word than as a cell of its own. + +## 13. Cache and SIMD geometry + +| lane | bytes per 16 selves | lines per 16 selves | selves per 64-byte line | +|---|---|---|---| +| u8 | 16 | ¼ | 64 | +| u16 | 32 | ½ | 32 | +| u32 | **64** | **1** (one AVX-512 register) | 16 | +| u64 | 128 | 2 | 8 | + +A 16-self cell is line-aligned only for u32 lanes. For u8 lanes, the unit that matches a line is the 64-row u64 word. + +Granularity does not change payload traffic. Every set-bit walk, u16 or u64, loads exactly the live rows, and therefore exactly the live lines. The `payload_lines` column is identical across them. So alignment alone buys nothing. Granularity only matters through the dense paths, where the full-cell path for u32 is one 64-byte line. + +I did not inspect the generated assembly. Whether the dense loops auto-vectorise is **unverified**; their 4–6× speed-up over set-bit walking (§6) is consistent with vectorisation but does not prove it. + +## 14. Falsifier outcomes + +| # | falsifier | outcome | +|---|---|---| +| F1 | u16 aperture consistently slower than u64 set-bit visitation | **TRIGGERED.** Visitation: 7× at the floor, 2–5× mid-density, ≈1× at 100 %. Ladder geomean 2.70×. Keep the law, execute in u64 words. | +| F2 | popcount mode switching costs more than it saves | **TRIGGERED for thresholds. NOT triggered for "word or quarter is full".** The `== MAX` dense switch wins 2.4–2.9× on runs, and costs at most 17 % on the worst layout (one survivor per word). The general popcount crossover is a narrow, payload-dependent band (§6). Use one set-bit path plus a full-word dense path. | +| F3 | coarse target histogram does not amortise | **TRIGGERED. Kill `MultiplicityAperture`.** It is never the fastest route (0 of 30, with run order rotated); the exact same-pass target mask and the cell bitmap both beat it (§10). | +| F4 | ordinal lists win materially and reproducibly | **NOT triggered.** Geomean 1.30× slower. Every apparent win failed to reproduce or was sub-µs. No earned exception. | +| F5 | payload size does not affect the benefit | **PARTLY TRIGGERED.** True at ≥ 10 % density (lines are all live). False at ≤ 1 % (115–469× for 128-byte records). Refinement: count touched lines, not bytes. | +| F6 | bounded extent already removes most of the work | **TRIGGERED only for dense masks.** Otherwise extent and mask compose (§11). | + +## 15. Final ruling + +**KEEP** + +- **The mask is the only carrier.** `[u64]` words are the schedule. Reading 16-bit cells is a free shift of the same bytes when a kernel wants it; there is no new type, no copy and no persisted view. +- **The subtraction law:** extent → zero words → set bits → payload. Measured on every operation tried: plain lane folds, VIA support, VIA multiplicity, CSR K propagation, and extent-plus-mask. +- **A full-word dense path** (`word == u64::MAX` → contiguous fold) inside set-bit walks. +- **The exact target mask produced in the same pass** as the schedule for the next round and for clearing `K_next`. +- **No selection vector**, which reconfirms the D-GATED-GATHER-0 ruling. + +**BUILD** (each as its own PR with parity tests; nothing here is built yet) + +- `mask_gather_u32` gated by a mask, walking u64 words and set bits. This is the D-GATED-GATHER-0 build, now with stronger support: VIA gating wins here too. +- A full-word dense fast path in the fold kernels' u64 walk (`group_walk`, `masked_sum_*`), to be validated first against the ndarray parity suite. +- An opt-in full-quarter (16-run) dense check, **only** if a workload with 16-aligned runs is measured to need it. On scattered layouts it costs 1.25–1.64×. +- For multi-hop K propagation: produce the next frontier mask in the same pass, and clear `K_next` by walking it. Not yet measured through mask-risc; this lab hand-rolled it. + +**KILL** + +- The u16 cell as the primary scheduling unit. +- `adapt16`, the per-cell skip/dense/set-bit policy (geomean 2.07×). +- Popcount-threshold mode switching. +- The `[u16; 4096]` coarse multiplicity histogram (`MultiplicityAperture`). +- Any new carrier type (`ApertureMask`, `SparseMask`, `SelectionVector`) and any new V4 opcode (`APERTURE`, `GATED_LOAD16`, `SPARSE_VIA`, `APERTURE_SUM`). + +## 16. The eight questions + +1. **Is `[u16; 4096]` a useful zero-copy aperture view of the 64K support mask?** It is zero-copy: a shift of the same 8 KiB, endianness-proof, and verified. As a *schedule* it is not useful. It is the wrong granularity to iterate, and only full 16-run detection uses it. +2. **Is 16-self granularity measurably better than 64-bit word scheduling?** No. It is 1.3–2.3× slower per layout (2.7× overall geomean). It wins only on data with full 16-runs that do not fill a word (islands16: 0.40 vs 0.82). +3. **Does the benefit increase with payload width?** Only at low density, ≤ 1 %, up to ~470×. At ≥ 10 % nearly every line is live and the ratio stays flat. Runtime tracks rows for lanes up to 16 B and lines from 32 B up. +4. **Can `mask(self)` reliably schedule LOAD and VIA without selection vectors?** Yes. The u64 gate wins or ties every reproducible case except one family: support at 100 % source density into concentrated targets (hotspot, zipf, many-to-one), where a branch-free full scan is 15–20 % faster. Nothing there is skippable, so this is not a selection-vector case either. +5. **Are bounded extent and aperture complementary?** Yes, and they multiply. The exception is a dense mask, where the extent alone suffices. +6. **Does a coarse target histogram help K / multiplicity propagation?** No; it is killed. What helps is the exact next-frontier mask written in the same pass, which removes the dense clearing and scanning (up to ~20×, fastest in 20 of 30 cases). +7. **Is any new carrier justified?** No. +8. **Is any new V4 opcode justified?** No. Every route here is a lowering of `LOAD` / `KEEP` / `VIA` / `SUM` / `GROUP_SUM`. The one real gap is executor-level: `Gather` has no `under`. That is a missing operand on an existing op, not a new opcode. + +## 17. Limits of this lab + +- Single machine, single thread, release build, hand-rolled routes. +- Only the "native count" in A1 runs through mask-risc. The rest are transcriptions of the production loop shape. +- N = 65 536 fits a u16 self. Larger universes were not tested, so the 4×-test argument against u16 cells is measured only here, though it can only grow with N. +- Sub-µs differences sit at the timer floor and are not interpreted. +- Timing noise: + - The host drifted by up to 2.4× on an identical pass during the review re-run, with no CPU steal (§11). + - A4 is a per-row minimum over 5 runs, and A6 a median over 4 rotated runs. + - A1, A2, A3 and A5 are single runs with a fixed route order. Their rulings rest on margins of 1.3× or more, and on geometric means over hundreds of cases. Individual rows can carry drift. +- Review corrections (PR #1377): + - cache-line counts now use the real allocation offset (§7); + - the 100 % extent case now covers the whole universe instead of its last 80 % (§11); + - A6 now rotates the route order (§10). +- No assembly inspection (§13). diff --git a/.claude/research/D-APERTURE-16-MATRIX.md b/.claude/research/D-APERTURE-16-MATRIX.md new file mode 100644 index 000000000..9fdbf1862 --- /dev/null +++ b/.claude/research/D-APERTURE-16-MATRIX.md @@ -0,0 +1,33 @@ +# D-APERTURE-16 — technique matrix + +Companion to `D-APERTURE-16-0.md`, measured 2026-10-07 by `aperture16_probe`. "Extra bytes" means bytes beyond the 8 KiB support mask that the technique allocates or writes. "Exact support" asks whether the technique answers which selves are live; "multiplicity" asks whether it carries counts. "Geomean" is relative to `u64` bit visit over the 464-case payload ladder. + +| Technique | Extra bytes | Exact support | Multiplicity | Measured | Verdict | +|---|---|---|---|---|---| +| u64 bit visit (skip zero word, walk set bits) | 0 | yes | via the fold it drives | baseline; production shape (`group_walk`) | **KEEP** — the schedule | +| u64 bit visit + full-word dense path | 0 | yes | via the fold | geomean 0.80; 2.4–2.9× on runs; worst layout +17 % | **KEEP / BUILD** into the fold kernels | +| u16 aperture bit visit | 0 (shift view) | yes | via the fold | geomean 2.70; empty-mask floor 7× slower | **KILL** as a scheduling unit | +| adaptive u16 aperture (per cell: skip / dense / set-bit) | 0 | yes | via the fold | geomean 2.01; 2–5× slower on scattered data | **KILL** | +| u64 → full u16 quarter dense (`adapt64>16`) | 0 | yes | via the fold | geomean 0.71; 0.40 on 16-islands, 1.25–1.64× slower on scattered layouts | **OPT-IN only**, behind a measured workload | +| branch-free full-quarter detect (`adapt64q`) | 0 | yes | via the fold | geomean 0.81; does not remove the scattered-data cost | **KILL** | +| popcount-threshold mode switch | 0 | yes | — | crossover is payload-dependent (dense wins only at k ≥ 13 of 16, never for 32-byte records) | **KILL** | +| ordinal list (`Vec`) | 2 · live | yes | via the fold | geomean 1.30; no reproducible win | **KILL** (no earned exception) | +| bounded extent (`partition_point → Cmp::Range`) composed with u64 words | 0 | yes | — | multiplies with the mask: 2.9 µs vs 29.1 (extent only) and 11.1 (mask only) | **KEEP**; needs an ordering witness for scalar lanes | +| coarse target bitmap (`target >> 4`, written during the exact pass) | 512 | no (cells only) | no | up to ~20× over clearing and scanning all of `K_next`; fastest in 10 of 30 cases, nearly all by ≤ 1 µs at ≤ 1 % density | **PARK**; superseded by the exact target mask | +| exact target mask written during the exact pass | 0 (it is the next frontier) | yes | no; it schedules clearing `K_next` | fastest in 20 of 30 cases (run order rotated), ~2× ahead of the cell bitmap at 10 % density; up to ~20× | **KEEP / BUILD** for multi-hop K propagation | +| u16 target histogram (`[u16; 4096]` pre-pass) | 8 192 | no | coarse counts, saturating | fastest in 0 of 30 cases; slower than the cell bitmap in 29 of 30, and the exception (one-to-one, 100 %) lost to the exact target mask | **KILL** (`MultiplicityAperture`) | +| full scan, branch-free | 0 | yes | yes | only for 100 %-density support into concentrated targets (15–20 %) | **KEEP** as the dense-end lowering | +| per-row branch on the mask bit | 0 | yes | yes | geomean 14.3; up to 218× slower | **KILL** | + +## The rule that survived + +```text +extent → [lo, hi) (ordered lane; binary search) +words → skip m[w] == 0 (one test per 64 selves) +dense → m[w] == u64::MAX (contiguous fold) +bits → walk set bits (trailing_zeros) +bytes → touch payload last (cost tracks live rows, or live lines from 32 B up) +target → write the next frontier mask in the same pass; it schedules the next round +``` + +No new carrier and no new V4 opcode. The u16 cell is a free shift of the same bytes, useful only for detecting full 16-runs. diff --git a/.claude/research/D-APERTURE-16-tables.md b/.claude/research/D-APERTURE-16-tables.md new file mode 100644 index 000000000..1620dfd13 --- /dev/null +++ b/.claude/research/D-APERTURE-16-tables.md @@ -0,0 +1,366 @@ +# D-APERTURE-16-0 — full measurement tables + +Generated from `aperture16_probe` output, 2026-10-07; see `D-APERTURE-16-0.md` for configuration and reading. Regenerated after review: A2 line counts use the real allocation offset, A4 keeps the requested extent inside the lane and takes a per-row minimum over 5 runs, A6 is a median over all four run positions. + +## A1 visitation (µs, median of 101) + +| layout | density | live | zero cells | full cells | zero words | full words | u64 visit | u16 visit | u64→u16 | per-row scan | ordinal list | u16/u64 | +|---|---|---|---|---|---|---|---|---|---|---|---|---| +| uniform | 1 | 65536 | 0 | 4096 | 0 | 1024 | 46.3 | 47.2 | 54.5 | 53.2 | 56.2 | 1.02 | +| uniform | 0.5 | 32768 | 0 | 0 | 0 | 0 | 25.1 | 49.9 | 35.6 | 53.2 | 29.8 | 1.99 | +| uniform | 0.1 | 6554 | 763 | 0 | 0 | 0 | 4.7 | 21.9 | 7.5 | 53.3 | 6.3 | 4.65 | +| uniform | 0.01 | 655 | 3488 | 0 | 542 | 0 | 0.9 | 3.9 | 2.2 | 53.2 | 1.5 | 4.46 | +| uniform | 0.001 | 66 | 4030 | 0 | 962 | 0 | 0.5 | 3.9 | 0.8 | 53.2 | 0.5 | 7.84 | +| clustered | 1 | 65536 | 0 | 4096 | 0 | 1024 | 46.3 | 47.2 | 54.5 | 53.2 | 56.2 | 1.02 | +| clustered | 0.5 | 32768 | 2003 | 1996 | 463 | 459 | 23.0 | 26.3 | 28.2 | 53.2 | 28.3 | 1.14 | +| clustered | 0.1 | 6554 | 3670 | 395 | 906 | 86 | 4.9 | 8.3 | 6.2 | 53.2 | 5.0 | 1.68 | +| clustered | 0.01 | 655 | 4054 | 38 | 1012 | 8 | 0.9 | 4.1 | 1.3 | 53.2 | 1.0 | 4.50 | +| clustered | 0.001 | 66 | 4091 | 4 | 1022 | 0 | 0.5 | 3.7 | 0.8 | 53.2 | 0.5 | 7.07 | +| single-range | 1 | 65536 | 0 | 4096 | 0 | 1024 | 46.3 | 47.2 | 65.1 | 53.3 | 56.2 | 1.02 | +| single-range | 0.5 | 32768 | 2047 | 2047 | 511 | 511 | 23.5 | 25.5 | 27.6 | 53.2 | 28.4 | 1.09 | +| single-range | 0.1 | 6554 | 3685 | 409 | 921 | 101 | 6.1 | 10.4 | 6.1 | 53.3 | 9.5 | 1.71 | +| single-range | 0.01 | 655 | 4054 | 40 | 1012 | 10 | 0.9 | 4.1 | 1.3 | 53.2 | 1.0 | 4.43 | +| single-range | 0.001 | 66 | 4090 | 4 | 1021 | 1 | 0.5 | 3.7 | 0.8 | 53.2 | 0.5 | 7.10 | +| tiny-runs | 1 | 65536 | 0 | 4096 | 0 | 1024 | 64.6 | 70.2 | 75.7 | 53.2 | 56.5 | 1.09 | +| tiny-runs | 0.5 | 32768 | 71 | 14 | 0 | 0 | 24.6 | 53.2 | 32.9 | 55.2 | 29.3 | 2.16 | +| tiny-runs | 0.1 | 6554 | 2179 | 0 | 104 | 0 | 4.7 | 23.7 | 7.2 | 53.2 | 6.3 | 5.02 | +| tiny-runs | 0.01 | 655 | 3850 | 0 | 815 | 0 | 0.8 | 3.9 | 1.6 | 77.4 | 1.3 | 4.64 | +| tiny-runs | 0.001 | 66 | 4071 | 0 | 1003 | 0 | 0.5 | 5.0 | 1.4 | 53.2 | 0.5 | 9.16 | +| islands16 | 1 | 65536 | 0 | 4096 | 0 | 1024 | 46.3 | 47.2 | 54.5 | 53.2 | 56.4 | 1.02 | +| islands16 | 0.5 | 32768 | 2048 | 2048 | 60 | 61 | 23.4 | 37.4 | 30.3 | 53.2 | 28.6 | 1.60 | +| islands16 | 0.1 | 6544 | 3687 | 409 | 672 | 0 | 8.4 | 12.8 | 11.4 | 111.7 | 11.7 | 1.52 | +| islands16 | 0.01 | 640 | 4056 | 40 | 984 | 0 | 1.4 | 6.7 | 2.5 | 111.7 | 2.1 | 4.74 | +| islands16 | 0.001 | 64 | 4092 | 4 | 1020 | 0 | 0.9 | 6.1 | 1.6 | 56.2 | 0.5 | 7.14 | +| alternating | 0.5 | 32768 | 0 | 0 | 0 | 0 | 23.2 | 24.2 | 27.3 | 53.2 | 28.1 | 1.04 | +| one-per-16 | 0.0625 | 4096 | 0 | 0 | 0 | 0 | 3.0 | 7.2 | 6.2 | 53.2 | 3.5 | 2.42 | +| one-per-64 | 0.015625 | 1024 | 3072 | 0 | 0 | 0 | 1.0 | 3.9 | 3.4 | 53.2 | 1.6 | 3.98 | + +## OCC local occupancy (µs; every one of the 4096 cells holds exactly k live bits) + +| payload | k | u16 set-bit | u16 dense masked | u64 set-bit | adapt16 | +|---|---|---|---|---|---| +| u8 | 0 | 3.0 | 3.1 | 0.5 | 3.0 | +| u8 | 1 | 5.3 | 33.9 | 3.0 | 8.0 | +| u8 | 2 | 8.3 | 33.9 | 5.4 | 13.3 | +| u8 | 4 | 14.7 | 33.9 | 10.7 | 18.4 | +| u8 | 8 | 28.5 | 33.9 | 21.4 | 33.9 | +| u8 | 10 | 35.9 | 34.0 | 26.7 | 42.5 | +| u8 | 12 | 43.9 | 34.0 | 32.0 | 51.7 | +| u8 | 13 | 47.8 | 34.0 | 34.7 | 56.2 | +| u8 | 15 | 55.7 | 34.0 | 42.1 | 65.4 | +| u8 | 16 | 72.8 | 46.5 | 67.0 | 11.6 | +| u32 | 0 | 3.8 | 4.9 | 1.3 | 6.8 | +| u32 | 1 | 10.8 | 46.4 | 4.9 | 13.4 | +| u32 | 2 | 14.0 | 41.6 | 5.6 | 13.3 | +| u32 | 4 | 15.0 | 34.1 | 11.0 | 17.4 | +| u32 | 8 | 28.6 | 34.1 | 21.5 | 31.5 | +| u32 | 10 | 36.0 | 34.1 | 26.9 | 39.1 | +| u32 | 12 | 44.1 | 34.1 | 32.1 | 81.1 | +| u32 | 13 | 90.8 | 45.9 | 65.6 | 105.3 | +| u32 | 15 | 87.1 | 43.0 | 68.2 | 61.5 | +| u32 | 16 | 61.7 | 34.5 | 43.5 | 10.8 | +| u64 | 0 | 3.0 | 3.0 | 0.5 | 3.0 | +| u64 | 1 | 5.5 | 32.9 | 3.7 | 7.8 | +| u64 | 2 | 8.8 | 32.9 | 5.5 | 11.0 | +| u64 | 4 | 15.3 | 32.9 | 10.8 | 18.0 | +| u64 | 8 | 29.6 | 32.9 | 21.5 | 32.4 | +| u64 | 10 | 36.6 | 37.0 | 26.8 | 39.6 | +| u64 | 12 | 45.4 | 32.9 | 31.7 | 47.7 | +| u64 | 13 | 48.6 | 32.9 | 34.2 | 52.0 | +| u64 | 15 | 57.1 | 32.9 | 39.2 | 60.4 | +| u64 | 16 | 61.1 | 32.9 | 41.7 | 10.5 | +| rec32 | 0 | 3.6 | 3.0 | 0.5 | 3.6 | +| rec32 | 1 | 11.1 | 113.1 | 6.5 | 13.2 | +| rec32 | 2 | 20.0 | 109.7 | 13.7 | 21.2 | +| rec32 | 4 | 43.9 | 109.5 | 33.6 | 43.9 | +| rec32 | 8 | 73.0 | 110.0 | 61.0 | 71.1 | +| rec32 | 10 | 81.8 | 109.9 | 66.6 | 83.6 | +| rec32 | 12 | 91.8 | 109.6 | 69.9 | 90.9 | +| rec32 | 13 | 96.1 | 109.5 | 73.4 | 95.9 | +| rec32 | 15 | 104.8 | 109.3 | 82.3 | 104.3 | +| rec32 | 16 | 108.8 | 122.2 | 83.2 | 57.6 | +| rec128 | 0 | 3.6 | 3.0 | 0.6 | 3.0 | +| rec128 | 1 | 21.1 | 281.8 | 16.3 | 22.8 | +| rec128 | 2 | 61.9 | 263.6 | 56.2 | 61.1 | +| rec128 | 4 | 139.1 | 276.2 | 121.0 | 123.2 | +| rec128 | 8 | 222.5 | 271.6 | 224.0 | 224.2 | +| rec128 | 10 | 381.8 | 327.2 | 303.2 | 271.5 | +| rec128 | 12 | 207.9 | 194.8 | 201.5 | 228.7 | +| rec128 | 13 | 373.0 | 300.3 | 375.1 | 345.5 | +| rec128 | 15 | 214.3 | 187.9 | 207.5 | 214.4 | +| rec128 | 16 | 220.2 | 184.4 | 205.8 | 183.9 | + +## A2 payload ladder (µs, median of 31; payload_lines counted from the real allocation offset) + +| layout | density | payload | live | live lines | full-branchless | u64-visit | u16-visit | adapt64 | adapt64>16 | adapt64q | ordinal-list | +|---|---|---|---|---|---|---|---|---|---|---|---| +| uniform | 1 | u8 | 65536 | 1024 | 48.7 | 43.0 | 61.7 | 2.1 | 1.8 | 5.6 | 115.7 | +| uniform | 0.5 | u8 | 32768 | 1024 | 84.3 | 37.5 | 62.4 | 41.1 | 55.5 | 51.6 | 74.3 | +| uniform | 0.1 | u8 | 6554 | 1024 | 48.7 | 4.5 | 11.9 | 5.0 | 6.8 | 8.7 | 8.2 | +| uniform | 0.01 | u8 | 655 | 482 | 48.7 | 1.1 | 3.3 | 1.1 | 2.2 | 2.0 | 1.6 | +| uniform | 0.001 | u8 | 66 | 62 | 48.8 | 0.7 | 3.2 | 0.6 | 0.8 | 0.7 | 0.5 | +| clustered | 1 | u8 | 65536 | 1024 | 48.6 | 42.7 | 60.3 | 2.1 | 6.9 | 5.6 | 81.0 | +| clustered | 0.5 | u8 | 32768 | 561 | 68.9 | 22.6 | 32.3 | 7.4 | 5.7 | 3.8 | 40.7 | +| clustered | 0.1 | u8 | 6554 | 118 | 48.7 | 5.0 | 9.0 | 1.4 | 1.0 | 1.2 | 7.6 | +| clustered | 0.01 | u8 | 655 | 12 | 48.7 | 1.0 | 3.6 | 0.6 | 0.7 | 0.6 | 1.1 | +| clustered | 0.001 | u8 | 66 | 2 | 48.7 | 0.7 | 3.1 | 0.6 | 0.6 | 0.5 | 0.5 | +| islands16 | 1 | u8 | 65536 | 1024 | 48.7 | 43.4 | 60.8 | 7.9 | 1.8 | 5.6 | 81.1 | +| islands16 | 0.5 | u8 | 32768 | 964 | 48.7 | 28.1 | 38.8 | 22.9 | 3.7 | 4.3 | 41.0 | +| islands16 | 0.1 | u8 | 6544 | 352 | 48.7 | 4.8 | 9.4 | 5.6 | 1.7 | 1.6 | 7.8 | +| islands16 | 0.01 | u8 | 640 | 40 | 48.7 | 0.9 | 3.8 | 1.0 | 0.7 | 0.6 | 1.1 | +| islands16 | 0.001 | u8 | 64 | 4 | 48.7 | 0.7 | 3.1 | 0.6 | 0.7 | 0.6 | 0.4 | +| uniform | 1 | u32 | 65536 | 4096 | 108.1 | 74.5 | 113.3 | 9.2 | 9.6 | 16.5 | 176.3 | +| uniform | 0.5 | u32 | 32768 | 4096 | 48.5 | 30.2 | 51.3 | 30.0 | 40.2 | 38.1 | 41.9 | +| uniform | 0.1 | u32 | 6554 | 3333 | 64.4 | 7.3 | 16.7 | 12.2 | 7.0 | 9.9 | 9.2 | +| uniform | 0.01 | u32 | 655 | 608 | 48.5 | 1.1 | 3.3 | 1.2 | 2.0 | 2.0 | 1.5 | +| uniform | 0.001 | u32 | 66 | 66 | 48.5 | 0.7 | 4.0 | 0.9 | 0.5 | 0.7 | 0.8 | +| clustered | 1 | u32 | 65536 | 4096 | 48.5 | 42.8 | 61.2 | 5.2 | 5.2 | 8.5 | 81.0 | +| clustered | 0.5 | u32 | 32768 | 2093 | 48.5 | 22.5 | 32.9 | 14.3 | 3.7 | 5.4 | 40.6 | +| clustered | 0.1 | u32 | 6554 | 426 | 48.5 | 4.9 | 9.1 | 1.6 | 3.2 | 4.7 | 7.6 | +| clustered | 0.01 | u32 | 655 | 42 | 48.4 | 1.0 | 3.7 | 1.0 | 0.4 | 0.6 | 1.9 | +| clustered | 0.001 | u32 | 66 | 5 | 59.0 | 0.7 | 3.1 | 0.6 | 1.0 | 0.8 | 0.5 | +| islands16 | 1 | u32 | 65536 | 4096 | 53.6 | 70.0 | 61.4 | 5.2 | 9.3 | 16.0 | 81.0 | +| islands16 | 0.5 | u32 | 32768 | 2048 | 48.4 | 25.3 | 38.4 | 26.0 | 5.1 | 6.1 | 41.1 | +| islands16 | 0.1 | u32 | 6544 | 409 | 48.5 | 4.9 | 9.6 | 5.9 | 6.2 | 7.9 | 8.1 | +| islands16 | 0.01 | u32 | 640 | 40 | 48.4 | 0.9 | 3.8 | 0.9 | 2.0 | 1.7 | 1.1 | +| islands16 | 0.001 | u32 | 64 | 4 | 48.5 | 0.7 | 3.1 | 0.6 | 0.4 | 0.6 | 0.4 | +| uniform | 1 | u64 | 65536 | 8193 | 48.3 | 41.7 | 61.1 | 6.9 | 6.9 | 7.6 | 84.3 | +| uniform | 0.5 | u64 | 32768 | 8166 | 410.8 | 49.7 | 66.9 | 50.7 | 40.5 | 36.1 | 92.8 | +| uniform | 0.1 | u64 | 6554 | 4643 | 145.0 | 9.4 | 16.6 | 9.7 | 14.6 | 16.7 | 18.3 | +| uniform | 0.01 | u64 | 655 | 633 | 53.7 | 1.1 | 3.3 | 1.2 | 1.8 | 2.0 | 1.8 | +| uniform | 0.001 | u64 | 66 | 65 | 49.3 | 1.1 | 4.3 | 1.1 | 0.5 | 0.7 | 1.0 | +| clustered | 1 | u64 | 65536 | 8193 | 48.3 | 41.7 | 61.3 | 6.9 | 6.9 | 7.6 | 84.3 | +| clustered | 0.5 | u64 | 32768 | 4148 | 49.5 | 22.3 | 32.9 | 7.2 | 4.6 | 5.0 | 42.2 | +| clustered | 0.1 | u64 | 6554 | 834 | 48.6 | 5.1 | 9.1 | 1.8 | 2.3 | 4.0 | 8.3 | +| clustered | 0.01 | u64 | 655 | 83 | 48.3 | 1.1 | 3.7 | 0.7 | 0.6 | 0.9 | 1.1 | +| clustered | 0.001 | u64 | 66 | 9 | 48.2 | 0.7 | 3.1 | 0.7 | 0.4 | 0.6 | 0.4 | +| islands16 | 1 | u64 | 65536 | 8193 | 48.3 | 41.7 | 61.3 | 6.9 | 6.9 | 7.6 | 84.3 | +| islands16 | 0.5 | u64 | 32768 | 5130 | 65.5 | 25.6 | 38.0 | 28.1 | 5.0 | 5.8 | 44.6 | +| islands16 | 0.1 | u64 | 6544 | 1195 | 53.0 | 5.2 | 9.6 | 5.0 | 2.7 | 4.3 | 9.1 | +| islands16 | 0.01 | u64 | 640 | 120 | 48.7 | 0.9 | 3.8 | 1.0 | 0.5 | 0.7 | 1.1 | +| islands16 | 0.001 | u64 | 64 | 12 | 48.3 | 0.7 | 3.1 | 0.7 | 0.4 | 0.7 | 0.4 | +| uniform | 1 | rec32 | 65536 | 32769 | 85.6 | 82.4 | 109.0 | 57.3 | 57.4 | 56.1 | 130.6 | +| uniform | 0.5 | rec32 | 32768 | 28661 | 85.8 | 59.8 | 81.9 | 60.5 | 67.6 | 62.8 | 85.0 | +| uniform | 0.1 | rec32 | 6554 | 8756 | 85.7 | 10.0 | 15.9 | 12.5 | 12.6 | 15.1 | 13.6 | +| uniform | 0.01 | rec32 | 655 | 984 | 86.2 | 1.6 | 5.0 | 1.8 | 2.3 | 3.0 | 2.1 | +| uniform | 0.001 | rec32 | 66 | 95 | 85.8 | 0.6 | 4.0 | 0.7 | 0.7 | 0.8 | 0.5 | +| clustered | 1 | rec32 | 65536 | 32769 | 85.6 | 82.5 | 108.7 | 57.6 | 57.6 | 56.7 | 130.2 | +| clustered | 0.5 | rec32 | 32768 | 16440 | 85.6 | 41.4 | 56.7 | 19.0 | 17.2 | 16.5 | 64.1 | +| clustered | 0.1 | rec32 | 6554 | 3293 | 85.9 | 8.4 | 14.3 | 4.3 | 3.7 | 6.8 | 12.8 | +| clustered | 0.01 | rec32 | 655 | 330 | 85.7 | 1.1 | 4.7 | 0.9 | 0.8 | 0.8 | 2.4 | +| clustered | 0.001 | rec32 | 66 | 34 | 85.8 | 0.5 | 3.7 | 0.6 | 0.5 | 0.6 | 0.5 | +| islands16 | 1 | rec32 | 65536 | 32769 | 85.6 | 82.3 | 108.6 | 55.2 | 54.5 | 56.2 | 130.2 | +| islands16 | 0.5 | rec32 | 32768 | 17418 | 173.6 | 80.1 | 98.4 | 75.8 | 34.4 | 34.7 | 116.3 | +| islands16 | 0.1 | rec32 | 6544 | 3649 | 170.5 | 15.4 | 24.6 | 15.8 | 5.3 | 6.1 | 23.1 | +| islands16 | 0.01 | rec32 | 640 | 360 | 171.1 | 2.5 | 10.5 | 2.6 | 1.5 | 1.7 | 3.4 | +| islands16 | 0.001 | rec32 | 64 | 36 | 172.3 | 1.0 | 8.6 | 1.3 | 1.3 | 1.3 | 1.4 | +| uniform | 1 | rec128 | 65536 | 131073 | 380.1 | 341.7 | 333.1 | 353.2 | 293.7 | 291.1 | 410.4 | +| uniform | 0.5 | rec128 | 32768 | 81967 | 341.2 | 252.1 | 243.1 | 226.1 | 250.7 | 287.6 | 296.5 | +| uniform | 0.1 | rec128 | 6554 | 18979 | 349.5 | 55.9 | 78.5 | 68.0 | 68.6 | 63.7 | 75.0 | +| uniform | 0.01 | rec128 | 655 | 1962 | 275.3 | 2.4 | 5.6 | 2.6 | 3.3 | 3.6 | 3.7 | +| uniform | 0.001 | rec128 | 66 | 197 | 279.3 | 0.6 | 4.1 | 0.8 | 0.8 | 0.9 | 0.9 | +| clustered | 1 | rec128 | 65536 | 131073 | 286.6 | 268.6 | 278.8 | 299.0 | 259.7 | 262.8 | 341.0 | +| clustered | 0.5 | rec128 | 32768 | 65591 | 241.1 | 99.1 | 104.0 | 92.1 | 123.7 | 122.1 | 162.9 | +| clustered | 0.1 | rec128 | 6554 | 13124 | 357.7 | 20.6 | 27.6 | 17.8 | 12.5 | 15.0 | 21.5 | +| clustered | 0.01 | rec128 | 655 | 1312 | 395.1 | 3.3 | 11.6 | 2.9 | 1.7 | 2.1 | 4.5 | +| clustered | 0.001 | rec128 | 66 | 133 | 365.9 | 1.2 | 6.9 | 1.5 | 1.4 | 1.4 | 1.6 | +| islands16 | 1 | rec128 | 65536 | 131073 | 238.2 | 214.5 | 224.3 | 206.6 | 203.6 | 260.8 | 283.0 | +| islands16 | 0.5 | rec128 | 32768 | 66570 | 255.0 | 127.6 | 136.1 | 126.5 | 130.9 | 130.1 | 152.4 | +| islands16 | 0.1 | rec128 | 6544 | 13465 | 286.5 | 18.8 | 23.2 | 20.1 | 18.2 | 17.1 | 23.5 | +| islands16 | 0.01 | rec128 | 640 | 1320 | 279.3 | 1.9 | 5.4 | 2.1 | 2.0 | 2.0 | 2.4 | +| islands16 | 0.001 | rec128 | 64 | 132 | 293.0 | 0.5 | 3.8 | 0.8 | 0.6 | 0.7 | 0.8 | + +## A3 VIA (µs, median of 101; includes clearing the target) + +| target | src density | semantics | full scan | u64 gate | u16 gate | ordinal | +|---|---|---|---|---|---|---| +| uniform | 1 | support | 109.3 | 79.7 | 85.1 | 137.8 | +| uniform | 1 | multiplicity | 180.4 | 130.8 | 135.8 | 233.8 | +| uniform | 0.5 | support | 106.3 | 45.1 | 65.6 | 73.6 | +| uniform | 0.5 | multiplicity | 430.9 | 76.3 | 88.3 | 113.1 | +| uniform | 0.1 | support | 106.2 | 9.9 | 14.8 | 16.4 | +| uniform | 0.1 | multiplicity | 197.7 | 33.1 | 34.2 | 37.7 | +| uniform | 0.01 | support | 194.0 | 3.8 | 6.7 | 3.5 | +| uniform | 0.01 | multiplicity | 153.7 | 19.8 | 23.7 | 20.7 | +| uniform | 0.001 | support | 108.2 | 1.8 | 6.0 | 1.7 | +| uniform | 0.001 | multiplicity | 145.9 | 18.6 | 23.0 | 18.5 | +| uniform | 0.0001 | support | 106.0 | 1.8 | 5.7 | 1.6 | +| uniform | 0.0001 | multiplicity | 146.7 | 18.5 | 22.8 | 18.4 | +| clustered | 1 | support | 105.4 | 81.6 | 88.6 | 131.4 | +| clustered | 1 | multiplicity | 147.8 | 109.5 | 117.7 | 182.6 | +| clustered | 0.5 | support | 106.8 | 46.9 | 71.6 | 69.4 | +| clustered | 0.5 | multiplicity | 549.4 | 122.6 | 138.4 | 178.1 | +| clustered | 0.1 | support | 259.4 | 21.7 | 32.5 | 32.8 | +| clustered | 0.1 | multiplicity | 282.1 | 32.3 | 33.6 | 37.1 | +| clustered | 0.01 | support | 105.4 | 2.7 | 6.8 | 3.6 | +| clustered | 0.01 | multiplicity | 137.1 | 19.7 | 23.6 | 20.7 | +| clustered | 0.001 | support | 105.9 | 1.8 | 6.0 | 1.7 | +| clustered | 0.001 | multiplicity | 131.2 | 18.6 | 23.0 | 18.7 | +| clustered | 0.0001 | support | 105.5 | 1.8 | 5.7 | 1.6 | +| clustered | 0.0001 | multiplicity | 129.8 | 18.5 | 22.8 | 18.3 | +| hotspot | 1 | support | 112.1 | 137.1 | 136.5 | 164.8 | +| hotspot | 1 | multiplicity | 218.6 | 110.6 | 117.0 | 179.5 | +| hotspot | 0.5 | support | 112.0 | 66.9 | 73.2 | 83.5 | +| hotspot | 0.5 | multiplicity | 586.7 | 122.4 | 138.1 | 179.2 | +| hotspot | 0.1 | support | 259.6 | 21.5 | 32.5 | 32.8 | +| hotspot | 0.1 | multiplicity | 183.1 | 30.7 | 32.3 | 35.7 | +| hotspot | 0.01 | support | 113.3 | 2.7 | 6.8 | 3.8 | +| hotspot | 0.01 | multiplicity | 132.8 | 19.6 | 23.6 | 20.5 | +| hotspot | 0.001 | support | 112.0 | 1.8 | 6.1 | 1.8 | +| hotspot | 0.001 | multiplicity | 128.1 | 18.6 | 23.0 | 18.5 | +| hotspot | 0.0001 | support | 112.0 | 1.8 | 5.7 | 1.6 | +| hotspot | 0.0001 | multiplicity | 127.8 | 18.5 | 22.8 | 18.4 | +| zipf | 1 | support | 122.7 | 152.0 | 147.9 | 178.4 | +| zipf | 1 | multiplicity | 146.7 | 110.3 | 119.5 | 181.9 | +| zipf | 0.5 | support | 122.2 | 74.5 | 73.6 | 89.8 | +| zipf | 0.5 | multiplicity | 433.7 | 70.5 | 92.3 | 100.2 | +| zipf | 0.1 | support | 125.1 | 18.0 | 15.8 | 20.7 | +| zipf | 0.1 | multiplicity | 205.2 | 30.7 | 32.3 | 35.6 | +| zipf | 0.01 | support | 122.0 | 2.7 | 6.9 | 4.0 | +| zipf | 0.01 | multiplicity | 135.6 | 19.5 | 23.6 | 20.7 | +| zipf | 0.001 | support | 124.6 | 1.8 | 5.9 | 1.8 | +| zipf | 0.001 | multiplicity | 130.5 | 18.6 | 23.0 | 18.7 | +| zipf | 0.0001 | support | 143.3 | 1.8 | 7.9 | 1.6 | +| zipf | 0.0001 | multiplicity | 130.5 | 18.5 | 22.8 | 18.5 | +| one-to-one | 1 | support | 178.6 | 79.2 | 149.6 | 138.6 | +| one-to-one | 1 | multiplicity | 172.0 | 130.8 | 139.2 | 223.4 | +| one-to-one | 0.5 | support | 107.1 | 44.8 | 70.4 | 109.2 | +| one-to-one | 0.5 | multiplicity | 550.8 | 123.4 | 139.6 | 183.5 | +| one-to-one | 0.1 | support | 253.7 | 21.1 | 32.0 | 32.1 | +| one-to-one | 0.1 | multiplicity | 292.5 | 53.4 | 59.9 | 60.6 | +| one-to-one | 0.01 | support | 231.7 | 2.6 | 6.8 | 3.5 | +| one-to-one | 0.01 | multiplicity | 154.3 | 19.8 | 23.9 | 20.7 | +| one-to-one | 0.001 | support | 107.1 | 1.8 | 6.3 | 1.7 | +| one-to-one | 0.001 | multiplicity | 149.5 | 18.6 | 23.0 | 18.6 | +| one-to-one | 0.0001 | support | 104.7 | 1.8 | 5.7 | 1.6 | +| one-to-one | 0.0001 | multiplicity | 149.2 | 18.5 | 22.8 | 18.3 | +| many-to-one | 1 | support | 136.2 | 157.6 | 162.8 | 199.5 | +| many-to-one | 1 | multiplicity | 156.0 | 142.7 | 148.9 | 202.7 | +| many-to-one | 0.5 | support | 135.5 | 76.3 | 110.8 | 154.4 | +| many-to-one | 0.5 | multiplicity | 432.8 | 113.9 | 136.2 | 165.3 | +| many-to-one | 0.1 | support | 137.8 | 15.8 | 15.9 | 21.0 | +| many-to-one | 0.1 | multiplicity | 193.5 | 32.6 | 33.5 | 38.2 | +| many-to-one | 0.01 | support | 136.8 | 2.7 | 6.9 | 4.1 | +| many-to-one | 0.01 | multiplicity | 224.4 | 23.5 | 23.7 | 20.5 | +| many-to-one | 0.001 | support | 135.8 | 2.8 | 13.2 | 2.9 | +| many-to-one | 0.001 | multiplicity | 247.9 | 28.4 | 38.5 | 28.6 | +| many-to-one | 0.0001 | support | 248.5 | 1.8 | 5.7 | 1.6 | +| many-to-one | 0.0001 | multiplicity | 145.8 | 18.4 | 22.7 | 18.3 | + +## A4 bounded extent × mask (µs, per-row minimum over 5 runs, each a median of 101; u32 payload, ordered i32 lane; window starts at min(span/5, span − width)) + +| mask density | extent frac | extent selves | live in extent | closed cells in extent | full universe | extent only | mask only (u16) | extent+u16 | extent+u64 | +|---|---|---|---|---|---|---|---|---|---| +| 1 | 1 | 65536 | 65536 | 0 | 101.6 | 58.0 | 83.5 | 50.6 | 42.9 | +| 1 | 0.5 | 32768 | 32768 | 0 | 101.7 | 29.1 | 70.2 | 25.4 | 21.9 | +| 1 | 0.1 | 6552 | 6552 | 0 | 101.7 | 5.9 | 60.9 | 5.1 | 4.4 | +| 1 | 0.01 | 652 | 652 | 0 | 101.6 | 0.6 | 58.7 | 0.6 | 0.5 | +| 1 | 0.001 | 64 | 64 | 0 | 101.6 | 0.1 | 58.5 | 0.1 | 0.1 | +| 0.5 | 1 | 65536 | 32768 | 0 | 101.6 | 58.0 | 65.4 | 47.1 | 29.1 | +| 0.5 | 0.5 | 32768 | 16336 | 0 | 101.6 | 29.1 | 57.2 | 15.0 | 12.6 | +| 0.5 | 0.1 | 6552 | 3273 | 0 | 101.7 | 5.9 | 54.7 | 2.8 | 2.4 | +| 0.5 | 0.01 | 652 | 345 | 0 | 101.6 | 0.6 | 54.0 | 0.3 | 0.3 | +| 0.5 | 0.001 | 64 | 31 | 0 | 101.6 | 0.1 | 54.9 | 0.1 | 0.1 | +| 0.1 | 1 | 65536 | 6554 | 757 | 101.6 | 58.0 | 12.2 | 10.0 | 5.5 | +| 0.1 | 0.5 | 32768 | 3321 | 349 | 101.6 | 29.1 | 11.1 | 5.1 | 2.9 | +| 0.1 | 0.1 | 6552 | 689 | 70 | 101.7 | 5.9 | 10.4 | 1.1 | 0.6 | +| 0.1 | 0.01 | 652 | 59 | 7 | 101.7 | 0.6 | 10.1 | 0.2 | 0.1 | +| 0.1 | 0.001 | 64 | 6 | 0 | 101.6 | 0.1 | 10.3 | 0.1 | 0.1 | +| 0.01 | 1 | 65536 | 655 | 3480 | 101.7 | 58.0 | 4.9 | 6.7 | 1.6 | +| 0.01 | 0.5 | 32768 | 330 | 1742 | 101.6 | 29.1 | 4.8 | 3.3 | 0.9 | +| 0.01 | 0.1 | 6552 | 65 | 350 | 101.8 | 5.9 | 4.7 | 0.7 | 0.2 | +| 0.01 | 0.01 | 652 | 8 | 33 | 101.6 | 0.6 | 4.7 | 0.1 | 0.1 | +| 0.01 | 0.001 | 64 | 0 | 4 | 101.7 | 0.1 | 4.7 | 0.1 | 0.1 | +| 0.001 | 1 | 65536 | 66 | 4030 | 101.7 | 58.0 | 4.5 | 6.5 | 1.1 | +| 0.001 | 0.5 | 32768 | 35 | 2013 | 101.6 | 29.1 | 4.5 | 3.3 | 0.7 | +| 0.001 | 0.1 | 6552 | 6 | 404 | 101.6 | 5.9 | 4.5 | 0.7 | 0.2 | +| 0.001 | 0.01 | 652 | 1 | 40 | 101.7 | 0.6 | 4.5 | 0.1 | 0.1 | +| 0.001 | 0.001 | 64 | 0 | 4 | 101.7 | 0.1 | 4.5 | 0.1 | 0.1 | + +## A5 K propagation over CSR (µs, median of 101; includes clearing K_next) + +| target | fanout | src density | live | edges | edge scan | src u64 gate | src u16 gate | ordinal | +|---|---|---|---|---|---|---|---|---| +| uniform | 1 | 1 | 65536 | 65536 | 186.3 | 181.6 | 186.5 | 242.0 | +| uniform | 1 | 0.1 | 6554 | 65536 | 231.5 | 36.4 | 43.9 | 42.6 | +| uniform | 1 | 0.01 | 655 | 65536 | 172.1 | 21.9 | 25.6 | 22.2 | +| uniform | 1 | 0.001 | 66 | 65536 | 166.4 | 18.7 | 23.3 | 20.0 | +| uniform | 4 | 1 | 65536 | 262144 | 766.7 | 374.9 | 382.1 | 477.4 | +| uniform | 4 | 0.1 | 6554 | 262144 | 778.6 | 62.9 | 76.2 | 66.7 | +| uniform | 4 | 0.01 | 655 | 262144 | 682.4 | 23.3 | 29.0 | 35.1 | +| uniform | 4 | 0.001 | 66 | 262144 | 654.5 | 19.3 | 24.0 | 20.8 | +| uniform | 16 | 1 | 65536 | 1048576 | 2984.1 | 1287.6 | 1270.6 | 1376.1 | +| uniform | 16 | 0.1 | 6554 | 1048576 | 2731.7 | 171.3 | 185.8 | 182.6 | +| uniform | 16 | 0.01 | 655 | 1048576 | 2616.1 | 33.8 | 39.2 | 35.0 | +| uniform | 16 | 0.001 | 66 | 1048576 | 2590.5 | 20.3 | 24.8 | 21.8 | +| clustered | 1 | 1 | 65536 | 65536 | 172.9 | 183.2 | 179.9 | 217.8 | +| clustered | 1 | 0.1 | 6554 | 65536 | 214.9 | 36.6 | 47.0 | 42.3 | +| clustered | 1 | 0.01 | 655 | 65536 | 157.5 | 20.7 | 25.9 | 23.6 | +| clustered | 1 | 0.001 | 66 | 65536 | 153.1 | 18.7 | 23.2 | 20.2 | +| clustered | 4 | 1 | 65536 | 262144 | 652.6 | 300.1 | 301.4 | 346.7 | +| clustered | 4 | 0.1 | 6554 | 262144 | 658.3 | 57.3 | 73.7 | 60.7 | +| clustered | 4 | 0.01 | 655 | 262144 | 594.0 | 23.1 | 28.6 | 24.4 | +| clustered | 4 | 0.001 | 66 | 262144 | 582.5 | 19.0 | 23.4 | 20.3 | +| clustered | 16 | 1 | 65536 | 1048576 | 2776.9 | 1004.1 | 944.7 | 1015.0 | +| clustered | 16 | 0.1 | 6554 | 1048576 | 2464.5 | 252.3 | 255.4 | 241.8 | +| clustered | 16 | 0.01 | 655 | 1048576 | 2423.6 | 45.2 | 55.9 | 34.9 | +| clustered | 16 | 0.001 | 66 | 1048576 | 2371.2 | 19.9 | 24.4 | 21.5 | +| hotspot | 1 | 1 | 65536 | 65536 | 233.2 | 213.2 | 182.7 | 216.2 | +| hotspot | 1 | 0.1 | 6554 | 65536 | 213.7 | 36.0 | 43.9 | 42.1 | +| hotspot | 1 | 0.01 | 655 | 65536 | 155.1 | 20.5 | 24.8 | 21.9 | +| hotspot | 1 | 0.001 | 66 | 65536 | 257.4 | 18.7 | 23.3 | 20.1 | +| hotspot | 4 | 1 | 65536 | 262144 | 707.8 | 318.0 | 303.0 | 367.5 | +| hotspot | 4 | 0.1 | 6554 | 262144 | 655.6 | 57.0 | 72.6 | 61.5 | +| hotspot | 4 | 0.01 | 655 | 262144 | 581.1 | 22.4 | 28.4 | 24.0 | +| hotspot | 4 | 0.001 | 66 | 262144 | 570.1 | 18.8 | 23.5 | 20.3 | +| hotspot | 16 | 1 | 65536 | 1048576 | 2857.2 | 1889.5 | 1939.3 | 979.8 | +| hotspot | 16 | 0.1 | 6554 | 1048576 | 2445.6 | 166.2 | 191.6 | 155.2 | +| hotspot | 16 | 0.01 | 655 | 1048576 | 2317.0 | 30.6 | 36.8 | 31.8 | +| hotspot | 16 | 0.001 | 66 | 1048576 | 2316.5 | 30.7 | 40.6 | 31.9 | +| many-to-one | 1 | 1 | 65536 | 65536 | 317.3 | 208.3 | 179.2 | 218.5 | +| many-to-one | 1 | 0.1 | 6554 | 65536 | 217.3 | 35.9 | 51.5 | 42.2 | +| many-to-one | 1 | 0.01 | 655 | 65536 | 160.6 | 20.3 | 25.5 | 21.8 | +| many-to-one | 1 | 0.001 | 66 | 65536 | 154.0 | 18.6 | 23.3 | 20.0 | +| many-to-one | 4 | 1 | 65536 | 262144 | 691.7 | 603.3 | 612.2 | 733.5 | +| many-to-one | 4 | 0.1 | 6554 | 262144 | 688.8 | 67.1 | 78.3 | 77.8 | +| many-to-one | 4 | 0.01 | 655 | 262144 | 606.7 | 22.9 | 28.6 | 24.3 | +| many-to-one | 4 | 0.001 | 66 | 262144 | 594.7 | 18.9 | 23.6 | 20.2 | +| many-to-one | 16 | 1 | 65536 | 1048576 | 2664.0 | 2329.3 | 2358.0 | 2602.8 | +| many-to-one | 16 | 0.1 | 6554 | 1048576 | 2585.5 | 239.5 | 260.9 | 265.8 | +| many-to-one | 16 | 0.01 | 655 | 1048576 | 2426.2 | 39.2 | 48.2 | 42.4 | +| many-to-one | 16 | 0.001 | 66 | 1048576 | 2505.0 | 20.8 | 25.5 | 22.2 | + +## A6 target-side scheduling (µs): median over 4 runs in which every route took every run position (`A6_ROT=0..3`), each run a median of 101 + +| target | src density | live | touched target cells | max contributions/cell | exact + full consume | exact + cell-seen bitmap | u16 histogram pre-pass | exact + target mask | worst position spread | +|---|---|---|---|---|---|---|---|---|---| +| uniform | 1 | 65536 | 4096 | 32 | 244.0 | 209.2 | 244.0 | 215.9 | 1.87× | +| uniform | 0.1 | 6554 | 3279 | 7 | 138.0 | 71.0 | 99.8 | 36.3 | 2.00× | +| uniform | 0.01 | 655 | 610 | 4 | 122.9 | 15.7 | 22.3 | 8.2 | 1.13× | +| uniform | 0.001 | 66 | 66 | 1 | 121.4 | 7.0 | 14.7 | 6.3 | 1.11× | +| uniform | 0.0001 | 7 | 7 | 1 | 119.5 | 5.6 | 12.8 | 5.7 | 1.00× | +| clustered | 1 | 65536 | 4096 | 34 | 217.4 | 218.4 | 246.8 | 171.8 | 2.09× | +| clustered | 0.1 | 6554 | 3310 | 7 | 162.1 | 69.7 | 131.8 | 36.7 | 2.18× | +| clustered | 0.01 | 655 | 607 | 2 | 141.0 | 21.2 | 32.0 | 10.6 | 2.23× | +| clustered | 0.001 | 66 | 66 | 1 | 159.6 | 11.7 | 19.4 | 8.5 | 2.25× | +| clustered | 0.0001 | 7 | 7 | 1 | 130.6 | 5.6 | 13.3 | 5.7 | 1.89× | +| hotspot | 1 | 65536 | 3265 | 14830 | 223.0 | 191.4 | 241.6 | 152.2 | 2.01× | +| hotspot | 0.1 | 6554 | 617 | 1499 | 166.9 | 31.6 | 69.2 | 28.4 | 1.78× | +| hotspot | 0.01 | 655 | 67 | 162 | 140.9 | 11.0 | 20.6 | 9.9 | 2.14× | +| hotspot | 0.001 | 66 | 10 | 19 | 137.1 | 6.8 | 19.5 | 10.1 | 2.21× | +| hotspot | 0.0001 | 7 | 5 | 2 | 138.4 | 10.7 | 17.3 | 8.6 | 2.39× | +| zipf | 1 | 65536 | 100 | 61376 | 305.2 | 209.4 | 261.1 | 208.1 | 2.07× | +| zipf | 0.1 | 6554 | 36 | 6166 | 152.5 | 31.4 | 67.1 | 25.3 | 2.00× | +| zipf | 0.01 | 655 | 12 | 615 | 140.8 | 10.0 | 20.0 | 10.8 | 2.12× | +| zipf | 0.001 | 66 | 3 | 60 | 139.0 | 7.8 | 19.1 | 8.3 | 1.88× | +| zipf | 0.0001 | 7 | 2 | 6 | 138.2 | 7.5 | 17.5 | 7.7 | 1.88× | +| one-to-one | 1 | 65536 | 4096 | 16 | 276.0 | 288.2 | 275.2 | 266.0 | 1.86× | +| one-to-one | 0.1 | 6554 | 3341 | 7 | 152.9 | 67.7 | 103.8 | 38.5 | 1.47× | +| one-to-one | 0.01 | 655 | 612 | 3 | 122.5 | 16.2 | 22.8 | 8.3 | 1.77× | +| one-to-one | 0.001 | 66 | 66 | 1 | 120.6 | 7.0 | 14.7 | 6.3 | 1.74× | +| one-to-one | 0.0001 | 7 | 7 | 1 | 119.6 | 5.6 | 12.8 | 5.7 | 1.80× | +| many-to-one | 1 | 65536 | 16 | 4096 | 244.9 | 153.6 | 292.2 | 149.8 | 1.09× | +| many-to-one | 0.1 | 6554 | 16 | 458 | 134.6 | 24.0 | 69.7 | 24.2 | 2.02× | +| many-to-one | 0.01 | 655 | 16 | 54 | 121.6 | 8.3 | 14.6 | 7.7 | 2.21× | +| many-to-one | 0.001 | 66 | 16 | 7 | 122.4 | 6.3 | 13.9 | 6.3 | 1.95× | +| many-to-one | 0.0001 | 7 | 7 | 1 | 122.2 | 5.6 | 12.8 | 5.7 | 2.28× | diff --git a/crates/lance-graph-benches/examples/aperture16_probe.rs b/crates/lance-graph-benches/examples/aperture16_probe.rs new file mode 100644 index 000000000..377ba6d4b --- /dev/null +++ b/crates/lance-graph-benches/examples/aperture16_probe.rs @@ -0,0 +1,1530 @@ +//! D-APERTURE-16-0: can the 64K self-space support mask be read as a +//! `[u16; 4096]` execution aperture, and does 16-self granularity schedule +//! work better than the 64-bit words the production kernels already walk? +//! +//! Universe: `self: u16`, 65 536 ordinals. Support mask: `[u64; 1024]` +//! (8 KiB). The u16 view is NOT a copy and NOT a cast: cell `c` is +//! `(words[c >> 2] >> ((c & 3) * 16)) as u16`, which is the same bits in +//! the same order on any endianness (on little-endian it is also exactly what +//! `slice::align_to::` would yield; checked in `views_agree`). +//! +//! Modes (all assert every route's answer equal before printing a time): +//! - `a1` : visitation only — the same mask walked as u64 words, as u16 +//! cells, adaptive u16 / u64, and a materialised ordinal list; +//! - `a2` : payload-width ladder (u8 … 128-byte records) — full scans vs +//! gated visitation; +//! - `occ`: one fixed occupancy per cell (0..=16 live) — set-bit visit vs a +//! dense masked 16-lane loop, the local crossover; +//! - `a3` : VIA `via[self] -> target` — support (`target |= 1`) and +//! multiplicity (`K_next[target] += K[self]`) separately; +//! - `a4` : bounded extent × aperture over an ordered lane; +//! - `a5` : K propagation over a CSR edge list, gated by source support; +//! - `a6` : coarse target cells — exact-only vs cell-seen bitmap during the +//! pass vs a `[u16; 4096]` histogram pre-pass, consumed by a next-frontier +//! build. +//! +//! ```bash +//! cargo run --release -p lance-graph-benches --example aperture16_probe -- a1 +//! ``` +//! +//! Output is TSV on stdout, one row per (pattern, payload, route). + +// The subject of this probe IS the row-ordinal geometry (`self`, `cell`, +// `word`), so its loops index by ordinal on purpose; an iterator would hide +// the coordinate the measurement is about. +#![allow(clippy::needless_range_loop)] + +use std::hint::black_box; +use std::time::Instant; + +use lance_graph_mask_risc::{execute_into, Foreign, LaneRef, Out, Planes, Scratch, Value}; +use lance_graph_quack::{lower, Agg, Col, Filter, Mask, Query}; + +const N: usize = 65_536; +const WORDS: usize = N / 64; // 1024 +const CELLS: usize = N / 16; // 4096 + +// ───────────────────────────── utilities ───────────────────────────── + +fn mix(mut x: u64) -> u64 { + x = x.wrapping_add(0x9E37_79B9_7F4A_7C15); + x = (x ^ (x >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + x = (x ^ (x >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + x ^ (x >> 31) +} + +fn reps() -> usize { + std::env::var("PROBE_REPS") + .ok() + .and_then(|s| s.parse().ok()) + .unwrap_or(101) +} + +/// Median wall time in ns of `f`, and its last result. +fn time(reps: usize, mut f: impl FnMut() -> T) -> (f64, T) { + let mut last = f(); // warm + let mut v: Vec = (0..reps) + .map(|_| { + let t0 = Instant::now(); + last = black_box(f()); + t0.elapsed().as_nanos() + }) + .collect(); + v.sort_unstable(); + (v[v.len() / 2] as f64, last) +} + +#[inline(always)] +fn bit(m: &[u64], i: usize) -> u64 { + (m[i >> 6] >> (i & 63)) & 1 +} + +/// Zero-copy u16 cell view of the u64 words. +#[inline(always)] +fn cell(m: &[u64], c: usize) -> u16 { + (m[c >> 2] >> ((c & 3) * 16)) as u16 +} + +// ───────────────────────────── layouts ───────────────────────────── + +#[derive(Clone, Copy, Debug)] +enum Layout { + Uniform, + Clustered, + SingleRange, + TinyRuns, + Islands, + Alternating, + OnePer16, + OnePer64, +} + +impl Layout { + fn name(self) -> &'static str { + match self { + Layout::Uniform => "uniform", + Layout::Clustered => "clustered", + Layout::SingleRange => "single-range", + Layout::TinyRuns => "tiny-runs", + Layout::Islands => "islands16", + Layout::Alternating => "alternating", + Layout::OnePer16 => "one-per-16", + Layout::OnePer64 => "one-per-64", + } + } + const DENSITY_LAYOUTS: [Layout; 5] = [ + Layout::Uniform, + Layout::Clustered, + Layout::SingleRange, + Layout::TinyRuns, + Layout::Islands, + ]; + const FIXED: [Layout; 3] = [Layout::Alternating, Layout::OnePer16, Layout::OnePer64]; +} + +const DENSITIES: [f64; 11] = [ + 1.0, 0.75, 0.5, 0.25, 0.1, 0.05, 0.01, 0.001, 0.0001, 0.00001, 0.000001, +]; + +fn set(m: &mut [u64], i: usize) { + m[i >> 6] |= 1 << (i & 63); +} + +/// Exactly `round(d·N)` live selves (fixed layouts ignore `d`). +fn build(layout: Layout, d: f64, seed: u64) -> Vec { + let mut m = vec![0u64; WORDS]; + let k = (d * N as f64).round() as usize; + match layout { + Layout::Uniform => { + // partial Fisher–Yates: exactly k distinct positions + let mut p: Vec = (0..N as u32).collect(); + for i in 0..k { + let j = i + (mix(seed ^ i as u64) as usize % (N - i)); + p.swap(i, j); + set(&mut m, p[i] as usize); + } + } + Layout::Clustered => { + // runs of 128..640 selves at random starts until k are live + let mut live = 0; + let mut s = seed; + while live < k { + s = mix(s); + let start = (s % N as u64) as usize; + let len = 128 + (mix(s ^ 7) % 512) as usize; + for i in start..(start + len).min(N) { + if live == k { + break; + } + if bit(&m, i) == 0 { + set(&mut m, i); + live += 1; + } + } + } + } + Layout::SingleRange => { + let start = (N - k) / 3; + for i in start..start + k { + set(&mut m, i); + } + } + Layout::TinyRuns => { + let mut live = 0; + let mut s = seed; + while live < k { + s = mix(s); + let start = (s % N as u64) as usize; + let len = 2 + (mix(s ^ 3) % 3) as usize; + for i in start..(start + len).min(N) { + if live == k { + break; + } + if bit(&m, i) == 0 { + set(&mut m, i); + live += 1; + } + } + } + } + Layout::Islands => { + // whole 16-cells fully dense, chosen at random; k rounded down to + // a multiple of 16 (a partial island would not be an island) + let cells = k / 16; + let mut p: Vec = (0..CELLS as u32).collect(); + for i in 0..cells { + let j = i + (mix(seed ^ i as u64) as usize % (CELLS - i)); + p.swap(i, j); + let c = p[i] as usize; + for s in 0..16 { + set(&mut m, c * 16 + s); + } + } + } + Layout::Alternating => m.iter_mut().for_each(|w| *w = 0x5555_5555_5555_5555), + Layout::OnePer16 => { + for c in 0..CELLS { + set(&mut m, c * 16 + (mix(seed ^ c as u64) % 16) as usize); + } + } + Layout::OnePer64 => { + for w in 0..WORDS { + set(&mut m, w * 64 + (mix(seed ^ w as u64) % 64) as usize); + } + } + } + m +} + +fn patterns() -> Vec<(Layout, f64, Vec)> { + let mut v = Vec::new(); + for l in Layout::DENSITY_LAYOUTS { + for d in DENSITIES { + v.push((l, d, build(l, d, 0xA9E7 ^ (d.to_bits())))); + } + } + for l in Layout::FIXED { + let m = build(l, 0.0, 0x51); + let d = live(&m) as f64 / N as f64; + v.push((l, d, m)); + } + v +} + +fn live(m: &[u64]) -> usize { + m.iter().map(|w| w.count_ones() as usize).sum() +} + +/// Geometry of one mask, the accounting the brief asks for. +#[derive(Clone, Copy)] +struct Geo { + live: usize, + zero_cells: usize, + full_cells: usize, + zero_words: usize, + full_words: usize, +} + +fn geo(m: &[u64]) -> Geo { + let mut g = Geo { + live: live(m), + zero_cells: 0, + full_cells: 0, + zero_words: 0, + full_words: 0, + }; + for c in 0..CELLS { + match cell(m, c) { + 0 => g.zero_cells += 1, + 0xFFFF => g.full_cells += 1, + _ => {} + } + } + for &w in m { + match w { + 0 => g.zero_words += 1, + u64::MAX => g.full_words += 1, + _ => {} + } + } + g +} + +/// Distinct 64-byte payload lines that hold at least one live self. +/// +/// `base` is the payload's address modulo 64. A `Vec

` is only aligned to +/// `align_of::

()` (≤ 8 here), so a record can straddle one more line than +/// an aligned layout would; the count uses the real offset. +fn live_lines(m: &[u64], width: usize, base: usize) -> usize { + let mut next_free = 0usize; // first line not yet counted + let mut n = 0; + for w in 0..WORDS { + let mut b = m[w]; + while b != 0 { + let i = w * 64 + b.trailing_zeros() as usize; + b &= b - 1; + let first = ((base + i * width) / 64).max(next_free); + let last = (base + i * width + width - 1) / 64; + if last >= first { + n += last - first + 1; + next_free = last + 1; + } + } + } + n +} + +// ───────────────────────────── payloads ───────────────────────────── + +trait Payload: Copy { + const W: usize; + fn make(i: u64) -> Self; + fn f(&self) -> u64; +} +impl Payload for u8 { + const W: usize = 1; + fn make(i: u64) -> Self { + mix(i) as u8 + } + #[inline(always)] + fn f(&self) -> u64 { + *self as u64 + } +} +impl Payload for u16 { + const W: usize = 2; + fn make(i: u64) -> Self { + mix(i) as u16 + } + #[inline(always)] + fn f(&self) -> u64 { + *self as u64 + } +} +impl Payload for u32 { + const W: usize = 4; + fn make(i: u64) -> Self { + (mix(i) as u32) >> 12 // < 2^20, so an i32 sum is exact + } + #[inline(always)] + fn f(&self) -> u64 { + *self as u64 + } +} +impl Payload for u64 { + const W: usize = 8; + fn make(i: u64) -> Self { + mix(i) + } + #[inline(always)] + fn f(&self) -> u64 { + *self + } +} +macro_rules! rec { + ($n:literal) => { + impl Payload for [u64; $n] { + const W: usize = 8 * $n; + fn make(i: u64) -> Self { + core::array::from_fn(|k| mix(i ^ ((k as u64) << 40))) + } + /// Touches every byte of the record. + #[inline(always)] + fn f(&self) -> u64 { + self.iter().fold(0u64, |a, x| a.wrapping_add(*x)) + } + } + }; +} +rec!(2); +rec!(4); +rec!(8); +rec!(16); + +// ───────────────────────────── routes ───────────────────────────── + +/// Scan every row, branch on its bit (the "consult mask per row" baseline). +fn scan_branch(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (i, x) in p.iter().enumerate() { + if bit(m, i) == 1 { + a = a.wrapping_add(x.f()); + } + } + a +} + +/// Load every row, AND with the bit (branchless; touches N × W). +fn full_branchless(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (i, x) in p.iter().enumerate() { + a = a.wrapping_add(x.f() & bit(m, i).wrapping_neg()); + } + a +} + +/// The production shape (ndarray `group_walk` / `masked_sum_i32`): skip a +/// zero word in one test, then visit set bits. +fn visit64(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + let i = w * 64 + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + a +} + +/// The aperture: test each u16 cell, visit set bits inside it. +fn visit16(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for c in 0..CELLS { + let mut b = cell(m, c); + while b != 0 { + let i = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + a +} + +/// u64 word test first (one test skips 64), u16 cells only inside live +/// words — the aperture read hierarchically. +fn visit64_16(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (w, &word) in m.iter().enumerate() { + if word == 0 { + continue; + } + for q in 0..4 { + let mut b = (word >> (q * 16)) as u16; + let base = w * 64 + q * 16; + while b != 0 { + let i = base + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + } + a +} + +/// Adaptive u16: skip / dense 16 / set-bit. +fn adapt16(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for c in 0..CELLS { + let b = cell(m, c); + if b == 0 { + continue; + } + let base = c * 16; + if b == 0xFFFF { + for x in &p[base..base + 16] { + a = a.wrapping_add(x.f()); + } + } else { + let mut b = b; + while b != 0 { + let i = base + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + } + a +} + +/// Adaptive u64: skip / dense 64 / set-bit. +fn adapt64(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (w, &word) in m.iter().enumerate() { + if word == 0 { + continue; + } + let base = w * 64; + if word == u64::MAX { + for x in &p[base..base + 64] { + a = a.wrapping_add(x.f()); + } + } else { + let mut b = word; + while b != 0 { + let i = base + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + } + a +} + +/// Hierarchical adaptive: u64 skip / dense 64; inside a partial word each +/// u16 quarter is skipped, run dense when full, set-bit visited otherwise. +fn adapt64_16(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (w, &word) in m.iter().enumerate() { + if word == 0 { + continue; + } + let base = w * 64; + if word == u64::MAX { + for x in &p[base..base + 64] { + a = a.wrapping_add(x.f()); + } + continue; + } + for q in 0..4 { + let b = (word >> (q * 16)) as u16; + let qb = base + q * 16; + if b == 0xFFFF { + for x in &p[qb..qb + 16] { + a = a.wrapping_add(x.f()); + } + } else { + let mut b = b; + while b != 0 { + let i = qb + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + } + } + a +} + +/// u64 schedule with branch-free full-run detection: the full u16 quarters +/// of a word are found without a branch, run dense, and every remaining +/// set bit is walked in ONE loop over the word (no per-quarter split). A +/// word with no full quarter costs one predictable test over `visit64`. +fn adapt64q(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for (w, &word) in m.iter().enumerate() { + if word == 0 { + continue; + } + let base = w * 64; + let mut full = 0u64; + for q in 0..4 { + let hit = (((word >> (q * 16)) & 0xFFFF) == 0xFFFF) as u64; + full |= hit.wrapping_neg() & (0xFFFF << (q * 16)); + } + if full != 0 { + for q in 0..4 { + if (full >> (q * 16)) & 1 == 1 { + for x in &p[base + q * 16..base + q * 16 + 16] { + a = a.wrapping_add(x.f()); + } + } + } + } + let mut b = word & !full; + while b != 0 { + let i = base + b.trailing_zeros() as usize; + b &= b - 1; + a = a.wrapping_add(p[i].f()); + } + } + a +} + +/// Selection vector: materialise ordinals, then fold. The build is inside +/// the timed region and its bytes are reported. +fn ordinal(m: &[u64], p: &[P], buf: &mut Vec) -> u64 { + buf.clear(); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + buf.push((w * 64) as u16 + b.trailing_zeros() as u16); + b &= b - 1; + } + } + buf.iter() + .fold(0u64, |a, &i| a.wrapping_add(p[i as usize].f())) +} + +/// Pure visitation checksum (no payload): count and ordinal sum. +fn vis_ck(i: usize, a: &mut (u64, u64)) { + a.0 += 1; + a.1 += i as u64; +} + +// ───────────────────────────── modes ───────────────────────────── + +fn mode_a1() { + let r = reps(); + println!("layout\tdensity\tlive\tzero_cells\tfull_cells\tzero_words\tfull_words\troute\tns\tmask_tests\ttemp_bytes"); + // native production Count over the same mask, for reference + let count_p = lower(&Query { + filter: Filter::plane(Mask(0)), + agg: Agg::Count, + }) + .unwrap(); + for (l, d, m) in patterns() { + let g = geo(&m); + let mut buf: Vec = Vec::with_capacity(N); + type Route<'a> = ( + &'static str, + Box (u64, u64) + 'a>, + usize, + usize, + ); + let routes: [Route; 6] = [ + ( + "u64-visit", + Box::new(|| { + let mut a = (0, 0); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + vis_ck(w * 64 + b.trailing_zeros() as usize, &mut a); + b &= b - 1; + } + } + a + }), + WORDS, + 0, + ), + ( + "u16-visit", + Box::new(|| { + let mut a = (0, 0); + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + vis_ck(c * 16 + b.trailing_zeros() as usize, &mut a); + b &= b - 1; + } + } + a + }), + CELLS, + 0, + ), + ( + "u64>u16-visit", + Box::new(|| { + let mut a = (0, 0); + for (w, &word) in m.iter().enumerate() { + if word == 0 { + continue; + } + for q in 0..4 { + let mut b = (word >> (q * 16)) as u16; + while b != 0 { + vis_ck(w * 64 + q * 16 + b.trailing_zeros() as usize, &mut a); + b &= b - 1; + } + } + } + a + }), + WORDS + 4 * (WORDS - g.zero_words), + 0, + ), + ( + "per-row-scan", + Box::new(|| { + let mut a = (0, 0); + for i in 0..N { + if bit(&m, i) == 1 { + vis_ck(i, &mut a); + } + } + a + }), + N, + 0, + ), + ( + "ordinal-list", + Box::new(|| { + buf.clear(); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + buf.push((w * 64) as u16 + b.trailing_zeros() as u16); + b &= b - 1; + } + } + let mut a = (0, 0); + for &i in buf.iter() { + vis_ck(i as usize, &mut a); + } + a + }), + WORDS, + 2 * g.live, + ), + ( + "native-count", + Box::new(|| { + let masks: [&[u64]; 1] = [&m]; + let planes = Planes { + n_rows: N, + masks: &masks, + lanes: &[], + }; + let mut s = Scratch::for_program(&count_p, N).unwrap(); + match execute_into(&count_p, &planes, &Foreign::NONE, &mut s, Out::None) + .unwrap() + { + Value::Count(c) => (c as u64, u64::MAX), + v => panic!("{v:?}"), + } + }), + WORDS, + 0, + ), + ]; + let mut want: Option<(u64, u64)> = None; + for (name, mut f, tests, temp) in routes { + let (ns, got) = time(r, &mut f); + if name == "native-count" { + assert_eq!(got.0, g.live as u64, "native count {l:?} {d}"); + } else if let Some(w) = want { + assert_eq!(got, w, "{name} {l:?} {d}"); + } else { + want = Some(got); + } + println!( + "{}\t{d}\t{}\t{}\t{}\t{}\t{}\t{name}\t{ns:.0}\t{tests}\t{temp}", + l.name(), + g.live, + g.zero_cells, + g.full_cells, + g.zero_words, + g.full_words + ); + } + } +} + +fn ladder_one(name: &str, r: usize, pats: &[(Layout, f64, Vec)]) { + let p: Vec

= (0..N as u64).map(P::make).collect(); + for (l, d, m) in pats { + let g = geo(m); + let mut buf: Vec = Vec::with_capacity(N); + let base = p.as_ptr() as usize % 64; + let lines = live_lines(m, P::W, base); + let all_lines = (base + N * P::W).div_ceil(64); + let want = scan_branch(m, &p); + type R<'a, P> = (&'static str, fn(&[u64], &[P]) -> u64); + let routes: [R

; 9] = [ + ("adapt64q", adapt64q::

), + ("adapt64>16", adapt64_16::

), + ("scan-branch", scan_branch::

), + ("full-branchless", full_branchless::

), + ("u64-visit", visit64::

), + ("u16-visit", visit16::

), + ("u64>u16-visit", visit64_16::

), + ("adapt16", adapt16::

), + ("adapt64", adapt64::

), + ]; + for (rn, f) in routes { + let (ns, got) = time(r, || f(black_box(m), black_box(&p))); + assert_eq!(got, want, "{rn} {name} {l:?} {d}"); + let (loads, ln) = match rn { + "scan-branch" | "full-branchless" => ( + if rn == "scan-branch" { g.live } else { N }, + if rn == "scan-branch" { + lines + } else { + all_lines + }, + ), + _ => (g.live, lines), + }; + println!( + "{}\t{d}\t{}\t{name}\t{}\t{rn}\t{ns:.0}\t{loads}\t{ln}\t0", + l.name(), + g.live, + P::W + ); + } + let (ns, got) = time(r, || ordinal(black_box(m), black_box(&p), &mut buf)); + assert_eq!(got, want, "ordinal {name} {l:?} {d}"); + println!( + "{}\t{d}\t{}\t{name}\t{}\tordinal-list\t{ns:.0}\t{}\t{lines}\t{}", + l.name(), + g.live, + P::W, + g.live, + 2 * g.live + ); + } +} + +fn mode_a2() { + let r = reps(); + println!("layout\tdensity\tlive\tpayload\twidth_B\troute\tns\tpayload_loads\tpayload_lines\ttemp_bytes"); + let pats = patterns(); + ladder_one::("u8", r, &pats); + ladder_one::("u16", r, &pats); + ladder_one::("u32", r, &pats); + ladder_one::("u64", r, &pats); + ladder_one::<[u64; 2]>("rec16", r, &pats); + ladder_one::<[u64; 4]>("rec32", r, &pats); + ladder_one::<[u64; 8]>("rec64", r, &pats); + ladder_one::<[u64; 16]>("rec128", r, &pats); +} + +/// Every cell holds exactly `k` live bits (random positions). +fn occ_mask(k: usize, seed: u64) -> Vec { + let mut m = vec![0u64; WORDS]; + for c in 0..CELLS { + let mut slots: [u8; 16] = core::array::from_fn(|i| i as u8); + for i in 0..k { + let j = i + (mix(seed ^ ((c * 16 + i) as u64)) as usize % (16 - i)); + slots.swap(i, j); + set(&mut m, c * 16 + slots[i] as usize); + } + } + m +} + +/// Dense masked 16-lane loop for every non-zero cell (branch-free inside). +fn dense16(m: &[u64], p: &[P]) -> u64 { + let mut a = 0u64; + for c in 0..CELLS { + let b = cell(m, c) as u64; + if b == 0 { + continue; + } + let base = c * 16; + for (s, x) in p[base..base + 16].iter().enumerate() { + a = a.wrapping_add(x.f() & ((b >> s) & 1).wrapping_neg()); + } + } + a +} + +fn occ_one(name: &str, r: usize) { + let p: Vec

= (0..N as u64).map(P::make).collect(); + for k in 0..=16 { + let m = occ_mask(k, 0x0CC ^ k as u64); + let want = scan_branch(&m, &p); + let (a, ga) = time(r, || visit16(black_box(&m), black_box(&p))); + let (b, gb) = time(r, || dense16(black_box(&m), black_box(&p))); + let (c, gc) = time(r, || visit64(black_box(&m), black_box(&p))); + let (e, ge) = time(r, || adapt16(black_box(&m), black_box(&p))); + assert_eq!((ga, gb, gc, ge), (want, want, want, want), "occ {name} {k}"); + println!("{name}\t{}\t{k}\t{a:.0}\t{b:.0}\t{c:.0}\t{e:.0}", P::W); + } +} + +fn mode_occ() { + let r = reps(); + println!("payload\twidth_B\tlive_per_cell\tu16_setbit_ns\tu16_dense_masked_ns\tu64_setbit_ns\tadapt16_ns"); + occ_one::("u8", r); + occ_one::("u32", r); + occ_one::("u64", r); + occ_one::<[u64; 4]>("rec32", r); + occ_one::<[u64; 16]>("rec128", r); +} + +// ── VIA: support and multiplicity, separately ── + +#[derive(Clone, Copy)] +enum Target { + Uniform, + Clustered, + Hotspot, + Zipf, + OneToOne, + ManyToOne, +} + +impl Target { + const ALL: [Target; 6] = [ + Target::Uniform, + Target::Clustered, + Target::Hotspot, + Target::Zipf, + Target::OneToOne, + Target::ManyToOne, + ]; + fn name(self) -> &'static str { + match self { + Target::Uniform => "uniform", + Target::Clustered => "clustered", + Target::Hotspot => "hotspot", + Target::Zipf => "zipf", + Target::OneToOne => "one-to-one", + Target::ManyToOne => "many-to-one", + } + } + fn of(self, s: usize) -> u16 { + let h = mix(s as u64 ^ 0x7A); + match self { + Target::Uniform => h as u16, + // the target stays within 256 of the source: local edges + Target::Clustered => (s as u64).wrapping_add(h % 256) as u16, + // 90 % of edges land in 64 hot targets + Target::Hotspot => { + if h % 10 < 9 { + (h >> 8) as u16 % 64 + } else { + (h >> 8) as u16 + } + } + // rank ∝ 1/u: inverse-transform of a heavy tail + Target::Zipf => { + let u = ((h >> 11) as f64 / (1u64 << 53) as f64).max(1e-12); + ((1.0 / u) as u64 % N as u64) as u16 + } + Target::OneToOne => (s as u16).wrapping_mul(40_503).wrapping_add(1), + Target::ManyToOne => (s >> 8) as u16, + } + } +} + +fn mode_a3() { + let r = reps(); + println!("target\tsrc_density\tlive\tsemantics\troute\tns\tvia_reads\tk_reads\ttarget_writes"); + let k: Vec = (0..N as u64).map(|i| (mix(i ^ 0x4B) % 7) as u32).collect(); + for t in Target::ALL { + let via: Vec = (0..N).map(|s| t.of(s)).collect(); + for d in [1.0, 0.5, 0.1, 0.01, 0.001, 0.0001] { + let m = build(Layout::Uniform, d, 0x5A ^ d.to_bits()); + let g = geo(&m); + let mut buf: Vec = Vec::with_capacity(N); + // ---- support: target_mask[via[s]] = 1 ---- + let mut tm = vec![0u64; WORDS]; + let sup = |tm: &mut [u64]| -> u64 { + tm.iter().map(|w| w.count_ones() as u64).sum::() + ^ tm.iter() + .fold(0u64, |a, w| a.wrapping_mul(31).wrapping_add(*w)) + }; + let mut want = None; + for route in ["full-scan", "u64-gate", "u16-gate", "ordinal"] { + let (ns, got) = time(r, || { + tm.iter_mut().for_each(|w| *w = 0); + match route { + "full-scan" => { + for s in 0..N { + let v = via[s] as usize; + tm[v >> 6] |= bit(&m, s) << (v & 63); + } + } + "u64-gate" => { + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + let v = via[w * 64 + b.trailing_zeros() as usize] as usize; + b &= b - 1; + tm[v >> 6] |= 1 << (v & 63); + } + } + } + "u16-gate" => { + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let v = via[c * 16 + b.trailing_zeros() as usize] as usize; + b &= b - 1; + tm[v >> 6] |= 1 << (v & 63); + } + } + } + _ => { + buf.clear(); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + buf.push((w * 64) as u16 + b.trailing_zeros() as u16); + b &= b - 1; + } + } + for &s in buf.iter() { + let v = via[s as usize] as usize; + tm[v >> 6] |= 1 << (v & 63); + } + } + } + sup(&mut tm) + }); + match want { + None => want = Some(got), + Some(w) => assert_eq!(got, w, "support {route}"), + } + let reads = if route == "full-scan" { N } else { g.live }; + println!( + "{}\t{d}\t{}\tsupport\t{route}\t{ns:.0}\t{reads}\t0\t{reads}", + t.name(), + g.live + ); + } + // ---- multiplicity: K_next[via[s]] += K[s] ---- + let mut kn = vec![0u32; N]; + let ck = |kn: &[u32]| { + kn.iter().enumerate().fold(0u64, |a, (i, x)| { + a.wrapping_add((*x as u64) * (i as u64 + 1)) + }) + }; + let mut want = None; + for route in ["full-scan", "u64-gate", "u16-gate", "ordinal"] { + let (ns, got) = time(r, || { + kn.iter_mut().for_each(|x| *x = 0); + match route { + "full-scan" => { + for s in 0..N { + kn[via[s] as usize] += k[s] & (bit(&m, s) as u32).wrapping_neg(); + } + } + "u64-gate" => { + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + let s = w * 64 + b.trailing_zeros() as usize; + b &= b - 1; + kn[via[s] as usize] += k[s]; + } + } + } + "u16-gate" => { + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + kn[via[s] as usize] += k[s]; + } + } + } + _ => { + buf.clear(); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + buf.push((w * 64) as u16 + b.trailing_zeros() as u16); + b &= b - 1; + } + } + for &s in buf.iter() { + kn[via[s as usize] as usize] += k[s as usize]; + } + } + } + ck(&kn) + }); + match want { + None => want = Some(got), + Some(w) => assert_eq!(got, w, "mult {route}"), + } + let reads = if route == "full-scan" { N } else { g.live }; + println!( + "{}\t{d}\t{}\tmultiplicity\t{route}\t{ns:.0}\t{reads}\t{reads}\t{reads}", + t.name(), + g.live + ); + } + } + } +} + +// ── bounded extent × aperture ── + +fn mode_a4() { + let r = reps(); + println!("mask_density\textent_frac\tlive_in_extent\tuniverse\textent_selves\tclosed_cells_in_extent\troute\tns\tts_reads\tpayload_loads"); + // non-decreasing ordered lane: ts = self / 4 (duplicates) + let ts: Vec = (0..N as i32).map(|i| i / 4).collect(); + let val: Vec = (0..N as u64).map(u32::make).collect(); + let span = ts[N - 1] + 1; + for md in [1.0, 0.5, 0.1, 0.01, 0.001] { + let m = build(Layout::Uniform, md, 0xE7 ^ md.to_bits()); + for ef in [1.0, 0.5, 0.1, 0.01, 0.001] { + // keep the requested width inside the lane: at ef = 1.0 the + // interval is the whole span, not a clipped tail of it. Only + // that case moves; every narrower window already fits at span/5. + let width = ((span as f64 * ef) as i32).max(1); + let a = (span / 5).min(span - width); + let b = a + width; + let want: u64 = (0..N) + .filter(|&i| bit(&m, i) == 1 && ts[i] >= a && ts[i] < b) + .map(|i| val[i] as u64) + .sum(); + let lo = ts.partition_point(|&t| t < a); + let hi = ts.partition_point(|&t| t < b); + let live_in: usize = (lo..hi).filter(|&i| bit(&m, i) == 1).count(); + let closed = (lo / 16..hi.div_ceil(16)) + .filter(|&c| cell(&m, c) == 0) + .count(); + // A: full universe — predicate every row, visit (m ∧ pred) + let (na, ga) = time(r, || { + let mut s = 0u64; + for i in 0..N { + let ok = (ts[i] >= a) & (ts[i] < b); + s += val[i] as u64 & ((bit(&m, i) & ok as u64).wrapping_neg()); + } + s + }); + // B: extent only — binary search, then every row of [lo, hi) + let (nb, gb) = time(r, || { + let lo = ts.partition_point(|&t| t < a); + let hi = ts.partition_point(|&t| t < b); + let mut s = 0u64; + for i in lo..hi { + s += val[i] as u64 & bit(&m, i).wrapping_neg(); + } + s + }); + // C: aperture only — visit live bits of the whole universe, + // test the ordered lane per live self + let (nc, gc) = time(r, || { + let mut s = 0u64; + for c in 0..CELLS { + let mut bb = cell(&m, c); + while bb != 0 { + let i = c * 16 + bb.trailing_zeros() as usize; + bb &= bb - 1; + if ts[i] >= a && ts[i] < b { + s += val[i] as u64; + } + } + } + s + }); + // D: extent + aperture — cells covering [lo, hi), edge cells + // clipped, set bits only + let (nd, gd) = time(r, || { + let lo = ts.partition_point(|&t| t < a); + let hi = ts.partition_point(|&t| t < b); + let mut s = 0u64; + if lo < hi { + let (c0, c1) = (lo / 16, (hi - 1) / 16); + for c in c0..=c1 { + let mut bb = cell(&m, c) as u32; + if c == c0 { + bb &= !0u32 << (lo % 16); + } + if c == c1 { + bb &= (1u32 << ((hi - 1) % 16 + 1)) - 1; + } + while bb != 0 { + let i = c * 16 + bb.trailing_zeros() as usize; + bb &= bb - 1; + s += val[i] as u64; + } + } + } + s + }); + // E: extent + u64 words — the same with 64-bit words + let (ne, ge) = time(r, || { + let lo = ts.partition_point(|&t| t < a); + let hi = ts.partition_point(|&t| t < b); + let mut s = 0u64; + if lo < hi { + let (w0, w1) = (lo / 64, (hi - 1) / 64); + for w in w0..=w1 { + let mut bb = m[w]; + if w == w0 { + bb &= !0u64 << (lo % 64); + } + if w == w1 && (hi - 1) % 64 != 63 { + bb &= (1u64 << ((hi - 1) % 64 + 1)) - 1; + } + while bb != 0 { + let i = w * 64 + bb.trailing_zeros() as usize; + bb &= bb - 1; + s += val[i] as u64; + } + } + } + s + }); + for (n, (ns, g, tsr, loads)) in [ + ("full-universe", (na, ga, N, N)), + ("extent-only", (nb, gb, 34, hi - lo)), + ("aperture-only", (nc, gc, live(&m), live(&m))), + ("extent+aperture16", (nd, gd, 34, live_in)), + ("extent+words64", (ne, ge, 34, live_in)), + ] { + assert_eq!(g, want, "{n} md={md} ef={ef}"); + println!( + "{md}\t{ef}\t{live_in}\t{N}\t{}\t{closed}\t{n}\t{ns:.0}\t{tsr}\t{loads}", + hi - lo + ); + } + } + } +} + +// ── K propagation over CSR edges ── + +struct Csr { + off: Vec, // N + 1 + dst: Vec, + src_of_edge: Vec, // for the edge-list full scan +} + +fn csr(t: Target, fan: usize) -> Csr { + let mut off = Vec::with_capacity(N + 1); + let mut dst = Vec::with_capacity(N * fan); + let mut soe = Vec::with_capacity(N * fan); + for s in 0..N { + off.push(dst.len() as u32); + for e in 0..fan { + dst.push( + t.of(s * fan + e) + ^ if matches!(t, Target::ManyToOne) { + 0 + } else { + e as u16 + }, + ); + soe.push(s as u16); + } + } + off.push(dst.len() as u32); + Csr { + off, + dst, + src_of_edge: soe, + } +} + +fn mode_a5() { + let r = reps(); + println!("target\tfanout\tsrc_density\tlive\tedges\troute\tns\tsrc_reads\tvia_reads\tk_reads\ttarget_writes"); + let k: Vec = (0..N as u64) + .map(|i| 1 + (mix(i ^ 0x4B) % 7) as u32) + .collect(); + for t in [ + Target::Uniform, + Target::Clustered, + Target::Hotspot, + Target::ManyToOne, + ] { + for fan in [1usize, 4, 16] { + let g = csr(t, fan); + let e = g.dst.len(); + for d in [1.0, 0.1, 0.01, 0.001] { + let m = build(Layout::Uniform, d, 0x6B ^ d.to_bits()); + let lv = live(&m); + let live_edges: usize = (0..N) + .filter(|&s| bit(&m, s) == 1) + .map(|s| (g.off[s + 1] - g.off[s]) as usize) + .sum(); + let mut kn = vec![0u32; N]; + let mut buf: Vec = Vec::with_capacity(N); + let ck = |kn: &[u32]| { + kn.iter().enumerate().fold(0u64, |a, (i, x)| { + a.wrapping_add((*x as u64) * (i as u64 + 1)) + }) + }; + let mut want = None; + for route in ["edge-scan", "src-u64-gate", "src-u16-gate", "ordinal"] { + let (ns, got) = time(r, || { + kn.iter_mut().for_each(|x| *x = 0); + match route { + "edge-scan" => { + for i in 0..e { + let s = g.src_of_edge[i] as usize; + kn[g.dst[i] as usize] += + k[s] & (bit(&m, s) as u32).wrapping_neg(); + } + } + "src-u64-gate" => { + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + let s = w * 64 + b.trailing_zeros() as usize; + b &= b - 1; + let ks = k[s]; + for &v in &g.dst[g.off[s] as usize..g.off[s + 1] as usize] { + kn[v as usize] += ks; + } + } + } + } + "src-u16-gate" => { + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + let ks = k[s]; + for &v in &g.dst[g.off[s] as usize..g.off[s + 1] as usize] { + kn[v as usize] += ks; + } + } + } + } + _ => { + buf.clear(); + for (w, &word) in m.iter().enumerate() { + let mut b = word; + while b != 0 { + buf.push((w * 64) as u16 + b.trailing_zeros() as u16); + b &= b - 1; + } + } + for &s in buf.iter() { + let s = s as usize; + let ks = k[s]; + for &v in &g.dst[g.off[s] as usize..g.off[s + 1] as usize] { + kn[v as usize] += ks; + } + } + } + } + ck(&kn) + }); + match want { + None => want = Some(got), + Some(w) => assert_eq!(got, w, "K {route}"), + } + let (sr, vr, kr, tw) = if route == "edge-scan" { + (e, e, e, e) + } else { + (lv, live_edges, lv, live_edges) + }; + println!( + "{}\t{fan}\t{d}\t{lv}\t{e}\t{route}\t{ns:.0}\t{sr}\t{vr}\t{kr}\t{tw}", + t.name() + ); + } + } + } + } +} + +// ── coarse target cells: does a target aperture pay for itself? ── + +fn mode_a6() { + let r = reps(); + println!("target\tsrc_density\tlive\ttouched_target_cells\tmax_cell_contrib\troute\tns\textra_pass_reads\tconsume_cells\trun_position"); + let k: Vec = (0..N as u64) + .map(|i| 1 + (mix(i ^ 0x4B) % 7) as u32) + .collect(); + let mut pattern = 0usize; + // A6_ROT shifts the rotation; running offsets 0..=3 puts every route in + // every run position for every pattern. + let rot_offset: usize = std::env::var("A6_ROT") + .ok() + .and_then(|s| s.parse().ok()) + .unwrap_or(0); + for t in Target::ALL { + let via: Vec = (0..N).map(|s| t.of(s)).collect(); + for d in [1.0, 0.1, 0.01, 0.001, 0.0001] { + let m = build(Layout::Uniform, d, 0x6C ^ d.to_bits()); + let lv = live(&m); + // reference: touched target cells, max contributions per cell + let mut hist_ref = vec![0u32; CELLS]; + for s in (0..N).filter(|&s| bit(&m, s) == 1) { + hist_ref[via[s] as usize >> 4] += 1; + } + let touched = hist_ref.iter().filter(|&&h| h > 0).count(); + let maxc = *hist_ref.iter().max().unwrap(); + // Each route: zero K_next, propagate under the source aperture, + // then CONSUME: build the next frontier mask (K_next > 0). + // The routes differ only in how much of K_next they zero and scan. + let mut kn = vec![0u32; N]; + let mut next = vec![0u64; WORDS]; + let mut seen = vec![0u64; CELLS / 64]; + let mut hist = vec![0u16; CELLS]; + let mut want = None; + // Rotate the route order per pattern so no route always runs + // first (cache state and clock drift would otherwise favour a + // fixed position); the order that ran is printed per row. + let mut order = [ + "exact+full-consume", + "exact+cell-seen", + "hist-prepass+exact", + "exact+target-mask", + ]; + order.rotate_left((pattern + rot_offset) % 4); + pattern += 1; + for (pos, route) in order.into_iter().enumerate() { + // K_next starts clean for every route; a route that only + // zeroes touched cells must leave it clean again. + kn.iter_mut().for_each(|x| *x = 0); + let (ns, got) = time(r, || { + next.iter_mut().for_each(|w| *w = 0); + match route { + "exact+full-consume" => { + kn.iter_mut().for_each(|x| *x = 0); + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + kn[via[s] as usize] += k[s]; + } + } + for (v, &x) in kn.iter().enumerate() { + next[v >> 6] |= ((x > 0) as u64) << (v & 63); + } + } + "exact+cell-seen" => { + seen.iter_mut().for_each(|w| *w = 0); + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + let v = via[s] as usize; + kn[v] += k[s]; + seen[v >> 10] |= 1 << ((v >> 4) & 63); + } + } + // consume only touched cells, and re-zero them + for (sw, &word) in seen.iter().enumerate() { + let mut b = word; + while b != 0 { + let tc = sw * 64 + b.trailing_zeros() as usize; + b &= b - 1; + let mut bits = 0u16; + for j in 0..16 { + let v = tc * 16 + j; + bits |= ((kn[v] > 0) as u16) << j; + kn[v] = 0; + } + next[tc >> 2] |= (bits as u64) << ((tc & 3) * 16); + } + } + } + "exact+target-mask" => { + // The exact next frontier is written in the same + // pass (K(u) ≥ 1 for every live u, so K_next > 0 + // iff touched); it is then the schedule for + // clearing K_next. No coarse object at all. + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + let v = via[s] as usize; + kn[v] += k[s]; + next[v >> 6] |= 1 << (v & 63); + } + } + for (w, &word) in next.iter().enumerate() { + let mut b = word; + while b != 0 { + kn[w * 64 + b.trailing_zeros() as usize] = 0; + b &= b - 1; + } + } + } + _ => { + hist.iter_mut().for_each(|h| *h = 0); + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + let h = &mut hist[via[s] as usize >> 4]; + *h = h.saturating_add(1); + } + } + for c in 0..CELLS { + let mut b = cell(&m, c); + while b != 0 { + let s = c * 16 + b.trailing_zeros() as usize; + b &= b - 1; + kn[via[s] as usize] += k[s]; + } + } + for (tc, &h) in hist.iter().enumerate() { + if h == 0 { + continue; + } + let mut bits = 0u16; + for j in 0..16 { + let v = tc * 16 + j; + bits |= ((kn[v] > 0) as u16) << j; + kn[v] = 0; + } + next[tc >> 2] |= (bits as u64) << ((tc & 3) * 16); + } + } + } + next.iter() + .fold(0u64, |a, w| a.wrapping_mul(31).wrapping_add(*w)) + }); + match want { + None => want = Some(got), + Some(w) => assert_eq!(got, w, "a6 {route} {} {d}", t.name()), + } + // The next-frontier checksum cannot see clearing: a route that + // skipped re-zeroing K_next would pass it AND time faster. The + // three schedule-driven routes must leave K_next clean. + if route != "exact+full-consume" { + assert!( + kn.iter().all(|&x| x == 0), + "a6 {route} left K_next dirty ({} {d})", + t.name() + ); + } + let (extra, consume) = match route { + "exact+full-consume" => (0, CELLS), + "exact+cell-seen" => (0, touched), + "exact+target-mask" => (0, WORDS), + _ => (lv, touched), + }; + println!( + "{}\t{d}\t{lv}\t{touched}\t{maxc}\t{route}\t{ns:.0}\t{extra}\t{consume}\t{pos}", + t.name() + ); + } + } + } +} + +/// The u16 view is the same bits as an in-place u16 reinterpretation. +fn views_agree() { + let m = build(Layout::Uniform, 0.3, 1); + // SAFETY: u64 -> u16 is a valid reinterpretation of initialised bytes; + // `align_to` returns empty prefix/suffix because u64 alignment ≥ u16. + let (pre, mid, suf) = unsafe { m.align_to::() }; + assert!(pre.is_empty() && suf.is_empty()); + if cfg!(target_endian = "little") { + for c in 0..CELLS { + assert_eq!(mid[c], cell(&m, c), "cell {c}"); + } + } + let _ = (Col(0), LaneRef::U32(&[])); +} + +fn main() { + views_agree(); + match std::env::args().nth(1).as_deref() { + Some("a1") => mode_a1(), + Some("a2") => mode_a2(), + Some("occ") => mode_occ(), + Some("a3") => mode_a3(), + Some("a4") => mode_a4(), + Some("a5") => mode_a5(), + Some("a6") => mode_a6(), + _ => eprintln!("usage: aperture16_probe a1|a2|occ|a3|a4|a5|a6"), + } +}