Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions .claude/board/entries/2026-10-07-aperture16-u64-word-schedule.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# 2026-10-07 — The subtraction law holds; its unit is the u64 word, not the u16 cell

## MEASURED

`crates/lance-graph-benches/examples/aperture16_probe.rs`, over a 65 536-self universe:
- 11 densities × 8 layouts × 8 payload widths × 10 routes, plus VIA, K propagation over CSR, extent × mask, and target-side scheduling;
- every route asserted equal before timing.

- The `[u16; 4096]` view of the 64K mask is zero-copy: it is a shift of the same 8 KiB, endianness-proof, and checked against `align_to`. As a schedule, though, it is **1.3–2.3× slower per layout** (2.7× over all ladder cases) than u64 set-bit visitation, and 7× slower on an empty mask (4096 tests against 1024).
- Adding a full-word dense path to the u64 walk (`adapt64`): geomean 0.80, 2.4–2.9× on runs, at most 17 % worse on any layout.
- Full 16-run detection (`adapt64>16`) wins on 16-aligned islands (0.40) but is 1.25–1.64× slower on scattered layouts. A branch-free form does not remove that cost.
- Extent and mask compose multiplicatively: 2.9 µs, against 29.1 for extent only and 11.1 for mask only. The extent alone suffices only when the mask is dense.
- The mask schedules VIA and K propagation with no selection vector. The ordinal list shows no reproducible win; every apparent win reversed on re-run.
- On the target side, writing the exact next-frontier mask in the same pass is up to ~20× faster than clearing and scanning `K_next`, and the fastest route in 20 of 30 cases with run order rotated. The `[u16; 4096]` histogram pre-pass is never the fastest route (0 of 30).

## CONSEQUENCE

- No new carrier and no new V4 opcode. The mask is authoritative; u64 words are the unit.
- BUILD candidates, unbuilt:
- gated `mask_gather_u32` (the D-GATED-GATHER-0 item);
- a full-word dense path in the fold kernels' u64 walk;
- next-frontier-mask-in-pass for multi-hop K.
- KILL: u16 cells as the scheduling unit; popcount-threshold switching; `MultiplicityAperture`.

## OPEN

- Single machine, single thread, hand-rolled routes. Only A1's count goes through mask-risc.
- Host timing drifted by up to 2.4× on an identical pass, with no CPU steal. A4 uses a per-row minimum over 5 runs and A6 a median over 4 rotated runs; the other modes are single runs.
- Review corrections (#1377): cache-line counts use the real allocation offset; the 100 % extent case covers the whole universe; A6 rotates route order.
- Assembly not inspected; whether the dense paths auto-vectorise is unverified.
- `quack::SemanticAperture` already uses the word "aperture" for a different concept (a facet-key care mask).

Report: `.claude/research/D-APERTURE-16-0.md`, `D-APERTURE-16-MATRIX.md`, `D-APERTURE-16-tables.md`.
3 changes: 2 additions & 1 deletion .claude/board/entries/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,13 +25,14 @@ index row, (3) no duplicate entry id. Checks 1 and 2 are deliberately
opposite directions; the stranding this convention prevents shows up in
exactly one of them, never both.

240 entries, 2026-08-06 .. 2026-10-07.
241 entries, 2026-08-06 .. 2026-10-07.

| date | entry id | finding | file |
|---|---|---|---|
| 2026-10-07 | `tinker-janus-fold-harvest` | TinkerPop bulk = K (GroupReduce), ONE_BULK = S, GValue pinning = bundle invalidation; JanusGraph slice = OrderedLaneWitness→Range (30–38×); no new V4 op | [2026-10-07-tinker-janus-fold-harvest.md](2026-10-07-tinker-janus-fold-harvest.md) |
| 2026-10-07 | `gated-gather-bit-schedule` | Production gather is 3.8× slow from a per-row branch; bit-gated visitation wins or ties at every density; ordinal vectors never win (R1 holds) | [2026-10-07-gated-gather-bit-schedule.md](2026-10-07-gated-gather-bit-schedule.md) |
| 2026-10-07 | `bind-bundle-gated-gather` | One bound query, five routes, identical answers; the semijoin is never gated and dominates; a gated mask gather beats a selection vector (R1 survives) | [2026-10-07-bind-bundle-gated-gather.md](2026-10-07-bind-bundle-gated-gather.md) |
| 2026-10-07 | `aperture16-u64-word-schedule` | The u16 view of the 64K mask is free but 1.3–2.3× slower per layout as a schedule than u64 words; extent × mask compose; the exact next-frontier mask beats a u16 target histogram; no new carrier or opcode | [2026-10-07-aperture16-u64-word-schedule.md](2026-10-07-aperture16-u64-word-schedule.md) |
| 2026-10-06 | `v3-v4-dual-reading-round3` | | [2026-10-06-v3-v4-dual-reading-round3.md](2026-10-06-v3-v4-dual-reading-round3.md) |
| 2026-10-06 | `D-ANCESTRY-K1-0` | | [2026-10-06-streamdto-circuit-ancestry.md](2026-10-06-streamdto-circuit-ancestry.md) |
| 2026-10-06 | `D-SPOG-W-0` | | [2026-10-06-spog-witness-sub-context.md](2026-10-06-spog-witness-sub-context.md) |
Expand Down
13 changes: 13 additions & 0 deletions .claude/knowledge/fold-execution-laws.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,19 @@ to compute it now · EXECUTE touches the bytes. No fifth verb was forced.
9. **Anti-lasagne.** A new IR must hold information that cannot live in the
semantic plan, the binding, V4, the bundle choice or the backend program.
None proposed so far does.
10. **Subtract from the known universe, in u64 words.** Scheduling is
extent → skip zero words → dense full words → walk set bits → touch
payload last. The unit is the **u64 mask word**, not a 16-self cell. The
u16 view of the same 8 KiB is a free shift, but as a schedule it is 1.3–2.3×
slower per layout (2.7× overall geomean), and 7× slower on an empty mask (4096 tests vs 1024). 16-cells earn
a role only in detecting full 16-runs inside a partial word, opt-in. Cost
tracks live rows for lanes up to 16 B and live cache lines from 32 B up, never N.
[MEASURED, `aperture16_probe`, D-APERTURE-16-0]
11. **The target side subtracts too: the next frontier IS the schedule.** For
K propagation, writing the exact next-frontier mask in the same pass and
clearing `K_next` by walking it beats clearing and scanning the dense array
by up to ~20× at sparse frontiers. A coarse `[u16; 4096]` target histogram
pre-pass is never the fastest route. [MEASURED, `aperture16_probe a6`]

## Measured state of the candidate V4 (R2IL)

Expand Down
Loading
Loading