Skip to content

fix(prefill): honor full_layer_prefill=False on the streaming engines - #111

Merged
linyubupa merged 1 commit into
mainfrom
fix/8b-prefill-full-layer-noop
Sep 20, 2026
Merged

linyubupa merged 1 commit into
mainfrom
fix/8b-prefill-full-layer-noop

Conversation

@linyubupa

Copy link
Copy Markdown
Collaborator

make_prefill_before_layer used full_n=0 both as "every layer" and as the unset/falsy value, so a hook built with full_n=0 always fell through to load_full_layer(). Both streaming engines installed that hook for every multi-token prefill regardless of full_layer_prefill, so on the tiers that ship prefill_hot=0 (edge0-8b) the flag was a silent no-op: disabling the whole-layer prefill still streamed all 128 experts of all 23 layers -- 4.06 GiB for a 27-token prompt, against 0.42 GiB for the routed experts alone.

  • hooks: take an explicit full_layer flag and skip the whole-layer load when it is off (the hot window, if any, still runs).
  • ling: pass full_layer=opts.full_layer_prefill, wire the hard-coded hot_n=0 to opts.prefill_hot, and install the hook only when the whole-layer prefill or the hot stack is actually requested.
  • qwen: the same gating.

Measured (M4 Pro, 24 GB, 27-token prompt, buffer cache evicted, MLX peak unchanged at 0.88-0.91 GiB):

full_layer_prefill=True 5.50 s 266,354 pageins (4.06 GiB) 23 whole-layer loads
full_layer_prefill=False 0.56 s 27,528 pageins (0.42 GiB) 0 whole-layer loads

Warm cache the two arms are level (0.24 s vs 0.34 s), so E3b stays the right default wherever the checkpoint fits in the page cache; the on-demand path is the lever for machines that cannot hold it (16 GB class, issue #110).

Tests: tests/test_prefill_hook.py covers the gating logic without weights, and a slow real-checkpoint regression in tests/test_e2e_slow.py asserts the on-demand arm never loads a whole layer while producing the same next token.

`make_prefill_before_layer` used `full_n=0` both as "every layer" and as the
unset/falsy value, so a hook built with `full_n=0` always fell through to
`load_full_layer()`.  Both streaming engines installed that hook for every
multi-token prefill regardless of `full_layer_prefill`, so on the tiers that
ship `prefill_hot=0` (edge0-8b) the flag was a silent no-op: disabling the
whole-layer prefill still streamed all 128 experts of all 23 layers -- 4.06 GiB
for a 27-token prompt, against 0.42 GiB for the routed experts alone.

- hooks: take an explicit `full_layer` flag and skip the whole-layer load when
  it is off (the hot window, if any, still runs).
- ling: pass `full_layer=opts.full_layer_prefill`, wire the hard-coded
  `hot_n=0` to `opts.prefill_hot`, and install the hook only when the
  whole-layer prefill or the hot stack is actually requested.
- qwen: the same gating.

Measured (M4 Pro, 24 GB, 27-token prompt, buffer cache evicted, MLX peak
unchanged at 0.88-0.91 GiB):

  full_layer_prefill=True   5.50 s   266,354 pageins (4.06 GiB)   23 whole-layer loads
  full_layer_prefill=False  0.56 s    27,528 pageins (0.42 GiB)    0 whole-layer loads

Warm cache the two arms are level (0.24 s vs 0.34 s), so E3b stays the right
default wherever the checkpoint fits in the page cache; the on-demand path is
the lever for machines that cannot hold it (16 GB class, issue #110).

Tests: `tests/test_prefill_hook.py` covers the gating logic without weights,
and a slow real-checkpoint regression in `tests/test_e2e_slow.py` asserts the
on-demand arm never loads a whole layer while producing the same next token.
@linyubupa
linyubupa merged commit afaf9d0 into main Sep 20, 2026
2 checks passed
pull Bot pushed a commit to vishalbelsare/Edge0 that referenced this pull request Sep 20, 2026
The on-demand prefill path was only reachable from Python (rebuilding the
tier's `LayerOptions` by hand), so the machines that need it most -- ones
whose page cache cannot hold the checkpoint -- had no supported way to
select it.  `Edge0-AI#111` made `full_layer_prefill=False` actually work; this
exposes it.

- `ModelConfig` gains `prefill_ondemand`; `from_pretrained` applies it to
  whatever preset the tier ships (`full_layer_prefill=False`), so it stays
  tier-generic and leaves the presets' measured defaults untouched.
- `edge0 demo|chat|serve --prefill-ondemand` maps onto it; the three
  subparsers now share one `_add_engine_flags` helper, and `main` is a thin
  wrapper over a testable `_build_parser`.

Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted):

  default (E3b)   5.33 s   265,965 pageins (4.06 GiB read)
  --prefill-ondemand  1.06 s    50,611 pageins (0.77 GiB read)
  warm-cache:     0.24 s vs 0.37 s (E3b stays the better default when the
                  checkpoint fits -- hence a flag, not a preset change)

Verified end to end: `edge0 chat --name edge0-8b --model-dir ... \
--prefill-ondemand --max-new 16` answers normally, and the slow suite
asserts the switch reaches the engine, loads no whole layer, and leaves the
next token identical.

An advisory whole-shard `madvise(MADV_WILLNEED)` before the whole-layer
prefill was tried in the same change and dropped: across evicted runs it
ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not
reproduce reliably enough to ship.  Left for a follow-up.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant