fix(prefill): honor full_layer_prefill=False on the streaming engines - #111
Merged
Merged
Conversation
`make_prefill_before_layer` used `full_n=0` both as "every layer" and as the unset/falsy value, so a hook built with `full_n=0` always fell through to `load_full_layer()`. Both streaming engines installed that hook for every multi-token prefill regardless of `full_layer_prefill`, so on the tiers that ship `prefill_hot=0` (edge0-8b) the flag was a silent no-op: disabling the whole-layer prefill still streamed all 128 experts of all 23 layers -- 4.06 GiB for a 27-token prompt, against 0.42 GiB for the routed experts alone. - hooks: take an explicit `full_layer` flag and skip the whole-layer load when it is off (the hot window, if any, still runs). - ling: pass `full_layer=opts.full_layer_prefill`, wire the hard-coded `hot_n=0` to `opts.prefill_hot`, and install the hook only when the whole-layer prefill or the hot stack is actually requested. - qwen: the same gating. Measured (M4 Pro, 24 GB, 27-token prompt, buffer cache evicted, MLX peak unchanged at 0.88-0.91 GiB): full_layer_prefill=True 5.50 s 266,354 pageins (4.06 GiB) 23 whole-layer loads full_layer_prefill=False 0.56 s 27,528 pageins (0.42 GiB) 0 whole-layer loads Warm cache the two arms are level (0.24 s vs 0.34 s), so E3b stays the right default wherever the checkpoint fits in the page cache; the on-demand path is the lever for machines that cannot hold it (16 GB class, issue #110). Tests: `tests/test_prefill_hook.py` covers the gating logic without weights, and a slow real-checkpoint regression in `tests/test_e2e_slow.py` asserts the on-demand arm never loads a whole layer while producing the same next token.
pull Bot
pushed a commit
to vishalbelsare/Edge0
that referenced
this pull request
Sep 20, 2026
The on-demand prefill path was only reachable from Python (rebuilding the tier's `LayerOptions` by hand), so the machines that need it most -- ones whose page cache cannot hold the checkpoint -- had no supported way to select it. `Edge0-AI#111` made `full_layer_prefill=False` actually work; this exposes it. - `ModelConfig` gains `prefill_ondemand`; `from_pretrained` applies it to whatever preset the tier ships (`full_layer_prefill=False`), so it stays tier-generic and leaves the presets' measured defaults untouched. - `edge0 demo|chat|serve --prefill-ondemand` maps onto it; the three subparsers now share one `_add_engine_flags` helper, and `main` is a thin wrapper over a testable `_build_parser`. Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted): default (E3b) 5.33 s 265,965 pageins (4.06 GiB read) --prefill-ondemand 1.06 s 50,611 pageins (0.77 GiB read) warm-cache: 0.24 s vs 0.37 s (E3b stays the better default when the checkpoint fits -- hence a flag, not a preset change) Verified end to end: `edge0 chat --name edge0-8b --model-dir ... \ --prefill-ondemand --max-new 16` answers normally, and the slow suite asserts the switch reaches the engine, loads no whole layer, and leaves the next token identical. An advisory whole-shard `madvise(MADV_WILLNEED)` before the whole-layer prefill was tried in the same change and dropped: across evicted runs it ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not reproduce reliably enough to ship. Left for a follow-up.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
make_prefill_before_layerusedfull_n=0both as "every layer" and as the unset/falsy value, so a hook built withfull_n=0always fell through toload_full_layer(). Both streaming engines installed that hook for every multi-token prefill regardless offull_layer_prefill, so on the tiers that shipprefill_hot=0(edge0-8b) the flag was a silent no-op: disabling the whole-layer prefill still streamed all 128 experts of all 23 layers -- 4.06 GiB for a 27-token prompt, against 0.42 GiB for the routed experts alone.full_layerflag and skip the whole-layer load when it is off (the hot window, if any, still runs).full_layer=opts.full_layer_prefill, wire the hard-codedhot_n=0toopts.prefill_hot, and install the hook only when the whole-layer prefill or the hot stack is actually requested.Measured (M4 Pro, 24 GB, 27-token prompt, buffer cache evicted, MLX peak unchanged at 0.88-0.91 GiB):
full_layer_prefill=True 5.50 s 266,354 pageins (4.06 GiB) 23 whole-layer loads
full_layer_prefill=False 0.56 s 27,528 pageins (0.42 GiB) 0 whole-layer loads
Warm cache the two arms are level (0.24 s vs 0.34 s), so E3b stays the right default wherever the checkpoint fits in the page cache; the on-demand path is the lever for machines that cannot hold it (16 GB class, issue #110).
Tests:
tests/test_prefill_hook.pycovers the gating logic without weights, and a slow real-checkpoint regression intests/test_e2e_slow.pyasserts the on-demand arm never loads a whole layer while producing the same next token.