feat(cli): --prefill-ondemand for demo / chat / serve - #112
Merged
Merged
Conversation
The on-demand prefill path was only reachable from Python (rebuilding the tier's `LayerOptions` by hand), so the machines that need it most -- ones whose page cache cannot hold the checkpoint -- had no supported way to select it. `#111` made `full_layer_prefill=False` actually work; this exposes it. - `ModelConfig` gains `prefill_ondemand`; `from_pretrained` applies it to whatever preset the tier ships (`full_layer_prefill=False`), so it stays tier-generic and leaves the presets' measured defaults untouched. - `edge0 demo|chat|serve --prefill-ondemand` maps onto it; the three subparsers now share one `_add_engine_flags` helper, and `main` is a thin wrapper over a testable `_build_parser`. Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted): default (E3b) 5.33 s 265,965 pageins (4.06 GiB read) --prefill-ondemand 1.06 s 50,611 pageins (0.77 GiB read) warm-cache: 0.24 s vs 0.37 s (E3b stays the better default when the checkpoint fits -- hence a flag, not a preset change) Verified end to end: `edge0 chat --name edge0-8b --model-dir ... \ --prefill-ondemand --max-new 16` answers normally, and the slow suite asserts the switch reaches the engine, loads no whole layer, and leaves the next token identical. An advisory whole-shard `madvise(MADV_WILLNEED)` before the whole-layer prefill was tried in the same change and dropped: across evicted runs it ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not reproduce reliably enough to ship. Left for a follow-up.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The on-demand prefill path was only reachable from Python (rebuilding the tier's
LayerOptionsby hand), so the machines that need it most -- ones whose page cache cannot hold the checkpoint -- had no supported way to select it.#111madefull_layer_prefill=Falseactually work; this exposes it.ModelConfiggainsprefill_ondemand;from_pretrainedapplies it to whatever preset the tier ships (full_layer_prefill=False), so it stays tier-generic and leaves the presets' measured defaults untouched.edge0 demo|chat|serve --prefill-ondemandmaps onto it; the three subparsers now share one_add_engine_flagshelper, andmainis a thin wrapper over a testable_build_parser.Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted):
default (E3b) 5.33 s 265,965 pageins (4.06 GiB read)
--prefill-ondemand 1.06 s 50,611 pageins (0.77 GiB read)
warm-cache: 0.24 s vs 0.37 s (E3b stays the better default when the
checkpoint fits -- hence a flag, not a preset change)
Verified end to end:
edge0 chat --name edge0-8b --model-dir ... \ --prefill-ondemand --max-new 16answers normally, and the slow suite asserts the switch reaches the engine, loads no whole layer, and leaves the next token identical.An advisory whole-shard
madvise(MADV_WILLNEED)before the whole-layer prefill was tried in the same change and dropped: across evicted runs it ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not reproduce reliably enough to ship. Left for a follow-up.