Skip to content

feat(cli): --prefill-ondemand for demo / chat / serve - #112

Merged
linyubupa merged 1 commit into
mainfrom
feat/prefill-ondemand-cli
Sep 20, 2026
Merged

linyubupa merged 1 commit into
mainfrom
feat/prefill-ondemand-cli

Conversation

@linyubupa

Copy link
Copy Markdown
Collaborator

The on-demand prefill path was only reachable from Python (rebuilding the tier's LayerOptions by hand), so the machines that need it most -- ones whose page cache cannot hold the checkpoint -- had no supported way to select it. #111 made full_layer_prefill=False actually work; this exposes it.

  • ModelConfig gains prefill_ondemand; from_pretrained applies it to whatever preset the tier ships (full_layer_prefill=False), so it stays tier-generic and leaves the presets' measured defaults untouched.
  • edge0 demo|chat|serve --prefill-ondemand maps onto it; the three subparsers now share one _add_engine_flags helper, and main is a thin wrapper over a testable _build_parser.

Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted):

default (E3b) 5.33 s 265,965 pageins (4.06 GiB read)
--prefill-ondemand 1.06 s 50,611 pageins (0.77 GiB read)
warm-cache: 0.24 s vs 0.37 s (E3b stays the better default when the
checkpoint fits -- hence a flag, not a preset change)

Verified end to end: edge0 chat --name edge0-8b --model-dir ... \ --prefill-ondemand --max-new 16 answers normally, and the slow suite asserts the switch reaches the engine, loads no whole layer, and leaves the next token identical.

An advisory whole-shard madvise(MADV_WILLNEED) before the whole-layer prefill was tried in the same change and dropped: across evicted runs it ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not reproduce reliably enough to ship. Left for a follow-up.

The on-demand prefill path was only reachable from Python (rebuilding the
tier's `LayerOptions` by hand), so the machines that need it most -- ones
whose page cache cannot hold the checkpoint -- had no supported way to
select it.  `#111` made `full_layer_prefill=False` actually work; this
exposes it.

- `ModelConfig` gains `prefill_ondemand`; `from_pretrained` applies it to
  whatever preset the tier ships (`full_layer_prefill=False`), so it stays
  tier-generic and leaves the presets' measured defaults untouched.
- `edge0 demo|chat|serve --prefill-ondemand` maps onto it; the three
  subparsers now share one `_add_engine_flags` helper, and `main` is a thin
  wrapper over a testable `_build_parser`.

Measured on edge0-8b (M4 Pro, 27-token prompt, buffer cache evicted):

  default (E3b)   5.33 s   265,965 pageins (4.06 GiB read)
  --prefill-ondemand  1.06 s    50,611 pageins (0.77 GiB read)
  warm-cache:     0.24 s vs 0.37 s (E3b stays the better default when the
                  checkpoint fits -- hence a flag, not a preset change)

Verified end to end: `edge0 chat --name edge0-8b --model-dir ... \
--prefill-ondemand --max-new 16` answers normally, and the slow suite
asserts the switch reaches the engine, loads no whole layer, and leaves the
next token identical.

An advisory whole-shard `madvise(MADV_WILLNEED)` before the whole-layer
prefill was tried in the same change and dropped: across evicted runs it
ranged from 1.16 s to 4.10 s against a 3.5-5.5 s baseline, i.e. it did not
reproduce reliably enough to ship.  Left for a follow-up.
@linyubupa
linyubupa merged commit fb4cd2c into main Sep 20, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant