Skip to content

qwen35: DSpark speculative decoding (Qwen3.8-27B drafters) - #625

Merged
davide221 merged 37 commits into
mainfrom
qwen38-dspark
Aug 24, 2026
Merged

qwen35: DSpark speculative decoding (Qwen3.8-27B drafters)#625
davide221 merged 37 commits into
mainfrom
qwen38-dspark

Conversation

@davide221

@davide221 davide221 commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

Adds DSpark and DFlash 2 drafter support to the qwen35 family spec-decode loop, targeting the published Qwen3.8-27B checkpoints (RadixArk/Qwen3.8-27B-DSpark: DFlash backbone + low-rank Markov bigram head + confidence head; z-lab/Qwen3.8-27B-DFlash2: DFlash backbone + grouped dynamic convs + candidate selector), plus the decode/prefill/tree optimizations listed below. Merged with current main (SpecLA #606, packed concurrent prefill #595); current validation is summarized below.

What's in here

Spec loop (qwen35_backend.cpp)

  • DSpark markov-corrected greedy chain (fused single-graph variant, non-fused fallback) when the drafter ships DSpark heads.
  • DFlash 2 candidate selector for the greedy chain, and selector-scored DDTree candidates (branch score = logp + selector compatibility with the branch's actual parent, log-softmax-normalized per position). Env gates, default-on: DFLASH_QWEN35_DSPARK, DFLASH_QWEN35_FUSED_DSPARK, DFLASH_QWEN35_DSPARK_TREE, DFLASH_QWEN35_DFLASH2_TREE.
  • Adaptive speculation policy: EMA of accepted drafts/step against the live measured spec/plain step-time ratio, plain-decode bursts below break-even (DFLASH_QWEN35_SPEC_STEP_RATIO, DFLASH_QWEN35_AR_BURST).
  • Target capture layers follow the drafter GGUF's dflash.target_layer_ids; dflash.mask_token_id from the GGUF wins over the family default. Either one wrong silently destroys acceptance.
  • --draft-block-size accepts 2..2x the checkpoint metadata (max 32). DFlash2's acceptance extrapolates past its trained block-8 horizon: block 16 is the recommended code/math config (HumanEval ten-prompt decode 143.6 -> 204.1 tok/s), byte-exact at every width because verification is exact. Widths past 16 hit a step-time cliff and stay rejected; long contexts (committed >= 8192) cap the verify width at 8, draft width untouched.

Draft side: DFlash 2 dynamic convs + selector in the draft graph and converter; optional YaRN rope keys parsed and honored; single-file DSpark releases convert directly.

ggml (HIP) kernels

  • Fused DeltaNet decode: ggml_ssm_conv_step, raw-gate gated_delta_net (gate tensor [dt_bias | A] built by the loader, src[9]/op_params[10], coexists with SpecLA's src[8]/op_params[2]), residual ADD+RMS_NORM+MUL fusion, in-place recurrent state. Launches per decode token 1929 -> ~1140.
  • Tree-mode support for the grouped-cols GDN kernel: DDTree verify no longer falls back to the generic per-token kernel (61-196 us/layer -> ~37 us); DFS parent-state reload verified to 8.4e-7 against the generic kernel.
  • 64x64 MMQ tiles for dense verify widths on RDNA (GGML_CUDA_MMQ_SMALL_TILE; Q4_K keeps main's 128x64), pathological mmq_x=32 tile skipped, FA vec short-KV split, legacy pool MAX_BUFFERS 256 -> 1024.
  • Chunked delta-net prefill path fixed and gated off by default (DFLASH27B_CHUNKED=1 opts in): chunk size 32 keeps the triangular solve on the fast warp kernel, ggml_cont on chunk slices avoids ROCm's pinned-host staging in cublasGemmBatchedEx (39 s -> 0.6 s per 512-token forward, verified ~1e-6). On gfx1201 the sequential fused GDN kernel still wins at ubatch 512, hence off.
  • DFLASH_PREFILL_TIMING=1 prints per-ubatch build/alloc/compute.

Numbers (R9700 / gfx1201, ROCm 7.2, greedy, 300-token gens)

Historical rows below were measured on the pure-IQ4_XS requant. The recommended serving target is now unsloth UD-IQ4_XS as downloaded (13.27 GiB): vs a Q8_0 reference (300x512-chunk wikitext-2 paired --kl-divergence; HumanEval-164 + GSM8K-200 through the serving path) it is Q8-class (+0.7% PPL, KLD 0.018, 94.1% top-1, HE 151/164, GSM8K 177/200 = the reference), and despite a 14% slower plain forward (AR 32.3 vs pure 37.4; sub-4-bit tensors hit slower kernels) it is the FASTEST code config end to end because acceptance with DFlash2 is higher on the less-degraded target: HE-10 block-16 decode 208.1 avg / 227.8 peak, e2e 156.2 (pure: 204.1 / 146.5). Same-file llama.cpp PR#27342 + the same drafter: 54.6 decode / 47.2 e2e, so 3.8x / 3.3x.

Decode, tok/s (code / prose / mixed):

config code prose mixed
plain decode (AR) 36.4 36.6 36.6
DFlash 2 chain (baseline) 111-121 62 123-124
DFlash 2 + DDTree budget 12 (--ddtree --ddtree-budget 12) 126 56 112
DFlash 2 + selector-scored DDTree 12 121 63 116
DSpark chain 48 32 40
upstream llama.cpp 9731ad3 (HIP, tg128) 32.5
hipfire (published, spec off) 36.2

The older raw/selector DDTree results remain useful workload data. The out-of-the-box default is the plain chain at the checkpoint's block 8; --draft-block-size 16 is the recommended code/math config, and block-8 SpecLA with budget 8 is a quality-validated alternative: repeated HumanEval aggregate decode was 163.9 versus 159.0 tok/s for block-8 chain (+3.1%). HumanEval+ was 145/164 versus 143/164, with two gains and no losses.

Prefill: 1036 / 1102 / 1038 tok/s at 512 / 2048 / 6000 prompt tokens (hipfire published 759 / 737 / 663; upstream llama.cpp 1318 / 1277 / 1123). Long context and thinking-mode numbers in the review thread; with max_tokens below hard_limit_reply_budget (4096) the reasoning budget hook force-closes thinking and the reply runs as plain decode.

Production serving config

dflash_server Qwen3.8-27B-UD-IQ4_XS.gguf --draft qwen38-dflash2-q8_0.gguf \
  --target-device hip:0 --draft-device hip:0 \
  --draft-block-size 16 \
  --cache-type-k q8_0 --cache-type-v q8_0

No env vars: the tuned decode path (F32 chain checkpoints, rollback threshold 1, q8 activation memo, KV rotation off for f16/q8_0 caches) is the default as of d461bad; DFLASH_SINGLE_CHAIN_CHECKPOINT_F32=0 restores the legacy F16 path (and ~1.1 GiB of VRAM at 128K ctx). Deploy notes: the disk prefix cache version bump to 2 invalidates every existing .dkv on first start, and finite --fa-window silently costs long-context recall on this model (leave it off; the server warns).

Greedy verification remains target-verified. In the block-8 chain/SpecLA A/B, all ten short replies and 133/164 full-suite replies were byte-identical; the remaining floating-point path differences improved two HumanEval+ tasks and regressed none. DSpark and the 3.6-DFlash drafter keep working (the z-lab Qwen3.6-27B drafter transfers to 3.8 unchanged).

Known limitations / follow-ups

  • The disk-cache identity salt resolves the K-rotation basis from the CLI --cache-type-k string; configs that reach f16/q8_0 via env shorthands (DFLASH27B_KV_K, DFLASH27B_KV_F16) or a family default evaluate the placeholder string instead. All documented configs are consistent; follow-up: feed the resolved type name into the salt.
  • The pre-made incoai/Qwen3.8-27B-DFlash2-GGUF drafter file (llama.cpp PR#27342 layout, arch dflash, different key/tensor names) is not loadable by this server; run server/scripts/convert_dflash_to_gguf.py on the z-lab checkpoint as the quickstart shows. Follow-up: accept the incoai layout via key/tensor aliases in the draft loader.
  • Verify MMQ runs at ~80% of practical bandwidth (457 GB/s at widths 4-16 vs 550 for the N=1 GEMV); it already uses int8 WMMA on RDNA4. Refuted: K-loop weight-tile double buffering (occupancy), 8-warp small tiles (write-back constraint), decode-once wide MMVQ (ties MMQ at N=8).
  • Tree mode still pays F32 node snapshots and un-fused conv copies (~6% of a tree step); a parent-aware fused conv step is the next tree lever.
  • Confidence-gate adaptive block length is wired but not validated; the chain runs with the gate off.
  • Prose acceptance is below break-even with every drafter; the adaptive policy runs prose mostly as plain decode.

Wire the DSpark drafter heads (low-rank Markov bigram correction +
confidence head) into the qwen35 spec-decode loop, so Qwen3.8-27B DSpark
drafters (e.g. RadixArk/Qwen3.8-27B-DSpark) run with full head support:

- spec loop: markov-corrected greedy chain (fused single-graph variant
  with non-fused fallback) replaces plain argmax projection when the
  drafter ships DSpark heads; DDTree candidate top-k gets the markov
  bias too. Env-gated: DFLASH_QWEN35_DSPARK, DFLASH_QWEN35_FUSED_DSPARK,
  DFLASH_QWEN35_DSPARK_TREE (all default on).
- target capture layers now follow the drafter GGUF's
  dflash.target_layer_ids instead of the evenly-spaced derivation; the
  Qwen3.8 drafter is trained on layers 4/16/28/40/52, not 1/16/31/46/61.
- draft loader: dflash.mask_token_id from the drafter GGUF wins over the
  family default (Qwen3.8 drafter uses 248077, default was 248070), and
  optional YaRN rope scaling keys are parsed into DraftWeights.
- draft graph: rope calls honor the drafter's YaRN config (previously
  hardcoded plain NEOX rope).
- Qwen35DFlashTarget exposes lm_head for the fused head path.
- convert_dflash_to_gguf.py: handle single-file DSpark releases (markov/
  confidence heads inline in model.safetensors), transformers>=5 nested
  rope_parameters and dflash_config.mask_token_id, and emit YaRN scaling
  metadata.

The confidence-gate adaptive block length is not wired yet (q_len sizes
the per-request step buffers); the chain runs with the gate off.
Verify/accept now run over v_len (the drafted chain's actual length)
instead of the buffer-sizing q_len, so the DSpark confidence gate's
adaptive block truncation is structurally supported. The gate itself
stays off by default (DFLASH_QWEN35_DSPARK_CONFIDENCE_THRESHOLD=0):
with the RadixArk Qwen3.8 drafter, any threshold in 0.1-0.5 truncates
to the same short chain regardless of value, so the confidence scores
coming out of the shared head path look mis-scaled and need a separate
investigation before the gate can help. threshold=0 is bench-verified
regression-free.
- ggml_ssm_conv_step: one kernel for the causal-conv decode/verify step
  (history window + silu(conv) + in-place history write-back + optional
  rollback window copy) replacing transpose/concat/ssm_conv/silu/cpy.
- ggml_gated_delta_net_set_raw_gates: the GDN kernel applies
  sigmoid(beta) and softplus(alpha + dt_bias) * A itself.
- ADD + RMS_NORM + MUL fusion (residual add materialized alongside the
  normalized output) in the CUDA/HIP graph evaluator.
- legacy pool MAX_BUFFERS 256 -> 1024: LUCE_Q8_MEMO holds ~300 pooled
  buffers per evaluation; a full pool freed in-flight buffers with
  cudaFree and produced illegal memory accesses on long prefills.
Rename the RDNA small-tile macro to GGML_CUDA_MMQ_SMALL_TILE and apply
it to IQ4_XS/Q4_K/Q5_K/Q6_K/Q8_0 in addition to the ROCmFPX formats.
At spec-decode verify widths (N<=16) the 128-row tile leaves a 5120-row
projection with only 40 blocks on a 64-CU gfx1201; 64x64/4-warp tiles
measured +12-23% on those shapes (verify step 43.8 -> 39.7 ms on
Qwen3.8-27B) at ~8% prefill cost.
- loader places attn_gate|attn_qkv and ssm_beta|ssm_alpha back to back
  and exposes zero-copy stacked aliases (L.wqkv_z, L.ssm_ba): one GEMV
  each instead of two (DFLASH_QWEN35_NO_STACK=1 disables).
- FFN uses ggml_swiglu_split so the backend fuses gate/up/GLU into one
  vector kernel at decode.
- DeltaNet block: single l2_norm over the q|k slab, ggml_ssm_conv_step,
  raw-gate gated_delta_net (in place, no state copy), no q/k head repeat
  (the kernel broadcasts). DFLASH_QWEN35_NO_FUSED_KERNELS=1 keeps the
  op-by-op graph for A/B.
- DFLASH_KV_ROTATE=0 skips the FWHT K/Q rotation (precision-neutral with
  q8_0/f16 caches, two fewer launches per attention layer).

Qwen3.8-27B IQ4_XS on R9700: plain decode 30.4 -> 33.8 tok/s with
identical greedy output.
- Qwen35AdaptiveSpecPolicy: EMA of accepted draft tokens per step; below
  0.8*(spec_step_ratio-1) the loop runs a burst of plain-decode steps
  (seed-only verify, no drafter/heads/snapshot/rollback, features still
  captured) and probes again afterwards. Env DFLASH_QWEN35_SPEC_STEP_RATIO
  (default 1.7, 0 disables) and DFLASH_QWEN35_AR_BURST (default 40).
  Low-acceptance prose 28.1 -> 32.4 tok/s, code/mixed unchanged.
- Confidence gate now uses the fused Markov graph and truncates on the
  host; DFLASH_QWEN35_DSPARK_CONF_DEBUG=1 prints per-position scores.
- spec-profile hooks for the chain path (project/snapshot/verify/
  rollback/feature).
launch_fattn was told the vec kernel consumes D keys per step; it walks
nthreads (128) per step, so a 256-key window at head_dim 256 ran as one
block per head. Passing nthreads lets it use two blocks per head plus the
combine pass: Qwen3.8-27B plain decode 34.3 -> 34.6 tok/s on R9700,
identical output.
The first spec step after a plain-decode burst updates the acceptance
EMA with alpha 0.5 so a stream that became predictable leaves plain
decode immediately; step ratio and start value keep the measured best
balance (45.7 / 31.8 / 40.4 tok/s code / prose / mixed).
The break-even acceptance now follows live EMAs of the spec-step and
plain-step wall times (default 1.9 until both are measured), so it is
right for any drafter block size (width-8 DSpark and width-16 DFlash
measure ~1.8 on gfx1201).
DFlash 2 (z-lab/inco, e.g. z-lab/Qwen3.8-27B-DFlash2) is the DFlash
backbone plus a grouped dynamic causal conv around attention and MLP in
every layer and a candidate selector head (top-k lm_head candidates per
block position, one path scored by a low-rank bigram form).

- converter: maps attention_conv/mlp_conv (base kernels F32, kernel
  projections) and candidate_selector tensors, emits dflash2.* metadata,
  reads block_size from dflash_config, emits SWA pattern for drafters
  with causal sliding layers.
- loader: DraftConvWeights per layer, DraftSelectorWeights, shape checks.
- draft graph: conv prepare/finish (two taps over the block, per-element
  base + per-group dynamic coefficient) in both the stateless and the
  cached-KV builders.
- selector chain: top-k via the target's GPU top-k (kMaxK 8 -> 16), one
  cached graph for hproj + codebook row gathers, host path search.
- spec loop uses the selector before the DSpark/argmax paths.

Qwen3.8-27B IQ4_XS on R9700, q8_0 drafter, greedy: 109.9 code / 50.7
prose / 111.8 mixed tok/s (DSpark drafter: 45.6 / 32.4 / 38.6);
avg 5.9-6.0 accepted tokens per 8-token block on code, ~2.7 on prose.
With the 64-row/4-warp tile the mmq_x=32 instantiation runs at 180 GB/s
on gfx1201 (17408x5120 IQ4_XS) against 443 GB/s at mmq_x=16 and 315 at
48, so N=17..32 batches (DDTree budgets, prefill remainders) took 2.4x
longer than N=16 or N=40. Choose the next tile instead.
Resolves the conflicts with SpecLA (#606) and the packed concurrent
prefill work (#595):
- SSM_CONV op_params[0]: 1 stays the SpecLA heavy-light conv, the dflash
  fused conv step now uses 2; the CUDA dispatcher and supports_op handle
  both.
- GDN raw-gate mode no longer uses src[8]/op_params[2] (compact-decode
  slot ids and the SpecLA marker): the loader builds one f32 [dt_bias | A]
  tensor per DeltaNet layer (TargetLayer::ssm_gate_ba, own small buffer)
  and ggml_gated_delta_net_set_raw_gates() attaches it as src[9] with
  op_params[10] = 1.
- gated_delta_net.cu launchers carry both the active-slot arguments and
  the raw-gate parameters.
- build_delta_net_block keeps main's token-axis segment structure; the
  stacked (z|qkv) and (beta|alpha) projections, the fused conv step, raw
  gates, single l2_norm over q|k and the head-broadcast shortcut apply on
  the plain single-sequence chain path only (ragged, compact-decode,
  SpecLA and chunked paths take main's materialized ops).
- MMQ tiles: Q4_K keeps main's 128x64 (LUCEBOX_RDNA_MMQ_Y); IQ4_XS, Q5_K,
  Q6_K, Q8_0 and the ROCmFPX formats keep the 64x64 small tile.
- Spec loop: DSpark Markov top-k stays available inside the non-conditional
  DDTree branch; SpecLA's conditional-draft path and draft-KV flag are
  preserved.

Verified on lucebox8 (R9700): AR 34.6 tok/s, DFlash2 110/54/115
(code/prose/mixed), DSpark 48/32/40.
@jkyamog

jkyamog commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

Bots been following this and testing it. They have noticed a regression TP/NVLink speed has dropped significantly vs old DFlash. They have a fix, but still cleaning and testing.

@jkyamog jkyamog mentioned this pull request Aug 20, 2026
The DDTree verify path fell back to the generic per-token GDN kernel
(61-196 us/layer on gfx1201) because the grouped-cols kernel had no
parent_ids handling. Port the DFS branch-transition state reload into
the grouped kernel: at parent_ids[t] != t-1 the register state shard
reloads from the parent's stored intermediate state (same-thread
read-after-write, no barrier), root-level siblings reset to the
pre-block state, and intermediates are written in tree mode so later
branches can read them.

Verified numerically against the generic tree kernel on a 13-node
branchy tree (max rel diff 8.4e-7, reduction-order noise only);
end-to-end DDTree budget-12 on the R9700 matches text output on
like-for-like runs at +2% tok/s.
DDTree branches were chosen by raw per-position top-k log-probs, ignoring
the DFlash2 selector entirely (it only improved the chain path). Factor
the chain selector into dflash2_score_candidates() + a host-side
branch-conditioned topk, and feed DDTree through build_ddtree_conditional:
each expansion scores candidates as logp + selector compatibility with the
branch's actual parent, log-softmax-normalized per position so cumulative
best-first comparisons across depths stay on a log-prob scale (without the
normalization the raw dot term mis-allocates the budget: code 126 -> 106).
DFLASH_QWEN35_DSPARK_TREE/raw top-k remain the fallback;
DFLASH_QWEN35_DFLASH2_TREE=0 disables.

R9700, budget 12 (code/prose/mixed tok/s): raw tree 126/56/112,
selector tree 121/63/116, chain 112/62/124. The selector tree no longer
collapses on low-acceptance content; chain remains the serving default.
The DFLASH27B_CHUNKED path was unusable on ROCm: a 512-token prefill
forward took 39 s. Two causes, both fixed:
- the [CS, CS] triangular solve at CS = 64 missed ggml-cuda's fast warp
  kernel (k <= 32) and fell into cublasStrsmBatched; CS = 32 keeps the
  solve on the fast path (graph cap raised to 32k nodes to match),
- the per-chunk slices are strided views, which pushed every chunk
  matmul into cublasGemmBatchedEx; on ROCm that API stages its device
  pointer arrays through per-call pinned host allocations (~1 ms of
  hipHostMalloc/hipFree per node). ggml_cont on the sliced operands
  restores the strided-batched fast path.

Result: 39 s -> 0.6 s per 512-token forward, output verified ~1e-6
against the sequential kernel at T = 64..2048 including padded chunks.
Still OFF by default: on gfx1201 the sequential fused GDN kernel wins
(514 ms vs 667 ms per forward; the ~20k-node chunk graph costs more in
launches than it saves in serialization). DFLASH27B_CHUNKED=1 opts in,
and the gate is per-call now, so enabling it no longer disables the
raw-gate fusion on the decode path as a side effect.

Also: env-gated DFLASH_PREFILL_TIMING=1 build/alloc/compute breakdown
per prefill ubatch, and drop the ROCMFP requant experiment script that
slipped into the merge commit (the format was refuted for this target).

R9700 regression check (pure-IQ4_XS target): AR 36.4-36.6, DFlash2 spec
111/62/123 code/prose/mixed, prefill 1036/1102/1038 tok/s @512/2048/6000.
The fp64 RoPE path (required for Qwen3.5-family freq_base=1e7, see the
fp32 precision wall note) computed pow(double, double) per element. On
RDNA4 that libcall made rope_multi the second-largest prefill kernel:
692 us per launch at n_tokens=512 vs 76 us for the fp32 upstream kernel,
~33 ms of a 514 ms 512-token prefill forward.

Replace pow() with binary exponentiation (<= 7 double multiplies for
exponent < 128), keeping the large-freq_base precision to within 1 ulp.

R9700: 512-token prefill forward 514 -> 414-423 ms (prefill ~996 ->
~1225 tok/s); DFlash2 spec decode 112.7/62.2/125.1 code/prose/mixed
(from 111/61.5/123.4) with per-position acceptance identical.
The dense hybrid types compile their MMQ instances with the 64x64
small tile (GGML_CUDA_MMQ_SMALL_TILE), which wins 12-23% at spec-decode
verify widths but re-streams the weights through narrow x-tiles at
prefill widths (measured +16-18% kernel time at N=512 vs the 128x128
upstream shape). The tile shape is baked into every mmq.cuh constexpr
via macros, so one TU can only hold one shape.

Add big-tile twin instances for IQ4_XS/Q4_K/Q5_K/Q6_K/Q8_0 that
re-include mmq.cuh inside namespace lucebox_mmq_big with no tile macro,
giving the 128x128 shape distinct symbols, plus bridge functions and a
runtime dispatch: RDNA4 + ncols_dst >= 256 takes the big tile
(measured crossover: small wins to N=64, tie at 128, big wins 16-18%
at 512); everything else keeps today's path. LUCE_MMQ_BIG_PREFILL=0
disables. gfx1151 behavior unchanged.

R9700, Qwen3.8-27B pure-IQ4_XS: 512-token prefill forward 414 -> 365-374 ms
(prefill ~1225 -> ~1385 tok/s, past upstream llama.cpp's 1366 pp512);
generated output hash-identical; spec decode unchanged at 112.6/62.4/125.4
code/prose/mixed (verify widths never take the big tile).
Plain decode streams every weight byte exactly once per token, so caching
the weight stream in L2 only evicts the activations and KV that other
kernels reuse. Add a nontemporal load variant (HIP sc0/sc1 bypass hints)
and use it for the IQ4_XS weight words in the MMVQ vec_dot; the q8_1
activation loads keep normal caching.

Scope notes from measurement (R9700): the MMVQ weight reads are
wave-contiguous full cache lines, so bypassing L2 is free there (GEMV
550 -> 554 GB/s, AR decode 36.6 -> 37.0 tok/s, +1%). The same hint in
the MMQ tile loader was measured 32% SLOWER (457 -> 310 GB/s at N=8:
tile blocks on different CUs share cache lines, and bypassing L2
amplifies DRAM traffic), so MMQ keeps cached loads. q8_0 qs is only
2-byte aligned and keeps get_int_b2.

Spec decode unchanged at 112.9/62.4/125.4; outputs identical.
@davide221

davide221 commented Aug 21, 2026

Copy link
Copy Markdown
Contributor Author

R9700 / gfx1201 correction and supported follow-up (539ca43):

The Qwen3.8 DFlash2 checkpoint declares block_size=8, and its published configuration uses seven speculative tokens per verification step. The earlier width-12/16 numbers were useful experiments, but those widths exceed the checkpoint's trained/published horizon and should not be production recommendations. The server now rejects --draft-block-size values above the drafter metadata; the flag remains available for safe down-tuning from 8.

Supported measurements on Qwen3.8-27B IQ4_XS + Q8_0 DFlash2:

Setting HE 10-prompt aggregate decode HumanEval+ pass@1
plain AR 36.9 tok/s
block-8 chain 159.0 tok/s 143/164
block-8 SpecLA, budget 8 163.9 tok/s 145/164

SpecLA's gain was reproducible across repeated A/B runs (+3.1% aggregate decode). It gained two HumanEval+ tasks and lost none; all 10 short A/B replies and 133/164 full-suite replies were byte-identical.

Production flags are now:

--specla --ddtree-budget 8

with the checkpoint metadata block size (no --draft-block-size override). The rebuilt gfx1201 server passed test_feature_gate, rejected a live width-12 launch with the checkpoint-specific [2, 8] bound, reproduced 164.45 aggregate tok/s, and is serving healthily on lucebox8.

DeanoC added a commit to GeometricAGI/geo-lucebox that referenced this pull request Aug 21, 2026
GQH has no MMVQ. The fused matvec hook was still gated on
LUCE_MMVQ_MAX_NCOLS (default 3), so an 8-wide DFlash2 verify dequantized
the whole weight matrix per step. On gfx1201 that was 6x slower than a
fused N=5 and made spec slower than AR (HE 24 tok/s at AL 0.94).

Stop using that cap for GQH. Exact-width arms now cover ncols 2..8
(DFlash2 default block size); 9..16 take the generic instantiation.
ctest -R gqh 25/25, including fused nvec=8.

Measured on R9700, GQH-Q3KXL + DFlash2, same flags as Luce-Org#625:
  HE-10  24.2 -> 96.1 tok/s aggregate (4.0x), AL unchanged
  code   23.9 -> 92.4, mixed 12.0 -> 46.9, prose 8.2 -> 29.6
@davide221

Copy link
Copy Markdown
Contributor Author

R9700 pure-AR follow-up (commit 739cf84)

  • Qwen3.8-27B IQ4_XS, pure AR, same HumanEval prompt: original 8-wave kernel 37.230 tok/s vs tuned 6-wave kernel 37.413 tok/s in a hot-state 4-run A/B (+0.49%).
  • The generated response was byte-identical to the original-kernel/cache baseline; four tuned repeats also produced the same SHA-256.
  • A full cache sweep found K=q8_0, V=f16 fastest: 37.21 tok/s before the wave change, vs 36.90 for q8/q8. Disabling LUCE_Q8_MEMO fell to 36.65.
  • Pure-AR profiling attributes about 89% of GPU kernel time to IQ4_XS MMVQ. Wider 64-bit weight loads and a two-block occupancy hint were neutral and were removed.
  • Existing classical quantizations did not approach IQ4_XS: Q4_0 32.76, Q2_K 32.04, IQ2_XXS 30.65, Q4_K_S 29.45 tok/s. Their extra decode/unpack work costs more than their bandwidth savings on gfx1201.

This does not claim 40 tok/s exact AR. From the measured bottleneck, 40 needs roughly another 7% end-to-end / about 8% in the already bandwidth-bound IQ4_XS kernel; that likely requires a new simple sub-4-bit format or hardware clock/power work, not another stock GGUF quant.

Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Add mixed AR and speculative execution for concurrent Qwen3.8 requests. Preserve PR Luce-Org#625 grouped GDN and SpecLA ownership while restricting raw-gate fusion to dense plain chains.

Record selector, cost-profile, transition, and request telemetry needed to profile later policy changes.
Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Document the PR Luce-Org#625 integration, benchmark protocol, two-replicate reference result, and the remaining policy work.
Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Document the PR Luce-Org#625 integration, benchmark protocol, two-replicate reference result, and the remaining policy work.
Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Document the PR Luce-Org#625 integration, benchmark protocol, two-replicate reference result, and the remaining policy work.
Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Document the PR Luce-Org#625 integration, benchmark protocol, two-replicate reference result, and the remaining policy work.
Graffioh added a commit to Graffioh/lucebox-hub that referenced this pull request Aug 21, 2026
Document the PR Luce-Org#625 integration, benchmark protocol, two-replicate reference result, and the remaining policy work.
One ggml_dflash_dyn_conv node (SSM_CONV variant 3) replaces the ~19-op
per-site expansion at all 20 apply sites of the draft step. rn intrinsics
keep the result bit-identical to the unfused node chain; validated on
Qwen3.8-27B DFlash2 (identical accepts and output hashes, HumanEval and
3-domain cells). DFLASH_DYN_CONV_FUSED=0 restores the unfused graph.
Greedy chain verification keeps output byte-identical to plain decode at
any draft width, so widening only risks acceptance depth. Measured on the
Qwen3.8-27B DFlash2 checkpoint (R9700): completions byte-identical across
widths 8/10/12/16/20/24; HumanEval ten-prompt decode 143.6 -> 204.1 tok/s
average (peaks above 230) and math 129.2 -> 152.6 at --draft-block-size 16;
past 16 the step time cliffs with no commit gain, so the guard now sits at
2x the horizon. test_feature_gate updated, README rewritten accordingly.
Review findings on the adaptive plain-decode bursts:
- bursts stay off under sampled-verify (a burst committed greedy argmax
  tokens without passing through the sampler chain)
- error paths no longer restore a KV/state snapshot that a burst step
  never took (the copy could be a whole burst stale, desyncing the
  recurrent state from the KV on the next request)
- pending tool-call hints end a burst immediately instead of being
  ignored for up to 40 tokens
- 1-token burst steps no longer count as q_len-wide draft steps in the
  accept-rate that steers the PFlash residency bandit
- the drafter GGUF's own mask token id is now honored when building the
  noise block (the loader-resolved field was written but never read)
- --draft-block-size and the legacy drafter YaRN config are re-applied
  after park/unpark, so neither silently reverts mid-process
- the cached DFlash2 selector projection graph is invalidated when draft
  weights are freed; it was keyed on the DraftWeights address, which is
  stable across park/unpark, so a reload left it computing on freed
  weight tensors
- geometric_draft_topk falls back to the CPU path for K values without
  an instantiated kernel (9-11, 13-15); the silent default returned
  success with uninitialized scratch as token ids
…arden drafter YaRN metadata

- the 64x64 small-tile MMQ shapes are a gfx1201 measurement and their
  big-tile twin dispatch is RDNA4-only, but the compile guard applied to
  RDNA3 too: gfx1151 silently regressed IQ4_XS/Q5_K/Q6_K/Q8_0 prefill.
  RDNA3 keeps stock shapes on host and device now.
- Metal/Vulkan reject the fused SSM_CONV variants and raw-gate GDN
  (src[9]) instead of running the stock kernels on the wrong operand
  layout; the CPU GDN path asserts on raw-gate mode; ggml_ssm_conv_step
  rejects tap counts its kernels do not instantiate
- the drafter GGUF YaRN path ignores a factor without
  original_context_length (corr-dims would collapse through log(0)); the
  converter accepts the legacy rope_scaling.type spelling and falls back
  to the top-level original_max_position_embeddings
mrciffa added 4 commits August 24, 2026 02:11
Long context. Verify attention runs through ggml's tile kernel, whose cost
scales with the number of query rows and with KV length. A wide draft block
is a large win on short prompts and a liability once the KV is big, because
the extra rows multiply a term that is itself growing. Measured on one
Radeon AI PRO R9700 with Qwen3.8-27B IQ4_XS + DFlash2 (decode tok/s):

    prompt   block 16   block 12   block 8
    26,728       44.2       44.0      42.5   <- crossover
    39,861       32.7       40.2      46.0
    81,365       23.3          -      27.8

Narrow the block past a context threshold (default 32768 tokens, override
with DFLASH_LONGCTX_BLOCK_TOKENS, 0 disables). The rule reads context length
only: it never inspects the observed acceptance rate, so the width chosen for
a given prompt is reproducible. Below the threshold nothing changes.

After, on the same box: HumanEval-10 is 138.59 tok/s with output sha
a4467e9d, unchanged; 26,728 tokens 43.9 (unchanged, below threshold); 39,861
tokens 32.7 -> 45.5 (+39%); 53,264 29.1 -> 33.0; 81,365 23.3 -> 25.5.

Review fixes in the same change:

- mmq.cuh: arch-gating the small MMQ tiles reverted the ROCmFPX instances on
  RDNA3 from the tuned 64x64/4-warp tile to the stock 128x128, contradicting
  the comment directly above which records that the small tile is faster on
  gfx1151 for those formats. The gate conflated two different reasons for the
  small tile. Split it: GGML_CUDA_MMQ_SMALL_TILE_RDNA4_ONLY marks the dense
  hybrid types, whose measurement is gfx1201 and whose big-tile twin dispatch
  is RDNA4-only; ROCmFPX keeps the small tile on both RDNA generations.
  Applied to the generator so the checked-in instances stay reproducible.

- rope.cu: rope_theta_fp64's binary-exponentiation loop never terminates for
  a negative exponent, because `e >>= 1` on a negative int is an arithmetic
  shift that never reaches 0, which would hang the GPU. No current caller
  passes one, but the helper takes a plain int and the pow() it replaced was
  total over that domain, so handle it.

- graph_builders.cpp: only the fused target-step builder was raised to 32768
  nodes. build_layer_step, build_layer_prefn_step and
  build_hybrid_full_layer_step construct the same delta-net blocks for
  layer-split and tensor-parallel placements and still capped at 16384, so a
  chunked delta-net prefill at a 512-token ubatch would abort there.

- dflash2_head.cpp: the thread_local SelectorGraph had no destructor, so a
  worker thread that drafted and then exited leaked its ggml context and the
  gallocr's device buffer. One leaked GPU allocation per thread that ever
  drafted.
The first cut put the threshold at 32768, which left a visible dip in the
decode curve: 13,148 tokens ran at 51.9 tok/s and 26,728 at 43.4, both still
at full width on an already-large KV, while 39,861 recovered to 45.5 once
narrowing engaged. Decode is commit depth over step time, and the engine's
own counters show the step is what moves: 82.0 ms at 26,728 against 81.1 ms
at 39,861, flat despite 1.5x the context, because dropping 17 verify rows to
9 offsets the extra KV almost exactly.

Measuring the threshold directly (decode tok/s, same prompts):

    prompt    wide (16)   narrowed (8)
     13,148        51.9           54.9
     26,728        43.4           46.9
     39,861        32.7           45.5

Narrowing wins everywhere it was tested above ~8K, so move the default down.
Worth noting that narrowing inside a block-16 drafter beats configuring the
drafter at block 8 outright (46.9 vs 42.5 at 26,728), because acceptance
holds up better: 0.414 against 0.375.

8192 is deliberately conservative rather than the lowest value measured. It
sits above every prompt in the short-context benchmarks this engine is tuned
for, so the high-acceptance completion workloads that make the wide block pay
are untouched. The prompts used here are question-answering over a long file,
which accepts far less than code completion does, and the two should not be
collapsed into one threshold on this evidence alone.

After, on one R9700: HumanEval-10 is 137.48 tok/s with output sha a4467e9d,
unchanged; the long curve reads 67.0 / 55.1 / 46.7 / 45.4 at 6,208 / 13,148 /
26,728 / 39,861 and no longer dips.
…eded

Running Qwen3.8-27B at the speeds this PR reports required four environment
variables on the command line. That is a bad default: the tuned path is the
one everybody should get, and requiring opt-in means the likely outcome is
someone running the slow path by accident and concluding the engine is slow.
Fold all four into the defaults and drop the env var this PR had added.

- Exact F32 chain checkpoints, and the rollback-from-one-accepted-token they
  enable, are now on by default (was DFLASH_SINGLE_CHAIN_CHECKPOINT_F32 plus
  DFLASH_FAST_ROLLBACK_THRESHOLD=1). Setting the former to 0 restores the
  legacy F16 replay path, including its threshold of 5.

- q8_1 activation memoisation is on by default (was LUCE_Q8_MEMO=1).
  LUCE_Q8_MEMO=0 opts out.

- Graph-level K rotation is off by default for the f16 and q8_0 caches we
  serve with, where it is precision-neutral and costs two extra launches per
  attention layer (was DFLASH_KV_ROTATE=0). Narrower cache types, where
  spreading outliers actually buys accuracy, keep it. DFLASH_KV_ROTATE=1
  forces it on.

- The long-context draft narrowing threshold added earlier in this PR is now
  a named constant rather than DFLASH_LONGCTX_BLOCK_TOKENS.

Also fixes the narrowing floor: it read std::max(DFLASH27B_DRAFT_BLOCK_SIZE,
q_len_cfg / 2), and that constant is 16, so max(16, 8) never narrowed and the
long-context gain was silently absent. Floor at the DFlash2 checkpoint's own
published block size of 8 instead.

Verified with a launch command carrying no environment variables at all:
HumanEval-10 138.00 tok/s with output sha a4467e9d, unchanged; long context
55.4 / 46.9 / 45.5 tok/s at 13,148 / 26,728 / 39,861, matching what the four
env vars produced.
Qwen3.8-27B with the DFlash2 block-diffusion drafter is the fastest thing
this engine serves and the configuration the recent work is measured on, so
it replaces Qwen3.6-27B in the speedup table, the drafter list, the
per-device settings matrix, and both quickstarts.

The speedup table now carries two Qwen3.8 rows rather than one, because a
single number invites the obvious objection. 6.1x is decode against plain
llama.cpp, the comparison the rest of that table uses. 3.5x is decode against
llama.cpp speculating with the same DFlash2 drafter through its own
--spec-type draft-dflash, which is the harder and fairer fight.

Run commands lose their environment variables: the tuned decode path is now
the default in the engine, so the documented command is the fast one.

Three Qwen3.6 mentions stay on purpose. Qwen3.6 35B-A3B is a different
model and still supported; the --paged-attention and
--target-split-fast-rollback rows cite the target those numbers were
measured on, and relabelling them would misattribute the measurement.
@davide221

Copy link
Copy Markdown
Contributor Author

@cubic-dev-ai review

@cubic-dev-ai

cubic-dev-ai Bot commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review

@davide221 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

11 issues found across 58 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="server/src/common/draft_block_size.h">

<violation number="1" location="server/src/common/draft_block_size.h:16">
P2: When a drafter GGUF declares a block size at least `INT_MAX / 2 + 1`, `2 * checkpoint_block_size` overflows before this predicate returns. Compare without multiplication or reject out-of-range metadata so malformed checkpoints cannot trigger undefined behavior during startup.</violation>
</file>

<file name="server/src/qwen35/gguf_target_loader.cpp">

<violation number="1" location="server/src/qwen35/gguf_target_loader.cpp:858">
P2: When a later load step fails after stacked/gate tensors are created, the function releases only `out.buf` and leaves `out.stack_ctx`/`out.gate_ctx`/`out.gate_buf` allocated. Extend rollback cleanup to free these auxiliary allocations on every early-return path, not just in `free_target_weights()`.</violation>
</file>

<file name="server/src/qwen35/qwen35_backend.cpp">

<violation number="1" location="server/src/qwen35/qwen35_backend.cpp:3872">
P2: `out_accept_rate` is now normalized by `n_spec_steps * q_len`, but offered candidates vary by step (`v_len`) and tree steps never increment `n_spec_steps`. This misreports acceptance and can mis-steer the residency bandit. Track and divide by the actual offered candidate count per step instead of fixed `q_len`.

(Based on your team's feedback about offered-candidate denominator for DSpark acceptance.) .</violation>
</file>

<file name="server/deps/llama.cpp/ggml/src/ggml-cuda/mmq_big.h">

<violation number="1" location="server/deps/llama.cpp/ggml/src/ggml-cuda/mmq_big.h:21">
P2: This `const void *` bridge causes the generated wrappers to dereference a live `::mmq_args` object as the distinct `lucebox_mmq_big::mmq_args` type, which violates strict-aliasing rules despite identical fields. Share `mmq_args` through a common header or copy the bytes into a namespaced object before invoking the template.</violation>
</file>

<file name="server/deps/llama.cpp/ggml/src/ggml-metal/ggml-metal-device.m">

<violation number="1" location="server/deps/llama.cpp/ggml/src/ggml-metal/ggml-metal-device.m:1186">
P1: When Qwen35 tree verification runs on Metal, `ggml_ssm_conv_tree` keeps mode 0, so this predicate still selects Metal. The Metal launcher ignores `src[2]` and computes the DFS-neighbor window instead of the parent-chain window; also require `op->src[2] == NULL` here.</violation>
</file>

<file name="server/deps/llama.cpp/ggml/src/ggml-cpu/ops.cpp">

<violation number="1" location="server/deps/llama.cpp/ggml/src/ggml-cpu/ops.cpp:9336">
P1: When a CPU backend executes `ggml_ssm_conv_tree`, this guard passes because tree mode keeps `op_params[0]` at zero, but the CPU kernel ignores `src[2]` and produces incorrect branch windows. Reject tree mode here by checking `dst->src[2]`, or implement the parent-chain convolution on CPU.</violation>

<violation number="2" location="server/deps/llama.cpp/ggml/src/ggml-cpu/ops.cpp:10724">
P1: When CPU receives a tree, persistent, or SpecLA GDN graph, this raw-gate guard passes because those modes use `src[6]`/`src[7]`, then the CPU kernel silently ignores their metadata. Reject unsupported `src[6]`/`src[7]` modes here or add equivalent CPU implementations.</violation>
</file>

<file name="server/deps/llama.cpp/ggml/src/ggml-cuda/mmq.cu">

<violation number="1" location="server/deps/llama.cpp/ggml/src/ggml-cuda/mmq.cu:84">
P2: MoE calls use `args.ncols_dst` as total routed rows, not the actual prefill width used for MMQ tiling. This enables the big kernel for small batches (for example, batch 128 with top-k 2), defeating the stated crossover guard and potentially regressing prefill performance; gate this dispatch on `args.ncols_max` instead.</violation>
</file>

<file name="server/src/server/server_main.cpp">

<violation number="1" location="server/src/server/server_main.cpp:314">
P2: `--draft-block-size` is hard-capped to 32 in argument parsing, but backend validation supports values up to `2 * checkpoint_block_size`. This can reject valid operator settings before model metadata is even loaded. Remove the fixed `> 32` cap here and leave upper-bound enforcement to the backend metadata check.</violation>
</file>

<file name="server/deps/llama.cpp/ggml/src/ggml-cuda/mmvq.cu">

<violation number="1" location="server/deps/llama.cpp/ggml/src/ggml-cuda/mmvq.cu:2700">
P1: When CUDA graph replay is enabled, this default turns on q8 memoization without preserving its buffers for the graph lifetime. The next graph-compute call frees `luce_q8_memo` before replay, leaving captured q8 kernel pointers vulnerable to pool reuse and incorrect output or memory corruption; keep these allocations alive until the graph is invalidated or disable this memo path for graph replay.</violation>
</file>

<file name="server/src/common/dflash2_head.cpp">

<violation number="1" location="server/src/common/dflash2_head.cpp:99">
P2: The cached selector graph holds raw pointers to the draft's selector tensors but is only invalidated by the address check, and `dflash2_selector_graph_invalidate()` is never called when draft weights are freed. `Qwen35Backend::free_drafter()` (qwen35_backend.cpp:1207) frees the draft without invalidating, so a park/unpark reload that lands DraftWeights at the same address reuses the thread-local graph with dangling tensor pointers (use-after-free / garbage scores). Call `dflash2_selector_graph_invalidate()` wherever the draft weights are freed (e.g., in `free_drafter()`).</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

if (op->op_params[0] != 0) {
return false;
}
return has_simdgroup_reduction;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When Qwen35 tree verification runs on Metal, ggml_ssm_conv_tree keeps mode 0, so this predicate still selects Metal. The Metal launcher ignores src[2] and computes the DFS-neighbor window instead of the parent-chain window; also require op->src[2] == NULL here.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At server/deps/llama.cpp/ggml/src/ggml-metal/ggml-metal-device.m, line 1186:

<comment>When Qwen35 tree verification runs on Metal, `ggml_ssm_conv_tree` keeps mode 0, so this predicate still selects Metal. The Metal launcher ignores `src[2]` and computes the DFS-neighbor window instead of the parent-chain window; also require `op->src[2] == NULL` here.</comment>

<file context>
@@ -1179,14 +1179,20 @@ bool ggml_metal_device_supports_op(ggml_metal_device_t dev, const struct ggml_te
+            if (op->op_params[0] != 0) {
+                return false;
+            }
+            return has_simdgroup_reduction;
         case GGML_OP_SSM_SCAN:
             return has_simdgroup_reduction;
</file context>
Suggested change
return has_simdgroup_reduction;
return has_simdgroup_reduction && op->src[2] == NULL;

Comment on lines +10724 to +10725
GGML_ASSERT(dst->op_params[10] == 0 && dst->src[9] == NULL &&
"raw-gate gated_delta_net is CUDA/HIP only");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When CPU receives a tree, persistent, or SpecLA GDN graph, this raw-gate guard passes because those modes use src[6]/src[7], then the CPU kernel silently ignores their metadata. Reject unsupported src[6]/src[7] modes here or add equivalent CPU implementations.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At server/deps/llama.cpp/ggml/src/ggml-cpu/ops.cpp, line 10724:

<comment>When CPU receives a tree, persistent, or SpecLA GDN graph, this raw-gate guard passes because those modes use `src[6]`/`src[7]`, then the CPU kernel silently ignores their metadata. Reject unsupported `src[6]`/`src[7]` modes here or add equivalent CPU implementations.</comment>

<file context>
@@ -10716,6 +10718,12 @@ static void ggml_compute_forward_gated_delta_net_f32(
+    // Raw-gate mode (loader-built [dt_bias|A] in src[9], op_params[10])
+    // is CUDA/HIP only; the CPU path would silently use the raw values
+    // as final gates.
+    GGML_ASSERT(dst->op_params[10] == 0 && dst->src[9] == NULL &&
+                "raw-gate gated_delta_net is CUDA/HIP only");
+
</file context>
Suggested change
GGML_ASSERT(dst->op_params[10] == 0 && dst->src[9] == NULL &&
"raw-gate gated_delta_net is CUDA/HIP only");
GGML_ASSERT(dst->op_params[10] == 0 && dst->src[9] == NULL &&
dst->src[6] == NULL && dst->src[7] == NULL &&
"raw-gate/tree/persistent gated_delta_net is not supported on CPU");

Comment on lines +9336 to 9337
GGML_ASSERT(ggml_get_op_params_i32(dst, 0) == 0 && "ggml_ssm_conv_step is not supported on CPU");
const ggml_tensor * src0 = dst->src[0]; // conv_x

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When a CPU backend executes ggml_ssm_conv_tree, this guard passes because tree mode keeps op_params[0] at zero, but the CPU kernel ignores src[2] and produces incorrect branch windows. Reject tree mode here by checking dst->src[2], or implement the parent-chain convolution on CPU.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At server/deps/llama.cpp/ggml/src/ggml-cpu/ops.cpp, line 9336:

<comment>When a CPU backend executes `ggml_ssm_conv_tree`, this guard passes because tree mode keeps `op_params[0]` at zero, but the CPU kernel ignores `src[2]` and produces incorrect branch windows. Reject tree mode here by checking `dst->src[2]`, or implement the parent-chain convolution on CPU.</comment>

<file context>
@@ -9332,6 +9332,8 @@ void ggml_compute_forward_flash_attn_back(
         const ggml_compute_params * params,
         ggml_tensor * dst) {
+    // dflash: the fused step mode (ggml_ssm_conv_step) is CUDA/HIP only
+    GGML_ASSERT(ggml_get_op_params_i32(dst, 0) == 0 && "ggml_ssm_conv_step is not supported on CPU");
     const ggml_tensor * src0 = dst->src[0]; // conv_x
     const ggml_tensor * src1 = dst->src[1]; // conv1d.weight
</file context>
Suggested change
GGML_ASSERT(ggml_get_op_params_i32(dst, 0) == 0 && "ggml_ssm_conv_step is not supported on CPU");
const ggml_tensor * src0 = dst->src[0]; // conv_x
GGML_ASSERT(ggml_get_op_params_i32(dst, 0) == 0 && dst->src[2] == NULL &&
"non-standard ssm_conv modes are not supported on CPU");

static const bool luce_q8_memo_on = []() {
const char * e = getenv("LUCE_Q8_MEMO");
return e && e[0] == '1' && e[1] == '\0';
return !(e && e[0] == '0' && e[1] == '\0');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: When CUDA graph replay is enabled, this default turns on q8 memoization without preserving its buffers for the graph lifetime. The next graph-compute call frees luce_q8_memo before replay, leaving captured q8 kernel pointers vulnerable to pool reuse and incorrect output or memory corruption; keep these allocations alive until the graph is invalidated or disable this memo path for graph replay.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At server/deps/llama.cpp/ggml/src/ggml-cuda/mmvq.cu, line 2700:

<comment>When CUDA graph replay is enabled, this default turns on q8 memoization without preserving its buffers for the graph lifetime. The next graph-compute call frees `luce_q8_memo` before replay, leaving captured q8 kernel pointers vulnerable to pool reuse and incorrect output or memory corruption; keep these allocations alive until the graph is invalidated or disable this memo path for graph replay.</comment>

<file context>
@@ -2688,9 +2691,13 @@ void ggml_cuda_mul_mat_vec_q(
     static const bool luce_q8_memo_on = []() {
         const char * e = getenv("LUCE_Q8_MEMO");
-        return e && e[0] == '1' && e[1] == '\0';
+        return !(e && e[0] == '0' && e[1] == '\0');
     }();
     const size_t q8_bytes = ne13*ne12 * ne11*ne10_padded * sizeof(block_q8_1)/QK8_1;
</file context>

Comment thread server/src/draft/draft_gguf_loader.cpp
const int n_rows_pred = 1 + n_cand * K;
SelectorGraph & g = selector_graph();
const uint64_t cur_gen = s_selector_generation.load(std::memory_order_relaxed);
if (!g.ctx || g.dw != &dw || g.backend != backend || g.n_cand != n_cand ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The cached selector graph holds raw pointers to the draft's selector tensors but is only invalidated by the address check, and dflash2_selector_graph_invalidate() is never called when draft weights are freed. Qwen35Backend::free_drafter() (qwen35_backend.cpp:1207) frees the draft without invalidating, so a park/unpark reload that lands DraftWeights at the same address reuses the thread-local graph with dangling tensor pointers (use-after-free / garbage scores). Call dflash2_selector_graph_invalidate() wherever the draft weights are freed (e.g., in free_drafter()).

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At server/src/common/dflash2_head.cpp, line 99:

<comment>The cached selector graph holds raw pointers to the draft's selector tensors but is only invalidated by the address check, and `dflash2_selector_graph_invalidate()` is never called when draft weights are freed. `Qwen35Backend::free_drafter()` (qwen35_backend.cpp:1207) frees the draft without invalidating, so a park/unpark reload that lands DraftWeights at the same address reuses the thread-local graph with dangling tensor pointers (use-after-free / garbage scores). Call `dflash2_selector_graph_invalidate()` wherever the draft weights are freed (e.g., in `free_drafter()`).</comment>

<file context>
@@ -0,0 +1,222 @@
+    const int n_rows_pred = 1 + n_cand * K;
+    SelectorGraph & g = selector_graph();
+    const uint64_t cur_gen = s_selector_generation.load(std::memory_order_relaxed);
+    if (!g.ctx || g.dw != &dw || g.backend != backend || g.n_cand != n_cand ||
+        g.K != K || g.gen != cur_gen) {
+        selector_graph_free(g);
</file context>

Comment thread server/deps/llama.cpp/ggml/src/ggml.c Outdated
Comment thread server/src/common/chain_rollback_policy.h
Comment thread server/src/server/server_main.cpp Outdated
Comment thread README.md Outdated
mrciffa added 2 commits August 24, 2026 12:31
Review of the previous commits found two ways the long-context narrowing
could break, both from narrowing q_len itself.

q_len is not just a host buffer size. The drafter graphs are built from
dw_.block_size (dflash_draft_kv.cpp:26, dflash_draft_graph.cpp:51), which
nothing narrowed, so the callers wrote only the first q_len rows of a
block_size-row input tensor while draft_kv_begin_step still marked every
column visible to every query row. The kept rows then attended to rows that
were either permanently zero (draft-KV path, which clears its buffer) or
uninitialised device memory with garbage RoPE positions (legacy path). Output
stayed correct because the target verifies, but the drafter was being fed
noise, and it still paid for a full-width forward, so the draft-side saving
never existed.

Worse, the IPC drafter validates the payload against a fixed block_size
(dflash_draft_ipc.cpp:227), so a narrowed noise_embed made propose() return
false and the request fail with DecodeFailed. Any prompt at or above the
threshold with --remote-draft would have errored out.

Cap the verify batch instead. The drafter proposes its full q_len block, so
every drafter buffer, graph and IPC contract keeps the shape it was built
with; the trailing drafted tokens are simply not checked this step. That is
where the win came from anyway: the target forward is what carries the
tile-kernel attention cost that grows with context.

Same numbers as before on one R9700, so nothing was lost: HumanEval-10
138.37 tok/s with output sha a4467e9d, and 54.9 / 46.5 / 45.4 tok/s at
13,148 / 26,728 / 39,861 tokens.

Also in this commit:

- The acceptance rate divided accepted tokens by n_spec_steps * q_len, which
  understated it once fewer than q_len positions were checked. Use the
  verified width.

- DISK_CACHE_VERSION 1 -> 2. The KV-rotation default change alters the basis
  K rows are stored in for f16/q8_0 caches, and the cache identity is derived
  from tensor names, types and shapes only, so a .dkv written by an older
  binary would have been accepted and read with an unrotated Q. Wrong
  attention over the whole cached prefix, silently.

- Document the VRAM cost of defaulting to F32 chain checkpoints: measured
  +1.11 GiB on Qwen3.8-27B at 128K context (22.47 vs 21.36 GiB total), with
  the opt-out named for when that headroom matters.
Flipping the tuned decode path to on-by-default left several comments
asserting the opposite, which is worse than no comment: the F32 checkpoint
site still said "preserve the established F16 default ... opt in with an
explicit environment flag", kvflash_qk.h still called K rotation the default
for Q8_0 when it is now off for exactly that type, and the q8_1 memoisation
comment still led with LUCE_Q8_MEMO=1.

Adds the two env vars the flips turned into opt-outs to the environment
index, with their new sense, since neither was listed.
mrciffa and others added 5 commits August 24, 2026 13:03
The version bump that rejects caches written in the old K-rotation basis
left this assertion pinned at 1, which failed the CUDA build job's ctest run
while every GPU job passed.
A finite --fa-window caps the full-attention layers to a sliding window, so
any content earlier than the window is invisible to them. Nothing reports
this. The model still answers, it simply cannot see the head of a long
prompt, so the failure looks like a model quality problem rather than a
configuration one.

Found while benchmarking long context on a Radeon AI PRO R9700 with
Qwen3.8-27B: with --fa-window 2048, a label planted at the top of the prompt
was recalled at 1,556 tokens and missed at every longer length, with no
diagnostic anywhere. Dropping the flag restored recall at all lengths up to
81K, and cost nothing measurable (44.4 vs 45.7 tok/s at a 27K prompt), so the
silent tradeoff was not even buying speed on this hardware.

Warn once, the first time a prompt actually outgrows the window, naming both
numbers so the cause is unambiguous. Verified on the same box: the warning
fires exactly once, on the 6,208-token request that does lose the label, and
does not fire for the 1,556-token request that retrieves it.

Co-authored-by: mrciffa <davide@lucebox.com>
…652)

* qwen35: span the whole KV pool for slot-mapped verify attention

With --kvflash the KV cache lives at pool slots and verify_batch builds its
attention mask in slot space over the entire pool. The flash-attention view,
however, was still sized from the logical context length (kv_start + n_tokens,
rounded to the 256 stride). Those two disagree: slot indices are not ordered
by logical position, so the view could end below slots the mask still marked
visible. Attention then read rows that were never written and the softmax row
degenerated, which surfaced as an argmax of -1 for every verify row past the
first.

The symptom was a hard failure. do_spec_decode saw the invalid seed, fell back
to plain decode, and that fallback inherited the same state and failed too, so
the request returned decode_failed. Reproduced on a Radeon AI PRO R9700 with
Qwen3.8-27B and --kvflash auto (16384-token pool): prompts of 1556 and 6208
tokens died on the first speculative step, while 13148 and 26728 happened to
survive because their logical extent covered the slots in use.

Span the pool instead. The mask is sized from the same pool and is what
restricts which slots are readable, so this is the bound that matches the
caller's contract. The condition is scoped to the slot-mapped path: a set_rows
KV write together with an explicit mask is a pair only kvflash verify
produces, since the non-kvflash step-invariant write requires no mask and the
paged path never reaches this branch.

Measured on the same box, --kvflash auto, block-16 DFlash2, greedy, prompts
that previously failed now complete: 1556 tokens 59.5 tok/s, 6208 tokens 43.3
tok/s, both recalling a label planted at the top of the context; 13148 and
26728 are unchanged at 39.0 and 38.6. Zero invalid-seed events across the
sweep. The default (non-kvflash) path is untouched: HumanEval-10 is 144.61
tok/s with output sha a4467e9d, identical to before the change.

* fix: address Cubic review findings for PR 652

* fix: address Cubic follow-up review

---------

Co-authored-by: mrciffa <davide@lucebox.com>
- spec-decode: burst steps now snapshot when a stall-floor or budget
  force-close hook is armed; those paths restore+replay mid-step and
  previously copied back state up to a whole burst stale (silent GDN
  corruption on long prose under the default adaptive policy)
- disk cache: fold the effective FWHT K-rotation basis into the identity
  salt (shared helper common/kv_rotation.h); a cache written under one
  basis is never adopted by a session using the other
- ggml-cpu: supports_op now rejects the CUDA-only SSM_CONV modes 2/3 and
  raw-gate GDN so hybrid placements route to GPU instead of asserting
- mmvq: skip the q8_1 activation memo while concurrent streams are active
  (fork/join graph optimizer could read a memo entry mid-quantize)
- draft loader: validate conv_group_size / selector_top_k / mask_token_id
  / selector codebook vocab instead of SIGFPE or device OOB on malformed
  GGUFs
- target loader: pair-reorder uses exact suffix matches; suffix-superset
  names (e.g. .weight.scale) no longer drop a tensor from the ordering
- long-ctx verify cap clamped to the checkpoint block so the accept-rate
  denominator matches the positions actually drafted
- --draft-block-size help/error text now states the 2..2x-metadata rule
- docs: ENVIRONMENT.md inventory for the new env vars, README harness
  example uses the DFlash2 drafter it downloads, stale comments fixed

Validated on lucebox8 (R9700, gfx1201): byte-identical spec-decode output
old vs new on code/prose/mixed at block 16, speeds unchanged;
test_server_unit and test_feature_gate pass.
…ickstart

- both quickstarts now download unsloth/Qwen3.8-27B-GGUF UD-IQ4_XS and
  serve it as downloaded: measured Q8_0-class (KLD 0.018, top-1 94.1%,
  HumanEval 151/164, GSM8K 177/200 = the Q8_0 reference) and the fastest
  code config end to end (HE-10 block-16 decode 208.1 avg / 227.8 peak,
  e2e 156.2) because its acceptance with DFlash2 more than repays the
  slower plain forward
- the drafter step now runs our converter on the z-lab checkpoint: the
  pre-made incoai GGUF is the llama.cpp PR#27342 layout (arch "dflash",
  different key/tensor names) and this server does not load it; the old
  quickstart shipped a --draft flag that failed at startup
- server/README: the specla/ddtree-8 config is a validated alternative,
  not the out-of-the-box default (that is the plain chain at block 8)
@davide221
davide221 marked this pull request as ready for review August 24, 2026 16:53
@davide221
davide221 merged commit f686c44 into main Aug 24, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants