Draft-target model speculative decoding verification engine with rejection sampling, acceptance rate telemetry, and dynamic speedup ratio estimation.
-
Updated
Sep 28, 2026 - Python
Draft-target model speculative decoding verification engine with rejection sampling, acceptance rate telemetry, and dynamic speedup ratio estimation.
Draft-target model speculative decoding verification engine with rejection sampling, acceptance rate telemetry, and dynamic speedup ratio estimation.
On-policy distillation for speculative decoding: train the draft model on the target's own rollouts, no offline feature cache. A fork of deepseek-ai/DeepSpec with DeepSeek-V4-Flash (DSpark) support.
Speculative decoding runtime with rejection sampling, adaptive gamma strategy, and provable correctness guarantees. Achieves 1.41x speedup on CPU with Qwen2-0.5B/1.5B pair. Draft model generates candidates, target model verifies in single forward pass. 31/31 tests passing.
Analytical benchmark for speculative decoding at batch sizes 1-64. Finds the gamma crossover where speculative decoding becomes slower than greedy, and shows that batch size does not degrade performance when continuous refill is active.
CLI for building and testing DFlash-style speculative decoding draft models.
Tree-based speculative decoding benchmarked against linear under equal verification budget — branch factor, depth, and draft quality sweep with fair node-level comparison.
Simulates speculative decoding to find the optimal speculation length K across 576 configurations (3 draft models x 8 K values x 6 acceptance rates x 4 cost ratios). Key findings: 6.06x max speedup, breakeven at cost_ratio=0.25, optimal K grows from 1-3 at 50% acceptance to 7-15 at 95% acceptance.
Speculative-draft training pipeline (DFLASH/DSPARK/DFLASH2), orchestrating the vllm-project/speculators engine end to end, with preset configs for Qwen-family models
Calibrated simulation benchmark for real-time LLM request routing, comparing complexity signals, output-length awareness, cost savings, and quality-risk trade-offs.
Measures real speculative decoding speedup using the official HuggingFace assistant_model API across 4 model pairs and output lengths up to 512 tokens. Best result: distilgpt2->gpt2-medium achieves 1.747x speedup at 512 output tokens. Validates that cost_ratio and output_length are the key parameters.
MTP drafter for lossless speculative decoding of Qwen3.8-27B-AEON on MLX. Weights on Hugging Face.
multiple tokens, and a verifier filters them using the main model’s confidence. Focuses on speed–accuracy tradeoffs, visualization, and modular design for easy benchmarking and research.
Measure speculative-decoding acceptance rate and expected speedup for a target/draft model pair
To associate your repository with the draft-model topic, visit your repo's landing page and select "manage topics."