Repository navigation
Fuse the transformer blocks of six diffusion families into stored checkpoints (0.11.0) - #16
Merged
Merged
Conversation
…ckpoints (0.11.0) orbitquant.fused groups the projections that read one input into one INT8-surrogate GEMM per block (Q|K|V, SwiGLU gate|up, both streams of a joint attention), runs norm and AdaLN modulation in the activation-quantization prologue and SwiGLU, sigmoid gates and gated residual updates in the GEMM epilogue, fuses Q/K normalization with RoPE and uses the INT8 Q.K^T / FP16 P.V attention kernel where a block's Q/K norm weights allow it. Families: Krea 2, FLUX.2, Ideogram 4, Qwen-Image 2.1, MiniMax-H3 and Boogu-Image. A fused checkpoint stores only the fused groups and records the layout in its quantization_config; from_pretrained and load_orbitquant_artifact rebuild the groups before the weights load. Groups may mix weight widths (low-bit boundary/interior protection) as row segments of one GEMM and may take 8-bit activations (per-token absmax INT8 of the rotated input). The activation quantizer covers RPBH blocks up to 16384.
…locks on valid rows Fused groups keep their weights as frozen parameters instead of buffers: diffusers' streamed group offloading moves only parameters back to the host, so block-level streaming kept every fused block on the GPU and ran out of memory under a VRAM cap. Krea 2 blocks run on the rows of valid tokens only (padded text rows are never attended to and are dropped at the output), find those rows once per forward, and the text fusion stack runs once per prompt instead of in every denoising step.
Diffusers casts every floating tensor of a checkpoint to the model dtype before the quantizer places it, which rounded the FP32 per-row scales of INT8 groups (and would round BF16 dense rows of an FP16 model). The quantizer now places fused group tensors from the stored values, so a fused checkpoint loads bit-identical to the model it was saved from.
…offloading of fused weights
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
New
orbitquant.fused: fused transformer blocks stored as checkpoints. Per block, the projections that read one input become one INT8-surrogate GEMM (attention Q|K|V, SwiGLU gate|up, both streams of a joint attention); RMSNorm/LayerNorm and AdaLN modulation (per tensor, or per row from a modulation table) run in the activation-quantization prologue; SwiGLU, sigmoid gates, GELU-tanh and gated residual updates (per column or per row from a table) run in the GEMM epilogue; Q/K normalization and RoPE (interleaved or rotate-half, optionally on a leading slice of the head) are one kernel; attention uses the INT8 Q·Kᵀ / FP16 P·V kernel where a block's Q/K RMSNorm weights stay within a spread limit and BF16 SDPA elsewhere.Families: Krea 2, FLUX.2 (double/single stream, reference-image KV cache), Ideogram 4, Qwen-Image 2.1 (block-causal prefill, KV-cached decode), MiniMax-H3 (per-row AdaLN tables, padding documents) and Boogu-Image (refiner, single/double stream).
fuse(model)converts a loaded per-projection model in place;save_pretrainedwrites a checkpoint holding only the fused groups and records the layout inquantization_config.fused_layout.from_pretrainedandload_orbitquant_artifactrebuild the empty groups before the weights load and finish the blocks after (prepare_skeleton/finalize).fuse_component_artifactconverts anorbitquant-v1component artifact.activation_bitsin the layout). Boogu-Image uses it for every group: its 3360 channels only allow a 32-wide rotation block, and 4-bit Q/K/V inputs produce visible texture artifacts.orbitquant.fuseddoes not import Triton; CPU-only code can read layouts.use_stream=True) moves only parameters back to the host, so buffers piled up on the GPU until it ran out of memory.fuse_component_artifactaccepts artifacts whose manifest lists files that a later repository edit removed; the fused manifest lists the files the fused copy ships.orbitquant.runtime.krea2from 0.10 is unchanged.Measurements
RTX 4060 Ti 16 GB (torch 2.10 + cu128, Triton 3.6, diffusers
80c7ed26unless noted), hot, per-projection W4A4 → fused W4A4 of the same published checkpoint:RTX 3090, MiniMax-H3 608×480×124 frames, 24 sigma points / 23 forwards, release runner: 6.06 → 4.0 s per forward resident; generation 175 → 116 s resident, and with block-level streamed group offloading (no
record_stream) 174 → 114 s at 5.1 GiB of GPU memory.Replays of real block inputs (fused vs per-projection, rel. L2): FLUX.2 0.002–0.027, Ideogram 4 0.001–0.017. A fused checkpoint loaded with
from_pretrainedrenders bit-identical images to the in-memory fused model (FLUX.2 10/10, Krea 2 6/6, Qwen-Image 2.1 6/6). Same-seed images move with the different rounding but keep their detail and legibility (paired contact sheets); reference edits stay close (FLUX.2: PSNR 32.7 dB, SSIM 0.985).Checks
uv run ruff check .,uv run pytest(CPU): pass.scripts/run_paper_methodology_checks.sh,scripts/run_hf_compat_checks.sh --mode current|release|dev: pass.tests/test_fused_family_kernels.py,tests/test_fused_dit_kernels.py,tests/test_krea2_runtime.py,tests/test_fused_layout.py: pass.