Skip to content

Run batch-1 Krea 2 forwards on the valid prompt rows (0.11.1) - #18

Merged
iamwavecut merged 1 commit into
mainfrom
perf/krea2-valid-text-rows
Oct 2, 2026
Merged

iamwavecut merged 1 commit into
mainfrom
perf/krea2-valid-text-rows

Conversation

@iamwavecut

Copy link
Copy Markdown
Owner

What changes

  • orbitquant.fused Krea 2: a batch-1 forward with a padded prompt drops the padded text rows once (embeddings and the matching RoPE positions) and runs without a mask, instead of gathering and scattering the valid rows in every block. The trimmed tensors of the last two prompts are reused, so the per-prompt text-fusion cache keeps hitting across steps (a CFG pipeline alternates two prompts). Batch > 1 and grad mode keep the per-block path.
  • Version 0.11.1.

The padded rows were never attended to and were dropped at the output, so the image rows attend to the same keys; the result differs from 0.11.0 only by kernel rounding (the masked and unmasked attention take different kernels).

Measurements

RTX 4060 Ti 16 GB, WaveCut/Krea-2-Turbo-OrbitQuant-W4A4 fused revision, 1248×832, 8 steps, the draw worker's zero-copy offload:

Runtime Denoising step Request (text encoder, steps, VAE)
OrbitQuant 0.10 Krea2FastRunner (previous revision) 1.037 s 11.35 s
0.11.0 pipeline call ~1.09 s 12.38 s
0.11.1 pipeline call 1.033 s 11.84 s

Per-step CUDA kernel time of the 0.10 runner and the 0.11.1 pipeline is identical (1.027 / 1.025 s per profiled step).

Checks

  • ruff check, CPU pytest: pass.
  • CUDA on an RTX 4060 Ti: tests/test_fused_family_kernels.py tests/test_fused_dit_kernels.py tests/test_fused_layout.py pass, including the new test_krea2_batch_one_forward_runs_on_the_valid_text_rows.

Krea2Pipeline pads every prompt to 512 text tokens and masks the padding. The fused blocks
skipped the padded rows with a gather and scatter in every block; a batch-1 forward now drops
them once, before the text fusion, and runs without a mask. The trimmed embeddings and
positions are reused for the same prompt, so the text-fusion cache stays warm across steps.
The padded rows got no attention weight and never reached the output, so the image rows see
the same keys.

RTX 4060 Ti, 1248x832, 8 steps, through the pipeline with zero-copy offload: a denoising step
takes 1.03 s, as with the OrbitQuant 0.10 runner, instead of about 1.09 s.
@iamwavecut
iamwavecut merged commit 7d9bfe8 into main Oct 2, 2026
3 checks passed
@iamwavecut
iamwavecut deleted the perf/krea2-valid-text-rows branch October 2, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant