[None][test] Unwaive 17 recovered perf-sanity test_e2e cases - #17967
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 12 included reviews per hour; 10 remain after this review. WalkthroughThe change adds one LagunaXS NVFP4 accuracy waiver and removes obsolete DGX_B200 and performance waivers from the integration-test waiver list. ChangesIntegration waivers
Estimated code review effort: 1 (Trivial) | ~2 minutes Merge Risk: ⚪ Minimal · up to This change only re-enables 17 previously waived performance tests after reported successful runs on the matching hardware; no actionable merge-blocking risk remains beyond normal checks and review. Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
brnguyen2
left a comment
There was a problem hiding this comment.
Suggestion on how to land this, plus two evidence asks.
Land these as post-merge first, then promote. Removing the waive puts each case straight back into its test-db stage, and for the ones whose entry is stage: pre_merge a single re-failure blocks everyone's PRs, not just yours. Since the recovery evidence here is one green run per case (see below), the low-risk path is: unwaive and flip the entry to post-merge in the same PR, let them soak on main for a week or two of post-merge runs, then a follow-up PR promotes the ones with a clean streak back to pre-merge. You keep the coverage you're restoring, and an intermittent case costs a post-merge alert instead of an L0 outage. Concretely, gemma4_26b_a4b_nvfp4_tp1_1k1k and llama8b_spec_bs1_128_128 are stage: pre_merge in l0_b200_perf_sanity.yml — those are the two worth moving; the rest are already post-merge, so for them "unwaive as-is" is fine risk-wise and only points 2/3 apply.
-
Evidence that these are actually fixed, not intermittently passing. The PR rests on one re-run of each case, with a functional pass bar (Total == Successful, Failed == 0, non-null throughput, no crash/OOM/hang). For perf-sanity and disagg cases, one green run doesn't distinguish "fixed" from "passed this time". Please post, per unwaived ID: the pipeline/job URLs, how many runs each got, and on which platform. A case with a single pass is exactly the one that should land post-merge rather than pre-merge.
-
Tie each unwaive to a root cause. None of the 8 cited bugs (6517846, 6572843, 6566777, 6374910, 6571410, 6571408, 6432948, 6153575) appears to have a fix identified. A bug that's still open with no fixing commit and a case that passed once is intermittency, not recovery. For each bug, please name what fixed it (commit/MR, or an infra/config change such as a driver/container/baseline update) and then move it to V2C so it isn't left open with nothing tracking it. Where you can't name a cause, say so explicitly — those are the strongest candidates for post-merge-only. 6153575 in particular is 105 days open and titled a perf regression in
super_ad_blackwell; a regression doesn't recover on its own, so something specific must have changed. -
Recovery criterion doesn't cover the pre-merge gate.
perf_regression_utils.process_and_upload_test_resultssetsfail_on_regression = not is_post_merge, so a pre-merge case fails on an OpenSearch baseline regression even when the run completes with full request accounting. Your stated bar (Total == Successful == num_prompts, Failed == 0, non-null throughput) doesn't exercise that comparison at all, which is another reason to land the two pre-merge cases as post-merge: post-merge runs upload a baseline without gating on it, so you get real signal before they can block anyone. If you'd rather keep them pre-merge, please/bot runso those two actually hit the regression gate first — 6571410/6571408 were both filed as pre-merge failures.
Happy to re-review once the staging decision and the run evidence are in.
|
/bot run --disable-fail-fast --stage-list "DGX_B200-PyTorch-PerfSanity-1,DGX_B200-PyTorch-1,DGX_B200-PyTorch-2,DGX_B200-PyTorch-3,DGX_B200-PyTorch-4,DGX_B200-PyTorch-5,DGX_B200-PyTorch-6,DGX_B200-PyTorch-7,DGX_B200-PyTorch-8,DGX_B200-PyTorch-9,DGX_B200-AutoDeploy-Post-Merge-1,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge*,GB200-4_GPUs-PyTorch-PerfSanity-Post-Merge*,GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge*,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge*,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE8-GPU32-Post-Merge*,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN4-NODE2-GPU8-Post-Merge*" |
8edb7f5 to
110fc67
Compare
|
/bot run --disable-fail-fast --stage-list "DGX_B200-PyTorch-PerfSanity-1,DGX_B200-PyTorch-1,DGX_B200-PyTorch-2,DGX_B200-PyTorch-3,DGX_B200-PyTorch-4,DGX_B200-PyTorch-5,DGX_B200-PyTorch-6,DGX_B200-PyTorch-7,DGX_B200-PyTorch-8,DGX_B200-PyTorch-9,DGX_B200-AutoDeploy-Post-Merge-1,DGX_B200-8_GPUs-PyTorch-PerfSanity-Post-Merge*,GB200-4_GPUs-PyTorch-PerfSanity-Post-Merge*,GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge*,GB200-8_GPUs-2_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU1-GEN1-NODE1-GPU4-Post-Merge*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-FUNCTIONAL-ONLY-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8*,GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE2-GPU8-Post-Merge*,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU2-GEN1-NODE8-GPU32-Post-Merge*,GB300-36_GPUs-9_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN4-NODE2-GPU8-Post-Merge*" |
|
PR_Github #67756 [ run ] triggered by Bot. Commit: |
|
PR_Github #67756 [ run ] completed with state
|
110fc67 to
0c96b06
Compare
|
/bot skip --comment "Unwaive a perf test, No need to run the whole CI pipeline" |
|
PR_Github #68316 [ skip ] triggered by Bot. Commit: |
Re-ran every currently-waived perf/test_perf_sanity.py::test_e2e case on GPUs and checked functional recovery (runs to completion with full request accounting: Total == Successful == num_prompts, 0 failed): GB300 on GB300 hardware, GB200 on GB200 hardware, B200 on DGX B200. Verified at main commit 3253b64. 17 of the 27 waived cases now pass -- their original failures were resolved elsewhere. Rebased onto main d0e8baa: 4 of those 17 waive lines had already been removed upstream, so this commit deletes the remaining 13; the merged result unwaives all 17. The 10 still-failing cases (DeepSeek-V4-Pro-fp4 host-RAM OOM on the multi-ctx-server topologies, glm5 tep8 max_num_tokens config, and the dsr1_fp4_v2 KV-cache-v2 scheduler deadlock) remain waived. Unwaived cases (bracket id -> parent NVBug): - aggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL (nvbugs/6581075) - aggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL (nvbugs/6517846) - aggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXL (nvbugs/6517846) - disagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL (nvbugs/6572843) - disagg_upload-e2e-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXL (nvbugs/6601537) - disagg_upload-e2e-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL (nvbugs/6566777) - disagg_upload-gen_only-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL (nvbugs/6581075) - disagg_upload-gen_only-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXL (nvbugs/6601537) - disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL (nvbugs/6581075) - disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL (nvbugs/6581075) - disagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXL (nvbugs/6581075) - aggr_upload-dynamo_gpt_oss_120b_fp4_blackwell-gpt_oss_fp4_tep4_adp_cutlass_8k1k (nvbugs/6374910) - disagg_upload-gen_only-gb200_gpt-oss-120b-fp4_8k1k_con1024_ctx1_tp1_gen1_tp4_eplb0_mtp0_ccb-NIXL (nvbugs/6581075) - aggr_upload-gemma4_26b_a4b_nvfp4_blackwell-gemma4_26b_a4b_nvfp4_tp1_1k1k (nvbugs/6571410) - aggr_upload-host_perf_llama8b_spec_decode-llama8b_spec_bs1_128_128 (nvbugs/6571408) - aggr_upload-deepseek_r1_fp8_blackwell-r1_fp8_tp8_mtp3_8k1k (nvbugs/6432948) - aggr_upload-super_ad_blackwell-super_ad_ws1_1k1k (nvbugs/6153575) Signed-off-by: chenfeiz0326 <203214996+chenfeiz0326@users.noreply.github.com>
0c96b06 to
e45d9d3
Compare
|
/bot skip --comment "Unwaive a perf test, No need to run the whole CI pipeline" |
|
PR_Github #68316 [ skip ] completed with state |
|
PR_Github #68320 [ skip ] triggered by Bot. Commit: |
|
PR_Github #68320 [ skip ] completed with state |
Summary
Re-ran every currently-waived
perf/test_perf_sanity.py::test_e2e[...]case on real GPUs to check whether the originally-waived failures have
recovered. A case counts as recovered only if it runs to completion with
full request accounting (
Total == Successful == num_prompts,Failed == 0,non-null throughput) and no crash/OOM/hang/timeout.
17 of 27 waived cases now pass, so this PR removes only those 17 waive
lines from
tests/integration/test_lists/waives.txt. The other 10 stillfail and stay waived.
maincommit3253b64043b9d8ceb0a3c802c5e76eb7c51af58d(harness read from the commit under test).
waives.txt, 17 deletions, 0 insertions).Unwaived cases (bracket id → parent NVBug)
GB300 (11)
aggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXLaggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXLaggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXLdisagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXLdisagg_upload-e2e-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXLdisagg_upload-e2e-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXLdisagg_upload-gen_only-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXLdisagg_upload-gen_only-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXLdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXLdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXLdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXLGB200 (2)
aggr_upload-dynamo_gpt_oss_120b_fp4_blackwell-gpt_oss_fp4_tep4_adp_cutlass_8k1kdisagg_upload-gen_only-gb200_gpt-oss-120b-fp4_8k1k_con1024_ctx1_tp1_gen1_tp4_eplb0_mtp0_ccb-NIXLB200 / DGX B200 (4)
full:DGX_B200/...aggr_upload-gemma4_26b_a4b_nvfp4_blackwell-gemma4_26b_a4b_nvfp4_tp1_1k1kfull:DGX_B200/...aggr_upload-host_perf_llama8b_spec_decode-llama8b_spec_bs1_128_128aggr_upload-deepseek_r1_fp8_blackwell-r1_fp8_tp8_mtp3_8k1kaggr_upload-super_ad_blackwell-super_ad_ws1_1k1kVerification logs
Recovery of each unwaived case was confirmed from its run log below. All logs are now stored on the internal repair-bot host under
/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/<GPU>/. GB300/GB200 rows carry thetrtllm-benchmarkaccounting log (Total == Successful, 0 failed); B200 rows carry the full CI-pytestslurm-<id>.outstream.GB300 — aws_cmh (11)
aggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c03-glm5-con1024-ctxonly.logaggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c04-glm5-con1-ctxonly.logaggr_upload-ctx_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c05-glm5-con512-ctxonly.logdisagg_upload-e2e-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c06-r1-128k8k-con256-e2e.logdisagg_upload-e2e-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c10-v4pro-con8-e2e.logdisagg_upload-e2e-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c11-glm5-con1024-e2e.logdisagg_upload-gen_only-gb300_deepseek-r1-fp4_128k8k_con256_ctx1_pp4_gen1_dep8_eplb0_mtp1_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c12-r1-128k8k-con256-genonly.logdisagg_upload-gen_only-gb300_deepseek-v4-pro-fp4_8k1k_con8_ctx1_dep4_gen4_tep8_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c16-v4pro-con8-genonly.logdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1024_ctx1_dep2_gen1_dep8_eplb256_mtp1_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c17-glm5-con1024-genonly.logdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con1_ctx1_dep2_gen1_tep8_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c18-glm5-con1-genonly.logdisagg_upload-gen_only-gb300_glm-5-fp4_8k1k_con512_ctx1_dep2_gen1_dep32_eplb0_mtp3_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB300/c19-glm5-con512-genonly.logGB200 — lyris (2)
aggr_upload-dynamo_gpt_oss_120b_fp4_blackwell-gpt_oss_fp4_tep4_adp_cutlass_8k1k/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB200/gb200-gptoss-aggr-8k1k-2734693.logdisagg_upload-gen_only-gb200_gpt-oss-120b-fp4_8k1k_con1024_ctx1_tp1_gen1_tp4_eplb0_mtp0_ccb-NIXL/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/GB200/gb200-gptoss-disagg-genonly-con1024-2735090.logB200 / DGX B200 — computelab (4)
aggr_upload-gemma4_26b_a4b_nvfp4_blackwell-gemma4_26b_a4b_nvfp4_tp1_1k1k/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/B200/b200-gemma-1k1k-3749341.logaggr_upload-host_perf_llama8b_spec_decode-llama8b_spec_bs1_128_128/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/B200/b200-llama8b-spec-3747851.logaggr_upload-deepseek_r1_fp8_blackwell-r1_fp8_tp8_mtp3_8k1k/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/B200/b200-dsr1-fp8-8k1k-3750394.logaggr_upload-super_ad_blackwell-super_ad_ws1_1k1k/home/scratch.chenfeiz_gpu/repo/20260818-waive-triage/logs/B200/b200-superad-1k1k-3749445.logStill waived (10, unchanged)
DeepSeek-V4-Pro-fp4 host-RAM OOM during weight load on the multi-ctx-server
topologies (con180/con666/con4301 → ctx3/ctx6/ctx12), the glm5
tep8_mtp3_8k1kmax_num_tokensconfig mismatch, and thedeepseek_r1_fp4_v2 dep4_mtp1_1k8kKV-cache-v2 scheduler deadlock. Fixes for these are tracked separately.
Test Coverage
Re-enables 17 perf-sanity
test_e2ecases in CI that were previously skipped.Dev Engineer Review
tests/integration/test_lists/waives.txtremoves 17 recoveredperf/test_perf_sanity.py::test_e2ewaivers.Verdict: sufficient.
QA Engineer Review
test-db/orqa/files were modified.tests/integration/test_lists/waives.txt.perf/test_perf_sanity.py::test_e2eon GB300, GB200, and DGX B200 hardware.Verdict: needs follow-up.