Skip to content

[https://nvbugs/6480621][test] Revert to 60-second KV transfer timeout for GB300 DeepSeek V4 Pro disaggregated perf-sanity - #17137

Merged
chienchunhung merged 1 commit into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6480621-timeout-60s-ci
Aug 28, 2026
Merged

[https://nvbugs/6480621][test] Revert to 60-second KV transfer timeout for GB300 DeepSeek V4 Pro disaggregated perf-sanity#17137
chienchunhung merged 1 commit into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6480621-timeout-60s-ci

Conversation

@chienchunhung

@chienchunhung chienchunhung commented Jul 31, 2026

Copy link
Copy Markdown
Collaborator

Summary

This PR updates one GB300 DeepSeek V4 Pro disaggregated perf-sanity target:

  • explicitly enables cache_transceiver_precheck;
  • sets GEN and CTX kv_transfer_timeout_ms from 600000 ms to 60000 ms;
  • keeps the exact gen-only test unwaived so the modified target is exercised.

PRs #17223 and #18185 have merged into main. The current PR head a4293f296272d202ffad8ebcbd4953066b377580 is rebased directly on main at 36138a2c8e0350d5f23bba41c1c12962a187009b. Merged #18185 supplies the DeepSeek V4 EPLB weight-loading OOM fix and removes this target's waiver.

Validation

Current-head targeted validation: PASSED

Command:

/bot run --disable-fail-fast --disable-reuse-test --only-multi-gpu-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

Exact test:

perf/test_perf_sanity.py::test_e2e[disagg_upload-gen_only-gb300_deepseek-v4-pro-fp4_8k1k_con180_ctx3_dep4_gen1_dep32_eplb384_mtp3_ccb-NIXL]

Evidence:

  • PR_Github #69570 and L0 #56888 ran against commit a4293f296 with test reuse disabled.
  • The exact test passed in 1089.791 seconds; its JUnit result has zero failures.
  • Slurm job 3325958 completed with exit code 0:0 across nvl72d217-T[01-11].
  • The disaggregated server and benchmark completed successfully using the Python transceiver, 60-second GEN/CTX KV-transfer timeouts, and the cache-transceiver precheck.
  • No host OOM, KV-transfer timeout, or data-corruption failure was observed.

This validates the current #17223 + #18185 + #17137 combination for the 11-node/44-GPU, 3-CTX, concurrency-180 CI proxy. It does not replace rerunning the original 8P1D, concurrency-1760 AgentPerf workload reported in NVBUG #6480621.

Related PRs

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63105 [ run ] triggered by Bot. Commit: 38ca970 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63106 [ run ] triggered by Bot. Commit: 38ca970 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63105 [ run ] completed with state ABORTED. Commit: 38ca970

Link to invocation

@chienchunhung chienchunhung changed the title [https://nvbugs/6480621][test] Validate 60-second KV transfer timeout [https://nvbugs/6480621][test] Revert to 60-second KV transfer timeout for GB300 DeepSeek V4 Pro disaggregated perf-sanity Jul 31, 2026
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63106 [ run ] completed with state SUCCESS. Commit: 38ca970
/LLM/main/L0_MergeRequest_PR pipeline #51197 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63144 [ run ] triggered by Bot. Commit: 38ca970 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63144 [ run ] completed with state FAILURE. Commit: 38ca970
/LLM/main/L0_MergeRequest_PR pipeline #51231 (Partly Tested) completed with status: 'ABORTED'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63150 [ run ] triggered by Bot. Commit: 38ca970 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63150 [ run ] completed with state FAILURE. Commit: 38ca970
/LLM/main/L0_MergeRequest_PR pipeline #51234 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63165 [ run ] triggered by Bot. Commit: 38ca970 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63165 [ run ] completed with state FAILURE. Commit: 38ca970
/LLM/main/L0_MergeRequest_PR pipeline #51249 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63555 [ run ] triggered by Bot. Commit: 3e6c120 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63555 [ run ] completed with state SUCCESS. Commit: 3e6c120
/LLM/main/L0_MergeRequest_PR pipeline #51521 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung
chienchunhung force-pushed the codex/nvbug-6480621-timeout-60s-ci branch from 3e6c120 to b7bd98a Compare August 3, 2026 21:40
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63564 [ run ] triggered by Bot. Commit: b7bd98a Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #63564 [ run ] completed with state SUCCESS. Commit: b7bd98a
/LLM/main/L0_MergeRequest_PR pipeline #51530 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@chienchunhung
chienchunhung marked this pull request as ready for review August 3, 2026 23:41
@chienchunhung
chienchunhung requested review from a team as code owners August 3, 2026 23:41
@chienchunhung
chienchunhung force-pushed the codex/nvbug-6480621-timeout-60s-ci branch from 6573ef3 to f66c303 Compare August 26, 2026 17:51
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69490 [ run ] triggered by Bot. Commit: f66c303 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69490 [ run ] completed with state SUCCESS. Commit: f66c303
/LLM/main/L0_MergeRequest_PR pipeline #56817 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung
chienchunhung force-pushed the codex/nvbug-6480621-timeout-60s-ci branch from f66c303 to 6b5d69c Compare August 26, 2026 21:38
@chienchunhung
chienchunhung requested a review from a team as a code owner August 26, 2026 21:38
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69539 [ run ] triggered by Bot. Commit: 6b5d69c Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69539 [ run ] completed with state FAILURE. Commit: 6b5d69c
/LLM/main/L0_MergeRequest_PR pipeline #56859 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@mikeiovine mikeiovine left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamp on behalf of runtime-devs, delegating review to @NVIDIA/trt-llm-torch-attention-devs

@chienchunhung
chienchunhung force-pushed the codex/nvbug-6480621-timeout-60s-ci branch from 6b5d69c to a4293f2 Compare August 27, 2026 00:18
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --disable-reuse-test --only-multi-gpu-test --stage-list "GB300-44_GPUs-11_Nodes-PyTorch-Disagg-PerfSanity-CTX3-NODE1-GPU4-GEN1-NODE8-GPU32-Post-Merge-2"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69570 [ run ] triggered by Bot. Commit: a4293f2 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69570 [ run ] completed with state SUCCESS. Commit: a4293f2
/LLM/main/L0_MergeRequest_PR pipeline #56888 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@chienchunhung
chienchunhung enabled auto-merge (squash) August 27, 2026 18:02
@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69783 [ run ] triggered by Bot. Commit: a4293f2 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69783 [ run ] completed with state SUCCESS. Commit: a4293f2
/LLM/main/L0_MergeRequest_PR pipeline #57078 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
@chienchunhung
chienchunhung force-pushed the codex/nvbug-6480621-timeout-60s-ci branch from a4293f2 to ea57698 Compare August 27, 2026 23:07
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69820 [ run ] triggered by Bot. Commit: ea57698 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #69820 [ run ] completed with state SUCCESS. Commit: ea57698
/LLM/main/L0_MergeRequest_PR pipeline #57112 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #70010 [ run ] triggered by Bot. Commit: ea57698 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #70010 [ run ] completed with state SUCCESS. Commit: ea57698
/LLM/main/L0_MergeRequest_PR pipeline #57289 completed with status: 'SUCCESS'

CI Report

Link to invocation

@chienchunhung
chienchunhung merged commit 3c4dc51 into NVIDIA:main Aug 28, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api-compatible Accepted LLM API contract change that is backwards-compatible ci: full pre-merge approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants