Skip to content

feat(training): add single-rank CUDA MPS mode - #2065

Open
TATP-233 wants to merge 1 commit into
develop/tensor-runtimefrom
feature/cuda-mps-single-rank
Open

TATP-233 wants to merge 1 commit into
develop/tensor-runtimefrom
feature/cuda-mps-single-rank

Conversation

@TATP-233

@TATP-233 TATP-233 commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Adds the explicit training.cuda_process_sharing: null | mps owner mode for the shared SAC/FastSAC/FlashSAC/WarpSAC off-policy path.
  • Adds a fail-closed UniLab training-runtime probe for Linux/NVIDIA CUDA, MJWarp, single-host/single-rank, same physical learner/collector GPU UUID, a live MPS control socket/FIFO, and a reachable control daemon. Explicit mps requests never fall back to independent CUDA contexts.
  • Adds an ADR and English/Chinese tensor-runtime and production documentation explaining daemon setup, evidence, scope, and limitations.

Linked Work

Validation

  • make test-all passed on the final local head before this PR was created or updated
  • Additional task-specific validation listed below

Commands actually run:

make check
make test
make test-all
uv run --no-sync pytest tests/training/test_cuda_process_sharing.py \
  tests/algos/test_offpolicy_double_buffer_runner.py \
  tests/utils/test_experiment_tracking.py -q
uv run --no-sync pytest tests/scripts/test_check_docs.py -q
cd docs/sphinx && \
  UNILAB_DOCS_SKIP_AUTODEC=1 uv run --no-project \
  --with-requirements requirements.txt sphinx-build -b html -n source build/html
CUDA_MPS_PIPE_DIRECTORY=/tmp/unilab-mps-quick/mps/pipe \
CUDA_MPS_LOG_DIRECTORY=/tmp/unilab-mps-quick/mps/log \
CUDA_VISIBLE_DEVICES=0 uv run --no-sync python - <<'PY'
from unilab.training.cuda_process_sharing import probe_cuda_process_sharing
print(probe_cuda_process_sharing(
    "mps", "cuda:0", "cuda:0", backend="mjwarp"
).manifest())
PY

Benchmark matrix on the final head: two tasks, default versus explicit MPS, two repeats, 300 iterations each; metrics averaged over the final 100 iterations:

Case Default steps/s MPS steps/s Throughput
FlashSAC / G1 Motion Tracking / MJWarp 102,379.93 120,742.50 +17.94%
SAC / G1 Walk Flat / MJWarp 78,738.39 108,954.83 +38.38%

All eight completed runs reached status=completed and 300/300 iterations. Every run had normal shutdown, no cleanup errors, replay published/release sequences equal, final replay occupancy zero, and zero dropped batches/early returns. MPS runs recorded validated manifest evidence with matching learner/collector UUIDs and the same control pipe/server PID.

Benchmark commands:

export CUDA_VISIBLE_DEVICES=0
uv run train --algo flashsac --task g1_motion_tracking --sim mjwarp \
  training.no_play=true training.export_onnx=false \
  training.nan_guard.enabled=false algo.max_iterations=300 \
  algo.save_interval=10000 training.log_dir=/absolute/path/run

export CUDA_MPS_PIPE_DIRECTORY=/absolute/path/mps/pipe
export CUDA_MPS_LOG_DIRECTORY=/absolute/path/mps/log
uv run train --algo flashsac --task g1_motion_tracking --sim mjwarp \
  training.cuda_process_sharing=mps \
  training.no_play=true training.export_onnx=false \
  training.nan_guard.enabled=false algo.max_iterations=300 \
  algo.save_interval=10000 training.log_dir=/absolute/path/run

uv run --no-sync python scripts/benchmark/rl/extract_offpolicy_metrics.py \
  /absolute/path/run --last 100 --json

Remote CI route:

  • Base is not main; remote CI is not scheduled. Local make test-all is the test gate. The later integration PR to main will run remote CI.

Impact

  • Backend impact: MJWarp only for mps; default behavior is unchanged for all backends
  • Platform impact: Linux/NVIDIA CUDA only for mps; default behavior is unchanged
  • Training effect expected: no training-semantic change; MPS changes GPU process execution sharing only. env_steps_per_sync is unchanged.

Artifacts

  • W&B: none
  • benchmark result: local /home/user/ws/unilabsim2/UniLab/benchmarks/cuda-mps-20261004/ (not committed; README.md and results.json record commands, host, metrics, and manifest evidence)
  • video / screenshot: none
  • ONNX / checkpoint: none

Checklist

  • Added or updated tests where needed
  • Updated docs if behavior or workflow changed
  • Linked the driving issue
  • Noted any follow-up work explicitly

Follow-up:

  • Multi-GPU DP and cross-rank daemon topology remain excluded and tracked by CUDA MPS under multi-GPU DP: evidence gaps and design constraints #2063.
  • Other CUDA backends and CUDA_MPS_ACTIVE_THREAD_PERCENTAGE remain evidence-gated follow-ups.
  • The local RTX 4090 benchmark passes the Phase 3 thresholds but is not the discussion's RTX 6000D reference-host evidence; a maintainer may still request that reference-host matrix.

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner October 4, 2026 17:25
@TATP-233

TATP-233 commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator Author

Tested PR #2065 on the requested server ssh 592, repository ~/unilabsim/UniLab.

Environment

  • Host: lyg2245
  • UniLab: 5cc7d997bb5c8670a73b4ba2d09e75a6f22e0132 (PR head)
  • Exact workspace siblings from tensor_runtime_workspace.json:
    • UniSim 5e835923e892e9f79213eac64565beb90b7bb074
    • unilab-rl aa2f8602aef05b32cf28d4e13716a37849b4f785
    • mjbatch-uni b3a3b480ff360cb8144e2557338b5040b5917900
  • GPU: one RTX 5090, UUID GPU-f4082127-7b08-5e35-212e-5bb52b872ac6
  • Driver: 580.95.05
  • Torch: 2.14.0+cu130, CUDA available, one visible device
  • Topology: single host, single rank, CUDA_VISIBLE_DEVICES=0
  • A user-owned isolated MPS daemon was started under benchmarks/cuda-mps-592-20261005/mps, used only for these tests, and stopped afterward.

The server had unrelated dirty UniLab work. It was preserved in stash pre-pr-2065-cuda-mps-existing-work (patch SHA-256 46a8307fcd99b26f58b6ecb57d0144f23f2d65d4d4852b90ee275c8eb73e3e8a) and was not included in this test.

Focused tests

An initial run against the server's older unpinned unilab_rl sibling (72a50b3) exposed 12 unrelated tensor-runtime-bound test failures. After checking out the exact workspace sibling pins, all focused tests passed:

75 passed in 9.51s

Command:

UNILAB_LOCAL_UNISIM="$HOME/unilabsim/unisim" CUDA_VISIBLE_DEVICES=0 \
uv run --no-sync pytest \
  tests/training/test_cuda_process_sharing.py \
  tests/algos/test_offpolicy_double_buffer_runner.py \
  tests/utils/test_experiment_tracking.py -q

Real MPS probe:

{
  "configured": "mps",
  "effective": "mps",
  "validated": true,
  "learner_device": "cuda:0",
  "collector_device": "cuda:0",
  "learner_gpu_uuid": "F40821277B085E35212E5BB52B872AC6",
  "collector_gpu_uuid": "F40821277B085E35212E5BB52B872AC6",
  "server_pid": 566566
}

Phase 3 benchmark

Two repeats per arm, 300 iterations each, metrics averaged over final 100 iterations.

FlashSAC / G1 Motion Tracking / MJWarp

Mode steps/s iteration ms learner train ms learner wait ms
Default 73,121.25 28.023 10.675 11.400
CUDA MPS 82,642.51 24.891 9.063 9.782
Delta +13.02% -11.18% -15.10% -14.20%

Threshold: ≥ +10%; passed.

SAC / G1 Walk Flat / MJWarp

Mode steps/s iteration ms learner train ms learner wait ms
Default 64,902.60 31.575 11.297 17.320
CUDA MPS 85,767.33 23.887 13.557 8.705
Delta +32.15% -24.35% +20.00% -49.74%

Threshold: ≥ +20%; passed.

Integrity evidence

All eight runs passed:

  • status=completed
  • 300/300 iterations
  • metric/runtime manifest schema 1
  • normal_completion
  • cleanup errors []
  • replay published sequence == release sequence
  • final replay occupancy 0
  • dropped batches 0
  • early returns 0
  • MPS manifest configured/effective mps, validated true
  • learner and collector GPU UUID identical
  • MPS run_config.json snapshots equal "mps"; baseline snapshots equal null

No residual compute processes remained after the MPS daemon was stopped.

Artifacts

Server-local, not committed:

~/unilabsim/UniLab/benchmarks/cuda-mps-592-20261005/

Contains results.json, run logs/timings, TensorBoard events, run_config.json, and run_summary.json for all eight runs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant