Repository navigation
Replies: 1 comment
|
The scope of this discussion has been narrowed to single host, single rank to match its evidence. The multi-GPU DP analysis — NCCL interaction, cross-rank fail-closed semantics, per-host daemon vs per-rank validation, memory-guard behavior under MPS, and a proposed DP acceptance gate — now lives in #2063. The design constraints here (per-rank config semantics, rank-local probe, per-rank manifest evidence) were kept explicitly so DP support remains an additive extension. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Productize CUDA MPS for rank-local MJWarp off-policy training
Summary
CUDA MPS should become an explicit supported runtime mode for MJWarp off-policy training rather than remaining an undocumented host-level setting.
UniLab already requires every rank to keep its learner and collector on the same rank-local GPU. Each side is still a separate process with its own CUDA context, so that required topology is arbitrated by coarse driver time-slicing. MPS lets the rank-local contexts share the GPU and overlap their kernels; it does not change placement or training semantics.
This discussion proposes the product direction, acceptance criteria, execution path, and open decisions. No code is proposed yet.
Scope of this discussion: single host, single rank. All evidence below is one GPU, one learner process, one collector process. Multi-GPU DP is out of scope here and is tracked in #2063; the design constraints below are chosen so that DP support remains an additive extension rather than a redesign.
Evidence
All measurements below were taken on the same host, same code head, one physical GPU,
training.no_play=true,training.export_onnx=false, with only the MPS environment changed. Values are the mean of the final 100 iterations. FlashSAC Motion Tracking used 150 iterations; FastSAC G1 Walk Flat used 200.Host and topology
cuda:0, collector backendcuda:0cuda:0UniLab deliberately requires rank-local learner/collector placement. The proposal builds on that contract rather than replacing it.
FlashSAC + G1 Motion Tracking + MJWarp
Relative improvement:
FastSAC / SAC + G1 Walk Flat + MJWarp
Relative improvement:
Why this works
The collector is not simply becoming faster. In FastSAC G1 Walk Flat with
env_steps_per_sync=4, for example:Under MPS, the collector step is slightly slower because learner kernels now genuinely compete with it on the GPU. The large win is end-to-end because learner work overlaps collector work instead of waiting through multi-context time-slicing.
This is the key distinction from the rejected CPU-priority approach:
torch.cuda.Stream(priority=...)The OS-priority experiment was closed without merge because it did not match the bottleneck: the measured gain on motion tracking was noise-level, G1 Walk Flat regressed by about 3%, and CPU was not saturated.
Scope boundary
Two runtime facts are already fixed by design and are not under discussion:
env_steps_per_syncis a training-semantics setting: it changes the number of collector ticks consumed per learner synchronization and therefore the algorithm’s update boundary.A third boundary is set by the evidence: everything measured and proposed here is single host, single rank (world_size = 1). DP behavior is intentionally unspecified in this discussion; see #2063.
MPS productization concerns only GPU execution sharing for the required rank-local topology. It does not introduce cross-GPU learner/collector placement and does not alter
env_steps_per_sync. Until #2063 produces DP evidence, the supported topology for this mode is one rank on one host.Non-goals
env_steps_per_sync; it changes the learner update boundary and is already excluded by the scope boundary above.Product direction
Principle
UniLab should expose an explicit runtime execution mode, not an implicit environment side channel.
Users should not need to know that MJWarp off-policy training is multi-process, that learner and collector have separate CUDA contexts, or that this topology can be improved by MPS. The owner configuration should describe the desired execution mode; UniLab should validate the host state and record the effective result in run evidence.
Proposed owner configuration
A single selection at the training owner level:
Semantics:
null: current behavior; the default remains unchanged.mps: request shared GPU execution through CUDA MPS for the rank-local training process tree.This is intentionally not a standalone backend switch and does not replace
training.sim_backend. It is an execution-sharing mode for the configured runtime.Requested behavior
When
cuda_process_sharing: mpsis selected:CUDA_MPS_PIPE_DIRECTORYpoints to a live control pipe;cuda:0(CUDA MPS under multi-GPU DP: evidence gaps and design constraints #2063).runtime_manifest.run_config.json.Runtime evidence
At minimum, the manifest should record:
{ "cuda_process_sharing": { "configured": "mps", "effective": "mps", "server_pid": 2768293, "control_pipe": "/tmp/mps-test/control", "learner_device": "cuda:0", "collector_device": "cuda:0", "validated": true } }If validation fails, there is no run; the failure should identify the first unmet prerequisite.
The evidence object is deliberately per-rank: a future DP extension records one such object per rank plus host-level server evidence, without changing this schema (#2063).
Documentation direction
Document the mode in the tensor-runtime production guide:
null;env_steps_per_syncremains a separate training-semantics owner setting and is not modified bycuda_process_sharing.Execution path
The work is split into reviewable phases. Each phase has concrete exit criteria; the next phase starts only after those criteria are met.
Phase 0 - Decision and scope record
Outcome: an issue or ADR records that
cuda_process_sharingis a supported runtime mode.The scope statement records UniLab's existing rank-local rule: each rank keeps its learner and collector on the same physical GPU.
Scope decision:
Suggested initial scope:
env_steps_per_syncremains a training-semantics owner decision and is not modified.Exit criteria:
Phase 1 - Probe and diagnostics library
Implement a small owner module, likely under
unilab/training/or an equivalent owner package, exposing:Responsibilities:
Tests should use fake filesystem/process results and cover:
mpswith no silent fallback.Exit criteria:
Phase 2 - Owner config and builder wiring
Add the config field to the relevant off-policy owner YAML and resolve it in the shared off-policy assembly path before constructing the learner or collector.
The probe must run before:
Record the evidence into the runtime manifest supplied by the runner.
Tests should cover:
null;mpsis resolved and passed to the builder path;Exit criteria:
mpsrequest cannot reach env creation.Phase 3 - Short-run benchmark gate
Before merging Phase 2, rerun the short benchmark matrix on one reference host.
Required cases:
For each case:
Suggested shape:
Acceptance thresholds:
Exit criteria:
Phase 4 - Production documentation and runbook
Document:
Provide both English and Chinese pages or sections, following the repository's docs policy.
Exit criteria:
Phase 5 - Optional resource partitioning investigation
Only after explicit MPS mode is stable, investigate
CUDA_MPS_ACTIVE_THREAD_PERCENTAGE.This would allow SM resource partitioning between clients. It should be treated as a separate contract because it changes the performance tradeoff:
Note that this percentage is a property of the MPS server — global across all clients and GPUs it serves. Per-rank partitioning under DP would require multiple daemons with distinct pipe directories; that tension is discussed in #2063.
Do not include it in the initial PR.
Phase 6 - Optional broader backend validation
After MJWarp is supported, evaluate whether other CUDA tensor backends benefit:
This should be evidence-driven. Do not expand the support matrix without benchmark data.
Explicitly rejected alternative: CPU priority configuration
We should not productize
learner_scheduling,buffer_scheduling, andcollector_schedulingbased on OS nice or affinity.Reasons:
The correct scheduling lever for this topology is GPU context sharing, not OS process priority.
Open questions
Daemon ownership.
Should UniLab only validate an existing MPS daemon, or should optional tooling start one under an explicitly user-owned pipe/log directory?
Default value.
Should
nullmean "do not use MPS", or should there be a third value such asautothat uses MPS when available?Recommendation: ship
null | mpsfirst. Addautoonly after a reliable host policy is defined.DP behavior.
DP is out of scope for this discussion. The only DP-related constraint accepted here is that the per-rank config semantics, rank-local probe, and per-rank evidence shape must not preclude the DP extension. The DP-specific problems — NCCL interaction, cross-rank fail-closed semantics, daemon layering, memory-budget guards — are tracked in CUDA MPS under multi-GPU DP: evidence gaps and design constraints #2063.
Container policy.
How should the mode interact with Docker, Slurm, Kubernetes, or restricted runners where MPS control is unavailable?
Recommendation: fail closed with a clear diagnostic when explicitly requested.
Metric label.
Should throughput evidence record whether MPS was active in a canonical scalar or only in the runtime manifest?
Recommendation: runtime manifest only initially.
Scope of algorithms.
Should this be shared off-policy runtime behavior or only MJWarp owners?
Recommendation: shared off-policy assembly, but backend validation restricts effective support to MJWarp initially.
CI.
Can CI validate anything beyond the fail-closed path without an MPS-enabled GPU runner?
Recommendation: CI covers validation logic with fakes; the performance gate remains a recorded manual/reference-host benchmark.
Suggested acceptance definition
The feature is complete when:
runtime_manifest.Reproduction sketch
Start an MPS control daemon in an explicitly owned directory:
Run the short benchmark:
Extract metrics:
Stop the daemon after testing:
All reactions