Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
---
orphan: true
---

# ADR-0013 CUDA MPS Single-Rank Execution Sharing

语言: 简体中文

- Status: Proposed
- Date: 2026-10-04
- Owners: UniLab training runtime maintainers
- Supersedes: None
- Superseded by: None

## Context

off-policy 训练已经要求每个 rank 的 learner 与 collector 使用同一张 rank-local GPU,
但二者仍是独立 CUDA context,由驱动按多 context 粗粒度分时仲裁。Discussion #1800
的参考主机证据显示,显式启用 NVIDIA CUDA MPS 可以让这两个 context 共享 GPU 并重叠
kernel,从而改善端到端吞吐;其证据范围是单主机、单 rank、MJWarp。

现有配置没有表达该 host execution mode,用户只能在部署层隐式设置环境变量。这既不能
fail closed,也没有可审计的 runtime evidence。

## Decision

- 新增共享 off-policy owner 设置
`training.cuda_process_sharing: null | mps`,默认 `null` 保持现状。
- 初始有效范围限定为 Linux、NVIDIA CUDA、`training.sim_backend=mjwarp`、单主机、
`world_size=1` 的 SAC/FastSAC/FlashSAC/WarpSAC 共享路径。
- `mps` 是显式请求:所有前置条件在 env probe、learner construction 与 collector
spawn 之前验证,失败时给出第一个未满足条件和启动既有 daemon 的命令,不 fallback。
- rank-local 物理一致性按 GPU UUID 判断,而不是 CUDA ordinal。
- UniLab 只验证既有 MPS control daemon 与 control socket/FIFO,不 start/stop/repair
daemon,不引入 `auto`、SM percentage 或 OS nice/affinity 配置。
- `run_config.json` 记录配置值;runtime manifest v1 增加 producer 诊断字段
`cuda_process_sharing`,记录 configured/effective、devices、UUID、control pipe、
server PID 与 validated 状态。该字段先保持 v1 诊断扩展,不提升为稳定公共字段。
- 多 GPU DP 继续由 #2063 单独决策;本 ADR 的 per-rank 参数与 evidence shape 保留
加法扩展空间。

## Stable Contracts

- Owner 配置:三个 off-policy 主配置中的 `training.cuda_process_sharing`。
- Fail-closed probe:`src/unilab/training/cuda_process_sharing.py`。
- Builder wiring:`src/unilab/scripts/train_offpolicy.py` 在构造 env/runner 前 probe,
并把 evidence 合并进 runner 的 runtime manifest。
- 测试与文档:probe fake、owner default、builder fail-closed、manifest evidence、
英文/中文生产指南。

## Alternatives Considered

- 继续要求用户只在 shell 设置 MPS 环境变量:无法验证拓扑,也没有结构化证据。
- 由 UniLab 启动/停止 daemon:把共享 host 服务生命周期混入训练进程,且无法安全服务
并发/多用户运行。
- CPU priority/affinity:Discussion #1800 的测量显示其不解决跨进程 CUDA context
arbitration,且引入未证实的公共配置复杂度。

## Consequences

- 显式 `mps` 请求失败时不会有默认多 context fallback;这避免静默性能和拓扑漂移。
- MPS daemon、pipe/log 目录与多用户隔离仍是部署者责任。
- 非 MJWarp backend 和 DP 请求会失败,不形成支持声明。
- `env_steps_per_sync` 仍是训练语义 owner 设置,不由 execution-sharing mode 修改。

## Evidence In Repo

- `src/unilab/training/cuda_process_sharing.py`
- `src/unilab/scripts/train_offpolicy.py`
- `src/unilab/conf/sac/config.yaml`
- `src/unilab/conf/flashsac/config.yaml`
- `src/unilab/conf/warpsac/config.yaml`
- `tests/training/test_cuda_process_sharing.py`
- `tests/algos/test_offpolicy_double_buffer_runner.py`
- `docs/sphinx/source/en/2-user_guide/1-training/7-tensor_runtime_production.md`
- `docs/sphinx/source/zh_CN/2-user_guide/1-training/7-tensor_runtime_production.md`

## Related Documents

- {doc}`ADR Index </adr/README>`
- {doc}`RL Infrastructure 开发标准 </zh_CN/4-developer_guide/0-index>`
- Discussion #1800
- Issue #2063
1 change: 1 addition & 0 deletions docs/sphinx/source/adr/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ orphan: true
| [ADR-0010 Fixed Model Variant Ownership Boundary](ADR-0010-fixed-model-variant-ownership-boundary.md) | Fixed variants / cross-repository boundary | Proposed |
| [ADR-0011 Torch-Only Manager-Based Runtime](ADR-0011-torch-only-manager-based-runtime.md) | Manager runtime / tensor lifecycle | Superseded |
| [ADR-0012 Sole Tensor Manager And Scoped Backends](ADR-0012-sole-tensor-manager-and-scoped-backends.md) | Manager runtime / backend scope | Proposed |
| [ADR-0013 CUDA MPS Single-Rank Execution Sharing](ADR-0013-cuda-mps-single-rank-execution-sharing.md) | Training runtime / execution sharing | Proposed |

## ADR Governance

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,12 +17,15 @@ autotuner and does not add multi-GPU scaling.
| `training.collector_metrics_interval` | `1` | Positive integer, at most `10000` |
| `training.replay_ingress_depth` | `2` | Positive integer, at most `16` |
| `training.replay_ingress_slot_rows` | `null` | `1` through `algo.num_envs`; `null` means `algo.num_envs` |
| `training.cuda_process_sharing` | `null` | `null` or explicit `mps` for single-rank MJWarp |
| Learner rows per synchronization | `algo.batch_size * algo.updates_per_step` | Constrained by the CUDA memory budget |

Values must be exact positive integers. Booleans, strings, floating-point
numbers, zero, and negative values fail closed. The G1 Motion Tracking / MJWarp
owner intentionally overrides only `collector_metrics_interval` to `100`; all
other tensor-runtime defaults remain as listed above.
Values other than `training.cuda_process_sharing` must be exact positive
integers. Booleans, strings, floating-point numbers, zero, and negative values
fail closed. `training.cuda_process_sharing` accepts only `null` or the explicit
string `mps`. The G1 Motion Tracking / MJWarp owner intentionally overrides only
`collector_metrics_interval` to `100`; all other tensor-runtime defaults remain
as listed above.

## Replay ingress tradeoffs

Expand Down Expand Up @@ -68,6 +71,8 @@ The runtime manifest records the effective evidence needed to audit a run:
- `inference_memory_budget` records the bounded CUDA inference-ring budget.
- `tensor_memory_budget` records the combined CUDA inference, replay storage,
replay ingress, learner batch, and workspace budget calculation.
- `cuda_process_sharing` records validated execution-sharing evidence when the
owner explicitly requests `mps`; it is absent for the default `null` mode.

The manifest is written before spawn when a budget decision is made. If an
unsafe combination is rejected, the error identifies the offending setting and
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,7 @@ the MJWarp owner inherits the MuJoCo FlashSAC owner and resolves to:
| Replay-ingress depth | `training.replay_ingress_depth=2` |
| Replay-ingress slot rows | `null`, resolving to `algo.num_envs` |
| Collector metric interval | `training.collector_metrics_interval=100` |
| CUDA process sharing | `training.cuda_process_sharing=null` |
| Scene | `src/unilab/assets/robots/g1/scene_flat.xml` |
| Motion | `motions/g1/dance1_subject2_part.npz` |

Expand Down Expand Up @@ -87,6 +88,56 @@ export CUDA_VISIBLE_DEVICES=<single-host-cuda-ordinal>
All CUDA ordinals inside trainer processes are relative to that mask. Do not
add more devices: M11 is single-GPU only.

## CUDA MPS Execution Sharing

`training.cuda_process_sharing` is an explicit execution-sharing mode for the
already-required rank-local topology. It is not a backend switch and does not
replace `--sim mjwarp` or owner YAML selection.

The default remains `null`: learner and collector keep independent CUDA contexts.
For the supported single-host, single-rank MJWarp off-policy topology, request
CUDA MPS with:

```bash
training.cuda_process_sharing=mps
```

Start an existing control daemon under user-owned directories before training:

```bash
export CUDA_MPS_PIPE_DIRECTORY=/absolute/path/mps/pipe
export CUDA_MPS_LOG_DIRECTORY=/absolute/path/mps/log
mkdir -p "$CUDA_MPS_PIPE_DIRECTORY" "$CUDA_MPS_LOG_DIRECTORY"
nvidia-cuda-mps-control -d
```

UniLab validates, but never starts or stops, that daemon. An `mps` request must
be Linux/NVIDIA CUDA, use MJWarp, resolve one physical GPU by UUID for learner
and collector, have `world_size=1`, and reach a live control socket/FIFO and
control daemon. Any unmet prerequisite fails before environment probing,
learner construction, or collector spawn; there is no silent multi-context
fallback. The error names the first prerequisite and the daemon start command.

`run_config.json` records the configured owner value. A valid run's
`run_summary.json` embeds `runtime_manifest.cuda_process_sharing` with the
configured/effective mode, learner and collector devices, physical UUID
evidence, control pipe, server PID, and validation state. The section is a
runtime-manifest v1 producer diagnostic, not a stable scalar contract.

MPS changes only GPU execution sharing. It does not change `env_steps_per_sync`,
learner/collector placement, or training semantics. Multi-GPU DP remains
unsupported until the separate DP gate and daemon topology decision are
completed. CUDA MPS is host- and deployment-dependent: in shared containers,
multi-user hosts, or restricted runners, control may be unavailable, and an
explicit request fails closed rather than silently degrading.

Stop the daemon after use:

```bash
export CUDA_MPS_PIPE_DIRECTORY=/absolute/path/mps/pipe
echo quit | nvidia-cuda-mps-control
```

Check the runtime that will execute the benchmark:

```bash
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,10 +17,12 @@ SAC、FlashSAC 与 WarpSAC 会在探测环境或构造 learner 之前解析所
| `training.collector_metrics_interval` | `1` | 正整数,最大 `10000` |
| `training.replay_ingress_depth` | `2` | 正整数,最大 `16` |
| `training.replay_ingress_slot_rows` | `null` | `1` 到 `algo.num_envs`;`null` 表示 `algo.num_envs` |
| `training.cuda_process_sharing` | `null` | `null` 或单 rank MJWarp 显式 `mps` |
| Learner 每次同步的行数 | `algo.batch_size * algo.updates_per_step` | 受 CUDA 显存预算约束 |

所有值必须是精确的正整数。布尔值、字符串、浮点数、零和负数都会 fail
closed。G1 Motion Tracking / MJWarp owner 仅有意将
除 `training.cuda_process_sharing` 外,所有值必须是精确的正整数。布尔值、字符串、
浮点数、零和负数都会 fail closed。`training.cuda_process_sharing` 只接受 `null`
或显式字符串 `mps`。G1 Motion Tracking / MJWarp owner 仅有意将
`collector_metrics_interval` 覆盖为 `100`;其他 tensor-runtime 默认值仍保持
上表所示。

Expand Down Expand Up @@ -64,6 +66,8 @@ Runtime manifest 记录审计 run 所需的有效证据:
- `inference_memory_budget` 记录有边界的 CUDA inference-ring 预算。
- `tensor_memory_budget` 记录 CUDA inference、replay storage、replay ingress、
learner batch 与 workspace 的组合预算计算。
- owner 显式请求 `mps` 时,`cuda_process_sharing` 记录已验证的
execution-sharing 证据;默认 `null` 模式下该字段不存在。

预算决策在 spawn 前写入 manifest。如果不安全的组合被拒绝,错误会指出具体
设置,且相关进程不会被启动。
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ MJWarp owner 继承 FlashSAC MuJoCo owner,并解析为:
| Replay-ingress depth | `training.replay_ingress_depth=2` |
| Replay-ingress slot rows | `null`,解析为 `algo.num_envs` |
| Collector metric interval | `training.collector_metrics_interval=100` |
| CUDA 进程共享 | `training.cuda_process_sharing=null` |
| Scene | `src/unilab/assets/robots/g1/scene_flat.xml` |
| Motion | `motions/g1/dance1_subject2_part.npz` |

Expand Down Expand Up @@ -82,6 +83,51 @@ export CUDA_VISIBLE_DEVICES=<single-host-cuda-ordinal>
trainer 进程内所有 CUDA ordinal 都相对于该 mask。不要添加更多设备:M11 仅
覆盖单 GPU。

## CUDA MPS 执行共享

`training.cuda_process_sharing` 是针对既有 rank-local 拓扑的显式
execution-sharing 模式。它不是 backend 开关,也不替代 `--sim mjwarp` 或 owner
YAML 选择。

默认 `null` 表示 learner 与 collector 保持独立 CUDA context。对受支持的单主机、
单 rank MJWarp off-policy 拓扑,使用以下设置请求 CUDA MPS:

```bash
training.cuda_process_sharing=mps
```

训练前在用户拥有的目录中启动既有 control daemon:

```bash
export CUDA_MPS_PIPE_DIRECTORY=/absolute/path/mps/pipe
export CUDA_MPS_LOG_DIRECTORY=/absolute/path/mps/log
mkdir -p "$CUDA_MPS_PIPE_DIRECTORY" "$CUDA_MPS_LOG_DIRECTORY"
nvidia-cuda-mps-control -d
```

UniLab 只验证、绝不 start/stop 该 daemon。`mps` 请求必须是 Linux/NVIDIA CUDA、
使用 MJWarp、learner 与 collector 按 UUID 解析到同一张物理 GPU、`world_size=1`,
且能访问 live control socket/FIFO 与 control daemon。任一条件不满足都会在
environment probe、learner construction 或 collector spawn 之前失败;没有静默
multi-context fallback。错误会指出第一个未满足条件以及 daemon 启动命令。

`run_config.json` 记录配置值。有效 run 的 `run_summary.json` 内嵌
`runtime_manifest.cuda_process_sharing`,包含 configured/effective 模式、learner
与 collector device、物理 UUID 证据、control pipe、server PID 和 validation 状态。
该 section 是 runtime-manifest v1 的 producer diagnostic,不是稳定标量契约。

MPS 只改变 GPU execution sharing,不改变 `env_steps_per_sync`、
learner/collector placement 或训练语义。多 GPU DP 在独立的 DP gate 与 daemon
拓扑决策完成前保持不支持。CUDA MPS 依赖 host 与部署方式:在共享容器、多用户
主机或受限 runner 中 control 可能不可用,显式请求会 fail closed,不会静默降级。

使用后停止 daemon:

```bash
export CUDA_MPS_PIPE_DIRECTORY=/absolute/path/mps/pipe
echo quit | nvidia-cuda-mps-control
```

检查实际执行 benchmark 的 runtime:

```bash
Expand Down
3 changes: 3 additions & 0 deletions src/unilab/conf/flashsac/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,9 @@ training:
wandb_notes: null
wandb_mode: null
sim_backend: mujoco
# null preserves independent CUDA contexts; 'mps' explicitly requests CUDA
# process sharing and fails closed if the host daemon is not already valid.
cuda_process_sharing: null
nan_guard:
enabled: true
buffer_size: 100
Expand Down
3 changes: 3 additions & 0 deletions src/unilab/conf/sac/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,9 @@ training:
wandb_notes: null
wandb_mode: null
sim_backend: mujoco
# null preserves independent CUDA contexts; 'mps' explicitly requests CUDA
# process sharing and fails closed if the host daemon is not already valid.
cuda_process_sharing: null
nan_guard:
enabled: true
buffer_size: 100
Expand Down
3 changes: 3 additions & 0 deletions src/unilab/conf/warpsac/config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,9 @@ training:
wandb_notes: null
wandb_mode: null
sim_backend: mujoco
# null preserves independent CUDA contexts; 'mps' explicitly requests CUDA
# process sharing and fails closed if the host daemon is not already valid.
cuda_process_sharing: null
nan_guard:
enabled: true
buffer_size: 100
Expand Down
46 changes: 38 additions & 8 deletions src/unilab/scripts/train_offpolicy.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@
configure_backend_process_device,
pin_genesis_device_before_cuda_init,
resolve_backend_env_device_id,
resolve_backend_process_device,
)
from unilab.training import (
assert_offpolicy_task_choice_matches_algo,
Expand All @@ -49,6 +50,10 @@
resolve_nan_guard_cfg,
should_run_playback,
)
from unilab.training.cuda_process_sharing import (
CudaProcessSharingEvidence,
probe_cuda_process_sharing,
)
from unilab.training.experiment import ExperimentTracker
from unilab.training.onnx_export import export_policy_onnx, verify_policy_onnx
from unilab.utils.checkpoint import (
Expand Down Expand Up @@ -151,12 +156,6 @@ def build_offpolicy_play_env_cfg_override(algo_name: str, cfg: DictConfig) -> di

def build_runner(algo_name: str, cfg: DictConfig, log_dir: str | None = None):
"""Build algorithm runner from unified Hydra config."""
env_factory = registry_env_factory(str(cfg.training.task_name), str(cfg.training.sim_backend))
from uni_rl.offpolicy.thread_budget import (
apply_torch_thread_runtime,
resolve_torch_thread_runtime,
)

# Cold-path DP CPU partition: each rank's collector owns one contiguous
# CPU block (single rank keeps the legacy unset behavior). The ids only
# reach the collector env override — never the num_envs=1 probe envs,
Expand All @@ -180,6 +179,24 @@ def build_runner(algo_name: str, cfg: DictConfig, log_dir: str | None = None):
)
if bound_device is not None:
rank_device = bound_device
collector_device = resolve_backend_process_device(
str(cfg.training.sim_backend),
rank_device,
)
cuda_process_sharing = probe_cuda_process_sharing(
getattr(cfg.training, "cuda_process_sharing", None),
rank_device,
collector_device,
backend=str(cfg.training.sim_backend),
world_size=dp_world_size,
)

env_factory = registry_env_factory(str(cfg.training.task_name), str(cfg.training.sim_backend))
from uni_rl.offpolicy.thread_budget import (
apply_torch_thread_runtime,
resolve_torch_thread_runtime,
)

# Every rank is single-GPU; rank-local visibility routes the backend payload.
env_cfg_override = apply_backend_env_device_override(
build_offpolicy_env_cfg_override(algo_name, cfg),
Expand Down Expand Up @@ -283,9 +300,23 @@ def build_runner(algo_name: str, cfg: DictConfig, log_dir: str | None = None):
else:
raise ValueError(f"Unsupported algo: {algo_name}")

_attach_cuda_process_sharing_manifest(runner, cuda_process_sharing)
return runner


def _attach_cuda_process_sharing_manifest(
runner: Any,
evidence: CudaProcessSharingEvidence,
) -> None:
"""Merge validated mode evidence into the producer-owned manifest."""

runtime_manifest = getattr(runner, "runtime_manifest", None)
if isinstance(runtime_manifest, dict):
runtime_manifest["cuda_process_sharing"] = evidence.manifest()
return
runner.runtime_manifest = {"cuda_process_sharing": evidence.manifest()}


def play_offpolicy(
algo_name: str,
cfg: DictConfig,
Expand Down Expand Up @@ -407,6 +438,7 @@ def play_offpolicy(

def main(cfg: DictConfig) -> None:
enable_faulthandler()
import torch

reject_removed_device_config(OmegaConf.select(cfg, "training.devices", default=None))
rank = current_dp_rank()
Expand Down Expand Up @@ -460,8 +492,6 @@ def main(cfg: DictConfig) -> None:
if rank == 0 and world_size > 1:
supervisor = DpRankSupervisor(world_size=world_size, log_dir=log_dir)

import torch

tracker = None
if not cfg.training.play_only and rank == 0:
tracker = ExperimentTracker(
Expand Down
Loading