Repository navigation
Conversation
|
Tested PR #2065 on the requested server Environment
The server had unrelated dirty UniLab work. It was preserved in stash Focused testsAn initial run against the server's older unpinned Command: UNILAB_LOCAL_UNISIM="$HOME/unilabsim/unisim" CUDA_VISIBLE_DEVICES=0 \
uv run --no-sync pytest \
tests/training/test_cuda_process_sharing.py \
tests/algos/test_offpolicy_double_buffer_runner.py \
tests/utils/test_experiment_tracking.py -qReal MPS probe: {
"configured": "mps",
"effective": "mps",
"validated": true,
"learner_device": "cuda:0",
"collector_device": "cuda:0",
"learner_gpu_uuid": "F40821277B085E35212E5BB52B872AC6",
"collector_gpu_uuid": "F40821277B085E35212E5BB52B872AC6",
"server_pid": 566566
}Phase 3 benchmarkTwo repeats per arm, 300 iterations each, metrics averaged over final 100 iterations. FlashSAC / G1 Motion Tracking / MJWarp
Threshold: ≥ +10%; passed. SAC / G1 Walk Flat / MJWarp
Threshold: ≥ +20%; passed. Integrity evidenceAll eight runs passed:
No residual compute processes remained after the MPS daemon was stopped. ArtifactsServer-local, not committed: Contains |
Summary
training.cuda_process_sharing: null | mpsowner mode for the shared SAC/FastSAC/FlashSAC/WarpSAC off-policy path.mpsrequests never fall back to independent CUDA contexts.Linked Work
develop/tensor-runtimecuda_process_sharingis added as a runtime-manifest v1 producer diagnostic, not a new stable manifest v2 field.Validation
make test-allpassed on the final local head before this PR was created or updatedCommands actually run:
Benchmark matrix on the final head: two tasks, default versus explicit MPS, two repeats, 300 iterations each; metrics averaged over the final 100 iterations:
All eight completed runs reached
status=completedand 300/300 iterations. Every run had normal shutdown, no cleanup errors, replay published/release sequences equal, final replay occupancy zero, and zero dropped batches/early returns. MPS runs recorded validated manifest evidence with matching learner/collector UUIDs and the same control pipe/server PID.Benchmark commands:
Remote CI route:
main; remote CI is not scheduled. Localmake test-allis the test gate. The later integration PR tomainwill run remote CI.Impact
mps; default behavior is unchanged for all backendsmps; default behavior is unchangedenv_steps_per_syncis unchanged.Artifacts
/home/user/ws/unilabsim2/UniLab/benchmarks/cuda-mps-20261004/(not committed;README.mdandresults.jsonrecord commands, host, metrics, and manifest evidence)Checklist
Follow-up:
CUDA_MPS_ACTIVE_THREAD_PERCENTAGEremain evidence-gated follow-ups.