Repository navigation
feat(macos): enable mujoco/mjbatch tensor-runtime off-policy training on macOS - #2061
Merged
TATP-233 merged 4 commits intoOct 4, 2026
Merged
Conversation
UniLab HEAD consumes the packed-reset public-width contract (SimBackend.get_public_state_widths / PublicStateWidths) and the selected-reset publication contract that only exist at the sibling branch tips; the stale manifest pins predated them and left the mujoco HOST_BRIDGE tensor reset dispatch without a viable packed commit (issue #2055). - unisim-core: 09d5caed -> 5e835923 - unilab-rl: 75680c23 -> 036fcba6 - mjbatch-uni: unchanged (tip == pin) Refs remain develop/tensor-runtime; the integration branches in the sibling repos are based on these exact heads. No main mutation, no PyPI publication.
The fail-closed CUDA-only platform gate (#1811) fires before the fake-runtime and mock-worker backend suites reach their mocks, so on CUDA-less hosts (macOS, CPU-only Linux) 58 adapter tests failed and 18 errored despite being designed to run without a GPU. Restore them by bypassing the gate inside those suites; the gate itself remains covered by tests/base/test_cuda_backend_platform_preflight.py. Also repair four smaller suite debts: - skip the CUDA-device observation naming test when CUDA is absent; - drop the never-committed outputs/velb_compare/micro_velb.py entry from the historical benchmark scope test; - skip the mjwarp soak command test when the mjwarp extra is missing; - skip the soak resource-monitor test off Linux (sampling reads /proc/<pid>/task). Non-slow suite on macOS arm64: 1801 passed, 59 skipped (was 1724 passed / 62 failed / 18 errors). No runtime or contract change. Part of issue #2055.
…f-CUDA Extend the FlashSAC g1_motion_tracking mujoco slow rollout with a selected-row reset through the packed host-bridge tensor command commit — the path training autoreset takes (issue #2055); it was the uncovered gap behind the mujoco motion-tracking reset failure. Also skip the two real-Newton G1 owner cases when CUDA is unavailable instead of failing on the platform gate's RuntimeError (they only caught ImportError before).
6 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes the UniLab portion of roadmap #2055: run the completed
npenv→torchenv runtime on macOS (Apple Silicon) for the MuJoCo/mjbatch
HOST_BRIDGE backend, with SAC +
g1_walk_flatand FlashSAC +g1_motion_trackingtraining on CPU physics + MPS learner.chore(workspace): realigntensor_runtime_workspace.jsonsibling pinswith the
develop/tensor-runtimebranch tips. UniLab HEAD already consumesthe packed-reset public-width contract
(
SimBackend.get_public_state_widths/PublicStateWidths) that onlyexists at the tips; the stale pins made
ResetStateTransaction.can_commit_packed()fail closed on mujoco, leavingthe host-bridge tensor reset dispatch (added in
24e57ffa) without a viablepacked commit and breaking
flashsac + g1_motion_tracking --sim mujocoatthe first autoreset. Sibling integration branches
dev/issue-2055-macos-tensor-runtimeare based on those exact tips; nosibling code change was required.
test(platform): the fake-runtime / mock-worker CUDA-only backend suitesbypass the factory platform gate (dedicated gate coverage stays in
tests/base/test_cuda_backend_platform_preflight.py), restoring 58 failed(CUDA observation naming, mjwarp-extra soak command, Linux-
/procsoakresource monitor) and removal of the never-committed
outputs/velb_compare/micro_velb.pybenchmark-scope entry.test(g1): new slow-lane regression case — selected-row reset through thepacked host-bridge tensor motion command commit on mujoco; newton G1 owner
cases skip off-CUDA.
Platform / behavior impact
macOS + MuJoCo (mjbatch) only. No change to runtime code, contracts, CUDA
behavior, Linux behavior, or training semantics: the only non-test change is
the workspace manifest realignment. No npenv-runtime / np-rng revert. No
mainordevelop/tensor-runtimemutation. No PyPI publication.Validation (exact commands on this head, macOS arm64, torch 2.14.0 MPS)
make check— pass (ruff, mypy 144 files, pyright 0 errors).make test(uv run pytest -m "not slow" -q) — 1801 passed, 59 skipped.make test-all— pass (check + coverage lane + benchmark smoke 35/35 and 36/36).uv run pytest tests/envs/test_env_configs.py -m slow -k mujoco— 5 passed (incl. new selected-reset case).uv run pytest tests/envs/locomotion/g1/test_g1_owner_contract.py -m slow— 17 passed, 4 skipped.uv run pytest tests/tasks/test_flashsac_owner_contract.py— 10 passed.algo.max_iterations=100,training.no_play=true):uv run train --algo sac --task g1_walk_flat --sim mujoco— 100/100, normal_completion, collector exit 0.uv run train --algo flashsac --task g1_motion_tracking --sim mujoco— 100/100, normal_completion, collector exit 0.uv run train --algo sac --task g1_walk_flat --sim mujoco training.play_render_mode=record(owner-default
algo.max_iterations=5000) — completed 5000/5000,10,260,480 env steps in 11.8 min (~14.3k steps/s), final mean reward
278.98 / best 303.60 / episode length 946.85,
normal_completion,collector exit 0, no residual processes; ONNX export verified
(max_diff 7.8e-07); 800-frame playback video rendered.
Raw evidence: issue #2055 comments.
Governance
Declared base:
develop/tensor-runtime. Related: #1706 (NpEnv removal),#2052 (Newton canonical tensor runtime),
5-contributing_workflow.mdtensor-runtime sibling dependency transition (profile stays local-editable;
publication remains out of scope). No new public contract, no ADR required:
the packed-reset width contract is declared by the sibling branch tips this
manifest now pins.