Skip to content

Add affine AutoEP checkpoint placement - #8544

Merged
delock merged 8 commits into
deepspeedai:masterfrom
jinyouzhi:autoep-affine-placement
Sep 24, 2026
Merged

delock merged 8 commits into
deepspeedai:masterfrom
jinyouzhi:autoep-affine-placement

Conversation

@jinyouzhi

Copy link
Copy Markdown
Contributor

Follow up #8385

This pull request introduces support for flexible expert placement in DeepSpeed's AutoEP (Automatic Expert Placement) system, enabling non-uniform, non-contiguous, and replicated expert layouts. The main changes add a versioned expert placement descriptor, validation logic, and integration into checkpoint consolidation and metadata validation. This lays the groundwork for more advanced expert scheduling and model parallelism strategies.

The most important changes are:

AutoEP Expert Placement Descriptor and Affine Map Lowering

  • Added a new module autoep_affine.py that defines the expert placement descriptor, validation, legacy uniform descriptor synthesis, and lowering to affine maps for sharded tensor reconstruction. This enables flexible, versioned expert placement beyond the legacy uniform contiguous layout.

Integration into Checkpoint Consolidation and Metadata

  • Updated autoep_universal.py to:
    • Accept and validate the new expert placement descriptor in layer metadata, relaxing the requirement that num_local_experts * ep_size == num_experts when a placement is provided.
    • Use the placement descriptor and affine map for reconstructing full expert tensors during checkpoint consolidation, supporting arbitrary expert layouts. [1] [2]
    • Validate placement consistency during expert file consolidation.
    • Import and use the new placement logic. [1] [2]

Metadata Validation Enhancements

  • Updated autoep_zero3_metadata.py to:
    • Import and use placement validation and legacy descriptor synthesis.
    • Track and validate placements for each layer entry, and check runtime layer consistency. [1] [2] [3]
    • Handle the presence or absence of the placement descriptor during partitioned metadata validation.

Documentation Updates

  • Expanded the affine IR specification (affine_ir_spec.md) to document the new AutoEP placement descriptor, its semantics, and its integration into the IR and runtime, clarifying the distinction between placement provenance and scheduling.

Bugfixes and Robustness

  • Fixed a potential bug in affine.py by skipping empty piece lists during tensor rebuilding, preventing errors when a rank has no assigned pieces.

Related: #8252, #8230.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@delock
delock self-requested a review September 17, 2026 06:40
@jinyouzhi
jinyouzhi marked this pull request as ready for review September 18, 2026 16:06
@delock

delock commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Hi @jinyouzhi thank you for your PR. This PR overall looks good to me, I have two questions:

  1. did you tested that the new checkpoint format can be correctly converted to UC during training?
  2. I see that UC info in checkpoint contains dict rather than affine map. Lower to affine map happens during convertion to UC checkpoint. For AutoTP, affine map is written in UC info in checkpoint. Is this difference a design choice?

@jinyouzhi

jinyouzhi commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Thank you for comments.
For 1, I am design e2e test to validate loss curve consistency ( 50 steps -> save/load -> 50 steps vs 100 steps)
For 2, aggree with your observation, AutoEP should align with AutoTP by persisting a versioned per-parameter affine map at save time. I will retain the placement descriptor for provenance and legacy fallback.
@delock

jinyouzhi and others added 2 commits September 22, 2026 07:25
…e-autoep-checkpoint-placement

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

# Conflicts:
#	deepspeed/module_inject/auto_ep_layer.py
Persist versioned per-parameter affine maps in AutoEP checkpoint metadata and use them for conversion and restore, while retaining placement descriptors for provenance and legacy checkpoints.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@jinyouzhi

Copy link
Copy Markdown
Contributor Author

@delock

I completed an end-to-end loss-resume validation using the DeepSpeedExamples script (deepspeedai/DeepSpeedExamples#1014 ) training/deepspeed_finetune_demo/run_autoep_affine_ir_checkpoint_experiment.sh , with Moonlight-16B-A3B on 4 GPUs (AutoEP=4, ZeRO-3 with CPU parameter/optimizer offload).

The experiment trains an uninterrupted 100-step baseline, saves a native checkpoint at step 50, then resumes both directly from the native checkpoint and after  ds_to_universal.py  conversion. The converter completed the AutoEP ZeRO-3 expert-state consolidation successfully, and all runs had finite loss values. For every resumed step (51–100), the native-resume and Universal-checkpoint-resume losses are exactly identical. Their shared difference from the uninterrupted baseline is  max_abs_error=0.1374  and  mean_abs_error=0.01052 ; because the native and Universal paths match at every step, this difference is not introduced by the Universal/AffineIR conversion.

I attached the loss plot below.

loss_comparison

Use explicit placement rank IDs for local expert lookup and reject checkpoints whose persisted affine map conflicts with placement provenance.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@@ -288,19 +336,27 @@ def consolidate_autoep_zero12_expert_states(temp_dir, output_dir, expert_param_i

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is consolidate_autoep_zero12_expert_states still needed?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

also if it is no longer called, does it mean this functionality is lost? Is it intentionally not called?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed. thanks

Consolidate ZeRO FP32 master expert weights and optimizer states after per-expert model files, and verify native and Universal loss-resume parity for ZeRO stages 1 and 2.

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@jinyouzhi

Copy link
Copy Markdown
Contributor Author

fix ZeRO1/2 issues and also have a test with ZeRO-2 @delock

loss_zero2

Remove the unreachable expp_rank optimizer consolidation helper; ZeRO-1/2 optimizer state is handled by the active shard merger. Keep expert-file FP32 fallbacks in the temporary directory and let the merger write each final tensor once.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@delock
delock enabled auto-merge September 24, 2026 03:53
auto-merge was automatically disabled September 24, 2026 05:02

Head branch was pushed to by a user without write access

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
@jinyouzhi
jinyouzhi force-pushed the autoep-affine-placement branch from 1a6df6a to 87a00c9 Compare September 24, 2026 05:22
@delock
delock added this pull request to the merge queue Sep 24, 2026
Merged via the queue into deepspeedai:master with commit 3571027 Sep 24, 2026
13 checks passed
pull Bot pushed a commit to AmirulAndalib/DeepSpeed that referenced this pull request Sep 29, 2026
Restore now takes this rank's shard from the map the layer published,
which is the same piece list conversion uses run in the opposite
direction. That completes the symmetry §7 of `affine_ir_spec.md`
describes and that conversion has had since deepspeedai#8385: until now the two
sides described the same layout in two different ways, and could
disagree.

Follows deepspeedai#8519 (layers emit the map) and deepspeedai#8575 (GPTBigCode and Yuan
describe themselves).

### This one is not additive

The map lives on the parameter, built by this job's layer. It is `M_t` —
it describes the topology being restored into, regardless of how the
source checkpoint was converted. So the new path is taken for **every**
AutoTP restore, not only for checkpoints that were converted through a
map. A parameter whose layer published no map falls back to the existing
category keys unchanged.

That is the correct reading of the spec rather than a shortcut: §6.1
keeps `M_t` out of the file precisely because the restoring job derives
it from its own layers. But it does mean the blast radius here is wider
than the three PRs before it, so the equivalence is worth checking
rather than assuming.

I ran `tests/unit/checkpoint/` and `tests/unit/runtime/tensor_parallel/`
(385 tests) against this branch and against upstream master in a
separate worktree. The failing set is **identical by name**, not merely
the same count — 75 failed, 187 passed either way, all of them FusedAdam
JIT, fp16-on-CPU, or the pre-existing `st_st_size` typo in
`test_convert_checkpoint.py`.

### The tests are built to discriminate

`test_restore_prefers_the_map_over_the_category_keys` gives one
parameter deliberately contradictory geometry: the category keys
describe a plain dim-0 split, so rank 0 would take the first half, while
the map hands rank 0 the second half. Which half rank 0 receives says
which side of the metadata restore actually consulted.

This matters because the old path covers the same layouts correctly.
Without a discriminating case the map branch could be deleted and every
existing test would stay green. I checked: removing the branch fails
that test, while `test_restore_falls_back_when_no_map_was_published`
keeps passing.

### Validation

112 passed on CPU/gloo (`DS_ACCELERATOR=cpu LOCAL_SIZE=4`): the affine
suite, the resume matrix, the coverage test,
`tests/unit/runtime/tensor_parallel/`, and deepspeedai#8544's AutoEP affine suite,
which this now sits on top of after the rebase.


### Scale powers

Restore resolves a partition once per state file, so each state takes
its own power of a piece's scale. The power table lives in affine.py as
SCALE_POWER_BY_STATE — ds_to_universal.py had its own copy since deepspeedai#8519,
so conversion and restore now read one definition rather than two that
could drift.

This also reaches the refusal built for scaled optimizer states, which
only triggers on a non-parameter power. No layer emits a scaled map
today, so it is ahead of the path rather than behind it.

Related: deepspeedai#8252, deepspeedai#8230.

cc @delock

---------

Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants