From 5b985ec7d7583db6751cb4780138beedc9893771 Mon Sep 17 00:00:00 2001 From: st-bang97 Date: Fri, 2 Oct 2026 06:54:30 +0000 Subject: [PATCH 1/2] Add Reflow ZeRO-3 fine-tuning examples Signed-off-by: st-bang97 --- training/DeepSpeed-Reflow/.gitignore | 20 + training/DeepSpeed-Reflow/README.md | 215 +++++++ training/DeepSpeed-Reflow/check_bitexact.sh | 117 ++++ .../finetune_llama-13b_1gpu.sh | 153 +++++ .../finetune_llama-70b_8gpu.sh | 157 +++++ .../finetune_llama-70b_8gpu_lion.sh | 158 +++++ .../finetune_llama-7b_1gpu.sh | 152 +++++ .../finetune_llama-8b_1gpu.sh | 152 +++++ .../DeepSpeed-Reflow/finetune_opt-30b_1gpu.sh | 155 +++++ .../DeepSpeed-Reflow/finetune_opt-30b_4gpu.sh | 155 +++++ .../finetune_opt-30b_4gpu_lion.sh | 134 +++++ .../DeepSpeed-Reflow/finetune_opt-30b_8gpu.sh | 155 +++++ .../finetune_opt-30b_8gpu_lion.sh | 134 +++++ .../finetune_opt-350m-1gpu.sh | 153 +++++ training/DeepSpeed-Reflow/finetune_zero3.py | 538 ++++++++++++++++++ training/DeepSpeed-Reflow/requirements.txt | 8 + training/DeepSpeed-Reflow/run_compare.sh | 232 ++++++++ .../DeepSpeed-Reflow/tests/test_examples.py | 229 ++++++++ 18 files changed, 3017 insertions(+) create mode 100644 training/DeepSpeed-Reflow/.gitignore create mode 100644 training/DeepSpeed-Reflow/README.md create mode 100755 training/DeepSpeed-Reflow/check_bitexact.sh create mode 100755 training/DeepSpeed-Reflow/finetune_llama-13b_1gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_llama-70b_8gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_llama-70b_8gpu_lion.sh create mode 100755 training/DeepSpeed-Reflow/finetune_llama-7b_1gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_llama-8b_1gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_opt-30b_1gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_opt-30b_4gpu.sh create mode 100755 training/DeepSpeed-Reflow/finetune_opt-30b_4gpu_lion.sh create mode 100755 training/DeepSpeed-Reflow/finetune_opt-30b_8gpu.sh create mode 100644 training/DeepSpeed-Reflow/finetune_opt-30b_8gpu_lion.sh create mode 100755 training/DeepSpeed-Reflow/finetune_opt-350m-1gpu.sh create mode 100644 training/DeepSpeed-Reflow/finetune_zero3.py create mode 100644 training/DeepSpeed-Reflow/requirements.txt create mode 100644 training/DeepSpeed-Reflow/run_compare.sh create mode 100644 training/DeepSpeed-Reflow/tests/test_examples.py diff --git a/training/DeepSpeed-Reflow/.gitignore b/training/DeepSpeed-Reflow/.gitignore new file mode 100644 index 000000000..c198614a2 --- /dev/null +++ b/training/DeepSpeed-Reflow/.gitignore @@ -0,0 +1,20 @@ +# Conda environment setup — local to our environment, not part of the public example. +# Public users install via requirements.txt; this captures our exact dev environment. +environment.yml +conda_env*.yml +*.conda.yml + +# Generated at runtime by the launcher scripts (DeepSpeed config + training outputs/logs). +*_config.json +*_output/ +__pycache__/ +wandb/ + +# JIT-built CPU-Adam/Lion op extensions (torch cpp_extension build dirs). +cpu_adam/ +cpu_lion/ +fused_adam/ + +# Comparison / benchmark output collected by run_compare.sh. +compare_out/ +bitexact_out/ diff --git a/training/DeepSpeed-Reflow/README.md b/training/DeepSpeed-Reflow/README.md new file mode 100644 index 000000000..e6ef92642 --- /dev/null +++ b/training/DeepSpeed-Reflow/README.md @@ -0,0 +1,215 @@ +# Reflow Fine-Tuning Examples + +Fine-tune large language models with [DeepSpeed](https://www.deepspeed.ai/) ZeRO Stage 3 + **Reflow**, an asynchronous CPU-offload optimizer for mixed-precision BF16 training. Reflow keeps the optimizer state and FP32 master weights on the CPU like ZeRO-Offload, but overlaps the CPU optimizer work with the backward pass instead of running it serially afterward. It uses the **same GPU memory** as ZeRO-Offload and **~12% less host RAM** (OPT-30B: ~549 vs ~625 GB) — gradients are offloaded in half precision (BF16) and promoted to FP32 inside the CPU kernel, so there is no CPU-side FP32 gradient buffer (see [Memory](#memory) for measured numbers). + +Reflow supports **Adam/AdamW and Lion** with BF16 model parameters and gradients, FP32 master weights, and FP32 optimizer states. The same scripts run the Reflow path or the plain ZeRO-Offload baseline. `check_bitexact.sh` compares deterministic per-step losses for a selected model and configuration. + +## Quick Start + +### 1. Install dependencies + +```bash +pip install -r requirements.txt +# The runtime change is currently proposed in deepspeedai/DeepSpeed#8719. +git clone https://github.com/deepspeedai/DeepSpeed.git +git -C DeepSpeed fetch origin pull/8719/head:reflow +git -C DeepSpeed switch reflow +pip install -e ./DeepSpeed +``` + +Reflow is a DeepSpeed runtime feature, so you also need a DeepSpeed build that includes it — the `deepspeed.runtime.reflow` module and the `reflow_*` CPU-Adam/Lion ops (JIT-built on first use). No custom modeling code is required: the examples fine-tune plain Hugging Face Transformers models (`--model_name`) on `tatsu-lab/alpaca` by default (override with `--dataset_name`). + + +The launch scripts default to FlashAttention 2. Install it separately after PyTorch: + +```bash +pip install flash-attn --no-build-isolation +``` + +Alternatively, set `ATTN_IMPLEMENTATION=sdpa` or `eager` when running a launch script; these use PyTorch attention and do not require `flash-attn`. Llama checkpoints may require Hugging Face access approval and authentication. + +### 2. Enable Reflow (one block) + +Add a `reflow` block to the `zero_optimization` section of a ZeRO Stage 3 config with CPU optimizer offload (the scripts generate this for you): + +```jsonc +"offload_optimizer": { + "device": "cpu", + "pin_memory": true +}, +"reflow": {} // the only addition Reflow requires +``` + +An empty block is enough; every Reflow setting (NUMA core binding, thread counts) is optional — see [Configuration](#configuration). Remove the block to fall back to plain ZeRO-Offload. + +### 3. Run a fine-tuning script + +Each script takes the mode (`reflow` or `zerooffload`) as the first argument and an optional **global batch size** as the second (it must be divisible by the GPU count): + +```bash +# bash