DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment
🌐 Project Page | 📄 Paper | 📚 arXiv
Sheng-Hao Liao, Shang-Fu Chen, Tai-Ming Huang, Wen-Huang Cheng, Kai-Lung Hua
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026
Drag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask- and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model's inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation.
DirectDrag/
├── directdrag_ui.py # Gradio web UI (main entry point)
├── directdrag.py # Thin runner: resizing, coordinate mapping, I/O
├── pipeline.py # DirectDragger — the core editing pipeline
├── utils/ # Library modules
│ ├── drag_utils.py # point tracking, motion supervision helpers
│ ├── continuous_drag.py # latent warpage / stretch operators
│ ├── lora_utils.py # per-image LoRA training
│ ├── readout_guidance/ # readout-guided feature alignment
│ ├── unet_drag/ # patched diffusers UNet (feature readout, KV control)
│ ├── eval_utils.py # IF (LPIPS) and MD metrics
│ └── gscore.py # GScore (LLM-as-judge) scoring + batch CLI
├── configs/rg_config.yaml # Readout guidance configuration
├── rg_weights/ # Readout guidance checkpoints (shipped with the repo)
├── data/
│ ├── demo_samples/ # Example images + their drag instructions
│ ├── DragBench/ # DragBench benchmark data
│ └── Drag100/ # Drag100 benchmark data
├── scripts/ # Setup, launcher, benchmark and evaluation scripts
├── docker/ # Dockerfile + docker-compose.yaml
├── notebooks/GScore.ipynb # Original GScore notebook
└── assets/ # Logo, figures, screenshots
| GPU | NVIDIA GPU with CUDA support |
| VRAM | ≥ 16 GB recommended (Stable Diffusion v1.5 at 512×512, with per-image LoRA training and latent optimization) |
| Disk | ~1 GB for this repository (includes the benchmark data and readout weights) + ~6 GB for the downloaded diffusion model cache |
- Linux, NVIDIA driver + CUDA
- Python 3.9 (Docker image) or 3.10 (native aarch64 setup)
- PyTorch 2.0.1, diffusers 0.24.0, transformers 4.27.0, Gradio 3.50.2 (see
requirements.txt)
The Stable Diffusion base model (runwayml/stable-diffusion-v1-5) and VAE
(stabilityai/sd-vae-ft-mse) are downloaded automatically from Hugging Face on first
run. Set HF_HOME if you want to control where that ~6 GB cache lives.
The image is frakw/direct-drag:latest, built for x86_64 + CUDA 11.8. It requires the
NVIDIA Container Toolkit.
git clone https://github.com/frakw/DirectDrag.git
cd DirectDrag
docker compose -f docker/docker-compose.yaml up -d # pulls frakw/direct-drag:latest
docker exec -it direct-drag-container bashThe repository is bind-mounted into the container at /root/DirectDrag, so the image
carries only the environment — code changes take effect without rebuilding.
Then, inside the container:
python directdrag_ui.pyThe UI is published on http://localhost:7862 (port 7860 inside the container).
To build the image locally instead of pulling it:
docker compose -f docker/docker-compose.yaml buildgit clone https://github.com/frakw/DirectDrag.git
cd DirectDrag
conda create -n DirectDrag python=3.9.20 -y
conda activate DirectDrag
pip install torch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2
pip install -r requirements.txtOr create the environment straight from the pinned spec:
conda env create -f environment.yaml
conda activate DirectDragaarch64 / Blackwell (GB10, DGX Spark) users: the Docker image above is x86_64 + CUDA 11.8 and cannot target these GPUs. Use the dedicated one-shot setup instead:
./scripts/setup_aarch64.shIt creates a Python 3.10 environment with a CUDA 13 PyTorch build and installs
scripts/requirements-aarch64-cu13.txt.
python directdrag_ui.pyOr use the helper script, which starts the server detached and manages the pidfile:
./scripts/run_ui.sh # start (default port 7862)
./scripts/run_ui.sh status # check
./scripts/run_ui.sh stop # stopGScore is an LLM-as-judge metric and calls the Gemini API. To use it, copy
.env.example to .env and fill in your key:
cp .env.example .env
# .env
# GEMINI_API_KEY=your-key-hereIF and MD run fully locally and need no key.
Drop an image into the Original Image & Click Points panel. Images larger than 512px on the longest side are downscaled automatically (aspect ratio preserved) to keep memory usage bounded.
Click on the image to place points. They alternate as pairs: the handle point (red) is what you grab, the target point (blue) is where it should end up, and a white arrow is drawn between them. Add as many pairs as you like. Use Undo Point / Undo Pair to correct mistakes.
No mask and no prompt are needed — DirectDrag infers the editable region from the point displacement itself.
The tabs under the main panels expose every knob:
| Tab | What it controls |
|---|---|
| DirectDrag Parameters | The three method components — Auto Soft Mask, Readout-Guided Feature Alignment, Latent Warpage — and their strengths. Turning them off reproduces the paper's ablations. |
| Drag Parameters | Latent learning rate, drag end time step, point-tracking iterations per step |
| Diffusion Model | Base diffusion model and VAE (SD v1.5 / SD 2.1 / SDXL, or a local checkpoint) |
| LoRA Parameters | LoRA cache directory, training steps, learning rate, batch size, rank |
| Advanced Parameters | Patch radii r1/r2, feature index, inversion strength, early-stopping limits |
Press Run. The first run on a given image trains a per-image LoRA (progress is shown
in the bar above the Drag Instruction panel); subsequent runs on the same image reuse the
cached LoRA and skip straight to editing. The 10 most recently used images are kept in
lora_tmp/, older ones are evicted automatically.
The edited image appears in the Dragged Image panel. Use the download card next to
Run to save it (the file is also written to gradio_results/).
The Drag Instruction panel shows the current points as JSON:
{
"points": [[192, 128], [240, 96]]
}Points are absolute pixel coordinates in the (possibly downscaled) image, listed as alternating handle/target pairs. You can:
- edit the JSON directly and press Import / Apply to move the points,
- Download Instruction to save it as a
.jsonfile, - drop a
.jsonfile into the picker to load and apply it in one step.
Dropping an instruction next to an image with the same base name in data/demo_samples/
(e.g. candy.jpg + candy.json) makes it a clickable example at the bottom of the page —
one click loads both the image and its points.
Tick IF, MD and/or GScore in the Evaluation panel and press Evaluate:
- IF — Image Fidelity,
1 − LPIPSbetween the input and the result (higher is better) - MD — Mean Distance between the dragged handle points and their targets, measured with DIFT correspondences (lower is better)
- GScore — an LLM-as-judge quality score out of 10 (requires
GEMINI_API_KEY)
IF and MD download their backing models the first time they are used, so the first
evaluation takes noticeably longer than later ones. MD in particular pulls
stabilityai/stable-diffusion-2-1, which is gated on Hugging Face — accept its
licence and run huggingface-cli login first, or point DIRECTDRAG_DIFT_MODEL in your
.env at a local copy of the model.
Run DirectDrag over a whole benchmark:
python scripts/run_directdrag_dragbench.py data/DragBench
python scripts/run_directdrag_drag100.py data/Drag100Then compute the metrics:
# IF + MD over a result folder
python scripts/run_eval_IF_MD.py --drag_bench_root data/DragBench --eval_root <result_dir>
# GScore (LLM-as-judge) over a result folder
python -m utils.gscore --dataset_path data/DragBench --result_path <result_dir>
# Drag100 DAI
python scripts/compute_drag100_DAI.pyscripts/extract_dragbench.py converts the original DragBench release into the folder
layout these scripts expect.
This work builds on DragDiffusion, GoodDrag, Readout Guidance and diffusers. The DragBench and Drag100 benchmarks come from DragDiffusion and GoodDrag respectively.
@inproceedings{liao2026directdrag,
title={DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment},
author={Liao, Sheng-Hao and Chen, Shang-Fu and Huang, Tai-Ming and Cheng, Wen-Huang and Hua, Kai-Lung},
booktitle={2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
pages={8252--8261},
year={2026},
organization={IEEE}
}Released under the Apache License 2.0.









