Skip to content

Repository files navigation

DirectDrag

DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment

🌐 Project Page  |  📄 Paper  |  📚 arXiv

Sheng-Hao Liao, Shang-Fu Chen, Tai-Ming Huang, Wen-Huang Cheng, Kai-Lung Hua

IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026


DirectDrag teaser

Abstract

Drag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask- and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model's inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation.

Framework

DirectDrag framework

Repository Structure

DirectDrag/
├── directdrag_ui.py        # Gradio web UI (main entry point)
├── directdrag.py           # Thin runner: resizing, coordinate mapping, I/O
├── pipeline.py             # DirectDragger — the core editing pipeline
├── utils/                  # Library modules
│   ├── drag_utils.py       #   point tracking, motion supervision helpers
│   ├── continuous_drag.py  #   latent warpage / stretch operators
│   ├── lora_utils.py       #   per-image LoRA training
│   ├── readout_guidance/   #   readout-guided feature alignment
│   ├── unet_drag/          #   patched diffusers UNet (feature readout, KV control)
│   ├── eval_utils.py       #   IF (LPIPS) and MD metrics
│   └── gscore.py           #   GScore (LLM-as-judge) scoring + batch CLI
├── configs/rg_config.yaml  # Readout guidance configuration
├── rg_weights/             # Readout guidance checkpoints (shipped with the repo)
├── data/
│   ├── demo_samples/       # Example images + their drag instructions
│   ├── DragBench/          # DragBench benchmark data
│   └── Drag100/            # Drag100 benchmark data
├── scripts/                # Setup, launcher, benchmark and evaluation scripts
├── docker/                 # Dockerfile + docker-compose.yaml
├── notebooks/GScore.ipynb  # Original GScore notebook
└── assets/                 # Logo, figures, screenshots

Requirements

Hardware

GPU NVIDIA GPU with CUDA support
VRAM ≥ 16 GB recommended (Stable Diffusion v1.5 at 512×512, with per-image LoRA training and latent optimization)
Disk ~1 GB for this repository (includes the benchmark data and readout weights) + ~6 GB for the downloaded diffusion model cache

Software

  • Linux, NVIDIA driver + CUDA
  • Python 3.9 (Docker image) or 3.10 (native aarch64 setup)
  • PyTorch 2.0.1, diffusers 0.24.0, transformers 4.27.0, Gradio 3.50.2 (see requirements.txt)

The Stable Diffusion base model (runwayml/stable-diffusion-v1-5) and VAE (stabilityai/sd-vae-ft-mse) are downloaded automatically from Hugging Face on first run. Set HF_HOME if you want to control where that ~6 GB cache lives.

Installation

Option A — Docker (recommended)

The image is frakw/direct-drag:latest, built for x86_64 + CUDA 11.8. It requires the NVIDIA Container Toolkit.

git clone https://github.com/frakw/DirectDrag.git
cd DirectDrag
docker compose -f docker/docker-compose.yaml up -d   # pulls frakw/direct-drag:latest
docker exec -it direct-drag-container bash

The repository is bind-mounted into the container at /root/DirectDrag, so the image carries only the environment — code changes take effect without rebuilding.

Then, inside the container:

python directdrag_ui.py

The UI is published on http://localhost:7862 (port 7860 inside the container).

To build the image locally instead of pulling it:

docker compose -f docker/docker-compose.yaml build

Option B — Native (conda + pip)

git clone https://github.com/frakw/DirectDrag.git
cd DirectDrag

conda create -n DirectDrag python=3.9.20 -y
conda activate DirectDrag

pip install torch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2
pip install -r requirements.txt

Or create the environment straight from the pinned spec:

conda env create -f environment.yaml
conda activate DirectDrag

aarch64 / Blackwell (GB10, DGX Spark) users: the Docker image above is x86_64 + CUDA 11.8 and cannot target these GPUs. Use the dedicated one-shot setup instead:

./scripts/setup_aarch64.sh

It creates a Python 3.10 environment with a CUDA 13 PyTorch build and installs scripts/requirements-aarch64-cu13.txt.

Launching the UI

python directdrag_ui.py

Or use the helper script, which starts the server detached and manages the pidfile:

./scripts/run_ui.sh            # start (default port 7862)
./scripts/run_ui.sh status     # check
./scripts/run_ui.sh stop       # stop

Optional: GScore API key

GScore is an LLM-as-judge metric and calls the Gemini API. To use it, copy .env.example to .env and fill in your key:

cp .env.example .env
# .env
# GEMINI_API_KEY=your-key-here

IF and MD run fully locally and need no key.

Usage

1. Upload an image

Drop an image into the Original Image & Click Points panel. Images larger than 512px on the longest side are downscaled automatically (aspect ratio preserved) to keep memory usage bounded.

Upload an image

2. Mark handle and target points

Click on the image to place points. They alternate as pairs: the handle point (red) is what you grab, the target point (blue) is where it should end up, and a white arrow is drawn between them. Add as many pairs as you like. Use Undo Point / Undo Pair to correct mistakes.

No mask and no prompt are needed — DirectDrag infers the editable region from the point displacement itself.

Mark handle and target points

3. (Optional) Tune parameters

The tabs under the main panels expose every knob:

Tab What it controls
DirectDrag Parameters The three method components — Auto Soft Mask, Readout-Guided Feature Alignment, Latent Warpage — and their strengths. Turning them off reproduces the paper's ablations.
Drag Parameters Latent learning rate, drag end time step, point-tracking iterations per step
Diffusion Model Base diffusion model and VAE (SD v1.5 / SD 2.1 / SDXL, or a local checkpoint)
LoRA Parameters LoRA cache directory, training steps, learning rate, batch size, rank
Advanced Parameters Patch radii r1/r2, feature index, inversion strength, early-stopping limits

Parameters

4. Run

Press Run. The first run on a given image trains a per-image LoRA (progress is shown in the bar above the Drag Instruction panel); subsequent runs on the same image reuse the cached LoRA and skip straight to editing. The 10 most recently used images are kept in lora_tmp/, older ones are evicted automatically.

Running

5. Get the result

The edited image appears in the Dragged Image panel. Use the download card next to Run to save it (the file is also written to gradio_results/).

Result

6. Save, share and replay an edit

The Drag Instruction panel shows the current points as JSON:

{
  "points": [[192, 128], [240, 96]]
}

Points are absolute pixel coordinates in the (possibly downscaled) image, listed as alternating handle/target pairs. You can:

  • edit the JSON directly and press Import / Apply to move the points,
  • Download Instruction to save it as a .json file,
  • drop a .json file into the picker to load and apply it in one step.

Dropping an instruction next to an image with the same base name in data/demo_samples/ (e.g. candy.jpg + candy.json) makes it a clickable example at the bottom of the page — one click loads both the image and its points.

Drag instruction

7. Evaluate the result

Tick IF, MD and/or GScore in the Evaluation panel and press Evaluate:

  • IF — Image Fidelity, 1 − LPIPS between the input and the result (higher is better)
  • MD — Mean Distance between the dragged handle points and their targets, measured with DIFT correspondences (lower is better)
  • GScore — an LLM-as-judge quality score out of 10 (requires GEMINI_API_KEY)

IF and MD download their backing models the first time they are used, so the first evaluation takes noticeably longer than later ones. MD in particular pulls stabilityai/stable-diffusion-2-1, which is gated on Hugging Face — accept its licence and run huggingface-cli login first, or point DIRECTDRAG_DIFT_MODEL in your .env at a local copy of the model.

Evaluation

Benchmarks

Run DirectDrag over a whole benchmark:

python scripts/run_directdrag_dragbench.py data/DragBench
python scripts/run_directdrag_drag100.py data/Drag100

Then compute the metrics:

# IF + MD over a result folder
python scripts/run_eval_IF_MD.py --drag_bench_root data/DragBench --eval_root <result_dir>

# GScore (LLM-as-judge) over a result folder
python -m utils.gscore --dataset_path data/DragBench --result_path <result_dir>

# Drag100 DAI
python scripts/compute_drag100_DAI.py

scripts/extract_dragbench.py converts the original DragBench release into the folder layout these scripts expect.

Acknowledgements

This work builds on DragDiffusion, GoodDrag, Readout Guidance and diffusers. The DragBench and Drag100 benchmarks come from DragDiffusion and GoodDrag respectively.

Citation

@inproceedings{liao2026directdrag,
  title={DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment},
  author={Liao, Sheng-Hao and Chen, Shang-Fu and Huang, Tai-Ming and Cheng, Wen-Huang and Hua, Kai-Lung},
  booktitle={2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
  pages={8252--8261},
  year={2026},
  organization={IEEE}
}

License

Released under the Apache License 2.0.

About

[WACV 2026] DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages