ATOD is a benchmark and evaluation framework for agentic task-oriented dialogue systems. It accompanies the AACL 2026 paper:
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
The repository contains the fixed 1,000-dialogue ATOD benchmark, the pipeline that generated it, and the code required to run ATOD-Eval.
data/ Fixed ATOD benchmark and dataset documentation
generation/ Six-stage synthetic dialogue generation pipeline
MemSys/ Agentic symbolic-vector memory evaluator
evaluation/ End-of-dialogue ATOD evaluation
ATODEval/ Dependency, completion, memory, and quality metrics
scripts/ Dataset validation and statistics utilities
tests/ Offline unit tests
The repository excludes baseline experiment harnesses, model weights, credentials, generated result caches, the upstream Schema-Guided Dialogue dataset, and paper/review materials.
The release contains:
| Split | Dialogues |
|---|---|
| Medium | 428 |
| Complex | 572 |
| Total | 1,000 |
Each file is a compact JSON array. Every dialogue contains a stable ID, a complexity label, a goal graph, dialogue turns, goal-status transitions, and complexity metadata. See data/README.md and DATASET_CARD.md.
Validate the release before use:
python scripts/validate_data.py
python scripts/compute_dataset_stats.pyPython 3.10 or newer is recommended.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtThe released evaluator uses the Bedrock runtime. Configure credentials and a
region through the standard AWS SDK settings, then provide a compatible model
ID either through ATOD_MODEL_ID or the --model-id option:
export ATOD_MODEL_ID="<model-id>"The exact model configuration used for the reported experiments is described in the paper.
generation/ contains the pipeline that produced the benchmark: goal extraction from the Schema-Guided Dialogue dataset, co-occurrence graph construction, random-walk trajectory sampling, model-based trajectory annotation and complexity labelling, dialogue generation with verifiers, and turn-level goal status annotation. Re-running it produces a new synthetic set, not the released data. See generation/README.md.
SGD_DIR=/path/to/dstc8-schema-guided-dialogue/train NUM_TRAJECTORIES=20 \
generation/run_pipeline.shAfter configuring model access, run a small sample:
python evaluation/evaluation_atod.py \
--max-samples 5The full evaluation makes multiple model calls per dialogue and may incur substantial cost. Start with --max-samples.
Offline tests (no model access needed):
python -m unittest discover -s testsATODEval/ implements the ATOD-Eval metrics defined in the paper: dGCR and NTC over decided goals (dgcr.py, ntc.py, no model calls), dependency-edge precision/recall/F1 (dependency.py), and the LLM-judged metrics with the prompt templates from the paper appendix: memory recall accuracy (memory_recall_accuracy.py, requires the evaluated system's per-turn goal states via --predictions-dir), proactivity effectiveness (proactivity_effectiveness.py, grounded × beneficial), turn-level relevance (turn_level_quality.py, score / 5) and dialogue-level coherence (dialogue_level_quality.py, native 1–5 scale).
python ATODEval/dgcr.py --complexity all
python ATODEval/proactivity_effectiveness.py --complexity medium --sample-size 5This code is being released solely for academic and scientific reproducibility purposes, in support of the methods and findings described in the associated publication. Pull requests are not being accepted in order to maintain the code exactly as it was used in the paper.
Unless otherwise noted, the code and ATOD data are released under the Creative Commons Attribution-NonCommercial 4.0 International license. Third-party materials retain their original licenses; see THIRD_PARTY_LICENSES.md.
If you use ATOD or ATOD-Eval, please cite:
@article{zhang2026atod,
title={ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems},
author={Zhang, Yifei and Nayyeri, Hooshang and Khaziev, Rinat and Yilmaz, Emine and Tur, Gokhan and Hakkani-T{\"u}r, Dilek and Thadakamalla, Hari P},
journal={arXiv preprint arXiv:2601.11854},
year={2026}
}The paper is available on arXiv. Citation metadata is also provided in CITATION.cff and will be updated when the final Anthology record becomes available.