Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AEGIS

Awareness-Enhanced Guidance for Iterative Safeguard

arXiv

AEGIS is an exploratory framework for studying span-guided multilingual text detoxification across English, Mandarin Chinese, and Korean. It separates a span-level detector from frozen generator backbones so that the effect of harmful-span, intensity, and target guidance can be examined without treating the framework as a state-of-the-art claim.

Resource Status
Paper arXiv:2607.17713
Code Detector training and guided-generation pipeline available
Data Use the official upstream datasets described in DATA.md
License MIT for code; upstream terms apply to data and models

Research question

When does explicit span-level guidance improve detoxification, and when does it change the trade-off between toxicity reduction and meaning preservation?

Components

Component Purpose
aegis/datasets/ Language-specific dataset processing and BIO supervision
aegis/training/ XLM-R detector training for English, Chinese, and Korean
aegis/generation/ Guided and unguided rewriting with frozen generators
results/ Compact detector training histories

Setup

git clone https://github.com/cosmic4dev/AEGIS.git
cd AEGIS
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Obtain the datasets described in DATA.md, then adapt their local paths through the loader arguments.

Usage

Train a detector:

python -m aegis.training.english_xlmr_train --epochs 8 --patience 3
python -m aegis.training.chinese_xlmr_train --epochs 8 --patience 3
python -m aegis.training.korean_xlmr_train --epochs 8 --patience 3

Generate matched guided and unguided rewrites:

python -m aegis.generation.generate \
  --input_csv /path/to/evaluation_input.csv \
  --output_csv outputs/rewrites.csv \
  --generator_model_name Qwen/Qwen3-8B

The input CSV requires sample_id, original_text, toxicity_strength, and harmful_span_texts columns.

Evidence boundary

The repository supports inspection of the framework and reproduction with properly obtained datasets. Span guidance is treated as a conditional control signal: its benefit may vary with language, generator, and evaluation metric. Raw datasets, generated-text pools, checkpoints, human-evaluation records, and manuscript or submission sources are intentionally excluded.

Related project

The focused guided-versus-unguided analysis is maintained separately in span-guided-detoxification.

License

Code is released under the MIT License. External datasets and models retain their original licenses.

Citation

@article{park2026aegis,
  title   = {{AEGIS}: Awareness-Enhanced Guidance for Iterative Safeguard},
  author  = {Park, Kyungwon and Lee, Sangmin and Chon, Heejae and Kang, Hyungu},
  journal = {arXiv preprint arXiv:2607.17713},
  year    = {2026},
  doi     = {10.48550/arXiv.2607.17713},
  url     = {https://arxiv.org/abs/2607.17713}
}

About

Exploratory framework for span-guided multilingual text detoxification across English, Chinese, and Korean.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages