Skip to content

Dataset Expansion and Training & Loss Formulation Hangul Pos Tokenizer - #887

Open
kahye wants to merge 2 commits into
ReaLLMASIC:masterfrom
kahye:pos-tokenizer-more-data
Open

Dataset Expansion and Training & Loss Formulation Hangul Pos Tokenizer#887
kahye wants to merge 2 commits into
ReaLLMASIC:masterfrom
kahye:pos-tokenizer-more-data

Conversation

@kahye

@kahye kahye commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator
  1. Dataset Expansion:
    • get_dataset.sh: Expanded corpus generation by integrating full OPUS-100 en-ko with all KLUE task splits (dp, ner, mrc, nli, re, sts,
    ynat).
    • lane_metadata.json: Updated total character count to 21,065,854.
  2. Training & Loss Formulation:
    • train_args.py: Added --pos_loss_weight CLI parameter.
    • train.py: Supported custom POS loss weighting (wₚₒₛ) alongside structural loss down-weighting (
w
 struct

), and added milestone epoch checkpoint savings (iterations 3472, 5786, 8100, 11572).

  1. Evaluation & Benchmark Suite:
    • run_four_capability_evals.py: Comprehensive 4-capability benchmark evaluation covering:
    1. Token-Level Extraction (KLUE-NER)
    2. Syntactic Parsing (KLUE-DP)
    3. Informal & Noisy Text Resilience (NSMC & UnSmile @ 80% noise)
    4. Rare Vocabulary / OOV Generalization (KorMedMCQA)
    • run_phonetic_slang_eval.py & run_vocab_tail_perplexity.py: Added --mc_ckpt and --base_ckpt explicit checkpoint path override flags.
  2. Experiment & Demo Runners:
    • run_option1_sweep.py: Automated training sweep and 4-capability evaluation across loss weights
w       ∈ [0.05,0.08,0.10]
 struct

.

• run_opt1_10ep_experiment.py: 10-epoch training and evaluation progression runner.
• run_all_epoch_evals.sh: Script to run evaluations across 3, 5, and 10 epoch checkpoints.

kahye added 2 commits August 11, 2026 19:18
…ility evaluation suite

- Expand Korean POS dataset in get_dataset.sh with OPUS-100 and KLUE task splits (DP, NER, MRC, NLI, RE, STS, YNAT) and update lane_metadata.json
- Add --pos_loss_weight argument in train_args.py and handle POS loss weighting and milestone checkpoint saving in train.py
- Add --mc_ckpt and --base_ckpt path override support to benchmarks/run_phonetic_slang_eval.py and benchmarks/run_vocab_tail_perplexity.py
- Add 4-capability evaluation benchmark suite (benchmarks/run_four_capability_evals.py) covering KLUE-NER, KLUE-DP, noisy text resilience (NSMC/UnSmile), and rare vocabulary/OOV (KorMedMCQA)
- Add evaluation runner demos/run_all_epoch_evals.sh, Option 1 sweep runner run_option1_sweep.py, and 10-epoch experiment runner run_opt1_10ep_experiment.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant