What eight iterations of a rule-generated corpus reveal about label provenance and metric validity in domain fine-tuning.
LoadBrief converts a free-text athlete monitoring summary — training load, heart-rate variability, wellness scores — into a structured, audience-conditioned load-management brief for an athlete, coach, or sports scientist. The corpus is generated by a rule-based simulator; the model is a LoRA fine-tune of Llama 3 8B Instruct.
- Paper: Loadbrief.pdf
- Model: tyhob/loadbrief
- Dataset: tyhob/loadbrief-50k
Built with Meta Llama 3. Model weights are subject to the Meta Llama 3 license.
This is a research artifact, not a medical device. The released corpus contains documented defects and the released model reproduces them. Neither should inform decisions about a real person's training or health. See Known defects below.
The original goal was a specialist model for athlete load management. It works: the final model reaches 0.960 exact risk-classification accuracy against a base model scoring 0.000.
Auditing that result is what the paper is actually about. Rule-generated data is not self-validating, and neither are the metrics used to evaluate models trained on it. "Correct by construction" guarantees only that a label agrees with the rule that produced it. It says nothing about whether that rule's output agrees with the data rendered alongside it, whether a declared generation parameter was reachable, or whether the metric scoring the result can recognize a correct answer when shown one. Each is a separate property needing a separate check, and in this project every one went unchecked through multiple training cycles.
Five findings, all measured against the released artifacts:
Both labels are per-scenario constants. risk_level and
overreaching_classification are written from the scenario definition, not
computed from the sampled signals. Across 40,000 records each of 19 scenarios
maps to exactly one of each, with zero variance. A TF-IDF bag-of-words classifier
recovers the risk label from the narrative at 0.950 — within one point of the 8B
fine-tune — so the headline accuracy cannot be evidence of clinical inference.
16.1% of records contradict themselves, in two disjoint families: 2,412 pair
critically suppressed HRV with a LOW risk header; 4,021 name one overreaching
class in the signal-integration narrative and a different one in the
classification section.
Three consistency guards pass every one of them, each for a different structural reason. Passing all three is not evidence of consistency.
The composite reward is mis-specified against its own reference data. Its 0.701 ceiling is an artifact of three of five components the ground-truth briefs cannot earn. An untuned model scores 0.422, leaving a usable range of 0.279.
The overreaching metric cannot read 45% of correct answers, and 0% of one class. What looked like a seven-revision capability loss was a generator change.
loadbrief/
├── README.md
├── LICENSE
├── requirements.txt
│
├── loadbrief_generator/ # the rule-based simulator
│ ├── config.py # thresholds: ACWR zones, HRV, wellness, registers
│ ├── athlete_generator.py
│ ├── simulator/ # scenarios, time series, metrics, data levels
│ ├── narrative/ # narrator + linguistic variation
│ ├── brief_generator/ # rule engine, signal synthesis, audience adapter
│ ├── quality/ # schema validator + quality filter
│ ├── label_agreement.py # severity-rank divergence filter
│ ├── conflict_rationale.py # override-rationale phrase banks
│ └── dataset/ # splits, HF export
│
├── training/
│ ├── m4_sft_training.py # SFT (all eight revisions, identical config)
│ └── m4_grpo_training.py # GRPO (v1 only; not used in the release)
│
├── evaluation/
│ ├── run_baselines.py # generate completions
│ ├── evaluate_all.py # rule metrics + reward + reference ceiling
│ ├── llm_judge.py # LLM-as-judge
│ ├── show_judge_strata.py # stratified analysis
│ └── format_dataset.py
│
├── audits/ # the checks this paper argues for
│ ├── reachability_audit.py # declared targets vs realized distribution
│ ├── overreaching_selftest.py # can the metric read its own ground truth?
│ ├── consistency_check.py # phrase-bank contradiction detector
│ ├── calibrate_judge.py # judge scored on ground-truth briefs
│ ├── run_full_calibration.sh # all scenarios x registers, resumable
│ ├── summarize_calibration.py # bootstrap CIs + model-vs-data gap
│ └── verify_paper_tables.py # recompute every table in the paper
│
└── results/ # per-revision metrics and judge ratings
Checkpoints and the full corpus are not committed; they are on Hugging Face, linked above.
Every table in the paper is recomputed from the released artifacts:
python3 audits/verify_paper_tables.py --root . --check-validatorIt discovers the corpus, evaluation outputs and judge ratings wherever they sit, re-derives each published value, and prints PASS/FAIL per check. Exit code 0 if all pass.
The individual audits run standalone:
# do scenarios produce the metric ranges they declare?
python3 audits/reachability_audit.py --data dataset_v8/train.jsonl \
--scenarios loadbrief_generator/simulator/scenarios.py
# can the overreaching metric read the corpus's own correct answers?
python3 audits/overreaching_selftest.py --root .
# does the judge score ground-truth briefs highly?
export GEMINI_API_KEY=...
./audits/run_full_calibration.sh gemini
python3 audits/summarize_calibration.py --calibration calibrationDocumented rather than repaired, so that the dataset and the released model remain consistent with each other. Full detail in the paper and the dataset card.
| Defect | Scope |
|---|---|
| Labels are per-scenario constants | all records |
Adverse HRV prose under a LOW header |
2,412 (6.0%) |
| Signal-integration vs classification mismatch | 4,021 (10.1%) |
CRITICAL unreachable above ACWR 0.8 |
all records |
| Overreaching label unreadable by string extraction | 45% of briefs |
| Athlete register carries less clinical content | all records |
recreational_minimal_data declared range unmet |
327 records |
early_overreaching label–signal divergence exempted, not repaired |
1,716 records |
If you train on this corpus, filter records whose output_coach contains both
"critically suppressed" and "RISK LEVEL: LOW", and score overreaching against
the label field directly rather than parsing it out of generated text.
Seven cheap checks, each of which would have caught a defect this project carried through multiple training cycles, none of them standard practice. In short: audit declared-target reachability; state label provenance explicitly; validate extraction metrics against reference outputs; decompose composite rewards against ground truth; calibrate model-based judges against ground truth; do not let a guard's scope be set by the data it inspects; and pair deterministic with model-based evaluation, treating neither as evidence about the other's blind spots. The paper develops each.
Code is MIT (see LICENSE). Model weights derive from Llama 3 and are governed
by the Meta Llama 3 license. The dataset is released separately under CC BY 4.0.