Skip to content
tyhobbsPublic

About

Fine-tuning Llama 3 8B to convert athlete-monitoring data into structured load-management briefs — grounded in published sports-science rules.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

LoadBrief

What eight iterations of a rule-generated corpus reveal about label provenance and metric validity in domain fine-tuning.

LoadBrief converts a free-text athlete monitoring summary — training load, heart-rate variability, wellness scores — into a structured, audience-conditioned load-management brief for an athlete, coach, or sports scientist. The corpus is generated by a rule-based simulator; the model is a LoRA fine-tune of Llama 3 8B Instruct.

Built with Meta Llama 3. Model weights are subject to the Meta Llama 3 license.

This is a research artifact, not a medical device. The released corpus contains documented defects and the released model reproduces them. Neither should inform decisions about a real person's training or health. See Known defects below.


What this project turned out to be about

The original goal was a specialist model for athlete load management. It works: the final model reaches 0.960 exact risk-classification accuracy against a base model scoring 0.000.

Auditing that result is what the paper is actually about. Rule-generated data is not self-validating, and neither are the metrics used to evaluate models trained on it. "Correct by construction" guarantees only that a label agrees with the rule that produced it. It says nothing about whether that rule's output agrees with the data rendered alongside it, whether a declared generation parameter was reachable, or whether the metric scoring the result can recognize a correct answer when shown one. Each is a separate property needing a separate check, and in this project every one went unchecked through multiple training cycles.

Five findings, all measured against the released artifacts:

Both labels are per-scenario constants. risk_level and overreaching_classification are written from the scenario definition, not computed from the sampled signals. Across 40,000 records each of 19 scenarios maps to exactly one of each, with zero variance. A TF-IDF bag-of-words classifier recovers the risk label from the narrative at 0.950 — within one point of the 8B fine-tune — so the headline accuracy cannot be evidence of clinical inference.

16.1% of records contradict themselves, in two disjoint families: 2,412 pair critically suppressed HRV with a LOW risk header; 4,021 name one overreaching class in the signal-integration narrative and a different one in the classification section.

Three consistency guards pass every one of them, each for a different structural reason. Passing all three is not evidence of consistency.

The composite reward is mis-specified against its own reference data. Its 0.701 ceiling is an artifact of three of five components the ground-truth briefs cannot earn. An untuned model scores 0.422, leaving a usable range of 0.279.

The overreaching metric cannot read 45% of correct answers, and 0% of one class. What looked like a seven-revision capability loss was a generator change.

Repository structure

loadbrief/
├── README.md
├── LICENSE
├── requirements.txt
│
├── loadbrief_generator/      # the rule-based simulator
│   ├── config.py             #   thresholds: ACWR zones, HRV, wellness, registers
│   ├── athlete_generator.py
│   ├── simulator/            #   scenarios, time series, metrics, data levels
│   ├── narrative/            #   narrator + linguistic variation
│   ├── brief_generator/      #   rule engine, signal synthesis, audience adapter
│   ├── quality/              #   schema validator + quality filter
│   ├── label_agreement.py    #   severity-rank divergence filter
│   ├── conflict_rationale.py #   override-rationale phrase banks
│   └── dataset/              #   splits, HF export
│
├── training/
│   ├── m4_sft_training.py    #   SFT (all eight revisions, identical config)
│   └── m4_grpo_training.py   #   GRPO (v1 only; not used in the release)
│
├── evaluation/
│   ├── run_baselines.py      #   generate completions
│   ├── evaluate_all.py       #   rule metrics + reward + reference ceiling
│   ├── llm_judge.py          #   LLM-as-judge
│   ├── show_judge_strata.py  #   stratified analysis
│   └── format_dataset.py
│
├── audits/                   # the checks this paper argues for
│   ├── reachability_audit.py       #   declared targets vs realized distribution
│   ├── overreaching_selftest.py    #   can the metric read its own ground truth?
│   ├── consistency_check.py        #   phrase-bank contradiction detector
│   ├── calibrate_judge.py          #   judge scored on ground-truth briefs
│   ├── run_full_calibration.sh     #   all scenarios x registers, resumable
│   ├── summarize_calibration.py    #   bootstrap CIs + model-vs-data gap
│   └── verify_paper_tables.py      #   recompute every table in the paper
│
└── results/                  # per-revision metrics and judge ratings

Checkpoints and the full corpus are not committed; they are on Hugging Face, linked above.

Reproducing the paper's tables

Every table in the paper is recomputed from the released artifacts:

python3 audits/verify_paper_tables.py --root . --check-validator

It discovers the corpus, evaluation outputs and judge ratings wherever they sit, re-derives each published value, and prints PASS/FAIL per check. Exit code 0 if all pass.

The individual audits run standalone:

# do scenarios produce the metric ranges they declare?
python3 audits/reachability_audit.py --data dataset_v8/train.jsonl \
    --scenarios loadbrief_generator/simulator/scenarios.py

# can the overreaching metric read the corpus's own correct answers?
python3 audits/overreaching_selftest.py --root .

# does the judge score ground-truth briefs highly?
export GEMINI_API_KEY=...
./audits/run_full_calibration.sh gemini
python3 audits/summarize_calibration.py --calibration calibration

Known defects in the released corpus

Documented rather than repaired, so that the dataset and the released model remain consistent with each other. Full detail in the paper and the dataset card.

Defect Scope
Labels are per-scenario constants all records
Adverse HRV prose under a LOW header 2,412 (6.0%)
Signal-integration vs classification mismatch 4,021 (10.1%)
CRITICAL unreachable above ACWR 0.8 all records
Overreaching label unreadable by string extraction 45% of briefs
Athlete register carries less clinical content all records
recreational_minimal_data declared range unmet 327 records
early_overreaching label–signal divergence exempted, not repaired 1,716 records

If you train on this corpus, filter records whose output_coach contains both "critically suppressed" and "RISK LEVEL: LOW", and score overreaching against the label field directly rather than parsing it out of generated text.

Recommendations

Seven cheap checks, each of which would have caught a defect this project carried through multiple training cycles, none of them standard practice. In short: audit declared-target reachability; state label provenance explicitly; validate extraction metrics against reference outputs; decompose composite rewards against ground truth; calibrate model-based judges against ground truth; do not let a guard's scope be set by the data it inspects; and pair deterministic with model-based evaluation, treating neither as evidence about the other's blind spots. The paper develops each.

License

Code is MIT (see LICENSE). Model weights derive from Llama 3 and are governed by the Meta Llama 3 license. The dataset is released separately under CC BY 4.0.

About

Fine-tuning Llama 3 8B to convert athlete-monitoring data into structured load-management briefs — grounded in published sports-science rules.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages