Repository navigation
healthcheck: fine-tune with a real warmup and 3 epochs - #4
Merged
Merged
Conversation
run_healthcheck() called CrossEncoder.fit without warmup_steps. Its default of 10000 kept a typical run (a few hundred steps) below a few percent of its learning rate, so the fine-tuned model was a near monotone shift of the base model and auc_tuned came back equal to auc_baseline. Warmup is now 10% of total steps and the default is 3 epochs, matching the hosted service since verifier-core 0.2.0 (cacheverifier-service #67). Offline on a 500-row later holdout: LmArena AUC 0.903 -> 0.932, AmazonHelp 0.551 -> 0.700. Runs take about 3x as long. 0.3.1.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
run_healthcheck()callsCrossEncoder.fit(train_dataloader, epochs=1)withoutwarmup_steps. sentence-transformers 3.x defaults toWarmupLinearwithwarmup_steps=10000. A typical Health Check is a few hundred steps, so the learning rate never gets past a few percent of 2e-5. The "fine-tuned" model ends up as a near-monotone shift of the base model, andauc_tunedcomes back equal toauc_baseline.The hosted service had the same bug, fixed in cacheverifier-service #67 (verifier-core 0.2.0). A local repro there gave a Spearman of 0.9997 between base and tuned scores on held-out rows.
Fix
TRAIN_WARMUP_FRACTION = 0.1: warmup is 10% of total steps, passed explicitly tofit.TRAIN_EPOCHS = 3: new default forrun_healthcheck()and--epochs, matching the hosted service.Evidence
Offline sweep, same chronological split as the hosted service, evaluated on a 500-row later holdout:
After the hosted fix was deployed, the production rerun gave 0.930 / 0.671 on the same holdout.
Tests
test_run_healthcheck_warms_up_over_a_tenth_of_its_steps: spies onCrossEncoder.fitand checksepochs == 3andwarmup_steps == 10%of total steps. It needs the healthcheck extra, like the existing end-to-end test.pytest tests/test_healthcheck.py: 8 passed (1 skipped by design with torch installed).ruff checkis clean.Note: open PR #3 (
--base-model) touches the same file and also bumps the version. It will need a rebase after this.