Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -326,6 +326,30 @@ expected to be extremely slow there. Each improved epoch atomically rewrites `mo
|
||||
and updates `training-state.json`, so progress is visible and an interrupted run retains
|
||||
the last selected checkpoint.
|
||||
|
||||
To continue a completed run without discarding its trained task heads, pass its selected
|
||||
`model/` directory through `--resume-from` and write to a new output directory.
|
||||
Continuation restores the backbone and all four heads strictly, then starts a fresh
|
||||
optimizer and learning-rate schedule; `--model` remains reserved for an untrained local
|
||||
upstream checkpoint. The first base run was still improving when its three-epoch schedule
|
||||
ended, so continue its selected checkpoint conservatively before changing architecture:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/venv/bin/python -u \
|
||||
ml/purpose-classifier/train_deep_mlx.py \
|
||||
--variant base \
|
||||
--resume-from \
|
||||
ml/purpose-classifier/outputs/purpose-deep-v1-base-mlx/model \
|
||||
--dataset-dir \
|
||||
ml/purpose-classifier/.artifacts/dataset-v1-history-first-prompts \
|
||||
--epochs 3 \
|
||||
--learning-rate 1e-5 \
|
||||
--early-stopping-patience 2 \
|
||||
--progress-steps 10 \
|
||||
--output-dir \
|
||||
ml/purpose-classifier/outputs/purpose-deep-v1-base-mlx-cont-3e \
|
||||
--overwrite-output
|
||||
```
|
||||
|
||||
Do not launch the large rung yet. It is justified only after base is evaluated on the
|
||||
frozen set; large must beat base by at least two hard-slice points, while deep itself must
|
||||
reach 97% scored overall and beat the shipping lite artifact by five hard-slice points.
|
||||
|
||||
Reference in New Issue
Block a user