Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 17:28:44 -07:00
parent 93e9c838bb
commit e4403b570d
7 changed files with 408 additions and 19 deletions
+32 -3
View File
@@ -72,7 +72,23 @@ platform selector, then install `requirements-base.txt`.
Training writes a local checkpoint, `calibration.json`, and `metrics.json` under
`outputs/purpose-lite-v1/`. It selects checkpoints and fits temperature on label-scorable
validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every `vague-eval`
record counts as an abstention miss even if its synthetic label happens to match.
record counts as an abstention miss even if its synthetic label happens to match. The
incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot
silently replace it with a regression. Validation early stopping defaults to two epochs
without an improvement greater than 0.05 points.
Continuation training accepts a local checkpoint. `--boundary-weight` is an opt-in,
validation-selected loss weight for the measured weakest slice; it does not add held-out
fixtures to training:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1/model \
--epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \
--boundary-weight 2 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \
--overwrite-output
```
For a wiring smoke test, use a small deterministic prefix:
@@ -104,8 +120,21 @@ ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
`export.py` emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/
normalization contract, golden tokenizations, shared calibration config, graph checks,
artifact hashes, and a size report. The int8 graph is the ≤25 MiB shipping candidate;
the fp16 graph remains the accelerator-oriented conversion input.
artifact hashes, and a size report. Its default 256-record quantization calibration sample
is deterministic and stratified by purpose, slice, and primary language; the export report
records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping
candidate; the fp16 graph remains the accelerator-oriented conversion input.
When scoring ONNX, add `--compare-pytorch` to measure artifact drift against
`--model-dir`. The report then includes overall, label-scorable, and per-slice label
agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label
transition:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
--compare-pytorch --no-gate
```
## Audit curation