Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -54,7 +54,9 @@ python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test
|
||||
Use a dedicated virtual environment. The base model is pinned to a specific
|
||||
`sentence-transformers/all-MiniLM-L6-v2` commit: a 6-layer, 384-dimensional encoder. The
|
||||
training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export
|
||||
can expose a fixed `1 x 128` runtime shape.
|
||||
can expose a fixed `1 x 128` runtime shape. Long prompts preserve both ends as
|
||||
`[CLS]` + 63 head tokens + `[SEP]` + 62 tail tokens + `[SEP]`; this keeps the ask when it
|
||||
follows a pasted log or stack trace while retaining enough leading context to interpret it.
|
||||
|
||||
```bash
|
||||
python3 -m venv ml/purpose-classifier/.venv
|
||||
@@ -68,8 +70,9 @@ accelerator training, install the platform's `torch==2.13.0` build using PyTorch
|
||||
platform selector, then install `requirements-base.txt`.
|
||||
|
||||
Training writes a local checkpoint, `calibration.json`, and `metrics.json` under
|
||||
`outputs/purpose-lite-v1/`. It fits one validation-only temperature and derives nested
|
||||
HIGH/MEDIUM/LOW cutoffs from calibrated top-one probability plus top-two margin.
|
||||
`outputs/purpose-lite-v1/`. It selects checkpoints and fits temperature on label-scorable
|
||||
validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every `vague-eval`
|
||||
record counts as an abstention miss even if its synthetic label happens to match.
|
||||
|
||||
For a wiring smoke test, use a small deterministic prefix:
|
||||
|
||||
@@ -85,10 +88,39 @@ ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py
|
||||
```
|
||||
|
||||
The command returns failure unless frozen accuracy is at least 95%, every purpose recall
|
||||
is at least 85%, and measured batch-one p95 is at most 20 ms. Use `--no-gate` only for
|
||||
diagnostic runs. Accelerator residency, ONNX export/quantization, tokenizer golden tests,
|
||||
tier-drift evaluation, and Core ML parity remain follow-on work. Before calling dataset
|
||||
work item 2 complete, also run a semantic embedding duplicate audit and record the planned
|
||||
10% human label spot-check; the current dependency-free word-trigram pass is deliberately
|
||||
conservative.
|
||||
The command returns failure unless label-scorable frozen accuracy is at least 95%, every
|
||||
purpose recall is at least 85%, at least 90% of the deliberately context-free
|
||||
`vague-eval` slice resolves LOW, every misroute stays within one routing cost tier, and
|
||||
measured batch-one p95 is at most 20 ms. Use `--no-gate` only for diagnostic runs.
|
||||
|
||||
## Export and score ONNX
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
|
||||
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
|
||||
--report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json
|
||||
```
|
||||
|
||||
`export.py` emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/
|
||||
normalization contract, golden tokenizations, shared calibration config, graph checks,
|
||||
artifact hashes, and a size report. The int8 graph is the ≤25 MiB shipping candidate;
|
||||
the fp16 graph remains the accelerator-oriented conversion input.
|
||||
|
||||
## Audit curation
|
||||
|
||||
Run the semantic embedding duplicate audit. It also emits the deterministic,
|
||||
purpose/slice/language-stratified 10% human label-and-difficulty review CSV:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py
|
||||
```
|
||||
|
||||
The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as
|
||||
`purpose-lite`. Similarity only proposes review candidates; it never edits source data or
|
||||
the frozen split automatically. The 18 word-trigram exclusions and the human-review
|
||||
completion rule are recorded in `data/curation-review-v1.json`; the semantic report is
|
||||
versioned as `data/semantic-audit-v1.json`.
|
||||
|
||||
The one-time, hardware-bound energy and accelerator-residency procedure is in
|
||||
`ENERGY_AND_RESIDENCY.md`.
|
||||
|
||||
Reference in New Issue
Block a user