Files
nucleic-purpose-classifier/README.md
T

384 lines
18 KiB
Markdown
Raw Normal View History

# Purpose classifier
This directory is the reproducible data and training pipeline for
`docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of
work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning,
temperature calibration, shared confidence thresholds, and the frozen-set accuracy,
recall, hard-slice, calibration, and latency report.
## Data contract
The canonical generated sources are listed in `data/generation-manifest.json`.
`round2-NN.jsonl` files are retained generation batches and intentionally duplicate
`purpose-prompts-round2.jsonl`; they are provenance, not additional training input.
`prepare_data.py`:
- validates the strict generated-record schema;
- removes exact and high-overlap word-trigram duplicates;
- fails for review if a high-overlap pair has conflicting labels;
- keeps shipped fixtures completely outside source data;
- holds every `vague-eval` record out of training;
- optionally applies a completed, versioned human-review ledger before splitting;
- stratifies by primary purpose, slice, and primary language; and
- verifies that the deterministic test partition still matches the versioned
`data/frozen-test-v1.jsonl`.
The frozen test set is the synthetic JSONL plus the 87 classifiable records in
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general`
records are excluded because `general` is deliberately not a model label. The exact
membership and hashes are locked in `data/dataset-v1-manifest.json`.
## Prepare
From the repository root:
```bash
python3 ml/purpose-classifier/validate-data.py
python3 ml/purpose-classifier/prepare_data.py
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'
```
For a newly generated raw 200-record batch, enable batch-shape checks explicitly with
`validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200`.
The generated train/validation copies land under `.artifacts/dataset-v1/` and are
gitignored. A source, curation, seed, or split-policy change that moves the frozen test
set fails closed. After reviewing such a change, intentionally version it with:
```bash
python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test
```
## Train purpose-lite
Use a dedicated virtual environment. The base model is pinned to a specific
`sentence-transformers/all-MiniLM-L6-v2` commit: a 6-layer, 384-dimensional encoder. The
training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export
can expose a fixed `1 x 128` runtime shape. Long prompts preserve both ends as
`[CLS]` + 63 head tokens + `[SEP]` + 62 tail tokens + `[SEP]`; this keeps the ask when it
follows a pasted log or stack trace while retaining enough leading context to interpret it.
```bash
python3 -m venv ml/purpose-classifier/.venv
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py
```
The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental
multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel
accelerator training, install the platform's `torch==2.13.0` build using PyTorch's
platform selector, then install `requirements-base.txt`.
Training writes a local checkpoint, `calibration.json`, and `metrics.json` under
`outputs/purpose-lite-v1/`. It selects checkpoints and fits temperature on label-scorable
validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every `vague-eval`
record counts as an abstention miss even if its synthetic label happens to match. The
incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot
silently replace it with a regression. Validation early stopping defaults to two epochs
without an improvement greater than 0.05 points.
Continuation training accepts a local checkpoint. `--boundary-weight` is an opt-in,
validation-selected loss weight for the measured weakest slice; it does not add held-out
fixtures to training:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1/model \
--epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \
--boundary-weight 2 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \
--overwrite-output
```
For QAT, `--quantization-aware` replaces the model's linear and embedding forwards with
straight-through fake quantization matching the shipping QDQ graph: per-tensor uint8
embeddings, per-channel symmetric int8 linear weights, and per-tensor uint8 activations.
Parameter names remain unchanged, so the selected checkpoint reopens as an ordinary
Transformers model and uses the same `export.py` path. Keep the incoming checkpoint as
epoch zero and select QAT only on validation. Training logs progress every 50 batches by
default (`--progress-steps 0` disables it), so a long CPU run remains observable:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--epochs 2 --learning-rate 1e-6 --warmup-ratio 0 \
--early-stopping-patience 1 --boundary-weight 2 --quantization-aware \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-qat1 \
--overwrite-output
```
On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at
94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable
PyTorch↔ONNX agreement. It was the pre-distillation quantized candidate and remained two correct
predictions below the 95% gate. A subsequent validation-selected `5e-7` epoch improved
int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%;
it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model
iteration should incorporate reviewed boundary data and be selected on a revised
validation/frozen dataset version.
To target only the remaining float→int8 decision drift, cache the float teacher in a
separate inference process and use its logits for QAT distillation. Keeping teacher and
student models out of the same process avoids doubling peak resident memory:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/cache_teacher.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--output ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--overwrite-output
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--distillation-cache \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--distillation-weight 0.9 --distillation-temperature 2 \
--distillation-selection-weight 0.5 --quantization-aware \
--epochs 2 --learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat \
--overwrite-output
```
The cache binds each logit row to normalized prompt hash plus expected label. Training
fails closed if either split changes. Selection combines label accuracy with float-teacher
agreement, retains the incoming checkpoint as epoch zero, and logs label/distillation loss
separately. A 64-record wiring run exercised cache loading, shuffled row alignment,
backpropagation, selection, and ordinary checkpoint reload. The current shared CPU runtime
then showed severe post-batch throttling, so no full candidate result is claimed from that
canary.
### Native Apple Silicon training with MLX
Use the MLX backend when training on Apple Silicon. It implements the same six-layer BERT
classifier, fixed head-tail tokenization, export-matched QAT graph, cached-teacher
distillation, validation selection, and early stopping with native MLX arrays. Fake
quantization is decomposed into Metal-supported round, clip, and straight-through-gradient
operations, avoiding PyTorch's unsupported MPS fake-quant operator. Selected weights are
written back with the original Hugging Face parameter names, so the existing PyTorch
`export.py` and `eval.py` paths remain unchanged.
Install the additional pinned dependency into the macOS virtual environment:
```bash
ml/purpose-classifier/venv/bin/python -m pip install \
-r ml/purpose-classifier/requirements-mlx.txt
```
Before the first full run on a new MLX or Transformers version, run the fail-closed parity
check. It requires exact fake-quant primitives, float-logit parity, matching QAT
predictions with bounded backend drift, healthy QAT gradients, and an exact Hugging Face
→ MLX → Hugging Face weight round trip:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/verify_mlx.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model
```
Then run the distilled QAT candidate natively on Metal:
```bash
ml/purpose-classifier/venv/bin/python -u ml/purpose-classifier/train_mlx.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--distillation-cache \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--distillation-weight 0.9 --distillation-temperature 2 \
--distillation-selection-weight 0.5 --quantization-aware \
--epochs 4 --early-stopping-patience 1 \
--learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
--progress-steps 1 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e \
--overwrite-output
```
The full Metal run stopped after epoch three and selected epoch two at 94.75% fake-quant
validation accuracy. Its 23,148,500-byte int8-QDQ export scores **95.20% frozen
(892/937)**, 95.19% macro recall, 94.17% scored-hard accuracy, and 98.08% scorable
PyTorch↔ONNX agreement. Every purpose recall is above 91%, and the vague-abstention and
routing-tier-drift gates pass. This is the current accuracy-qualified shipping candidate;
latency and energy/residency still require measurement on the target Apple and Windows
accelerator runtimes.
### Convert and validate Core ML
Core ML Tools no longer maintains the legacy ONNX converter, so the Apple artifact is
converted directly from the selected Hugging Face checkpoint. `convert_coreml.py` uses a
fixed-shape export-only BERT forward to avoid dynamic Transformers masking helpers, checks
that forward against Transformers before conversion, writes an ML Program package, and
records hashes for every package file.
Install the pinned converter in the macOS environment and create the package:
```bash
ml/purpose-classifier/venv/bin/python -m pip install \
-r ml/purpose-classifier/requirements-coreml.txt
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/convert_coreml.py \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--output \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--overwrite-output
```
The direct float16 package is the conversion baseline, not the accepted Apple artifact.
On the first physical Apple-Silicon run it scored 94.98% (890/937), two correct decisions
behind the accepted ONNX graph, with 97.97% scorable label agreement. Calibrate a
Core ML-native W8A8 candidate with the same deterministic 256-record sample and QDQ policy
as the ONNX exporter:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/quantize_coreml.py \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-fp16.mlpackage \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--output \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--overwrite-output
```
Activation calibration writes its temporary packages under the candidate output directory
and removes each package immediately after prediction; this avoids Core ML Tools retaining
one full weight copy per calibration step until process exit. It prints progress while it
runs. The candidate uses per-tensor asymmetric uint8 activations,
per-channel symmetric int8 linear weights, and per-tensor asymmetric uint8 embedding
weights. Activation quantization is limited to floating-point linear operations; applying
Core ML Tools' global policy also selects integer embedding-index additions and produces
an invalid quantize operation. It fails the command if the resulting package exceeds
25 MiB.
Run the frozen gate with CPU+Neural Engine placement and compare labels directly with the
accepted int8 ONNX artifact. Gated Core ML evaluation fails closed without
`--compare-onnx`, and requires at least 99.5% scorable label agreement:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/eval.py \
--model-dir \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/model \
--calibration \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/calibration.json \
--coreml-model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--coreml-compute-units cpu-and-ne \
--compare-onnx \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/export/purpose-lite-v1-int8-qdq.onnx \
--report \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/ane-frozen-eval.json
```
Record the compute plan separately; this reports both operation-count and estimated-cost
ANE shares. Repeat evaluation with `--coreml-compute-units cpu-only --no-gate` before the
energy comparison in `ENERGY_AND_RESIDENCY.md`:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/inspect_coreml.py \
--model \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/purpose-lite-v1-w8a8.mlpackage \
--compute-units cpu-and-ne \
--report \
ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx-4e/coreml/ane-compute-plan.json
```
For a wiring smoke test, use a small deterministic prefix:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--epochs 1 --max-train-records 64 --max-validation-records 64 \
--output-dir ml/purpose-classifier/outputs/smoke --overwrite-output
```
## Evaluate
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py
```
The command returns failure unless label-scorable frozen accuracy is at least 95%, every
purpose recall is at least 85%, at least 90% of the deliberately context-free
`vague-eval` slice resolves LOW, every misroute stays within one routing cost tier, and
measured batch-one p95 is at most 20 ms. Use `--no-gate` only for diagnostic runs.
## Export and score ONNX
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
--report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json
```
`export.py` emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/
normalization contract, golden tokenizations, shared calibration config, graph checks,
artifact hashes, and a size report. Its default 256-record quantization calibration sample
is deterministic and stratified by purpose, slice, and primary language; the export report
records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping
candidate; the fp16 graph remains the accelerator-oriented conversion input.
When scoring ONNX, add `--compare-pytorch` to measure artifact drift against
`--model-dir`. The report then includes overall, label-scorable, and per-slice label
agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label
transition:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
--compare-pytorch --no-gate
```
## Audit curation
Run the semantic embedding duplicate audit. It also emits the deterministic,
purpose/slice/language-stratified 10% human label-and-difficulty review CSV:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py
```
The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as
`purpose-lite`. Similarity only proposes review candidates; it never edits source data or
the frozen split automatically. The 18 word-trigram exclusions and the human-review
completion rule are recorded in `data/curation-review-v1.json`; the semantic report is
versioned as `data/semantic-audit-v1.json`.
## Optional human review
The dataset owner accepted the curated generated labels and difficulty metadata as-is on
2026-07-31, so the blank 1,219-row review sample is not a training or rollout blocker. It
remains available as an optional future audit. Check its progress without running the
embedding audit again:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py
```
Mark each row `accept`, `relabel`, or `reject`. `accept` and `reject` leave the four
`reviewed*` fields blank; `reject` requires notes. For `relabel`, blank reviewed fields
retain their generated value, `<none>` clears a secondary purpose, and notes are required.
If a secondary purpose is added or removed, set `reviewedSlice` consistently (`mixed`
when a secondary is present). The validator rejects stale generated columns, missing or
duplicate sample rows, invalid label combinations, and partially completed rows.
When every row has a human decision, write the versionable ledger:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py --finalize
```
Build an isolated candidate split first:
```bash
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/prepare_data.py \
--human-review ml/purpose-classifier/data/human-review-v1.json \
--output-dir ml/purpose-classifier/.artifacts/reviewed-candidate \
--frozen-test ml/purpose-classifier/.artifacts/reviewed-frozen-candidate.jsonl \
--manifest ml/purpose-classifier/.artifacts/reviewed-manifest-candidate.json \
--refresh-frozen-test
```
Inspect the ledger, decision summary, candidate manifest, and split diff. Only then rerun
the same command with the three candidate-path overrides removed to intentionally replace
the versioned frozen dataset and manifest.
`--regenerate` recreates a blank CSV in the current schema and is only appropriate before
review begins.
The one-time, hardware-bound energy and accelerator-residency procedure is in
`ENERGY_AND_RESIDENCY.md`.