2026-07-30 03:48:00 -07:00
|
|
|
# Purpose classifier
|
|
|
|
|
|
|
|
|
|
This directory is the reproducible data and training pipeline for
|
|
|
|
|
`docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of
|
|
|
|
|
work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning,
|
|
|
|
|
temperature calibration, shared confidence thresholds, and the frozen-set accuracy,
|
|
|
|
|
recall, hard-slice, calibration, and latency report.
|
|
|
|
|
|
|
|
|
|
## Data contract
|
|
|
|
|
|
|
|
|
|
The canonical generated sources are listed in `data/generation-manifest.json`.
|
|
|
|
|
`round2-NN.jsonl` files are retained generation batches and intentionally duplicate
|
|
|
|
|
`purpose-prompts-round2.jsonl`; they are provenance, not additional training input.
|
|
|
|
|
|
|
|
|
|
`prepare_data.py`:
|
|
|
|
|
|
|
|
|
|
- validates the strict generated-record schema;
|
|
|
|
|
- removes exact and high-overlap word-trigram duplicates;
|
|
|
|
|
- fails for review if a high-overlap pair has conflicting labels;
|
|
|
|
|
- keeps shipped fixtures completely outside source data;
|
|
|
|
|
- holds every `vague-eval` record out of training;
|
|
|
|
|
- stratifies by primary purpose, slice, and primary language; and
|
|
|
|
|
- verifies that the deterministic test partition still matches the versioned
|
|
|
|
|
`data/frozen-test-v1.jsonl`.
|
|
|
|
|
|
|
|
|
|
The frozen test set is the synthetic JSONL plus the 87 classifiable records in
|
|
|
|
|
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general`
|
|
|
|
|
records are excluded because `general` is deliberately not a model label. The exact
|
|
|
|
|
membership and hashes are locked in `data/dataset-v1-manifest.json`.
|
|
|
|
|
|
|
|
|
|
## Prepare
|
|
|
|
|
|
|
|
|
|
From the repository root:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
python3 ml/purpose-classifier/validate-data.py
|
|
|
|
|
python3 ml/purpose-classifier/prepare_data.py
|
|
|
|
|
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
For a newly generated raw 200-record batch, enable batch-shape checks explicitly with
|
|
|
|
|
`validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200`.
|
|
|
|
|
|
|
|
|
|
The generated train/validation copies land under `.artifacts/dataset-v1/` and are
|
|
|
|
|
gitignored. A source, curation, seed, or split-policy change that moves the frozen test
|
|
|
|
|
set fails closed. After reviewing such a change, intentionally version it with:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
## Train purpose-lite
|
|
|
|
|
|
|
|
|
|
Use a dedicated virtual environment. The base model is pinned to a specific
|
|
|
|
|
`sentence-transformers/all-MiniLM-L6-v2` commit: a 6-layer, 384-dimensional encoder. The
|
|
|
|
|
training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export
|
2026-07-30 05:04:34 -07:00
|
|
|
can expose a fixed `1 x 128` runtime shape. Long prompts preserve both ends as
|
|
|
|
|
`[CLS]` + 63 head tokens + `[SEP]` + 62 tail tokens + `[SEP]`; this keeps the ask when it
|
|
|
|
|
follows a pasted log or stack trace while retaining enough leading context to interpret it.
|
2026-07-30 03:48:00 -07:00
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
python3 -m venv ml/purpose-classifier/.venv
|
|
|
|
|
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental
|
|
|
|
|
multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel
|
|
|
|
|
accelerator training, install the platform's `torch==2.13.0` build using PyTorch's
|
|
|
|
|
platform selector, then install `requirements-base.txt`.
|
|
|
|
|
|
|
|
|
|
Training writes a local checkpoint, `calibration.json`, and `metrics.json` under
|
2026-07-30 05:04:34 -07:00
|
|
|
`outputs/purpose-lite-v1/`. It selects checkpoints and fits temperature on label-scorable
|
|
|
|
|
validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every `vague-eval`
|
2026-07-30 17:28:44 -07:00
|
|
|
record counts as an abstention miss even if its synthetic label happens to match. The
|
|
|
|
|
incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot
|
|
|
|
|
silently replace it with a regression. Validation early stopping defaults to two epochs
|
|
|
|
|
without an improvement greater than 0.05 points.
|
|
|
|
|
|
|
|
|
|
Continuation training accepts a local checkpoint. `--boundary-weight` is an opt-in,
|
|
|
|
|
validation-selected loss weight for the measured weakest slice; it does not add held-out
|
|
|
|
|
fixtures to training:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
|
|
|
|
|
--model ml/purpose-classifier/outputs/purpose-lite-v1/model \
|
|
|
|
|
--epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \
|
|
|
|
|
--boundary-weight 2 \
|
|
|
|
|
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \
|
|
|
|
|
--overwrite-output
|
|
|
|
|
```
|
2026-07-30 03:48:00 -07:00
|
|
|
|
|
|
|
|
For a wiring smoke test, use a small deterministic prefix:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
|
|
|
|
|
--epochs 1 --max-train-records 64 --max-validation-records 64 \
|
|
|
|
|
--output-dir ml/purpose-classifier/outputs/smoke --overwrite-output
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
## Evaluate
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py
|
|
|
|
|
```
|
|
|
|
|
|
2026-07-30 05:04:34 -07:00
|
|
|
The command returns failure unless label-scorable frozen accuracy is at least 95%, every
|
|
|
|
|
purpose recall is at least 85%, at least 90% of the deliberately context-free
|
|
|
|
|
`vague-eval` slice resolves LOW, every misroute stays within one routing cost tier, and
|
|
|
|
|
measured batch-one p95 is at most 20 ms. Use `--no-gate` only for diagnostic runs.
|
|
|
|
|
|
|
|
|
|
## Export and score ONNX
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
|
|
|
|
|
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
|
|
|
|
|
--report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
`export.py` emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/
|
|
|
|
|
normalization contract, golden tokenizations, shared calibration config, graph checks,
|
2026-07-30 17:28:44 -07:00
|
|
|
artifact hashes, and a size report. Its default 256-record quantization calibration sample
|
|
|
|
|
is deterministic and stratified by purpose, slice, and primary language; the export report
|
|
|
|
|
records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping
|
|
|
|
|
candidate; the fp16 graph remains the accelerator-oriented conversion input.
|
|
|
|
|
|
|
|
|
|
When scoring ONNX, add `--compare-pytorch` to measure artifact drift against
|
|
|
|
|
`--model-dir`. The report then includes overall, label-scorable, and per-slice label
|
|
|
|
|
agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label
|
|
|
|
|
transition:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
|
|
|
|
|
--onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
|
|
|
|
|
--compare-pytorch --no-gate
|
|
|
|
|
```
|
2026-07-30 05:04:34 -07:00
|
|
|
|
|
|
|
|
## Audit curation
|
|
|
|
|
|
|
|
|
|
Run the semantic embedding duplicate audit. It also emits the deterministic,
|
|
|
|
|
purpose/slice/language-stratified 10% human label-and-difficulty review CSV:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as
|
|
|
|
|
`purpose-lite`. Similarity only proposes review candidates; it never edits source data or
|
|
|
|
|
the frozen split automatically. The 18 word-trigram exclusions and the human-review
|
|
|
|
|
completion rule are recorded in `data/curation-review-v1.json`; the semantic report is
|
|
|
|
|
versioned as `data/semantic-audit-v1.json`.
|
|
|
|
|
|
|
|
|
|
The one-time, hardware-bound energy and accelerator-residency procedure is in
|
|
|
|
|
`ENERGY_AND_RESIDENCY.md`.
|