Files
nucleic-purpose-classifier/README.md
T

8.6 KiB

Purpose classifier

This directory is the reproducible data and training pipeline for docs/PURPOSE_CLASSIFIER.md. The current slice covers work item 2 and the first part of work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning, temperature calibration, shared confidence thresholds, and the frozen-set accuracy, recall, hard-slice, calibration, and latency report.

Data contract

The canonical generated sources are listed in data/generation-manifest.json. round2-NN.jsonl files are retained generation batches and intentionally duplicate purpose-prompts-round2.jsonl; they are provenance, not additional training input.

prepare_data.py:

  • validates the strict generated-record schema;
  • removes exact and high-overlap word-trigram duplicates;
  • fails for review if a high-overlap pair has conflicting labels;
  • keeps shipped fixtures completely outside source data;
  • holds every vague-eval record out of training;
  • stratifies by primary purpose, slice, and primary language; and
  • verifies that the deterministic test partition still matches the versioned data/frozen-test-v1.jsonl.

The frozen test set is the synthetic JSONL plus the 87 classifiable records in Tests/NucleicCoreTests/Fixtures/purpose-prompts.json. The fixture file's five general records are excluded because general is deliberately not a model label. The exact membership and hashes are locked in data/dataset-v1-manifest.json.

Prepare

From the repository root:

python3 ml/purpose-classifier/validate-data.py
python3 ml/purpose-classifier/prepare_data.py
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'

For a newly generated raw 200-record batch, enable batch-shape checks explicitly with validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200.

The generated train/validation copies land under .artifacts/dataset-v1/ and are gitignored. A source, curation, seed, or split-policy change that moves the frozen test set fails closed. After reviewing such a change, intentionally version it with:

python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test

Train purpose-lite

Use a dedicated virtual environment. The base model is pinned to a specific sentence-transformers/all-MiniLM-L6-v2 commit: a 6-layer, 384-dimensional encoder. The training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export can expose a fixed 1 x 128 runtime shape. Long prompts preserve both ends as [CLS] + 63 head tokens + [SEP] + 62 tail tokens + [SEP]; this keeps the ask when it follows a pasted log or stack trace while retaining enough leading context to interpret it.

python3 -m venv ml/purpose-classifier/.venv
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py

The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel accelerator training, install the platform's torch==2.13.0 build using PyTorch's platform selector, then install requirements-base.txt.

Training writes a local checkpoint, calibration.json, and metrics.json under outputs/purpose-lite-v1/. It selects checkpoints and fits temperature on label-scorable validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every vague-eval record counts as an abstention miss even if its synthetic label happens to match. The incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot silently replace it with a regression. Validation early stopping defaults to two epochs without an improvement greater than 0.05 points.

Continuation training accepts a local checkpoint. --boundary-weight is an opt-in, validation-selected loss weight for the measured weakest slice; it does not add held-out fixtures to training:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --model ml/purpose-classifier/outputs/purpose-lite-v1/model \
  --epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \
  --boundary-weight 2 \
  --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \
  --overwrite-output

For QAT, --quantization-aware replaces the model's linear and embedding forwards with straight-through fake quantization matching the shipping QDQ graph: per-tensor uint8 embeddings, per-channel symmetric int8 linear weights, and per-tensor uint8 activations. Parameter names remain unchanged, so the selected checkpoint reopens as an ordinary Transformers model and uses the same export.py path. Keep the incoming checkpoint as epoch zero and select QAT only on validation. Training logs progress every 50 batches by default (--progress-steps 0 disables it), so a long CPU run remains observable:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
  --epochs 2 --learning-rate 1e-6 --warmup-ratio 0 \
  --early-stopping-patience 1 --boundary-weight 2 --quantization-aware \
  --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-qat1 \
  --overwrite-output

On dataset v1, that validation-selected run produced a 23,148,500-byte int8 graph at 94.88% frozen accuracy (889/937), 94.46% scored-hard accuracy, and 98.19% scorable PyTorch↔ONNX agreement. It is the current quantized candidate, but remains two correct predictions below the 95% gate. A subsequent validation-selected 5e-7 epoch improved int8 validation accuracy from 93.31% to 93.71% but regressed frozen accuracy to 94.34%; it is rejected. Do not continue optimizer-only QAT sweeps on this split. The next model iteration should incorporate reviewed boundary data and be selected on a revised validation/frozen dataset version.

For a wiring smoke test, use a small deterministic prefix:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --epochs 1 --max-train-records 64 --max-validation-records 64 \
  --output-dir ml/purpose-classifier/outputs/smoke --overwrite-output

Evaluate

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py

The command returns failure unless label-scorable frozen accuracy is at least 95%, every purpose recall is at least 85%, at least 90% of the deliberately context-free vague-eval slice resolves LOW, every misroute stays within one routing cost tier, and measured batch-one p95 is at most 20 ms. Use --no-gate only for diagnostic runs.

Export and score ONNX

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
  --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
  --report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json

export.py emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/ normalization contract, golden tokenizations, shared calibration config, graph checks, artifact hashes, and a size report. Its default 256-record quantization calibration sample is deterministic and stratified by purpose, slice, and primary language; the export report records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping candidate; the fp16 graph remains the accelerator-oriented conversion input.

When scoring ONNX, add --compare-pytorch to measure artifact drift against --model-dir. The report then includes overall, label-scorable, and per-slice label agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label transition:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
  --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
  --compare-pytorch --no-gate

Audit curation

Run the semantic embedding duplicate audit. It also emits the deterministic, purpose/slice/language-stratified 10% human label-and-difficulty review CSV:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py

The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as purpose-lite. Similarity only proposes review candidates; it never edits source data or the frozen split automatically. The 18 word-trigram exclusions and the human-review completion rule are recorded in data/curation-review-v1.json; the semantic report is versioned as data/semantic-audit-v1.json.

The one-time, hardware-bound energy and accelerator-residency procedure is in ENERGY_AND_RESIDENCY.md.