Files
nucleic-purpose-classifier/README.md
T

7.0 KiB

Purpose classifier

This directory is the reproducible data and training pipeline for docs/PURPOSE_CLASSIFIER.md. The current slice covers work item 2 and the first part of work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning, temperature calibration, shared confidence thresholds, and the frozen-set accuracy, recall, hard-slice, calibration, and latency report.

Data contract

The canonical generated sources are listed in data/generation-manifest.json. round2-NN.jsonl files are retained generation batches and intentionally duplicate purpose-prompts-round2.jsonl; they are provenance, not additional training input.

prepare_data.py:

  • validates the strict generated-record schema;
  • removes exact and high-overlap word-trigram duplicates;
  • fails for review if a high-overlap pair has conflicting labels;
  • keeps shipped fixtures completely outside source data;
  • holds every vague-eval record out of training;
  • stratifies by primary purpose, slice, and primary language; and
  • verifies that the deterministic test partition still matches the versioned data/frozen-test-v1.jsonl.

The frozen test set is the synthetic JSONL plus the 87 classifiable records in Tests/NucleicCoreTests/Fixtures/purpose-prompts.json. The fixture file's five general records are excluded because general is deliberately not a model label. The exact membership and hashes are locked in data/dataset-v1-manifest.json.

Prepare

From the repository root:

python3 ml/purpose-classifier/validate-data.py
python3 ml/purpose-classifier/prepare_data.py
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'

For a newly generated raw 200-record batch, enable batch-shape checks explicitly with validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200.

The generated train/validation copies land under .artifacts/dataset-v1/ and are gitignored. A source, curation, seed, or split-policy change that moves the frozen test set fails closed. After reviewing such a change, intentionally version it with:

python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test

Train purpose-lite

Use a dedicated virtual environment. The base model is pinned to a specific sentence-transformers/all-MiniLM-L6-v2 commit: a 6-layer, 384-dimensional encoder. The training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export can expose a fixed 1 x 128 runtime shape. Long prompts preserve both ends as [CLS] + 63 head tokens + [SEP] + 62 tail tokens + [SEP]; this keeps the ask when it follows a pasted log or stack trace while retaining enough leading context to interpret it.

python3 -m venv ml/purpose-classifier/.venv
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py

The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel accelerator training, install the platform's torch==2.13.0 build using PyTorch's platform selector, then install requirements-base.txt.

Training writes a local checkpoint, calibration.json, and metrics.json under outputs/purpose-lite-v1/. It selects checkpoints and fits temperature on label-scorable validation records. When deriving nested HIGH/MEDIUM/LOW cutoffs, every vague-eval record counts as an abstention miss even if its synthetic label happens to match. The incoming checkpoint is scored and retained as epoch zero, so a continuation run cannot silently replace it with a regression. Validation early stopping defaults to two epochs without an improvement greater than 0.05 points.

Continuation training accepts a local checkpoint. --boundary-weight is an opt-in, validation-selected loss weight for the measured weakest slice; it does not add held-out fixtures to training:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --model ml/purpose-classifier/outputs/purpose-lite-v1/model \
  --epochs 3 --learning-rate 3e-6 --warmup-ratio 0 \
  --boundary-weight 2 \
  --output-dir ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune \
  --overwrite-output

For a wiring smoke test, use a small deterministic prefix:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --epochs 1 --max-train-records 64 --max-validation-records 64 \
  --output-dir ml/purpose-classifier/outputs/smoke --overwrite-output

Evaluate

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py

The command returns failure unless label-scorable frozen accuracy is at least 95%, every purpose recall is at least 85%, at least 90% of the deliberately context-free vague-eval slice resolves LOW, every misroute stays within one routing cost tier, and measured batch-one p95 is at most 20 ms. Use --no-gate only for diagnostic runs.

Export and score ONNX

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/export.py
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
  --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
  --report ml/purpose-classifier/outputs/purpose-lite-v1/export/int8-frozen-eval.json

export.py emits fixed-shape opset-17 fp16 and int8-QDQ graphs, a tokenizer/ normalization contract, golden tokenizations, shared calibration config, graph checks, artifact hashes, and a size report. Its default 256-record quantization calibration sample is deterministic and stratified by purpose, slice, and primary language; the export report records the seed, distribution, and prompt hashes. The int8 graph is the ≤25 MiB shipping candidate; the fp16 graph remains the accelerator-oriented conversion input.

When scoring ONNX, add --compare-pytorch to measure artifact drift against --model-dir. The report then includes overall, label-scorable, and per-slice label agreement plus every correct→incorrect, incorrect→correct, and changed-wrong-label transition:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py \
  --onnx-model ml/purpose-classifier/outputs/purpose-lite-v1/export/purpose-lite-v1-int8-qdq.onnx \
  --compare-pytorch --no-gate

Audit curation

Run the semantic embedding duplicate audit. It also emits the deterministic, purpose/slice/language-stratified 10% human label-and-difficulty review CSV:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/audit_data.py

The semantic pass uses the same commit-pinned MiniLM encoder and fixed 128-token input as purpose-lite. Similarity only proposes review candidates; it never edits source data or the frozen split automatically. The 18 word-trigram exclusions and the human-review completion rule are recorded in data/curation-review-v1.json; the semantic report is versioned as data/semantic-audit-v1.json.

The one-time, hardware-bound energy and accelerator-residency procedure is in ENERGY_AND_RESIDENCY.md.