Files
nucleic-purpose-classifier/README.md
T

4.1 KiB

Purpose classifier

This directory is the reproducible data and training pipeline for docs/PURPOSE_CLASSIFIER.md. The current slice covers work item 2 and the first part of work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning, temperature calibration, shared confidence thresholds, and the frozen-set accuracy, recall, hard-slice, calibration, and latency report.

Data contract

The canonical generated sources are listed in data/generation-manifest.json. round2-NN.jsonl files are retained generation batches and intentionally duplicate purpose-prompts-round2.jsonl; they are provenance, not additional training input.

prepare_data.py:

  • validates the strict generated-record schema;
  • removes exact and high-overlap word-trigram duplicates;
  • fails for review if a high-overlap pair has conflicting labels;
  • keeps shipped fixtures completely outside source data;
  • holds every vague-eval record out of training;
  • stratifies by primary purpose, slice, and primary language; and
  • verifies that the deterministic test partition still matches the versioned data/frozen-test-v1.jsonl.

The frozen test set is the synthetic JSONL plus the 87 classifiable records in Tests/NucleicCoreTests/Fixtures/purpose-prompts.json. The fixture file's five general records are excluded because general is deliberately not a model label. The exact membership and hashes are locked in data/dataset-v1-manifest.json.

Prepare

From the repository root:

python3 ml/purpose-classifier/validate-data.py
python3 ml/purpose-classifier/prepare_data.py
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'

For a newly generated raw 200-record batch, enable batch-shape checks explicitly with validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200.

The generated train/validation copies land under .artifacts/dataset-v1/ and are gitignored. A source, curation, seed, or split-policy change that moves the frozen test set fails closed. After reviewing such a change, intentionally version it with:

python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test

Train purpose-lite

Use a dedicated virtual environment. The base model is pinned to a specific sentence-transformers/all-MiniLM-L6-v2 commit: a 6-layer, 384-dimensional encoder. The training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export can expose a fixed 1 x 128 runtime shape.

python3 -m venv ml/purpose-classifier/.venv
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py

The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel accelerator training, install the platform's torch==2.13.0 build using PyTorch's platform selector, then install requirements-base.txt.

Training writes a local checkpoint, calibration.json, and metrics.json under outputs/purpose-lite-v1/. It fits one validation-only temperature and derives nested HIGH/MEDIUM/LOW cutoffs from calibrated top-one probability plus top-two margin.

For a wiring smoke test, use a small deterministic prefix:

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
  --epochs 1 --max-train-records 64 --max-validation-records 64 \
  --output-dir ml/purpose-classifier/outputs/smoke --overwrite-output

Evaluate

ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py

The command returns failure unless frozen accuracy is at least 95%, every purpose recall is at least 85%, and measured batch-one p95 is at most 20 ms. Use --no-gate only for diagnostic runs. Accelerator residency, ONNX export/quantization, tokenizer golden tests, tier-drift evaluation, and Core ML parity remain follow-on work. Before calling dataset work item 2 complete, also run a semantic embedding duplicate audit and record the planned 10% human label spot-check; the current dependency-free word-trigram pass is deliberately conservative.