4.1 KiB
Purpose classifier
This directory is the reproducible data and training pipeline for
docs/PURPOSE_CLASSIFIER.md. The current slice covers work item 2 and the first part of
work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning,
temperature calibration, shared confidence thresholds, and the frozen-set accuracy,
recall, hard-slice, calibration, and latency report.
Data contract
The canonical generated sources are listed in data/generation-manifest.json.
round2-NN.jsonl files are retained generation batches and intentionally duplicate
purpose-prompts-round2.jsonl; they are provenance, not additional training input.
prepare_data.py:
- validates the strict generated-record schema;
- removes exact and high-overlap word-trigram duplicates;
- fails for review if a high-overlap pair has conflicting labels;
- keeps shipped fixtures completely outside source data;
- holds every
vague-evalrecord out of training; - stratifies by primary purpose, slice, and primary language; and
- verifies that the deterministic test partition still matches the versioned
data/frozen-test-v1.jsonl.
The frozen test set is the synthetic JSONL plus the 87 classifiable records in
Tests/NucleicCoreTests/Fixtures/purpose-prompts.json. The fixture file's five general
records are excluded because general is deliberately not a model label. The exact
membership and hashes are locked in data/dataset-v1-manifest.json.
Prepare
From the repository root:
python3 ml/purpose-classifier/validate-data.py
python3 ml/purpose-classifier/prepare_data.py
python3 -m unittest discover -s ml/purpose-classifier/tests -p 'test_*.py'
For a newly generated raw 200-record batch, enable batch-shape checks explicitly with
validate-data.py path/to/batch.jsonl --batch-size 200 --expected-total 200.
The generated train/validation copies land under .artifacts/dataset-v1/ and are
gitignored. A source, curation, seed, or split-policy change that moves the frozen test
set fails closed. After reviewing such a change, intentionally version it with:
python3 ml/purpose-classifier/prepare_data.py --refresh-frozen-test
Train purpose-lite
Use a dedicated virtual environment. The base model is pinned to a specific
sentence-transformers/all-MiniLM-L6-v2 commit: a 6-layer, 384-dimensional encoder. The
training collator always pads/truncates to 128 tokens so the later ONNX/Core ML export
can expose a fixed 1 x 128 runtime shape.
python3 -m venv ml/purpose-classifier/.venv
ml/purpose-classifier/.venv/bin/pip install -r ml/purpose-classifier/requirements.txt
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py
The default requirements use PyTorch's CPU-only wheel on Linux, avoiding an accidental
multi-gigabyte CUDA install in CI and development containers. For NVIDIA, AMD, or Intel
accelerator training, install the platform's torch==2.13.0 build using PyTorch's
platform selector, then install requirements-base.txt.
Training writes a local checkpoint, calibration.json, and metrics.json under
outputs/purpose-lite-v1/. It fits one validation-only temperature and derives nested
HIGH/MEDIUM/LOW cutoffs from calibrated top-one probability plus top-two margin.
For a wiring smoke test, use a small deterministic prefix:
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/train.py \
--epochs 1 --max-train-records 64 --max-validation-records 64 \
--output-dir ml/purpose-classifier/outputs/smoke --overwrite-output
Evaluate
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/eval.py
The command returns failure unless frozen accuracy is at least 95%, every purpose recall
is at least 85%, and measured batch-one p95 is at most 20 ms. Use --no-gate only for
diagnostic runs. Accelerator residency, ONNX export/quantization, tokenizer golden tests,
tier-drift evaluation, and Core ML parity remain follow-on work. Before calling dataset
work item 2 complete, also run a semantic embedding duplicate audit and record the planned
10% human label spot-check; the current dependency-free word-trigram pass is deliberately
conservative.