Merge nucleic/fuzzy-dewy-urchin-hpvd into dev
This commit is contained in:
@@ -1,10 +1,10 @@
|
||||
# Purpose classifier
|
||||
|
||||
This directory is the reproducible data and training pipeline for
|
||||
`docs/PURPOSE_CLASSIFIER.md`. The current slice covers work item 2 and the first part of
|
||||
work item 3: deterministic curation/splitting, a frozen v1 eval set, MiniLM fine-tuning,
|
||||
temperature calibration, shared confidence thresholds, and the frozen-set accuracy,
|
||||
recall, hard-slice, calibration, and latency report.
|
||||
`docs/PURPOSE_CLASSIFIER.md`. The current dataset is the promoted Sol-high v2 reset. Its
|
||||
dataset lineage, completed lite/deep experiments, artifact hashes, decisions, and bounded
|
||||
next steps are locked in `experiments/sol-high-v2.json`. Older v1 results below are kept
|
||||
as implementation history and are not label-compatible candidates for Sol-high v2.
|
||||
|
||||
## Data contract
|
||||
|
||||
@@ -24,8 +24,8 @@ The canonical generated sources are listed in `data/generation-manifest.json`.
|
||||
- verifies that the deterministic test partition still matches the versioned
|
||||
`data/frozen-test-v1.jsonl`.
|
||||
|
||||
The frozen test set is the synthetic JSONL plus the 87 classifiable records in
|
||||
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's five `general`
|
||||
The Sol-high v2 frozen test set is the synthetic JSONL plus the 89 classifiable records in
|
||||
`Tests/NucleicCoreTests/Fixtures/purpose-prompts.json`. The fixture file's three `general`
|
||||
records are excluded because `general` is deliberately not a model label. The exact
|
||||
membership and hashes are locked in `data/dataset-v1-manifest.json`.
|
||||
|
||||
@@ -100,6 +100,41 @@ The purpose-lite command starts from the pinned MiniLM revision. The purpose-dee
|
||||
starts from the pinned ModernBERT base variant; neither command supplies a prior classifier
|
||||
checkpoint or continuation flag.
|
||||
|
||||
### Sol-high v2 promotion and training result
|
||||
|
||||
Promotion completed on 2026-08-02/03 with 12,164 combined training records, 1,208
|
||||
validation records, and 1,119 synthetic test records. It explicitly excluded the reviewed
|
||||
real-population conflicts at labeled source lines 1201 and 1481. The ignored combined
|
||||
training manifest is bound by SHA-256 in the versioned experiment ledger without exposing
|
||||
private prompt text.
|
||||
|
||||
| Run | Selected validation | Frozen evaluation | Decision |
|
||||
|---|---:|---:|---|
|
||||
| lite, from pretrained | 91.28% primary | 93.29% overall; 88.84% hard | rejected |
|
||||
| lite, boundary continuation | 92.04% primary | 93.72% overall; 91.16% hard | rejected; 13 correct decisions short of 95% |
|
||||
| deep base | 73.39% primary; 67.50% hard | not opened | rejected on validation |
|
||||
| deep distilled continuation | 77.64% primary; 72.92% hard | not opened | rejected on validation |
|
||||
|
||||
The boundary continuation is useful as the selected lite validation baseline and cached
|
||||
teacher, but it is not a shipping candidate. Do not tune another lite continuation against
|
||||
this frozen set, start QAT/export, chain another deep continuation, or start ModernBERT
|
||||
large/purpose-max.
|
||||
|
||||
The next result-bearing work is a validation-only representation probe: compare the
|
||||
current pretrained prediction-head CLS representation with attention-masked mean pooling
|
||||
using a regularized primary-only linear head. A pooling change earns one full
|
||||
ModernBERT-base run only if it gains at least three points on validation overall or hard,
|
||||
reaches at least 85% overall, and causes no per-purpose collapse. Before any deep frozen
|
||||
look, the revised model must come within two points of lite overall and beat lite by at
|
||||
least three points on the identical validation hard slice. If the probe fails, compare a
|
||||
sentence-trained encoder with the same protocol or remove the deep tier; increasing model
|
||||
size is not the next variable.
|
||||
|
||||
For lite, use validation-only error analysis and data improvements, then reserve a new
|
||||
independent holdout before the next candidate cycle. QAT, export, target-runtime parity,
|
||||
latency, energy, and residency resume only after a new float candidate qualifies. See
|
||||
`experiments/sol-high-v2.json` for exact metrics, hashes, and gates.
|
||||
|
||||
## Label local Nucleic history
|
||||
|
||||
`export_nucleic_prompts.py` extracts the first `userText` event from every local
|
||||
@@ -333,9 +368,10 @@ The full Metal run stopped after epoch three and selected epoch two at 94.75% fa
|
||||
validation accuracy. Its 23,148,500-byte int8-QDQ export scores **95.20% frozen
|
||||
(892/937)**, 95.19% macro recall, 94.17% scored-hard accuracy, and 98.08% scorable
|
||||
PyTorch↔ONNX agreement. Every purpose recall is above 91%, and the vague-abstention and
|
||||
routing-tier-drift gates pass. This is the current accuracy-qualified shipping candidate;
|
||||
latency and energy/residency still require measurement on the target Apple and Windows
|
||||
accelerator runtimes.
|
||||
routing-tier-drift gates pass. This became the accepted artifact for the historical v1
|
||||
label contract; it is not a Sol-high v2 candidate. Its remaining latency and
|
||||
energy/residency work is relevant only as runtime evidence unless a new candidate adopts
|
||||
the same export path.
|
||||
|
||||
### First-prompt history augmentation experiment
|
||||
|
||||
@@ -390,10 +426,15 @@ float checkpoint improves frozen v1 to **95.09% (891/937)**, two correct decisio
|
||||
the accepted baseline's float checkpoint. That gain does not survive export: the matched
|
||||
23,148,500-byte int8-QDQ graph scores **94.66% (887/937)**, 94.66% macro recall, and
|
||||
94.17% scored-hard accuracy. It is five correct decisions behind the accepted int8
|
||||
baseline and fails the 95% shipping gate. Preserve the artifact as a rejected experiment;
|
||||
`purpose-lite-v1-distilled-qat-mlx-4e` remains the candidate of record.
|
||||
baseline and fails the 95% shipping gate. Preserve the artifact as a rejected historical
|
||||
experiment; `purpose-lite-v1-distilled-qat-mlx-4e` remains the accepted v1 artifact only.
|
||||
|
||||
## Train purpose-deep
|
||||
## Purpose-deep implementation and historical v1 experiments
|
||||
|
||||
The Sol-high v2 base and distilled runs described above supersede the experimental
|
||||
sequence in this section. Keep the commands and results below for implementation lineage;
|
||||
do not run another continuation or the large rung from them. The representation probe in
|
||||
`experiments/sol-high-v2.json` is the current deep next step.
|
||||
|
||||
The next classifier tier is an MLX-native ModernBERT multi-task model. It keeps the
|
||||
primary eight-way purpose output and jointly learns secondary purpose, a mixed-intent
|
||||
@@ -513,9 +554,9 @@ distillation regression cannot overwrite the 80.31% candidate. Teacher agreement
|
||||
reported for diagnosis but does not enter deep checkpoint selection; overall and hard
|
||||
primary label accuracy remain the only selection inputs.
|
||||
|
||||
Do not launch the large rung yet. It is justified only after base is evaluated on the
|
||||
frozen set; large must beat base by at least two hard-slice points, while deep itself must
|
||||
reach 97% scored overall and beat the shipping lite artifact by five hard-slice points.
|
||||
The historical large-rung rule required base to qualify first, large to beat base by at
|
||||
least two hard-slice points, and deep to reach 97% scored overall while beating lite by
|
||||
five hard-slice points. Neither the v1 nor Sol-high v2 base result unlocked that rung.
|
||||
|
||||
### Convert and validate Core ML
|
||||
|
||||
|
||||
Reference in New Issue
Block a user