Merge nucleic/upbeat-yarn-seal-ekbe into dev

This commit is contained in:
2026-08-01 21:26:31 -07:00
parent 9ecb82c983
commit 8f5acb7755
5 changed files with 1073 additions and 8 deletions
+52
View File
@@ -29,6 +29,58 @@ The frozen test set is the synthetic JSONL plus the 87 classifiable records in
records are excluded because `general` is deliberately not a model label. The exact
membership and hashes are locked in `data/dataset-v1-manifest.json`.
## Full Sol-high label reset
`rebuild_sol_high.py` is the end-to-end workflow for intentionally invalidating the old
labels and rebuilding both classifiers from one teacher. It freezes and relabels every
canonical public prompt, every shipped fixture, local Nucleic first-prompt history, and
the pinned SWE-chat candidates with `gpt-5.6-sol` at high reasoning effort. The teacher
settings are not configurable in this workflow. Batches default to eight prompts and
failed structured responses are retried up to ten times.
The reset is staged below the ignored `.artifacts/sol-high-reset/` directory. Snapshot
hashes prevent a resumed run from silently mixing input revisions or teacher settings.
Run these commands from the repository root with the purpose-classifier environment:
```bash
"$PY" ml/purpose-classifier/rebuild_sol_high.py snapshot
"$PY" ml/purpose-classifier/rebuild_sol_high.py label \
--python "$PY"
"$PY" ml/purpose-classifier/rebuild_sol_high.py status
```
Rerunning `label` resumes every existing decision log and does not relabel completed
lines. Once `status` reports `"complete": true`, promote the staged result explicitly:
```bash
"$PY" ml/purpose-classifier/rebuild_sol_high.py promote \
--confirm overwrite-all-labels-with-sol-high
```
Promotion first validates and curates the complete staged population. It then replaces
the two canonical source files, regenerates all round-two mirrors, relabels shipped
fixtures (teacher-rejected fixtures become the runtime `general` fallback), rebuilds the
frozen public evaluation split, and replaces `.artifacts/dataset-v1/` with the combined
training dataset. Nucleic history and SWE-chat are deduplicated against public evaluation
data and added to training only; `vague-eval` records remain optimization exclusions.
The prior files are moved into a timestamped ignored backup before replacement.
The old human/semantic review assertions are marked superseded because they were tied to
the invalidated label population. Consequently, Sol-high v2 evaluation measures agreement
with the new teacher and is not directly comparable with the old v1 scores.
Print the two clean, from-pretrained-base training commands after promotion:
```bash
"$PY" ml/purpose-classifier/rebuild_sol_high.py train-commands
```
The purpose-lite command starts from the pinned MiniLM revision. The purpose-deep command
starts from the pinned ModernBERT base variant; neither command supplies a prior classifier
checkpoint or continuation flag.
## Label local Nucleic history
`export_nucleic_prompts.py` extracts the first `userText` event from every local