Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -19,6 +19,7 @@ The canonical generated sources are listed in `data/generation-manifest.json`.
|
||||
- fails for review if a high-overlap pair has conflicting labels;
|
||||
- keeps shipped fixtures completely outside source data;
|
||||
- holds every `vague-eval` record out of training;
|
||||
- optionally applies a completed, versioned human-review ledger before splitting;
|
||||
- stratifies by primary purpose, slice, and primary language; and
|
||||
- verifies that the deterministic test partition still matches the versioned
|
||||
`data/frozen-test-v1.jsonl`.
|
||||
@@ -177,5 +178,45 @@ the frozen split automatically. The 18 word-trigram exclusions and the human-rev
|
||||
completion rule are recorded in `data/curation-review-v1.json`; the semantic report is
|
||||
versioned as `data/semantic-audit-v1.json`.
|
||||
|
||||
## Complete the human review
|
||||
|
||||
The deterministic CSV currently contains 1,219 blank review rows. Check progress without
|
||||
running the embedding audit again:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py
|
||||
```
|
||||
|
||||
Mark each row `accept`, `relabel`, or `reject`. `accept` and `reject` leave the four
|
||||
`reviewed*` fields blank; `reject` requires notes. For `relabel`, blank reviewed fields
|
||||
retain their generated value, `<none>` clears a secondary purpose, and notes are required.
|
||||
If a secondary purpose is added or removed, set `reviewedSlice` consistently (`mixed`
|
||||
when a secondary is present). The validator rejects stale generated columns, missing or
|
||||
duplicate sample rows, invalid label combinations, and partially completed rows.
|
||||
|
||||
When every row has a human decision, write the versionable ledger:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/review_data.py --finalize
|
||||
```
|
||||
|
||||
Build an isolated candidate split first:
|
||||
|
||||
```bash
|
||||
ml/purpose-classifier/.venv/bin/python ml/purpose-classifier/prepare_data.py \
|
||||
--human-review ml/purpose-classifier/data/human-review-v1.json \
|
||||
--output-dir ml/purpose-classifier/.artifacts/reviewed-candidate \
|
||||
--frozen-test ml/purpose-classifier/.artifacts/reviewed-frozen-candidate.jsonl \
|
||||
--manifest ml/purpose-classifier/.artifacts/reviewed-manifest-candidate.json \
|
||||
--refresh-frozen-test
|
||||
```
|
||||
|
||||
Inspect the ledger, decision summary, candidate manifest, and split diff. Only then rerun
|
||||
the same command with the three candidate-path overrides removed to intentionally replace
|
||||
the versioned frozen dataset and manifest.
|
||||
|
||||
`--regenerate` recreates a blank CSV in the current schema and is only appropriate before
|
||||
review begins.
|
||||
|
||||
The one-time, hardware-bound energy and accelerator-residency procedure is in
|
||||
`ENERGY_AND_RESIDENCY.md`.
|
||||
|
||||
Reference in New Issue
Block a user