Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 20:14:37 -07:00
parent 09f98cdd00
commit 90092332db
8 changed files with 1668 additions and 0 deletions
+43
View File
@@ -145,6 +145,49 @@ backpropagation, selection, and ordinary checkpoint reload. The current shared C
then showed severe post-batch throttling, so no full candidate result is claimed from that
canary.
### Native Apple Silicon training with MLX
Use the MLX backend when training on Apple Silicon. It implements the same six-layer BERT
classifier, fixed head-tail tokenization, export-matched QAT graph, cached-teacher
distillation, validation selection, and early stopping with native MLX arrays. Fake
quantization is decomposed into Metal-supported round, clip, and straight-through-gradient
operations, avoiding PyTorch's unsupported MPS fake-quant operator. Selected weights are
written back with the original Hugging Face parameter names, so the existing PyTorch
`export.py` and `eval.py` paths remain unchanged.
Install the additional pinned dependency into the macOS virtual environment:
```bash
ml/purpose-classifier/venv/bin/python -m pip install \
-r ml/purpose-classifier/requirements-mlx.txt
```
Before the first full run on a new MLX or Transformers version, run the fail-closed parity
check. It requires exact fake-quant primitives, float-logit parity, matching QAT
predictions with bounded backend drift, healthy QAT gradients, and an exact Hugging Face
→ MLX → Hugging Face weight round trip:
```bash
ml/purpose-classifier/venv/bin/python ml/purpose-classifier/verify_mlx.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model
```
Then run the distilled QAT candidate natively on Metal:
```bash
ml/purpose-classifier/venv/bin/python -u ml/purpose-classifier/train_mlx.py \
--model ml/purpose-classifier/outputs/purpose-lite-v1-boundary-tune/model \
--distillation-cache \
ml/purpose-classifier/outputs/purpose-lite-v1-boundary-teacher.pt \
--distillation-weight 0.9 --distillation-temperature 2 \
--distillation-selection-weight 0.5 --quantization-aware \
--epochs 2 --early-stopping-patience 1 \
--learning-rate 1e-6 --warmup-ratio 0 --boundary-weight 1 \
--progress-steps 1 \
--output-dir ml/purpose-classifier/outputs/purpose-lite-v1-distilled-qat-mlx \
--overwrite-output
```
For a wiring smoke test, use a small deterministic prefix:
```bash