Merge nucleic/sleek-ember-seal-uady into dev

This commit is contained in:
2026-07-30 21:08:05 -07:00
parent 9e462a79fc
commit 40176d0c76
4 changed files with 373 additions and 3 deletions
+9 -1
View File
@@ -19,7 +19,8 @@ placement.
## macOS
- Generate the ML Program with `convert_coreml.py`, then run the gated `eval.py
- Generate the float ML Program with `convert_coreml.py`, calibrate the W8A8 candidate
with `quantize_coreml.py`, then run the gated `eval.py
--coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ...` command from
the README. Preserve the conversion manifest and frozen report with the model metrics.
- Run `inspect_coreml.py` with `--compute-units cpu-and-ne`; preserve its full operation
@@ -36,6 +37,13 @@ placement.
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
GPU operation-share floor before it can activate a deep model.
The first float16 candidate is a diagnostic baseline only: it achieved 94.98% scored
accuracy and 97.97% agreement with the accepted ONNX artifact, so it fails the rollout
accuracy and parity gates despite 1.51 ms p95 latency. Its compute plan preferred the ANE
for 150/165 operations with placement information (90.91%) and 59.71% of estimated cost.
Do not spend energy-measurement time on that rejected package; repeat placement, latency,
and energy measurements on the first accuracy-qualified W8A8 package.
## Windows
- Record the Windows ML execution provider and assigned device after AOT compilation.