Merge nucleic/sleek-ember-seal-uady into dev
This commit is contained in:
@@ -19,7 +19,8 @@ placement.
|
||||
|
||||
## macOS
|
||||
|
||||
- Generate the ML Program with `convert_coreml.py`, then run the gated `eval.py
|
||||
- Generate the float ML Program with `convert_coreml.py`, calibrate the W8A8 candidate
|
||||
with `quantize_coreml.py`, then run the gated `eval.py
|
||||
--coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ...` command from
|
||||
the README. Preserve the conversion manifest and frozen report with the model metrics.
|
||||
- Run `inspect_coreml.py` with `--compute-units cpu-and-ne`; preserve its full operation
|
||||
@@ -36,6 +37,13 @@ placement.
|
||||
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
|
||||
GPU operation-share floor before it can activate a deep model.
|
||||
|
||||
The first float16 candidate is a diagnostic baseline only: it achieved 94.98% scored
|
||||
accuracy and 97.97% agreement with the accepted ONNX artifact, so it fails the rollout
|
||||
accuracy and parity gates despite 1.51 ms p95 latency. Its compute plan preferred the ANE
|
||||
for 150/165 operations with placement information (90.91%) and 59.71% of estimated cost.
|
||||
Do not spend energy-measurement time on that rejected package; repeat placement, latency,
|
||||
and energy measurements on the first accuracy-qualified W8A8 package.
|
||||
|
||||
## Windows
|
||||
|
||||
- Record the Windows ML execution provider and assigned device after AOT compilation.
|
||||
|
||||
Reference in New Issue
Block a user