2026-07-30 05:04:34 -07:00
|
|
|
# Purpose-classifier energy and accelerator-residency check
|
|
|
|
|
|
|
|
|
|
This is the required per-model-version hardware spot check from
|
|
|
|
|
`docs/PURPOSE_CLASSIFIER.md` §8.5. It is deliberately not CI: the numbers are meaningful
|
|
|
|
|
only on named physical hardware with a stable power source, thermal state, and runtime
|
|
|
|
|
placement.
|
|
|
|
|
|
|
|
|
|
## Common protocol
|
|
|
|
|
|
|
|
|
|
1. Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC
|
|
|
|
|
state, ambient power mode, and runtime compute policy.
|
|
|
|
|
2. Warm the model with 100 batch-one classifications, then classify the same fixed
|
|
|
|
|
1,000-prompt sequence. Tokenization is included. Disable network activity and other
|
|
|
|
|
foreground workloads.
|
|
|
|
|
3. Run at least five trials per compute policy, alternating policy order. Report median
|
|
|
|
|
wall time, p95 inference latency, average package power, and joules per inference.
|
|
|
|
|
4. Capture the runtime placement evidence alongside the power trace. A fast CPU fallback
|
|
|
|
|
is still a residency failure for `purpose-deep`.
|
|
|
|
|
|
|
|
|
|
## macOS
|
|
|
|
|
|
2026-07-30 21:08:05 -07:00
|
|
|
- Generate the float ML Program with `convert_coreml.py`, calibrate the W8A8 candidate
|
|
|
|
|
with `quantize_coreml.py`, then run the gated `eval.py
|
2026-07-30 20:56:46 -07:00
|
|
|
--coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ...` command from
|
|
|
|
|
the README. Preserve the conversion manifest and frozen report with the model metrics.
|
|
|
|
|
- Run `inspect_coreml.py` with `--compute-units cpu-and-ne`; preserve its full operation
|
|
|
|
|
report and record both `neuralEngineOperationShare` and
|
|
|
|
|
`neuralEngineEstimatedCostShare`. A gated evaluation without the ONNX comparison is
|
|
|
|
|
invalid.
|
2026-07-30 05:04:34 -07:00
|
|
|
- Run the release Core ML artifact once with `.cpuAndNeuralEngine`, recording the
|
|
|
|
|
`MLComputePlan` ANE operation share and the artifact's required floor.
|
|
|
|
|
- Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
|
|
|
|
|
- Sample the 1,000-inference region with `sudo powermetrics --samplers cpu_power,gpu_power`
|
|
|
|
|
at one-second intervals. Store the raw trace outside git and put its summarized table in
|
|
|
|
|
the model metrics report.
|
|
|
|
|
- The default ANE policy is accepted only when it uses less average power and fewer joules
|
|
|
|
|
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
|
|
|
|
|
GPU operation-share floor before it can activate a deep model.
|
|
|
|
|
|
2026-07-30 21:08:05 -07:00
|
|
|
The first float16 candidate is a diagnostic baseline only: it achieved 94.98% scored
|
|
|
|
|
accuracy and 97.97% agreement with the accepted ONNX artifact, so it fails the rollout
|
|
|
|
|
accuracy and parity gates despite 1.51 ms p95 latency. Its compute plan preferred the ANE
|
|
|
|
|
for 150/165 operations with placement information (90.91%) and 59.71% of estimated cost.
|
|
|
|
|
Do not spend energy-measurement time on that rejected package; repeat placement, latency,
|
|
|
|
|
and energy measurements on the first accuracy-qualified W8A8 package.
|
|
|
|
|
|
2026-07-30 05:04:34 -07:00
|
|
|
## Windows
|
|
|
|
|
|
|
|
|
|
- Record the Windows ML execution provider and assigned device after AOT compilation.
|
|
|
|
|
Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only
|
|
|
|
|
valid for lite.
|
|
|
|
|
- On battery compare `MAX_EFFICIENCY` with an explicit CPU session. On AC compare the
|
|
|
|
|
selected NPU, otherwise the validated vendor GPU EP, with CPU.
|
|
|
|
|
- Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented
|
|
|
|
|
NPU/GPU telemetry. Include package/device average power and energy per inference in the
|
|
|
|
|
model metrics report.
|
|
|
|
|
- Verify that unplugging skips the discrete-GPU rung and that an idle session unloads
|
|
|
|
|
without keeping the dGPU awake.
|
|
|
|
|
|
|
|
|
|
The report is incomplete if it gives latency without placement evidence and energy, or if
|
|
|
|
|
it compares different prompt sequences between policies.
|