3.0 KiB
3.0 KiB
Purpose-classifier energy and accelerator-residency check
This is the required per-model-version hardware spot check from
docs/PURPOSE_CLASSIFIER.md §8.5. It is deliberately not CI: the numbers are meaningful
only on named physical hardware with a stable power source, thermal state, and runtime
placement.
Common protocol
- Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC state, ambient power mode, and runtime compute policy.
- Warm the model with 100 batch-one classifications, then classify the same fixed 1,000-prompt sequence. Tokenization is included. Disable network activity and other foreground workloads.
- Run at least five trials per compute policy, alternating policy order. Report median wall time, p95 inference latency, average package power, and joules per inference.
- Capture the runtime placement evidence alongside the power trace. A fast CPU fallback
is still a residency failure for
purpose-deep.
macOS
- Generate the ML Program with
convert_coreml.py, then run the gatedeval.py --coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ...command from the README. Preserve the conversion manifest and frozen report with the model metrics. - Run
inspect_coreml.pywith--compute-units cpu-and-ne; preserve its full operation report and record bothneuralEngineOperationShareandneuralEngineEstimatedCostShare. A gated evaluation without the ONNX comparison is invalid. - Run the release Core ML artifact once with
.cpuAndNeuralEngine, recording theMLComputePlanANE operation share and the artifact's required floor. - Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
- Sample the 1,000-inference region with
sudo powermetrics --samplers cpu_power,gpu_powerat one-second intervals. Store the raw trace outside git and put its summarized table in the model metrics report. - The default ANE policy is accepted only when it uses less average power and fewer joules
per inference than CPU-only. On AC, a
.allGPU retry must separately meet its declared GPU operation-share floor before it can activate a deep model.
Windows
- Record the Windows ML execution provider and assigned device after AOT compilation. Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only valid for lite.
- On battery compare
MAX_EFFICIENCYwith an explicit CPU session. On AC compare the selected NPU, otherwise the validated vendor GPU EP, with CPU. - Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented NPU/GPU telemetry. Include package/device average power and energy per inference in the model metrics report.
- Verify that unplugging skips the discrete-GPU rung and that an idle session unloads without keeping the dGPU awake.
The report is incomplete if it gives latency without placement evidence and energy, or if it compares different prompt sequences between policies.