Files

3.5 KiB

Purpose-classifier energy and accelerator-residency check

This is the required per-model-version hardware spot check from docs/PURPOSE_CLASSIFIER.md §8.5. It is deliberately not CI: the numbers are meaningful only on named physical hardware with a stable power source, thermal state, and runtime placement.

Common protocol

  1. Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC state, ambient power mode, and runtime compute policy.
  2. Warm the model with 100 batch-one classifications, then classify the same fixed 1,000-prompt sequence. Tokenization is included. Disable network activity and other foreground workloads.
  3. Run at least five trials per compute policy, alternating policy order. Report median wall time, p95 inference latency, average package power, and joules per inference.
  4. Capture the runtime placement evidence alongside the power trace. A fast CPU fallback is still a residency failure for purpose-deep.

macOS

  • Generate the float ML Program with convert_coreml.py, calibrate the W8A8 candidate with quantize_coreml.py, then run the gated eval.py --coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ... command from the README. Preserve the conversion manifest and frozen report with the model metrics.
  • Run inspect_coreml.py with --compute-units cpu-and-ne; preserve its full operation report and record both neuralEngineOperationShare and neuralEngineEstimatedCostShare. A gated evaluation without the ONNX comparison is invalid.
  • Run the release Core ML artifact once with .cpuAndNeuralEngine, recording the MLComputePlan ANE operation share and the artifact's required floor.
  • Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
  • Sample the 1,000-inference region with sudo powermetrics --samplers cpu_power,gpu_power at one-second intervals. Store the raw trace outside git and put its summarized table in the model metrics report.
  • The default ANE policy is accepted only when it uses less average power and fewer joules per inference than CPU-only. On AC, a .all GPU retry must separately meet its declared GPU operation-share floor before it can activate a deep model.

The first float16 candidate is a diagnostic baseline only: it achieved 94.98% scored accuracy and 97.97% agreement with the accepted ONNX artifact, so it fails the rollout accuracy and parity gates despite 1.51 ms p95 latency. Its compute plan preferred the ANE for 150/165 operations with placement information (90.91%) and 59.71% of estimated cost. Do not spend energy-measurement time on that rejected package; repeat placement, latency, and energy measurements on the first accuracy-qualified W8A8 package.

Windows

  • Record the Windows ML execution provider and assigned device after AOT compilation. Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only valid for lite.
  • On battery compare MAX_EFFICIENCY with an explicit CPU session. On AC compare the selected NPU, otherwise the validated vendor GPU EP, with CPU.
  • Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented NPU/GPU telemetry. Include package/device average power and energy per inference in the model metrics report.
  • Verify that unplugging skips the discrete-GPU rung and that an idle session unloads without keeping the dGPU awake.

The report is incomplete if it gives latency without placement evidence and energy, or if it compares different prompt sequences between policies.