Files
nucleic-purpose-classifier/ENERGY_AND_RESIDENCY.md
T

62 lines
3.5 KiB
Markdown

# Purpose-classifier energy and accelerator-residency check
This is the required per-model-version hardware spot check from
`docs/PURPOSE_CLASSIFIER.md` §8.5. It is deliberately not CI: the numbers are meaningful
only on named physical hardware with a stable power source, thermal state, and runtime
placement.
## Common protocol
1. Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC
state, ambient power mode, and runtime compute policy.
2. Warm the model with 100 batch-one classifications, then classify the same fixed
1,000-prompt sequence. Tokenization is included. Disable network activity and other
foreground workloads.
3. Run at least five trials per compute policy, alternating policy order. Report median
wall time, p95 inference latency, average package power, and joules per inference.
4. Capture the runtime placement evidence alongside the power trace. A fast CPU fallback
is still a residency failure for `purpose-deep`.
## macOS
- Generate the float ML Program with `convert_coreml.py`, calibrate the W8A8 candidate
with `quantize_coreml.py`, then run the gated `eval.py
--coreml-model ... --coreml-compute-units cpu-and-ne --compare-onnx ...` command from
the README. Preserve the conversion manifest and frozen report with the model metrics.
- Run `inspect_coreml.py` with `--compute-units cpu-and-ne`; preserve its full operation
report and record both `neuralEngineOperationShare` and
`neuralEngineEstimatedCostShare`. A gated evaluation without the ONNX comparison is
invalid.
- Run the release Core ML artifact once with `.cpuAndNeuralEngine`, recording the
`MLComputePlan` ANE operation share and the artifact's required floor.
- Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
- Sample the 1,000-inference region with `sudo powermetrics --samplers cpu_power,gpu_power`
at one-second intervals. Store the raw trace outside git and put its summarized table in
the model metrics report.
- The default ANE policy is accepted only when it uses less average power and fewer joules
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
GPU operation-share floor before it can activate a deep model.
The first float16 candidate is a diagnostic baseline only: it achieved 94.98% scored
accuracy and 97.97% agreement with the accepted ONNX artifact, so it fails the rollout
accuracy and parity gates despite 1.51 ms p95 latency. Its compute plan preferred the ANE
for 150/165 operations with placement information (90.91%) and 59.71% of estimated cost.
Do not spend energy-measurement time on that rejected package; repeat placement, latency,
and energy measurements on the first accuracy-qualified W8A8 package.
## Windows
- Record the Windows ML execution provider and assigned device after AOT compilation.
Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only
valid for lite.
- On battery compare `MAX_EFFICIENCY` with an explicit CPU session. On AC compare the
selected NPU, otherwise the validated vendor GPU EP, with CPU.
- Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented
NPU/GPU telemetry. Include package/device average power and energy per inference in the
model metrics report.
- Verify that unplugging skips the discrete-GPU rung and that an idle session unloads
without keeping the dGPU awake.
The report is incomplete if it gives latency without placement evidence and energy, or if
it compares different prompt sequences between policies.