Files
nucleic-purpose-classifier/ENERGY_AND_RESIDENCY.md
T

47 lines
2.5 KiB
Markdown
Raw Normal View History

# Purpose-classifier energy and accelerator-residency check
This is the required per-model-version hardware spot check from
`docs/PURPOSE_CLASSIFIER.md` §8.5. It is deliberately not CI: the numbers are meaningful
only on named physical hardware with a stable power source, thermal state, and runtime
placement.
## Common protocol
1. Record the artifact SHA-256, app/build commit, machine model, OS build, battery/AC
state, ambient power mode, and runtime compute policy.
2. Warm the model with 100 batch-one classifications, then classify the same fixed
1,000-prompt sequence. Tokenization is included. Disable network activity and other
foreground workloads.
3. Run at least five trials per compute policy, alternating policy order. Report median
wall time, p95 inference latency, average package power, and joules per inference.
4. Capture the runtime placement evidence alongside the power trace. A fast CPU fallback
is still a residency failure for `purpose-deep`.
## macOS
- Run the release Core ML artifact once with `.cpuAndNeuralEngine`, recording the
`MLComputePlan` ANE operation share and the artifact's required floor.
- Repeat with a diagnostic CPU-only configuration on the same Mac and power source.
- Sample the 1,000-inference region with `sudo powermetrics --samplers cpu_power,gpu_power`
at one-second intervals. Store the raw trace outside git and put its summarized table in
the model metrics report.
- The default ANE policy is accepted only when it uses less average power and fewer joules
per inference than CPU-only. On AC, a `.all` GPU retry must separately meet its declared
GPU operation-share floor before it can activate a deep model.
## Windows
- Record the Windows ML execution provider and assigned device after AOT compilation.
Deep requires QNN NPU, OpenVINO NPU/GPU, NvTensorRT-RTX, or MIGraphX; bare CPU is only
valid for lite.
- On battery compare `MAX_EFFICIENCY` with an explicit CPU session. On AC compare the
selected NPU, otherwise the validated vendor GPU EP, with CPU.
- Capture the 1,000-inference window with HWiNFO sensor logging or the vendor's documented
NPU/GPU telemetry. Include package/device average power and energy per inference in the
model metrics report.
- Verify that unplugging skips the discrete-GPU rung and that an idle session unloads
without keeping the dGPU awake.
The report is incomplete if it gives latency without placement evidence and energy, or if
it compares different prompt sequences between policies.