Inference efficiency and cost
Energy per request
Implemented
efficiency.energy_per_requestDefinition
Energy consumed per request in watt-hours, from power readings sampled at a fixed interval over the run (trapezoidal integration) or from a measured total, divided by the requests served.
Formula
E = ∫ P dt / 3600 / n_requests
Range: [0, ∞) Wh
Inputs and outputs
- power_watts: see the signature of es.energy_per_request
- interval_s: see the signature of es.energy_per_request
- n_requests: see the signature of es.energy_per_request
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.energy_per_request([300, 320, 310], interval_s=1.0, n_requests=10)References
- Luccioni S, Jernite Y, Strubell E. Power hungry processing: watts driving the cost of AI deployment? ACM FAccT. 2024:85-99.