Inference efficiency and cost
Requests per second
Implemented
efficiency.requests_per_secondDefinition
Completed requests per second over the measurement window, from request completion times.
Formula
(n − 1) / (t_last − t_first)
Range: [0, ∞)
Inputs and outputs
- completion_times: see the signature of es.requests_per_second
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- With ``start_time`` (when the load test began) the rate is n / (t_last − start_time).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.requests_per_second([0.1, 0.4, 0.9, 1.3])References
- Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.