Skip to content
EvalSuite
Documentation menu

Inference efficiency and cost

Requests per second

Implementedefficiency.requests_per_second

Definition

Completed requests per second over the measurement window, from request completion times.

Formula

(n − 1) / (t_last − t_first)

Range: [0, ∞)

Inputs and outputs

  • completion_times: see the signature of es.requests_per_second

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • With ``start_time`` (when the load test began) the rate is n / (t_last − start_time).

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.requests_per_second([0.1, 0.4, 0.9, 1.3])

References

  1. Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.

Implementation status