Inference efficiency and cost
Throughput / tokens per second
Implemented
efficiency.throughputDefinition
Output tokens generated per second: system throughput over the measurement window (total tokens / wall-clock span), and the mean per-request decode speed when request durations are given.
Formula
Σ output tokens / (t_last_end − t_first_start)
Range: [0, ∞) tokens/s
Inputs and outputs
- output_tokens: see the signature of es.throughput
- start_times: see the signature of es.throughput
- end_times: see the signature of es.throughput
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.throughput(output_tokens=[200, 300], start_times=[0, 0.5], end_times=[4, 5])References
- Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with PagedAttention. SOSP. 2023:611-626.
- Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.