Inference efficiency and cost
Time per Output Token (TPOT)
Implemented
efficiency.time_per_output_tokenDefinition
Mean time between output tokens after the first (inter-token latency), per request: (t_end − t_first) / (output tokens − 1).
Formula
TPOT = (t_end − t_first) / (n_out − 1)
Range: [0, ∞) s
Inputs and outputs
- first_token_times: see the signature of es.time_per_output_token
- end_times: see the signature of es.time_per_output_token
- output_tokens: see the signature of es.time_per_output_token
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.time_per_output_token(first_token_times=[0.2], end_times=[2.2], output_tokens=[101])References
- Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with PagedAttention. SOSP. 2023:611-626.
- Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.