Skip to content
EvalSuite
Documentation menu

Inference efficiency and cost

Time to First Token (TTFT)

Implementedefficiency.time_to_first_token

Definition

Time from sending a request to receiving the first output token; mean with p50/p90/p95/p99.

Formula

TTFT = t_first_token − t_request

Range: [0, ∞) s

Inputs and outputs

  • request_times: see the signature of es.time_to_first_token
  • first_token_times: see the signature of es.time_to_first_token

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.time_to_first_token(request_times=[0.0, 1.0], first_token_times=[0.21, 1.35])

References

  1. Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with PagedAttention. SOSP. 2023:611-626.
  2. Reddi VJ, Cheng C, Kanter D, et al. MLPerf inference benchmark. ISCA. 2020:446-459.

Implementation status