Inference efficiency and cost
p50 / p95 / p99 latency
Implemented
efficiency.latency_percentilesDefinition
End-to-end request latency percentiles; tail latency (p95, p99) governs user experience at scale. The value is p95; mean, p50, p90 and p99 are in ``params``.
Formula
p-th percentile of latency (linear interpolation)
Range: [0, ∞)
Inputs and outputs
- latencies: see the signature of es.latency_percentiles
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.latency_percentiles([0.8, 1.1, 0.9, 4.2], percentile=95)References
- Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80.