Skip to content
EvalSuite
Documentation menu

Inference efficiency and cost

p50 / p95 / p99 latency

Implementedefficiency.latency_percentiles

Definition

End-to-end request latency percentiles; tail latency (p95, p99) governs user experience at scale. The value is p95; mean, p50, p90 and p99 are in ``params``.

Formula

p-th percentile of latency (linear interpolation)

Range: [0, ∞)

Inputs and outputs

  • latencies: see the signature of es.latency_percentiles

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.latency_percentiles([0.8, 1.1, 0.9, 4.2], percentile=95)

References

  1. Dean J, Barroso LA. The tail at scale. Commun ACM. 2013;56(2):74-80.

Implementation status