Skip to content
EvalSuite
Documentation menu

Text generation

Perplexity

Implementedtext-generation.perplexity

Definition

Exponentiated token-average negative log-likelihood, pooled over all tokens. Only comparable between models that share a tokenizer.

Formula

PPL = exp(−(1/T) Σ_t log p(x_t | x_<t))

Range: [1, ∞)

Inputs and outputs

  • token_logprobs: per-token natural-log probabilities, one array per sequence

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.perplexity(token_logprobs)

References

  1. Jelinek F, Mercer RL, Bahl LR, Baker JK. Perplexity—a measure of the difficulty of speech recognition tasks. JASA. 1977;62(S1):S63.

Implementation status