Skip to content
EvalSuite
Documentation menu

Text generation

MAUVE

Implementedtext-generation.mauve

Definition

Gap between the distribution of generated text and of human text: both sets of feature vectors are quantized together (L2 normalisation, PCA to 90% variance, k-means), and MAUVE is the area under the divergence frontier of the two histograms. 1 means indistinguishable.

Formula

area under {(exp(−c·KL(Q‖R_λ)), exp(−c·KL(P‖R_λ))) : R_λ = λP + (1−λ)Q}

Range: (0, 1]

Inputs and outputs

  • reference_features: one feature vector per human text
  • generated_features: one feature vector per generated text

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.mauve(human_features, model_features, random_state=0)

References

  1. Pillutla K, Swayamdipta S, Zellers R, Thickstun J, Welleck S, Choi Y, Harchaoui Z. MAUVE: measuring the gap between neural text and human text using divergence frontiers. NeurIPS. 2021.
  2. Pillutla K, Liu L, Thickstun J, Welleck S, Swayamdipta S, Zellers R, Oh S, Choi Y, Harchaoui Z. MAUVE scores for generative models: theory and practice. JMLR. 2023;24(356):1-92.

Implementation status