Skip to content
EvalSuite
Documentation menu

Text generation and semantic similarity

Available since v0.4.0

references come first (like y_true), predictions second. Each example can have one reference string or a list of references. BLEU, chrF and TER are corpus-level on a 0–100 scale with sacreBLEU's tokenization, smoothing and multi-reference rules, so scores are comparable with published results. ROUGE follows Google's rouge-score (0–1, best reference per example), METEOR follows NLTK, CIDEr-D follows the coco-caption implementation.

Python
import evalsuite as es

references = [["the cat sat on the mat", "a cat on a mat"], ["a dog ran"], ["hello world"]]
predictions = ["the cat sat on a mat", "dog ran", "hello there world"]

es.bleu(references, predictions)                 # corpus BLEU (sacrebleu.corpus_bleu)
es.sentence_bleu(references, predictions, average=None)
es.chrf(references, predictions, word_order=2)   # chrF++
es.ter(references, predictions)
es.rouge_l(references, predictions)              # also rouge_1, rouge_2, rouge_lsum
es.meteor(references, predictions)               # Porter stemming; pass synonyms=... for WordNet
es.distinct_n(predictions, n=2)
es.self_bleu(predictions)

METEOR's default stemmer needs NLTK (pip install "evalsuite-python[llm]"). WordNet synonyms need a corpus download, so they are off unless you pass a synonyms function.

Language-model likelihood

Python
import numpy as np
token_logprobs = [np.log([0.5, 0.25, 0.8]), np.log([0.9, 0.6])]   # natural-log probability of each token
es.perplexity(token_logprobs)          # pooled over all tokens
es.cross_entropy(token_logprobs, base=2)  # bits per token

Perplexity is only comparable between models that share a tokenizer.

Embedding-based metrics

Pass the embeddings your encoder produced; EvalSuite computes the metric exactly.

Python
rng = np.random.default_rng(0)
ref_tokens = [rng.normal(size=(5, 16)) for _ in range(3)]    # token embeddings per reference
pred_tokens = [rng.normal(size=(4, 16)) for _ in range(3)]
es.bertscore(ref_tokens, pred_tokens)                          # greedy cosine matching, optional IDF weights
es.moverscore(ref_tokens, pred_tokens)                         # exact transport (same as POT's ot.emd2)
es.embedding_similarity(rng.normal(size=(3, 16)), rng.normal(size=(3, 16)), metric="cosine")

human, model = rng.normal(size=(200, 16)), rng.normal(0.3, 1, size=(200, 16))
es.mauve(human, model, random_state=0)                         # 1 = indistinguishable distributions

MAUVE quantizes both feature sets together (L2 normalisation, PCA to 90% of the variance, k-means) and measures the area under the divergence frontier; the divergence step is identical to mauve-text, and results depend on the clustering, so report random_state.

Metrics

  • BLEUImplemented

    Corpus-level geometric mean of clipped n-gram precisions (n = 1..4) times a brevity penalty, computed exactly as sacreBLEU (13a tokenization, exponential smoothing).

    es.bleu

  • chrF / chrF++Implemented

    F-beta score (β = 2) over character n-grams (n = 1..6), plus word uni- and bigrams for chrF++ (word_order=2); robust for morphologically rich languages. Computed exactly as sacreBLEU.

    es.chrf

  • CIDEr-DImplemented

    Consensus with several references: cosine similarity of TF-IDF weighted n-gram vectors (n = 1..4), clipped to the reference counts and damped by a Gaussian length penalty; IDF comes from the references of the whole evaluated set (the coco-caption implementation).

    es.cider

  • Average negative log-probability the model assigns to the observed tokens, pooled over all tokens of all sequences.

    es.cross_entropy

  • Distinct-nImplemented

    Number of distinct n-grams divided by the total number of n-grams across all predictions (whitespace tokens); a simple lexical-diversity indicator.

    es.distinct_n

  • MAUVEImplemented

    Gap between the distribution of generated text and of human text: both sets of feature vectors are quantized together (L2 normalisation, PCA to 90% variance, k-means), and MAUVE is the area under the divergence frontier of the two histograms. 1 means indistinguishable.

    es.mauve

  • METEORImplemented

    Unigram alignment between prediction and reference by exact, stemmed and (optionally) synonym matches; harmonic mean weighted towards recall, penalised for fragmented alignments. Best reference per example, averaged over examples (NLTK's meteor_score).

    es.meteor

  • PerplexityImplemented

    Exponentiated token-average negative log-likelihood, pooled over all tokens. Only comparable between models that share a tokenizer.

    es.perplexity

  • ROUGE-1Implemented

    Unigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.

    es.rouge_1

  • ROUGE-2Implemented

    Bigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.

    es.rouge_2

  • ROUGE-LImplemented

    Longest common subsequence between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.

    es.rouge_l

  • ROUGE-LsumImplemented

    Summary-level union LCS over sentences (one sentence per line) between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.

    es.rouge_lsum

  • Self-BLEUImplemented

    Average sentence BLEU of each prediction against all other predictions as references; higher means the outputs are more alike (less diverse).

    es.self_bleu

  • Sentence BLEUImplemented

    BLEU computed for each example separately (exponential smoothing, effective order) and averaged; use for per-example scores, not for reporting corpus quality.

    es.sentence_bleu

  • TERImplemented

    Translation edit rate: minimum number of insertions, deletions, substitutions and block shifts to turn the prediction into the closest reference, divided by the average reference length (Tercom, as in sacreBLEU). Lower is better.

    es.ter

  • BERTScoreImplemented

    Greedy matching of contextual token embeddings by cosine similarity: precision averages, over prediction tokens, the best similarity to any reference token; recall does the reverse; F1 combines them. Optional IDF weights and baseline rescaling as in the original implementation.

    es.bertscore

  • Similarity or distance between the sentence embedding of each prediction and of its reference: cosine similarity, or Euclidean / Manhattan distance. Values depend on the embedding model.

    es.embedding_similarity

  • Scores from any learned evaluation model or judge you supply, per example, so they get the same confidence intervals, model comparison and reports as every other metric.

    es.model_score

  • MoverScoreImplemented

    One minus the earth mover's distance between the IDF-weighted token embeddings of prediction and reference, with Euclidean transport cost between L2-normalised embeddings (MoverScore v2 style).

    es.moverscore