Text generation and semantic similarity
Available since v0.4.0references come first (like y_true), predictions second. Each example can have one reference string or a list of references. BLEU, chrF and TER are corpus-level on a 0–100 scale with sacreBLEU's tokenization, smoothing and multi-reference rules, so scores are comparable with published results. ROUGE follows Google's rouge-score (0–1, best reference per example), METEOR follows NLTK, CIDEr-D follows the coco-caption implementation.
import evalsuite as es
references = [["the cat sat on the mat", "a cat on a mat"], ["a dog ran"], ["hello world"]]
predictions = ["the cat sat on a mat", "dog ran", "hello there world"]
es.bleu(references, predictions) # corpus BLEU (sacrebleu.corpus_bleu)
es.sentence_bleu(references, predictions, average=None)
es.chrf(references, predictions, word_order=2) # chrF++
es.ter(references, predictions)
es.rouge_l(references, predictions) # also rouge_1, rouge_2, rouge_lsum
es.meteor(references, predictions) # Porter stemming; pass synonyms=... for WordNet
es.distinct_n(predictions, n=2)
es.self_bleu(predictions)METEOR's default stemmer needs NLTK (pip install "evalsuite-python[llm]"). WordNet synonyms need a corpus download, so they are off unless you pass a synonyms function.
Language-model likelihood
import numpy as np
token_logprobs = [np.log([0.5, 0.25, 0.8]), np.log([0.9, 0.6])] # natural-log probability of each token
es.perplexity(token_logprobs) # pooled over all tokens
es.cross_entropy(token_logprobs, base=2) # bits per tokenPerplexity is only comparable between models that share a tokenizer.
Embedding-based metrics
Pass the embeddings your encoder produced; EvalSuite computes the metric exactly.
rng = np.random.default_rng(0)
ref_tokens = [rng.normal(size=(5, 16)) for _ in range(3)] # token embeddings per reference
pred_tokens = [rng.normal(size=(4, 16)) for _ in range(3)]
es.bertscore(ref_tokens, pred_tokens) # greedy cosine matching, optional IDF weights
es.moverscore(ref_tokens, pred_tokens) # exact transport (same as POT's ot.emd2)
es.embedding_similarity(rng.normal(size=(3, 16)), rng.normal(size=(3, 16)), metric="cosine")
human, model = rng.normal(size=(200, 16)), rng.normal(0.3, 1, size=(200, 16))
es.mauve(human, model, random_state=0) # 1 = indistinguishable distributionsMAUVE quantizes both feature sets together (L2 normalisation, PCA to 90% of the variance, k-means) and measures the area under the divergence frontier; the divergence step is identical to mauve-text, and results depend on the clustering, so report random_state.
Metrics
| Metric | Description | Status | API |
|---|---|---|---|
| BLEU | Corpus-level geometric mean of clipped n-gram precisions (n = 1..4) times a brevity penalty, computed exactly as sacreBLEU (13a tokenization, exponential smoothing). | Implemented | es.bleu |
| chrF / chrF++ | F-beta score (β = 2) over character n-grams (n = 1..6), plus word uni- and bigrams for chrF++ (word_order=2); robust for morphologically rich languages. Computed exactly as sacreBLEU. | Implemented | es.chrf |
| CIDEr-D | Consensus with several references: cosine similarity of TF-IDF weighted n-gram vectors (n = 1..4), clipped to the reference counts and damped by a Gaussian length penalty; IDF comes from the references of the whole evaluated set (the coco-caption implementation). | Implemented | es.cider |
| Cross-entropy (negative log-likelihood) | Average negative log-probability the model assigns to the observed tokens, pooled over all tokens of all sequences. | Implemented | es.cross_entropy |
| Distinct-n | Number of distinct n-grams divided by the total number of n-grams across all predictions (whitespace tokens); a simple lexical-diversity indicator. | Implemented | es.distinct_n |
| MAUVE | Gap between the distribution of generated text and of human text: both sets of feature vectors are quantized together (L2 normalisation, PCA to 90% variance, k-means), and MAUVE is the area under the divergence frontier of the two histograms. 1 means indistinguishable. | Implemented | es.mauve |
| METEOR | Unigram alignment between prediction and reference by exact, stemmed and (optionally) synonym matches; harmonic mean weighted towards recall, penalised for fragmented alignments. Best reference per example, averaged over examples (NLTK's meteor_score). | Implemented | es.meteor |
| Perplexity | Exponentiated token-average negative log-likelihood, pooled over all tokens. Only comparable between models that share a tokenizer. | Implemented | es.perplexity |
| ROUGE-1 | Unigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example. | Implemented | es.rouge_1 |
| ROUGE-2 | Bigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example. | Implemented | es.rouge_2 |
| ROUGE-L | Longest common subsequence between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example. | Implemented | es.rouge_l |
| ROUGE-Lsum | Summary-level union LCS over sentences (one sentence per line) between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example. | Implemented | es.rouge_lsum |
| Self-BLEU | Average sentence BLEU of each prediction against all other predictions as references; higher means the outputs are more alike (less diverse). | Implemented | es.self_bleu |
| Sentence BLEU | BLEU computed for each example separately (exponential smoothing, effective order) and averaged; use for per-example scores, not for reporting corpus quality. | Implemented | es.sentence_bleu |
| TER | Translation edit rate: minimum number of insertions, deletions, substitutions and block shifts to turn the prediction into the closest reference, divided by the average reference length (Tercom, as in sacreBLEU). Lower is better. | Implemented | es.ter |
- BLEUImplemented
Corpus-level geometric mean of clipped n-gram precisions (n = 1..4) times a brevity penalty, computed exactly as sacreBLEU (13a tokenization, exponential smoothing).
es.bleu
- chrF / chrF++Implemented
F-beta score (β = 2) over character n-grams (n = 1..6), plus word uni- and bigrams for chrF++ (word_order=2); robust for morphologically rich languages. Computed exactly as sacreBLEU.
es.chrf
- CIDEr-DImplemented
Consensus with several references: cosine similarity of TF-IDF weighted n-gram vectors (n = 1..4), clipped to the reference counts and damped by a Gaussian length penalty; IDF comes from the references of the whole evaluated set (the coco-caption implementation).
es.cider
- Cross-entropy (negative log-likelihood)Implemented
Average negative log-probability the model assigns to the observed tokens, pooled over all tokens of all sequences.
es.cross_entropy
- Distinct-nImplemented
Number of distinct n-grams divided by the total number of n-grams across all predictions (whitespace tokens); a simple lexical-diversity indicator.
es.distinct_n
- MAUVEImplemented
Gap between the distribution of generated text and of human text: both sets of feature vectors are quantized together (L2 normalisation, PCA to 90% variance, k-means), and MAUVE is the area under the divergence frontier of the two histograms. 1 means indistinguishable.
es.mauve
- METEORImplemented
Unigram alignment between prediction and reference by exact, stemmed and (optionally) synonym matches; harmonic mean weighted towards recall, penalised for fragmented alignments. Best reference per example, averaged over examples (NLTK's meteor_score).
es.meteor
- PerplexityImplemented
Exponentiated token-average negative log-likelihood, pooled over all tokens. Only comparable between models that share a tokenizer.
es.perplexity
- ROUGE-1Implemented
Unigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.
es.rouge_1
- ROUGE-2Implemented
Bigram overlap between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.
es.rouge_2
- ROUGE-LImplemented
Longest common subsequence between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.
es.rouge_l
- ROUGE-LsumImplemented
Summary-level union LCS over sentences (one sentence per line) between prediction and reference after lowercasing and splitting on non-alphanumerics (as Google's rouge-score); F-measure averaged over examples, best reference per example.
es.rouge_lsum
- Self-BLEUImplemented
Average sentence BLEU of each prediction against all other predictions as references; higher means the outputs are more alike (less diverse).
es.self_bleu
- Sentence BLEUImplemented
BLEU computed for each example separately (exponential smoothing, effective order) and averaged; use for per-example scores, not for reporting corpus quality.
es.sentence_bleu
- TERImplemented
Translation edit rate: minimum number of insertions, deletions, substitutions and block shifts to turn the prediction into the closest reference, divided by the average reference length (Tercom, as in sacreBLEU). Lower is better.
es.ter
| Metric | Description | Status | API |
|---|---|---|---|
| BERTScore | Greedy matching of contextual token embeddings by cosine similarity: precision averages, over prediction tokens, the best similarity to any reference token; recall does the reverse; F1 combines them. Optional IDF weights and baseline rescaling as in the original implementation. | Implemented | es.bertscore |
| Embedding similarity | Similarity or distance between the sentence embedding of each prediction and of its reference: cosine similarity, or Euclidean / Manhattan distance. Values depend on the embedding model. | Implemented | es.embedding_similarity |
| Model-based score (COMET, BLEURT, BARTScore, AlignScore, ...) | Scores from any learned evaluation model or judge you supply, per example, so they get the same confidence intervals, model comparison and reports as every other metric. | Implemented | es.model_score |
| MoverScore | One minus the earth mover's distance between the IDF-weighted token embeddings of prediction and reference, with Euclidean transport cost between L2-normalised embeddings (MoverScore v2 style). | Implemented | es.moverscore |
- BERTScoreImplemented
Greedy matching of contextual token embeddings by cosine similarity: precision averages, over prediction tokens, the best similarity to any reference token; recall does the reverse; F1 combines them. Optional IDF weights and baseline rescaling as in the original implementation.
es.bertscore
- Embedding similarityImplemented
Similarity or distance between the sentence embedding of each prediction and of its reference: cosine similarity, or Euclidean / Manhattan distance. Values depend on the embedding model.
es.embedding_similarity
Scores from any learned evaluation model or judge you supply, per example, so they get the same confidence intervals, model comparison and reports as every other metric.
es.model_score
- MoverScoreImplemented
One minus the earth mover's distance between the IDF-weighted token embeddings of prediction and reference, with Euclidean transport cost between L2-normalised embeddings (MoverScore v2 style).
es.moverscore