Factuality and QA
Exact match
Implemented
factuality.exact_matchDefinition
Share of predictions identical to a reference after normalization (SQuAD: lowercase, no punctuation or articles, single spaces); with several references, any match counts.
Formula
mean_i max_r [norm(prediction_i) = norm(reference_ir)]
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.exact_match(references, predictions) # SQuAD normalisationReferences
- Rajpurkar P, Zhang J, Lopyrev K, Liang P. SQuAD: 100,000+ questions for machine comprehension of text. EMNLP. 2016:2383-2392.