Factuality and QA
Token F1
Implemented
factuality.token_f1Definition
Harmonic mean of token precision and recall between the normalized prediction and reference (bag of words, SQuAD); best reference per example, averaged over examples.
Formula
F1 = 2·P·R/(P + R), P = |common|/|pred tokens|, R = |common|/|ref tokens|
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.token_f1(references, predictions)References
- Rajpurkar P, Zhang J, Lopyrev K, Liang P. SQuAD: 100,000+ questions for machine comprehension of text. EMNLP. 2016:2383-2392.