Skip to content
EvalSuite
Documentation menu

Factuality and QA

Token F1

Implementedfactuality.token_f1

Definition

Harmonic mean of token precision and recall between the normalized prediction and reference (bag of words, SQuAD); best reference per example, averaged over examples.

Formula

F1 = 2·P·R/(P + R), P = |common|/|pred tokens|, R = |common|/|ref tokens|

Range: [0, 1]

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • predictions: one model output per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.token_f1(references, predictions)

References

  1. Rajpurkar P, Zhang J, Lopyrev K, Liang P. SQuAD: 100,000+ questions for machine comprehension of text. EMNLP. 2016:2383-2392.

Implementation status