Skip to content
EvalSuite
Documentation menu

Reasoning

Benchmark accuracy

Implementedreasoning.benchmark_accuracy

Definition

Share of benchmark questions answered correctly after extracting the final answer with the benchmark's convention (GSM8K '####' or last number, MATH \boxed{}, multiple-choice letter) and comparing with the gold answer.

Formula

mean_i [extract(output_i) = gold_i]

Range: [0, 1]

Inputs and outputs

  • references: one reference string (or a list of references) per example
  • outputs: model outputs (strings)

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.benchmark_accuracy(gsm8k_solutions, outputs, style="gsm8k")

References

  1. Cobbe K, Kosaraju V, Bavarian M, et al. Training verifiers to solve math word problems. arXiv:2110.14168. 2021.
  2. Hendrycks D, Burns C, Kadavath S, et al. Measuring mathematical problem solving with the MATH dataset. NeurIPS Datasets and Benchmarks. 2021.

Implementation status