Reasoning
Majority-vote accuracy
Implemented
reasoning.majority_vote_accuracyDefinition
Accuracy of the most frequent answer among several samples per question (self-consistency); ties go to the answer seen first.
Formula
mean_i [mode(samples_i) matches a reference]
Range: [0, 1]
Inputs and outputs
- references: one reference string (or a list of references) per example
- samples: sampled answers per question
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.majority_vote_accuracy(answers, samples)References
- Wang X, Wei J, Schuurmans D, et al. Self-consistency improves chain of thought reasoning in language models. ICLR. 2023.