Structured output and tools
Tool-call precision, recall and F1
Implemented
structured-output.tool_call_f1Definition
Matching of predicted to expected calls (same tool, and same arguments unless match='name'), pooled over examples: precision over predicted calls, recall over expected calls.
Formula
F1 = 2PR/(P+R), P = matched/predicted, R = matched/expected
Range: [0, 1]
Inputs and outputs
- expected_calls: expected tool calls per example
- predicted_calls: predicted tool calls per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.tool_call_f1(expected_calls, predicted_calls)References
- Patil SG, Mao H, Yan F, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.