Skip to content
EvalSuite
Documentation menu

Structured output and tools

Tool-call precision, recall and F1

Implementedstructured-output.tool_call_f1

Definition

Matching of predicted to expected calls (same tool, and same arguments unless match='name'), pooled over examples: precision over predicted calls, recall over expected calls.

Formula

F1 = 2PR/(P+R), P = matched/predicted, R = matched/expected

Range: [0, 1]

Inputs and outputs

  • expected_calls: expected tool calls per example
  • predicted_calls: predicted tool calls per example

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.tool_call_f1(expected_calls, predicted_calls)

References

  1. Patil SG, Mao H, Yan F, et al. The Berkeley Function Calling Leaderboard (BFCL): from tool use to agentic evaluation of large language models. ICML. 2025.

Implementation status