Agents and tool use
Planning accuracy / plan adherence
Implemented
agents.plan_adherenceDefinition
Agreement between the steps executed and a reference plan, in order: precision and recall of the longest common subsequence, their F1 (the value), and the exact-plan match rate.
Formula
P = LCS / |executed|, R = LCS / |plan|, F1 = 2PR / (P + R)
Range: [0, 1]
Inputs and outputs
- plans: see the signature of es.plan_adherence
- executed: see the signature of es.plan_adherence
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``plans`` and ``executed``: per task, a list of step identifiers (tool names, action labels, ...).
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.plan_adherence([["search", "read", "answer"]], [["search", "answer"]])References
- Ma C, Zhang J, Zhu Z, et al. AgentBoard: an analytical evaluation board of multi-turn LLM agents. NeurIPS Datasets and Benchmarks. 2024.