Structured output and tools
Multi-turn instruction retention
Implemented
structured-output.instruction_retentionDefinition
Share of later turns in which instructions given earlier in the conversation are still satisfied, pooled over conversations.
Formula
later-turn checks passed / later-turn checks
Range: [0, 1]
Inputs and outputs
- turn_checks: per conversation, checks in later turns
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.instruction_retention(turn_checks)References
- Kwan WC, Zeng X, Jiang Y, et al. MT-Eval: a multi-turn capabilities evaluation benchmark for large language models. EMNLP. 2024.