Skip to content
EvalSuite
Documentation menu

Structured output and tools

Multi-turn instruction retention

Implementedstructured-output.instruction_retention

Definition

Share of later turns in which instructions given earlier in the conversation are still satisfied, pooled over conversations.

Formula

later-turn checks passed / later-turn checks

Range: [0, 1]

Inputs and outputs

  • turn_checks: per conversation, checks in later turns

Returns: MetricResult (float, or per-example array with average=None)

Assumptions

No assumptions beyond valid, aligned inputs of the documented types.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.4.0
import evalsuite as es

es.instruction_retention(turn_checks)

References

  1. Kwan WC, Zeng X, Jiang Y, et al. MT-Eval: a multi-turn capabilities evaluation benchmark for large language models. EMNLP. 2024.

Implementation status