Text generation
CIDEr-D
Implemented
text-generation.ciderDefinition
Consensus with several references: cosine similarity of TF-IDF weighted n-gram vectors (n = 1..4), clipped to the reference counts and damped by a Gaussian length penalty; IDF comes from the references of the whole evaluated set (the coco-caption implementation).
Formula
10 · mean_n mean_refs [Σ min(g_h, g_r)·g_r / (‖g_h‖‖g_r‖)] · exp(−Δ²/2σ²)
Range: [0, 10]
Inputs and outputs
- references: one reference string (or a list of references) per example
- predictions: one model output per example
Returns: MetricResult (float, or per-example array with average=None)
Assumptions
No assumptions beyond valid, aligned inputs of the documented types.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.cider(caption_references, captions)References
- Vedantam R, Zitnick CL, Parikh D. CIDEr: consensus-based image description evaluation. CVPR. 2015:4566-4575.