Code generation
CodeBLEU
Implemented
code.codebleuDefinition
Weighted sum of BLEU, keyword-weighted n-gram match, AST subtree match and data-flow match between generated and reference code (Ren et al.). The two n-gram terms follow the reference implementation exactly; the syntax and data-flow terms use Python's own ``ast`` instead of tree-sitter, so they are defined for Python source.
Formula
CodeBLEU = α·BLEU + β·BLEU_weight + γ·Match_ast + δ·Match_df
Range: [0, 1]
Inputs and outputs
- references: see the signature of es.codebleu
- predictions: see the signature of es.codebleu
Returns: MetricResult (value plus counts, intervals and breakdowns in params)
Assumptions
- ``references``: one reference program (or a list of references) per prediction. Components are in ``params``. When no reference has any data flow the data-flow term counts as 1, as in the reference implementation.
Limitations
No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.
Python API
import evalsuite as es
es.codebleu(["def add(a, b):\n return a + b"], ["def add(x, y):\n return x + y"])References
- Ren S, Guo D, Lu S, et al. CodeBLEU: a method for automatic evaluation of code synthesis. arXiv:2009.10297. 2020.