Skip to content
EvalSuite
Documentation menu

Code generation

CodeBLEU

Implementedcode.codebleu

Definition

Weighted sum of BLEU, keyword-weighted n-gram match, AST subtree match and data-flow match between generated and reference code (Ren et al.). The two n-gram terms follow the reference implementation exactly; the syntax and data-flow terms use Python's own ``ast`` instead of tree-sitter, so they are defined for Python source.

Formula

CodeBLEU = α·BLEU + β·BLEU_weight + γ·Match_ast + δ·Match_df

Range: [0, 1]

Inputs and outputs

  • references: see the signature of es.codebleu
  • predictions: see the signature of es.codebleu

Returns: MetricResult (value plus counts, intervals and breakdowns in params)

Assumptions

  • ``references``: one reference program (or a list of references) per prediction. Components are in ``params``. When no reference has any data flow the data-flow term counts as 1, as in the reference implementation.

Limitations

No metric-specific limitations are documented yet. Interpret the value alongside the task, data, and other metrics.

Python API

PythonSince v0.5.0
import evalsuite as es

es.codebleu(["def add(a, b):\n    return a + b"], ["def add(x, y):\n    return x + y"])

References

  1. Ren S, Guo D, Lu S, et al. CodeBLEU: a method for automatic evaluation of code synthesis. arXiv:2009.10297. 2020.

Implementation status