Skip to content
EvalSuite
Documentation menu

CLI reference

Available in v0.1.0

Installing the package adds an evalsuite command (also available as python -m evalsuite). It reads CSV, TSV, Parquet or JSON files with one row per observation, or CSV from standard input with -.

Shellv0.1.0
evalsuite --version
evalsuite metrics                      # list every metric
evalsuite info classification.mcc      # definition, formula, range, references

evalsuite evaluate predictions.csv --y-true label --y-pred pred --y-prob prob
evalsuite evaluate predictions.csv --y-true label --y-pred pred -o report.html

evalsuite report predictions.csv --y-true label --y-pred pred -f markdown
evalsuite compare predictions.csv --y-true label --pred A=pred_a --pred B=pred_b
evalsuite plot roc predictions.csv --y-true label --y-prob prob -o roc.png

evalsuite benchmark --quick
Shellv0.2.0
evalsuite diagnostic predictions.csv --y-true label --y-pred pred
evalsuite calibration predictions.csv --y-true label --y-prob prob
evalsuite plot decision predictions.csv --y-true label --y-prob prob -o decision.png
evalsuite benchmark --suite clinical
Shellv0.3.0
evalsuite segmentation truth.npy pred.npy --num-classes 3 --ignore-index 255
evalsuite segmentation truth_dir/ pred_dir/ --spacing 0.8,0.8 --aggregate image --plot per_class.png
evalsuite detection instances_val.json detections.json -o detection.html --plot pr.png
evalsuite benchmark --suite vision
evalsuite benchmark                    # overall summary first, then every case in alphabetical order
Shellv0.4.0
evalsuite text outputs.csv --prediction output --reference reference
evalsuite text outputs.jsonl --prediction output --reference ref1 --reference ref2 --metrics bleu,chrf,rouge_l -o text.html
evalsuite benchmark --suite llm

evalsuite text reads CSV, TSV, Parquet or JSON-lines files with one example per row; repeat --reference for several reference columns (empty cells are skipped).

Segmentation masks are read from .npy or .npz files, or from folders of PNG masks with the vision extra. Detection reads COCO JSON: the ground-truth annotations file and a results list.

Behaviour

  • --help is available on every command.
  • The output format comes from --format (text, json, csv, markdown, latex, html) or from the extension of --output.
  • Bad input prints a one-line error naming the problem and exits with code 2; there are no tracebacks for user errors.
  • The CLI never executes shell commands or user-supplied code.