CLI reference
Available in v0.1.0Installing the package adds an evalsuite command (also available as python -m evalsuite). It reads CSV, TSV, Parquet or JSON files with one row per observation, or CSV from standard input with -.
evalsuite --version
evalsuite metrics # list every metric
evalsuite info classification.mcc # definition, formula, range, references
evalsuite evaluate predictions.csv --y-true label --y-pred pred --y-prob prob
evalsuite evaluate predictions.csv --y-true label --y-pred pred -o report.html
evalsuite report predictions.csv --y-true label --y-pred pred -f markdown
evalsuite compare predictions.csv --y-true label --pred A=pred_a --pred B=pred_b
evalsuite plot roc predictions.csv --y-true label --y-prob prob -o roc.png
evalsuite benchmark --quickevalsuite diagnostic predictions.csv --y-true label --y-pred pred
evalsuite calibration predictions.csv --y-true label --y-prob prob
evalsuite plot decision predictions.csv --y-true label --y-prob prob -o decision.png
evalsuite benchmark --suite clinicalevalsuite segmentation truth.npy pred.npy --num-classes 3 --ignore-index 255
evalsuite segmentation truth_dir/ pred_dir/ --spacing 0.8,0.8 --aggregate image --plot per_class.png
evalsuite detection instances_val.json detections.json -o detection.html --plot pr.png
evalsuite benchmark --suite vision
evalsuite benchmark # overall summary first, then every case in alphabetical orderevalsuite text outputs.csv --prediction output --reference reference
evalsuite text outputs.jsonl --prediction output --reference ref1 --reference ref2 --metrics bleu,chrf,rouge_l -o text.html
evalsuite benchmark --suite llmevalsuite text reads CSV, TSV, Parquet or JSON-lines files with one example per row; repeat --reference for several reference columns (empty cells are skipped).
Segmentation masks are read from .npy or .npz files, or from folders of PNG masks with the vision extra. Detection reads COCO JSON: the ground-truth annotations file and a results list.
Behaviour
--helpis available on every command.- The output format comes from
--format(text,json,csv,markdown,latex,html) or from the extension of--output. - Bad input prints a one-line error naming the problem and exits with code 2; there are no tracebacks for user errors.
- The CLI never executes shell commands or user-supplied code.