Skip to content
EvalSuite
Documentation menu

API reference

evalsuite-python v0.4.0

Everything is available from the top-level package: import evalsuite as es. Run help(es.f1) or es.metric_info("classification.f1") for full documentation of any function.

Python API

FunctionsPurposeSince
es.evaluateEvery applicable metric in one call, inputs validated oncev0.1.0
es.accuracy, es.precision, es.recall, es.f1, es.fbeta, es.specificity, es.npv, es.balanced_accuracy, es.mcc, es.cohen_kappa, es.jaccard, es.hamming_loss, es.confusion_matrixLabel-based classification metricsv0.1.0
es.roc_auc, es.average_precision, es.log_loss, es.brier_score, es.top_k_accuracy, es.roc_curve, es.pr_curveScore-based classification metrics and curvesv0.1.0
es.calibration_curve, es.expected_calibration_errorCalibrationv0.1.0
es.mae, es.mse, es.rmse, es.r2, es.adjusted_r2, es.mape, es.smape, es.msle, es.rmsle, es.median_absolute_error, es.explained_variance, es.max_error, es.mean_bias_error, es.huber_loss, es.quantile_loss, es.rae, es.rseRegression metricsv0.1.0
es.compareModel comparison with intervals, paired tests and correctionsv0.1.0
es.bootstrap_ci, es.proportion_ci, es.accuracy_ci, es.roc_auc_ciConfidence intervalsv0.1.0
es.mcnemar_test, es.delong_test, es.paired_bootstrap_testPaired testsv0.1.0
es.cohens_d, es.hedges_g, es.cliffs_delta, es.adjust_pvaluesEffect sizes and multiple-testing correctionv0.1.0
es.classification_reportPer-class report tablev0.1.0
es.plot.roc, es.plot.pr, es.plot.calibration, es.plot.confusion_matrix, es.plot.residuals, es.plot.comparisonFigures (needs the plot extra)v0.1.0
es.list_metrics, es.metric_infoMetric registryv0.1.0
es.sensitivity, es.ppv, es.lr_positive, es.lr_negative, es.diagnostic_odds_ratio, es.youden_j, es.net_benefit, es.diagnostic_report, es.decision_curveClinical evaluationv0.2.0
es.maximum_calibration_error, es.calibration_slope, es.calibration_intercept, es.hosmer_lemeshow, es.calibration_reportCalibrationv0.2.0
es.t_test, es.paired_t_test, es.mann_whitney_test, es.wilcoxon_test, es.kruskal_wallis_test, es.friedman_test, es.shapiro_wilk_test, es.chi_square_test, es.fisher_exact_test, es.cramers_vHypothesis tests and effect sizesv0.2.0
es.plot.decision_curveDecision curve figurev0.2.0
es.dice, es.iou, es.miou, es.pixel_accuracy, es.mean_pixel_accuracy, es.boundary_iou, es.hausdorff_distance, es.average_surface_distance, es.segmentation_confusion, es.per_image_scores, es.segmentation_reportSemantic segmentationv0.3.0
es.box_iou, es.detection_report, es.mean_average_precision, es.average_precision_detection, es.detection_pr_curve, es.from_cocoObject detection (COCO protocol)v0.3.0
es.plot.segmentation, es.plot.per_class, es.plot.detection_prVision figuresv0.3.0
es.bleu, es.sentence_bleu, es.chrf, es.ter, es.rouge_1, es.rouge_2, es.rouge_l, es.rouge_lsum, es.meteor, es.cider, es.perplexity, es.cross_entropy, es.distinct_n, es.self_bleu, es.mauve, es.text_reportText generationv0.4.0
es.bertscore, es.embedding_similarity, es.moverscore, es.model_scoreSemantic and learned metrics (your embeddings or scorer)v0.4.0
es.exact_match, es.token_f1, es.faithfulness, es.hallucination_rate, es.groundedness, es.citation_precision, es.citation_recall, es.claim_verification_accuracy, es.knowledge_consistency, es.answer_correctness, es.answer_relevance, es.abstention_accuracyFactuality and question answeringv0.4.0
es.win_rate, es.bradley_terry, es.elo_ratings, es.krippendorff_alpha, es.fleiss_kappa, es.judge_agreement, es.position_consistency, es.verbosity_bias, es.self_preference_bias, es.rubric_scoreLLM-as-a-judge and preferencesv0.4.0
es.pass_at_k, es.majority_vote_accuracy, es.benchmark_accuracy, es.extract_answerReasoningv0.4.0
es.precision_at_k, es.recall_at_k, es.hit_rate_at_k, es.mrr, es.mean_average_precision_at_k, es.ndcg_at_k, es.context_precision, es.context_recall, es.context_relevance, es.latency_summary, es.task_success_rate, es.failure_attributionRetrieval and RAGv0.4.0
es.json_validity, es.json_schema_compliance, es.validate_json_schema, es.extract_json, es.xml_validity, es.required_field_accuracy, es.tool_selection_accuracy, es.tool_argument_accuracy, es.tool_call_f1, es.api_call_success_rate, es.instruction_compliance_rate, es.constraint_satisfaction_rate, es.format_compliance, es.instruction_retention, es.extra_content_rateStructured output, tools, instructionsv0.4.0
es.plot.ratings, es.plot.win_matrix, es.plot.text_scoresLLM figuresv0.4.0

Common signature

Pythonv0.1.0
metric(
  y_true,
  y_pred,                # or y_prob for score-based metrics
  *,
  average="auto",        # binary, micro, macro, weighted, samples, None
  labels=None,
  pos_label=None,
  sample_weight=None,
  zero_division="warn",  # or 0, 1, np.nan
)

Every metric returns a MetricResult that behaves like a number (float(r), f"{r:.3f}", comparisons) and records the parameters used.

HTTP API

Live

The FastAPI service lives in backend/ of the website repository. It provides accounts, personal API keys and authenticated evaluation. Evaluation runs on the released EvalSuite package, and every response names the version in engine.label and engine.version. The API is hosted at https://evalsuite-api.onrender.com.

Authentication

  1. Create an account on the website. You must be 18 or over; registration asks for your name, date of birth, role, country and intended use (organisation is optional).
  2. Sign in and open the dashboard, then create your API key. Each account has one active key. It is shown once; only a hash is stored. To replace it, use Rotate key: the old key stops working immediately and a new one is issued.
  3. Send it with every request, either as X-API-Key: <key> or Authorization: Bearer <key>.

Keys can be revoked instantly from the dashboard. A key can call evaluation endpoints but cannot create or revoke other keys.

ShellLive
curl https://evalsuite-api.onrender.com/api/v1/evaluate \
-H "X-API-Key: $EVALSUITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "task": "binary-classification",
  "yTrue": [1, 0, 1, 1, 0],
  "yPred": [1, 0, 0, 1, 0],
  "yProb": [0.9, 0.2, 0.4, 0.8, 0.1],
  "metrics": ["classification.accuracy", "classification.roc_auc"],
  "confidence": {"method": "wilson", "level": 0.95, "nBootstrap": 1000, "randomState": 42}
}'
Pythonclient.pyLive
import os
import requests

response = requests.post(
  "https://evalsuite-api.onrender.com/api/v1/evaluate",
  headers={"X-API-Key": os.environ["EVALSUITE_API_KEY"]},
  json={
      "task": "regression",
      "yTrue": [3.1, 2.4, 5.0],
      "yPred": [2.9, 2.8, 4.6],
      "metrics": ["regression.mae", "regression.rmse"],
      "confidence": {"method": "none", "level": 0.95, "nBootstrap": 1000, "randomState": 0},
  },
  timeout=30,
)
response.raise_for_status()
print(response.json())

Endpoints

MethodPathAuthPurpose
GET/api/v1/healthnoneLiveness
GET/api/v1/health/readynoneReadiness (database and rate-limit store)
GET/api/v1/metricsnoneMetric registry
GET/api/v1/metrics/{metric_id}noneOne metric definition
POST/api/v1/auth/registernoneCreate an account
POST/api/v1/auth/loginnoneGet a short-lived session token
GET/api/v1/auth/mesessionCurrent account
DELETE/api/v1/auth/mesessionDelete the account and everything linked to it
POST/api/v1/auth/password/forgotnoneEmail a one-time reset link
POST/api/v1/auth/password/resetnoneSet a new password with that link
POST/api/v1/auth/password/changesessionChange password; other sessions end
GET, POST/api/v1/keyssessionList keys, or create the account's single active key (18+)
POST/api/v1/keys/{id}/rotatesessionRevoke the active key and issue its replacement
DELETE/api/v1/keys/{id}sessionRevoke a key
GET/api/v1/usagesessionRequest counts
POST/api/v1/downloads/requestnone (email required)Record an email for security notices and get a 10-minute signed download link
POST/api/v1/downloads/unsubscribeunsubscribe tokenStop security notices
POST/api/v1/evaluateAPI key or sessionRun an evaluation
POST/api/v1/plot, /report, /bootstrapPlanned

Sessions and passwords

Sessions are short-lived and tied to the password they were issued under: changing or resetting the password signs out every other session immediately. Reset links are single-use, expire after 30 minutes by default, and carry the token in the URL fragment so it never reaches server logs. The forgot-password endpoint answers identically whether or not an email is registered.

Limits and errors

Requests are limited to 1 MB, 5,000 observations and 2,000 bootstrap resamples, and evaluation is rate-limited per account (default 60 per minute; 429 with Retry-After). Every error has the shape {"error": {"code": "...", "message": "..."}}.

Data handling

Submitted data is processed in memory and never stored or logged. Only counters (time, endpoint, number of observations) are kept to show your usage. Do not send identifiable patient information.