API reference
evalsuite-python v0.4.0Everything is available from the top-level package: import evalsuite as es. Run help(es.f1) or es.metric_info("classification.f1") for full documentation of any function.
Python API
| Functions | Purpose | Since |
|---|---|---|
es.evaluate | Every applicable metric in one call, inputs validated once | v0.1.0 |
es.accuracy, es.precision, es.recall, es.f1, es.fbeta, es.specificity, es.npv, es.balanced_accuracy, es.mcc, es.cohen_kappa, es.jaccard, es.hamming_loss, es.confusion_matrix | Label-based classification metrics | v0.1.0 |
es.roc_auc, es.average_precision, es.log_loss, es.brier_score, es.top_k_accuracy, es.roc_curve, es.pr_curve | Score-based classification metrics and curves | v0.1.0 |
es.calibration_curve, es.expected_calibration_error | Calibration | v0.1.0 |
es.mae, es.mse, es.rmse, es.r2, es.adjusted_r2, es.mape, es.smape, es.msle, es.rmsle, es.median_absolute_error, es.explained_variance, es.max_error, es.mean_bias_error, es.huber_loss, es.quantile_loss, es.rae, es.rse | Regression metrics | v0.1.0 |
es.compare | Model comparison with intervals, paired tests and corrections | v0.1.0 |
es.bootstrap_ci, es.proportion_ci, es.accuracy_ci, es.roc_auc_ci | Confidence intervals | v0.1.0 |
es.mcnemar_test, es.delong_test, es.paired_bootstrap_test | Paired tests | v0.1.0 |
es.cohens_d, es.hedges_g, es.cliffs_delta, es.adjust_pvalues | Effect sizes and multiple-testing correction | v0.1.0 |
es.classification_report | Per-class report table | v0.1.0 |
es.plot.roc, es.plot.pr, es.plot.calibration, es.plot.confusion_matrix, es.plot.residuals, es.plot.comparison | Figures (needs the plot extra) | v0.1.0 |
es.list_metrics, es.metric_info | Metric registry | v0.1.0 |
es.sensitivity, es.ppv, es.lr_positive, es.lr_negative, es.diagnostic_odds_ratio, es.youden_j, es.net_benefit, es.diagnostic_report, es.decision_curve | Clinical evaluation | v0.2.0 |
es.maximum_calibration_error, es.calibration_slope, es.calibration_intercept, es.hosmer_lemeshow, es.calibration_report | Calibration | v0.2.0 |
es.t_test, es.paired_t_test, es.mann_whitney_test, es.wilcoxon_test, es.kruskal_wallis_test, es.friedman_test, es.shapiro_wilk_test, es.chi_square_test, es.fisher_exact_test, es.cramers_v | Hypothesis tests and effect sizes | v0.2.0 |
es.plot.decision_curve | Decision curve figure | v0.2.0 |
es.dice, es.iou, es.miou, es.pixel_accuracy, es.mean_pixel_accuracy, es.boundary_iou, es.hausdorff_distance, es.average_surface_distance, es.segmentation_confusion, es.per_image_scores, es.segmentation_report | Semantic segmentation | v0.3.0 |
es.box_iou, es.detection_report, es.mean_average_precision, es.average_precision_detection, es.detection_pr_curve, es.from_coco | Object detection (COCO protocol) | v0.3.0 |
es.plot.segmentation, es.plot.per_class, es.plot.detection_pr | Vision figures | v0.3.0 |
es.bleu, es.sentence_bleu, es.chrf, es.ter, es.rouge_1, es.rouge_2, es.rouge_l, es.rouge_lsum, es.meteor, es.cider, es.perplexity, es.cross_entropy, es.distinct_n, es.self_bleu, es.mauve, es.text_report | Text generation | v0.4.0 |
es.bertscore, es.embedding_similarity, es.moverscore, es.model_score | Semantic and learned metrics (your embeddings or scorer) | v0.4.0 |
es.exact_match, es.token_f1, es.faithfulness, es.hallucination_rate, es.groundedness, es.citation_precision, es.citation_recall, es.claim_verification_accuracy, es.knowledge_consistency, es.answer_correctness, es.answer_relevance, es.abstention_accuracy | Factuality and question answering | v0.4.0 |
es.win_rate, es.bradley_terry, es.elo_ratings, es.krippendorff_alpha, es.fleiss_kappa, es.judge_agreement, es.position_consistency, es.verbosity_bias, es.self_preference_bias, es.rubric_score | LLM-as-a-judge and preferences | v0.4.0 |
es.pass_at_k, es.majority_vote_accuracy, es.benchmark_accuracy, es.extract_answer | Reasoning | v0.4.0 |
es.precision_at_k, es.recall_at_k, es.hit_rate_at_k, es.mrr, es.mean_average_precision_at_k, es.ndcg_at_k, es.context_precision, es.context_recall, es.context_relevance, es.latency_summary, es.task_success_rate, es.failure_attribution | Retrieval and RAG | v0.4.0 |
es.json_validity, es.json_schema_compliance, es.validate_json_schema, es.extract_json, es.xml_validity, es.required_field_accuracy, es.tool_selection_accuracy, es.tool_argument_accuracy, es.tool_call_f1, es.api_call_success_rate, es.instruction_compliance_rate, es.constraint_satisfaction_rate, es.format_compliance, es.instruction_retention, es.extra_content_rate | Structured output, tools, instructions | v0.4.0 |
es.plot.ratings, es.plot.win_matrix, es.plot.text_scores | LLM figures | v0.4.0 |
Common signature
metric(
y_true,
y_pred, # or y_prob for score-based metrics
*,
average="auto", # binary, micro, macro, weighted, samples, None
labels=None,
pos_label=None,
sample_weight=None,
zero_division="warn", # or 0, 1, np.nan
)Every metric returns a MetricResult that behaves like a number (float(r), f"{r:.3f}", comparisons) and records the parameters used.
HTTP API
LiveThe FastAPI service lives in backend/ of the website repository. It provides accounts, personal API keys and authenticated evaluation. Evaluation runs on the released EvalSuite package, and every response names the version in engine.label and engine.version. The API is hosted at https://evalsuite-api.onrender.com.
Authentication
- Create an account on the website. You must be 18 or over; registration asks for your name, date of birth, role, country and intended use (organisation is optional).
- Sign in and open the dashboard, then create your API key. Each account has one active key. It is shown once; only a hash is stored. To replace it, use Rotate key: the old key stops working immediately and a new one is issued.
- Send it with every request, either as
X-API-Key: <key>orAuthorization: Bearer <key>.
Keys can be revoked instantly from the dashboard. A key can call evaluation endpoints but cannot create or revoke other keys.
curl https://evalsuite-api.onrender.com/api/v1/evaluate \
-H "X-API-Key: $EVALSUITE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "binary-classification",
"yTrue": [1, 0, 1, 1, 0],
"yPred": [1, 0, 0, 1, 0],
"yProb": [0.9, 0.2, 0.4, 0.8, 0.1],
"metrics": ["classification.accuracy", "classification.roc_auc"],
"confidence": {"method": "wilson", "level": 0.95, "nBootstrap": 1000, "randomState": 42}
}'import os
import requests
response = requests.post(
"https://evalsuite-api.onrender.com/api/v1/evaluate",
headers={"X-API-Key": os.environ["EVALSUITE_API_KEY"]},
json={
"task": "regression",
"yTrue": [3.1, 2.4, 5.0],
"yPred": [2.9, 2.8, 4.6],
"metrics": ["regression.mae", "regression.rmse"],
"confidence": {"method": "none", "level": 0.95, "nBootstrap": 1000, "randomState": 0},
},
timeout=30,
)
response.raise_for_status()
print(response.json())Endpoints
| Method | Path | Auth | Purpose |
|---|---|---|---|
| GET | /api/v1/health | none | Liveness |
| GET | /api/v1/health/ready | none | Readiness (database and rate-limit store) |
| GET | /api/v1/metrics | none | Metric registry |
| GET | /api/v1/metrics/{metric_id} | none | One metric definition |
| POST | /api/v1/auth/register | none | Create an account |
| POST | /api/v1/auth/login | none | Get a short-lived session token |
| GET | /api/v1/auth/me | session | Current account |
| DELETE | /api/v1/auth/me | session | Delete the account and everything linked to it |
| POST | /api/v1/auth/password/forgot | none | Email a one-time reset link |
| POST | /api/v1/auth/password/reset | none | Set a new password with that link |
| POST | /api/v1/auth/password/change | session | Change password; other sessions end |
| GET, POST | /api/v1/keys | session | List keys, or create the account's single active key (18+) |
| POST | /api/v1/keys/{id}/rotate | session | Revoke the active key and issue its replacement |
| DELETE | /api/v1/keys/{id} | session | Revoke a key |
| GET | /api/v1/usage | session | Request counts |
| POST | /api/v1/downloads/request | none (email required) | Record an email for security notices and get a 10-minute signed download link |
| POST | /api/v1/downloads/unsubscribe | unsubscribe token | Stop security notices |
| POST | /api/v1/evaluate | API key or session | Run an evaluation |
| POST | /api/v1/plot, /report, /bootstrap | Planned |
Sessions and passwords
Sessions are short-lived and tied to the password they were issued under: changing or resetting the password signs out every other session immediately. Reset links are single-use, expire after 30 minutes by default, and carry the token in the URL fragment so it never reaches server logs. The forgot-password endpoint answers identically whether or not an email is registered.
Limits and errors
Requests are limited to 1 MB, 5,000 observations and 2,000 bootstrap resamples, and evaluation is rate-limited per account (default 60 per minute; 429 with Retry-After). Every error has the shape {"error": {"code": "...", "message": "..."}}.
Data handling
Submitted data is processed in memory and never stored or logged. Only counters (time, endpoint, number of observations) are kept to show your usage. Do not send identifiable patient information.