Evaluating Machine Learning Models
Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.
MCP get_skill({ skillId: "model-evaluation-suite-859a0c30" })Use this skill with your agent
Create a free account and connect via MCP
# Model Evaluation Suite Use this skill when the model exists and the question is whether it is good enough. ## Overview This skill focuses on choosing and interpreting the right evaluation metrics for the problem, then comparing candidate models or thresholds. ## When to Use This Skill - Comparing candidate models with consistent metrics - Reviewing precision/recall/F1/AUC, regression error, calibration, or ranking quality - Stress-testing validation strategy before deployment or publication ## Not For / Boundaries - Building the training pipeline itself: use `scikit-learn` for classical modeling or `ml-pipeline-workflow` for end-to-end workflow ownership - Engineering features: use `preprocessing-data-with-automated-pipelines` - Checking train/test contamination: use `ml-data-leakage-guard` ## Typical Outputs - Metric suite recommendations - Model comparison tables - Notes on threshold tradeoffs, calibration, and validation weaknesses ## Related Skills - `scikit-learn` for class-level error breakdowns and confusion matrices - `scientific-reporting` when the evaluation must become a deliverable
Related Skills
More skills in Data, AI & Research
Ablation Planner
Use when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
Ablation Planner
Use when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
About
Provides information about the bitwize-music plugin, its version, and its creator. Use when the user asks about the plugin, its purpose, version, or capabilities.
Ab Test Analysis
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.
Academic Search
Search and analyze academic literature. Find papers, understand research methodologies, and synthesize academic findings for research projects.
Adaptyv
How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
Explore Other Categories
Skills from other categories with shared topics
⚙️ MuAPI Platform Utilities
Setup and utility scripts for muapi.ai — configure API keys, test connectivity, and poll for async generation results
✏️ MuAPI Media Editing & Enhancement
Edit and enhance images and videos with AI via muapi.ai — prompt-based editing, upscaling, background removal, face swap, lipsync, video effects, and more
🍌 Nano-Banana Expert Skill (Gemini 3 Style)
Reasoning-driven image generation using structured creative briefs (Gemini 3 style) — generates high-fidelity images via muapi.ai with logic-based prompting