Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan evaluation strategies for machine-learning product changes. Use when deciding between offline evaluation, interleaving, online A/B tests, multi-armed bandits, or model filtering for ranking, recommendation, search, personalization, or other ML-powered user experiences.
.claude/skills/hashgraph-online-ml-experiment-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -42% | 0% |
| case-10 | ✓→✗ | ▼ Worse | 19% | 0% |
| case-20 | ✓→✓ | = Same ✓ | 20% | 0% |
Use this skill to choose how to evaluate machine-learning product changes before they consume live experiment traffic or affect users. It focuses on offline evaluation, offline-online correlation, interleaving, model filtering, and when classic A/B testing or adaptive strategies are justified.
Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from Chapter 4 on offline evaluation, offline-online correlation, multi-armed bandits, and interleaving for rankers.
Related skills:
experiment-sensitivity-optimization for reducing live variants and traffic.adaptive-experimentation-strategy for bandits and dynamic allocation.ab-test-design-brief for standard online A/B test planning.| Need | Read | |------|------| | ML evaluation concepts | references/core/knowledge.md | | Selection and validation rules | references/core/rules.md | | Evaluation strategy examples | references/core/examples.md | | Step-by-step evaluation plan | workflows/choose-ml-evaluation-strategy.md |
users.
needed and infrastructure can support it.
markdown# ML Evaluation Strategy ## Model Decision [What model or ranking decision must be made.] ## Recommended Evaluation Path [Offline only | Offline then A/B | Interleaving | A/B test | Adaptive strategy] ## Why - Product risk: - Offline signal available: - Online evidence needed: - Traffic or capacity constraint: ## Metrics | Metric | Offline/Online | Role | Concern | |--------|----------------|------|---------| ## Implementation Notes - Data needed: - Logging needed: - Correlation check: - Rollout guardrails:
understood.
where attribution can be logged.
observability, and operational ownership.
Other measured skills in the registry, with their headline benchmark lift.