Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assess whether experiment results are credible enough to influence product decisions. Use when checking false positive or false negative risk, underpowered metrics, suspiciously large lifts, replication needs, meta-analysis, stratified sampling, covariate adjustment, or whether A/B test insights should be trusted.
.claude/skills/hashgraph-online-trustworthy-experiment-insights/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -8% | 0% |
Use this skill to decide whether an experiment result is believable enough to shape a product or engineering decision. It focuses on false positives, false negatives, power, replication, meta-analysis, stratified sampling, covariate adjustment, and suspicious result review.
Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from Chapter 6 on false positives and negatives, meta-analysis, metric sensitivity, stratified random sampling, covariate adjustments, replication, longer runs, and statistical power.
Related skills:
ab-test-results-readout for standard experiment reporting.experiment-sensitivity-optimization for improving precision before orduring experiment design.
experiment-verification-monitoring for operational validity checks.| Need | Read | |------|------| | Insight-quality concepts | references/core/knowledge.md | | Credibility and follow-up rules | references/core/rules.md | | Result-review scenarios | references/core/examples.md | | Step-by-step credibility review | workflows/review-experiment-credibility.md |
weak prior, or contradiction with prior experiments.
or over-broad metric choice.
markdown# Experiment Insight Credibility Review ## Result Under Review [Experiment, metric, observed result, and proposed decision.] ## Credibility Assessment [Trust | Trust with caveats | Replicate | Extend | Investigate | Do not trust] ## Evidence | Check | Finding | Risk | |-------|---------|------| ## Follow-Up - Replication needed: - Longer run needed: - Meta-analysis/comparison: - Variance reduction opportunity: ## Decision Guidance [What decision can be made now, and what should wait.]
population, metric, design, and timing.
health first.
Other measured skills in the registry, with their headline benchmark lift.