Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
.claude/skills/yogsoth-ai-validity-probing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-01 | ✓→✗ | ▼ Worse | -52% | 0% |
| case-02 | ✓→✗ | ▼ Worse | -48% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -50% | 0% |
| case-08 | ✓→✗ | ▼ Worse | -14% | 0% |
Deep investigation of construct validity for individual benchmarks. Determines whether a benchmark actually measures the capability it claims to measure, or whether high scores can be achieved through shortcuts, artifacts, or unrelated competencies.
Produce a construct validity assessment that identifies the gap between what a benchmark claims to measure and what it actually measures. Expose confounds, shortcuts, and alternative explanations for high performance.
| Resource | Floor | Target | |----------|-------|--------| | Benchmarks probed | 2 | 3 | | Papers read | 30 | 40 | | Web searches | 20 | 30 |
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks probed | 0 | 3 | PENDING |
| Papers fetched | 0 | 40 | PENDING |
| Papers read | 0 | 30 | PENDING |
| Web searches | 0 | 30 | PENDING |
| Construct validity assessments | 0 | 3 | PENDING |
| Artifact detection runs | 0 | 3 | PENDING |
| Alternative explanation catalogs | 0 | 3 | PENDING |
| Convergent validity checks | 0 | 3 | PENDING |
</HARD-GATE>Cannot exit until 80% of all targets met.
a. Collect the original benchmark paper + all critique/analysis papers b. Run construct-validity-assessment: map claimed capability to actual task requirements c. Run artifact-detection tactic: probe for shortcuts and spurious correlations d. Run metric-decomposition: identify what signals the metric actually rewards e. Check convergent validity: do models that score high also perform well on related tasks? f. Check discriminant validity: do models that score high fail on tasks they shouldn't if the capability were real?
yamlvalidity_report: benchmark_name: string claimed_capability: string actual_measurement: string # what it really measures validity_verdict: valid|partially_valid|questionable|invalid evidence_strength: strong|moderate|weak confounds_identified: - confound: string severity: high|medium|low evidence: string shortcuts_found: - shortcut: string exploit_method: string performance_gain: string convergent_validity: pass|partial|fail discriminant_validity: pass|partial|fail recommendations: list[string]
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | artifact-detection | Detect annotation artifacts and shortcuts in benchmarks | | evaluation-protocol-comparison | Compare implementation differences of same benchmark across papers |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | benchmark-synthesis | Produce final structured audit report | | construct-validity-assessment | Evaluate whether benchmark measures its claimed capability | | contamination-audit | Detect train-test data leakage and memorization artifacts | | metric-decomposition | Decompose composite metrics into constituent signals, analyze polarity and ceiling effects |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 16,086 | 55,809 | +247% | 1 | 1 | 0% | 1,651 | 9,156 | +455% | 0 | 0 | — |
case-10 | pass→pass | 22,776 | 28,611 | +26% | 1 | 1 | 0% | 2,746 | 5,141 | +87% | 0 | 0 | — |
case-01 | pass→fail | 26,816 | 21,322 | -20% | 1 | 1 | 0% | 3,803 | 1,829 | -52% | 0 | 0 | — |
case-02 | pass→fail | 22,611 | 21,421 | -5% | 1 | 1 | 0% | 3,714 | 1,947 | -48% | 0 | 0 | — |
case-03 | fail→fail | 38,590 | 13,903 | -64% | 1 | 1 | 0% | 5,497 | 1,939 | -65% | 0 | 0 | — |
case-04 | fail→pass | 17,199 | 22,693 | +32% | 1 | 1 | 0% | 2,385 | 4,171 | +75% | 0 | 0 | — |
case-16 | pass→pass | 20,424 | 26,769 | +31% | 1 | 1 | 0% | 2,492 | 4,313 | +73% | 0 | 0 | — |
case-05 | pass→fail | 19,000 | 7,445 | -61% | 1 | 1 | 0% | 3,838 | 1,937 | -50% | 0 | 0 | — |
case-06 | pass→pass | 14,939 | 82,186 | +450% | 1 | 1 | 0% | 2,374 | 5,181 | +118% | 0 | 0 | — |
case-07 | pass→pass | 21,871 | 42,540 | +95% | 1 | 1 | 0% | 2,533 | 6,986 | +176% | 0 | 0 | — |
case-08 | pass→fail | 19,577 | 48,197 | +146% | 1 | 1 | 0% | 2,132 | 1,826 | -14% | 0 | 0 | — |
case-09 | pass→pass | 22,906 | 40,600 | +77% | 1 | 1 | 0% | 2,199 | 8,241 | +275% | 0 | 0 | — |
case-11 | pass→pass | 19,896 | 11,843 | -40% | 1 | 1 | 0% | 2,369 | 1,961 | -17% | 0 | 0 | — |
case-12 | pass→pass | 14,041 | 24,729 | +76% | 1 | 1 | 0% | 1,421 | 4,150 | +192% | 0 | 0 | — |
case-13 | pass→fail | 24,964 | 13,290 | -47% | 1 | 1 | 0% | 3,067 | 2,252 | -27% | 0 | 0 | — |
case-14 | pass→pass | 13,966 | 53,598 | +284% | 1 | 1 | 0% | 2,228 | 9,029 | +305% | 0 | 0 | — |
case-15 | pass→pass | 21,564 | 41,073 | +90% | 1 | 1 | 0% | 2,647 | 6,860 | +159% | 0 | 0 | — |
case-17 | pass→fail | 13,665 | 21,998 | +61% | 1 | 1 | 0% | 1,341 | 1,797 | +34% | 0 | 0 | — |
case-18 | pass→fail | 23,374 | 12,354 | -47% | 1 | 1 | 0% | 2,837 | 2,070 | -27% | 0 | 0 | — |
case-19 | pass→pass | 13,850 | 61,251 | +342% | 1 | 1 | 0% | 2,413 | 9,160 | +280% | 0 | 0 | — |
case-20 | pass→pass | 7,059 | 19,435 | +175% | 1 | 1 | 0% | 1,218 | 3,337 | +174% | 0 | 0 | — |
case-21 | pass→fail | 17,434 | 21,238 | +22% | 1 | 1 | 0% | 1,880 | 1,677 | -11% | 0 | 0 | — |
case-23 | pass→pass | 13,413 | 18,702 | +39% | 1 | 1 | 0% | 1,520 | 4,301 | +183% | 0 | 0 | — |
case-24 | pass→fail | 15,002 | 20,293 | +35% | 1 | 1 | 0% | 2,204 | 2,047 | -7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 18 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -60 percentage points is the difference between those two pass rates over the 18 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.