Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Comprehensively identify all relevant methods for a task — 50 methods, 60 web searches budget
.claude/skills/yogsoth-ai-method-inventory/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✓→✓ | = Same ✓ | 110% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 127% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 77% | 0% |
| case-01 | ✗→✗ | = Same ✗ | -83% | 0% |
| case-02 | ✗→✗ | = Same ✗ | 15% | 0% |
Systematically discover and catalog every relevant method (published, preprint, industry) that has been evaluated on the target task. Combines leaderboard scraping, citation chain traversal, and keyword-based academic search to ensure no significant method is missed.
| Resource | Floor | Target | |----------|-------|--------| | Methods discovered | 30 | 50 | | Web searches | 40 | 60 | | Papers consulted | 20 | 40 |
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Methods discovered | 0 | 50 | BLOCKED |
| Web searches used | 0 | 60 | — |
| Papers consulted | 0 | 40 | — |
| Leaderboard sources | 0 | 5 | — |
| Citation chains traced | 0 | 10 | — |
</HARD-GATE>Cannot exit until methods_discovered >= 40 (80% of target).
json{ "task": "string", "domain": "string", "methods": [ { "name": "string", "aliases": ["string"], "year": 2024, "authors": "string", "venue": "string", "family": "string", "key_innovation": "string", "paper_id": "string", "source": "leaderboard|paper|citation_chain|preprint" } ], "coverage_assessment": { "leaderboard_sources_checked": 0, "citation_chains_traced": 0, "survey_papers_consulted": 0, "confidence": "high|medium|low" } }
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | leaderboard-harvesting | Systematically collect performance data from platforms and papers |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | method-discovery | Identify all relevant methods via literature, leaderboards, citation chains |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 41,814 | 18,143 | -57% | 1 | 1 | 0% | 8,323 | 1,452 | -83% | 0 | 0 | — |
case-02 | fail→fail | 48,201 | 46,813 | -3% | 1 | 1 | 0% | 8,301 | 9,508 | +15% | 0 | 0 | — |
case-03 | fail→fail | 26,857 | 23,389 | -13% | 1 | 1 | 0% | 4,802 | 1,961 | -59% | 0 | 0 | — |
case-04 | pass→pass | 27,205 | 46,618 | +71% | 1 | 1 | 0% | 4,277 | 8,974 | +110% | 0 | 0 | — |
case-05 | pass→pass | 18,929 | 46,106 | +144% | 1 | 1 | 0% | 3,952 | 8,973 | +127% | 0 | 0 | — |
case-06 | pass→pass | 30,638 | 42,572 | +39% | 1 | 1 | 0% | 5,067 | 8,954 | +77% | 0 | 0 | — |
case-07 | fail→fail | 31,429 | 13,832 | -56% | 1 | 1 | 0% | 5,376 | 1,294 | -76% | 0 | 0 | — |
case-08 | fail→fail | 20,822 | 15,707 | -25% | 1 | 1 | 0% | 4,436 | 1,388 | -69% | 0 | 0 | — |
case-09 | fail→fail | 25,279 | 10,313 | -59% | 1 | 1 | 0% | 3,972 | 1,396 | -65% | 0 | 0 | — |
case-10 | fail→fail | 22,068 | 53,402 | +142% | 1 | 1 | 0% | 3,679 | 8,963 | +144% | 0 | 0 | — |
case-11 | fail→fail | 37,626 | 17,398 | -54% | 1 | 1 | 0% | 7,171 | 1,194 | -83% | 0 | 0 | — |
case-12 | fail→fail | 18,020 | 17,928 | -1% | 1 | 1 | 0% | 2,484 | 1,454 | -41% | 0 | 0 | — |
case-13 | fail→fail | 28,715 | 56,148 | +96% | 1 | 1 | 0% | 5,349 | 8,954 | +67% | 0 | 0 | — |
case-14 | fail→fail | 33,587 | 19,413 | -42% | 1 | 1 | 0% | 5,997 | 1,565 | -74% | 0 | 0 | — |
case-15 | fail→fail | 18,618 | 8,620 | -54% | 1 | 1 | 0% | 3,886 | 1,297 | -67% | 0 | 0 | — |
case-16 | fail→fail | 28,641 | 13,455 | -53% | 1 | 1 | 0% | 4,826 | 1,176 | -76% | 0 | 0 | — |
case-17 | fail→fail | 20,838 | 18,358 | -12% | 1 | 1 | 0% | 4,094 | 1,207 | -71% | 0 | 0 | — |
case-18 | fail→fail | 49,042 | 13,373 | -73% | 1 | 1 | 0% | 5,186 | 1,352 | -74% | 0 | 0 | — |
case-19 | fail→fail | 24,343 | 30,241 | +24% | 1 | 1 | 0% | 4,949 | 1,385 | -72% | 0 | 0 | — |
case-20 | fail→fail | 32,045 | 17,900 | -44% | 1 | 1 | 0% | 5,905 | 1,360 | -77% | 0 | 0 | — |
case-21 | fail→fail | 17,848 | 34,504 | +93% | 1 | 1 | 0% | 2,872 | 5,436 | +89% | 0 | 0 | — |
case-22 | fail→fail | 12,746 | 14,823 | +16% | 1 | 1 | 0% | 3,454 | 1,505 | -56% | 0 | 0 | — |
case-23 | fail→fail | 27,371 | 16,819 | -39% | 1 | 1 | 0% | 4,254 | 1,295 | -70% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 8 counted toward the lift figure. The other 15 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 8 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.