Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: evidence-strength-based AHRQ PiCMe assessment — drive gap prioritization with the quality of literature evidence
.claude/skills/yogsoth-ai-evidence-based-prioritization/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 41% | 0% |
Evidence-strength-based prioritization: use the six dimensions of the AHRQ PiCMe framework to systematically assess the quality of the literature evidence behind each gap, surfacing the gaps where evidence is weakest and impact is greatest.
Core principle: a gap's priority depends not only on how important it is, but also on how weak the existing evidence is. The weaker the evidence and the higher the importance, the higher the priority.
The AHRQ PiCMe six-dimension assessment framework:
Final score = importance score × (1 − evidence sufficiency). Gaps with more sufficient evidence get lower priority (because others are already working on them).
Key insight: this framework naturally favors "neglected important problems" over "popular but already crowded problems."
| Tier | Number of gaps | PiCMe dimensions | Literature check | Final output | |------|---------|-----------|---------|---------| | S | 3–8 | all 6 dimensions | ≥2 supporting references per gap | ranking table + evidence-void report | | M | 9–15 | all 6 dimensions | ≥3 supporting references per gap | ranking table + evidence-void report + attack suggestions for top 3 gaps | | L | 16–20 | all 6 dimensions | ≥5 supporting references per gap | ranking table + detailed evidence map + attack suggestions for top 5 gaps |
gap-normalization SOP: standardize the gap format and extract the list of supporting references for each gapahrq-picme-assessment SOP: run the six-dimension assessment on each gapimportance-scoring SOP: assess importance independently (not influenced by evidence strength)scoring-matrix-construction tactic: build a gap × PiCMe-dimension matrixpriority-synthesis SOP: produce the final ranking + evidence-void summaryAfter each round, record:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | hypothesis-formation-scoring-matrix-construction | Tactic: orchestrate multi-dimensional scoring SOPs to build a comprehensive assessment matrix for all gaps |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | gap-normalization | SOP: Unify gaps from different sources into the standard GapRecord format | | priority-synthesis | SOP: synthesize all scoring data into a final gap priority list and attack-path suggestions |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 37,017 | 24,212 | -35% | 1 | 1 | 0% | 6,233 | 5,319 | -15% | 0 | 0 | — |
case-02 | fail→pass | 35,903 | 32,824 | -9% | 1 | 1 | 0% | 6,226 | 7,138 | +15% | 0 | 0 | — |
case-03 | fail→fail | 35,227 | 33,175 | -6% | 1 | 1 | 0% | 6,224 | 7,136 | +15% | 0 | 0 | — |
case-04 | pass→pass | 18,743 | 32,176 | +72% | 1 | 1 | 0% | 3,848 | 6,895 | +79% | 0 | 0 | — |
case-05 | pass→fail | 21,089 | 31,662 | +50% | 1 | 1 | 0% | 3,805 | 7,103 | +87% | 0 | 0 | — |
case-06 | pass→fail | 15,279 | 23,766 | +56% | 1 | 1 | 0% | 2,945 | 5,253 | +78% | 0 | 0 | — |
case-07 | pass→pass | 14,180 | 2,872 | -80% | 1 | 1 | 0% | 2,115 | 1,411 | -33% | 0 | 0 | — |
case-08 | fail→pass | 26,329 | 2,752 | -90% | 1 | 1 | 0% | 1,770 | 1,389 | -22% | 0 | 0 | — |
case-09 | pass→fail | 14,868 | 6,216 | -58% | 1 | 1 | 0% | 2,069 | 1,926 | -7% | 0 | 0 | — |
case-10 | pass→pass | 8,590 | 5,206 | -39% | 1 | 1 | 0% | 1,613 | 1,821 | +13% | 0 | 0 | — |
case-11 | pass→pass | 15,339 | 8,609 | -44% | 1 | 1 | 0% | 2,474 | 2,345 | -5% | 0 | 0 | — |
case-12 | pass→pass | 9,983 | 7,516 | -25% | 1 | 1 | 0% | 1,562 | 1,989 | +27% | 0 | 0 | — |
case-13 | pass→pass | 8,228 | 6,521 | -21% | 1 | 1 | 0% | 1,199 | 1,865 | +56% | 0 | 0 | — |
case-14 | pass→pass | 5,987 | 3,133 | -48% | 1 | 1 | 0% | 880 | 1,293 | +47% | 0 | 0 | — |
case-15 | fail→pass | 9,161 | 6,557 | -28% | 1 | 1 | 0% | 1,347 | 1,973 | +46% | 0 | 0 | — |
case-16 | pass→fail | 10,753 | 3,087 | -71% | 1 | 1 | 0% | 1,677 | 1,475 | -12% | 0 | 0 | — |
case-17 | pass→pass | 7,490 | 3,347 | -55% | 1 | 1 | 0% | 1,183 | 1,377 | +16% | 0 | 0 | — |
case-18 | fail→pass | 11,084 | 8,683 | -22% | 1 | 1 | 0% | 1,653 | 2,338 | +41% | 0 | 0 | — |
case-19 | fail→pass | 10,008 | 3,569 | -64% | 1 | 1 | 0% | 1,546 | 1,519 | -2% | 0 | 0 | — |
case-20 | pass→pass | 13,454 | 9,297 | -31% | 1 | 1 | 0% | 1,991 | 2,221 | +12% | 0 | 0 | — |
case-21 | fail→fail | 15,301 | 4,308 | -72% | 1 | 1 | 0% | 2,513 | 1,678 | -33% | 0 | 0 | — |
case-22 | fail→fail | 16,113 | 12,142 | -25% | 1 | 1 | 0% | 2,374 | 2,725 | +15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 21 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.