Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Choose the right LLM for a task by trading off quality, cost, latency, and constraints. Use when asked which model to use, whether to upgrade/downgrade a model, how to cut LLM costs without hurting quality, or to justify a model choice. Produces a recommendation with the decision criteria, a per-option comparison, a routing strategy (cheap-by-default, escalate when needed), and how to validate the choice with an eval.
.claude/skills/mohitagw15856-model-selection-advisor/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -34% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 21% | 0% |
The right model is rarely "the biggest one" or "the cheapest one" — it's the smallest model that clears the task's quality bar within its latency and cost budget, with a path to escalate the hard cases. This skill makes that trade-off explicit and defensible, and ties it to an eval so the choice is measured, not vibes.
Given "what model should I use for summarising support tickets?", deliver a concrete recommendation anyway — infer the task's difficulty, volume, and latency sensitivity, label the assumptions, and recommend. Never hand back "it depends" with no pick; give a default and the condition under which you'd change it.
Ask for these only if they aren't already provided (else infer and label):
1. Decision criteria — the 3–5 factors that actually decide it here, ranked (e.g. reasoning depth > latency > cost), with why.
2. Option comparison — the realistic candidates scored against the criteria. Keep it provider-agnostic in method; name a default family (e.g. the Claude family — a small/fast tier, a balanced tier, a frontier tier) and reason by tier, not a single hardcoded model, so the advice survives model releases.
| Option (tier) | Quality on this task | Latency | Relative cost | Fit | |---|---|---|---|---| | Small/fast | clears bar for easy cases | low | $ | default for the bulk | | Balanced | clears bar for most cases | med | $$ | when small misses | | Frontier | clears the hardest cases | higher | $$$ | escalation / eval judge |
3. Recommendation — the default model/tier, in one sentence, with the single reason.
4. Routing strategy — cheap-by-default with escalation: run the small tier first, detect low-confidence or hard cases (length, ambiguity, a validator/judge failing), and escalate those to a stronger tier. This usually beats picking one model for everything on both cost and quality.
5. Validation — how to confirm the choice: a small eval set scored per tier (pair with eval-rubric-designer and ai-eval-plan), and a cost/latency estimate at real volume (pair with llm-cost-latency-budget).
Model-selection practice — quality/cost/latency trade-offs, tiered routing with escalation, and eval-driven validation.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 18,840 | 22,523 | +20% | 1 | 1 | 0% | 2,982 | 3,874 | +30% | 0 | 0 | — |
case-01 | fail→fail | 20,609 | 20,175 | -2% | 1 | 1 | 0% | 3,875 | 4,252 | +10% | 0 | 0 | — |
case-02 | fail→fail | 21,502 | 18,052 | -16% | 1 | 1 | 0% | 4,005 | 4,168 | +4% | 0 | 0 | — |
case-03 | fail→fail | 23,762 | 25,325 | +7% | 1 | 1 | 0% | 4,193 | 4,064 | -3% | 0 | 0 | — |
case-05 | pass→pass | 26,449 | 29,496 | +12% | 1 | 1 | 0% | 4,581 | 7,164 | +56% | 0 | 0 | — |
case-06 | pass→pass | 12,902 | 18,327 | +42% | 1 | 1 | 0% | 2,902 | 4,896 | +69% | 0 | 0 | — |
case-07 | pass→pass | 21,245 | 15,946 | -25% | 1 | 1 | 0% | 3,836 | 3,740 | -3% | 0 | 0 | — |
case-08 | fail→pass | 17,065 | 19,006 | +11% | 1 | 1 | 0% | 2,953 | 4,422 | +50% | 0 | 0 | — |
case-09 | fail→pass | 31,503 | 16,839 | -47% | 1 | 1 | 0% | 6,195 | 4,058 | -34% | 0 | 0 | — |
case-10 | pass→fail | 18,900 | 17,211 | -9% | 1 | 1 | 0% | 3,384 | 4,187 | +24% | 0 | 0 | — |
case-11 | fail→pass | 21,486 | 15,663 | -27% | 1 | 1 | 0% | 3,941 | 3,754 | -5% | 0 | 0 | — |
case-12 | fail→pass | 17,227 | 15,556 | -10% | 1 | 1 | 0% | 3,093 | 3,989 | +29% | 0 | 0 | — |
case-13 | fail→fail | 19,657 | 17,303 | -12% | 1 | 1 | 0% | 3,617 | 4,089 | +13% | 0 | 0 | — |
case-14 | pass→pass | 29,966 | 19,886 | -34% | 1 | 1 | 0% | 5,610 | 4,514 | -20% | 0 | 0 | — |
case-15 | fail→fail | 24,584 | 19,790 | -20% | 1 | 1 | 0% | 4,017 | 4,163 | +4% | 0 | 0 | — |
case-16 | fail→pass | 16,243 | 15,220 | -6% | 1 | 1 | 0% | 3,050 | 3,681 | +21% | 0 | 0 | — |
case-17 | fail→fail | 19,807 | 14,949 | -25% | 1 | 1 | 0% | 3,584 | 3,604 | +1% | 0 | 0 | — |
case-18 | fail→fail | 19,877 | 16,508 | -17% | 1 | 1 | 0% | 3,358 | 3,823 | +14% | 0 | 0 | — |
case-19 | fail→pass | 19,694 | 17,529 | -11% | 1 | 1 | 0% | 3,225 | 4,230 | +31% | 0 | 0 | — |
case-20 | pass→pass | 26,992 | 17,675 | -35% | 1 | 1 | 0% | 5,019 | 3,954 | -21% | 0 | 0 | — |
case-21 | fail→fail | 19,352 | 21,179 | +9% | 1 | 1 | 0% | 3,201 | 4,651 | +45% | 0 | 0 | — |
case-22 | pass→pass | 16,285 | 18,657 | +15% | 1 | 1 | 0% | 2,797 | 4,050 | +45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.