Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan the migration of an LLM feature from one model to another without breaking production. Use when a model is being deprecated, a newer model looks better or cheaper, or when asked how to upgrade models safely, run shadow traffic, or set rollback criteria for a model change. Produces a phased migration plan with eval gates, shadow/canary stages, prompt-adaptation notes, and rollback triggers. For choosing which model in the first place use model-selection-advisor.
.claude/skills/mohitagw15856-model-migration-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 24% | 0% |
A model swap changes every output of your feature at once. This skill plans the migration like the risky deploy it is: eval first, shadow second, canary third — with numbers, not vibes, deciding each promotion.
Ask for (if not already provided):
prompt-regression-suite) or at minimum golden examples; if none exist, phase 0 is building onePhase 0 — Baseline. Freeze a regression suite against the current model. Without a baseline, "the new model is fine" is unfalsifiable. Record current cost, latency (p50/p95), and quality scores.
Phase 1 — Offline eval. Run the suite against the target model with the prompt as-is, then with adapted prompts. Promotion criteria: pass rate ≥ baseline, no canary failures, cost/latency within budget. Expect to iterate here — most "model regressions" are prompt-fit issues.
Phase 2 — Shadow. Mirror a sample of real traffic to the new model; log, never serve. Compare distributions: refusal rate, output length, format-violation rate, judge scores on a sample. Duration: long enough to cover weekly traffic patterns.
Phase 3 — Canary. Serve the new model to 1-5]% of traffic behind a flag, tagged in analytics. Watch the same metrics plus user-visible signals (regenerate rate, thumbs-down, support tickets). Widen in steps; each step has the same promotion criteria.
Phase 4 — Full cutover + cleanup. 100% traffic, old model kept warm behind the flag for period], then removed. Update model pins everywhere (including the eval judge if it referenced the old model), and re-baseline the regression suite on the new model.
Between model generations, re-check: instruction-following strictness (newer models often follow the letter, exposing sloppy prompts), format compliance (JSON/markdown habits differ), verbosity defaults, refusal boundaries, tool-calling style, and system-prompt sensitivity. Adapt the prompt per model rather than writing to the lowest common denominator — keep per-model prompt versions if both run simultaneously.
Why now: driver + deadline]. Blast radius: traffic, audience, cost of a bad output].
| Phase | Gate to pass | Duration | Owner | |---|---|---|---| | 0 Baseline | suite frozen; cost/latency recorded | | | | 1 Offline eval | criteria] | | | | 2 Shadow | criteria] | | | | 3 Canary x]% → y]% | criteria] | | | | 4 Cutover + cleanup | criteria] | | |
Prompt adaptations found/expected: list]
Rollback: triggers numbers]; mechanism flag]; owner who].
Cost/latency forecast: current] → projected], at traffic].
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 50,372 | 23,797 | -53% | 1 | 1 | 0% | 4,856 | 5,335 | +10% | 0 | 0 | — |
case-02 | fail→fail | 30,739 | 22,811 | -26% | 1 | 1 | 0% | 6,090 | 5,447 | -11% | 0 | 0 | — |
case-03 | fail→pass | 34,047 | 29,169 | -14% | 1 | 1 | 0% | 6,258 | 6,054 | -3% | 0 | 0 | — |
case-04 | pass→pass | 21,559 | 17,428 | -19% | 1 | 1 | 0% | 3,133 | 3,995 | +28% | 0 | 0 | — |
case-05 | pass→pass | 12,160 | 10,455 | -14% | 1 | 1 | 0% | 2,200 | 2,792 | +27% | 0 | 0 | — |
case-06 | fail→fail | 13,388 | 14,024 | +5% | 1 | 1 | 0% | 2,432 | 3,583 | +47% | 0 | 0 | — |
case-07 | pass→pass | 13,271 | 11,708 | -12% | 1 | 1 | 0% | 2,335 | 3,231 | +38% | 0 | 0 | — |
case-08 | fail→fail | 12,062 | 11,911 | -1% | 1 | 1 | 0% | 2,108 | 3,211 | +52% | 0 | 0 | — |
case-09 | fail→pass | 14,605 | 11,698 | -20% | 1 | 1 | 0% | 2,541 | 3,109 | +22% | 0 | 0 | — |
case-10 | fail→fail | 10,349 | 11,077 | +7% | 1 | 1 | 0% | 1,932 | 3,040 | +57% | 0 | 0 | — |
case-11 | pass→pass | 13,704 | 12,357 | -10% | 1 | 1 | 0% | 2,317 | 3,360 | +45% | 0 | 0 | — |
case-12 | fail→pass | 12,706 | 12,672 | -0% | 1 | 1 | 0% | 2,282 | 3,444 | +51% | 0 | 0 | — |
case-13 | fail→pass | 17,984 | 15,241 | -15% | 1 | 1 | 0% | 3,103 | 3,863 | +24% | 0 | 0 | — |
case-14 | pass→pass | 12,271 | 8,992 | -27% | 1 | 1 | 0% | 2,127 | 2,599 | +22% | 0 | 0 | — |
case-15 | fail→pass | 13,662 | 13,507 | -1% | 1 | 1 | 0% | 2,580 | 3,484 | +35% | 0 | 0 | — |
case-16 | fail→pass | 10,684 | 7,945 | -26% | 1 | 1 | 0% | 1,756 | 2,434 | +39% | 0 | 0 | — |
case-17 | pass→pass | 9,987 | 9,247 | -7% | 1 | 1 | 0% | 1,773 | 2,707 | +53% | 0 | 0 | — |
case-18 | pass→pass | 12,036 | 9,849 | -18% | 1 | 1 | 0% | 2,188 | 2,831 | +29% | 0 | 0 | — |
case-19 | fail→pass | 11,572 | 9,301 | -20% | 1 | 1 | 0% | 2,056 | 2,685 | +31% | 0 | 0 | — |
case-20 | pass→pass | 18,823 | 17,239 | -8% | 1 | 1 | 0% | 3,336 | 4,328 | +30% | 0 | 0 | — |
case-21 | pass→pass | 15,004 | 12,436 | -17% | 1 | 1 | 0% | 2,723 | 3,381 | +24% | 0 | 0 | — |
case-22 | pass→pass | 15,713 | 16,729 | +6% | 1 | 1 | 0% | 2,997 | 4,006 | +34% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.