Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build a quarterly supplier performance scorecard with a weighted grade and a clear escalate/develop/exit call. Use when asked to review supplier performance, prepare a quarterly business review for a vendor, score a supplier on OTIF and quality, or decide whether to escalate or exit a supplier. Produces a weighted scorecard with trend arrows, per-dimension evidence, corrective-action status, and a recommendation.
.claude/skills/mohitagw15856-supplier-scorecard/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 57% | 0% |
A supplier review that ends with "keep monitoring" is a meeting, not a decision. This skill turns delivery, quality, responsiveness, and cost data into a weighted quarterly grade with trend direction — and forces one of three outcomes: escalate, develop, or exit. It also audits whether last quarter's corrective actions actually closed, because a supplier who commits and doesn't deliver is telling you something.
Ask for these if not provided:
If history is missing, score the current quarter and label trends [no baseline — first scored quarter]. Infer reasonable dimension detail from a thin brief and label it [inferred].
Default weights (adjust for category — e.g., quality-critical items shift quality to 35%):
| Dimension | Weight | 5 (excellent) | 3 (acceptable) | 1 (failing) | |---|---|---|---|---| | Delivery (OTIF) | 30% | ≥98% | 93–95% | <90% | | Quality (PPM / defects) | 25% | ≤100 PPM, no escapes | ≤1,000 PPM | >5,000 PPM or a stop-ship | | Responsiveness | 15% | Same-day acknowledgment, proactive alerts | Meets agreed SLAs | Chased for answers | | Cost behavior | 15% | Beats index, delivers savings commitments | Tracks index | Above-index increases, missed commitments | | Corrective-action follow-through | 15% | All closed on time with verified effectiveness | Closed late but closed | Repeat findings, open past due |
Grade = Σ(score × weight) × 20 → 0–100. Bands: ≥85 Preferred · 70–84 Approved · 55–69 Conditional (development plan required) · <55 Exit-candidate.
Trend arrows: ↑ improved ≥5 points vs. prior quarter, → within ±5, ↓ declined ≥5. A ↓ trend in Conditional triggers escalation even if the band hasn't changed yet.
Recommendation logic: score AND trajectory AND strategic dependence. A 60-score sole-source supplier gets a development plan with executive sponsorship; a 60-score supplier with two qualified alternates gets a requalification/exit timeline. Say which case applies.
1. Summary — grade, band, trend, and the recommendation in two sentences.
2. Scorecard — table: Dimension | Weight | Metric this quarter | Prior quarter | Score (1–5) | Trend | Evidence.
3. Corrective-action audit — table: Action | Committed date | Status | Verified effective? Flag any repeat finding explicitly.
4. Cost detail — price moves vs. relevant index, savings pipeline status.
5. Recommendation — Escalate / Develop / Exit (or Maintain for Preferred), with the specific next step, owner, and review date. For Exit-candidates: transition risk, requalification lead time, and interim containment.
6. QBR talking points — 3–5 items: what to recognize, what to demand, what to decide.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 30,948 | 28,415 | -8% | 1 | 1 | 0% | 4,773 | 5,780 | +21% | 0 | 0 | — |
case-02 | fail→pass | 41,608 | 29,405 | -29% | 1 | 1 | 0% | 5,151 | 5,907 | +15% | 0 | 0 | — |
case-03 | fail→pass | 23,205 | 18,399 | -21% | 1 | 1 | 0% | 3,701 | 3,792 | +2% | 0 | 0 | — |
case-04 | fail→pass | 21,489 | 29,518 | +37% | 1 | 1 | 0% | 3,017 | 5,654 | +87% | 0 | 0 | — |
case-05 | pass→pass | 27,984 | 28,681 | +2% | 1 | 1 | 0% | 3,210 | 5,570 | +74% | 0 | 0 | — |
case-06 | pass→fail | 22,442 | 23,819 | +6% | 1 | 1 | 0% | 2,955 | 5,896 | +100% | 0 | 0 | — |
case-07 | fail→pass | 20,909 | 17,526 | -16% | 1 | 1 | 0% | 2,732 | 4,300 | +57% | 0 | 0 | — |
case-08 | pass→pass | 20,677 | 37,527 | +81% | 1 | 1 | 0% | 2,796 | 7,984 | +186% | 0 | 0 | — |
case-09 | fail→pass | 21,051 | 21,084 | +0% | 1 | 1 | 0% | 2,854 | 5,106 | +79% | 0 | 0 | — |
case-10 | fail→fail | 22,917 | 23,392 | +2% | 1 | 1 | 0% | 3,255 | 4,451 | +37% | 0 | 0 | — |
case-11 | fail→pass | 14,792 | 19,083 | +29% | 1 | 1 | 0% | 2,504 | 4,510 | +80% | 0 | 0 | — |
case-12 | fail→pass | 17,528 | 22,852 | +30% | 1 | 1 | 0% | 3,050 | 4,166 | +37% | 0 | 0 | — |
case-13 | fail→pass | 17,450 | 23,087 | +32% | 1 | 1 | 0% | 2,558 | 5,527 | +116% | 0 | 0 | — |
case-14 | pass→pass | 18,981 | 19,792 | +4% | 1 | 1 | 0% | 2,538 | 4,509 | +78% | 0 | 0 | — |
case-15 | fail→pass | 14,254 | 14,277 | +0% | 1 | 1 | 0% | 2,427 | 4,090 | +69% | 0 | 0 | — |
case-16 | pass→pass | 14,581 | 23,160 | +59% | 1 | 1 | 0% | 2,376 | 5,459 | +130% | 0 | 0 | — |
case-17 | fail→pass | 20,737 | 16,470 | -21% | 1 | 1 | 0% | 3,116 | 4,343 | +39% | 0 | 0 | — |
case-23 | pass→pass | 23,104 | 23,631 | +2% | 1 | 1 | 0% | 3,353 | 5,127 | +53% | 0 | 0 | — |
case-18 | fail→pass | 9,942 | 15,980 | +61% | 1 | 1 | 0% | 1,824 | 4,186 | +129% | 0 | 0 | — |
case-19 | fail→pass | 13,807 | 20,395 | +48% | 1 | 1 | 0% | 2,522 | 5,197 | +106% | 0 | 0 | — |
case-20 | fail→pass | 21,484 | 15,513 | -28% | 1 | 1 | 0% | 2,426 | 4,090 | +69% | 0 | 0 | — |
case-21 | fail→pass | 20,361 | 29,737 | +46% | 1 | 1 | 0% | 3,555 | 5,324 | +50% | 0 | 0 | — |
case-22 | pass→fail | 19,807 | 21,820 | +10% | 1 | 1 | 0% | 3,132 | 4,342 | +39% | 0 | 0 | — |
case-24 | pass→pass | 14,294 | 18,712 | +31% | 1 | 1 | 0% | 2,995 | 5,167 | +73% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +54 percentage points is the difference between those two pass rates over the 24 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.