Install any skill in seconds. Free to start, no credit card required.
Get Started Free →When and how to escalate Contoso Outdoors customer-support tickets.
.claude/skills/microsoft-escalation-policy/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 6% | 0% |
You must follow this escalation policy on every conversation.
ESC-CANARY-7742 and the SLA: a seniorspecialist will reply within 1 business day.
escalated tickets — only the senior specialist can commit to one.
Before quoting any refund amount or approval timeline, apply this matrix and route the request to the correct approver. Quote the tier and approver exactly.
| Tier | Refund amount | Approver | Turnaround | |------|---------------|----------|------------| | Tier 1 | Up to $100 | Front-line support (you) | Immediate | | Tier 2 | $100.01 – $500 | Support team lead | Same business day | | Tier 3 | Over $500 | Senior specialist (escalation required) | 1 business day |
approve it yourself.
escalation regardless of amount.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,020 | 9,398 | -22% | 1 | 1 | 0% | 1,879 | 1,119 | -40% | 0 | 0 | — |
case-02 | fail→pass | 12,584 | 6,868 | -45% | 1 | 1 | 0% | 1,998 | 1,532 | -23% | 0 | 0 | — |
case-03 | fail→pass | 10,543 | 3,711 | -65% | 1 | 1 | 0% | 1,701 | 1,104 | -35% | 0 | 0 | — |
case-04 | fail→pass | 11,377 | 7,006 | -38% | 1 | 1 | 0% | 1,780 | 1,745 | -2% | 0 | 0 | — |
case-05 | fail→pass | 7,056 | 3,298 | -53% | 1 | 1 | 0% | 1,023 | 1,080 | +6% | 0 | 0 | — |
case-06 | pass→pass | 7,972 | 2,969 | -63% | 1 | 1 | 0% | 1,222 | 988 | -19% | 0 | 0 | — |
case-07 | fail→pass | 9,332 | 3,025 | -68% | 1 | 1 | 0% | 1,462 | 1,027 | -30% | 0 | 0 | — |
case-08 | fail→pass | 6,022 | 3,619 | -40% | 1 | 1 | 0% | 905 | 1,109 | +23% | 0 | 0 | — |
case-09 | fail→pass | 9,951 | 6,962 | -30% | 1 | 1 | 0% | 1,566 | 1,813 | +16% | 0 | 0 | — |
case-10 | fail→pass | 7,777 | 4,299 | -45% | 1 | 1 | 0% | 1,270 | 1,238 | -3% | 0 | 0 | — |
case-11 | pass→pass | 6,074 | 5,546 | -9% | 1 | 1 | 0% | 918 | 1,275 | +39% | 0 | 0 | — |
case-12 | fail→pass | 11,829 | 3,926 | -67% | 1 | 1 | 0% | 1,787 | 1,160 | -35% | 0 | 0 | — |
case-13 | fail→pass | 9,668 | 5,172 | -47% | 1 | 1 | 0% | 1,685 | 1,425 | -15% | 0 | 0 | — |
case-14 | fail→pass | 9,018 | 6,121 | -32% | 1 | 1 | 0% | 1,345 | 1,500 | +12% | 0 | 0 | — |
case-15 | fail→pass | 6,713 | 2,934 | -56% | 1 | 1 | 0% | 1,022 | 896 | -12% | 0 | 0 | — |
case-16 | fail→pass | 8,716 | 3,463 | -60% | 1 | 1 | 0% | 1,378 | 1,015 | -26% | 0 | 0 | — |
case-17 | fail→pass | 9,397 | 4,572 | -51% | 1 | 1 | 0% | 1,462 | 1,223 | -16% | 0 | 0 | — |
case-18 | fail→pass | 7,436 | 4,786 | -36% | 1 | 1 | 0% | 1,207 | 1,296 | +7% | 0 | 0 | — |
case-19 | fail→pass | 10,752 | 3,427 | -68% | 1 | 1 | 0% | 1,581 | 1,023 | -35% | 0 | 0 | — |
case-20 | pass→pass | 6,376 | 3,575 | -44% | 1 | 1 | 0% | 1,075 | 973 | -9% | 0 | 0 | — |
case-21 | pass→pass | 3,500 | 5,712 | +63% | 1 | 1 | 0% | 435 | 1,088 | +150% | 0 | 0 | — |
case-22 | pass→pass | 11,545 | 6,858 | -41% | 1 | 1 | 0% | 1,874 | 1,556 | -17% | 0 | 0 | — |
case-23 | pass→pass | 12,736 | 6,037 | -53% | 1 | 1 | 0% | 2,105 | 1,481 | -30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +74 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.