Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Compare vendors on a matrix that decides instead of decorates — the criteria weighted before the demos (so the shiny demo can't rewrite them), the evidence-based scoring with the marketing-vs-verified flags, the total-cost row that includes switching, and the reference-check questions that get honest answers. Use when asked compare these vendors/tools, build the selection matrix, the demo wowed us now what, or make this procurement decision defensible. Produces the weighted matrix, the scoring e
.claude/skills/mohitagw15856-vendor-comparison-matrix/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 40% | 0% |
Vendor selections get decided by the best demo and then justified by a matrix built afterward — criteria reverse-engineered to bless the favorite, which is how the shiny interface wins over the boring integration that actually mattered. The honest matrix inverts the order: criteria and weights are set before the demos (from the requirements, with the must-have/nice-to-have line drawn), scores cite evidence (the competitive-scan-lite claimed/observed/verified flags), the cost row is total cost (license + implementation + training + the switching cost both ways), and references get called with questions designed to pierce the happy-customer screen.
Ask for these if not provided:
Must-haves: pass/fail per vendor · Nice-to-haves × weights]
| Criterion (weight) | A] | B] | Incumbent] | |---|---|---|---| Every cell: score + 📢/👁/✅ flag · the flag-count summary per vendor]
Per vendor: license + implementation + training + integration + switch-in + exit-later = all-in over term]]
The script's piercing set · the off-list candidate found · findings per call]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 50,423 | 39,831 | -21% | 1 | 1 | 0% | 6,198 | 7,157 | +15% | 0 | 0 | — |
case-02 | fail→pass | 63,781 | 34,965 | -45% | 1 | 1 | 0% | 7,821 | 5,823 | -26% | 0 | 0 | — |
case-03 | fail→pass | 41,420 | 40,766 | -2% | 1 | 1 | 0% | 5,821 | 6,042 | +4% | 0 | 0 | — |
case-04 | pass→pass | 14,296 | 22,943 | +60% | 1 | 1 | 0% | 1,824 | 3,518 | +93% | 0 | 0 | — |
case-05 | pass→pass | 30,785 | 64,119 | +108% | 1 | 1 | 0% | 4,207 | 4,993 | +19% | 0 | 0 | — |
case-06 | pass→pass | 15,597 | 22,187 | +42% | 1 | 1 | 0% | 1,784 | 3,489 | +96% | 0 | 0 | — |
case-07 | fail→fail | 25,059 | 35,349 | +41% | 1 | 1 | 0% | 2,998 | 4,893 | +63% | 0 | 0 | — |
case-08 | fail→pass | 24,783 | 29,971 | +21% | 1 | 1 | 0% | 2,531 | 4,519 | +79% | 0 | 0 | — |
case-09 | fail→pass | 17,112 | 14,800 | -14% | 1 | 1 | 0% | 1,980 | 2,774 | +40% | 0 | 0 | — |
case-10 | fail→pass | 40,970 | 23,809 | -42% | 1 | 1 | 0% | 7,210 | 4,635 | -36% | 0 | 0 | — |
case-11 | fail→pass | 25,052 | 26,747 | +7% | 1 | 1 | 0% | 2,874 | 4,794 | +67% | 0 | 0 | — |
case-12 | pass→pass | 25,342 | 23,958 | -5% | 1 | 1 | 0% | 2,906 | 4,137 | +42% | 0 | 0 | — |
case-13 | fail→pass | 24,349 | 27,594 | +13% | 1 | 1 | 0% | 2,635 | 3,849 | +46% | 0 | 0 | — |
case-14 | pass→pass | 18,421 | 21,925 | +19% | 1 | 1 | 0% | 2,478 | 3,692 | +49% | 0 | 0 | — |
case-15 | pass→pass | 19,373 | 22,870 | +18% | 1 | 1 | 0% | 2,096 | 3,937 | +88% | 0 | 0 | — |
case-16 | fail→fail | 24,305 | 24,302 | -0% | 1 | 1 | 0% | 2,867 | 3,887 | +36% | 0 | 0 | — |
case-17 | fail→pass | 22,650 | 29,493 | +30% | 1 | 1 | 0% | 2,981 | 5,193 | +74% | 0 | 0 | — |
case-18 | pass→pass | 30,378 | 30,660 | +1% | 1 | 1 | 0% | 3,181 | 5,337 | +68% | 0 | 0 | — |
case-19 | fail→pass | 26,383 | 23,250 | -12% | 1 | 1 | 0% | 3,387 | 4,957 | +46% | 0 | 0 | — |
case-20 | fail→pass | 13,011 | 17,643 | +36% | 1 | 1 | 0% | 1,987 | 3,062 | +54% | 0 | 0 | — |
case-21 | fail→pass | 47,665 | 28,044 | -41% | 1 | 1 | 0% | 8,234 | 5,989 | -27% | 0 | 0 | — |
case-22 | pass→pass | 17,451 | 22,189 | +27% | 1 | 1 | 0% | 2,828 | 3,911 | +38% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.