Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluate a new tool before it joins the stack — the problem-first framing (tools answer needs, not demos), the trial designed with success criteria upfront, the stack-fit check (integration, overlap, the tool-sprawl tax), and the security/data review sized to the stakes. Use when asked should we buy this tool, evaluate this software for the team, we have three tools that do this already, or run a proper trial before committing. Produces the need statement, the trial design with pre-set criteria,
.claude/skills/mohitagw15856-tool-procurement-eval/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 71% | 0% |
Tools enter stacks backwards: someone sees a demo, gets excited, and the "evaluation" becomes a justification ritual (vendor-comparison-matrix fights this at the compare stage; this skill fights it at the door). The forward order: the need stated first (which problem, whose, costing what — purchase-justification arithmetic), the stack-fit check before the trial (does something we own already do this? — the overlap audit that kills half of tool requests honestly), the trial designed with success criteria written before day one (or the trial's warm feelings decide), and the security/data review sized to what the tool touches — because the fun tool that ingests customer data is a compliance decision wearing a productivity costume.
Ask for these if not provided:
Problem/owner/cost · must-haves · nice-to-haves — dated pre-demo]
Overlap: (owned tools × coverage %) · the config-gap finding if applicable · integration reality · the sprawl tax lines]
Duration · pilots (enthusiast + skeptic named) · the pre-written criteria · decision date · the security gate status before real data]
Adopt: owner/rollout/renewal-row · Decline: the logged reason · either way: in the decision log]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 33,727 | 31,343 | -7% | 1 | 1 | 0% | 3,945 | 4,988 | +26% | 0 | 0 | — |
case-02 | fail→pass | 39,148 | 27,595 | -30% | 1 | 1 | 0% | 5,355 | 5,253 | -2% | 0 | 0 | — |
case-03 | fail→pass | 25,687 | 25,374 | -1% | 1 | 1 | 0% | 3,265 | 4,694 | +44% | 0 | 0 | — |
case-04 | pass→pass | 18,499 | 17,736 | -4% | 1 | 1 | 0% | 1,726 | 3,172 | +84% | 0 | 0 | — |
case-05 | pass→pass | 16,900 | 18,143 | +7% | 1 | 1 | 0% | 1,700 | 2,879 | +69% | 0 | 0 | — |
case-06 | pass→pass | 16,959 | 17,909 | +6% | 1 | 1 | 0% | 1,529 | 3,209 | +110% | 0 | 0 | — |
case-07 | pass→pass | 18,369 | 17,806 | -3% | 1 | 1 | 0% | 1,888 | 2,974 | +58% | 0 | 0 | — |
case-08 | pass→pass | 22,198 | 20,489 | -8% | 1 | 1 | 0% | 2,787 | 3,700 | +33% | 0 | 0 | — |
case-09 | pass→pass | 22,887 | 23,532 | +3% | 1 | 1 | 0% | 2,519 | 3,972 | +58% | 0 | 0 | — |
case-10 | pass→pass | 22,599 | 22,719 | +1% | 1 | 1 | 0% | 2,576 | 4,052 | +57% | 0 | 0 | — |
case-11 | fail→fail | 41,363 | 25,356 | -39% | 1 | 1 | 0% | 3,033 | 4,416 | +46% | 0 | 0 | — |
case-12 | pass→pass | 14,805 | 15,867 | +7% | 1 | 1 | 0% | 1,619 | 3,051 | +88% | 0 | 0 | — |
case-13 | fail→pass | 35,163 | 24,187 | -31% | 1 | 1 | 0% | 2,729 | 3,779 | +38% | 0 | 0 | — |
case-14 | fail→pass | 16,370 | 26,883 | +64% | 1 | 1 | 0% | 2,581 | 4,425 | +71% | 0 | 0 | — |
case-15 | pass→pass | 19,763 | 24,537 | +24% | 1 | 1 | 0% | 2,202 | 4,028 | +83% | 0 | 0 | — |
case-16 | pass→pass | 21,760 | 20,985 | -4% | 1 | 1 | 0% | 2,514 | 3,699 | +47% | 0 | 0 | — |
case-17 | pass→pass | 14,472 | 22,453 | +55% | 1 | 1 | 0% | 1,918 | 3,483 | +82% | 0 | 0 | — |
case-18 | fail→pass | 23,128 | 25,603 | +11% | 1 | 1 | 0% | 2,340 | 3,791 | +62% | 0 | 0 | — |
case-19 | pass→pass | 20,961 | 11,163 | -47% | 1 | 1 | 0% | 1,937 | 2,917 | +51% | 0 | 0 | — |
case-20 | pass→pass | 18,071 | 18,819 | +4% | 1 | 1 | 0% | 1,932 | 3,411 | +77% | 0 | 0 | — |
case-21 | fail→pass | 20,389 | 22,292 | +9% | 1 | 1 | 0% | 2,418 | 3,789 | +57% | 0 | 0 | — |
case-22 | pass→pass | 17,412 | 15,458 | -11% | 1 | 1 | 0% | 2,387 | 2,919 | +22% | 0 | 0 | — |
case-23 | pass→pass | 20,606 | 20,474 | -1% | 1 | 1 | 0% | 2,269 | 3,659 | +61% | 0 | 0 | — |
case-24 | pass→pass | 16,676 | 9,812 | -41% | 1 | 1 | 0% | 1,872 | 2,856 | +53% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +29 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.