Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluate, select, and contract with vendors and SaaS tools. Use this skill when comparing alternatives, running an RFP, scoring vendors against criteria, negotiating contracts, planning a switch, or assessing a vendor's risk. Triggers on vendor evaluation, RFP, vendor selection, build vs buy, SaaS evaluation, vendor scorecard, vendor comparison, contract negotiation, vendor switch, procurement. Also triggers when a renewal is coming up or when a tool isn't meeting expectations.
.claude/skills/rampstackco-vendor-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 95% | 0% |
Pick the right tool or service, negotiate fair terms, and avoid the lock-in traps. Stack-agnostic. Applies to SaaS, infrastructure providers, agencies, and any external dependency.
cost-optimization)A structured vendor evaluation. Skip phases at your peril.
Before looking at vendors, define what you actually need.
The temptation: skip this and start demoing. Vendors are happy to show off; you end up choosing what looks shiny rather than what fits.
Before evaluating vendors, decide whether you should build instead.
Build when:
Buy when:
Most teams over-build. The rule of thumb: buy unless there's a strong reason to build. Then question even that strong reason.
Cast a wider net than feels comfortable, then narrow.
Sources:
Cast wide first. Aim for 5-8 candidates. Then narrow to 2-4 finalists for deep evaluation.
Use a scorecard. Without one, you'll be swayed by demo theatrics or who has the friendliest sales rep.
Scorecard dimensions (weight by your situation):
Functional fit (40%): Does it do what you need? Edge cases handled? UX quality. Workflow fit.
Technical fit (15%): Integration with your stack. API quality and completeness. Data export and portability. Performance at your scale. Self-hosted, hybrid, or SaaS-only.
Operational fit (10%): Onboarding effort. Training and adoption. Documentation quality. Support quality (test by submitting a ticket). SLAs.
Security and compliance (10%): SOC 2, ISO 27001, HIPAA, etc., as applicable. Data residency. Encryption at rest and in transit. Access controls and audit logs. Penetration test results (ask). Subprocessors.
Vendor health (10%): Years in business. Funding and runway (or revenue if private). Customer base size and similar customers. Public references. Roadmap visibility.
Cost (10%): License or subscription cost. Implementation and onboarding cost. Training cost. Integration cost. Opportunity cost (in-house resource time). Switching cost (in case of failure).
Lock-in risk (5%): Data export quality. Standard formats vs proprietary. Migration paths to alternatives. Open standards alignment. Contract escape clauses.
Score each finalist 1-5 on each dimension. Multiply by weight. Sum.
The score isn't gospel. It surfaces the tradeoffs.
Most enterprise contracts are negotiable. Most aren't negotiated.
What's negotiable:
Common negotiation moves:
What to avoid:
Write a one-page brief: what we need, why, success criteria, constraints, stakeholders.
Honestly answer the build/buy question. Document the rationale.
Wider net first, narrowed via desk research:
Eliminate obvious misfits. Land on 2-4 finalists.
For each finalist:
Don't be charmed by the polished demo. Try it with your real workflow.
Critical for any vendor handling sensitive data:
This can take weeks for enterprise vendors. Start early.
Apply the scorecard. Do this collaboratively with stakeholders.
The scoring conversation matters more than the final number. It surfaces disagreement (one person scored UX 5, another scored 2: why?).
With the apparent winner:
Contract signing is the start, not the end. Plan:
Record:
This is gold for the next renewal or the next similar evaluation.
Skipping the needs definition. Demoing first. Buying what's shiny. Realizing 6 months in that the actual need wasn't met.
Single-source decisions. Talking to one vendor; deciding. No comparison. Probably overpaying or under-fitting.
Charisma-driven decisions. Buying based on the sales rep's likability. The product is what you'll use for years; the rep won't be there.
Reference calls that the vendor curated. Of course their references love them. Find references the vendor didn't suggest.
Glossing over security. Security review skipped because of timeline pressure. Then a breach. Slow down or accept the risk explicitly.
Demos that don't match the use case. Their default demo, not yours. Always do a use-case demo.
Trial that doesn't simulate real usage. A trial with synthetic data tells you the product works in synthetic conditions. Use real (or close to real) data.
Negotiating only on price. Terms, SLAs, and exit clauses matter more for long-term satisfaction than 5% price.
Auto-renewal without notice tracking. Renewal happens; rate goes up 15%. No one was watching. Track renewals; review with notice.
Lock-in without exit plan. Tightly integrating into a vendor's proprietary surface. When you want to leave, you can't. Plan exit at the start.
Multi-year contract for an unproven vendor. Save the multi-year for vendors you trust. New vendor: shorter term, evaluate after.
No internal champion. Tool selected; no one drives adoption. Tool sits unused. Identify the champion before signing.
Negotiating after a verbal commitment. "Yes, we want to buy" means they have less reason to negotiate. Keep options open until terms are settled.
Ignoring red flags in security review. Vendor's security responses are evasive or incomplete. Treat as a no.
Comparing apples to oranges. Vendors price differently (per user, per usage, flat). Build a comparable cost model at your scale.
A vendor evaluation document includes:
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/evaluation-rubric.md - Scoring template with weighted dimensions, 1-5 scale criteria for each dimension, and a worked vendor-comparison example.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 34,782 | 36,468 | +5% | 1 | 1 | 0% | 6,277 | 8,785 | +40% | 0 | 0 | — |
case-02 | fail→pass | 37,571 | 35,681 | -5% | 1 | 1 | 0% | 6,277 | 8,785 | +40% | 0 | 0 | — |
case-03 | fail→pass | 35,987 | 33,420 | -7% | 1 | 1 | 0% | 6,262 | 8,770 | +40% | 0 | 0 | — |
case-04 | fail→fail | 18,393 | 20,948 | +14% | 1 | 1 | 0% | 3,185 | 6,037 | +90% | 0 | 0 | — |
case-05 | fail→pass | 13,826 | 9,874 | -29% | 1 | 1 | 0% | 2,518 | 4,073 | +62% | 0 | 0 | — |
case-06 | fail→fail | 60,048 | 49,414 | -18% | 1 | 1 | 0% | 4,623 | 3,377 | -27% | 0 | 0 | — |
case-07 | pass→pass | 16,347 | 15,268 | -7% | 1 | 1 | 0% | 2,390 | 4,694 | +96% | 0 | 0 | — |
case-08 | fail→fail | 17,384 | 18,706 | +8% | 1 | 1 | 0% | 2,628 | 5,235 | +99% | 0 | 0 | — |
case-09 | fail→pass | 15,763 | 10,181 | -35% | 1 | 1 | 0% | 2,341 | 4,001 | +71% | 0 | 0 | — |
case-10 | fail→pass | 14,469 | 12,116 | -16% | 1 | 1 | 0% | 2,196 | 4,288 | +95% | 0 | 0 | — |
case-11 | fail→pass | 13,523 | 7,010 | -48% | 1 | 1 | 0% | 2,026 | 3,596 | +77% | 0 | 0 | — |
case-12 | pass→pass | 16,267 | 18,141 | +12% | 1 | 1 | 0% | 2,352 | 5,174 | +120% | 0 | 0 | — |
case-13 | fail→pass | 16,052 | 11,869 | -26% | 1 | 1 | 0% | 2,320 | 4,190 | +81% | 0 | 0 | — |
case-14 | fail→pass | 15,558 | 11,200 | -28% | 1 | 1 | 0% | 2,397 | 4,160 | +74% | 0 | 0 | — |
case-15 | fail→pass | 13,918 | 13,663 | -2% | 1 | 1 | 0% | 2,045 | 4,386 | +114% | 0 | 0 | — |
case-16 | pass→pass | 14,047 | 6,495 | -54% | 1 | 1 | 0% | 1,877 | 3,381 | +80% | 0 | 0 | — |
case-17 | fail→fail | 12,193 | 9,627 | -21% | 1 | 1 | 0% | 1,815 | 4,002 | +120% | 0 | 0 | — |
case-18 | fail→pass | 15,854 | 8,882 | -44% | 1 | 1 | 0% | 2,482 | 3,846 | +55% | 0 | 0 | — |
case-19 | fail→pass | 16,233 | 11,968 | -26% | 1 | 1 | 0% | 2,339 | 4,076 | +74% | 0 | 0 | — |
case-20 | pass→pass | 17,030 | 10,743 | -37% | 1 | 1 | 0% | 2,539 | 3,987 | +57% | 0 | 0 | — |
case-21 | fail→pass | 17,562 | 6,909 | -61% | 1 | 1 | 0% | 2,583 | 3,525 | +36% | 0 | 0 | — |
case-22 | fail→pass | 16,436 | 13,530 | -18% | 1 | 1 | 0% | 2,355 | 4,466 | +90% | 0 | 0 | — |
case-23 | fail→pass | 18,332 | 16,547 | -10% | 1 | 1 | 0% | 2,546 | 4,747 | +86% | 0 | 0 | — |
case-24 | fail→pass | 18,807 | 15,370 | -18% | 1 | 1 | 0% | 2,648 | 4,635 | +75% | 0 | 0 | — |
case-25 | pass→pass | 13,160 | 10,164 | -23% | 1 | 1 | 0% | 1,909 | 4,011 | +110% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +60 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.