Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Score a subcontractor's performance across schedule reliability, quality, safety, paperwork, and change-order behaviour with weighted anchors. Use when asked to evaluate a sub, build a subcontractor scorecard, decide whether to rebid or rehire a trade, review sub performance for prequalification, or justify removing a sub from the bid list. Produces a weighted scorecard with per-dimension anchored ratings, evidence notes, and an award/retention recommendation.
.claude/skills/mohitagw15856-subcontractor-scorecard/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 43% | 0% |
Every GC has a mental list of subs they'd hire again and subs they'd rather eat the higher bid to avoid. This skill turns that instinct into a defensible record: five weighted dimensions, anchored 1–5 scales so two reviewers score alike, evidence per rating, and a recommendation you can put in front of a preconstruction meeting — or a sub's principal — without it reading as a grudge.
Ask for what's missing; from partial data, score what's evidenced and mark unscored dimensions [insufficient data] rather than guessing:
Score each dimension 1–5 against the anchors, then weight:
| Dimension | Weight | 1 (fails) | 3 (solid) | 5 (excellent) | |---|---|---|---|---| | Schedule reliability | 30% | Chronic slips, ghost crews, drives the critical path late | Hits most dates; slips flagged early with recovery plan | Hits dates, staffs to plan, accelerates when asked without drama | | Quality / rework | 25% | Repeated failed inspections, punch counts far above trade norm, back-charged rework | Normal punch volume, closes items promptly, rare rework | First-time-quality culture; punch list light and closed fast | | Safety | 20% | Recordable(s) from ignored controls; fights the safety program | Compliant; participates in briefings; near-misses reported | Brings hazards to you first; crews self-police; clean record | | Paperwork discipline | 10% | Chases required for waivers/certs; closeout drags months | Mostly on time with reminders | Billing, waivers, and closeout docs arrive right, first time | | Change-order behaviour | 15% | Weaponises COs — lowball bid, then claims on every RFI | Prices changes fairly with backup; negotiates in good faith | Flags cost issues before they're changes; transparent pricing |
Composite = Σ(rating × weight) × 20, giving 0–100. Map to a recommendation:
Safety scores of 1–2 cap the overall recommendation at Conditional regardless of composite — a sub who hurts people isn't "preferred" at any price.
1. Composite score & recommendation tier — with the one-paragraph justification. 2. Dimension table — | Dimension | Weight | Rating (1–5) | Evidence | 3. Trend note — improving, stable, or declining vs. prior projects, if history given. 4. Conditions / feedback points — specific, evidence-backed items for the sub conversation. 5. Data gaps — dimensions scored on thin evidence, flagged [insufficient data].
[insufficient data] and say what record-keeping would fix it| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,901 | 23,362 | +7% | 1 | 1 | 0% | 3,097 | 4,659 | +50% | 0 | 0 | — |
case-02 | fail→pass | 17,092 | 16,062 | -6% | 1 | 1 | 0% | 3,569 | 4,306 | +21% | 0 | 0 | — |
case-03 | fail→pass | 20,116 | 26,240 | +30% | 1 | 1 | 0% | 2,337 | 4,490 | +92% | 0 | 0 | — |
case-04 | pass→pass | 23,527 | 30,946 | +32% | 1 | 1 | 0% | 3,394 | 4,869 | +43% | 0 | 0 | — |
case-05 | pass→pass | 28,081 | 21,348 | -24% | 1 | 1 | 0% | 3,851 | 4,657 | +21% | 0 | 0 | — |
case-06 | pass→pass | 43,384 | 35,695 | -18% | 1 | 1 | 0% | 6,902 | 5,070 | -27% | 0 | 0 | — |
case-07 | fail→pass | 12,001 | 14,563 | +21% | 1 | 1 | 0% | 2,442 | 2,961 | +21% | 0 | 0 | — |
case-08 | fail→fail | 16,718 | 8,558 | -49% | 1 | 1 | 0% | 1,839 | 2,729 | +48% | 0 | 0 | — |
case-09 | pass→pass | 17,944 | 18,916 | +5% | 1 | 1 | 0% | 2,025 | 3,151 | +56% | 0 | 0 | — |
case-10 | pass→pass | 16,898 | 22,193 | +31% | 1 | 1 | 0% | 2,085 | 3,524 | +69% | 0 | 0 | — |
case-15 | pass→pass | 14,603 | 12,414 | -15% | 1 | 1 | 0% | 2,381 | 3,320 | +39% | 0 | 0 | — |
case-11 | fail→fail | 18,119 | 13,862 | -23% | 1 | 1 | 0% | 1,929 | 2,852 | +48% | 0 | 0 | — |
case-12 | fail→fail | 13,249 | 9,871 | -25% | 1 | 1 | 0% | 1,412 | 3,049 | +116% | 0 | 0 | — |
case-13 | pass→pass | 9,596 | 18,473 | +93% | 1 | 1 | 0% | 1,334 | 3,047 | +128% | 0 | 0 | — |
case-14 | fail→fail | 10,419 | 17,688 | +70% | 1 | 1 | 0% | 1,555 | 3,389 | +118% | 0 | 0 | — |
case-16 | fail→fail | 17,153 | 17,579 | +2% | 1 | 1 | 0% | 2,675 | 3,454 | +29% | 0 | 0 | — |
case-17 | fail→fail | 11,806 | 14,854 | +26% | 1 | 1 | 0% | 2,028 | 2,915 | +44% | 0 | 0 | — |
case-18 | pass→pass | 16,486 | 12,295 | -25% | 1 | 1 | 0% | 2,252 | 3,343 | +48% | 0 | 0 | — |
case-19 | fail→fail | 17,084 | 16,424 | -4% | 1 | 1 | 0% | 3,008 | 3,964 | +32% | 0 | 0 | — |
case-20 | pass→pass | 11,015 | 10,588 | -4% | 1 | 1 | 0% | 1,874 | 3,274 | +75% | 0 | 0 | — |
case-21 | pass→pass | 13,760 | 24,761 | +80% | 1 | 1 | 0% | 2,176 | 4,447 | +104% | 0 | 0 | — |
case-22 | fail→fail | 11,862 | 9,192 | -23% | 1 | 1 | 0% | 2,323 | 2,917 | +26% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.