Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a rigorous candidate reference check that surfaces real signal. Use when asked to prepare or conduct reference calls for a job candidate, design reference questions, or build a reference-check rubric. Produces a structured question set, probing follow-ups, a red/yellow/green scoring rubric, and the legal guardrails — designed to get past 'they were great' without leading the referee.
.claude/skills/mohitagw15856-reference-check-script/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 21% | 0% |
A reference check fails when it's a formality — you call, they say "great to work with," you tick the box. This skill writes a call that actually de-risks the hire: questions that make it safe and easy for a referee to be honest, follow-ups that open the gap between polite and true, and a rubric that turns what you hear into a defensible signal.
Given the role and what you need to validate, write the full script — tailor the questions to the specific concerns (the shaky interview signal, the seniority stretch, the manager-vs-IC question). Default to a ~20–25 minute call structure.
Ask for (if not provided, else infer and label the assumption):
A 2–3 sentence opener that sets candor: confidential, no scripted "sales pitch" needed, you're calibrating not gatekeeping. The single most important move — referees mirror the tone you set.
Grouped, with the intent of each noted:
The second question that opens the gap: "Can you give me an example?" · "How did that compare to others at that level?" · "What would their harshest fair critic say?" · silence (let them fill it).
| Signal | 🟢 Green | 🟡 Yellow | 🔴 Red | |---|---|---|---| | Specificity | concrete examples | generic praise | evasive / can't recall | | Rehire | enthusiastic, unprompted | qualified | hesitation or no | | Growth honesty | candid, coachable | vague | defensive / dodged |
Plus a one-line overall read and any follow-up to run (e.g. a back-channel if listed refs are all glowing-but-generic).
Ask about job performance only. Do not ask about age, health/disability, family/pregnancy, religion, national origin, or other protected characteristics. Keep it consistent across candidates so it's comparable and defensible.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 29,201 | 30,036 | +3% | 1 | 1 | 0% | 3,712 | 4,160 | +12% | 0 | 0 | — |
case-02 | fail→fail | 28,458 | 26,529 | -7% | 1 | 1 | 0% | 3,843 | 4,158 | +8% | 0 | 0 | — |
case-03 | fail→pass | 28,569 | 22,115 | -23% | 1 | 1 | 0% | 3,945 | 3,861 | -2% | 0 | 0 | — |
case-04 | pass→fail | 15,906 | 64,114 | +303% | 1 | 1 | 0% | 2,213 | 4,887 | +121% | 0 | 0 | — |
case-05 | pass→fail | 11,579 | 31,466 | +172% | 1 | 1 | 0% | 2,102 | 5,608 | +167% | 0 | 0 | — |
case-06 | pass→pass | 18,788 | 27,915 | +49% | 1 | 1 | 0% | 2,678 | 3,857 | +44% | 0 | 0 | — |
case-07 | fail→fail | 26,056 | 23,809 | -9% | 1 | 1 | 0% | 4,099 | 4,612 | +13% | 0 | 0 | — |
case-08 | fail→fail | 21,031 | 23,607 | +12% | 1 | 1 | 0% | 3,263 | 4,390 | +35% | 0 | 0 | — |
case-09 | fail→pass | 21,057 | 24,316 | +15% | 1 | 1 | 0% | 2,722 | 4,100 | +51% | 0 | 0 | — |
case-10 | fail→fail | 20,267 | 18,931 | -7% | 1 | 1 | 0% | 2,821 | 3,738 | +33% | 0 | 0 | — |
case-11 | fail→fail | 21,872 | 20,545 | -6% | 1 | 1 | 0% | 3,367 | 4,013 | +19% | 0 | 0 | — |
case-12 | fail→fail | 32,576 | 22,638 | -31% | 1 | 1 | 0% | 4,269 | 4,152 | -3% | 0 | 0 | — |
case-13 | fail→fail | 18,346 | 17,581 | -4% | 1 | 1 | 0% | 2,574 | 3,625 | +41% | 0 | 0 | — |
case-14 | fail→pass | 28,527 | 58,154 | +104% | 1 | 1 | 0% | 3,642 | 4,173 | +15% | 0 | 0 | — |
case-15 | fail→pass | 37,557 | 34,048 | -9% | 1 | 1 | 0% | 3,162 | 3,830 | +21% | 0 | 0 | — |
case-16 | fail→pass | 25,182 | 23,731 | -6% | 1 | 1 | 0% | 3,400 | 4,122 | +21% | 0 | 0 | — |
case-17 | fail→fail | 30,322 | 24,739 | -18% | 1 | 1 | 0% | 3,577 | 3,886 | +9% | 0 | 0 | — |
case-18 | fail→pass | 25,683 | 26,645 | +4% | 1 | 1 | 0% | 3,817 | 4,523 | +18% | 0 | 0 | — |
case-19 | fail→pass | 15,915 | 19,250 | +21% | 1 | 1 | 0% | 2,473 | 3,820 | +54% | 0 | 0 | — |
case-20 | fail→fail | 21,820 | 19,149 | -12% | 1 | 1 | 0% | 3,316 | 3,813 | +15% | 0 | 0 | — |
case-21 | fail→pass | 22,911 | 20,777 | -9% | 1 | 1 | 0% | 3,055 | 3,595 | +18% | 0 | 0 | — |
case-22 | fail→pass | 26,173 | 17,854 | -32% | 1 | 1 | 0% | 3,118 | 4,068 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.