Install any skill in seconds. Free to start, no credit card required.
Get Started Free →When the user needs to design an interview process, create interview questions, build scorecards, calibrate interviewers, or evaluate candidates for a role.
.claude/skills/mkurman-interview-kit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 102% | 0% |
A complete interview kit document containing:
Each score level must include 1-2 concrete behavioral anchors specific to the role being evaluated.
Structure behavioral questions to elicit complete, pattern-revealing answers:
Example: "Tell me about a time you had to ship something with significant technical debt. What was the situation, what did you decide, and how did it play out? Would you make the same call again?"
Use three inputs to triangulate:
Always present comp as a range with a target midpoint, not a single number.
job-description — Use the JD's competency requirements as input for the interview loopsourcing-outreach — Align outreach messaging with the interview process so candidates know what to expectPrompt: "Design an interview loop for a senior backend engineer. 15-person startup."
Good output snippet:
## Interview Loop — Senior Backend Engineer
### Competencies to Evaluate
1. System design & architecture (technical depth)
2. Code quality & testing practices (craft)
3. Debugging & production thinking (operational maturity)
4. Technical communication (collaboration)
5. Ownership & initiative (startup fit)
### Stage 1: Founder Screen (30 min)
- Evaluate: Motivation, communication, logistics
- Questions:
- "What's drawing you to an early-stage company right now?"
- "Walk me through the most impactful project you led in the last year."
- Scorecard: 1-4 on communication, motivation, startup-fit
### Stage 2: Technical Deep-Dive (60 min)
- Evaluate: System design, code quality
- Format: Live system design discussion + code review exercise
- Scorecard: 1-4 on architecture thinking, code craft, trade-off reasoningPrompt: "Our interviewers keep disagreeing on candidates."
Good output snippet:
This usually means you lack structured evaluation criteria. Three-step fix:
1. Define 4-5 competencies per role with written behavioral descriptions
2. Give each interviewer a scorecard to fill out independently BEFORE debrief
3. In the debrief, discuss only scores that diverge by 2+ points —
focus on evidence, not impressions
The goal is calibrated, evidence-based evaluation — not consensus.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→fail | 16,143 | 18,865 | +17% | 1 | 1 | 0% | 2,450 | 4,563 | +86% | 0 | 0 | — |
case-01 | fail→fail | 34,298 | 36,192 | +6% | 1 | 1 | 0% | 5,138 | 7,440 | +45% | 0 | 0 | — |
case-02 | fail→fail | 40,495 | 33,652 | -17% | 1 | 1 | 0% | 6,226 | 6,950 | +12% | 0 | 0 | — |
case-03 | fail→fail | 20,247 | 27,033 | +34% | 1 | 1 | 0% | 3,222 | 6,217 | +93% | 0 | 0 | — |
case-04 | fail→pass | 24,405 | 23,155 | -5% | 1 | 1 | 0% | 3,584 | 5,435 | +52% | 0 | 0 | — |
case-05 | fail→pass | 16,360 | 20,025 | +22% | 1 | 1 | 0% | 2,421 | 4,667 | +93% | 0 | 0 | — |
case-07 | fail→pass | 18,720 | 18,596 | -1% | 1 | 1 | 0% | 2,726 | 4,689 | +72% | 0 | 0 | — |
case-08 | fail→pass | 16,105 | 20,735 | +29% | 1 | 1 | 0% | 2,450 | 4,961 | +102% | 0 | 0 | — |
case-09 | fail→pass | 18,661 | 25,144 | +35% | 1 | 1 | 0% | 2,726 | 5,516 | +102% | 0 | 0 | — |
case-10 | fail→fail | 13,872 | 9,546 | -31% | 1 | 1 | 0% | 2,083 | 3,244 | +56% | 0 | 0 | — |
case-11 | pass→pass | 17,490 | 18,782 | +7% | 1 | 1 | 0% | 2,590 | 4,536 | +75% | 0 | 0 | — |
case-12 | fail→fail | 22,023 | 20,696 | -6% | 1 | 1 | 0% | 3,313 | 4,873 | +47% | 0 | 0 | — |
case-13 | fail→pass | 15,571 | 19,621 | +26% | 1 | 1 | 0% | 2,336 | 4,839 | +107% | 0 | 0 | — |
case-14 | pass→pass | 13,143 | 17,759 | +35% | 1 | 1 | 0% | 2,117 | 4,558 | +115% | 0 | 0 | — |
case-15 | pass→pass | 15,996 | 18,600 | +16% | 1 | 1 | 0% | 2,510 | 4,533 | +81% | 0 | 0 | — |
case-16 | fail→fail | 15,526 | 16,064 | +3% | 1 | 1 | 0% | 2,319 | 4,068 | +75% | 0 | 0 | — |
case-17 | fail→fail | 19,983 | 31,645 | +58% | 1 | 1 | 0% | 3,264 | 6,804 | +108% | 0 | 0 | — |
case-18 | pass→pass | 14,756 | 25,522 | +73% | 1 | 1 | 0% | 2,311 | 5,497 | +138% | 0 | 0 | — |
case-19 | pass→pass | 15,576 | 16,719 | +7% | 1 | 1 | 0% | 2,301 | 4,326 | +88% | 0 | 0 | — |
case-20 | pass→pass | 13,151 | 13,132 | -0% | 1 | 1 | 0% | 2,066 | 3,700 | +79% | 0 | 0 | — |
case-21 | pass→pass | 14,570 | 14,069 | -3% | 1 | 1 | 0% | 2,229 | 3,892 | +75% | 0 | 0 | — |
case-22 | pass→pass | 17,868 | 21,733 | +22% | 1 | 1 | 0% | 2,887 | 5,170 | +79% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.