Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used when the user asks to "design interview processes", "create hiring pipelines", "calibrate interview loops", "generate interview questions", "design competency matrices", "analyze interviewer bias", "create scoring rubrics", "build question banks", or "optimize hiring systems". Use for designing role-specific interview loops, competency assessments, and hiring calibration systems.
.claude/skills/alirezarezvani-interview-system-designer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -63% | 0% |
Comprehensive interview loop planning and calibration support for role-based hiring systems.
Use this skill to create structured interview loops, standardize question quality, and keep hiring signal consistent across interviewers.
bash# Generate a loop plan for a role and level python3 scripts/interview_planner.py --role "Senior Software Engineer" --level senior # JSON output for integration with internal tooling python3 scripts/interview_planner.py --role "Product Manager" --level mid --json
scripts/interview_planner.py to generate a baseline loop.references/interview-frameworks.mdreferences/bias_mitigation_checklist.mdreferences/competency_matrix_templates.mdreferences/debrief_facilitation_guide.md| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 14,177 | 11,498 | -19% | 1 | 1 | 0% | 3,230 | 2,973 | -8% | 0 | 0 | — |
case-08 | fail→pass | 3,988 | 1,346 | -66% | 1 | 1 | 0% | 654 | 626 | -4% | 0 | 0 | — |
case-01 | fail→fail | 15,395 | 20,591 | +34% | 1 | 1 | 0% | 3,073 | 4,408 | +43% | 0 | 0 | — |
case-02 | fail→pass | 14,340 | 7,593 | -47% | 1 | 1 | 0% | 2,935 | 1,978 | -33% | 0 | 0 | — |
case-03 | fail→pass | 10,428 | 4,213 | -60% | 1 | 1 | 0% | 1,793 | 1,218 | -32% | 0 | 0 | — |
case-04 | pass→pass | 5,411 | 3,871 | -28% | 1 | 1 | 0% | 1,025 | 1,086 | +6% | 0 | 0 | — |
case-05 | pass→pass | 12,932 | 11,564 | -11% | 1 | 1 | 0% | 2,149 | 2,374 | +10% | 0 | 0 | — |
case-06 | pass→pass | 9,884 | 6,699 | -32% | 1 | 1 | 0% | 1,650 | 1,504 | -9% | 0 | 0 | — |
case-07 | fail→pass | 14,390 | 10,035 | -30% | 1 | 1 | 0% | 2,468 | 2,109 | -15% | 0 | 0 | — |
case-09 | fail→pass | 9,305 | 1,346 | -86% | 1 | 1 | 0% | 1,663 | 609 | -63% | 0 | 0 | — |
case-10 | fail→pass | 11,793 | 2,487 | -79% | 1 | 1 | 0% | 2,081 | 852 | -59% | 0 | 0 | — |
case-11 | fail→pass | 11,308 | 3,322 | -71% | 1 | 1 | 0% | 1,924 | 970 | -50% | 0 | 0 | — |
case-12 | pass→pass | 12,690 | 11,733 | -8% | 1 | 1 | 0% | 2,092 | 2,423 | +16% | 0 | 0 | — |
case-13 | pass→pass | 11,753 | 10,064 | -14% | 1 | 1 | 0% | 2,009 | 2,159 | +7% | 0 | 0 | — |
case-14 | pass→pass | 12,702 | 14,120 | +11% | 1 | 1 | 0% | 2,036 | 2,600 | +28% | 0 | 0 | — |
case-15 | pass→pass | 10,552 | 10,078 | -4% | 1 | 1 | 0% | 1,781 | 2,001 | +12% | 0 | 0 | — |
case-16 | pass→pass | 13,763 | 14,335 | +4% | 1 | 1 | 0% | 2,455 | 2,843 | +16% | 0 | 0 | — |
case-17 | pass→pass | 8,808 | 6,744 | -23% | 1 | 1 | 0% | 1,544 | 1,473 | -5% | 0 | 0 | — |
case-18 | pass→pass | 12,979 | 14,879 | +15% | 1 | 1 | 0% | 2,275 | 3,024 | +33% | 0 | 0 | — |
case-19 | pass→pass | 13,699 | 14,290 | +4% | 1 | 1 | 0% | 2,204 | 2,659 | +21% | 0 | 0 | — |
case-20 | pass→pass | 12,721 | 14,819 | +16% | 1 | 1 | 0% | 2,289 | 2,822 | +23% | 0 | 0 | — |
case-21 | pass→pass | 16,251 | 16,630 | +2% | 1 | 1 | 0% | 3,115 | 3,769 | +21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.