Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design calibrated interview loops, competency-based question banks, and hiring calibration. Use when designing interview processes, creating hiring pipelines, generating scoring rubrics, analyzing interviewer bias, or building question banks.
.claude/skills/borghei-interview-system-designer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 83% | 0% |
Design role-specific interview loops, generate competency-based question banks with scoring rubrics, and detect interviewer bias through statistical calibration analysis.
Before designing, confirm these inputs. If any is unknown or vague, ASK — do not assume:
loop_designer.py vs question_bank_generator.py vs hiring_calibrator.py)--role/--level)--competencies and the question-bank scope)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
The Python tools live at the skill root (not in scripts/). All support --help, JSON/text output.
| Tool | Purpose | Command | |------|---------|---------| | loop_designer.py | Generate a calibrated interview loop (rounds, time, scorecards) | python loop_designer.py --role "Senior Software Engineer" --level senior --team platform --output loops/ | | question_bank_generator.py | Generate competency-based questions with rubrics + calibration examples | python question_bank_generator.py --role "Frontend Engineer" --competencies react,typescript,system-design --num-questions 30 | | hiring_calibrator.py | Detect bias/calibration drift across interviewers and periods | python hiring_calibrator.py --input interview_data.json --analysis-type comprehensive --trend-analysis |
Load the reference that matches the task — keep this file lean and pull detail on demand:
This skill covers:
This skill does NOT cover:
hr-operations/talent-acquisitionhr-operations/hr-business-partnerhr-operations/people-analyticsengineering/codebase-onboarding| Skill | Integration | Data Flow | |-------|-------------|-----------| | hr-operations/talent-acquisition | Feed designed interview loops and scorecards into the talent acquisition pipeline for end-to-end hiring execution | Loop JSON output → talent acquisition workflow input | | hr-operations/people-analytics | Supply calibration reports and interviewer performance data for workforce-level hiring analytics | Calibrator JSON reports → people analytics dashboards | | engineering/codebase-onboarding | Hand off hired candidate profiles and assessed competency gaps to onboarding plan generation | Scorecard results → onboarding skill-gap inputs | | hr-operations/hr-business-partner | Provide interview quality metrics and pass-rate data to support hiring bar discussions with HR leadership | Calibration trend data → HRBP quarterly reviews | | product-team | Align PM interview loop competencies with the product team's competency frameworks and role leveling guides | Competency matrix → PM loop designer --competencies input | | engineering/pr-review-expert | Use coding round evaluation criteria to inform code review standards for new hires during their ramp period | Scoring rubric technical criteria → PR review checklist alignment |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,470 | 39,291 | +61% | 1 | 1 | 0% | 4,036 | 7,609 | +89% | 0 | 0 | — |
case-02 | fail→fail | 39,803 | 39,841 | +0% | 1 | 1 | 0% | 6,197 | 7,614 | +23% | 0 | 0 | — |
case-03 | fail→pass | 17,122 | 3,985 | -77% | 1 | 1 | 0% | 2,565 | 2,046 | -20% | 0 | 0 | — |
case-04 | fail→pass | 16,517 | 5,420 | -67% | 1 | 1 | 0% | 2,584 | 2,222 | -14% | 0 | 0 | — |
case-05 | fail→pass | 22,442 | 23,085 | +3% | 1 | 1 | 0% | 3,347 | 5,243 | +57% | 0 | 0 | — |
case-06 | fail→fail | 20,131 | 4,397 | -78% | 1 | 1 | 0% | 2,919 | 2,062 | -29% | 0 | 0 | — |
case-07 | fail→pass | 7,121 | 3,108 | -56% | 1 | 1 | 0% | 1,096 | 1,901 | +73% | 0 | 0 | — |
case-08 | fail→pass | 7,456 | 4,075 | -45% | 1 | 1 | 0% | 1,126 | 2,059 | +83% | 0 | 0 | — |
case-09 | fail→pass | 6,684 | 4,233 | -37% | 1 | 1 | 0% | 1,019 | 2,011 | +97% | 0 | 0 | — |
case-10 | fail→pass | 12,598 | 5,537 | -56% | 1 | 1 | 0% | 2,004 | 2,250 | +12% | 0 | 0 | — |
case-11 | fail→pass | 12,761 | 4,003 | -69% | 1 | 1 | 0% | 1,806 | 2,000 | +11% | 0 | 0 | — |
case-12 | fail→pass | 5,635 | 2,595 | -54% | 1 | 1 | 0% | 846 | 1,849 | +119% | 0 | 0 | — |
case-13 | fail→pass | 11,013 | 6,036 | -45% | 1 | 1 | 0% | 1,522 | 2,383 | +57% | 0 | 0 | — |
case-14 | fail→pass | 6,986 | 2,822 | -60% | 1 | 1 | 0% | 985 | 1,939 | +97% | 0 | 0 | — |
case-15 | pass→pass | 27,243 | 44,468 | +63% | 1 | 1 | 0% | 3,954 | 7,592 | +92% | 0 | 0 | — |
case-16 | fail→pass | 11,868 | 14,118 | +19% | 1 | 1 | 0% | 1,854 | 3,387 | +83% | 0 | 0 | — |
case-17 | fail→pass | 18,158 | 14,698 | -19% | 1 | 1 | 0% | 2,533 | 3,458 | +37% | 0 | 0 | — |
case-18 | fail→pass | 13,220 | 8,686 | -34% | 1 | 1 | 0% | 1,903 | 2,832 | +49% | 0 | 0 | — |
case-19 | pass→pass | 16,428 | 14,403 | -12% | 1 | 1 | 0% | 2,288 | 3,609 | +58% | 0 | 0 | — |
case-20 | pass→pass | 8,381 | 6,685 | -20% | 1 | 1 | 0% | 1,277 | 2,299 | +80% | 0 | 0 | — |
case-21 | pass→pass | 7,868 | 3,516 | -55% | 1 | 1 | 0% | 1,149 | 1,971 | +72% | 0 | 0 | — |
case-22 | fail→pass | 6,902 | 3,179 | -54% | 1 | 1 | 0% | 892 | 1,925 | +116% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +68 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.