Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when drafting interview questions to assess a named competency (e.g. conflict resolution, ownership, dealing with ambiguity) — produces open, past-behavior questions that pull real STAR evidence, each with a follow-up probe, instead of yes/no, leading, or speculative "how would you" prompts. Do NOT use for legal-risk vetting of drafted questions (interview-question-compliance), writing up notes into a scorecard (interview-scorecard-format), or screening a resume against a rubric (candidate-screening).
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 412% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 416% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 412% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 411% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 415% | 0% |
Draft the questions an interviewer asks to assess a specific competency. Asked to "come up with questions to gauge how a candidate handles X", base models know what a behavioral question is — but they don't default to one. They mix in three weaker forms that don't produce evaluable evidence:
describe the textbook answer without ever having lived it. You learn what they think good looks like, not what they did.
with one word, and the word is always yes.
question hands them the conclusion you were supposed to test.
This skill produces the fourth form: open, past-behavior questions that make the candidate narrate a real event, so you can score what they actually did.
competency, trait, or behavior for a role.
(interview-question-compliance); turning interview notes into a scorecard (interview-scorecard-format); or scoring a resume against minimum bars (candidate-screening).
Ask what the candidate did, not what they would do, believe, or are. Every question points at a specific past event the candidate lived through, and is phrased so they cannot answer it in one word or read the desired answer off the question.
Each question must pass all four:
"Tell me about a time…", "Describe a situation where…", "Walk me through a specific instance when…". Never "How would you…", "What would you do if…", "Imagine…" — those invite a rehearsed ideal, not evidence.
reopen it. "Are you decisive?" → "Tell me about a decision you had to make with incomplete information."
your projects so organized?" assumes the trait; ask instead "Describe a time you had more work than time — how did you decide what got done?" Don't embed the value word ("great", "smoothly", "successfully") in the ask.
generic filler ("Tell me about yourself", "What's your greatest weakness?"). If you swapped in a different competency and the question still fit, it's too generic.
A good question opens the door; the follow-up probes get the evidence. STAR = Situation, Task, Action, Result. Candidates over-narrate the Situation and skip the Action and Result — the two parts that actually reveal the competency. Pair every lead question with probes that force those out:
differently?"
Give at least one probe per lead question, and make the Action/Result probes explicit — that's where a "we" answer gets pinned down to an "I".
> Competency to assess: handling conflict with a peer, for a senior engineer.
Weak (what the base tends to emit):
Strong:
disagreement?"
afterward?"
what happened."
For the named competency, produce 3–5 lead questions, each on its own line, each with its probes nested underneath. Don't grade the answers, don't invent the candidate's responses, and don't restate the competency name as a question ("Are you good at conflict?"). If the ask names a role, tune the situations to that role's world (a manager's peers, a salesperson's accounts) — but keep every question a past-behavior, open, non-leading one.
Other measured skills in the registry, with their headline benchmark lift.