Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Socratic questioning to crystallize vague requirements into a testable spec. Use when input is ambiguous, when the user asks for "deep interview", or before dev-autopilot if the brief is too broad. Pairs with @echo-analyst.
.claude/skills/evolution-foundation-dev-deep-interview/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 344% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 183% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -5% | 0% |
Derived from oh-my-claudecode (MIT, Yeachan Heo). Adapted for the EvoNexus Engineering Layer.
Deep Interview transforms vague ideas into concrete, testable specifications through Socratic questioning. It is the front gate of any high-stakes work — refusing to proceed until ambiguity is below an acceptable threshold.
dev-autopilot when input lacks file paths, function names, or concrete anchorsDrive ambiguity below 20% before any code or plan is generated. Output a spec that an executor agent can implement without further clarification.
Read the user's input. Score it on these dimensions (1-5 each, 5 = clear):
If average ≥ 4: skip interview, proceed. If average < 4: enter interview loop.
Ask one question at a time using AskUserQuestion (with 2-4 multiple-choice options when possible). Each question must:
Common question types:
Stop the loop when all dimensions ≥ 4 OR after 8 questions (whichever first — long interviews lose user patience).
Write the spec to workspace/projects/specs/[C]deep-interview-{name}.md with this structure:
markdown# Deep Interview Spec — {topic} **Date:** {iso} **Ambiguity score:** {avg}/5 ## Context [1-2 sentences on the problem and why it matters] ## In Scope - [item 1] - [item 2] ## Out of Scope - [explicit non-goals] ## Success Criteria - [testable criterion 1] - [testable criterion 2] ## Constraints - [tech / business / time] ## Open Questions - [items where ambiguity remains, with risk level] ## Suggested Next Step - `dev-autopilot` (if all dimensions ≥ 4) - `dev-plan` (if some dimensions still < 4 but you want to start scoping) - Manual implementation (if the spec is small enough)
@scout-explorer to look them up.@echo-analyst — for deeper requirements gap analysis after the interviewdev-autopilot — natural next step once spec is readydev-plan — if you want to scope further before execution| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | fail→pass | 8,265 | 2,277 | -72% | 1 | 1 | 0% | 1,270 | 1,410 | +11% | 0 | 0 | — |
case-01 | fail→fail | 7,851 | 3,720 | -53% | 1 | 1 | 0% | 1,304 | 1,607 | +23% | 0 | 0 | — |
case-02 | fail→pass | 3,154 | 6,301 | +100% | 1 | 1 | 0% | 446 | 1,979 | +344% | 0 | 0 | — |
case-03 | fail→pass | 3,803 | 4,883 | +28% | 1 | 1 | 0% | 616 | 1,746 | +183% | 0 | 0 | — |
case-04 | pass→pass | 7,687 | 14,056 | +83% | 1 | 1 | 0% | 1,486 | 3,044 | +105% | 0 | 0 | — |
case-05 | pass→pass | 12,586 | 10,709 | -15% | 1 | 1 | 0% | 2,838 | 3,182 | +12% | 0 | 0 | — |
case-06 | pass→pass | 5,933 | 4,989 | -16% | 1 | 1 | 0% | 879 | 1,755 | +100% | 0 | 0 | — |
case-07 | fail→fail | 9,220 | 4,085 | -56% | 1 | 1 | 0% | 1,444 | 1,638 | +13% | 0 | 0 | — |
case-08 | fail→pass | 10,107 | 9,739 | -4% | 1 | 1 | 0% | 1,561 | 2,533 | +62% | 0 | 0 | — |
case-09 | pass→pass | 13,001 | 8,285 | -36% | 1 | 1 | 0% | 2,476 | 2,460 | -1% | 0 | 0 | — |
case-10 | pass→pass | 12,438 | 10,952 | -12% | 1 | 1 | 0% | 1,947 | 2,756 | +42% | 0 | 0 | — |
case-12 | fail→pass | 13,189 | 5,139 | -61% | 1 | 1 | 0% | 1,990 | 1,896 | -5% | 0 | 0 | — |
case-13 | fail→pass | 15,208 | 2,711 | -82% | 1 | 1 | 0% | 2,261 | 1,481 | -34% | 0 | 0 | — |
case-14 | fail→fail | 11,134 | 6,316 | -43% | 1 | 1 | 0% | 1,926 | 1,926 | 0% | 0 | 0 | — |
case-15 | fail→pass | 10,461 | 2,759 | -74% | 1 | 1 | 0% | 1,580 | 1,505 | -5% | 0 | 0 | — |
case-16 | fail→pass | 5,787 | 1,697 | -71% | 1 | 1 | 0% | 812 | 1,240 | +53% | 0 | 0 | — |
case-17 | pass→pass | 4,026 | 2,983 | -26% | 1 | 1 | 0% | 597 | 1,452 | +143% | 0 | 0 | — |
case-18 | pass→pass | 4,000 | 3,418 | -15% | 1 | 1 | 0% | 531 | 1,531 | +188% | 0 | 0 | — |
case-19 | fail→pass | 9,494 | 1,917 | -80% | 1 | 1 | 0% | 1,472 | 1,289 | -12% | 0 | 0 | — |
case-20 | fail→pass | 7,999 | 1,532 | -81% | 1 | 1 | 0% | 1,190 | 1,197 | +1% | 0 | 0 | — |
case-21 | fail→pass | 9,531 | 1,830 | -81% | 1 | 1 | 0% | 1,490 | 1,312 | -12% | 0 | 0 | — |
case-22 | pass→pass | 8,464 | 4,173 | -51% | 1 | 1 | 0% | 1,326 | 1,651 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.