Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate evidenced opportunities or
.claude/skills/boshu2-idea-genie/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 204% | 0% |
One canonical root for idea work: elicit an evidence-grounded opportunity portfolio, or challenge a consequential idea with sealed independent perspectives. Both modes explore and advise; neither selects, schedules, tracks, implements, or validates work.
| Trigger phrases | Mode | Output contract | |---|---|---| | "idea genie", "what should we build", "supported opportunities" | elicit (single genie) | idea-portfolio.v1 via scripts/validate-output.sh | | "challenge this idea", "compare independent proposals", "stress-test a one-way door" | duel (adversarial challenge) | idea-challenge.v1 via scripts/validate-challenge.sh |
Elicitation is the entry mode. Dueling is an optional escalation for a consequential choice, typically consuming an idea-portfolio.v1 or a framed question.
Generate a small portfolio of evidenced options.
capabilities, and one normal or edge scenario.
idea-portfolio.v1, then return it to the caller or Plan.An empty no-new-work portfolio is valid. Plan alone may incorporate a selected option into the existing bead or caller intent.
Produce independent challenges for a consequential choice. The result is advisory evidence for Plan. It never decides whether a plan is ready and never turns a later optional Premortem challenge into an approval gate.
proposals from anchoring on earlier ones.
that synthesis might otherwise erase.
messaging service, council, or model-family rule.
tracker state because this strategy supplies evidence rather than lifecycle authority.
identifiers. Each produces its perspective before any is revealed. When the caller pins perspectives to model profiles, record each perspective's model_identity (see the agent-native model-dispatch recipe); a duel may use two distinct models on request. Sealed generation is unchanged: no perspective may see another before reveal. Unavailable profiles → disclose and continue single-model.
system fit, failure modes, and cost.
and minority reasoning.
idea-challenge.v1, validate it, and pass the artifact to Plan as oneoptional input alongside research and operator intent.
For a cheap two-way door, emit the lightweight packet directly after one fresh challenge. Do not manufacture panel ceremony.
.agents/scratch/ideas/<run-id>/idea-challenge.jsonidea-challenge.v1 JSON with route-specific fields enforced bythe validator
skills/idea-genie/scripts/validate-challenge.sh <idea-challenge.json>
handoff.owner is exactly plan; Plan may accept,reject, or combine the advisory evidence
perspective by named dimensions.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | fail→pass | 21,943 | 25,419 | +16% | 1 | 1 | 0% | 3,227 | 4,181 | +30% | 0 | 0 | — |
case-01 | fail→pass | 24,980 | 20,562 | -18% | 1 | 1 | 0% | 4,377 | 4,626 | +6% | 0 | 0 | — |
case-02 | fail→pass | 25,748 | 31,378 | +22% | 1 | 1 | 0% | 3,871 | 6,777 | +75% | 0 | 0 | — |
case-03 | fail→pass | 28,951 | 24,138 | -17% | 1 | 1 | 0% | 4,438 | 4,809 | +8% | 0 | 0 | — |
case-04 | fail→pass | 3,978 | 4,722 | +19% | 1 | 1 | 0% | 536 | 1,632 | +204% | 0 | 0 | — |
case-05 | fail→pass | 2,748 | 5,381 | +96% | 1 | 1 | 0% | 353 | 1,823 | +416% | 0 | 0 | — |
case-06 | pass→pass | 7,325 | 7,155 | -2% | 1 | 1 | 0% | 1,040 | 2,039 | +96% | 0 | 0 | — |
case-07 | fail→fail | 20,168 | 6,450 | -68% | 1 | 1 | 0% | 3,059 | 1,380 | -55% | 0 | 0 | — |
case-08 | pass→pass | 17,562 | 22,502 | +28% | 1 | 1 | 0% | 2,654 | 4,693 | +77% | 0 | 0 | — |
case-09 | fail→pass | 16,668 | 13,570 | -19% | 1 | 1 | 0% | 2,625 | 3,265 | +24% | 0 | 0 | — |
case-10 | fail→pass | 13,336 | 7,762 | -42% | 1 | 1 | 0% | 1,950 | 2,315 | +19% | 0 | 0 | — |
case-11 | fail→pass | 2,478 | 16,057 | +548% | 1 | 1 | 0% | 278 | 3,982 | +1332% | 0 | 0 | — |
case-12 | pass→pass | 16,383 | 18,779 | +15% | 1 | 1 | 0% | 2,430 | 4,259 | +75% | 0 | 0 | — |
case-13 | pass→pass | 15,793 | 23,491 | +49% | 1 | 1 | 0% | 2,272 | 4,856 | +114% | 0 | 0 | — |
case-14 | pass→pass | 21,779 | 16,774 | -23% | 1 | 1 | 0% | 3,205 | 3,566 | +11% | 0 | 0 | — |
case-15 | pass→pass | 16,581 | 10,907 | -34% | 1 | 1 | 0% | 2,663 | 2,803 | +5% | 0 | 0 | — |
case-16 | fail→pass | 20,671 | 14,723 | -29% | 1 | 1 | 0% | 2,866 | 3,125 | +9% | 0 | 0 | — |
case-17 | fail→pass | 14,712 | 10,779 | -27% | 1 | 1 | 0% | 2,157 | 2,661 | +23% | 0 | 0 | — |
case-19 | fail→pass | 15,788 | 12,607 | -20% | 1 | 1 | 0% | 1,643 | 3,118 | +90% | 0 | 0 | — |
case-20 | fail→pass | 18,566 | 31,130 | +68% | 1 | 1 | 0% | 3,087 | 6,263 | +103% | 0 | 0 | — |
case-21 | fail→pass | 22,452 | 27,603 | +23% | 1 | 1 | 0% | 3,295 | 3,986 | +21% | 0 | 0 | — |
case-22 | pass→pass | 21,708 | 17,067 | -21% | 1 | 1 | 0% | 3,172 | 3,739 | +18% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +64 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.