Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optional 3-5 direction picker for users who explicitly ask to compare visual directions.
.claude/skills/nexu-io-direction-picker/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | -41% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -45% | 0% |
Generative work benefits from explicit divergence before it converges. This atom defines how to present 3–5 distinct visual / structural / tonal directions when the user explicitly asks to see or compare direction options. Only in that case, emit one inline <question-form> with a direction-cards question. The submitted choice returns as the next user message.
The presence of this atom or the plan stage does not trigger a picker. Do not emit direction cards proactively. When the user has not explicitly requested options, infer a fitting direction from the brief, active design system, and known context, then continue.
When a picker was explicitly requested, the atom completes when the submitted form answer contains a direction id. The agent's next turn must lock onto that direction — backtracking forces a fresh devloop iteration of the picker stage.
(every direction must be a defensible standalone bet).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-19 | fail→pass | 19,100 | 11,577 | -39% | 1 | 1 | 0% | 3,004 | 1,770 | -41% | 0 | 0 | — |
case-06 | pass→pass | 16,216 | 9,478 | -42% | 1 | 1 | 0% | 2,334 | 2,068 | -11% | 0 | 0 | — |
case-01 | fail→pass | 16,077 | 9,367 | -42% | 1 | 1 | 0% | 2,634 | 1,858 | -29% | 0 | 0 | — |
case-02 | fail→pass | 21,722 | 8,193 | -62% | 1 | 1 | 0% | 3,223 | 1,535 | -52% | 0 | 0 | — |
case-03 | fail→pass | 26,878 | 11,317 | -58% | 1 | 1 | 0% | 3,673 | 2,216 | -40% | 0 | 0 | — |
case-04 | pass→pass | 52,690 | 85,253 | +62% | 1 | 1 | 0% | 8,238 | 8,485 | +3% | 0 | 0 | — |
case-05 | pass→pass | 39,934 | 20,320 | -49% | 1 | 1 | 0% | 6,074 | 4,183 | -31% | 0 | 0 | — |
case-07 | pass→pass | 25,990 | 21,012 | -19% | 1 | 1 | 0% | 3,410 | 3,656 | +7% | 0 | 0 | — |
case-08 | fail→pass | 22,189 | 10,438 | -53% | 1 | 1 | 0% | 3,058 | 1,674 | -45% | 0 | 0 | — |
case-09 | fail→pass | 22,587 | 14,022 | -38% | 1 | 1 | 0% | 3,063 | 2,679 | -13% | 0 | 0 | — |
case-10 | fail→pass | 11,502 | 10,701 | -7% | 1 | 1 | 0% | 2,049 | 1,708 | -17% | 0 | 0 | — |
case-11 | fail→pass | 13,906 | 8,451 | -39% | 1 | 1 | 0% | 1,806 | 1,621 | -10% | 0 | 0 | — |
case-12 | fail→pass | 20,217 | 8,281 | -59% | 1 | 1 | 0% | 2,736 | 1,582 | -42% | 0 | 0 | — |
case-13 | pass→pass | 9,161 | 37,370 | +308% | 1 | 1 | 0% | 1,399 | 6,366 | +355% | 0 | 0 | — |
case-14 | fail→pass | 32,203 | 8,751 | -73% | 1 | 1 | 0% | 3,157 | 1,447 | -54% | 0 | 0 | — |
case-15 | fail→pass | 19,007 | 30,588 | +61% | 1 | 1 | 0% | 2,663 | 1,748 | -34% | 0 | 0 | — |
case-16 | fail→pass | 22,349 | 9,350 | -58% | 1 | 1 | 0% | 3,603 | 1,826 | -49% | 0 | 0 | — |
case-17 | fail→pass | 22,024 | 9,137 | -59% | 1 | 1 | 0% | 3,007 | 1,778 | -41% | 0 | 0 | — |
case-18 | fail→pass | 22,667 | 8,686 | -62% | 1 | 1 | 0% | 3,416 | 1,469 | -57% | 0 | 0 | — |
case-20 | fail→pass | 17,238 | 7,592 | -56% | 1 | 1 | 0% | 2,677 | 1,290 | -52% | 0 | 0 | — |
case-21 | fail→pass | 24,191 | 9,542 | -61% | 1 | 1 | 0% | 3,237 | 1,288 | -60% | 0 | 0 | — |
case-22 | fail→pass | 22,145 | 8,484 | -62% | 1 | 1 | 0% | 2,964 | 1,798 | -39% | 0 | 0 | — |
case-23 | fail→pass | 21,015 | 9,142 | -56% | 1 | 1 | 0% | 3,623 | 1,688 | -53% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +78 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.