Install any skill in seconds. Free to start, no credit card required.
Get Started Free →One-time setup that gathers design context for your project and saves it to your AI config file. Run once to establish persistent design guidelines.
.claude/skills/bilal140202-teach-impeccable/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | -41% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -70% | 0% |
| case-21 | ✓→✗ | ▼ Worse | -49% | 0% |
| case-13 | ✓→✗ | ▼ Worse | -68% | 0% |
| case-22 | ✓→✗ | ▼ Worse | -70% | 0% |
Gather design context for this project, then persist it for all future sessions.
Before asking questions, thoroughly scan the project to discover what you can:
Note what you've learned and what remains unclear.
STOP and call the AskUserQuestionTool to clarify. Focus only on what you couldn't infer from the codebase:
Skip questions where the answer is already clear from the codebase exploration.
Synthesize your findings and the user's answers into a ## Design Context section:
markdown## Design Context ### Users [Who they are, their context, the job to be done] ### Brand Personality [Voice, tone, 3-word personality, emotional goals] ### Aesthetic Direction [Visual tone, references, anti-references, theme] ### Design Principles [3-5 principles derived from the conversation that should guide all design decisions]
Write this section to CLAUDE.md in the project root. If the file exists, append or update the Design Context section.
Confirm completion and summarize the key design principles that will now guide all future work.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,373 | 2,344 | -82% | 1 | 1 | 0% | 2,087 | 693 | -67% | 0 | 0 | — |
case-02 | fail→fail | 13,016 | 7,370 | -43% | 1 | 1 | 0% | 2,107 | 1,137 | -46% | 0 | 0 | — |
case-03 | fail→fail | 2,392 | 4,042 | +69% | 1 | 1 | 0% | 328 | 706 | +115% | 0 | 0 | — |
case-04 | fail→fail | 6,811 | 2,924 | -57% | 1 | 1 | 0% | 1,055 | 898 | -15% | 0 | 0 | — |
case-05 | fail→fail | 18,916 | 3,113 | -84% | 1 | 1 | 0% | 3,575 | 653 | -82% | 0 | 0 | — |
case-06 | pass→fail | 12,344 | 3,843 | -69% | 1 | 1 | 0% | 2,327 | 697 | -70% | 0 | 0 | — |
case-07 | fail→fail | 8,094 | 1,930 | -76% | 1 | 1 | 0% | 1,241 | 652 | -47% | 0 | 0 | — |
case-08 | fail→fail | 11,826 | 2,853 | -76% | 1 | 1 | 0% | 1,819 | 674 | -63% | 0 | 0 | — |
case-09 | fail→fail | 4,056 | 2,537 | -37% | 1 | 1 | 0% | 648 | 705 | +9% | 0 | 0 | — |
case-10 | fail→fail | 9,833 | 7,796 | -21% | 1 | 1 | 0% | 1,517 | 1,623 | +7% | 0 | 0 | — |
case-11 | fail→fail | 12,584 | 3,825 | -70% | 1 | 1 | 0% | 1,914 | 769 | -60% | 0 | 0 | — |
case-21 | pass→fail | 7,522 | 3,699 | -51% | 1 | 1 | 0% | 1,480 | 754 | -49% | 0 | 0 | — |
case-12 | fail→fail | 13,070 | 4,108 | -69% | 1 | 1 | 0% | 1,929 | 760 | -61% | 0 | 0 | — |
case-13 | pass→fail | 12,196 | 18,579 | +52% | 1 | 1 | 0% | 2,086 | 674 | -68% | 0 | 0 | — |
case-14 | fail→fail | 7,442 | 2,370 | -68% | 1 | 1 | 0% | 1,236 | 658 | -47% | 0 | 0 | — |
case-15 | fail→fail | 9,254 | 2,290 | -75% | 1 | 1 | 0% | 1,565 | 633 | -60% | 0 | 0 | — |
case-22 | pass→fail | 10,078 | 2,947 | -71% | 1 | 1 | 0% | 2,451 | 733 | -70% | 0 | 0 | — |
case-16 | fail→fail | 8,163 | 5,407 | -34% | 1 | 1 | 0% | 1,407 | 1,193 | -15% | 0 | 0 | — |
case-17 | fail→fail | 3,958 | 2,251 | -43% | 1 | 1 | 0% | 589 | 673 | +14% | 0 | 0 | — |
case-18 | fail→pass | 11,495 | 6,049 | -47% | 1 | 1 | 0% | 1,986 | 1,181 | -41% | 0 | 0 | — |
case-19 | fail→fail | 8,056 | 3,420 | -58% | 1 | 1 | 0% | 1,292 | 713 | -45% | 0 | 0 | — |
case-20 | fail→fail | 3,532 | 5,832 | +65% | 1 | 1 | 0% | 556 | 811 | +46% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 4 counted toward the lift figure. The other 18 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -14 percentage points is the difference between those two pass rates over the 4 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.