Install any skill in seconds. Free to start, no credit card required.
Get Started Free →사용자와 대화하며 디자인 선호도를 파악하고, oh-my-design 서베이를 통해 DESIGN.md를 생성합니다. '디자인 시스템 만들어줘', 'DESIGN.md 생성', '디자인 잡아줘', 'UI 스타일 정해줘' 등 디자인 시스템 구성이 필요할 때 트리거됩니다.
.claude/skills/kwakseongjae-omd-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -83% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -58% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -63% | 0% |
CORE_V2_LEGACY_WRITER_BLOCKED
This legacy survey flow cannot author a new project design system. It mapped preferences into the retired frontmatter/numbered-section format and could not produce the evidence-bound Core v2 package.
For a new project or autonomous build, route the original request to omd:autopilot. For design-system-only work, route it to the current omd:init workflow. Preserve the user's original brief and preferences; do not ask them to repeat information already present in the conversation.
Do not decode an old survey result into DESIGN.md, copy a catalog reference into the project root, or hand-author a partial replacement. New output must use the clean-top, seven-anchor standalone Core v2 projection. Only an explicitly adopted profile: portable-core manifest with exact graph/projection hashes may make the System Graph canonical; a migration candidate remains non-authoritative.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 6,134 | 5,077 | -17% | 1 | 1 | 0% | 1,046 | 1,147 | +10% | 0 | 0 | — |
case-02 | fail→pass | 26,388 | 8,743 | -67% | 1 | 1 | 0% | 6,105 | 1,030 | -83% | 0 | 0 | — |
case-03 | fail→pass | 7,278 | 5,332 | -27% | 1 | 1 | 0% | 1,263 | 1,132 | -10% | 0 | 0 | — |
case-04 | pass→pass | 7,334 | 7,210 | -2% | 1 | 1 | 0% | 1,335 | 1,626 | +22% | 0 | 0 | — |
case-05 | pass→pass | 4,500 | 4,360 | -3% | 1 | 1 | 0% | 724 | 1,012 | +40% | 0 | 0 | — |
case-06 | pass→pass | 7,962 | 5,977 | -25% | 1 | 1 | 0% | 1,637 | 1,454 | -11% | 0 | 0 | — |
case-07 | fail→pass | 10,116 | 2,533 | -75% | 1 | 1 | 0% | 1,590 | 662 | -58% | 0 | 0 | — |
case-08 | fail→pass | 12,393 | 2,639 | -79% | 1 | 1 | 0% | 1,852 | 692 | -63% | 0 | 0 | — |
case-21 | fail→pass | 8,528 | 2,604 | -69% | 1 | 1 | 0% | 1,382 | 651 | -53% | 0 | 0 | — |
case-09 | fail→pass | 9,264 | 6,356 | -31% | 1 | 1 | 0% | 1,523 | 1,418 | -7% | 0 | 0 | — |
case-10 | fail→pass | 15,232 | 3,403 | -78% | 1 | 1 | 0% | 2,482 | 859 | -65% | 0 | 0 | — |
case-11 | fail→pass | 10,858 | 3,467 | -68% | 1 | 1 | 0% | 2,515 | 883 | -65% | 0 | 0 | — |
case-12 | fail→pass | 7,796 | 1,436 | -82% | 1 | 1 | 0% | 1,229 | 450 | -63% | 0 | 0 | — |
case-22 | fail→pass | 12,540 | 4,369 | -65% | 1 | 1 | 0% | 1,786 | 994 | -44% | 0 | 0 | — |
case-13 | fail→pass | 14,418 | 2,958 | -79% | 1 | 1 | 0% | 2,486 | 718 | -71% | 0 | 0 | — |
case-14 | fail→pass | 18,987 | 1,691 | -91% | 1 | 1 | 0% | 3,105 | 477 | -85% | 0 | 0 | — |
case-15 | fail→pass | 16,655 | 2,164 | -87% | 1 | 1 | 0% | 2,624 | 605 | -77% | 0 | 0 | — |
case-16 | pass→pass | 9,925 | 2,997 | -70% | 1 | 1 | 0% | 1,622 | 558 | -66% | 0 | 0 | — |
case-17 | fail→pass | 10,722 | 3,260 | -70% | 1 | 1 | 0% | 1,733 | 806 | -53% | 0 | 0 | — |
case-18 | pass→pass | 9,818 | 2,199 | -78% | 1 | 1 | 0% | 1,560 | 590 | -62% | 0 | 0 | — |
case-19 | fail→pass | 4,943 | 4,991 | +1% | 1 | 1 | 0% | 734 | 1,140 | +55% | 0 | 0 | — |
case-20 | fail→pass | 14,549 | 1,585 | -89% | 1 | 1 | 0% | 2,205 | 480 | -78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +77 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/20/2026 | +26% |
Other measured skills in the registry, with their headline benchmark lift.