Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build or audit a design system including component library, design tokens, naming conventions, contribution model, and documentation. Use this skill whenever the user wants to build a design system, audit an existing system, define design tokens at the system level, structure a component library, or set up design system governance. Triggers on design system, component library, design tokens, atomic design, atoms, molecules, organisms, design system documentation, Storybook, Figma library, system
.claude/skills/rampstackco-design-system/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-15 | ✓→✗ | ▼ Worse | 81% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 122% | 0% |
Build, evolve, or audit a design system. Stack-agnostic in principle. Implementation is stack-specific (Figma, Storybook, code library, etc.) but the structure and governance principles transfer.
This skill is for building the system. For applying a system to specific pages or components, use design-standards. For brand visual identity, use brand-identity.
design-standards)brand-identity)frontend-component-build)brand-style-guide)If brand identity is undefined, run brand-identity first.
A complete design system has five layers, stacked. Each layer feeds the layer above.
The atomic decisions. Color, type, spacing, radius, shadow, motion, breakpoints.
Why this layer matters:
Output:
design-standards/references/design-tokens-template.md)Common patterns:
color-blue-600 (base) + color-text-link (semantic). Components reference semantic tokens. Theme changes update semantic tokens, not base.The smallest functional building blocks. Buttons, inputs, labels, badges, icons, links, dividers.
Per element, document:
Output:
Combinations of elements that form recognizable UI patterns. Cards, alerts, modals, navigation, forms, data tables, headers, footers.
Per component:
Output:
Larger structures that combine components. Sign-in flow, settings page, dashboard layout, marketing page sections.
Per pattern:
Output:
How the system gets used, contributed to, and maintained.
Documentation includes:
Governance includes:
A design system has multiple deliverables. Typically:
For a design system audit, output is a markdown report at design-system-audit.md:
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/system-architecture.md - The four-layer model (tokens, primitives, patterns, templates) and how to decide where new work belongs.references/system-audit-template.md - Template for auditing an existing design system.references/governance-playbook.md - Contribution model, ownership, and decision process for an active system.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | pass→pass | 9,437 | 7,875 | -17% | 1 | 1 | 0% | 1,457 | 3,235 | +122% | 0 | 0 | — |
case-01 | fail→pass | 34,483 | 31,337 | -9% | 1 | 1 | 0% | 5,402 | 6,911 | +28% | 0 | 0 | — |
case-02 | fail→fail | 32,197 | 31,667 | -2% | 1 | 1 | 0% | 6,213 | 8,245 | +33% | 0 | 0 | — |
case-03 | pass→pass | 18,978 | 24,692 | +30% | 1 | 1 | 0% | 3,413 | 6,548 | +92% | 0 | 0 | — |
case-04 | fail→fail | 24,637 | 23,866 | -3% | 1 | 1 | 0% | 3,646 | 6,116 | +68% | 0 | 0 | — |
case-05 | pass→pass | 24,934 | 27,083 | +9% | 1 | 1 | 0% | 5,230 | 8,199 | +57% | 0 | 0 | — |
case-06 | fail→pass | 16,160 | 12,787 | -21% | 1 | 1 | 0% | 2,655 | 4,373 | +65% | 0 | 0 | — |
case-08 | fail→fail | 11,891 | 10,330 | -13% | 1 | 1 | 0% | 1,743 | 3,557 | +104% | 0 | 0 | — |
case-09 | fail→pass | 12,720 | 3,698 | -71% | 1 | 1 | 0% | 2,165 | 2,600 | +20% | 0 | 0 | — |
case-10 | fail→fail | 16,415 | 15,951 | -3% | 1 | 1 | 0% | 2,524 | 4,695 | +86% | 0 | 0 | — |
case-11 | pass→pass | 16,487 | 10,864 | -34% | 1 | 1 | 0% | 2,659 | 3,765 | +42% | 0 | 0 | — |
case-12 | pass→pass | 18,631 | 15,445 | -17% | 1 | 1 | 0% | 2,775 | 4,269 | +54% | 0 | 0 | — |
case-13 | pass→pass | 12,058 | 12,068 | +0% | 1 | 1 | 0% | 1,861 | 3,802 | +104% | 0 | 0 | — |
case-14 | pass→pass | 18,196 | 13,491 | -26% | 1 | 1 | 0% | 2,949 | 4,113 | +39% | 0 | 0 | — |
case-15 | pass→fail | 11,378 | 8,912 | -22% | 1 | 1 | 0% | 1,892 | 3,420 | +81% | 0 | 0 | — |
case-16 | pass→pass | 9,836 | 3,891 | -60% | 1 | 1 | 0% | 1,584 | 2,641 | +67% | 0 | 0 | — |
case-17 | pass→pass | 14,816 | 14,989 | +1% | 1 | 1 | 0% | 2,224 | 4,275 | +92% | 0 | 0 | — |
case-18 | pass→pass | 14,090 | 12,467 | -12% | 1 | 1 | 0% | 2,201 | 3,927 | +78% | 0 | 0 | — |
case-19 | pass→pass | 15,699 | 15,475 | -1% | 1 | 1 | 0% | 2,279 | 4,384 | +92% | 0 | 0 | — |
case-20 | pass→pass | 12,808 | 7,224 | -44% | 1 | 1 | 0% | 1,924 | 3,208 | +67% | 0 | 0 | — |
case-21 | pass→pass | 13,797 | 13,917 | +1% | 1 | 1 | 0% | 2,090 | 4,145 | +98% | 0 | 0 | — |
case-22 | pass→pass | 12,233 | 16,379 | +34% | 1 | 1 | 0% | 1,926 | 4,217 | +119% | 0 | 0 | — |
case-23 | fail→fail | 16,382 | 13,458 | -18% | 1 | 1 | 0% | 2,533 | 4,078 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.