Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Identify, categorize, and prioritize accumulated design inconsistencies and structural problems across a product.
.claude/skills/owl-listener-design-debt-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 47% | 0% |
You are an expert in systematically identifying and triaging design debt before it becomes structural.
You conduct design debt audits that surface inconsistencies, outdated patterns, accessibility gaps, and structural problems — and produce a prioritized remediation plan that teams can act on.
Design debt is any gap between the current state of the product and the standard it should meet. Categories:
For each screen or component, tag:
Score debt items using: Severity × Frequency / Effort Surface a short list of high-priority items — the debt that's causing the most harm per unit of effort to fix.
Maintain a living document (not a one-time audit) tracking:
Review the register quarterly; update severity as the product changes.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 19,262 | 22,062 | +15% | 1 | 1 | 0% | 3,800 | 4,903 | +29% | 0 | 0 | — |
case-02 | fail→fail | 23,022 | 24,615 | +7% | 1 | 1 | 0% | 3,531 | 4,951 | +40% | 0 | 0 | — |
case-03 | pass→pass | 10,813 | 11,374 | +5% | 1 | 1 | 0% | 1,912 | 2,814 | +47% | 0 | 0 | — |
case-04 | pass→pass | 11,266 | 9,560 | -15% | 1 | 1 | 0% | 1,838 | 2,544 | +38% | 0 | 0 | — |
case-05 | pass→pass | 8,436 | 9,242 | +10% | 1 | 1 | 0% | 1,586 | 2,518 | +59% | 0 | 0 | — |
case-06 | pass→pass | 10,301 | 8,300 | -19% | 1 | 1 | 0% | 1,847 | 2,534 | +37% | 0 | 0 | — |
case-07 | fail→fail | 12,405 | 13,513 | +9% | 1 | 1 | 0% | 2,180 | 3,060 | +40% | 0 | 0 | — |
case-08 | fail→fail | 11,408 | 11,325 | -1% | 1 | 1 | 0% | 1,911 | 2,936 | +54% | 0 | 0 | — |
case-09 | fail→fail | 15,372 | 13,888 | -10% | 1 | 1 | 0% | 2,551 | 3,128 | +23% | 0 | 0 | — |
case-10 | pass→pass | 8,032 | 13,278 | +65% | 1 | 1 | 0% | 1,331 | 2,803 | +111% | 0 | 0 | — |
case-11 | pass→pass | 9,311 | 8,526 | -8% | 1 | 1 | 0% | 1,465 | 2,190 | +49% | 0 | 0 | — |
case-12 | pass→pass | 16,463 | 16,935 | +3% | 1 | 1 | 0% | 2,625 | 3,700 | +41% | 0 | 0 | — |
case-13 | pass→pass | 13,297 | 11,307 | -15% | 1 | 1 | 0% | 2,061 | 2,987 | +45% | 0 | 0 | — |
case-14 | pass→pass | 11,594 | 11,085 | -4% | 1 | 1 | 0% | 1,851 | 2,664 | +44% | 0 | 0 | — |
case-24 | pass→pass | 14,316 | 16,556 | +16% | 1 | 1 | 0% | 2,698 | 3,954 | +47% | 0 | 0 | — |
case-15 | pass→pass | 11,737 | 11,740 | +0% | 1 | 1 | 0% | 1,926 | 2,731 | +42% | 0 | 0 | — |
case-16 | fail→pass | 11,011 | 10,854 | -1% | 1 | 1 | 0% | 1,754 | 2,636 | +50% | 0 | 0 | — |
case-17 | fail→fail | 15,897 | 15,134 | -5% | 1 | 1 | 0% | 2,584 | 3,308 | +28% | 0 | 0 | — |
case-18 | pass→pass | 8,295 | 9,446 | +14% | 1 | 1 | 0% | 1,460 | 2,337 | +60% | 0 | 0 | — |
case-19 | fail→pass | 9,280 | 10,516 | +13% | 1 | 1 | 0% | 1,572 | 2,620 | +67% | 0 | 0 | — |
case-20 | fail→pass | 11,279 | 9,599 | -15% | 1 | 1 | 0% | 1,889 | 2,624 | +39% | 0 | 0 | — |
case-21 | fail→pass | 12,444 | 12,448 | +0% | 1 | 1 | 0% | 2,041 | 2,815 | +38% | 0 | 0 | — |
case-22 | pass→pass | 11,787 | 9,258 | -21% | 1 | 1 | 0% | 2,039 | 2,382 | +17% | 0 | 0 | — |
case-23 | fail→fail | 8,559 | 3,069 | -64% | 1 | 1 | 0% | 1,487 | 1,358 | -9% | 0 | 0 | — |
case-25 | pass→pass | 10,614 | 10,798 | +2% | 1 | 1 | 0% | 2,515 | 3,191 | +27% | 0 | 0 | — |
case-26 | pass→pass | 8,005 | 8,096 | +1% | 1 | 1 | 0% | 1,571 | 2,384 | +52% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 26 cases were attempted. The headline lift of +15 percentage points is the difference between those two pass rates over the 26 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.