Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when reviewing a Gutenberg change to a WordPress Design System package or its public contract, including `@wordpress/components`, `@wordpress/ui`, or `@wordpress/theme`; do not use to implement the change or review a consumer-only application.
.claude/skills/wordpress-design-system-code-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-14 | ✓→✗ | ▼ Worse | -30% | 0% |
and the applicable package source guidance.
post-change state, target source is its baseline, and MCP is supplementary current-design context.
For an internal change, verify contract preservation and focused coverage, then skip the public-only work below. For a public change:
as sufficient;
old and new accepted values, semantics, states, interaction, and styling in a compact contract table; complete the comparison even after finding one valid defect;
For either classification, use the public guide's package completion gate and inspect only the surfaces applicable to the change.
If a published package can run with a dependency supplied separately by WordPress, apply the package-runtime-compatibility skill before concluding the review.
Use browser evidence when source or class assertions cannot establish visual, focus, motion, or layout parity.
Before reporting a finding, identify the exact changed line, affected public contract or behaviour, target-source or consumer evidence, and concrete impact. Treat incomplete diff context as a verification gap unless the complete patch proves the defect. Apply the same evidence and precision standard even when another valid defect already exists.
Recheck every finding against the complete diff and source. Separate defects, verification gaps, and optional follow-ups. Report material findings with proportional severity and the smallest coherent direction; report no findings when the evidence exposes none.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,952 | 11,377 | -33% | 1 | 1 | 0% | 1,757 | 1,452 | -17% | 0 | 0 | — |
case-02 | fail→pass | 12,694 | 8,479 | -33% | 1 | 1 | 0% | 1,972 | 995 | -50% | 0 | 0 | — |
case-03 | fail→fail | 17,974 | 10,626 | -41% | 1 | 1 | 0% | 1,560 | 1,193 | -24% | 0 | 0 | — |
case-04 | pass→pass | 13,709 | 12,280 | -10% | 1 | 1 | 0% | 1,949 | 2,218 | +14% | 0 | 0 | — |
case-05 | pass→pass | 17,724 | 15,209 | -14% | 1 | 1 | 0% | 1,657 | 1,806 | +9% | 0 | 0 | — |
case-06 | pass→pass | 16,063 | 11,801 | -27% | 1 | 1 | 0% | 2,136 | 1,467 | -31% | 0 | 0 | — |
case-07 | pass→pass | 15,493 | 10,619 | -31% | 1 | 1 | 0% | 1,391 | 1,155 | -17% | 0 | 0 | — |
case-08 | pass→pass | 13,247 | 7,330 | -45% | 1 | 1 | 0% | 1,907 | 1,580 | -17% | 0 | 0 | — |
case-09 | pass→pass | 10,072 | 12,320 | +22% | 1 | 1 | 0% | 1,429 | 1,439 | +1% | 0 | 0 | — |
case-10 | pass→pass | 11,914 | 10,544 | -11% | 1 | 1 | 0% | 1,755 | 1,293 | -26% | 0 | 0 | — |
case-11 | pass→pass | 12,525 | 6,031 | -52% | 1 | 1 | 0% | 1,209 | 1,296 | +7% | 0 | 0 | — |
case-12 | pass→pass | 14,378 | 10,021 | -30% | 1 | 1 | 0% | 1,299 | 1,193 | -8% | 0 | 0 | — |
case-13 | fail→pass | 17,393 | 10,784 | -38% | 1 | 1 | 0% | 1,705 | 1,510 | -11% | 0 | 0 | — |
case-14 | pass→fail | 15,002 | 3,571 | -76% | 1 | 1 | 0% | 1,338 | 932 | -30% | 0 | 0 | — |
case-15 | pass→pass | 16,867 | 12,320 | -27% | 1 | 1 | 0% | 2,569 | 2,259 | -12% | 0 | 0 | — |
case-16 | pass→pass | 18,622 | 14,068 | -24% | 1 | 1 | 0% | 1,855 | 1,672 | -10% | 0 | 0 | — |
case-17 | pass→pass | 15,412 | 14,178 | -8% | 1 | 1 | 0% | 2,000 | 1,802 | -10% | 0 | 0 | — |
case-18 | fail→pass | 17,976 | 11,467 | -36% | 1 | 1 | 0% | 1,877 | 1,276 | -32% | 0 | 0 | — |
case-19 | pass→pass | 32,324 | 37,465 | +16% | 1 | 1 | 0% | 6,742 | 6,749 | +0% | 0 | 0 | — |
case-20 | fail→fail | 12,117 | 10,847 | -10% | 1 | 1 | 0% | 1,063 | 816 | -23% | 0 | 0 | — |
case-21 | fail→fail | 15,553 | 10,612 | -32% | 1 | 1 | 0% | 1,559 | 1,323 | -15% | 0 | 0 | — |
case-22 | fail→fail | 17,207 | 2,535 | -85% | 1 | 1 | 0% | 1,708 | 869 | -49% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.