Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Explain GRACE 4 methodology, .grace artifacts, semantic anchors, change lifecycle, verification, and migration boundaries.
.claude/skills/osovv-grace-explainer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -47% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -41% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -61% | 0% |
<skill> <core_model> GRACE 4 uses .grace as the durable project model:
.grace/context stores requirements, technology, principles, deployment, and UX constraints..grace/graph stores graph indexes and routed graph documents with GD-*, M-*, and DF-* tags..grace/verification stores verification indexes and routed V-M-* entries..grace/changes stores active and archived C-* change bundles with GraceChangeSpec, optional non-normative design context, and GraceChangePlan.</core_model>
<workflow>
grace-init creates the .grace skeleton.grace-spec creates an active change spec and waits for approval.grace-plan creates assertions, scopes, and T-* implementation tasks.grace-execute runs sequential or parallel-safe mode from the approved plan.grace lint and grace status provide validation and health evidence.grace-migrate; the CLI validates the result but does not convert legacy docs directly.</workflow> </skill>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,004 | 12,238 | -6% | 1 | 1 | 0% | 2,075 | 1,467 | -29% | 0 | 0 | — |
case-02 | fail→pass | 13,774 | 4,704 | -66% | 1 | 1 | 0% | 2,198 | 1,157 | -47% | 0 | 0 | — |
case-03 | fail→pass | 17,579 | 8,496 | -52% | 1 | 1 | 0% | 2,550 | 1,747 | -31% | 0 | 0 | — |
case-04 | fail→pass | 14,152 | 5,423 | -62% | 1 | 1 | 0% | 2,219 | 1,312 | -41% | 0 | 0 | — |
case-05 | fail→pass | 10,451 | 2,230 | -79% | 1 | 1 | 0% | 1,834 | 707 | -61% | 0 | 0 | — |
case-06 | fail→pass | 10,713 | 1,625 | -85% | 1 | 1 | 0% | 1,949 | 527 | -73% | 0 | 0 | — |
case-07 | fail→pass | 9,302 | 1,551 | -83% | 1 | 1 | 0% | 1,617 | 553 | -66% | 0 | 0 | — |
case-08 | fail→pass | 10,710 | 2,096 | -80% | 1 | 1 | 0% | 1,726 | 638 | -63% | 0 | 0 | — |
case-09 | fail→pass | 6,575 | 2,599 | -60% | 1 | 1 | 0% | 1,086 | 691 | -36% | 0 | 0 | — |
case-10 | pass→pass | 10,753 | 3,853 | -64% | 1 | 1 | 0% | 1,848 | 948 | -49% | 0 | 0 | — |
case-11 | fail→pass | 6,223 | 1,500 | -76% | 1 | 1 | 0% | 1,236 | 539 | -56% | 0 | 0 | — |
case-12 | fail→pass | 7,351 | 2,876 | -61% | 1 | 1 | 0% | 1,229 | 798 | -35% | 0 | 0 | — |
case-13 | fail→pass | 12,970 | 4,673 | -64% | 1 | 1 | 0% | 2,047 | 1,103 | -46% | 0 | 0 | — |
case-14 | fail→pass | 8,523 | 1,713 | -80% | 1 | 1 | 0% | 1,412 | 516 | -63% | 0 | 0 | — |
case-15 | fail→pass | 7,514 | 1,456 | -81% | 1 | 1 | 0% | 1,271 | 488 | -62% | 0 | 0 | — |
case-16 | fail→pass | 8,657 | 1,882 | -78% | 1 | 1 | 0% | 1,457 | 614 | -58% | 0 | 0 | — |
case-17 | fail→pass | 14,350 | 2,008 | -86% | 1 | 1 | 0% | 2,591 | 595 | -77% | 0 | 0 | — |
case-18 | pass→pass | 5,372 | 2,872 | -47% | 1 | 1 | 0% | 974 | 768 | -21% | 0 | 0 | — |
case-19 | fail→pass | 12,214 | 4,363 | -64% | 1 | 1 | 0% | 2,251 | 1,014 | -55% | 0 | 0 | — |
case-20 | fail→pass | 7,098 | 3,934 | -45% | 1 | 1 | 0% | 1,277 | 977 | -23% | 0 | 0 | — |
case-21 | pass→pass | 13,487 | 9,520 | -29% | 1 | 1 | 0% | 2,934 | 2,368 | -19% | 0 | 0 | — |
case-22 | pass→pass | 12,903 | 11,982 | -7% | 1 | 1 | 0% | 2,544 | 2,637 | +4% | 0 | 0 | — |
case-23 | pass→pass | 10,464 | 9,169 | -12% | 1 | 1 | 0% | 2,018 | 2,087 | +3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +78 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.