Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use for standard or complex new work before coding or planning. Also handles vague goals — clarifies before designing.
.claude/skills/hashgraph-online-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 17% | 0% |
<gate> No code until user approves the spec. </gate>
Goal too vague to name what to build, for whom, or what success looks like? Ask one question per turn until it's concrete. Don't propose solutions until then. Working notes can hold hypotheses, experiments, ruled-out directions (spike code → temporary worktree).
Clarify in dependency order. Resolve facts from the repo/tools; ask only the current decision frontier requiring user judgment. Don't ask downstream questions before prerequisite decisions or map the full dependency tree. Stop when implementation-affecting contract, data, failure, and test decisions are decided or deferred.
Goal clear? Propose 2-3 approaches with trade-offs; recommend one. Then write the spec.
If topology=multi-module (triage announcement or coordinator spec/plan declaration), read ../references/multi-module.md before designing. Create the coordinator spec with the declaration block (topology: multi-module, change-set, coordinator, repos) - this declaration is the on-disk mode marker downstream skills rely on.
A spec answers the open questions for THIS change. Typical:
Do spec idiomatically. Record convention — stack best practices + project conventions for this change (see ../references/quality.md); tdd/review verify against it.
No question → no section. Don't fill "Risks" / "Non-goals" if empty.
Use declarations, not narrative:
contract: <interface>
invariant: <what must hold>
test: <how we'll know>
convention: <stack best practices + project conventions — see ../references/quality.md>
deferred: <not deciding now>Reference code by path; never paste it.
Before handoff, close only decisions that affect implementation: contract, data, failure, test. Unresolved Working notes in those areas become decisions, deferred, or questions.
docs/staging/specs/YYYY-MM-DD-<topic>.md## Working notes: scratch, open questions, hypotheses, ruled-out directions (stripped at ship).docs/ROADMAP.md already exists?
docs/ROADMAP.md.docs/ROADMAP.md does not exist?
docs/ROADMAP.md:markdown
Stubs are intent, not commitment; update before expanding.
If roadmap exists or was created, reference the current milestone in staging spec:
milestone: M1 (see docs/ROADMAP.md)After spec is written to disk and before handing off to plan, inspect the spec to decide how many reviewers to dispatch. Read ../references/reviewers.md for the trigger table and reviewer charters.
Count the triggers that match the spec. If only trigger 1 fires (the baseline), dispatch a single spec-compliance reviewer — this is the default behavior, same cost as today. If multiple triggers fire, dispatch each reviewer as an independent subagent in parallel.
All reviewers receive the spec path (not the spec content — let them read it). Collect findings, then synthesize: evaluate each finding critically, resolve conflicts between reviewers, spot what was missed, judge severity, and patch the spec once. Present to user: what was found, fixed, deferred — and why.
Report to user: which triggers fired, which reviewers ran, what was found and fixed.
If the user decides not to proceed after clarification, stop here. No spec, no plan, no ship. Record reason briefly in working notes. If exploration produced a knowledge artifact (protocol spec, RE findings, data structure map), save it to docs/decisions/ via archive.
<gate>
docs/staging/specs/YYYY-MM-DD-<topic>.md must exist on disk before handing off to plan. For multi-module work, the coordinator spec and every affected module spec must exist in their owning repositories.</gate>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 46,683 | 13,554 | -71% | 1 | 1 | 0% | 6,604 | 1,535 | -77% | 0 | 0 | — |
case-02 | fail→fail | 8,047 | 6,052 | -25% | 1 | 1 | 0% | 1,425 | 1,565 | +10% | 0 | 0 | — |
case-03 | fail→fail | 12,641 | 15,796 | +25% | 1 | 1 | 0% | 1,265 | 1,353 | +7% | 0 | 0 | — |
case-04 | fail→fail | 18,388 | 4,529 | -75% | 1 | 1 | 0% | 2,481 | 1,567 | -37% | 0 | 0 | — |
case-05 | fail→fail | 8,783 | 14,499 | +65% | 1 | 1 | 0% | 477 | 1,433 | +200% | 0 | 0 | — |
case-06 | fail→fail | 28,103 | 14,494 | -48% | 1 | 1 | 0% | 3,228 | 1,461 | -55% | 0 | 0 | — |
case-07 | fail→pass | 11,594 | 8,441 | -27% | 1 | 1 | 0% | 2,081 | 1,731 | -17% | 0 | 0 | — |
case-08 | fail→pass | 16,862 | 2,674 | -84% | 1 | 1 | 0% | 1,894 | 1,502 | -21% | 0 | 0 | — |
case-09 | fail→fail | 14,873 | 9,659 | -35% | 1 | 1 | 0% | 2,044 | 1,927 | -6% | 0 | 0 | — |
case-10 | pass→pass | 19,362 | 13,440 | -31% | 1 | 1 | 0% | 2,994 | 2,510 | -16% | 0 | 0 | — |
case-15 | fail→pass | 11,937 | 2,534 | -79% | 1 | 1 | 0% | 1,770 | 1,556 | -12% | 0 | 0 | — |
case-11 | pass→pass | 12,522 | 9,566 | -24% | 1 | 1 | 0% | 1,695 | 1,877 | +11% | 0 | 0 | — |
case-12 | fail→pass | 12,640 | 8,963 | -29% | 1 | 1 | 0% | 2,295 | 1,735 | -24% | 0 | 0 | — |
case-13 | fail→fail | 19,172 | 11,166 | -42% | 1 | 1 | 0% | 2,325 | 1,609 | -31% | 0 | 0 | — |
case-14 | pass→fail | 13,982 | 13,631 | -3% | 1 | 1 | 0% | 2,284 | 1,666 | -27% | 0 | 0 | — |
case-16 | pass→pass | 10,900 | 8,729 | -20% | 1 | 1 | 0% | 1,085 | 1,720 | +59% | 0 | 0 | — |
case-17 | fail→pass | 17,542 | 4,178 | -76% | 1 | 1 | 0% | 1,614 | 1,891 | +17% | 0 | 0 | — |
case-18 | fail→pass | 4,506 | 3,112 | -31% | 1 | 1 | 0% | 654 | 1,498 | +129% | 0 | 0 | — |
case-19 | pass→pass | 9,593 | 7,578 | -21% | 1 | 1 | 0% | 1,677 | 1,587 | -5% | 0 | 0 | — |
case-20 | pass→pass | 9,500 | 7,441 | -22% | 1 | 1 | 0% | 1,587 | 1,522 | -4% | 0 | 0 | — |
case-21 | pass→pass | 19,640 | 4,379 | -78% | 1 | 1 | 0% | 2,251 | 1,677 | -25% | 0 | 0 | — |
case-22 | fail→pass | 16,240 | 7,926 | -51% | 1 | 1 | 0% | 1,677 | 1,594 | -5% | 0 | 0 | — |
case-23 | fail→pass | 16,356 | 3,770 | -77% | 1 | 1 | 0% | 1,501 | 1,593 | +6% | 0 | 0 | — |
case-24 | fail→pass | 15,589 | 10,348 | -34% | 1 | 1 | 0% | 1,318 | 1,489 | +13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 16 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +33 percentage points is the difference between those two pass rates over the 16 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.