Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at ship time to merge the spec into the living documentation, delete the staging spec and plan, and ask how to finish.
.claude/skills/hashgraph-online-archive/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -34% | 0% |
Always in sync with code. Current ground truth only — atemporal, no history, no plans.
README.md — for users: what it is, who for, how to usedocs/tech-spec.md — for developers/agents: current system state (core backbone, ≤300 lines)docs/specs/<topic>.md — modular subsystem/protocol details; referenced by pathdocs/ROADMAP.md — direction (≥3 milestones or long-term)CHANGELOG.md — version history (ship maintains)docs/decisions/ — architectural rationale (context / choice / ruled-out), append-only> Standard: Zero history, maximum truth density, instant scannability. > - The Razor: Past rationale belongs in docs/decisions/ or CHANGELOG.md; docs/tech-spec.md holds exclusively active ground truth.
docs/specs/<topic>.md, referenced by a one-line declaration in tech-spec.md.Declarations only.
purpose: <what problem this solves — one sentence>
user: <who uses this>
use-case: <key scenarios, one line each>
architecture: <structural shape — one line, or see docs/architecture.md>
stack: <language, runtime, frameworks, key deps>
entry: <where execution starts>
contract: <public APIs / interfaces that must not break — stability set; full surface by doc-coverage>
flow: <name>: <trigger> → <steps> → <output>
(complex — branching/async/multi-actor: one-line summary here, diagram in docs/specs/<flow>.md)
invariant: <what must always hold>
constraint: <limits, warnings from code>
convention: <naming, file structure, test patterns, lint/format/typecheck tools, error-handling, security baseline>
milestone: <current milestone> (see docs/ROADMAP.md)If topology=multi-module (triage announcement or coordinator spec/plan declaration), read ../references/multi-module.md. Merge each module's artifacts in its own repository first. Then merge shared contracts and integration facts in the coordinator, record the revision set, and archive it last. Keep module-local facts in module repos.
<gate> Before proceeding: (1) verify tdd/subagents have completed all tasks listed in the plan; (2) user has approved shipping/archiving this change (ship disposition or an explicit archive go-ahead). </gate>
docs/specs/<topic>.md.## Roadmap): update roadmap independently; do not duplicate into tech-spec.docs/decisions/YYYY-MM-DD-<topic>.md as context / choice / ruled-out.<gate> Before deleting staging spec/plan, verify Living Doc North Star: (1) Atemporal (zero history/dates/supersessions)? (2) Truth-dense (binding facts only)? (3) Modular (subsystems >15 lines in docs/specs/)? (4) User approved merged living-doc. </gate>
docs/staging/specs/YYYY-MM-DD-<topic>.md — content absorbed; Git has the history.docs/staging/plans/YYYY-MM-DD-<topic>.md — plans don't belong on main.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | fail→pass | 17,689 | 10,290 | -42% | 1 | 1 | 0% | 2,038 | 1,861 | -9% | 0 | 0 | — |
case-01 | fail→fail | 15,647 | 16,881 | +8% | 1 | 1 | 0% | 238 | 1,442 | +506% | 0 | 0 | — |
case-02 | fail→fail | 6,581 | 5,937 | -10% | 1 | 1 | 0% | 1,131 | 1,355 | +20% | 0 | 0 | — |
case-03 | fail→fail | 5,086 | 10,207 | +101% | 1 | 1 | 0% | 302 | 1,234 | +309% | 0 | 0 | — |
case-04 | pass→pass | 24,841 | 19,321 | -22% | 1 | 1 | 0% | 3,546 | 3,384 | -5% | 0 | 0 | — |
case-05 | pass→pass | 10,343 | 8,439 | -18% | 1 | 1 | 0% | 1,725 | 2,284 | +32% | 0 | 0 | — |
case-06 | pass→pass | 17,333 | 18,944 | +9% | 1 | 1 | 0% | 2,764 | 3,286 | +19% | 0 | 0 | — |
case-07 | fail→pass | 16,540 | 17,508 | +6% | 1 | 1 | 0% | 1,874 | 1,689 | -10% | 0 | 0 | — |
case-08 | fail→pass | 11,775 | 7,326 | -38% | 1 | 1 | 0% | 1,937 | 1,260 | -35% | 0 | 0 | — |
case-09 | fail→pass | 15,429 | 11,407 | -26% | 1 | 1 | 0% | 1,631 | 2,140 | +31% | 0 | 0 | — |
case-10 | pass→pass | 17,958 | 4,270 | -76% | 1 | 1 | 0% | 2,004 | 1,705 | -15% | 0 | 0 | — |
case-11 | fail→pass | 13,327 | 7,531 | -43% | 1 | 1 | 0% | 2,154 | 1,411 | -34% | 0 | 0 | — |
case-12 | pass→pass | 10,373 | 1,751 | -83% | 1 | 1 | 0% | 1,655 | 1,194 | -28% | 0 | 0 | — |
case-13 | fail→pass | 8,680 | 2,881 | -67% | 1 | 1 | 0% | 1,340 | 1,473 | +10% | 0 | 0 | — |
case-14 | fail→pass | 12,859 | 9,535 | -26% | 1 | 1 | 0% | 1,419 | 1,840 | +30% | 0 | 0 | — |
case-15 | fail→pass | 4,982 | 2,970 | -40% | 1 | 1 | 0% | 804 | 1,513 | +88% | 0 | 0 | — |
case-16 | fail→fail | 11,134 | 1,826 | -84% | 1 | 1 | 0% | 1,903 | 1,290 | -32% | 0 | 0 | — |
case-17 | pass→pass | 16,212 | 7,367 | -55% | 1 | 1 | 0% | 1,830 | 1,420 | -22% | 0 | 0 | — |
case-18 | fail→pass | 12,004 | 3,681 | -69% | 1 | 1 | 0% | 1,912 | 1,667 | -13% | 0 | 0 | — |
case-19 | fail→pass | 16,054 | 2,462 | -85% | 1 | 1 | 0% | 1,930 | 1,397 | -28% | 0 | 0 | — |
case-21 | pass→pass | 9,691 | 6,805 | -30% | 1 | 1 | 0% | 734 | 1,276 | +74% | 0 | 0 | — |
case-22 | fail→pass | 14,843 | 5,848 | -61% | 1 | 1 | 0% | 2,253 | 2,102 | -7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.