Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Start here for Spec Kitty. Orient CLI users and supported agent-harness users; choose the right command, skill family, and recovery path.
.claude/skills/priivacy-ai-spk-start-here/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 89% | 9 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -49% | 0% |
Orient the user by where they are working: command line or supported agent harness.
spec-kitty CLI commands. spk-* skills are not CLIcommands.
/spec-kitty.* in slash-command hosts, $spec-kitty.<command> in Codex, or the host's skill syntax for skill hosts.
spk-* skills are agent operating guides. Use them by name in askill-aware harness or let natural-language intent trigger them.
spk-admin-setup-doctor.Spec Kitty has three visible layers:
/spec-kitty.* slash commands and CLI commands that create oradvance mission artifacts.
spk-* operating guides that teach an agent how to use the product.demand by the runtime or by specialist skills.
Do not turn this into a full tutorial. Identify entry mode first, then route to the smallest useful workflow.
spk-admin-setup-doctor.spk-start-first-feature.spk-start-command-map.spk-start-agent-surface.spk-run-next.spk-run-program-orchestrate.spk-run-review-wp, then spk-gate-accept.spk-team-sync or spk-team-tracker.spk-doctrine-* family.spk-doctrine-show-me.spk-meta-skill-map.State the route, load only the routed skill, and continue. Prefer commands and runtime outputs over guessing from files. If the user asks for a plan, give the shortest workflow that gets them to the next concrete Spec Kitty action.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 38,569 | 35,048 | -9% | 1 | 1 | 0% | 1,395 | 1,424 | +2% | 0 | 0 | — |
case-02 | fail→pass | 17,832 | 8,226 | -54% | 1 | 1 | 0% | 2,862 | 2,032 | -29% | 0 | 0 | — |
case-03 | fail→pass | 14,802 | 6,251 | -58% | 1 | 1 | 0% | 2,164 | 1,702 | -21% | 0 | 0 | — |
case-04 | fail→pass | 6,550 | 5,683 | -13% | 1 | 1 | 0% | 1,011 | 1,470 | +45% | 0 | 0 | — |
case-05 | fail→pass | 9,565 | 2,903 | -70% | 1 | 1 | 0% | 1,436 | 1,106 | -23% | 0 | 0 | — |
case-06 | fail→pass | 15,454 | 4,017 | -74% | 1 | 1 | 0% | 2,604 | 1,336 | -49% | 0 | 0 | — |
case-07 | fail→pass | 6,988 | 1,895 | -73% | 1 | 1 | 0% | 1,142 | 882 | -23% | 0 | 0 | — |
case-13 | fail→pass | 10,370 | 4,441 | -57% | 1 | 1 | 0% | 1,683 | 1,348 | -20% | 0 | 0 | — |
case-08 | fail→pass | 6,866 | 3,148 | -54% | 1 | 1 | 0% | 924 | 1,138 | +23% | 0 | 0 | — |
case-09 | fail→pass | 9,752 | 3,159 | -68% | 1 | 1 | 0% | 1,543 | 1,033 | -33% | 0 | 0 | — |
case-10 | fail→pass | 10,812 | 4,182 | -61% | 1 | 1 | 0% | 1,556 | 1,296 | -17% | 0 | 0 | — |
case-11 | fail→pass | 12,350 | 2,558 | -79% | 1 | 1 | 0% | 2,066 | 979 | -53% | 0 | 0 | — |
case-12 | fail→pass | 9,275 | 2,804 | -70% | 1 | 1 | 0% | 1,369 | 977 | -29% | 0 | 0 | — |
case-14 | fail→pass | 11,258 | 5,546 | -51% | 1 | 1 | 0% | 1,739 | 1,481 | -15% | 0 | 0 | — |
case-15 | fail→pass | 9,223 | 3,326 | -64% | 1 | 1 | 0% | 1,499 | 1,108 | -26% | 0 | 0 | — |
case-16 | pass→pass | 9,427 | 4,727 | -50% | 1 | 1 | 0% | 1,586 | 1,380 | -13% | 0 | 0 | — |
case-17 | pass→pass | 9,957 | 3,723 | -63% | 1 | 1 | 0% | 1,550 | 1,151 | -26% | 0 | 0 | — |
case-18 | fail→fail | 7,575 | 1,523 | -80% | 1 | 1 | 0% | 1,287 | 794 | -38% | 0 | 0 | — |
case-19 | fail→pass | 12,905 | 5,034 | -61% | 1 | 1 | 0% | 2,209 | 1,521 | -31% | 0 | 0 | — |
case-20 | fail→pass | 10,040 | 2,062 | -79% | 1 | 1 | 0% | 1,556 | 889 | -43% | 0 | 0 | — |
case-21 | fail→pass | 10,248 | 3,294 | -68% | 1 | 1 | 0% | 1,788 | 1,242 | -31% | 0 | 0 | — |
case-22 | pass→pass | 10,923 | 9,317 | -15% | 1 | 1 | 0% | 2,042 | 2,372 | +16% | 0 | 0 | — |
case-23 | pass→fail | 11,537 | 10,038 | -13% | 1 | 1 | 0% | 2,278 | 2,450 | +8% | 0 | 0 | — |
case-24 | pass→fail | 5,899 | 4,874 | -17% | 1 | 1 | 0% | 1,170 | 1,519 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +63 percentage points is the difference between those two pass rates over the 24 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +73% |
Other measured skills in the registry, with their headline benchmark lift.