Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when engineering decisions span ideation, design, development, testing, release, operations, maintenance; API/reliability/security/data/doc lifecycle before process skills
.claude/skills/hashgraph-online-staff-engineer-mode/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 164% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 433% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 598% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 514% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 352% | 0% |
ONE PRIMARY SPECIALIST BY DEFAULT; INFER ROUTING CONTEXT BEFORE WITHHOLDINGLoading many specialists means routing failed.
Engineering architecture, reliability, operations, security, delivery, data, platform, client, AI/ML, accessibility, cost, readiness, rollout, migration, incidents, doc lifecycle, control records, and API/service contracts start here.
Do not invoke broad process skills or host orchestration first. Route through SEM, load the specialist, then use other tools for sub-decisions.
Build, design, reliability, HA, rollout, service review, launch, and incident prompts are engineering work. Doc lifecycle covers owner, truth, freshness, operational accuracy, missing guidance, or archive; routine copy cleanup does not route. Workflow/process/plan are artifacts, not bypasses.
To load a specialist, Read <specialist-root>/<slug>.md. Resolve <specialist-root> in this order:
SPECIALIST_ROOT= appears in session context (Claude Code, Cursor, OpenCode), use it.specialists directory beside the loaded plugin checkout.~/.codex/staff-engineer-mode/specialistsspecialists directory next to the loaded GEMINI.mdspecialists/ directory at the router skill's install root.Four rules, all mandatory:
Skill staff-engineer-mode:<slug> returns Unknown skill and is a routing failure.Classify by artifact, phase, surface, and risk; users need not know specialist names.
Resolve <template-root> from TEMPLATE_ROOT= in session context, otherwise from skills/_shared/assets/templates at the install root containing specialists/. Read its README.md, then choose the smallest owned template set covering the requested artifacts: one template for narrow work and multiple only for distinct artifacts in broad work. Read risk-exception-register.md only when the specialist invokes shared risk acceptance.
Treat Required Outputs as an artifact menu. For narrow requests, render the smallest core artifact plus risk-triggered sections. For broad, high-impact, or comprehensive work, expand each selected template. Keep artifacts user-visible.
Command attempts are event-policy exceptions. Before commits/amends, read agent-pr-review, stage once, inspect the exact staged diff, show the review, record the receipt separately, then commit separately. Do not combine stage/ack/commit/push or add AI attribution. Before release actions, read release-build-reproducibility and production-readiness-review, show both review artifacts, record the receipt separately, then run the release command separately.
Infer from prompt, branch context, conversation, and loaded context. Do not read new repo files before selecting a specialist, and do not ask for intake fields.
Pick primary and secondary only from this exact list. Never invent, shorten, or paraphrase a slug.
accessibility-gates, agent-pr-review, ai-coding-governance, api-design-and-compatibility,
architecture-decisions, backup-and-recovery, caching-and-derived-data,
client-application-security, code-readability-for-agents, configuration-and-automation-safety,
container-runtime-and-orchestration,
cost-aware-reliability, cryptography-and-key-lifecycle, database-operations, data-contracts,
data-lineage-and-provenance, data-pipeline-reliability, dependency-and-code-hygiene, dependency-resilience,
dev-environment-parity, distributed-data-and-consistency, documentation-lifecycle,
edge-traffic-and-ddos-defense, engineering-control-evidence, event-workflows,
experimentation-and-metric-guardrails, feature-flag-lifecycle, fleet-upgrades,
high-availability-design, identity-and-secrets, incident-response-and-postmortems,
infrastructure-and-policy-as-code, input-validation-and-injection-defense,
internal-service-networking, llm-application-security,
llm-evaluation, llm-serving-cost-and-latency, migration-and-deprecation,
ml-reliability-and-evaluation, mobile-release-engineering,
multi-region-and-data-residency, observability-and-alerting,
oncall-health, operational-ownership-transfer, performance-and-capacity,
persistent-connection-systems, platform-golden-paths, privacy-and-data-lifecycle,
production-readiness-review, progressive-delivery, release-build-reproducibility,
resilience-experiments, resilience-requirements, scheduled-job-reliability, secure-sdlc-and-threat-modeling,
service-decommission-and-sunset, slo-and-error-budgets,
software-supply-chain-security, state-machine-correctness, tenant-isolation,
test-data-engineering, testing-and-quality-gates, vulnerability-management,
web-release-gatesagent-pr-review before staging or inspection.primary (and any secondary) verbatim from the Bundled Specialist Slugs list above; if no listed slug fits, withhold routing instead of inventing or paraphrasing one.engineering-control-evidence only for cross-surface mappings, scorecards, exceptions, or control packs.Select one primary when context is enough. Recommend at most one secondary follow-up. Broad requests become a short sequence, not a pile of loaded specialists.
production-readiness-review.incident-response-and-postmortems first, even if root cause seems elsewhere.Treat "review" as a verb until the artifact proves otherwise.
agent-pr-review; general PR, branch, patch, last commit, staged change, or diff review before merge routes there, including tests-pass or deletion-behavior checks.dependency-and-code-hygiene.documentation-lifecycle, even when code and packaging are also inventoried.production-readiness-review.single_primary: output has one primary specialist unless routing is withheld.secondary_cap: output has no more than one secondary specialist.capability_translation: tool, vendor, or framework names are translated into capability language before routing and not repeated in route fields.scope_check: out-of-scope requests are reframed or declined without specialist names.ambiguity_check: ambiguous prompts infer the discriminating artifact before routing; withheld routes expose no specialist names, candidate routes, confidence labels, drafts, or intake questions.intent_inference: rationale identifies the requested artifact and phase before naming a skill.Load references/routing-matrix.md.
production-readiness-review; mobile startup/crash/offline -> mobile-release-engineering; canary metrics -> progressive-delivery; incidents -> incident-response-and-postmortems.agent-pr-review; surface-specific PRs route narrow.architecture-decisions; ownership transfer/handoff -> operational-ownership-transfer; AI repo legibility -> code-readability-for-agents; retry/timeout/fallback/overload -> dependency-resilience.high-availability-design; residency/geo-routing/replication-aware region placement -> multi-region-and-data-residency; fault injection -> resilience-experiments; telemetry -> observability-and-alerting; alert toil or recurring manual runbook work -> oncall-health.resilience-requirements; game days -> resilience-experiments; proven topology -> high-availability-design.container-runtime-and-orchestration; reconnect/heartbeat/fanout -> persistent-connection-systems; raw headroom -> performance-and-capacity.distributed-data-and-consistency; invariant models, counterexamples, and protocol state validation -> state-machine-correctness even for a distributed protocol; restore/corruption recovery -> backup-and-recovery; DB execution/query/schema regression -> database-operations.api-design-and-compatibility alone; broad consumer retirement/no-new-usage -> migration-and-deprecation; cross-surface schemas -> data-contracts; fixtures, events, cache, lineage, and pipelines stay distinct.event-workflows; non-event timeout/retry -> dependency-resilience; schema -> data-contracts; pipeline-freshness/replay -> data-pipeline-reliability; provenance: data-lineage-and-provenance.feature-flag-lifecycle; runtime config mutation -> configuration-and-automation-safety; desired-state drift/reconcile -> infrastructure-and-policy-as-code.release-build-reproducibility; env drift -> dev-environment-parity; provenance/signing/builder isolation -> software-supply-chain-security.migration-and-deprecation; terminal teardown/no-resurrection -> service-decommission-and-sunset; model promotion/drift -> ml-reliability-and-evaluation.production-readiness-review; staged exposure/rollback -> progressive-delivery; build artifact identity -> release-build-reproducibility; browser/mobile gates route client-specific.cryptography-and-key-lifecycle alone; broad threat/control selection -> secure-sdlc-and-threat-modeling; sink injection, identity/secrets, supply-chain trust, deployed flaws, tenant/privacy, and LLM risk route narrow.llm-application-security; eval, retrieval-grounded, or agent task-run checks -> llm-evaluation; serving cost/latency/token/cache/fallback budgets -> llm-serving-cost-and-latency; generic model-provider retry, timeout, circuit-breaker, or overload policy -> dependency-resilience; ML serving reliability -> ml-reliability-and-evaluation.edge-traffic-and-ddos-defense; private service trust-domain/peer-name/identity/transport/routing -> internal-service-networking; access/credentials -> identity-and-secrets; dependency calls -> dependency-resilience.test-data-engineering; CI/merge gates -> testing-and-quality-gates; environment drift -> dev-environment-parity.dependency-and-code-hygiene; fleet waves/support windows -> fleet-upgrades; supply-chain trust stays separate.cost-aware-reliability; raw headroom -> performance-and-capacity; LLM token/tail cost -> llm-serving-cost-and-latency.documentation-lifecycle; routine doc copy edits do not route; AI agent rules -> ai-coding-governance; control packs -> engineering-control-evidence.production-readiness-review is used for any broad prompt without a readiness event.| Mistake | Correction | | --- | --- | | Keyword matching | Infer artifact, phase, surface, and risk. | | Loading every related specialist | Choose one primary; list at most one follow-up. | | Treating tools as domains | Translate tools to capabilities. | | Asking intake too soon | Infer from prompt, repo, files, branch context, and conversation first. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→pass | 19,368 | 12,994 | -33% | 1 | 1 | 0% | 2,188 | 5,773 | +164% | 0 | 0 | — |
case-05 | fail→pass | 11,500 | 10,681 | -7% | 1 | 1 | 0% | 978 | 5,213 | +433% | 0 | 0 | — |
case-01 | fail→fail | 25,571 | 24,281 | -5% | 1 | 1 | 0% | 3,540 | 5,433 | +53% | 0 | 0 | — |
case-02 | fail→fail | 11,685 | 28,642 | +145% | 1 | 1 | 0% | 1,199 | 6,045 | +404% | 0 | 0 | — |
case-03 | fail→fail | 14,681 | 23,629 | +61% | 1 | 1 | 0% | 1,023 | 5,443 | +432% | 0 | 0 | — |
case-06 | fail→pass | 10,727 | 12,294 | +15% | 1 | 1 | 0% | 820 | 5,721 | +598% | 0 | 0 | — |
case-07 | fail→fail | 14,173 | 21,580 | +52% | 1 | 1 | 0% | 1,366 | 5,097 | +273% | 0 | 0 | — |
case-08 | fail→fail | 15,535 | 19,093 | +23% | 1 | 1 | 0% | 1,681 | 5,041 | +200% | 0 | 0 | — |
case-09 | pass→pass | 17,478 | 22,527 | +29% | 1 | 1 | 0% | 1,965 | 5,069 | +158% | 0 | 0 | — |
case-10 | fail→pass | 12,323 | 32,261 | +162% | 1 | 1 | 0% | 1,264 | 7,767 | +514% | 0 | 0 | — |
case-11 | pass→fail | 15,825 | 21,670 | +37% | 1 | 1 | 0% | 1,670 | 5,553 | +233% | 0 | 0 | — |
case-12 | fail→fail | 17,815 | 18,911 | +6% | 1 | 1 | 0% | 2,056 | 5,073 | +147% | 0 | 0 | — |
case-13 | fail→fail | 17,491 | 34,257 | +96% | 1 | 1 | 0% | 2,171 | 6,793 | +213% | 0 | 0 | — |
case-14 | pass→fail | 25,557 | 24,391 | -5% | 1 | 1 | 0% | 3,110 | 5,411 | +74% | 0 | 0 | — |
case-15 | fail→fail | 18,065 | 24,950 | +38% | 1 | 1 | 0% | 2,085 | 6,046 | +190% | 0 | 0 | — |
case-16 | fail→fail | 15,152 | 18,473 | +22% | 1 | 1 | 0% | 1,490 | 5,008 | +236% | 0 | 0 | — |
case-17 | fail→fail | 26,116 | 18,749 | -28% | 1 | 1 | 0% | 3,215 | 5,419 | +69% | 0 | 0 | — |
case-18 | fail→pass | 12,784 | 11,084 | -13% | 1 | 1 | 0% | 1,184 | 5,352 | +352% | 0 | 0 | — |
case-19 | fail→fail | 24,885 | 18,956 | -24% | 1 | 1 | 0% | 3,485 | 5,146 | +48% | 0 | 0 | — |
case-20 | fail→fail | 16,869 | 24,617 | +46% | 1 | 1 | 0% | 1,778 | 5,767 | +224% | 0 | 0 | — |
case-21 | fail→fail | 14,682 | 22,634 | +54% | 1 | 1 | 0% | 1,489 | 5,177 | +248% | 0 | 0 | — |
case-22 | fail→fail | 19,410 | 22,378 | +15% | 1 | 1 | 0% | 2,220 | 5,061 | +128% | 0 | 0 | — |
case-23 | pass→fail | 35,105 | 19,887 | -43% | 1 | 1 | 0% | 4,666 | 5,056 | +8% | 0 | 0 | — |
case-24 | pass→fail | 25,123 | 19,845 | -21% | 1 | 1 | 0% | 3,226 | 5,020 | +56% | 0 | 0 | — |
case-25 | pass→fail | 28,368 | 20,632 | -27% | 1 | 1 | 0% | 3,573 | 4,853 | +36% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 6 counted toward the lift figure. The other 19 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 6 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.