Install any skill in seconds. Free to start, no credit card required.
Get Started Free →To delegate, prompt and manage subagents. MUST activate to spawn subagent with a quality prompt. MANDATORY unless trivial one-liner.
.claude/skills/griddynamics-orchestration/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 13% | 0% |
| case-16 | ✓→✗ | ▼ Worse | -27% | 0% |
<orchestration>
<context>
Prerequisites: USE SKILL hitl, load-project-context
</context>
<request_sizing>
assets/o-team-manager.md.If requested directly: use session-level EXECUTION_CONTROLLER (plan ⊃ phases ⊃ steps ⊃ tasks) and APPLY SKILL FILE assets/o-session-execution-controller.md.
Complexity may shift one band; re-size as reality changes — discovery, surprises, clarification, target already done.
</request_sizing>
<process>
Dispatch:
<subagent_prompt_template> — terse, factual, specific, DRY.Routing:
Quality:
produce → check cycles, orchestrator-gated — loop or accept. Check = fresh eyes: separate subagent, different model if possible; never self-review. Compose per piece: {implement · design · tests → run} → review (spec first, code quality second) · complete → validate · produce → refute (adversarial) · author → user annotates. Validate incrementally + at flow end — close flow with a validation task.Plan mode:
MUST USE SKILL <name> entries (workflow, skills), incorporate plan + specs, define the implementation workflow — mini-loops, phases, steps, subagent + model per step — in MoSCoW, same directive language you were given.</process>
<subagent_prompt_template output-reformat="expand into proper md sections">
Syntax: <x> fill · {a|b} pick one · [..] optional · * always include.
You are <role/specialization>. {Lightweight|Full} subagent.
Tasks*: <list>
Scope*: root <path> [git worktree] · DO <in-scope + expected outputs> · DO NOT <out-of-scope · read-only · untouchable — no improvising beyond scope>
[Constraints: <naming · patterns · case sensitivity>]
Checklist*: <ACs · NFRs · FRs · open-ended · Severity-based · Unlimited by item count · Domain Specific · Tasks Specific>
Skills*: MUST USE SKILL `subagent-directives`[, `load-project-context`, <required>] · [RECOMMEND USE SKILL <skill>]
Original request*: <verbatim + agreed clarifications — carry through every step>
Context*: <all it needs — refs · files · decisions; subagent starts with ONLY `bootstrap-alwayson.md` + this prompt>
Output specs*: message <content + format — unambiguous, so orchestrator can verify> · [files: <high volume → unique path per subagent + format>] · MUST return: results, summary, side effects, anomalies, discoveries, contract changes, deviations, inconsistencies, insights
Evidence specs*: <proofs you demand back — per claim: deep links + line ranges + brief quotes; facts != assumptions>
[Process Requirements: <require processing one-by-one, small-set-by-small-set, group-by-group, ordering, do not reading ALL source files at once, etc.>]
[<free-form: anything not covered>]Orchestrator decides: include load-project-context only when needed — omit for self-contained tasks or exact-file prompts.
</subagent_prompt_template>
</orchestration>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 28,200 | 3,875 | -86% | 1 | 1 | 0% | 6,214 | 1,796 | -71% | 0 | 0 | — |
case-11 | fail→fail | 15,804 | 9,197 | -42% | 1 | 1 | 0% | 2,292 | 2,875 | +25% | 0 | 0 | — |
case-21 | pass→pass | 8,076 | 7,729 | -4% | 1 | 1 | 0% | 1,452 | 2,887 | +99% | 0 | 0 | — |
case-22 | pass→pass | 6,760 | 8,997 | +33% | 1 | 1 | 0% | 1,145 | 2,927 | +156% | 0 | 0 | — |
case-02 | fail→fail | 30,468 | 4,295 | -86% | 1 | 1 | 0% | 6,223 | 1,794 | -71% | 0 | 0 | — |
case-03 | fail→fail | 5,693 | 4,043 | -29% | 1 | 1 | 0% | 244 | 1,753 | +618% | 0 | 0 | — |
case-04 | fail→fail | 11,391 | 5,375 | -53% | 1 | 1 | 0% | 1,834 | 1,693 | -8% | 0 | 0 | — |
case-05 | pass→pass | 13,086 | 8,181 | -37% | 1 | 1 | 0% | 2,031 | 2,954 | +45% | 0 | 0 | — |
case-06 | fail→fail | 10,729 | 3,948 | -63% | 1 | 1 | 0% | 1,776 | 1,672 | -6% | 0 | 0 | — |
case-07 | pass→pass | 14,343 | 12,554 | -12% | 1 | 1 | 0% | 2,112 | 3,337 | +58% | 0 | 0 | — |
case-08 | pass→pass | 13,398 | 3,951 | -71% | 1 | 1 | 0% | 2,165 | 2,021 | -7% | 0 | 0 | — |
case-09 | pass→fail | 15,582 | 9,668 | -38% | 1 | 1 | 0% | 2,510 | 2,837 | +13% | 0 | 0 | — |
case-10 | fail→pass | 8,106 | 2,562 | -68% | 1 | 1 | 0% | 1,253 | 1,795 | +43% | 0 | 0 | — |
case-12 | pass→pass | 5,897 | 2,897 | -51% | 1 | 1 | 0% | 877 | 1,909 | +118% | 0 | 0 | — |
case-13 | pass→pass | 8,029 | 6,236 | -22% | 1 | 1 | 0% | 1,170 | 2,345 | +100% | 0 | 0 | — |
case-14 | fail→pass | 8,640 | 1,827 | -79% | 1 | 1 | 0% | 1,302 | 1,687 | +30% | 0 | 0 | — |
case-15 | fail→fail | 18,271 | 4,606 | -75% | 1 | 1 | 0% | 2,867 | 1,675 | -42% | 0 | 0 | — |
case-16 | pass→fail | 14,627 | 9,402 | -36% | 1 | 1 | 0% | 2,228 | 1,623 | -27% | 0 | 0 | — |
case-17 | pass→pass | 5,852 | 3,310 | -43% | 1 | 1 | 0% | 812 | 1,865 | +130% | 0 | 0 | — |
case-18 | fail→pass | 12,883 | 6,953 | -46% | 1 | 1 | 0% | 1,930 | 2,409 | +25% | 0 | 0 | — |
case-19 | pass→pass | 13,843 | 16,111 | +16% | 1 | 1 | 0% | 2,196 | 3,800 | +73% | 0 | 0 | — |
case-20 | pass→fail | 15,679 | 5,911 | -62% | 1 | 1 | 0% | 2,814 | 1,626 | -42% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 14 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 14 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.