Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Delegate a coding task to the GitHub Copilot CLI (`copilot`) as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to delegate implementation work to Copilot - phrasings like "have Copilot implement X", "delegate this to copilot", "run it through Copilot CLI", or "use copilot to implement/fix/refactor" - or wants to run a queue of coding tasks through Copilot while staying the reviewer. DO NOT USE for tasks small enough to do inline, or when the
.claude/skills/amelnagdy-copilot-delegate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 76% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -31% | 0% |
You are the orchestrator. Delegate a bounded coding task to a separate implementer — the GitHub Copilot CLI — then review what it produced and land it yourself. You write the brief and own the judgment; the implementer makes changes in its own session in a clean working tree; you verify and commit.
The loop needs only a shell command and file access, so any comparable orchestrator can drive it.
copilot CLI is not installed or authenticated.(MXC-based, controlled via the /sandbox command and settings, disabled by default) — this relay does not configure them. --read-only only disables edit tools (--mode plan); shell commands still run. If project files must not change at all, dispatch against a clean or isolated worktree.
copilot (npm install -g @github/copilot; the CLI requires Node 22+, the relayitself runs on Node 18+ — the relay probes copilot version).
copilot login (interactive web/device flow), or setCOPILOT_GITHUB_TOKEN / GH_TOKEN / GITHUB_TOKEN in the environment.
copilot version succeeds.--cd at, the target git repository.Copilot picks a default model (auto). To choose another, pass --model <name>. The relay accepts letters, digits, and . _ : / - only (the value reaches a shell on Windows).
Copilot supports a reasoning effort dial: --effort <level> with values low, medium, high, xhigh, or max. The relay rejects any other value before dispatch.
Run these five steps per task. Steps 1, 4, and 5 require judgment; 2 and 3 are mechanical.
Copilot sees only the text you send. It cannot read your conversation: the brief must stand alone with the goal, current state, what to change, what to leave untouched, the project's real gates, and a report contract. Keep each brief to a single task. Write it to a file and pass it as the relay's --brief. See references/writing-the-brief.md.
Use the bundled relay. It runs copilot -p with --output-format json --no-color --stream off, captures the JSONL event stream, and writes result.json.
bashnode "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo # choose a model: add --model <name> # set reasoning effort: add --effort <level> # read-only planning pass: add --read-only (forces --mode plan) # full tool autonomy: add --allow-all-tools # hard time limit (watchdog): add --timeout 2h (the 30m default suits brief runs) # resume a session: add --session <id> or --resume-last # see all options: node .../relay.mjs --help
The child's cwd pins the workspace. The relay writes artifacts under the system temp dir by default and never commits. See references/dispatch-and-poll.md.
The relay blocks until copilot finishes. Run it with the orchestrator's background-command facility, or background it in the shell and poll for result.json. A pre-run usage error exits 2 and writes no result; a missing copilot exits 127 and writes status: "copilot_unavailable".
Completion means the process exited and result.json exists — trust process state and the working tree, not the progress display. Copilot's final assistant message is the finalMessage field of result.json.
touchedFiles.See references/review-and-land.md.
If the work is good, commit it. The relay never commits — the diff and result.json are the record; run git status and git diff first to confirm exactly what changed. If the group has a PR flow, make the commit and push a branch; let human review happen. If the diff is wrong or incomplete, re-dispatch a corrected brief in a fresh run and review again.
Without --allow-all-tools, copilot auto-denies tool calls in headless mode: the process exits 0 but the relay detects the denial events and reports status: "failed" with the CLI's own error message and a hint to pass --allow-all-tools. This is the honest default — the orchestrator sees the failure rather than a silent no-op.
--allow-all-tools explicitly grants full tool autonomy. --read-only selects --mode plan, which disables edit tools so project files can't be changed by direct edits; it works without --allow-all-tools. Shell commands still run in plan mode, so it guards against edits, not against everything. The two flags are mutually exclusive.
Copilot also exposes sandbox controls, but they are upstream-experimental (MXC-based, controlled via the /sandbox command and settings, disabled by default). This relay does not configure them.
Delegation is something the human opts into. Once briefed, copilot works as a tool you approved use of. The boundary is: do not accept conclusions from the self-report; verify everything on disk. For anything touching credentials, production data, or irreversible operations, stop and ask the human first instead of encoding it in a brief.
brief delivery.
result.json, and failure recovery.
the diff done, at the end of a run.
constraint carry-forward, progress tracking, and the final coherence pass.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,784 | 9,596 | -61% | 1 | 1 | 0% | 3,945 | 2,024 | -49% | 0 | 0 | — |
case-02 | fail→fail | 6,471 | 9,252 | +43% | 1 | 1 | 0% | 306 | 2,018 | +559% | 0 | 0 | — |
case-03 | fail→fail | 6,039 | 15,923 | +164% | 1 | 1 | 0% | 300 | 1,973 | +558% | 0 | 0 | — |
case-04 | pass→pass | 10,225 | 25,990 | +154% | 1 | 1 | 0% | 1,215 | 1,973 | +62% | 0 | 0 | — |
case-05 | fail→pass | 7,851 | 6,400 | -18% | 1 | 1 | 0% | 1,235 | 2,310 | +87% | 0 | 0 | — |
case-06 | fail→pass | 15,806 | 8,452 | -47% | 1 | 1 | 0% | 2,326 | 2,527 | +9% | 0 | 0 | — |
case-07 | fail→pass | 11,300 | 5,350 | -53% | 1 | 1 | 0% | 1,515 | 2,189 | +44% | 0 | 0 | — |
case-08 | fail→fail | 41,250 | 6,644 | -84% | 1 | 1 | 0% | 1,541 | 2,686 | +74% | 0 | 0 | — |
case-09 | pass→pass | 19,027 | 9,677 | -49% | 1 | 1 | 0% | 2,804 | 2,879 | +3% | 0 | 0 | — |
case-10 | fail→pass | 32,272 | 3,598 | -89% | 1 | 1 | 0% | 1,148 | 2,016 | +76% | 0 | 0 | — |
case-11 | fail→pass | 19,977 | 5,017 | -75% | 1 | 1 | 0% | 3,232 | 2,243 | -31% | 0 | 0 | — |
case-12 | fail→pass | 12,687 | 3,268 | -74% | 1 | 1 | 0% | 2,080 | 2,074 | -0% | 0 | 0 | — |
case-13 | fail→pass | 30,489 | 2,919 | -90% | 1 | 1 | 0% | 1,698 | 1,963 | +16% | 0 | 0 | — |
case-14 | fail→pass | 38,618 | 7,096 | -82% | 1 | 1 | 0% | 1,630 | 2,655 | +63% | 0 | 0 | — |
case-15 | fail→pass | 11,209 | 3,580 | -68% | 1 | 1 | 0% | 1,697 | 1,900 | +12% | 0 | 0 | — |
case-16 | fail→pass | 12,334 | 9,297 | -25% | 1 | 1 | 0% | 2,037 | 1,915 | -6% | 0 | 0 | — |
case-17 | fail→pass | 11,575 | 125,951 | +988% | 1 | 1 | 0% | 1,508 | 1,886 | +25% | 0 | 0 | — |
case-18 | fail→pass | 8,237 | 13,249 | +61% | 1 | 1 | 0% | 1,275 | 2,330 | +83% | 0 | 0 | — |
case-19 | pass→pass | 32,159 | 38,343 | +19% | 1 | 1 | 0% | 2,654 | 3,528 | +33% | 0 | 0 | — |
case-20 | pass→pass | 14,232 | 2,980 | -79% | 1 | 1 | 0% | 2,145 | 1,885 | -12% | 0 | 0 | — |
case-21 | pass→pass | 7,462 | 6,377 | -15% | 1 | 1 | 0% | 1,114 | 2,140 | +92% | 0 | 0 | — |
case-22 | pass→pass | 45,167 | 9,852 | -78% | 1 | 1 | 0% | 1,825 | 3,068 | +68% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.