Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan, delegate, review, and recover coding-agent implementation work with tight scope control, acceptance criteria, and validation.
.claude/skills/davepoon-coding-agent-pm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 31% | 0% |
Use this skill when a human is preparing, delegating, reviewing, or recovering implementation work done by Claude Code, Codex, OpenClaw, or another coding agent.
markdown# Agent Brief ## Outcome Ship: ## Scope Owned files/modules: Do not edit: Forbidden actions: - Do not run destructive git commands. - Do not revert unrelated user changes. ## Context Relevant files: Existing patterns to follow: ## Acceptance Criteria - [ ] - [ ] - [ ] ## Validation Run: Manual checks: ## Final Response Must Include - Files changed - Validation run and result - Remaining risks or follow-ups
Review in this order:
If the agent drifted, stop new edits and ask it to map each change back to the original acceptance criteria.
Use this when a coding agent's work is hard to trust:
textPause implementation. Map every changed file and every material change back to the original acceptance criteria. For each change, mark it as required, optional, or unrelated. Do not edit files until this map is complete.
User: "Have a coding agent add OAuth login."
Better assignment:
markdownOutcome: Add GitHub OAuth login to the existing auth flow. Scope: - Own: src/auth/*, src/routes/login.tsx, tests/auth/* - Do not edit: billing, onboarding, database migrations unless a failing test proves it is required Acceptance criteria: - [ ] Existing email login still works. - [ ] GitHub login creates or links a user through the existing auth service. - [ ] Auth errors render through the existing form error pattern. - [ ] Tests cover success, denied OAuth, and existing linked account. Validation: - npm test -- auth - npm run lint
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,167 | 8,568 | -35% | 1 | 1 | 0% | 2,063 | 2,177 | +6% | 0 | 0 | — |
case-02 | fail→pass | 17,978 | 7,642 | -57% | 1 | 1 | 0% | 1,460 | 1,730 | +18% | 0 | 0 | — |
case-03 | fail→fail | 15,476 | 7,110 | -54% | 1 | 1 | 0% | 2,300 | 1,948 | -15% | 0 | 0 | — |
case-04 | pass→pass | 10,009 | 7,529 | -25% | 1 | 1 | 0% | 1,574 | 2,060 | +31% | 0 | 0 | — |
case-05 | fail→pass | 10,906 | 9,018 | -17% | 1 | 1 | 0% | 1,644 | 2,162 | +32% | 0 | 0 | — |
case-06 | fail→pass | 11,645 | 8,435 | -28% | 1 | 1 | 0% | 1,944 | 1,949 | +0% | 0 | 0 | — |
case-07 | pass→pass | 11,798 | 7,082 | -40% | 1 | 1 | 0% | 1,814 | 1,871 | +3% | 0 | 0 | — |
case-08 | pass→pass | 11,645 | 6,920 | -41% | 1 | 1 | 0% | 1,730 | 1,846 | +7% | 0 | 0 | — |
case-09 | pass→pass | 22,924 | 7,719 | -66% | 1 | 1 | 0% | 1,595 | 1,926 | +21% | 0 | 0 | — |
case-10 | pass→pass | 10,172 | 5,297 | -48% | 1 | 1 | 0% | 1,537 | 1,546 | +1% | 0 | 0 | — |
case-11 | pass→pass | 12,745 | 12,077 | -5% | 1 | 1 | 0% | 2,168 | 2,681 | +24% | 0 | 0 | — |
case-12 | fail→pass | 10,423 | 8,220 | -21% | 1 | 1 | 0% | 1,559 | 2,038 | +31% | 0 | 0 | — |
case-13 | pass→pass | 13,431 | 7,072 | -47% | 1 | 1 | 0% | 2,107 | 1,872 | -11% | 0 | 0 | — |
case-14 | pass→pass | 10,695 | 6,820 | -36% | 1 | 1 | 0% | 1,655 | 1,878 | +13% | 0 | 0 | — |
case-15 | pass→pass | 14,309 | 10,420 | -27% | 1 | 1 | 0% | 2,105 | 2,499 | +19% | 0 | 0 | — |
case-16 | pass→pass | 15,284 | 4,910 | -68% | 1 | 1 | 0% | 2,338 | 1,483 | -37% | 0 | 0 | — |
case-17 | fail→pass | 15,599 | 8,057 | -48% | 1 | 1 | 0% | 2,624 | 2,089 | -20% | 0 | 0 | — |
case-18 | pass→pass | 10,879 | 4,006 | -63% | 1 | 1 | 0% | 1,657 | 1,434 | -13% | 0 | 0 | — |
case-19 | pass→pass | 17,417 | 9,974 | -43% | 1 | 1 | 0% | 2,699 | 2,297 | -15% | 0 | 0 | — |
case-20 | pass→pass | 15,869 | 15,932 | +0% | 1 | 1 | 0% | 2,841 | 3,514 | +24% | 0 | 0 | — |
case-21 | pass→pass | 16,035 | 14,807 | -8% | 1 | 1 | 0% | 2,420 | 3,173 | +31% | 0 | 0 | — |
case-22 | pass→pass | 16,150 | 13,612 | -16% | 1 | 1 | 0% | 2,411 | 2,891 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.