Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Executes an approved plan with one primary implementation stream by default, using bounded parallel sidecars only when the write scopes are truly disjoint. Supports default Claude execution or an explicit Codex executor option. Automatically reviews the result for completeness and intent fidelity. Use after a plan is approved.
.claude/skills/dcouple-implement/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -16% | 0% |
Execute the approved plan directly in this Codex session.
Codex is the primary implementation authority in this workflow. If you also have a separate Claude workflow available, treat it as an optional parallel second-opinion lane rather than the primary executor.
./tmp/ready-plans/Source Artifacts, read the brief / intent artifact andresearch dossier before coding
Treat sources of truth as:
If no separate brief exists, treat the plan's Intent / Why, Locked Decisions, Known Mismatches / Assumptions, and success criteria as the minimum intent source of truth.
Before implementing, scan the plan for commands that must not be run automatically:
Collect them into a Manual Steps list and surface them before proceeding.
Schema / migration handling is done later after review. Do not handle it here.
Default to one primary implementation stream.
Only split work when all of the following are true:
Keep these with the primary stream unless there is an unusually clean reason not to:
Plan Delta note to the planactually wired and still preserves the intended outcome
Run these quality checks during the work when feasible:
bashnpm run typecheck npm run lint
After implementation, always run a review pass against the standards in implementation-reviewer.
Minimum gate:
Preferred gate:
If you are operating alongside a separate Claude workflow, you may use that second lane in parallel. If not, perform an additional adversarial Codex review focused on:
Do not surface questions until all active review lanes are complete and their findings are merged.
Split findings into:
Apply straightforward fixes directly, then rerun the review gate when needed.
After review gates are complete and auto-fixable issues are resolved, check if schema.ts was modified:
bashgit diff origin/main --name-only | grep schema.ts
If schema changed:
npm run db:diff:devIf schema did not change, skip this step silently.
Once all tasks pass review, brief intent is preserved, and the implementation is complete, move the plan from ./tmp/ready-plans/ to ./tmp/done-plans/.
Only move the plan when all tasks are confirmed complete.
Present the final result with:
If the review found issues, offer to fix them before the user commits.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,631 | 3,795 | -18% | 1 | 1 | 0% | 257 | 1,196 | +365% | 0 | 0 | — |
case-02 | fail→fail | 5,554 | 4,781 | -14% | 1 | 1 | 0% | 338 | 1,230 | +264% | 0 | 0 | — |
case-03 | fail→fail | 4,303 | 3,985 | -7% | 1 | 1 | 0% | 211 | 1,210 | +473% | 0 | 0 | — |
case-04 | fail→pass | 6,859 | 2,867 | -58% | 1 | 1 | 0% | 1,154 | 1,294 | +12% | 0 | 0 | — |
case-05 | fail→pass | 9,562 | 2,378 | -75% | 1 | 1 | 0% | 1,444 | 1,440 | -0% | 0 | 0 | — |
case-06 | pass→pass | 9,662 | 3,907 | -60% | 1 | 1 | 0% | 1,744 | 1,772 | +2% | 0 | 0 | — |
case-07 | fail→pass | 14,957 | 4,929 | -67% | 1 | 1 | 0% | 2,448 | 1,906 | -22% | 0 | 0 | — |
case-08 | pass→fail | 7,060 | 5,175 | -27% | 1 | 1 | 0% | 1,218 | 1,312 | +8% | 0 | 0 | — |
case-09 | fail→pass | 8,880 | 4,677 | -47% | 1 | 1 | 0% | 1,375 | 1,828 | +33% | 0 | 0 | — |
case-10 | pass→pass | 8,377 | 2,397 | -71% | 1 | 1 | 0% | 1,481 | 1,322 | -11% | 0 | 0 | — |
case-11 | fail→pass | 10,549 | 2,616 | -75% | 1 | 1 | 0% | 1,676 | 1,410 | -16% | 0 | 0 | — |
case-12 | fail→pass | 8,463 | 2,223 | -74% | 1 | 1 | 0% | 1,459 | 1,348 | -8% | 0 | 0 | — |
case-13 | pass→pass | 3,437 | 1,566 | -54% | 1 | 1 | 0% | 600 | 1,242 | +107% | 0 | 0 | — |
case-14 | fail→pass | 5,884 | 1,978 | -66% | 1 | 1 | 0% | 1,000 | 1,314 | +31% | 0 | 0 | — |
case-15 | fail→fail | 11,785 | 5,106 | -57% | 1 | 1 | 0% | 1,985 | 1,262 | -36% | 0 | 0 | — |
case-16 | fail→fail | 4,334 | 4,788 | +10% | 1 | 1 | 0% | 684 | 1,380 | +102% | 0 | 0 | — |
case-17 | pass→fail | 7,122 | 2,870 | -60% | 1 | 1 | 0% | 1,129 | 1,284 | +14% | 0 | 0 | — |
case-18 | fail→fail | 15,340 | 9,514 | -38% | 1 | 1 | 0% | 2,375 | 1,945 | -18% | 0 | 0 | — |
case-19 | fail→pass | 18,185 | 14,700 | -19% | 1 | 1 | 0% | 3,141 | 3,517 | +12% | 0 | 0 | — |
case-20 | fail→fail | 7,771 | 5,929 | -24% | 1 | 1 | 0% | 1,243 | 1,349 | +9% | 0 | 0 | — |
case-21 | pass→fail | 5,325 | 3,008 | -44% | 1 | 1 | 0% | 922 | 1,482 | +61% | 0 | 0 | — |
case-22 | pass→fail | 2,300 | 2,271 | -1% | 1 | 1 | 0% | 382 | 1,365 | +257% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 16 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 16 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.