Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute implementation tasks from design documents using markdown checkboxes. Use when (1) implementing features from feature-design-assistant output, (2) resuming interrupted work, (3) batch executing tasks. Triggers on 'start implementation', 'run tasks', 'resume'.
.claude/skills/davila7-task-execution-engine/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -6% | 0% |
Execute implementation tasks directly from design documents. Tasks are managed as markdown checkboxes - no separate session files needed.
bash# Get next task python3 scripts/task_manager.py next --file <design.md> # Mark task completed python3 scripts/task_manager.py done --file <design.md> --task "Task Title" # Mark task failed python3 scripts/task_manager.py fail --file <design.md> --task "Task Title" --reason "..." # Show status python3 scripts/task_manager.py status --file <design.md>
Tasks are written as markdown checkboxes in the design document:
markdown## Implementation Tasks - [ ] **Create User model** `priority:1` `phase:model` - files: src/models/user.py, tests/models/test_user.py - [ ] User model has email and password_hash fields - [ ] Email validation implemented - [ ] Password hashing uses bcrypt - [ ] **Implement JWT utils** `priority:2` `phase:model` - files: src/utils/jwt.py - [ ] generate_token() creates valid JWT - [ ] verify_token() validates JWT - [ ] **Create auth API** `priority:3` `phase:api` `deps:Create User model,Implement JWT utils` - files: src/api/auth.py - [ ] POST /register endpoint - [ ] POST /login endpoint
See references/task-format.md for full format specification.
LOOP until no tasks remain:
1. GET next task (task_manager.py next)
2. READ task details (files, criteria)
3. IMPLEMENT the task
4. VERIFY acceptance criteria
5. UPDATE status (task_manager.py done/fail)
6. CONTINUECompleted task:
markdown- [x] **Create User model** `priority:1` `phase:model` ✅ - files: src/models/user.py - [x] User model has email field - [x] Password hashing implemented
Failed task:
markdown- [x] **Create User model** `priority:1` `phase:model` ❌ - files: src/models/user.py - [ ] User model has email field - reason: Missing database configuration
To resume interrupted work, simply run again with the same design file:
/feature-pipeline docs/designs/xxx.mdThe task manager will find the first uncompleted task and continue from there.
This skill is typically triggered after /feature-analyzer completes:
User: /feature-analyzer implement user auth
Claude: [designs feature, generates task list]
Design saved to docs/designs/2026-01-02-user-auth.md
Ready to start implementation?
User: Yes / 开始实现
Claude: [executes tasks via task-execution-engine]| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,979 | 4,239 | +7% | 1 | 1 | 0% | 191 | 957 | +401% | 0 | 0 | — |
case-02 | fail→fail | 4,568 | 4,574 | +0% | 1 | 1 | 0% | 232 | 1,088 | +369% | 0 | 0 | — |
case-03 | fail→fail | 4,462 | 5,726 | +28% | 1 | 1 | 0% | 245 | 1,033 | +322% | 0 | 0 | — |
case-04 | fail→pass | 6,460 | 1,472 | -77% | 1 | 1 | 0% | 1,191 | 1,061 | -11% | 0 | 0 | — |
case-05 | fail→pass | 8,492 | 1,489 | -82% | 1 | 1 | 0% | 1,525 | 1,058 | -31% | 0 | 0 | — |
case-06 | fail→pass | 7,329 | 2,672 | -64% | 1 | 1 | 0% | 1,355 | 1,288 | -5% | 0 | 0 | — |
case-07 | fail→pass | 7,685 | 3,685 | -52% | 1 | 1 | 0% | 1,263 | 1,426 | +13% | 0 | 0 | — |
case-08 | fail→pass | 6,464 | 2,645 | -59% | 1 | 1 | 0% | 1,252 | 1,173 | -6% | 0 | 0 | — |
case-09 | fail→pass | 7,625 | 2,984 | -61% | 1 | 1 | 0% | 1,256 | 1,132 | -10% | 0 | 0 | — |
case-10 | fail→pass | 7,108 | 2,963 | -58% | 1 | 1 | 0% | 1,159 | 1,348 | +16% | 0 | 0 | — |
case-11 | fail→pass | 4,933 | 3,593 | -27% | 1 | 1 | 0% | 945 | 1,388 | +47% | 0 | 0 | — |
case-12 | fail→pass | 8,975 | 5,183 | -42% | 1 | 1 | 0% | 1,612 | 1,870 | +16% | 0 | 0 | — |
case-13 | fail→pass | 7,714 | 3,170 | -59% | 1 | 1 | 0% | 1,318 | 1,339 | +2% | 0 | 0 | — |
case-14 | pass→pass | 11,446 | 3,973 | -65% | 1 | 1 | 0% | 2,054 | 1,404 | -32% | 0 | 0 | — |
case-15 | pass→pass | 8,208 | 1,510 | -82% | 1 | 1 | 0% | 1,423 | 1,023 | -28% | 0 | 0 | — |
case-16 | fail→pass | 5,958 | 1,553 | -74% | 1 | 1 | 0% | 1,165 | 948 | -19% | 0 | 0 | — |
case-17 | pass→pass | 7,704 | 1,713 | -78% | 1 | 1 | 0% | 1,482 | 1,006 | -32% | 0 | 0 | — |
case-18 | fail→pass | 6,952 | 1,982 | -71% | 1 | 1 | 0% | 1,221 | 1,116 | -9% | 0 | 0 | — |
case-19 | fail→pass | 6,378 | 2,200 | -66% | 1 | 1 | 0% | 1,095 | 1,146 | +5% | 0 | 0 | — |
case-20 | pass→pass | 2,656 | 1,482 | -44% | 1 | 1 | 0% | 388 | 1,069 | +176% | 0 | 0 | — |
case-21 | pass→pass | 3,687 | 2,992 | -19% | 1 | 1 | 0% | 806 | 1,270 | +58% | 0 | 0 | — |
case-22 | pass→pass | 3,288 | 2,335 | -29% | 1 | 1 | 0% | 644 | 1,140 | +77% | 0 | 0 | — |
case-23 | pass→fail | 8,703 | 2,849 | -67% | 1 | 1 | 0% | 1,591 | 1,206 | -24% | 0 | 0 | — |
case-24 | fail→pass | 8,411 | 3,930 | -53% | 1 | 1 | 0% | 1,384 | 1,448 | +5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 21 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +54 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.