Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage tasks via dex CLI. Use when breaking down complex work, tracking implementation items, or persisting context across sessions.
.claude/skills/dcramer-dex/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 149% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 9% | 0% |
Use dex directly for all commands. If not on PATH, use npx @zeeg/dex instead.
bashcommand -v dex &>/dev/null && echo "use: dex" || echo "use: npx @zeeg/dex"
Dex tasks are tickets - structured artifacts with comprehensive context:
Think: "Would someone understand the what, why, and how from this task alone?"
Never reference dex task IDs in external artifacts (commits, PRs, docs). Task IDs like abc123 become meaningless once tasks are completed. Describe the work itself, not the task that tracked it.
Use dex when:
Skip dex when:
Some AI agents (like Claude Code) have built-in task tools. These are session-only and not the same as dex.
| | dex | Built-in Task Tools | | --------------- | ------------------------------------- | ------------------- | | Persistence | Files in .dex/ | Session-only | | Context | Rich (description + context + result) | Basic | | Hierarchy | 3-level (epic → task → subtask) | Flat |
Use dex for persistent work. Use built-in task tools for ephemeral in-session tracking only.
bashdex create "Short name" --description "Full implementation context"
Description should include: what needs to be done, why, implementation approach, and acceptance criteria. See examples.md for good/bad examples.
bashdex list # Pending tasks dex list --ready # Unblocked tasks dex show <id> # Full details
bashdex complete <id> --result "What was accomplished" --commit <sha>
GitHub/Shortcut-linked tasks require either --commit <sha> or --no-commit:
--commit <sha> when you have code changes (issue closes when merged)--no-commit for non-code tasks like planning or design (issue stays open)Always verify before completing. Results must include evidence: test counts, build status, manual testing outcomes. See verification.md for the full checklist.
bashdex edit <id> --description "Updated description" dex delete <id>
For full CLI reference including blockers, see cli-reference.md.
Tasks have two text fields:
dex list)--full)When you run dex show <id>, the description may be truncated. The CLI will hint at --full if there's more content.
When picking up a task, gather all relevant context:
bashdex show <id> --full # Full task details dex show <parent-id> --full # Parent context (if applicable) dex show <blocker-id> --full # What blockers accomplished
Before starting, verify you can answer:
If any answer is unclear:
Proceed without full context when:
Three levels: Epic (large initiative) → Task (significant work) → Subtask (atomic step).
Choosing the right level:
bash# Create subtask under parent dex create --parent <id> "Subtask name" --description "..."
For detailed hierarchy guidance, see hierarchies.md.
Complete tasks immediately after implementing AND verifying:
Your result must include explicit verification evidence. Don't just describe what you did—prove it works. See verification.md.
When a task is linked to a GitHub issue (shown in dex show output), include issue references in commit messages:
Fixes #NRefs #NCheck dex show <id> for GitHub issue info before committing. The "(via parent)" indicator means use Refs, direct metadata means use Fixes.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,383 | 3,636 | -73% | 1 | 1 | 0% | 2,290 | 1,662 | -27% | 0 | 0 | — |
case-02 | fail→fail | 5,880 | 5,224 | -11% | 1 | 1 | 0% | 938 | 1,748 | +86% | 0 | 0 | — |
case-03 | fail→fail | 5,006 | 5,962 | +19% | 1 | 1 | 0% | 893 | 1,842 | +106% | 0 | 0 | — |
case-04 | fail→pass | 4,391 | 2,001 | -54% | 1 | 1 | 0% | 731 | 1,820 | +149% | 0 | 0 | — |
case-05 | fail→pass | 6,196 | 2,630 | -58% | 1 | 1 | 0% | 941 | 1,948 | +107% | 0 | 0 | — |
case-06 | fail→pass | 10,173 | 2,642 | -74% | 1 | 1 | 0% | 1,720 | 1,971 | +15% | 0 | 0 | — |
case-07 | pass→pass | 7,371 | 1,953 | -74% | 1 | 1 | 0% | 1,150 | 1,763 | +53% | 0 | 0 | — |
case-08 | fail→pass | 8,704 | 2,342 | -73% | 1 | 1 | 0% | 1,333 | 1,898 | +42% | 0 | 0 | — |
case-09 | pass→pass | 6,398 | 7,902 | +24% | 1 | 1 | 0% | 1,051 | 2,190 | +108% | 0 | 0 | — |
case-10 | fail→pass | 11,438 | 4,284 | -63% | 1 | 1 | 0% | 2,001 | 2,184 | +9% | 0 | 0 | — |
case-11 | fail→pass | 6,831 | 1,505 | -78% | 1 | 1 | 0% | 1,118 | 1,697 | +52% | 0 | 0 | — |
case-12 | fail→pass | 8,627 | 1,510 | -82% | 1 | 1 | 0% | 1,397 | 1,721 | +23% | 0 | 0 | — |
case-13 | fail→pass | 7,770 | 2,193 | -72% | 1 | 1 | 0% | 1,316 | 1,824 | +39% | 0 | 0 | — |
case-14 | fail→pass | 8,915 | 2,500 | -72% | 1 | 1 | 0% | 1,468 | 1,867 | +27% | 0 | 0 | — |
case-15 | pass→pass | 10,716 | 3,761 | -65% | 1 | 1 | 0% | 1,718 | 2,079 | +21% | 0 | 0 | — |
case-16 | pass→pass | 17,475 | 12,334 | -29% | 1 | 1 | 0% | 2,831 | 3,553 | +26% | 0 | 0 | — |
case-17 | fail→fail | 8,796 | 4,562 | -48% | 1 | 1 | 0% | 1,513 | 2,288 | +51% | 0 | 0 | — |
case-18 | pass→pass | 8,288 | 4,552 | -45% | 1 | 1 | 0% | 1,245 | 2,269 | +82% | 0 | 0 | — |
case-19 | fail→fail | 12,825 | 2,318 | -82% | 1 | 1 | 0% | 2,182 | 1,830 | -16% | 0 | 0 | — |
case-20 | fail→pass | 8,146 | 1,577 | -81% | 1 | 1 | 0% | 1,477 | 1,747 | +18% | 0 | 0 | — |
case-21 | fail→pass | 9,829 | 1,745 | -82% | 1 | 1 | 0% | 1,487 | 1,735 | +17% | 0 | 0 | — |
case-22 | pass→pass | 9,184 | 1,686 | -82% | 1 | 1 | 0% | 1,434 | 1,658 | +16% | 0 | 0 | — |
case-23 | fail→pass | 10,729 | 1,490 | -86% | 1 | 1 | 0% | 1,710 | 1,659 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.