Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Document codebase as-is with thoughts directory for historical context
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 105% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 347% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 201% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 206% | 0% |
You are tasked with conducting comprehensive research across the codebase to answer user questions by spawning parallel sub-agents and synthesizing their findings.
When this command is invoked, respond with:
I'm ready to research the codebase. Please provide your research question or area of interest, and I'll analyze it thoroughly by exploring relevant components and connections.Then wait for the user's research query.
For codebase research:
IMPORTANT: All agents are documentarians, not critics. They will describe what exists without suggesting improvements or identifying issues.
For thoughts directory:
For web research (only if user explicitly asks):
For Linear tickets (if relevant):
The key is to use these agents intelligently:
hack/spec_metadata.sh script to generate all relevant metadatathoughts/shared/research/YYYY-MM-DD-ENG-XXXX-description.mdYYYY-MM-DD-ENG-XXXX-description.md where:2025-01-08-ENG-1478-parent-child-tracking.md2025-01-08-authentication-flow.mdmkdir -p thoughts/shared/researchmarkdown --- date: Current date and time with timezone in ISO format] researcher: Researcher name from thoughts status] git_commit: Current commit hash] branch: Current branch name] repository: Repository name] topic: "User's Question/Topic]" tags: research, codebase, relevant-component-names] status: complete last_updated: Current date in YYYY-MM-DD format] last_updated_by: Researcher name] ---
# Research: User's Question/Topic]
Date: Current date and time with timezone from step 4] Researcher: Researcher name from thoughts status] Git Commit: Current commit hash from step 4] Branch: Current branch name from step 4] Repository: Repository name]
## Research Question Original user query]
## Summary High-level documentation of what was found, answering the user's question by describing what exists]
## Detailed Findings
### Component/Area 1]
### Component/Area 2] ...
## Code References
path/to/file.py:123 - Description of what's thereanother/file.ts:45-67 - Description of the code block## Architecture Documentation Current patterns, conventions, and design implementations found in the codebase]
## Historical Context (from thoughts/) Relevant insights from thoughts/ directory with references]
thoughts/shared/something.md - Historical decision about Xthoughts/local/notes.md - Past exploration of YNote: Paths exclude "searchable/" even if found there
## Related Research Links to other research documents in thoughts/shared/research/]
## Open Questions Any areas that need further investigation]
git branch --show-current and git statusgh repo view --json owner,namehttps://github.com/{owner}/{repo}/blob/{commit}/{file}#L{line}last_updated and last_updated_by to reflect the updatelast_updated_note: "Added follow-up research for [brief description]" to frontmatter## Follow-up Research [timestamp]thoughts/searchable/allison/old_stuff/notes.md → thoughts/allison/old_stuff/notes.mdthoughts/searchable/shared/prs/123.md → thoughts/shared/prs/123.mdthoughts/searchable/global/shared/templates.md → thoughts/global/shared/templates.mdlast_updated, git_commit)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | fail→pass | 8,428 | 2,687 | -68% | 1 | 1 | 0% | 1,465 | 3,010 | +105% | 0 | 0 | — |
case-01 | fail→fail | 5,285 | 11,556 | +119% | 1 | 1 | 0% | 350 | 3,126 | +793% | 0 | 0 | — |
case-02 | fail→fail | 4,264 | 8,946 | +110% | 1 | 1 | 0% | 250 | 3,046 | +1118% | 0 | 0 | — |
case-03 | fail→fail | 4,674 | 12,611 | +170% | 1 | 1 | 0% | 278 | 3,186 | +1046% | 0 | 0 | — |
case-04 | pass→fail | 12,301 | 11,483 | -7% | 1 | 1 | 0% | 2,629 | 4,834 | +84% | 0 | 0 | — |
case-05 | pass→fail | 18,844 | 5,759 | -69% | 1 | 1 | 0% | 3,521 | 3,362 | -5% | 0 | 0 | — |
case-06 | pass→fail | 13,272 | 8,919 | -33% | 1 | 1 | 0% | 2,725 | 3,231 | +19% | 0 | 0 | — |
case-07 | pass→fail | 10,607 | 7,511 | -29% | 1 | 1 | 0% | 1,668 | 2,905 | +74% | 0 | 0 | — |
case-08 | fail→pass | 3,560 | 1,909 | -46% | 1 | 1 | 0% | 643 | 2,876 | +347% | 0 | 0 | — |
case-09 | fail→pass | 5,235 | 2,619 | -50% | 1 | 1 | 0% | 1,007 | 3,028 | +201% | 0 | 0 | — |
case-10 | pass→pass | 4,610 | 3,229 | -30% | 1 | 1 | 0% | 982 | 3,127 | +218% | 0 | 0 | — |
case-11 | pass→pass | 2,813 | 3,638 | +29% | 1 | 1 | 0% | 558 | 3,279 | +488% | 0 | 0 | — |
case-12 | fail→pass | 8,136 | 6,201 | -24% | 1 | 1 | 0% | 1,340 | 3,657 | +173% | 0 | 0 | — |
case-13 | fail→fail | 10,579 | 3,021 | -71% | 1 | 1 | 0% | 1,711 | 3,055 | +79% | 0 | 0 | — |
case-15 | pass→pass | 4,535 | 2,822 | -38% | 1 | 1 | 0% | 954 | 3,020 | +217% | 0 | 0 | — |
case-16 | fail→pass | 10,682 | 6,085 | -43% | 1 | 1 | 0% | 1,207 | 3,699 | +206% | 0 | 0 | — |
case-17 | fail→pass | 11,028 | 5,526 | -50% | 1 | 1 | 0% | 1,935 | 3,412 | +76% | 0 | 0 | — |
case-18 | pass→pass | 7,665 | 2,459 | -68% | 1 | 1 | 0% | 1,346 | 2,816 | +109% | 0 | 0 | — |
case-19 | pass→pass | 9,253 | 3,023 | -67% | 1 | 1 | 0% | 1,524 | 3,047 | +100% | 0 | 0 | — |
case-20 | fail→pass | 6,616 | 1,670 | -75% | 1 | 1 | 0% | 1,125 | 2,756 | +145% | 0 | 0 | — |
case-21 | fail→pass | 4,465 | 2,812 | -37% | 1 | 1 | 0% | 848 | 3,017 | +256% | 0 | 0 | — |
case-22 | fail→fail | 5,898 | 3,170 | -46% | 1 | 1 | 0% | 1,024 | 2,965 | +190% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 18 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 7/29/2026 | +38% |
Other measured skills in the registry, with their headline benchmark lift.