Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a disciplined, multi-source research investigation for a high-stakes question or decision — fan-out web search across many channels, parallel sub-agents, source triangulation (each claim backed by ≥3 independent sources), an adversarial review pass, and every source saved to its own file with verbatim quotes for reuse. Use when a low-quality answer is expensive: strategy work, comparing N products/methods/markets, validating a hypothesis with external data, or mapping how a field works. NOT
.claude/skills/alirezarezvani-deep-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 169% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 166% | 0% |
Turn "research this topic" into an auditable, reusable investigation instead of a one-shot wall of text. The output is a folder you can return to in a month: every claim traces to a specific source file, the plan documents why each choice was made, and a refresh protocol lets you update it later without re-running everything.
This is the heavy, methodical end of research. It is not a fast overview — it is the workflow you reach for when getting the answer wrong costs more than the tokens spent getting it right.
A router-style research skill (keyword-classify → delegate → short sequential search → markdown brief) is optimal when you need an answer fast and the decision risk is low. deep-research is the opposite trade: it pays for rigor. Use it when the answer feeds a strategy, an irreversible decision, a published artifact, or a hypothesis you need to actually test — situations where a shallow fallback would be a liability.
Concretely, deep-research adds what a fast overview does not: falsifiable hypotheses up front, parallel sub-agent fan-out across many channels, triangulation with explicit source-type diversity, a mandatory adversarial pass, per-source files with verbatim quotes, and a refresh_targets.md for delta-updates later.
Depth scales with the task — shallow runs the core phases inline; medium/deep add capability discovery, verification, and refresh targets.
| # | Phase | What it does | |---|-------|--------------| | 1 | Reframe | Rewrite the question, fix the underlying decision, state 2–4 falsifiable hypotheses | | 2 | Genre & blocks | Pick the report genre (qa / explainer / decision / landscape / validation / custom) and its building blocks | | 3 | Plan | Write plan.md: scope, structure, sourcing strategy, opposition queries, risk register, stop-criteria | | 3.5 | Capability discovery | Audit available API keys/channels in the environment; map subtopics to sources; fall back to HTML where needed | | 4 | Search (loop) | Dispatch sources → launch sub-agents in parallel → fetch & dedup → save each to sources/NN.md; re-evaluate between rounds | | 5 | Score & triangulate | Rate every source on Credibility / Recency / Bias; require ≥3 independent, differently-typed sources per thesis | | 6 | Synthesize + adversarial | Assemble the report from blocks, run 4 self-critique questions, add steel-manned counter-arguments | | 6.5 | Verify | Lightweight citation check before closing | | 7 | Refresh targets | Extract entities / numbers / hypotheses into refresh_targets.md — the entry point for future updates |
These are what separate a documented investigation from a confident guess:
sources/NN_slug.md with metadata, verbatim quotes, and scores. No dangling claim — every assertion links back to a specific file. An empty fetch produces an empty claim, never a fabricated citation.refresh_targets.md; an update <slug> run produces a delta (new entrants, entity changes, refreshed numbers, adversarial triggers) instead of replaying the whole investigation.findings/FN.md plus a sources.csv index — research compounds across questions instead of starting from zero each time.<root>/<slug>/
├── plan.md # scope, sourcing strategy, risk register, changelog
├── sources.csv # index of every source with scores
├── sources/
│ ├── 01_<slug>.md # one file = one source (metadata + verbatim quotes)
│ └── ...
├── findings/ # atomic, reusable theses (larger investigations)
│ └── F1_<short>.md
├── refresh_targets.md # what to watch on update (medium/deep)
├── diffs/
│ └── YYYY-MM-DD_delta.md # delta from an `update <slug>` run
└── YYYY-MM-DD_<genre>.md # final reportsources/ into one file. Per-source files are what make findings searchable and reusable across investigations.deep-research is the heavyweight alternative when rigor matters more than speed.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 34,916 | 34,935 | +0% | 1 | 1 | 0% | 6,244 | 7,871 | +26% | 0 | 0 | — |
case-02 | fail→fail | 34,567 | 34,068 | -1% | 1 | 1 | 0% | 6,233 | 7,860 | +26% | 0 | 0 | — |
case-03 | fail→fail | 34,305 | 33,207 | -3% | 1 | 1 | 0% | 6,228 | 7,855 | +26% | 0 | 0 | — |
case-04 | pass→pass | 6,658 | 6,662 | +0% | 1 | 1 | 0% | 1,079 | 2,777 | +157% | 0 | 0 | — |
case-05 | pass→pass | 15,635 | 28,524 | +82% | 1 | 1 | 0% | 2,999 | 6,488 | +116% | 0 | 0 | — |
case-06 | pass→fail | 30,952 | 34,341 | +11% | 1 | 1 | 0% | 5,342 | 7,788 | +46% | 0 | 0 | — |
case-07 | fail→fail | 29,822 | 33,050 | +11% | 1 | 1 | 0% | 5,447 | 7,835 | +44% | 0 | 0 | — |
case-08 | pass→pass | 5,021 | 6,181 | +23% | 1 | 1 | 0% | 690 | 2,527 | +266% | 0 | 0 | — |
case-09 | fail→pass | 19,033 | 15,190 | -20% | 1 | 1 | 0% | 3,251 | 4,492 | +38% | 0 | 0 | — |
case-10 | pass→pass | 16,178 | 15,535 | -4% | 1 | 1 | 0% | 3,057 | 4,482 | +47% | 0 | 0 | — |
case-11 | fail→pass | 12,843 | 29,703 | +131% | 1 | 1 | 0% | 2,229 | 5,990 | +169% | 0 | 0 | — |
case-12 | fail→pass | 11,654 | 10,008 | -14% | 1 | 1 | 0% | 2,065 | 3,498 | +69% | 0 | 0 | — |
case-13 | fail→pass | 11,654 | 11,345 | -3% | 1 | 1 | 0% | 2,191 | 3,876 | +77% | 0 | 0 | — |
case-14 | fail→pass | 11,021 | 20,242 | +84% | 1 | 1 | 0% | 1,920 | 5,101 | +166% | 0 | 0 | — |
case-15 | pass→pass | 21,653 | 20,570 | -5% | 1 | 1 | 0% | 4,241 | 5,547 | +31% | 0 | 0 | — |
case-16 | fail→pass | 14,519 | 9,333 | -36% | 1 | 1 | 0% | 2,479 | 3,237 | +31% | 0 | 0 | — |
case-17 | fail→pass | 7,904 | 6,598 | -17% | 1 | 1 | 0% | 1,549 | 2,828 | +83% | 0 | 0 | — |
case-18 | fail→pass | 18,035 | 4,130 | -77% | 1 | 1 | 0% | 1,070 | 2,351 | +120% | 0 | 0 | — |
case-19 | pass→pass | 10,274 | 8,365 | -19% | 1 | 1 | 0% | 2,003 | 3,172 | +58% | 0 | 0 | — |
case-20 | fail→pass | 5,725 | 4,827 | -16% | 1 | 1 | 0% | 973 | 2,665 | +174% | 0 | 0 | — |
case-21 | pass→pass | 12,953 | 13,038 | +1% | 1 | 1 | 0% | 2,357 | 4,015 | +70% | 0 | 0 | — |
case-22 | fail→pass | 10,165 | 10,785 | +6% | 1 | 1 | 0% | 1,811 | 3,591 | +98% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.