Install any skill in seconds. Free to start, no credit card required.
Get Started Free →COMPREHENSIVE top-tier research for topics where breadth, source quality, and auditable methodology matter. Use when the user needs 40+ sources, publication-grade synthesis, due diligence, policy or market analysis, or any high-stakes research where missing key dimensions would materially weaken the answer. This skill is especially valuable when the prompt has multiple dimensions that must all be covered faithfully. NOT for quick or medium-depth briefs; prefer gpt-research or native web_search/w
.claude/skills/valtterimelkko-deep-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 175% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 152% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 345% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 331% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 232% | 0% |
A comprehensive workflow for conducting broad, source-rich, auditable research using iterative waves of subagents.
Time: ~6–10 minutes Sources: 40–100+ Best for: academic rigor, due diligence, comparative analysis, high-stakes synthesis
The strength of this skill is breadth with control: it covers many dimensions without letting any single subagent become the whole research system.
The main agent is the research conductor.
Subagents are used for:
The main agent is responsible for:
Use this skill as the heaviest general-purpose research tier:
web_search / web_fetch — one-off current facts, a single page, or a few URLsgpt-research — quick / medium-depth synthesis across roughly 10-20 sourcesdeep-research (this skill) — 40+ sources, explicit methodology, high-stakes synthesislast30days — when the main need is recent discourse across social/web platforms rather than broad general-web synthesisIf the user simply says “research this”, reserve this skill for tasks where the consequence of missing a dimension is materially important.
Validate that the normal research path is viable, then define the scope.
Track at least:
Do not turn the manifest into bureaucratic overhead. It is a lightweight control table.
Cast a wide net cheaply and safely.
Scouts are discovery agents, not full analysts. They should identify promising sources and report them compactly.
Launch 3–6 parallel scout agents, each covering a distinct dimension or source class.
Examples:
Each scout should:
Use wording along these lines:
textUse your default acquisition path only. Do NOT invoke browser automation, playwright, or other non-default fallback workflows. Browser automation is NOT a bot-protection bypass strategy in this workflow. If a source is blocked by Cloudflare, anti-bot protection, rate limits, login walls, paywalls, or fetch failures, record the exact URL, barrier type, source class, and failure reason, then continue. Your primary deliverable is this response text. File writes are optional secondary artifacts. At the end, explicitly report: - created_files: [...] - modified_files: [...] - blocked_urls: [{url, barrier_type, reason, source_class}] - unblocked_sources: [{title, url, publication_date_or_recency, relevance_note, credibility_estimate}] - follow_up_queries: [...] If no files were created, say so explicitly.
markdown## Scout Summary Dimension: [name] Status: adequate | thin | blocked-heavy ## Candidate Sources 1. [Title] — URL - Relevance: ... - Credibility: High/Medium/Low - Date/Recency: ... ## Follow-up Queries - ... ## Blocked URLs - URL — barrier type — reason ## Files - created_files: [...] - modified_files: [...]
Do the deep reading and extraction on curated source batches.
Analysts receive curated URLs. They should not wander broadly. However, if a curated batch is missing a necessary primary source, a key dissenting source, or another narrowly defined source needed to evaluate a material claim, the analyst may add up to 1–2 tightly scoped supplementary sources and must label them clearly as supplementary recovery sources.
Give each analyst ~4–7 URLs or a similarly narrow, clearly bounded source batch.
Each analyst should:
Use wording along these lines:
textUse your default acquisition path only. Do NOT invoke browser automation, playwright, or other non-default fallback workflows. Browser automation is NOT a bot-protection bypass strategy in this workflow. If any curated URL is blocked or fails to load, report it exactly and continue with the remaining sources. Do not broaden scope unless explicitly instructed, except for at most 1–2 tightly scoped supplementary recovery sources when needed for a material claim. Your primary deliverable is this response text. File writes are optional secondary artifacts. At the end, explicitly report: - created_files: [...] - modified_files: [...] - blocked_urls: [{url, barrier_type, reason, source_class, estimated_importance}] - supplementary_sources: [{title, url, why_needed}] - follow_up_queries: [...]
markdown## Analyst Findings Batch: [name] Coverage: strong | moderate | weak ## Key Findings - ... ## Important Claims and Support - Claim: ... - Sources: ... - Cross-check status: ... ## Source Verification Notes - Source: ... - Author/Organisation: ... - Date/Recency: ... - Type: primary / secondary / commentary - Methodology present?: yes/no/not applicable - Bias/conflict notes: ... ## Caveats / Contradictions - ... ## Supplementary Recovery Sources - ... ## Follow-up Queries - ... ## Blocked URLs - URL — barrier type — reason — estimated importance ## Files - created_files: [...] - modified_files: [...]
Turn analysed batches into a coherent answer.
Decide whether the research is strong enough to stand on its own.
Before finalising, explicitly check:
Do not keep digging forever. Move toward report assembly unless a critical gap remains.
This happens after the standard workflow is complete, not mid-workflow.
Assess whether blocked/failed sources materially weaken the result.
Track, at minimum:
Class blocked sources as:
Recommend stronger fallback options primarily when blocked sources include:
that are not adequately replaced by unblocked alternatives
Also consider stronger fallback recommendations when blocked sources are:
Report to the user:
Do not dump a giant raw URL list into the final user-facing summary unless asked.
Fallback workflows are considered only at the end of the normal deep-research workflow.
Recommend camofox first when blocked sources remain important and the problem looks like a browser-style anti-bot / JS-rendering barrier, especially when:
Recommend NotebookLM when blocked sources remain important and the recovery problem is broad rather than a short list of individual pages, especially when:
NotebookLM is usually the next fallback when the blocked problem is broad or when multiple important blocked URLs can be recovered through a user-participatory notebook workflow.
Use NotebookLM to perform deep research on:
This is appropriate when blocked coverage affects a broad slice of the research.
If only a modest number of blocked URLs matter, recommend adding those URLs directly as notebook sources, then querying the notebook.
notebooklm skillRecommend residential proxy only when:
residential-proxy skillPrefer balanced forward progress:
The final report should include all of the following:
Major claims in the final report should be traceable to cited sources. Do not rely on uncited synthesis for material conclusions.
Include a concise methodology / limitations table in the final report, for example:
| Item | Status | |------|--------| | Dimensions covered | X / Y | | Sources used | N | | Blocked URLs | N | | Critical blocked sources | N | | Research confidence | High / Medium / Low | | Fallback recommended | None / Camofox / NotebookLM / Residential proxy |
This skill wins by keeping the overall workflow broad while making each subagent more focused and auditable.
Use broad scouting first, promote the best sources into deeper analysis, keep a clear blocked-source registry, and only consider fallback workflows after the normal research pass is complete.
Default stance: keep moving, surface blocked-source significance clearly, try targeted camofox recovery before paid proxy escalation, and involve the user before using higher-friction fallback paths.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | pass→pass | 8,736 | 7,221 | -17% | 1 | 1 | 0% | 1,255 | 5,227 | +316% | 0 | 0 | — |
case-01 | fail→fail | 79,499 | 27,149 | -66% | 1 | 1 | 0% | 6,257 | 5,127 | -18% | 0 | 0 | — |
case-02 | fail→fail | 41,366 | 14,136 | -66% | 1 | 1 | 0% | 6,247 | 5,104 | -18% | 0 | 0 | — |
case-03 | fail→fail | 36,109 | 14,909 | -59% | 1 | 1 | 0% | 6,245 | 4,980 | -20% | 0 | 0 | — |
case-04 | fail→pass | 13,028 | 12,566 | -4% | 1 | 1 | 0% | 1,976 | 5,442 | +175% | 0 | 0 | — |
case-05 | fail→pass | 13,950 | 9,932 | -29% | 1 | 1 | 0% | 2,263 | 5,704 | +152% | 0 | 0 | — |
case-06 | fail→pass | 6,615 | 2,792 | -58% | 1 | 1 | 0% | 1,038 | 4,620 | +345% | 0 | 0 | — |
case-08 | fail→pass | 11,688 | 5,327 | -54% | 1 | 1 | 0% | 1,171 | 5,048 | +331% | 0 | 0 | — |
case-09 | fail→pass | 9,634 | 5,294 | -45% | 1 | 1 | 0% | 1,469 | 4,884 | +232% | 0 | 0 | — |
case-10 | pass→pass | 11,056 | 7,835 | -29% | 1 | 1 | 0% | 1,767 | 5,171 | +193% | 0 | 0 | — |
case-11 | pass→pass | 12,574 | 12,642 | +1% | 1 | 1 | 0% | 1,961 | 6,227 | +218% | 0 | 0 | — |
case-12 | fail→pass | 15,764 | 8,292 | -47% | 1 | 1 | 0% | 2,888 | 5,593 | +94% | 0 | 0 | — |
case-13 | fail→pass | 9,611 | 2,121 | -78% | 1 | 1 | 0% | 1,547 | 4,472 | +189% | 0 | 0 | — |
case-14 | pass→pass | 13,384 | 3,741 | -72% | 1 | 1 | 0% | 1,660 | 4,662 | +181% | 0 | 0 | — |
case-15 | pass→pass | 18,249 | 16,690 | -9% | 1 | 1 | 0% | 2,979 | 6,667 | +124% | 0 | 0 | — |
case-16 | pass→pass | 17,910 | 6,609 | -63% | 1 | 1 | 0% | 1,578 | 5,197 | +229% | 0 | 0 | — |
case-17 | pass→fail | 8,415 | 4,123 | -51% | 1 | 1 | 0% | 1,321 | 4,527 | +243% | 0 | 0 | — |
case-18 | pass→pass | 9,247 | 10,891 | +18% | 1 | 1 | 0% | 1,557 | 5,196 | +234% | 0 | 0 | — |
case-19 | fail→pass | 23,782 | 8,629 | -64% | 1 | 1 | 0% | 2,093 | 4,935 | +136% | 0 | 0 | — |
case-20 | pass→fail | 5,541 | 8,819 | +59% | 1 | 1 | 0% | 939 | 4,538 | +383% | 0 | 0 | — |
case-21 | pass→fail | 9,697 | 10,584 | +9% | 1 | 1 | 0% | 1,497 | 4,758 | +218% | 0 | 0 | — |
case-22 | fail→fail | 4,175 | 10,901 | +161% | 1 | 1 | 0% | 670 | 5,026 | +650% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.