Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Full DeepScientist research pipeline: scout → baseline → idea → experiment → analysis → optimize → write → review → finalize. End-to-end autonomous research lifecycle.
.claude/skills/ds-full-pipeline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | — | — |
| case-14 | ✗→✓ | ▲ Improved | — | — |
| case-02 | ✗→✓ | ▲ Improved | — | — |
| case-19 | ✗→✓ | ▲ Improved | — | — |
| case-04 | ✗→✓ | ▲ Improved | — | — |
End-to-end autonomous research workflow for: $ARGUMENTS
This skill chains all DeepScientist research stages into a single pipeline:
/ds-scout → /ds-baseline → /ds-idea → /ds-experiment → /ds-analysis-campaign → /ds-optimize → /ds-write → /ds-review → /ds-finalizeFrame the research problem, survey literature, identify datasets/metrics, discover existing baselines.
/ds-scout "$ARGUMENTS"Output: Problem framing, literature map, baseline shortlist, evaluation contract.
🚦 Gate 1: Present the research landscape to the user. Wait for confirmation before proceeding.
Reproduce or import the most relevant baseline from Stage 1's shortlist.
/ds-baselineOutput: Working baseline with verified metrics, comparability contract.
Generate concrete research hypotheses based on the literature gaps and baseline analysis.
/ds-ideaOutput: Ranked candidate ideas with selection rationale.
🚦 Gate 2: Present top ideas to the user. Wait for confirmation of which idea to pursue.
Implement and run the main experiment for the selected idea.
/ds-experimentOutput: Experiment code, results, evidence artifacts.
Run follow-up experiments: ablations, robustness checks, error analysis.
/ds-analysis-campaignOutput: Ablation results, robustness data, writing-facing evidence slices.
If results are promising but not yet strong enough, run algorithm-first iterative improvement.
/ds-optimizeSkip this stage if main experiment results already meet the success criteria.
Draft the paper from accepted evidence.
/ds-writeOutput: LaTeX paper draft with figures and references.
Run an independent skeptical audit of the draft.
/ds-reviewOutput: Review report with severity-graded feedback.
If review identifies critical issues → fix and re-review (max 2 rounds).
Consolidate final claims, limitations, and recommendations.
/ds-finalizeOutput: Final paper, summary state, resume packet.
| Stage | Duration | Autonomous? | |-------|----------|-------------| | 1. Scout | 20-40 min | Wait for Gate 1 | | 2. Baseline | 15-60 min | Yes | | 3. Idea | 15-30 min | Wait for Gate 2 | | 4. Experiment | 30 min - hours | Yes | | 5. Analysis | 30-60 min | Yes | | 6. Optimize | 0-60 min | Yes (optional) | | 7. Write | 30-60 min | Yes | | 8. Review | 15-30 min | Yes | | 9. Finalize | 10-20 min | Yes |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-17 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-15 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-16 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-14 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-02 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-13 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-19 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-04 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-20 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-03 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-09 | pass→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-18 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-21 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-11 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-06 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-12 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-07 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-08 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-01 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-05 | fail→pass | — | — | — | — | — | — | — | — | — | — | — | — |
case-10 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
case-22 | fail→fail | — | — | — | — | — | — | — | — | — | — | — | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
The per-case answers from this run were removed by the retention sweep, so the case table below shows the verdicts without the text either arm produced. The counts above were recorded at the time and are unaffected. Answers are now kept for 180 days.
Other measured skills in the registry, with their headline benchmark lift.