Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys to GPU, collects initial results. Use when user says "实现实验", "implement experiments", "bridge", "从计划到跑实验", "deploy the plan", or has an experiment plan ready to execute.
.claude/skills/wanshuiyin-experiment-bridge/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 174% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 266% | 0% |
Implement and deploy experiments from plan: $ARGUMENTS
This skill bridges Workflow 1 (idea discovery + method refinement) and Workflow 2 (auto review loop). It takes the experiment plan and turns it into running experiments with initial results.
Workflow 1 output: This skill: Workflow 2 input:
refine-logs/EXPERIMENT_PLAN.md → implement → deploy → collect → initial results ready
refine-logs/EXPERIMENT_TRACKER.md code /run-experiment for /auto-review-loop
refine-logs/FINAL_PROPOSAL.mdfalse to review code before deploying.false to skip.true, prefer idea-stage/IDEA_CANDIDATES.md over the full idea-stage/IDEA_REPORT.md, and append completed runs to EXPERIMENT_LOG.md.> Override: /experiment-bridge "EXPERIMENT_PLAN.md" — compact: true, base repo: https://github.com/org/project
This skill expects one or more of:
refine-logs/EXPERIMENT_PLAN.md (best) — claim-driven experiment roadmap from /experiment-planrefine-logs/EXPERIMENT_TRACKER.md — run-by-run execution tablerefine-logs/FINAL_PROPOSAL.md — method description for implementation contextidea-stage/IDEA_CANDIDATES.md — compact idea summary (preferred when COMPACT = true) (fall back to `./IDEA_CANDIDATES.md` if not found)idea-stage/IDEA_REPORT.md — fallback if refine-logs don't exist (fall back to `./IDEA_REPORT.md` if not found)If none exist, ask the user what experiments to implement.
Read EXPERIMENT_PLAN.md and extract:
FINAL_PROPOSAL.md — what exactly to implementPresent a brief summary:
📋 Experiment plan loaded:
- Milestones: [N] (sanity → baseline → main → ablation)
- Must-run experiments: [N]
- Nice-to-have: [N]
- Estimated GPU-hours: [X]
Proceeding to implementation.Research-contract fallback: if idea-stage/docs/research_contract.md does not exist yet, create it now from templates/RESEARCH_CONTRACT_TEMPLATE.md using the selected idea + claims from the experiment plan — downstream /result-to-claim and /ablation-planner read it as the claims source.
If BASE_REPO is set — clone the repo first:
bashgit clone <BASE_REPO> base_repo/
For each milestone (in order), write the experiment scripts:
base_repo/) for existing experiment scripts, model code, and data loaders. Reuse as much as possible.Skip this step if CODE_REVIEW is false.
Before deploying, send the experiment code to a secondary Codex reviewer with xhigh reasoning:
textspawn_agent: model: gpt-6-astra reasoning_effort: xhigh message: | Review the following experiment implementation for correctness. ## Experiment Plan [paste key sections from EXPERIMENT_PLAN.md] ## Method Description [paste from FINAL_PROPOSAL.md] ## Implementation [paste the experiment scripts or exact file paths plus relevant snippets] Check for: 1. Does the code correctly implement the method described in the proposal? 2. Are all hyperparameters from the plan reflected in the code? 3. Are there logic bugs: wrong loss, wrong data split, missing eval, leakage, metric mismatch? 4. Is the evaluation metric computed against ground truth, not another model's output? 5. Are seeds, result paths, logging, and failure handling sufficient for reproducible experiments? Output: - BLOCKING issues that must be fixed before deployment - NON-BLOCKING issues that can wait - Suggested patches or checks === SCOPE LIMITS (these bound what you PROPOSE, never what you look for) === Report anything that is actually wrong here — including a rare-looking case, if this repo actually produces it. Then keep the fix in scope: 1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is welcome; over-defense is not. Assume a cooperating operator on their own machine — a malicious local user is NOT in the threat model. 2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes. Reporting a real defect in hashing code that already exists is fine. 3. NO speculative machinery: do not add feature flags, migration frameworks, compat layers, wrappers, pins, or similar mechanisms unless evidence shows a current repo defect they fix or an explicit existing invariant they must preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels, not evidence. Point to the failing path/artifact or invariant, and check the proposal's factual premises, such as whether a named package version exists. 4. NO corner-case obsession: exotic encodings, symlink races, RTL text and millisecond races are out of scope unless you can show the case arises here. 5. Where a rubric or checklist is genuinely needed, do not over-mechanize judgement. A clear sentence a human reads beats a scored table nobody maintains. Exception: code that runs remote commands, starts a network service, or installs an MCP server runs on the user's machine with their credentials — trust-boundary findings there are in scope and the default is strict. Say plainly when something is correct. Do not manufacture findings.
If BLOCKING issues are found, fix them and re-run this review once before Phase 3. Save the reviewer response and any fixes in refine-logs/EXPERIMENT_CODE_REVIEW.md. If reviewer delegation is unavailable, run the same checklist locally and mark the review [local-only].
Before deploying the full experiment suite, run the sanity-stage experiment:
/run-experiment [sanity experiment command]Wait for completion. Verify:
If sanity fails → READ the traceback/stderr/logs first, then fix the code and re-run — never re-run unchanged hoping for a different outcome. (The same read-the-primary-artifact discipline applies to surprising REVIEWER verdicts: see shared-references/review-tracing.md § Debugging With Traces.) After 1–2 failed patches on the same failure, discard and reimplement the failing script cleanly from the plan — a peer move to another patch, not a last resort; delete only the attempt's own code, never the plan / tracker / data / results (per external-cadence.md, "Let a broken attempt restart, not just patch"). Two clean reimplements failing the same way put the plan or the environment in question — report that explicitly. Do not proceed to full deployment with broken code.
If the same sanity failure repeats, trigger a second opinion: summarize the plan, code diff, command, logs, backend, and failure, then ask a fresh Codex reviewer agent for a rescue diagnosis. Apply only concrete fixes grounded in the logs.
Deploy experiments following the plan's milestone order. Route by job count and dependencies:
/run-experiment [experiment commands]For large batches (≥10 jobs), multi-seed sweeps, or teacher→student phase dependencies, use the queue scheduler:
/experiment-queue [grid spec or manifest]Auto-routing rule: if any milestone in EXPERIMENT_PLAN.md declares ≥10 jobs or declares phase dependencies, route that milestone to /experiment-queue; otherwise use /run-experiment. /experiment-queue adds OOM-aware retry with backoff, stale-screen cleanup, wave-transition race prevention, phase dependency enforcement, and crash-safe state persistence in queue_state.json.
For each milestone:
/run-experiment, or max_parallel from the queue manifest for /experiment-queue)/monitor-experiment to track progress; if /experiment-queue is active, monitor queue_state.jsonBackend lifecycle rules:
auto_destroy is configured, write the exact cleanup command before launch.🚦 Checkpoint (if AUTO_DEPLOY = false):
🔧 Code implementation complete. Ready to deploy:
Milestone 0 (sanity): [status — passed/pending]
Milestone 1 (baseline): [N experiments, ~X GPU-hours]
Milestone 2 (main method): [N experiments, ~X GPU-hours]
Milestone 3 (ablations): [N experiments, ~X GPU-hours]
Total estimated: ~X GPU-hours on [N] GPUs
Deploy now? Or review the code first?As experiments complete:
/training-check to detect NaN, loss divergence, plateaus, or overfitting. If W&B is not configured, skip silently.refine-logs/EXPERIMENT_TRACKER.md — fill in Status and Notes columnsmarkdown# Initial Experiment Results **Date**: [today] **Plan**: refine-logs/EXPERIMENT_PLAN.md ## Results by Milestone ### M0: Sanity — PASSED - [result] ### M1: Baselines | Run | System | Key Metric | Status | |-----|--------|-----------|--------| | R001 | baseline_1 | X.XX | DONE | ### M2: Main Method | Run | System | Key Metric | Status | |-----|--------|-----------|--------| | R003 | our_method | X.XX | DONE | ### M3: Ablations ... ## Summary - [X/Y] must-run experiments completed - Main result: [positive/negative/inconclusive] - Ready for /auto-review-loop: [YES/NO] ## Next Step → /auto-review-loop "[topic]"
Skip entirely if COMPACT is false.
Append each completed experiment to EXPERIMENT_LOG.md:
markdown## [Run ID] — [timestamp] - **System**: [method name] - **Config**: [key hyperparameters] - **Result**: [primary metric = X.XX] - **Verdict**: [positive / negative / inconclusive] - **Reproduce**: `python train.py --config configs/run_id.yaml --seed 42`
After main experiments (M2) complete with positive results, invoke /ablation-planner to design ablation studies:
refine-logs/EXPERIMENT_PLAN.md and refine-logs/EXPERIMENT_TRACKER.mdIf /ablation-planner is unavailable, skip silently.
Present final status:
🔬 Experiment bridge complete:
- Implemented: [N] experiment scripts
- Deployed: [N] experiments on [M] GPUs
- Completed: [X/Y] must-run, [A/B] nice-to-have
- Main result: [one sentence]
Results: refine-logs/EXPERIMENT_RESULTS.md
Tracker: refine-logs/EXPERIMENT_TRACKER.md
Ready for Workflow 2:
→ /auto-review-loop "[topic]"> Follow these shared protocols for all output files: > - Output Versioning Protocol — write timestamped file first, then copy to fixed name > - Output Manifest Protocol — log every output to MANIFEST.md > - Output Language Protocol — respect the project's language setting
cat << 'EOF' > file) to write in chunks. Do NOT ask the user for permission — just do it silently.EXPERIMENT_TRACKER.md should reflect real status after each run completes./idea-discovery "direction" ← Workflow 1: find + refine + plan
/experiment-bridge ← you are here (Workflow 1.5: implement + deploy)
/auto-review-loop "topic" ← Workflow 2: review + iterate
/paper-writing "NARRATIVE_REPORT.md" ← Workflow 3: write the paper
Or use /research-pipeline for the full end-to-end flow (includes this bridge).| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→pass | 15,993 | 8,877 | -44% | 1 | 1 | 0% | 1,571 | 4,733 | +201% | 0 | 0 | — |
case-01 | fail→fail | 14,638 | 16,890 | +15% | 1 | 1 | 0% | 262 | 4,546 | +1635% | 0 | 0 | — |
case-07 | fail→pass | 19,368 | 11,054 | -43% | 1 | 1 | 0% | 2,340 | 5,066 | +116% | 0 | 0 | — |
case-02 | fail→fail | 19,958 | 15,490 | -22% | 1 | 1 | 0% | 1,053 | 4,483 | +326% | 0 | 0 | — |
case-03 | fail→fail | 15,343 | 14,848 | -3% | 1 | 1 | 0% | 1,246 | 4,372 | +251% | 0 | 0 | — |
case-04 | fail→fail | 15,033 | 9,518 | -37% | 1 | 1 | 0% | 1,524 | 4,891 | +221% | 0 | 0 | — |
case-05 | fail→pass | 20,772 | 13,131 | -37% | 1 | 1 | 0% | 2,581 | 5,452 | +111% | 0 | 0 | — |
case-06 | fail→pass | 16,007 | 8,271 | -48% | 1 | 1 | 0% | 1,688 | 4,627 | +174% | 0 | 0 | — |
case-08 | fail→fail | 17,107 | 11,029 | -36% | 1 | 1 | 0% | 1,780 | 5,070 | +185% | 0 | 0 | — |
case-09 | fail→pass | 22,710 | 15,099 | -34% | 1 | 1 | 0% | 2,475 | 5,639 | +128% | 0 | 0 | — |
case-10 | pass→pass | 13,911 | 8,754 | -37% | 1 | 1 | 0% | 1,299 | 4,754 | +266% | 0 | 0 | — |
case-11 | fail→pass | 13,743 | 9,247 | -33% | 1 | 1 | 0% | 1,309 | 4,787 | +266% | 0 | 0 | — |
case-12 | fail→pass | 24,190 | 9,231 | -62% | 1 | 1 | 0% | 3,080 | 4,957 | +61% | 0 | 0 | — |
case-14 | pass→pass | 10,207 | 10,155 | -1% | 1 | 1 | 0% | 763 | 4,933 | +547% | 0 | 0 | — |
case-15 | fail→pass | 14,789 | 8,346 | -44% | 1 | 1 | 0% | 1,887 | 4,698 | +149% | 0 | 0 | — |
case-16 | pass→pass | 16,375 | 9,438 | -42% | 1 | 1 | 0% | 1,833 | 4,895 | +167% | 0 | 0 | — |
case-17 | fail→fail | 9,511 | 6,860 | -28% | 1 | 1 | 0% | 748 | 4,359 | +483% | 0 | 0 | — |
case-18 | fail→pass | 13,321 | 9,272 | -30% | 1 | 1 | 0% | 1,286 | 4,902 | +281% | 0 | 0 | — |
case-19 | fail→pass | 14,104 | 10,774 | -24% | 1 | 1 | 0% | 1,404 | 4,997 | +256% | 0 | 0 | — |
case-20 | fail→pass | 12,370 | 7,363 | -40% | 1 | 1 | 0% | 1,092 | 4,484 | +311% | 0 | 0 | — |
case-21 | fail→fail | 40,962 | 15,648 | -62% | 1 | 1 | 0% | 6,575 | 4,334 | -34% | 0 | 0 | — |
case-22 | fail→fail | 10,464 | 16,064 | +54% | 1 | 1 | 0% | 697 | 4,382 | +529% | 0 | 0 | — |
case-23 | fail→fail | 17,089 | 15,275 | -11% | 1 | 1 | 0% | 1,678 | 4,390 | +162% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +43 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/24/2026 | +41% |
| gemini-3.6-flash | verified | 8/17/2026 | +50% |
| gemini-3.6-flash | verified | 8/11/2026 | +41% |
Other measured skills in the registry, with their headline benchmark lift.