Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.
.claude/skills/wanshuiyin-monitor-experiment/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 201% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-15 | ✓→✗ | ▼ Worse | 7% | 0% |
Monitor: $ARGUMENTS
First identify the backend from AGENTS.md, run notes, or launch summary: local, SSH, Vast.ai, or Modal. Monitor the backend that was actually used; do not assume a plain SSH screen session when the run was launched through Vast.ai or Modal.
bashssh <server> "screen -ls"
For Vast.ai, also check instance state, SSH reachability, hourly cost, and whether auto_destroy is pending. For Modal, check the Modal run/app logs, function status, timeout, volume outputs, and cloud cost exposure.
For each screen session, capture the last N lines:
bashssh <server> "screen -S <name> -X hardcopy /tmp/screen_<name>.txt && tail -50 /tmp/screen_<name>.txt"
If hardcopy fails, check for log files or tee output.
bashssh <server> "ls -lt <results_dir>/*.json 2>/dev/null | head -20"
If JSON results exist, fetch and parse them:
bashssh <server> "cat <results_dir>/<latest>.json"
wandb: true in AGENTS.md)If the project enables W&B, pull metrics before interpreting results. Prefer W&B as the source of training curves and recent eval state, while still checking logs for crashes.
List recent runs:
bashpython3 - <<'PY' import wandb api = wandb.Api() for run in api.runs("<entity>/<project>", per_page=20): print(run.name, run.state, run.url) PY
Pull recent history for a specific run:
bashpython3 - <<'PY' import wandb api = wandb.Api() run = api.run("<entity>/<project>/<run_id>") for row in run.history(samples=50, keys=["train/loss", "eval/loss", "eval/accuracy", "train/lr"]): print(row) print("summary:", dict(run.summary)) PY
If W&B is configured but unavailable, report the connectivity problem and fall back to screen/log/json evidence. Do not interpret missing W&B data as experiment failure by itself.
Always include W&B dashboard links (run.url) when available so later review and paper-writing agents can inspect the exact training curves.
Present results in a comparison table:
| Experiment | Metric | Delta vs Baseline | Status |
|-----------|--------|-------------------|--------|
| Baseline | X.XX | — | done |
| Method A | X.XX | +Y.Y | done |After results are collected, check ~/.codex/feishu.json:
experiment_done notification: results summary table, delta vs baseline"off": skip entirely (no-op)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,620 | 4,927 | -72% | 1 | 1 | 0% | 3,182 | 1,073 | -66% | 0 | 0 | — |
case-02 | fail→fail | 10,363 | 5,851 | -44% | 1 | 1 | 0% | 1,892 | 1,094 | -42% | 0 | 0 | — |
case-03 | fail→fail | 3,992 | 4,846 | +21% | 1 | 1 | 0% | 443 | 1,023 | +131% | 0 | 0 | — |
case-04 | fail→fail | 6,248 | 4,922 | -21% | 1 | 1 | 0% | 1,051 | 1,116 | +6% | 0 | 0 | — |
case-05 | fail→fail | 4,882 | 4,874 | -0% | 1 | 1 | 0% | 841 | 1,115 | +33% | 0 | 0 | — |
case-06 | pass→pass | 8,154 | 3,037 | -63% | 1 | 1 | 0% | 1,506 | 1,401 | -7% | 0 | 0 | — |
case-07 | fail→fail | 8,259 | 6,508 | -21% | 1 | 1 | 0% | 1,200 | 1,118 | -7% | 0 | 0 | — |
case-08 | pass→pass | 9,919 | 2,990 | -70% | 1 | 1 | 0% | 1,579 | 1,266 | -20% | 0 | 0 | — |
case-09 | fail→pass | 6,053 | 2,878 | -52% | 1 | 1 | 0% | 927 | 1,280 | +38% | 0 | 0 | — |
case-10 | pass→pass | 5,964 | 2,209 | -63% | 1 | 1 | 0% | 856 | 1,119 | +31% | 0 | 0 | — |
case-11 | fail→pass | 2,430 | 2,569 | +6% | 1 | 1 | 0% | 419 | 1,261 | +201% | 0 | 0 | — |
case-12 | fail→pass | 9,045 | 3,065 | -66% | 1 | 1 | 0% | 1,349 | 1,335 | -1% | 0 | 0 | — |
case-13 | pass→pass | 13,392 | 7,749 | -42% | 1 | 1 | 0% | 2,133 | 2,093 | -2% | 0 | 0 | — |
case-14 | pass→pass | 8,378 | 4,993 | -40% | 1 | 1 | 0% | 1,271 | 1,564 | +23% | 0 | 0 | — |
case-15 | pass→fail | 6,057 | 5,778 | -5% | 1 | 1 | 0% | 1,006 | 1,080 | +7% | 0 | 0 | — |
case-16 | pass→pass | 6,043 | 3,414 | -44% | 1 | 1 | 0% | 1,065 | 1,421 | +33% | 0 | 0 | — |
case-17 | pass→pass | 6,500 | 3,657 | -44% | 1 | 1 | 0% | 1,008 | 1,375 | +36% | 0 | 0 | — |
case-18 | fail→pass | 13,013 | 6,991 | -46% | 1 | 1 | 0% | 2,228 | 2,025 | -9% | 0 | 0 | — |
case-19 | pass→pass | 7,858 | 6,542 | -17% | 1 | 1 | 0% | 1,626 | 2,042 | +26% | 0 | 0 | — |
case-20 | pass→pass | 11,552 | 13,954 | +21% | 1 | 1 | 0% | 2,075 | 3,047 | +47% | 0 | 0 | — |
case-21 | pass→pass | 9,717 | 8,170 | -16% | 1 | 1 | 0% | 1,857 | 2,406 | +30% | 0 | 0 | — |
case-22 | pass→pass | 7,295 | 3,255 | -55% | 1 | 1 | 0% | 1,189 | 1,357 | +14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 15 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.