Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Periodically check WandB metrics during training to catch problems early (NaN, loss divergence, idle GPUs). Avoids wasting GPU hours on broken runs. Use when training is running and you want automated health checks.
.claude/skills/wanshuiyin-training-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 439% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 28% | 0% |
Periodically read WandB metrics during training to catch problems early. Do not wait until training finishes to discover it was a waste of GPU time.
> ⏱ This skill is correctly cron-wired (see below): it polls > machine-checkable training health (NaN / divergence / idle GPU) — the additive > external-wait shape in > shared-references/external-cadence.md. > The occasional Codex call for an ambiguous metric is a one-shot check per > tick, not a multi-round verdict loop, so it stays additive — it never grows > into a wrapped verdict skill.
entity/project/run_id)gpt-6-astra — used via Codex MCP for ambiguous cases onlypythonimport wandb api = wandb.Api() run = api.run("<entity>/<project>/<run_id>") history = run.history()
If WandB is unreachable (API error, network issue), fall back to reading the log file directly via SSH:
bashssh server "tail -100 /path/to/training.log"
Check these signals:
| Signal | Judgment | Action | |--------|----------|--------| | NaN/Inf in loss | Clearly bad | Stop training, investigate | | Loss diverging (increasing for >N steps) | Clearly bad | Stop training, investigate | | Eval metrics significantly worse than baseline | Clearly bad | Stop training, investigate | | Loss decreasing, metrics improving | Clearly fine | Continue, increase check interval | | Loss flat but not diverging | Unsure | → Step 3 (Codex judgment) | | Metrics noisy, can't tell trend | Unsure | → Step 3 (Codex judgment) | | Slightly worse than baseline but still early | Unsure | → Step 3 (Codex judgment) |
Only escalate to Codex when the signal is ambiguous. For clearly good or clearly bad signals, act directly.
mcp__codex__codex:
model: gpt-6-astra
config: {"model_reasoning_effort": "xhigh"}
prompt: |
TRAINING HEALTH CHECK — need your judgment on ambiguous metrics.
Run: <entity>/<project>/<run_id>
Current epoch/step: X / Y total
Training loss (last 10 checkpoints): [values]
Eval metrics (last 3 evals): [values]
Baseline reference: [numbers from paper/reproduction]
What I'm unsure about: [specific concern]
Please respond with exactly one of:
- STOP: clearly problematic, should kill training
- CONTINUE: looks fine, check again next interval
- WAIT: not enough data to judge, check again sooner| Decision | Action | |----------|--------| | Stop | Kill the training session. Save the WandB run URL, key metrics, and reason for stopping. Log to project notes for debugging. | | Continue | Do nothing. Will be invoked again at next interval (increase interval if consistently healthy). | | Wait | Do nothing but keep the current short interval (don't increase). |
Training-check and watchdog.py operate at different levels:
| Layer | Tool | What it checks | Frequency | |-------|------|----------------|-----------| | Process health | watchdog.py | Session alive? GPU active? | Every 60s (continuous) | | Training quality | training-check | Loss trend? Metrics improving? | Every 10-60 min (periodic) |
Use both together:
After training is confirmed stable:
CronCreate (recurring, every 10 minutes initially):
"Run /training-check for wandb run <entity>/<project>/<run_id>"As the check interval increases, delete the old CronCreate job and create a new one with the longer interval.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→pass | 14,922 | 8,792 | -41% | 1 | 1 | 0% | 1,531 | 2,027 | +32% | 0 | 0 | — |
case-01 | fail→fail | 19,683 | 16,798 | -15% | 1 | 1 | 0% | 324 | 1,843 | +469% | 0 | 0 | — |
case-02 | pass→fail | 64,233 | 46,266 | -28% | 1 | 1 | 0% | 5,744 | 1,877 | -67% | 0 | 0 | — |
case-03 | fail→fail | 22,210 | 58,615 | +164% | 1 | 1 | 0% | 2,852 | 1,826 | -36% | 0 | 0 | — |
case-04 | fail→pass | 9,235 | 15,357 | +66% | 1 | 1 | 0% | 561 | 3,021 | +439% | 0 | 0 | — |
case-05 | pass→fail | 10,538 | 14,982 | +42% | 1 | 1 | 0% | 858 | 1,566 | +83% | 0 | 0 | — |
case-06 | pass→pass | 21,598 | 20,397 | -6% | 1 | 1 | 0% | 2,738 | 3,984 | +46% | 0 | 0 | — |
case-07 | pass→pass | 14,523 | 10,353 | -29% | 1 | 1 | 0% | 1,406 | 2,214 | +57% | 0 | 0 | — |
case-08 | pass→pass | 11,173 | 10,248 | -8% | 1 | 1 | 0% | 937 | 2,278 | +143% | 0 | 0 | — |
case-09 | fail→fail | 19,194 | 20,398 | +6% | 1 | 1 | 0% | 2,229 | 2,048 | -8% | 0 | 0 | — |
case-10 | fail→pass | 22,790 | 9,470 | -58% | 1 | 1 | 0% | 2,776 | 2,280 | -18% | 0 | 0 | — |
case-11 | fail→pass | 21,924 | 11,597 | -47% | 1 | 1 | 0% | 2,662 | 2,353 | -12% | 0 | 0 | — |
case-12 | fail→pass | 19,724 | 10,458 | -47% | 1 | 1 | 0% | 2,202 | 2,100 | -5% | 0 | 0 | — |
case-13 | fail→pass | 26,387 | 8,823 | -67% | 1 | 1 | 0% | 1,544 | 1,977 | +28% | 0 | 0 | — |
case-14 | pass→pass | 16,890 | 7,696 | -54% | 1 | 1 | 0% | 1,663 | 1,828 | +10% | 0 | 0 | — |
case-15 | fail→pass | 15,504 | 9,489 | -39% | 1 | 1 | 0% | 1,540 | 2,112 | +37% | 0 | 0 | — |
case-16 | fail→pass | 17,246 | 9,486 | -45% | 1 | 1 | 0% | 1,883 | 2,168 | +15% | 0 | 0 | — |
case-17 | pass→pass | 16,548 | 12,142 | -27% | 1 | 1 | 0% | 1,852 | 2,433 | +31% | 0 | 0 | — |
case-18 | fail→fail | 14,592 | 7,605 | -48% | 1 | 1 | 0% | 1,315 | 1,800 | +37% | 0 | 0 | — |
case-19 | pass→pass | 14,984 | 11,115 | -26% | 1 | 1 | 0% | 1,632 | 2,179 | +34% | 0 | 0 | — |
case-21 | fail→pass | 14,800 | 7,873 | -47% | 1 | 1 | 0% | 1,542 | 1,891 | +23% | 0 | 0 | — |
case-22 | pass→pass | 14,961 | 9,622 | -36% | 1 | 1 | 0% | 1,665 | 2,117 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/11/2026 | +41% |
Other measured skills in the registry, with their headline benchmark lift.