Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. Use when training models across multiple GPUs, setting up distributed training, or running fine-tuning jobs on CoreWeave H100 clusters. Trigger with phrases like "coreweave training", "coreweave multi-gpu", "distributed training coreweave", "fine-tune on coreweave".
.claude/skills/jeremylongshore-coreweave-core-workflow-b/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -22% | 0% |
> Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Run distributed GPU training on CoreWeave: single-node multi-GPU and multi-node training with PyTorch DDP, Slurm-on-Kubernetes, and shared storage.
yaml# training-job.yaml apiVersion: batch/v1 kind: Job metadata: name: llm-finetune spec: template: spec: restartPolicy: Never containers: - name: trainer image: ghcr.io/myorg/trainer:latest command: ["torchrun"] args: - "--nproc_per_node=8" - "train.py" - "--model_name=meta-llama/Llama-3.1-8B" - "--batch_size=4" - "--epochs=3" resources: limits: nvidia.com/gpu: "8" memory: 512Gi cpu: "64" volumeMounts: - name: data mountPath: /data - name: checkpoints mountPath: /checkpoints volumes: - name: data persistentVolumeClaim: claimName: training-data - name: checkpoints persistentVolumeClaim: claimName: model-checkpoints affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: gpu.nvidia.com/class operator: In values: ["A100_NVLINK_A100_SXM4_80GB"]
yaml# storage.yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: name: training-data spec: accessModes: ["ReadWriteMany"] resources: requests: storage: 500Gi storageClassName: shared-hdd-ord1 --- apiVersion: v1 kind: PersistentVolumeClaim metadata: name: model-checkpoints spec: accessModes: ["ReadWriteMany"] resources: requests: storage: 200Gi storageClassName: shared-ssd-ord1
bash# Watch training logs kubectl logs -f job/llm-finetune # Check GPU utilization kubectl exec -it $(kubectl get pod -l job-name=llm-finetune -o name) -- nvidia-smi # Check training metrics kubectl exec -it $(kubectl get pod -l job-name=llm-finetune -o name) -- \ cat /checkpoints/training_log.json | tail -5
| Error | Cause | Solution | |-------|-------|----------| | NCCL timeout | Network issue between GPUs | Use NVLink nodes (SXM4/SXM5) | | OOMKilled | Batch size too large | Reduce batch size or use gradient accumulation | | Checkpoint save failed | PVC full | Increase storage or prune old checkpoints | | Job evicted | Preemption | Use on-demand nodes for training |
For troubleshooting, see coreweave-common-errors.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 17,818 | 8,849 | -50% | 1 | 1 | 0% | 3,988 | 2,934 | -26% | 0 | 0 | — |
case-02 | fail→fail | 16,316 | 12,637 | -23% | 1 | 1 | 0% | 3,165 | 3,340 | +6% | 0 | 0 | — |
case-03 | fail→pass | 12,135 | 7,260 | -40% | 1 | 1 | 0% | 2,531 | 2,443 | -3% | 0 | 0 | — |
case-04 | pass→pass | 14,442 | 15,052 | +4% | 1 | 1 | 0% | 2,982 | 4,056 | +36% | 0 | 0 | — |
case-05 | pass→pass | 19,564 | 16,091 | -18% | 1 | 1 | 0% | 4,344 | 4,486 | +3% | 0 | 0 | — |
case-06 | pass→pass | 14,284 | 11,454 | -20% | 1 | 1 | 0% | 2,857 | 3,424 | +20% | 0 | 0 | — |
case-07 | fail→pass | 7,363 | 2,239 | -70% | 1 | 1 | 0% | 1,579 | 1,430 | -9% | 0 | 0 | — |
case-08 | fail→fail | 12,937 | 8,181 | -37% | 1 | 1 | 0% | 2,401 | 2,598 | +8% | 0 | 0 | — |
case-09 | pass→pass | 5,144 | 1,934 | -62% | 1 | 1 | 0% | 897 | 1,211 | +35% | 0 | 0 | — |
case-10 | pass→pass | 11,489 | 5,414 | -53% | 1 | 1 | 0% | 1,949 | 1,866 | -4% | 0 | 0 | — |
case-11 | fail→pass | 11,780 | 7,220 | -39% | 1 | 1 | 0% | 2,086 | 2,188 | +5% | 0 | 0 | — |
case-12 | pass→pass | 13,198 | 12,190 | -8% | 1 | 1 | 0% | 2,343 | 3,235 | +38% | 0 | 0 | — |
case-13 | fail→fail | 6,721 | 4,875 | -27% | 1 | 1 | 0% | 1,257 | 1,985 | +58% | 0 | 0 | — |
case-14 | pass→pass | 11,644 | 4,390 | -62% | 1 | 1 | 0% | 2,391 | 1,802 | -25% | 0 | 0 | — |
case-15 | fail→fail | 15,135 | 12,261 | -19% | 1 | 1 | 0% | 2,754 | 3,177 | +15% | 0 | 0 | — |
case-16 | fail→pass | 8,914 | 2,072 | -77% | 1 | 1 | 0% | 1,636 | 1,274 | -22% | 0 | 0 | — |
case-17 | fail→pass | 9,368 | 6,657 | -29% | 1 | 1 | 0% | 1,843 | 2,309 | +25% | 0 | 0 | — |
case-18 | pass→pass | 7,197 | 4,286 | -40% | 1 | 1 | 0% | 1,250 | 1,588 | +27% | 0 | 0 | — |
case-19 | fail→fail | 30,320 | 7,855 | -74% | 1 | 1 | 0% | 1,485 | 2,406 | +62% | 0 | 0 | — |
case-20 | fail→pass | 8,322 | 2,848 | -66% | 1 | 1 | 0% | 1,635 | 1,562 | -4% | 0 | 0 | — |
case-21 | pass→pass | 15,296 | 8,794 | -43% | 1 | 1 | 0% | 2,580 | 2,435 | -6% | 0 | 0 | — |
case-22 | pass→pass | 6,693 | 2,305 | -66% | 1 | 1 | 0% | 1,251 | 1,394 | +11% | 0 | 0 | — |
case-23 | pass→pass | 4,026 | 2,936 | -27% | 1 | 1 | 0% | 724 | 1,407 | +94% | 0 | 0 | — |
case-24 | fail→pass | 9,314 | 2,115 | -77% | 1 | 1 | 0% | 1,723 | 1,366 | -21% | 0 | 0 | — |
case-25 | pass→pass | 7,328 | 4,317 | -41% | 1 | 1 | 0% | 1,365 | 1,724 | +26% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 24 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.