Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Collect CoreWeave cluster diagnostics for support tickets. Use when preparing a support case, collecting GPU node status, or documenting pod failures. Trigger with phrases like "coreweave debug", "coreweave support", "coreweave diagnostics", "collect coreweave logs".
.claude/skills/jeremylongshore-coreweave-debug-bundle/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 51% | 0% |
> Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Collect GPU node health, Kubernetes pod status, event logs, and API connectivity into a single diagnostic archive for CoreWeave support tickets. This bundle captures cluster-level resource allocation, failed pod logs, GPU device plugin state, and network reachability so support engineers can diagnose infrastructure issues without requesting additional information. Useful when GPU pods are stuck pending, inference workloads OOM, or node autoscaling behaves unexpectedly.
bash#!/bin/bash set -euo pipefail BUNDLE="debug-coreweave-$(date +%Y%m%d-%H%M%S)" mkdir -p "$BUNDLE" # Environment check echo "=== CoreWeave Debug Bundle ===" | tee "$BUNDLE/summary.txt" echo "Generated: $(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$BUNDLE/summary.txt" echo "COREWEAVE_API_KEY: ${COREWEAVE_API_KEY:+[SET]}" >> "$BUNDLE/summary.txt" echo "KUBECONFIG: ${KUBECONFIG:-default}" >> "$BUNDLE/summary.txt" echo "kubectl: $(kubectl version --client --short 2>/dev/null || echo 'not found')" >> "$BUNDLE/summary.txt" # API connectivity HTTP=$(curl -s -o /dev/null -w "%{http_code}" -H "Authorization: Bearer ${COREWEAVE_API_KEY}" \ https://api.coreweave.com/v1/namespaces 2>/dev/null || echo "000") echo "API Status: HTTP $HTTP" >> "$BUNDLE/summary.txt" # Cluster state kubectl get nodes -o wide > "$BUNDLE/nodes.txt" 2>&1 || true kubectl get pods --all-namespaces -o wide > "$BUNDLE/pods.txt" 2>&1 || true kubectl get events --sort-by=.lastTimestamp > "$BUNDLE/events.txt" 2>&1 || true # GPU allocation and device plugin status kubectl describe nodes | grep -A10 "Allocated resources" > "$BUNDLE/gpu-allocation.txt" 2>&1 || true kubectl get pods -n kube-system -l k8s-app=nvidia-device-plugin -o wide > "$BUNDLE/gpu-plugin.txt" 2>&1 || true # Failed pod logs for pod in $(kubectl get pods --field-selector=status.phase=Failed -o name 2>/dev/null); do kubectl logs "$pod" --tail=200 > "$BUNDLE/$(basename "$pod")-logs.txt" 2>&1 || true done # Rate limit headers curl -s -D "$BUNDLE/rate-headers.txt" -o /dev/null \ -H "Authorization: Bearer ${COREWEAVE_API_KEY}" \ https://api.coreweave.com/v1/namespaces 2>/dev/null || true tar -czf "$BUNDLE.tar.gz" "$BUNDLE" && rm -rf "$BUNDLE" echo "Bundle: $BUNDLE.tar.gz"
bashtar -xzf debug-coreweave-*.tar.gz cat debug-coreweave-*/summary.txt # API + env status at a glance grep -i "error\|fail\|oom" debug-coreweave-*/events.txt # Critical events cat debug-coreweave-*/gpu-allocation.txt # GPU resource pressure
| Symptom | Check in Bundle | Fix | |---------|----------------|-----| | GPU pods stuck Pending | gpu-allocation.txt shows 0 allocatable GPUs | Request quota increase or switch to available GPU type | | OOMKilled on inference pod | events.txt for OOMKilled entries | Increase memory limits in pod spec; check model size vs allocated RAM | | Node NotReady | nodes.txt status column | Check events.txt for kubelet issues; contact CoreWeave if persistent | | API returns 401 | summary.txt shows HTTP 401 | Regenerate API key at CoreWeave dashboard; verify COREWEAVE_API_KEY is set | | NVIDIA device plugin missing | gpu-plugin.txt empty or error | Verify namespace kube-system has device plugin DaemonSet; redeploy if missing |
typescriptasync function checkCoreWeave(): Promise<void> { const key = process.env.COREWEAVE_API_KEY; if (!key) { console.error("[FAIL] COREWEAVE_API_KEY not set"); return; } const res = await fetch("https://api.coreweave.com/v1/namespaces", { headers: { Authorization: `Bearer ${key}` }, }); console.log(`[${res.ok ? "OK" : "FAIL"}] API: HTTP ${res.status}`); const limit = res.headers.get("x-ratelimit-remaining"); if (limit) console.log(`[INFO] Rate limit remaining: ${limit}`); } checkCoreWeave();
See coreweave-common-errors for GPU scheduling and Kubernetes troubleshooting patterns.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 23,775 | 15,379 | -35% | 1 | 1 | 0% | 5,259 | 4,948 | -6% | 0 | 0 | — |
case-02 | fail→pass | 11,646 | 5,769 | -50% | 1 | 1 | 0% | 2,210 | 2,374 | +7% | 0 | 0 | — |
case-03 | pass→pass | 10,931 | 7,142 | -35% | 1 | 1 | 0% | 2,450 | 2,929 | +20% | 0 | 0 | — |
case-04 | fail→pass | 11,689 | 3,552 | -70% | 1 | 1 | 0% | 2,005 | 1,830 | -9% | 0 | 0 | — |
case-05 | fail→pass | 15,164 | 10,176 | -33% | 1 | 1 | 0% | 2,719 | 3,101 | +14% | 0 | 0 | — |
case-06 | fail→pass | 7,957 | 3,345 | -58% | 1 | 1 | 0% | 1,335 | 1,851 | +39% | 0 | 0 | — |
case-07 | pass→pass | 9,809 | 5,557 | -43% | 1 | 1 | 0% | 1,859 | 2,287 | +23% | 0 | 0 | — |
case-08 | pass→pass | 12,055 | 4,421 | -63% | 1 | 1 | 0% | 2,181 | 2,132 | -2% | 0 | 0 | — |
case-09 | fail→pass | 7,116 | 3,706 | -48% | 1 | 1 | 0% | 1,332 | 2,005 | +51% | 0 | 0 | — |
case-10 | pass→pass | 3,873 | 2,530 | -35% | 1 | 1 | 0% | 697 | 1,693 | +143% | 0 | 0 | — |
case-11 | pass→pass | 2,949 | 2,234 | -24% | 1 | 1 | 0% | 413 | 1,616 | +291% | 0 | 0 | — |
case-12 | pass→pass | 6,492 | 2,960 | -54% | 1 | 1 | 0% | 1,150 | 1,832 | +59% | 0 | 0 | — |
case-18 | fail→pass | 7,663 | 2,646 | -65% | 1 | 1 | 0% | 1,322 | 1,770 | +34% | 0 | 0 | — |
case-13 | fail→pass | 7,409 | 2,330 | -69% | 1 | 1 | 0% | 1,362 | 1,626 | +19% | 0 | 0 | — |
case-14 | pass→fail | 13,386 | 6,958 | -48% | 1 | 1 | 0% | 2,478 | 2,604 | +5% | 0 | 0 | — |
case-15 | pass→pass | 4,993 | 3,122 | -37% | 1 | 1 | 0% | 855 | 1,755 | +105% | 0 | 0 | — |
case-16 | pass→pass | 9,101 | 5,759 | -37% | 1 | 1 | 0% | 977 | 1,765 | +81% | 0 | 0 | — |
case-17 | fail→pass | 23,836 | 5,917 | -75% | 1 | 1 | 0% | 1,399 | 2,338 | +67% | 0 | 0 | — |
case-19 | pass→pass | 3,739 | 1,295 | -65% | 1 | 1 | 0% | 606 | 1,391 | +130% | 0 | 0 | — |
case-20 | pass→pass | 8,451 | 7,901 | -7% | 1 | 1 | 0% | 1,556 | 2,727 | +75% | 0 | 0 | — |
case-21 | pass→pass | 8,542 | 8,346 | -2% | 1 | 1 | 0% | 1,564 | 2,833 | +81% | 0 | 0 | — |
case-22 | pass→pass | 11,594 | 9,516 | -18% | 1 | 1 | 0% | 2,213 | 3,137 | +42% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.