Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a CFO-grokkable, dollar-ranked FinOps report. CoreWeave ships no cost dashboard and no billing API, so the spend view is built from PromQL against its managed Grafana. Use when a user asks why their CoreWeave GPU bill is high, wants to find wasted GPU spend or idle reservations, or needs a GPU FinOps cost report.
.claude/skills/jeremylongshore-coreweave-gpu-cost-leak-hunter/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 132% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 90% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 122% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 151% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 146% | 0% |
> Community-contributed. Not affiliated with, endorsed by, or sponsored by > CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Audits a CoreWeave GPU cluster for real-dollar cost leaks — idle reserved capacity, GPUs on the wrong SKU, allocated-but-idle instances, and steady on-demand spend that should be committed — then emits a CFO-grokkable, dollar-ranked FinOps report.
CoreWeave ships no cost dashboard and no billing API (usage-monitoring docs]um]). There is also no single "dollars" metric — spend is reconstructed by querying usage from CoreWeave's managed Grafana in PromQL and multiplying each resource's usage by its rate-card price. This skill does exactly that, then ranks the leaks by monthly dollar impact.
The math is deterministic: PromQL returns usage counts, and the bundled scripts/rank-and-report.py does every multiplication, sum, and ranking — the agent never eyeballs a number. Two of the four categories are billed waste (Confirmed); the other two are a right-sizing model (Estimated) and a commitment decision (At-risk), labeled so a CFO never reads a modeled number as recoverable cash. Deep domain knowledge lives in references/, loaded only when a leak needs it.
to a member of the admin, metrics, or write group in the CoreWeave Cloud Console (usage-monitoring docs]um]). This is the hard dependency; Step 1 probes it and fails fast if the group is missing.
$CW_PROM_URL (the Grafana data-sourceproxy, e.g. https://grafana.ORG.coreweave.com/api/datasources/proxy/uid/UID) and a bearer token in $CW_TOKEN for curl.
kubeconfig for the cluster (CoreWeave-issued) so kubectl get can corroboratelive GPU allocation and node labels.
rates are supplied to the ranker from references/gpu-right-sizing.md (dated snapshot of coreweave.com/pricing]pr]) or the customer's contract.
jq and python3 for parsing query JSON and running the ranker.Authentication. All auth comes from the environment ($CW_PROM_URL, $CW_TOKEN, $KUBECONFIG) — no secrets are hardcoded. Grafana enforces the group membership above on every query.
The pipeline is detect → price → rank → report. PromQL returns usage; the dollar arithmetic runs in scripts/; deep knowledge loads from references/ on demand:
Probe billing:instance:total before anything else. An HTTP 401/403 or empty result means the token's principal is not in admin/metrics/write — STOP and report it; do not continue into the scans.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" \ --data-urlencode 'query=count(billing:instance:total)' \ "$CW_PROM_URL/api/v1/query" | jq -r '.status, (.data.result | length)'
If status is not success with a non-empty result, load references/promql-billing-setup.md and report the missing group access verbatim. Stop here.
Reconstruct 30-day GPU node-hours per instance type. CoreWeave has no dollars metric, so this returns usage — the ranker multiplies by the rate card. Write the JSON to the working dir for the ranker.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" \ --data-urlencode 'query=sum by (instance_type) (sum_over_time(billing:instance:total[30d:1h]))' \ "$CW_PROM_URL/api/v1/query" > "$OUT/baseline.json"
The rate card and the per-category PromQL live in references/gpu-cost-leak-categories.md. Load it now — the four scans below reference its recording-rule notes.
Reserved GPUs bill at the committed rate whether used or not. A reserved GPU sitting below a utilization floor is confirmed waste — you paid for it and it did no work. Cross reserved allocation (billing_gpu, filtered by the reservation label) against SM-active from DCGM.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \ 'query=sum by (instance_type,node) (avg_over_time(billing_gpu{reservation!=""}[30d:1h])) and on(node) (avg by (node) (avg_over_time(DCGM_FI_PROF_SM_ACTIVE[30d:1h])) < 0.05)' \ "$CW_PROM_URL/api/v1/query" > "$OUT/leak1-idle-reserved.json"
The reservation label key is provider-specific — confirm yours with kubectl get nodes --show-labels. Waste = idle reserved GPU-hours × committed rate (ranker input).
H100/H200 running small-model (~7B–30B) inference is over-paying: for that regime L40S is cheaper per token (directional — see gpu-right-sizing.md). Flag those instance-hours; the ranker re-prices them at the L40S rate.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \ 'query=sum by (instance_type) (sum_over_time(billing:instance:total{instance_type=~".*(h100|h200).*"}[30d:1h]))' \ "$CW_PROM_URL/api/v1/query" > "$OUT/leak2-wrong-gpu.json"
This is Estimated: the rate delta is exact rate-card math, but throughput equivalence on L40S is a model. Confirm the served model size with the cluster owner before acting; FP8 serving needs Hopper/Ada, not Ampere (see gpu-right-sizing.md).
On-demand GPUs that are allocated (billing) but running at low SM-utilization / low MFU bill the full on-demand rate for no work — confirmed billed waste, the GPU twin of an idle cluster.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \ 'query=(avg by (node,instance_type) (avg_over_time(DCGM_FI_PROF_SM_ACTIVE[30d:1h])) < 0.05) and on(node) (sum by (node) (avg_over_time(billing_gpu{reservation=""}[30d:1h])) > 0)' \ "$CW_PROM_URL/api/v1/query" > "$OUT/leak3-idle-ondemand.json"
Corroborate with kubectl get pods -A --field-selector=status.phase=Running to confirm nothing is actually scheduled on the flagged node. Waste = idle on-demand GPU-hours × on-demand rate.
A stable on-demand floor — GPUs of one type always running across the window — is paying on-demand for capacity a commitment discounts up to 60% (pricing]pr]). Measure the always-on floor with min_over_time.
bashcurl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \ 'query=min_over_time(sum by (instance_type) (billing:instance:total{reservation=""})[30d:1h])' \ "$CW_PROM_URL/api/v1/query" > "$OUT/leak4-commit-gap.json"
This is At-risk: the up-to-60% saving is pending a commitment decision, and a commitment is itself a paid obligation — see the over-reservation caution in gpu-cost-leak-categories.md. Savings = floor GPU-hours × on-demand rate × discount.
Assemble one leak object per category from the PromQL usage results plus the rate card, then pipe them to the deterministic ranker — the LLM does NOT do the arithmetic. Because CoreWeave exposes no dollars metric, each object carries usage_gpu_hours and its rate-card rate; the ranker multiplies usage × rate itself (and applies the re-price or discount factor). Each object's kind (confirmed / estimated / at-risk) tells the renderer to split the headline confirmed-vs-pending, rank descending by monthly dollars, and stamp a Confidence column.
bashOUT="${OUT:-$(pwd)/cost-leak-out}" && mkdir -p "$OUT" # Each Step wrote a leak-N.json {category, root_cause, fix, kind, usage_gpu_hours, # rate_usd_per_gpu_hour, ...}; the ranker does usage × rate deterministically. jq -s '.' "$OUT"/leak-*.json | \ python3 scripts/rank-and-report.py \ --monthly-spend 180000 --window-end "$WINDOW_END" \ --out "$OUT/cost-leak-report.md"
Render the output using the verbatim template in references/cfo-output-format.md. Use Glob to collect the per-leak JSON, Write the report, and Edit it to rescale the headline spend on request.
confirmed and unconfirmed dollars under one verb — A $180K/month CoreWeave GPU cluster is burning ~$40K/month (confirmed), plus up to ~$29K/month pending review — each with a /year companion.
fix), one row per category, highest dollar impact first, each fix a single change.
found them, and the underlying $/GPU-hour rates for the cluster engineer.
| Error | Cause | Solution | |-------|-------|----------| | HTTP 401/403 on /api/v1/query | Token principal not in admin/metrics/write | Run Step 1; report the group requirement from promql-billing-setup.md. Stop. | | Empty result for billing:instance:total | Wrong data-source proxy UID, or org has no billing metrics enabled | Verify $CW_PROM_URL points at the Grafana Prometheus proxy; confirm in Grafana Explore. | | DCGM_* series absent | DCGM exporter not scraped on the node pool | Skip Leaks 1/3 utilization filter for that pool; note "utilization unavailable" rather than reporting $0. | | reservation label missing | Provider label key differs per org | Confirm the reservation/committed label with kubectl get nodes --show-labels; substitute it in the query. | | Ranker prints ~$0/month confirmed | A kind value was mis-cased and dropped from the sum | The ranker normalizes case; verify each leak object's kind is one of the three tiers. |
Runs the full pipeline. The access probe passes, the four scans return rows, and the ranker emits a split, confidence-stamped report:
text### A $180K/month CoreWeave GPU cluster is burning **~$44,986/month** (confirmed), plus up to **~$29,110/month** pending review Trailing 30 days ending 2026-06-22. Confirmed **~$540K/year**; up to **~$349K/year** more pending review. Spend is reconstructed from PromQL against CoreWeave's managed Grafana (no billing API). Every line below is one change. | # | Where it's leaking | $/month | Confidence | The fix | |---|---|--:|---|---| | 1 | **Idle reserved GPUs** — reserved capacity billing around the clock below a utilization floor | **$26,280** | Confirmed | Right-size or release the reservation | | 2 | **Allocated-but-idle on-demand GPUs** — nodes up at <5% SM-active, paying full rate for no work | **$18,706** | Confirmed | Scale-to-zero / deschedule the idle nodes | | 3 | **H100/H200 on small-model inference** — L40S is cheaper per token in the 7B–30B regime | **$16,629** | Estimated | Move small inference to L40S | | 4 | **Steady on-demand that should be committed** — an always-on floor paying on-demand | **$12,481** | At-risk | Commit the stable floor (up to 60% off) | **The #1 line alone — idle reserved gpus (confirmed) — is ~$315K/year, fixed in one setting.**
User asks "are we paying for idle reserved GPUs?" The skill runs Step 3 only, crosses billing_gpu{reservation!=""} against DCGM_FI_PROF_SM_ACTIVE, and reports each reserved node below the floor with its 30-day committed spend.
references/gpu-cost-leak-categories.md — the four leak categories: definition, PromQL, root cause, the one fix.references/cfo-output-format.md — verbatim CFO report template + the never-sum invariant.references/gpu-right-sizing.md — L40S/H100/H200/A100/L40 decision table + the FP8 rule (figures flagged directional).references/promql-billing-setup.md — the billing metrics + group access, cited.coreweave-cost-tuning authors cost-control config; this skill detects leaks and dollarizes them.um]: https://docs.coreweave.com/docs/observability/usage-monitoring pr]: https://www.coreweave.com/pricing
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,989 | 14,216 | -32% | 1 | 1 | 0% | 3,803 | 4,251 | +12% | 0 | 0 | — |
case-02 | fail→fail | 30,734 | 6,413 | -79% | 1 | 1 | 0% | 6,000 | 4,431 | -26% | 0 | 0 | — |
case-03 | fail→fail | 13,507 | 7,170 | -47% | 1 | 1 | 0% | 2,500 | 4,290 | +72% | 0 | 0 | — |
case-04 | pass→pass | 10,897 | 8,343 | -23% | 1 | 1 | 0% | 2,290 | 5,635 | +146% | 0 | 0 | — |
case-05 | pass→pass | 12,424 | 9,472 | -24% | 1 | 1 | 0% | 2,485 | 5,600 | +125% | 0 | 0 | — |
case-06 | pass→pass | 13,087 | 14,628 | +12% | 1 | 1 | 0% | 3,044 | 7,229 | +137% | 0 | 0 | — |
case-07 | fail→pass | 12,515 | 5,750 | -54% | 1 | 1 | 0% | 2,122 | 4,917 | +132% | 0 | 0 | — |
case-08 | pass→pass | 15,949 | 9,474 | -41% | 1 | 1 | 0% | 2,796 | 5,729 | +105% | 0 | 0 | — |
case-09 | fail→pass | 16,573 | 7,796 | -53% | 1 | 1 | 0% | 2,668 | 5,077 | +90% | 0 | 0 | — |
case-10 | fail→pass | 13,881 | 8,240 | -41% | 1 | 1 | 0% | 2,349 | 5,206 | +122% | 0 | 0 | — |
case-11 | fail→pass | 9,676 | 3,896 | -60% | 1 | 1 | 0% | 1,802 | 4,520 | +151% | 0 | 0 | — |
case-12 | fail→fail | 13,721 | 4,810 | -65% | 1 | 1 | 0% | 790 | 4,482 | +467% | 0 | 0 | — |
case-13 | fail→pass | 12,482 | 7,983 | -36% | 1 | 1 | 0% | 2,105 | 5,173 | +146% | 0 | 0 | — |
case-14 | fail→pass | 13,656 | 2,666 | -80% | 1 | 1 | 0% | 1,526 | 4,319 | +183% | 0 | 0 | — |
case-15 | fail→pass | 13,617 | 7,551 | -45% | 1 | 1 | 0% | 2,228 | 5,172 | +132% | 0 | 0 | — |
case-16 | fail→pass | 12,275 | 6,335 | -48% | 1 | 1 | 0% | 2,148 | 4,875 | +127% | 0 | 0 | — |
case-17 | pass→pass | 10,757 | 7,873 | -27% | 1 | 1 | 0% | 1,933 | 5,277 | +173% | 0 | 0 | — |
case-22 | pass→pass | 13,652 | 9,997 | -27% | 1 | 1 | 0% | 2,300 | 5,427 | +136% | 0 | 0 | — |
case-18 | pass→pass | 7,836 | 3,146 | -60% | 1 | 1 | 0% | 1,173 | 4,335 | +270% | 0 | 0 | — |
case-19 | fail→fail | 11,535 | 4,259 | -63% | 1 | 1 | 0% | 1,864 | 4,539 | +144% | 0 | 0 | — |
case-20 | fail→pass | 6,988 | 2,092 | -70% | 1 | 1 | 0% | 1,145 | 4,214 | +268% | 0 | 0 | — |
case-21 | fail→pass | 5,943 | 3,477 | -41% | 1 | 1 | 0% | 1,035 | 4,441 | +329% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.