Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deploy inference services on CoreWeave with Helm charts and Kustomize. Use when deploying multi-model inference, managing GPU deployments at scale, or templating CoreWeave manifests. Trigger with phrases like "deploy coreweave", "coreweave helm", "coreweave kustomize", "coreweave deployment patterns".
.claude/skills/jeremylongshore-coreweave-deploy-integration/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 9% | 0% |
> Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Deploy GPU-accelerated inference services on CoreWeave Kubernetes (CKS). This skill covers containerizing inference workloads with NVIDIA CUDA base images, configuring GPU resource limits and node affinity for A100/H100 scheduling, setting up health checks that validate GPU availability and model loading, and executing rolling updates that respect GPU node draining. CoreWeave's scheduler requires explicit GPU resource requests to place pods on the correct hardware tier.
dockerfileFROM nvidia/cuda:12.4.0-runtime-ubuntu22.04 AS base RUN apt-get update && apt-get install -y --no-install-recommends \ python3 python3-pip curl && rm -rf /var/lib/apt/lists/* WORKDIR /app COPY requirements.txt ./ RUN pip3 install --no-cache-dir -r requirements.txt FROM base RUN groupadd -r app && useradd -r -g app app COPY --chown=app:app src/ ./src/ COPY --chown=app:app models/ ./models/ USER app EXPOSE 8080 HEALTHCHECK --interval=30s --timeout=10s --retries=3 \ CMD curl -f http://localhost:8080/health || exit 1 CMD ["python3", "src/server.py"]
bashexport COREWEAVE_API_KEY="cw_xxxxxxxxxxxx" export COREWEAVE_NAMESPACE="tenant-my-org" export MODEL_NAME="meta-llama/Llama-3.1-8B-Instruct" export GPU_TYPE="A100_PCIE_80GB" export GPU_COUNT="1" export LOG_LEVEL="info" export PORT="8080"
typescriptimport express from 'express'; import { execSync } from 'child_process'; const app = express(); app.get('/health', async (req, res) => { try { const gpuInfo = execSync('nvidia-smi --query-gpu=name,memory.used --format=csv,noheader').toString().trim(); const modelLoaded = globalThis.modelReady === true; if (!modelLoaded) throw new Error('Model not loaded'); res.json({ status: 'healthy', gpu: gpuInfo, model: process.env.MODEL_NAME, timestamp: new Date().toISOString() }); } catch (error) { res.status(503).json({ status: 'unhealthy', error: (error as Error).message }); } });
bashdocker build -t registry.coreweave.com/my-org/inference-svc:latest . docker push registry.coreweave.com/my-org/inference-svc:latest
yaml# k8s/deployment.yaml resources: limits: nvidia.com/gpu: 1 cpu: "4" memory: "48Gi" nodeSelector: gpu.nvidia.com/class: A100_PCIE_80GB
bashkubectl apply -f k8s/deployment.yaml -n tenant-my-org
bashkubectl get pods -n tenant-my-org -l app=inference-svc curl -s http://inference-svc.tenant-my-org.svc.cluster.local:8080/health | jq .
bashkubectl set image deployment/inference-svc \ inference=registry.coreweave.com/my-org/inference-svc:v2 \ -n tenant-my-org kubectl rollout status deployment/inference-svc -n tenant-my-org --timeout=600s
| Issue | Cause | Fix | |-------|-------|-----| | Pending pod stuck | No GPU nodes available for requested type | Check kubectl describe node for allocatable GPUs or switch GPU tier | | OOMKilled | Model exceeds GPU memory | Reduce model size, enable quantization, or request larger GPU | | nvidia-smi not found | Missing NVIDIA device plugin | Verify CoreWeave namespace has GPU operator installed | | 401 Unauthorized | Invalid API key or expired credentials | Regenerate key in CoreWeave dashboard | | Slow rolling update | GPU nodes take time to drain | Set terminationGracePeriodSeconds: 300 in deployment spec |
See coreweave-webhooks-events.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | fail→fail | 13,345 | 7,280 | -45% | 1 | 1 | 0% | 2,265 | 2,526 | +12% | 0 | 0 | — |
case-01 | fail→pass | 15,764 | 10,228 | -35% | 1 | 1 | 0% | 3,334 | 3,416 | +2% | 0 | 0 | — |
case-02 | fail→fail | 17,960 | 10,834 | -40% | 1 | 1 | 0% | 3,723 | 3,475 | -7% | 0 | 0 | — |
case-03 | fail→pass | 15,212 | 11,299 | -26% | 1 | 1 | 0% | 3,195 | 3,683 | +15% | 0 | 0 | — |
case-04 | pass→pass | 18,063 | 15,956 | -12% | 1 | 1 | 0% | 3,517 | 4,424 | +26% | 0 | 0 | — |
case-05 | pass→pass | 11,118 | 10,054 | -10% | 1 | 1 | 0% | 2,331 | 3,288 | +41% | 0 | 0 | — |
case-06 | pass→pass | 11,975 | 10,877 | -9% | 1 | 1 | 0% | 2,757 | 3,777 | +37% | 0 | 0 | — |
case-07 | fail→pass | 12,987 | 2,855 | -78% | 1 | 1 | 0% | 2,636 | 1,678 | -36% | 0 | 0 | — |
case-08 | fail→fail | 23,815 | 4,880 | -80% | 1 | 1 | 0% | 2,072 | 2,039 | -2% | 0 | 0 | — |
case-15 | fail→fail | 11,715 | 8,813 | -25% | 1 | 1 | 0% | 2,109 | 2,857 | +35% | 0 | 0 | — |
case-09 | fail→pass | 6,607 | 2,798 | -58% | 1 | 1 | 0% | 1,430 | 1,779 | +24% | 0 | 0 | — |
case-10 | fail→pass | 8,792 | 2,982 | -66% | 1 | 1 | 0% | 1,611 | 1,750 | +9% | 0 | 0 | — |
case-11 | fail→fail | 13,496 | 5,380 | -60% | 1 | 1 | 0% | 2,650 | 2,216 | -16% | 0 | 0 | — |
case-12 | fail→pass | 10,134 | 6,722 | -34% | 1 | 1 | 0% | 1,919 | 2,485 | +29% | 0 | 0 | — |
case-13 | fail→pass | 11,141 | 4,225 | -62% | 1 | 1 | 0% | 2,141 | 2,083 | -3% | 0 | 0 | — |
case-16 | fail→pass | 16,317 | 6,274 | -62% | 1 | 1 | 0% | 2,680 | 2,326 | -13% | 0 | 0 | — |
case-17 | fail→pass | 11,798 | 1,714 | -85% | 1 | 1 | 0% | 2,180 | 1,448 | -34% | 0 | 0 | — |
case-18 | pass→pass | 14,938 | 12,491 | -16% | 1 | 1 | 0% | 2,623 | 3,233 | +23% | 0 | 0 | — |
case-19 | fail→pass | 6,935 | 3,220 | -54% | 1 | 1 | 0% | 1,351 | 1,726 | +28% | 0 | 0 | — |
case-20 | fail→pass | 3,969 | 2,726 | -31% | 1 | 1 | 0% | 724 | 1,618 | +123% | 0 | 0 | — |
case-21 | fail→pass | 7,635 | 2,660 | -65% | 1 | 1 | 0% | 1,617 | 1,637 | +1% | 0 | 0 | — |
case-22 | pass→pass | 8,721 | 2,729 | -69% | 1 | 1 | 0% | 1,745 | 1,764 | +1% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.