Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deploy KServe InferenceService on CoreWeave with autoscaling and GPU scheduling. Use when serving ML models with KServe, configuring scale-to-zero, or deploying production inference endpoints on CoreWeave. Trigger with phrases like "coreweave inference service", "coreweave kserve", "coreweave model serving", "deploy model on coreweave".
.claude/skills/jeremylongshore-coreweave-core-workflow-a/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 28% | 0% |
> Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Deploy production inference services on CoreWeave using KServe InferenceService with GPU scheduling, autoscaling, and scale-to-zero. CKS natively integrates with KServe for serverless GPU inference.
coreweave-install-auth setupyaml# inference-service.yaml apiVersion: serving.kserve.io/v1beta1 kind: InferenceService metadata: name: llama-inference annotations: autoscaling.knative.dev/class: "kpa.autoscaling.knative.dev" autoscaling.knative.dev/metric: "concurrency" autoscaling.knative.dev/target: "1" autoscaling.knative.dev/minScale: "1" autoscaling.knative.dev/maxScale: "5" spec: predictor: minReplicas: 1 maxReplicas: 5 containers: - name: kserve-container image: vllm/vllm-openai:latest args: - "--model" - "meta-llama/Llama-3.1-8B-Instruct" - "--port" - "8080" ports: - containerPort: 8080 protocol: TCP resources: limits: nvidia.com/gpu: "1" memory: 48Gi cpu: "8" requests: nvidia.com/gpu: "1" memory: 32Gi cpu: "4" env: - name: HUGGING_FACE_HUB_TOKEN valueFrom: secretKeyRef: name: hf-token key: token affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: gpu.nvidia.com/class operator: In values: ["A100_PCIE_80GB"]
bashkubectl apply -f inference-service.yaml kubectl get inferenceservice llama-inference -w
yaml# For dev/staging -- scale down to zero when idle metadata: annotations: autoscaling.knative.dev/minScale: "0" # Scale to zero autoscaling.knative.dev/maxScale: "3" autoscaling.knative.dev/scaleDownDelay: "5m"
bash# Get inference URL INFERENCE_URL=$(kubectl get inferenceservice llama-inference \ -o jsonpath='{.status.url}') curl -X POST "${INFERENCE_URL}/v1/chat/completions" \ -H "Content-Type: application/json" \ -d '{"model": "meta-llama/Llama-3.1-8B-Instruct", "messages": [{"role": "user", "content": "Hello!"}]}'
| Error | Cause | Solution | |-------|-------|----------| | InferenceService not ready | GPU not available | Check node capacity and affinity | | Scale-to-zero cold start | First request after idle | Set minScale: 1 for production | | Model loading timeout | Large model download | Pre-cache model in PVC | | OOMKilled | Model too large | Use multi-GPU or quantized model |
For GPU training workloads, see coreweave-core-workflow-b.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,561 | 7,714 | -53% | 1 | 1 | 0% | 3,621 | 2,799 | -23% | 0 | 0 | — |
case-02 | fail→fail | 12,079 | 7,892 | -35% | 1 | 1 | 0% | 2,196 | 2,596 | +18% | 0 | 0 | — |
case-03 | fail→pass | 16,685 | 8,929 | -46% | 1 | 1 | 0% | 3,457 | 3,016 | -13% | 0 | 0 | — |
case-04 | fail→fail | 17,527 | 6,462 | -63% | 1 | 1 | 0% | 3,133 | 2,217 | -29% | 0 | 0 | — |
case-05 | fail→pass | 8,570 | 2,594 | -70% | 1 | 1 | 0% | 1,649 | 1,349 | -18% | 0 | 0 | — |
case-06 | pass→pass | 6,531 | 2,862 | -56% | 1 | 1 | 0% | 1,207 | 1,434 | +19% | 0 | 0 | — |
case-07 | pass→pass | 7,800 | 1,614 | -79% | 1 | 1 | 0% | 1,508 | 1,251 | -17% | 0 | 0 | — |
case-08 | fail→pass | 6,587 | 1,692 | -74% | 1 | 1 | 0% | 987 | 1,261 | +28% | 0 | 0 | — |
case-09 | pass→pass | 5,295 | 1,554 | -71% | 1 | 1 | 0% | 862 | 1,290 | +50% | 0 | 0 | — |
case-10 | pass→pass | 2,554 | 1,971 | -23% | 1 | 1 | 0% | 416 | 1,304 | +213% | 0 | 0 | — |
case-11 | fail→pass | 5,996 | 1,398 | -77% | 1 | 1 | 0% | 958 | 1,226 | +28% | 0 | 0 | — |
case-12 | fail→pass | 4,828 | 1,803 | -63% | 1 | 1 | 0% | 776 | 1,251 | +61% | 0 | 0 | — |
case-13 | fail→pass | 3,406 | 1,853 | -46% | 1 | 1 | 0% | 452 | 1,274 | +182% | 0 | 0 | — |
case-14 | pass→pass | 4,005 | 1,927 | -52% | 1 | 1 | 0% | 709 | 1,260 | +78% | 0 | 0 | — |
case-15 | pass→pass | 2,883 | 2,089 | -28% | 1 | 1 | 0% | 546 | 1,334 | +144% | 0 | 0 | — |
case-16 | pass→pass | 3,595 | 2,101 | -42% | 1 | 1 | 0% | 654 | 1,351 | +107% | 0 | 0 | — |
case-17 | pass→pass | 4,215 | 2,225 | -47% | 1 | 1 | 0% | 806 | 1,297 | +61% | 0 | 0 | — |
case-18 | fail→pass | 10,338 | 2,581 | -75% | 1 | 1 | 0% | 1,866 | 1,296 | -31% | 0 | 0 | — |
case-19 | fail→pass | 16,121 | 9,273 | -42% | 1 | 1 | 0% | 2,633 | 2,684 | +2% | 0 | 0 | — |
case-20 | pass→pass | 12,311 | 4,937 | -60% | 1 | 1 | 0% | 2,149 | 1,921 | -11% | 0 | 0 | — |
case-21 | pass→pass | 15,453 | 2,183 | -86% | 1 | 1 | 0% | 2,516 | 1,292 | -49% | 0 | 0 | — |
case-22 | pass→pass | 14,803 | 13,212 | -11% | 1 | 1 | 0% | 3,110 | 3,748 | +21% | 0 | 0 | — |
case-23 | pass→pass | 11,211 | 6,205 | -45% | 1 | 1 | 0% | 2,229 | 2,049 | -8% | 0 | 0 | — |
case-24 | pass→pass | 11,425 | 8,970 | -21% | 1 | 1 | 0% | 2,176 | 2,701 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +38 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.