Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Multi-cloud orchestration for ML workloads with automatic cost optimization. Use when you need to run training or batch jobs across multiple clouds, leverage spot instances with auto-recovery, or optimize GPU costs across providers.
.claude/skills/openlair-skypilot-multi-cloud-orchestration/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 96% | 0% |
| case-12 | ✓→✓ | = Same ✓ | 178% | 0% |
| case-11 | ✓→✓ | = Same ✓ | 102% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 151% | 0% |
Comprehensive guide to running ML workloads across clouds with automatic cost optimization using SkyPilot.
Use SkyPilot when:
Key features:
Use alternatives instead:
bashpip install "skypilot[aws,gcp,azure,kubernetes]" # Verify cloud credentials sky check
Create hello.yaml:
yamlresources: accelerators: T4:1 run: | nvidia-smi echo "Hello from SkyPilot!"
Launch:
bashsky launch -c hello hello.yaml # SSH to cluster ssh hello # Terminate sky down hello
yaml# Task name (optional) name: my-task # Resource requirements resources: cloud: aws # Optional: auto-select if omitted region: us-west-2 # Optional: auto-select if omitted accelerators: A100:4 # GPU type and count cpus: 8+ # Minimum CPUs memory: 32+ # Minimum memory (GB) use_spot: true # Use spot instances disk_size: 256 # Disk size (GB) # Number of nodes for distributed training num_nodes: 2 # Working directory (synced to ~/sky_workdir) workdir: . # Setup commands (run once) setup: | pip install -r requirements.txt # Run commands run: | python train.py
| Command | Purpose | |---------|---------| | sky launch | Launch cluster and run task | | sky exec | Run task on existing cluster | | sky status | Show cluster status | | sky stop | Stop cluster (preserve state) | | sky down | Terminate cluster | | sky logs | View task logs | | sky queue | Show job queue | | sky jobs launch | Launch managed job | | sky serve up | Deploy serving endpoint |
yaml# NVIDIA GPUs accelerators: T4:1 accelerators: L4:1 accelerators: A10G:1 accelerators: L40S:1 accelerators: A100:4 accelerators: A100-80GB:8 accelerators: H100:8 # Cloud-specific accelerators: V100:4 # AWS/GCP accelerators: TPU-v4-8 # GCP TPUs
yamlresources: accelerators: H100: 8 A100-80GB: 8 A100: 8 any_of: - cloud: gcp - cloud: aws - cloud: azure
yamlresources: accelerators: A100:8 use_spot: true spot_recovery: FAILOVER # Auto-recover on preemption
bash# Launch new cluster sky launch -c mycluster task.yaml # Run on existing cluster (skip setup) sky exec mycluster another_task.yaml # Interactive SSH ssh mycluster # Stream logs sky logs mycluster
yamlresources: accelerators: A100:4 autostop: idle_minutes: 30 down: true # Terminate instead of stop
bash# Set autostop via CLI sky autostop mycluster -i 30 --down
bash# All clusters sky status # Detailed view sky status -a
yamlresources: accelerators: A100:8 num_nodes: 4 # 4 nodes × 8 GPUs = 32 GPUs total setup: | pip install torch torchvision run: | torchrun \ --nnodes=$SKYPILOT_NUM_NODES \ --nproc_per_node=$SKYPILOT_NUM_GPUS_PER_NODE \ --node_rank=$SKYPILOT_NODE_RANK \ --master_addr=$(echo "$SKYPILOT_NODE_IPS" | head -n1) \ --master_port=12355 \ train.py
| Variable | Description | |----------|-------------| | SKYPILOT_NODE_RANK | Node index (0 to num_nodes-1) | | SKYPILOT_NODE_IPS | Newline-separated IP addresses | | SKYPILOT_NUM_NODES | Total number of nodes | | SKYPILOT_NUM_GPUS_PER_NODE | GPUs per node |
bashrun: | if [ "${SKYPILOT_NODE_RANK}" == "0" ]; then python orchestrate.py fi
bash# Launch managed job with spot recovery sky jobs launch -n my-job train.yaml
yamlname: training-job file_mounts: /checkpoints: name: my-checkpoints store: s3 mode: MOUNT resources: accelerators: A100:8 use_spot: true run: | python train.py \ --checkpoint-dir /checkpoints \ --resume-from-latest
bash# List jobs sky jobs queue # View logs sky jobs logs my-job # Cancel job sky jobs cancel my-job
yamlworkdir: ./my-project # Synced to ~/sky_workdir file_mounts: /data/config.yaml: ./config.yaml ~/.vimrc: ~/.vimrc
yamlfile_mounts: # Mount S3 bucket /datasets: source: s3://my-bucket/datasets mode: MOUNT # Stream from S3 # Copy GCS bucket /models: source: gs://my-bucket/models mode: COPY # Pre-fetch to disk # Cached mount (fast writes) /outputs: name: my-outputs store: s3 mode: MOUNT_CACHED
| Mode | Description | Best For | |------|-------------|----------| | MOUNT | Stream from cloud | Large datasets, read-heavy | | COPY | Pre-fetch to disk | Small files, random access | | MOUNT_CACHED | Cache with async upload | Checkpoints, outputs |
yaml# service.yaml service: readiness_probe: /health replica_policy: min_replicas: 1 max_replicas: 10 target_qps_per_replica: 2.0 resources: accelerators: A100:1 run: | python -m vllm.entrypoints.openai.api_server \ --model meta-llama/Llama-2-7b-chat-hf \ --port 8000
bash# Deploy sky serve up -n my-service service.yaml # Check status sky serve status # Get endpoint sky serve status my-service
yamlservice: replica_policy: min_replicas: 1 max_replicas: 10 target_qps_per_replica: 2.0 upscale_delay_seconds: 60 downscale_delay_seconds: 300 load_balancing_policy: round_robin
yaml# SkyPilot finds cheapest option resources: accelerators: A100:8 # No cloud specified - auto-select cheapest
bash# Show optimizer decision sky launch task.yaml --dryrun
yamlresources: accelerators: A100:8 any_of: - cloud: gcp region: us-central1 - cloud: aws region: us-east-1 - cloud: azure
yamlenvs: HF_TOKEN: $HF_TOKEN # Inherited from local env WANDB_API_KEY: $WANDB_API_KEY # Or use secrets secrets: - HF_TOKEN - WANDB_API_KEY
yamlname: llm-finetune file_mounts: /checkpoints: name: finetune-checkpoints store: s3 mode: MOUNT_CACHED resources: accelerators: A100:8 use_spot: true setup: | pip install transformers accelerate run: | python train.py \ --checkpoint-dir /checkpoints \ --resume
yamlname: hp-sweep-${RUN_ID} envs: RUN_ID: 0 LEARNING_RATE: 1e-4 BATCH_SIZE: 32 resources: accelerators: A100:1 use_spot: true run: | python train.py \ --lr $LEARNING_RATE \ --batch-size $BATCH_SIZE \ --run-id $RUN_ID
bash# Launch multiple jobs for i in {1..10}; do sky jobs launch sweep.yaml \ --env RUN_ID=$i \ --env LEARNING_RATE=$(python -c "import random; print(10**random.uniform(-5,-3))") done
bash# SSH to cluster ssh mycluster # View logs sky logs mycluster # Check job queue sky queue mycluster # View managed job logs sky jobs logs my-job
| Issue | Solution | |-------|----------| | Quota exceeded | Request quota increase, try different region | | Spot preemption | Use sky jobs launch for auto-recovery | | Slow file sync | Use MOUNT_CACHED mode for outputs | | GPU not available | Use any_of for fallback clouds |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 7,420 | 4,609 | -38% | 1 | 1 | 0% | 1,323 | 3,673 | +178% | 0 | 0 | — |
case-01 | fail→pass | 12,166 | 7,821 | -36% | 1 | 1 | 0% | 2,743 | 4,702 | +71% | 0 | 0 | — |
case-11 | pass→pass | 10,612 | 6,523 | -39% | 1 | 1 | 0% | 2,119 | 4,289 | +102% | 0 | 0 | — |
case-02 | fail→pass | 9,530 | 5,461 | -43% | 1 | 1 | 0% | 2,086 | 4,093 | +96% | 0 | 0 | — |
case-03 | pass→pass | 10,781 | 10,385 | -4% | 1 | 1 | 0% | 1,900 | 4,761 | +151% | 0 | 0 | — |
case-04 | pass→pass | 8,314 | 5,408 | -35% | 1 | 1 | 0% | 1,413 | 3,878 | +174% | 0 | 0 | — |
case-05 | pass→pass | 8,793 | 8,360 | -5% | 1 | 1 | 0% | 1,768 | 4,441 | +151% | 0 | 0 | — |
case-06 | pass→pass | 9,604 | 4,224 | -56% | 1 | 1 | 0% | 1,962 | 3,739 | +91% | 0 | 0 | — |
case-07 | pass→pass | 6,529 | 4,368 | -33% | 1 | 1 | 0% | 1,350 | 3,703 | +174% | 0 | 0 | — |
case-08 | pass→pass | 3,871 | 3,315 | -14% | 1 | 1 | 0% | 718 | 3,575 | +398% | 0 | 0 | — |
case-09 | pass→pass | 10,633 | 7,366 | -31% | 1 | 1 | 0% | 2,329 | 4,628 | +99% | 0 | 0 | — |
case-10 | pass→pass | 8,620 | 7,284 | -15% | 1 | 1 | 0% | 1,719 | 4,341 | +153% | 0 | 0 | — |
case-13 | pass→pass | 9,884 | 8,689 | -12% | 1 | 1 | 0% | 1,947 | 4,549 | +134% | 0 | 0 | — |
case-14 | pass→pass | 12,041 | 4,164 | -65% | 1 | 1 | 0% | 2,176 | 3,779 | +74% | 0 | 0 | — |
case-15 | pass→pass | 7,914 | 4,564 | -42% | 1 | 1 | 0% | 1,539 | 4,013 | +161% | 0 | 0 | — |
case-16 | pass→pass | 6,952 | 2,471 | -64% | 1 | 1 | 0% | 1,209 | 3,345 | +177% | 0 | 0 | — |
case-17 | pass→pass | 7,500 | 5,200 | -31% | 1 | 1 | 0% | 1,456 | 3,912 | +169% | 0 | 0 | — |
case-18 | pass→pass | 4,376 | 3,882 | -11% | 1 | 1 | 0% | 783 | 3,499 | +347% | 0 | 0 | — |
case-19 | pass→pass | 2,741 | 3,270 | +19% | 1 | 1 | 0% | 313 | 3,203 | +923% | 0 | 0 | — |
case-20 | pass→pass | 6,745 | 2,563 | -62% | 1 | 1 | 0% | 1,260 | 3,361 | +167% | 0 | 0 | — |
case-21 | pass→pass | 6,293 | 4,811 | -24% | 1 | 1 | 0% | 1,159 | 3,713 | +220% | 0 | 0 | — |
case-22 | pass→pass | 8,602 | 4,716 | -45% | 1 | 1 | 0% | 1,735 | 3,758 | +117% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.