Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Serverless GPU cloud for ML jobs and model APIs.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 360% | 0% |
Comprehensive guide to running ML workloads on Modal's serverless GPU cloud platform.
Use Modal when:
Key features:
Use alternatives instead:
bashpip install modal modal setup # Opens browser for authentication
pythonimport modal app = modal.App("hello-gpu") @app.function(gpu="T4") def gpu_info(): import subprocess return subprocess.run(["nvidia-smi"], capture_output=True, text=True).stdout @app.local_entrypoint() def main(): print(gpu_info.remote())
Run: modal run hello_gpu.py
pythonimport modal app = modal.App("text-generation") image = modal.Image.debian_slim().pip_install("transformers", "torch", "accelerate") @app.cls(gpu="A10G", image=image) class TextGenerator: @modal.enter() def load_model(self): from transformers import pipeline self.pipe = pipeline("text-generation", model="gpt2", device=0) @modal.method() def generate(self, prompt: str) -> str: return self.pipe(prompt, max_length=100)[0]["generated_text"] @app.local_entrypoint() def main(): print(TextGenerator().generate.remote("Hello, world"))
| Component | Purpose | |-----------|---------| | App | Container for functions and resources | | Function | Serverless function with compute specs | | Cls | Class-based functions with lifecycle hooks | | Image | Container image definition | | Volume | Persistent storage for models/data | | Secret | Secure credential storage |
| Command | Description | |---------|-------------| | modal run script.py | Execute and exit | | modal serve script.py | Development with live reload | | modal deploy script.py | Persistent cloud deployment |
| GPU | VRAM | Best For | |-----|------|----------| | T4 | 16GB | Budget inference, small models | | L4 | 24GB | Inference, Ada Lovelace arch | | A10G | 24GB | Training/inference, 3.3x faster than T4 | | L40S | 48GB | Recommended for inference (best cost/perf) | | A100-40GB | 40GB | Large model training | | A100-80GB | 80GB | Very large models | | H100 | 80GB | Fastest, FP8 + Transformer Engine | | H200 | 141GB | Auto-upgrade from H100, 4.8TB/s bandwidth | | B200 | Latest | Blackwell architecture |
python# Single GPU @app.function(gpu="A100") # Specific memory variant @app.function(gpu="A100-80GB") # Multiple GPUs (up to 8) @app.function(gpu="H100:4") # GPU with fallbacks @app.function(gpu=["H100", "A100", "L40S"]) # Any available GPU @app.function(gpu="any")
python# Basic image with pip image = modal.Image.debian_slim(python_version="3.11").pip_install( "torch==2.1.0", "transformers==4.36.0", "accelerate" ) # From CUDA base image = modal.Image.from_registry( "nvidia/cuda:12.1.0-cudnn8-devel-ubuntu22.04", add_python="3.11" ).pip_install("torch", "transformers") # With system packages image = modal.Image.debian_slim().apt_install("git", "ffmpeg").pip_install("whisper")
pythonvolume = modal.Volume.from_name("model-cache", create_if_missing=True) @app.function(gpu="A10G", volumes={"/models": volume}) def load_model(): import os model_path = "/models/llama-7b" if not os.path.exists(model_path): model = download_model() model.save_pretrained(model_path) volume.commit() # Persist changes return load_from_path(model_path)
python@app.function() @modal.fastapi_endpoint(method="POST") def predict(text: str) -> dict: return {"result": model.predict(text)}
pythonfrom fastapi import FastAPI web_app = FastAPI() @web_app.post("/predict") async def predict(text: str): return {"result": await model.predict.remote.aio(text)} @app.function() @modal.asgi_app() def fastapi_app(): return web_app
| Decorator | Use Case | |-----------|----------| | @modal.fastapi_endpoint() | Simple function → API | | @modal.asgi_app() | Full FastAPI/Starlette apps | | @modal.wsgi_app() | Django/Flask apps | | @modal.web_server(port) | Arbitrary HTTP servers |
python@app.function() @modal.batched(max_batch_size=32, wait_ms=100) async def batch_predict(inputs: list[str]) -> list[dict]: # Inputs automatically batched return model.batch_predict(inputs)
bash# Create secret modal secret create huggingface HF_TOKEN=hf_xxx
python@app.function(secrets=[modal.Secret.from_name("huggingface")]) def download_model(): import os token = os.environ["HF_TOKEN"]
python@app.function(schedule=modal.Cron("0 0 * * *")) # Daily midnight def daily_job(): pass @app.function(schedule=modal.Period(hours=1)) def hourly_job(): pass
python# Modal 1.0 autoscaler params: scaledown_window (was container_idle_timeout). # Input concurrency moved to the @modal.concurrent decorator. @app.function(scaledown_window=300) # Keep warm 5 min @modal.concurrent(max_inputs=10) # Handle concurrent requests per container def inference(): pass
python@app.cls(gpu="A100") class Model: @modal.enter() # Run once at container start def load(self): self.model = load_model() # Load during warm-up @modal.method() def predict(self, x): return self.model(x)
python@app.function() def process_item(item): return expensive_computation(item) @app.function() def run_parallel(): items = list(range(1000)) # Fan out to parallel containers results = list(process_item.map(items)) return results
python@app.function( gpu="A100", memory=32768, # 32GB RAM cpu=4, # 4 CPU cores timeout=3600, # 1 hour max scaledown_window=120, # Keep warm 2 min (was container_idle_timeout) retries=3, # Retry on failure max_containers=10, # Max concurrent containers (was concurrency_limit) min_containers=1, # Keep N containers warm (was keep_warm) ) def my_function(): pass
> Modal 1.0 autoscaler renames (see the migration guide): > - container_idle_timeout → scaledown_window > - concurrency_limit → max_containers > - keep_warm → min_containers > - allow_concurrent_inputs=N → the @modal.concurrent(max_inputs=N) decorator
python# Test locally if __name__ == "__main__": result = my_function.local() # View logs # modal app logs my-app
| Issue | Solution | |-------|----------| | Cold start latency | Increase scaledown_window, use @modal.enter() | | GPU OOM | Use larger GPU (A100-80GB), enable gradient checkpointing | | Image build fails | Pin dependency versions, check CUDA compatibility | | Timeout errors | Increase timeout, add checkpointing |
Other measured skills in the registry, with their headline benchmark lift.