Install any skill in seconds. Free to start, no credit card required.
Get Started Free →High-level PyTorch framework with Trainer class, automatic distributed training (DDP/FSDP/DeepSpeed), callbacks system, and minimal boilerplate. Scales from laptop to supercomputer with same code. Use when you want clean training loops with built-in best practices.
.claude/skills/openlair-pytorch-lightning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 125% | 0% |
| case-12 | ✓→✗ | ▼ Worse | 79% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 58% | 0% |
PyTorch Lightning organizes PyTorch code to eliminate boilerplate while maintaining flexibility.
Installation:
bashpip install lightning
Convert PyTorch to Lightning (3 steps):
pythonimport lightning as L import torch from torch import nn from torch.utils.data import DataLoader, Dataset # Step 1: Define LightningModule (organize your PyTorch code) class LitModel(L.LightningModule): def __init__(self, hidden_size=128): super().__init__() self.model = nn.Sequential( nn.Linear(28 * 28, hidden_size), nn.ReLU(), nn.Linear(hidden_size, 10) ) def training_step(self, batch, batch_idx): x, y = batch y_hat = self.model(x) loss = nn.functional.cross_entropy(y_hat, y) self.log('train_loss', loss) # Auto-logged to TensorBoard return loss def configure_optimizers(self): return torch.optim.Adam(self.parameters(), lr=1e-3) # Step 2: Create data train_loader = DataLoader(train_dataset, batch_size=32) # Step 3: Train with Trainer (handles everything else!) trainer = L.Trainer(max_epochs=10, accelerator='gpu', devices=2) model = LitModel() trainer.fit(model, train_loader)
That's it! Trainer handles:
Original PyTorch code:
pythonmodel = MyModel() optimizer = torch.optim.Adam(model.parameters()) model.to('cuda') for epoch in range(max_epochs): for batch in train_loader: batch = batch.to('cuda') optimizer.zero_grad() loss = model(batch) loss.backward() optimizer.step()
Lightning version:
pythonclass LitModel(L.LightningModule): def __init__(self): super().__init__() self.model = MyModel() def training_step(self, batch, batch_idx): loss = self.model(batch) # No .to('cuda') needed! return loss def configure_optimizers(self): return torch.optim.Adam(self.parameters()) # Train trainer = L.Trainer(max_epochs=10, accelerator='gpu') trainer.fit(LitModel(), train_loader)
Benefits: 40+ lines → 15 lines, no device management, automatic distributed
pythonclass LitModel(L.LightningModule): def __init__(self): super().__init__() self.model = MyModel() def training_step(self, batch, batch_idx): x, y = batch y_hat = self.model(x) loss = nn.functional.cross_entropy(y_hat, y) self.log('train_loss', loss) return loss def validation_step(self, batch, batch_idx): x, y = batch y_hat = self.model(x) val_loss = nn.functional.cross_entropy(y_hat, y) acc = (y_hat.argmax(dim=1) == y).float().mean() self.log('val_loss', val_loss) self.log('val_acc', acc) def test_step(self, batch, batch_idx): x, y = batch y_hat = self.model(x) test_loss = nn.functional.cross_entropy(y_hat, y) self.log('test_loss', test_loss) def configure_optimizers(self): return torch.optim.Adam(self.parameters(), lr=1e-3) # Train with validation trainer = L.Trainer(max_epochs=10) trainer.fit(model, train_loader, val_loader) # Test trainer.test(model, test_loader)
Automatic features:
python# Same code as single GPU! model = LitModel() # 8 GPUs with DDP (automatic!) trainer = L.Trainer( accelerator='gpu', devices=8, strategy='ddp' # Or 'fsdp', 'deepspeed' ) trainer.fit(model, train_loader)
Launch:
bash# Single command, Lightning handles the rest python train.py
No changes needed:
num_nodes=2)pythonfrom lightning.pytorch.callbacks import ModelCheckpoint, EarlyStopping, LearningRateMonitor # Create callbacks checkpoint = ModelCheckpoint( monitor='val_loss', mode='min', save_top_k=3, filename='model-{epoch:02d}-{val_loss:.2f}' ) early_stop = EarlyStopping( monitor='val_loss', patience=5, mode='min' ) lr_monitor = LearningRateMonitor(logging_interval='epoch') # Add to Trainer trainer = L.Trainer( max_epochs=100, callbacks=[checkpoint, early_stop, lr_monitor] ) trainer.fit(model, train_loader, val_loader)
Result:
pythonclass LitModel(L.LightningModule): # ... (training_step, etc.) def configure_optimizers(self): optimizer = torch.optim.Adam(self.parameters(), lr=1e-3) # Cosine annealing scheduler = torch.optim.lr_scheduler.CosineAnnealingLR( optimizer, T_max=100, eta_min=1e-5 ) return { 'optimizer': optimizer, 'lr_scheduler': { 'scheduler': scheduler, 'interval': 'epoch', # Update per epoch 'frequency': 1 } } # Learning rate auto-logged! trainer = L.Trainer(max_epochs=100) trainer.fit(model, train_loader)
Use PyTorch Lightning when:
Key advantages:
Use alternatives instead:
Issue: Loss not decreasing
Check data and model setup:
python# Add to training_step def training_step(self, batch, batch_idx): if batch_idx == 0: print(f"Batch shape: {batch[0].shape}") print(f"Labels: {batch[1]}") loss = ... return loss
Issue: Out of memory
Reduce batch size or use gradient accumulation:
pythontrainer = L.Trainer( accumulate_grad_batches=4, # Effective batch = batch_size × 4 precision='bf16' # Or 'fp16', reduces memory 50% )
Issue: Validation not running
Ensure you pass val_loader:
python# WRONG trainer.fit(model, train_loader) # CORRECT trainer.fit(model, train_loader, val_loader)
Issue: DDP spawns multiple processes unexpectedly
Lightning auto-detects GPUs. Explicitly set devices:
python# Test on CPU first trainer = L.Trainer(accelerator='cpu', devices=1) # Then GPU trainer = L.Trainer(accelerator='gpu', devices=1)
Callbacks: See references/callbacks.md for EarlyStopping, ModelCheckpoint, custom callbacks, and callback hooks.
Distributed strategies: See references/distributed.md for DDP, FSDP, DeepSpeed ZeRO integration, multi-node setup.
Hyperparameter tuning: See references/hyperparameter-tuning.md for integration with Optuna, Ray Tune, and WandB sweeps.
Precision options:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,855 | 7,970 | -33% | 1 | 1 | 0% | 2,404 | 4,132 | +72% | 0 | 0 | — |
case-02 | fail→pass | 14,278 | 9,567 | -33% | 1 | 1 | 0% | 2,938 | 4,459 | +52% | 0 | 0 | — |
case-03 | pass→pass | 11,539 | 7,365 | -36% | 1 | 1 | 0% | 2,547 | 4,015 | +58% | 0 | 0 | — |
case-04 | pass→pass | 8,884 | 7,666 | -14% | 1 | 1 | 0% | 1,956 | 4,105 | +110% | 0 | 0 | — |
case-05 | pass→pass | 9,988 | 5,126 | -49% | 1 | 1 | 0% | 1,977 | 3,453 | +75% | 0 | 0 | — |
case-06 | pass→pass | 8,272 | 3,728 | -55% | 1 | 1 | 0% | 1,532 | 3,260 | +113% | 0 | 0 | — |
case-07 | pass→pass | 9,448 | 6,211 | -34% | 1 | 1 | 0% | 2,032 | 3,865 | +90% | 0 | 0 | — |
case-08 | pass→pass | 6,208 | 3,795 | -39% | 1 | 1 | 0% | 1,165 | 3,274 | +181% | 0 | 0 | — |
case-09 | pass→pass | 5,052 | 5,962 | +18% | 1 | 1 | 0% | 1,095 | 3,647 | +233% | 0 | 0 | — |
case-10 | pass→pass | 11,750 | 9,836 | -16% | 1 | 1 | 0% | 2,282 | 4,416 | +94% | 0 | 0 | — |
case-11 | pass→pass | 7,773 | 5,045 | -35% | 1 | 1 | 0% | 1,455 | 3,447 | +137% | 0 | 0 | — |
case-12 | pass→fail | 9,443 | 5,182 | -45% | 1 | 1 | 0% | 1,898 | 3,388 | +79% | 0 | 0 | — |
case-13 | pass→pass | 12,270 | 8,935 | -27% | 1 | 1 | 0% | 2,610 | 4,177 | +60% | 0 | 0 | — |
case-14 | pass→pass | 7,909 | 4,608 | -42% | 1 | 1 | 0% | 1,650 | 3,299 | +100% | 0 | 0 | — |
case-15 | pass→pass | 10,313 | 8,002 | -22% | 1 | 1 | 0% | 1,908 | 4,091 | +114% | 0 | 0 | — |
case-16 | pass→pass | 13,535 | 11,970 | -12% | 1 | 1 | 0% | 2,492 | 4,904 | +97% | 0 | 0 | — |
case-17 | pass→pass | 16,970 | 10,950 | -35% | 1 | 1 | 0% | 3,707 | 4,973 | +34% | 0 | 0 | — |
case-18 | pass→pass | 10,760 | 6,278 | -42% | 1 | 1 | 0% | 2,298 | 3,723 | +62% | 0 | 0 | — |
case-19 | pass→pass | 12,196 | 7,108 | -42% | 1 | 1 | 0% | 2,277 | 3,992 | +75% | 0 | 0 | — |
case-20 | fail→pass | 8,357 | 7,029 | -16% | 1 | 1 | 0% | 1,825 | 4,101 | +125% | 0 | 0 | — |
case-21 | pass→pass | 7,737 | 5,337 | -31% | 1 | 1 | 0% | 1,685 | 3,772 | +124% | 0 | 0 | — |
case-22 | pass→pass | 9,408 | 9,328 | -1% | 1 | 1 | 0% | 2,219 | 4,670 | +110% | 0 | 0 | — |
case-23 | pass→pass | 3,680 | 2,390 | -35% | 1 | 1 | 0% | 716 | 2,980 | +316% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.