Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Educational GPT implementation in ~300 lines. Reproduces GPT-2 (124M) on OpenWebText. Clean, hackable code for learning transformers. By Andrej Karpathy. Perfect for understanding GPT architecture from scratch. Train on Shakespeare (CPU) or OpenWebText (multi-GPU).
.claude/skills/openlair-nanogpt/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -43% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 44% | 0% |
nanoGPT is a simplified GPT implementation designed for learning and experimentation.
Installation:
bashpip install torch numpy transformers datasets tiktoken wandb tqdm
Train on Shakespeare (CPU-friendly):
bash# Prepare data python data/shakespeare_char/prepare.py # Train (5 minutes on CPU) python train.py config/train_shakespeare_char.py # Generate text python sample.py --out_dir=out-shakespeare-char
Output:
ROMEO:
What say'st thou? Shall I speak, and be a man?
JULIET:
I am afeard, and yet I'll speak; for thou art
One that hath been a man, and yet I know not
What thou art.Complete training pipeline:
bash# Step 1: Prepare data (creates train.bin, val.bin) python data/shakespeare_char/prepare.py # Step 2: Train small model python train.py config/train_shakespeare_char.py # Step 3: Generate text python sample.py --out_dir=out-shakespeare-char
Config (config/train_shakespeare_char.py):
python# Model config n_layer = 6 # 6 transformer layers n_head = 6 # 6 attention heads n_embd = 384 # 384-dim embeddings block_size = 256 # 256 char context # Training config batch_size = 64 learning_rate = 1e-3 max_iters = 5000 eval_interval = 500 # Hardware device = 'cpu' # Or 'cuda' compile = False # Set True for PyTorch 2.0
Training time: ~5 minutes (CPU), ~1 minute (GPU)
Multi-GPU training on OpenWebText:
bash# Step 1: Prepare OpenWebText (takes ~1 hour) python data/openwebtext/prepare.py # Step 2: Train GPT-2 124M with DDP (8 GPUs) torchrun --standalone --nproc_per_node=8 \ train.py config/train_gpt2.py # Step 3: Sample from trained model python sample.py --out_dir=out
Config (config/train_gpt2.py):
python# GPT-2 (124M) architecture n_layer = 12 n_head = 12 n_embd = 768 block_size = 1024 dropout = 0.0 # Training batch_size = 12 gradient_accumulation_steps = 5 * 8 # Total batch ~0.5M tokens learning_rate = 6e-4 max_iters = 600000 lr_decay_iters = 600000 # System compile = True # PyTorch 2.0
Training time: ~4 days (8× A100)
Start from OpenAI checkpoint:
python# In train.py or config init_from = 'gpt2' # Options: gpt2, gpt2-medium, gpt2-large, gpt2-xl # Model loads OpenAI weights automatically python train.py config/finetune_shakespeare.py
Example config (config/finetune_shakespeare.py):
python# Start from GPT-2 init_from = 'gpt2' # Dataset dataset = 'shakespeare_char' batch_size = 1 block_size = 1024 # Fine-tuning learning_rate = 3e-5 # Lower LR for fine-tuning max_iters = 2000 warmup_iters = 100 # Regularization weight_decay = 1e-1
Train on your own text:
python# data/custom/prepare.py import numpy as np # Load your data with open('my_data.txt', 'r') as f: text = f.read() # Create character mappings chars = sorted(list(set(text))) stoi = {ch: i for i, ch in enumerate(chars)} itos = {i: ch for i, ch in enumerate(chars)} # Tokenize data = np.array([stoi[ch] for ch in text], dtype=np.uint16) # Split train/val n = len(data) train_data = data[:int(n*0.9)] val_data = data[int(n*0.9):] # Save train_data.tofile('data/custom/train.bin') val_data.tofile('data/custom/val.bin')
Train:
bashpython data/custom/prepare.py python train.py --dataset=custom
Use nanoGPT when:
Simplicity advantages:
model.pytrain.pyUse alternatives instead:
Issue: CUDA out of memory
Reduce batch size or context length:
pythonbatch_size = 1 # Reduce from 12 block_size = 512 # Reduce from 1024 gradient_accumulation_steps = 40 # Increase to maintain effective batch
Issue: Training too slow
Enable compilation (PyTorch 2.0+):
pythoncompile = True # 2× speedup
Use mixed precision:
pythondtype = 'bfloat16' # Or 'float16'
Issue: Poor generation quality
Train longer:
pythonmax_iters = 10000 # Increase from 5000
Lower temperature:
python# In sample.py temperature = 0.7 # Lower from 1.0 top_k = 200 # Add top-k sampling
Issue: Can't load GPT-2 weights
Install transformers:
bashpip install transformers
Check model name:
pythoninit_from = 'gpt2' # Valid: gpt2, gpt2-medium, gpt2-large, gpt2-xl
Model architecture: See references/architecture.md for GPT block structure, multi-head attention, and MLP layers explained simply.
Training loop: See references/training.md for learning rate schedule, gradient accumulation, and distributed data parallel setup.
Data preparation: See references/data.md for tokenization strategies (character-level vs BPE) and binary format details.
Performance:
compile=True: 2× speedupdtype=bfloat16: 50% memory reduction| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 13,009 | 7,990 | -39% | 1 | 1 | 0% | 2,508 | 3,601 | +44% | 0 | 0 | — |
case-02 | pass→pass | 10,725 | 6,271 | -42% | 1 | 1 | 0% | 2,013 | 3,235 | +61% | 0 | 0 | — |
case-03 | pass→pass | 12,262 | 6,528 | -47% | 1 | 1 | 0% | 2,038 | 3,153 | +55% | 0 | 0 | — |
case-04 | pass→pass | 5,367 | 3,279 | -39% | 1 | 1 | 0% | 1,126 | 2,429 | +116% | 0 | 0 | — |
case-05 | pass→pass | 4,376 | 1,807 | -59% | 1 | 1 | 0% | 785 | 2,375 | +203% | 0 | 0 | — |
case-06 | pass→pass | 5,637 | 3,130 | -44% | 1 | 1 | 0% | 1,060 | 2,614 | +147% | 0 | 0 | — |
case-07 | fail→fail | 11,895 | 7,528 | -37% | 1 | 1 | 0% | 2,468 | 3,639 | +47% | 0 | 0 | — |
case-08 | pass→pass | 4,252 | 2,413 | -43% | 1 | 1 | 0% | 800 | 2,622 | +228% | 0 | 0 | — |
case-09 | pass→pass | 7,784 | 5,599 | -28% | 1 | 1 | 0% | 1,430 | 3,074 | +115% | 0 | 0 | — |
case-10 | pass→pass | 10,680 | 8,034 | -25% | 1 | 1 | 0% | 2,124 | 3,582 | +69% | 0 | 0 | — |
case-11 | pass→pass | 6,835 | 2,517 | -63% | 1 | 1 | 0% | 1,250 | 2,479 | +98% | 0 | 0 | — |
case-12 | fail→pass | 6,947 | 3,026 | -56% | 1 | 1 | 0% | 1,301 | 2,592 | +99% | 0 | 0 | — |
case-13 | fail→fail | 6,051 | 4,434 | -27% | 1 | 1 | 0% | 1,128 | 2,928 | +160% | 0 | 0 | — |
case-14 | pass→pass | 8,160 | 2,864 | -65% | 1 | 1 | 0% | 1,423 | 2,553 | +79% | 0 | 0 | — |
case-15 | pass→pass | 8,483 | 2,183 | -74% | 1 | 1 | 0% | 529 | 2,410 | +356% | 0 | 0 | — |
case-16 | pass→pass | 5,303 | 2,984 | -44% | 1 | 1 | 0% | 1,027 | 2,611 | +154% | 0 | 0 | — |
case-17 | fail→fail | 10,849 | 5,251 | -52% | 1 | 1 | 0% | 1,923 | 3,073 | +60% | 0 | 0 | — |
case-18 | pass→pass | 3,973 | 1,686 | -58% | 1 | 1 | 0% | 630 | 2,324 | +269% | 0 | 0 | — |
case-19 | fail→pass | 12,892 | 2,490 | -81% | 1 | 1 | 0% | 2,394 | 2,363 | -1% | 0 | 0 | — |
case-20 | fail→pass | 12,485 | 1,710 | -86% | 1 | 1 | 0% | 2,406 | 2,318 | -4% | 0 | 0 | — |
case-21 | pass→pass | 8,626 | 2,964 | -66% | 1 | 1 | 0% | 1,475 | 2,533 | +72% | 0 | 0 | — |
case-22 | fail→pass | 20,559 | 1,994 | -90% | 1 | 1 | 0% | 4,232 | 2,415 | -43% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.