Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
.claude/skills/openlair-optimizing-attention-flash/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 122% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 383% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 54% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 152% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 153% | 0% |
Flash Attention provides 2-4x speedup and 10-20x memory reduction for transformer attention through IO-aware tiling and recomputation.
PyTorch native (easiest, PyTorch 2.2+):
pythonimport torch import torch.nn.functional as F q = torch.randn(2, 8, 512, 64, device='cuda', dtype=torch.float16) # [batch, heads, seq, dim] k = torch.randn(2, 8, 512, 64, device='cuda', dtype=torch.float16) v = torch.randn(2, 8, 512, 64, device='cuda', dtype=torch.float16) # Automatically uses Flash Attention if available out = F.scaled_dot_product_attention(q, k, v)
flash-attn library (more features):
bashpip install flash-attn --no-build-isolation
pythonfrom flash_attn import flash_attn_func # q, k, v: [batch, seqlen, nheads, headdim] out = flash_attn_func(q, k, v, dropout_p=0.0, causal=True)
Copy this checklist:
Flash Attention Integration:
- [ ] Step 1: Check PyTorch version (≥2.2)
- [ ] Step 2: Enable Flash Attention backend
- [ ] Step 3: Verify speedup with profiling
- [ ] Step 4: Test accuracy matches baselineStep 1: Check PyTorch version
bashpython -c "import torch; print(torch.__version__)" # Should be ≥2.2.0
If <2.2, upgrade:
bashpip install --upgrade torch
Step 2: Enable Flash Attention backend
Replace standard attention:
python# Before (standard attention) attn_weights = torch.softmax(q @ k.transpose(-2, -1) / math.sqrt(d_k), dim=-1) out = attn_weights @ v # After (Flash Attention) import torch.nn.functional as F out = F.scaled_dot_product_attention(q, k, v, attn_mask=mask)
Force Flash Attention backend:
pythonwith torch.backends.cuda.sdp_kernel( enable_flash=True, enable_math=False, enable_mem_efficient=False ): out = F.scaled_dot_product_attention(q, k, v)
Step 3: Verify speedup with profiling
pythonimport torch.utils.benchmark as benchmark def test_attention(use_flash): q, k, v = [torch.randn(2, 8, 2048, 64, device='cuda', dtype=torch.float16) for _ in range(3)] if use_flash: with torch.backends.cuda.sdp_kernel(enable_flash=True): return F.scaled_dot_product_attention(q, k, v) else: attn = (q @ k.transpose(-2, -1) / 8.0).softmax(dim=-1) return attn @ v # Benchmark t_flash = benchmark.Timer(stmt='test_attention(True)', globals=globals()) t_standard = benchmark.Timer(stmt='test_attention(False)', globals=globals()) print(f"Flash: {t_flash.timeit(100).mean:.3f}s") print(f"Standard: {t_standard.timeit(100).mean:.3f}s")
Expected: 2-4x speedup for sequences >512 tokens.
Step 4: Test accuracy matches baseline
python# Compare outputs q, k, v = [torch.randn(1, 8, 512, 64, device='cuda', dtype=torch.float16) for _ in range(3)] # Flash Attention out_flash = F.scaled_dot_product_attention(q, k, v) # Standard attention attn_weights = torch.softmax(q @ k.transpose(-2, -1) / 8.0, dim=-1) out_standard = attn_weights @ v # Check difference diff = (out_flash - out_standard).abs().max() print(f"Max difference: {diff:.6f}") # Should be <1e-3 for float16
For multi-query attention, sliding window, or H100 FP8.
Copy this checklist:
flash-attn Library Setup:
- [ ] Step 1: Install flash-attn library
- [ ] Step 2: Modify attention code
- [ ] Step 3: Enable advanced features
- [ ] Step 4: Benchmark performanceStep 1: Install flash-attn library
bash# NVIDIA GPUs (CUDA 12.0+) pip install flash-attn --no-build-isolation # Verify installation python -c "from flash_attn import flash_attn_func; print('Success')"
Step 2: Modify attention code
pythonfrom flash_attn import flash_attn_func # Input: [batch_size, seq_len, num_heads, head_dim] # Transpose from [batch, heads, seq, dim] if needed q = q.transpose(1, 2) # [batch, seq, heads, dim] k = k.transpose(1, 2) v = v.transpose(1, 2) out = flash_attn_func( q, k, v, dropout_p=0.1, causal=True, # For autoregressive models window_size=(-1, -1), # No sliding window softmax_scale=None # Auto-scale ) out = out.transpose(1, 2) # Back to [batch, heads, seq, dim]
Step 3: Enable advanced features
Multi-query attention (shared K/V across heads):
pythonfrom flash_attn import flash_attn_func # q: [batch, seq, num_q_heads, dim] # k, v: [batch, seq, num_kv_heads, dim] # Fewer KV heads out = flash_attn_func(q, k, v) # Automatically handles MQA
Sliding window attention (local attention):
python# Only attend to window of 256 tokens before/after out = flash_attn_func( q, k, v, window_size=(256, 256), # (left, right) window causal=True )
Step 4: Benchmark performance
pythonimport torch from flash_attn import flash_attn_func import time q, k, v = [torch.randn(4, 4096, 32, 64, device='cuda', dtype=torch.float16) for _ in range(3)] # Warmup for _ in range(10): _ = flash_attn_func(q, k, v) # Benchmark torch.cuda.synchronize() start = time.time() for _ in range(100): out = flash_attn_func(q, k, v) torch.cuda.synchronize() end = time.time() print(f"Time per iteration: {(end-start)/100*1000:.2f}ms") print(f"Memory allocated: {torch.cuda.max_memory_allocated()/1e9:.2f}GB")
For maximum performance on H100 GPUs.
FP8 Setup:
- [ ] Step 1: Verify H100 GPU available
- [ ] Step 2: Install flash-attn with FP8 support
- [ ] Step 3: Convert inputs to FP8
- [ ] Step 4: Run with FP8 attentionStep 1: Verify H100 GPU
bashnvidia-smi --query-gpu=name --format=csv # Should show "H100" or "H800"
Step 2: Install flash-attn with FP8 support
bashpip install flash-attn --no-build-isolation # FP8 support included for H100
Step 3: Convert inputs to FP8
pythonimport torch q = torch.randn(2, 4096, 32, 64, device='cuda', dtype=torch.float16) k = torch.randn(2, 4096, 32, 64, device='cuda', dtype=torch.float16) v = torch.randn(2, 4096, 32, 64, device='cuda', dtype=torch.float16) # Convert to float8_e4m3 (FP8) q_fp8 = q.to(torch.float8_e4m3fn) k_fp8 = k.to(torch.float8_e4m3fn) v_fp8 = v.to(torch.float8_e4m3fn)
Step 4: Run with FP8 attention
pythonfrom flash_attn import flash_attn_func # FlashAttention-3 automatically uses FP8 kernels on H100 out = flash_attn_func(q_fp8, k_fp8, v_fp8) # Result: ~1.2 PFLOPS, 1.5-2x faster than FP16
Use Flash Attention when:
Use alternatives instead:
Issue: ImportError: cannot import flash_attn
Install with no-build-isolation flag:
bashpip install flash-attn --no-build-isolation
Or install CUDA toolkit first:
bashconda install cuda -c nvidia pip install flash-attn --no-build-isolation
Issue: Slower than expected (no speedup)
Flash Attention benefits increase with sequence length:
Check sequence length is sufficient.
Issue: RuntimeError: CUDA error
Verify GPU supports Flash Attention:
pythonimport torch print(torch.cuda.get_device_capability()) # Should be ≥(7, 5) for Turing+
Flash Attention requires:
Issue: Accuracy degradation
Check dtype is float16 or bfloat16 (not float32):
pythonq = q.to(torch.float16) # Or torch.bfloat16
Flash Attention uses float16/bfloat16 for speed. Float32 not supported.
Integration with HuggingFace Transformers: See references/transformers-integration.md for enabling Flash Attention in BERT, GPT, Llama models.
Performance benchmarks: See references/benchmarks.md for detailed speed and memory comparisons across GPUs and sequence lengths.
Algorithm details: See references/algorithm.md for tiling strategy, recomputation, and IO complexity analysis.
Advanced features: See references/advanced-features.md for rotary embeddings, ALiBi, paged KV cache, and custom attention masks.
Not supported: V100 (Volta), CPU inference
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 16,586 | 13,280 | -20% | 1 | 1 | 0% | 3,978 | 6,114 | +54% | 0 | 0 | — |
case-02 | pass→pass | 9,801 | 7,317 | -25% | 1 | 1 | 0% | 1,911 | 4,817 | +152% | 0 | 0 | — |
case-03 | pass→pass | 9,190 | 7,185 | -22% | 1 | 1 | 0% | 1,922 | 4,868 | +153% | 0 | 0 | — |
case-04 | pass→pass | 9,636 | 11,792 | +22% | 1 | 1 | 0% | 2,064 | 4,625 | +124% | 0 | 0 | — |
case-05 | pass→pass | 8,109 | 8,230 | +1% | 1 | 1 | 0% | 1,533 | 4,952 | +223% | 0 | 0 | — |
case-06 | pass→pass | 7,657 | 6,471 | -15% | 1 | 1 | 0% | 1,463 | 4,459 | +205% | 0 | 0 | — |
case-07 | pass→pass | 8,553 | 5,888 | -31% | 1 | 1 | 0% | 1,642 | 4,355 | +165% | 0 | 0 | — |
case-08 | pass→pass | 12,871 | 9,034 | -30% | 1 | 1 | 0% | 2,392 | 5,005 | +109% | 0 | 0 | — |
case-09 | fail→fail | 11,439 | 15,905 | +39% | 1 | 1 | 0% | 2,149 | 6,241 | +190% | 0 | 0 | — |
case-10 | fail→fail | 15,298 | 10,454 | -32% | 1 | 1 | 0% | 2,613 | 5,251 | +101% | 0 | 0 | — |
case-11 | pass→pass | 6,742 | 2,173 | -68% | 1 | 1 | 0% | 545 | 3,583 | +557% | 0 | 0 | — |
case-12 | pass→pass | 12,919 | 7,309 | -43% | 1 | 1 | 0% | 2,242 | 4,530 | +102% | 0 | 0 | — |
case-13 | pass→pass | 2,751 | 2,546 | -7% | 1 | 1 | 0% | 509 | 3,675 | +622% | 0 | 0 | — |
case-14 | fail→fail | 5,491 | 5,637 | +3% | 1 | 1 | 0% | 1,033 | 4,219 | +308% | 0 | 0 | — |
case-15 | fail→pass | 9,808 | 5,249 | -46% | 1 | 1 | 0% | 1,867 | 4,148 | +122% | 0 | 0 | — |
case-16 | pass→pass | 7,536 | 6,409 | -15% | 1 | 1 | 0% | 1,433 | 4,518 | +215% | 0 | 0 | — |
case-17 | pass→pass | 11,908 | 7,980 | -33% | 1 | 1 | 0% | 2,303 | 4,746 | +106% | 0 | 0 | — |
case-18 | pass→pass | 14,620 | 10,028 | -31% | 1 | 1 | 0% | 2,878 | 5,107 | +77% | 0 | 0 | — |
case-19 | pass→pass | 15,088 | 8,398 | -44% | 1 | 1 | 0% | 2,592 | 4,724 | +82% | 0 | 0 | — |
case-20 | pass→pass | 14,526 | 9,086 | -37% | 1 | 1 | 0% | 2,936 | 5,112 | +74% | 0 | 0 | — |
case-21 | pass→pass | 7,370 | 4,143 | -44% | 1 | 1 | 0% | 902 | 3,979 | +341% | 0 | 0 | — |
case-22 | fail→pass | 4,247 | 3,227 | -24% | 1 | 1 | 0% | 769 | 3,716 | +383% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.