Install any skill in seconds. Free to start, no credit card required.
Get Started Free →RNN+Transformer hybrid with O(n) inference. Linear time, infinite context, no KV cache. Train like GPT (parallel), infer like RNN (sequential). Linux Foundation AI project. Production at Windows, Office, NeMo. RWKV-7 (March 2025). Models up to 14B parameters.
.claude/skills/openlair-rwkv-architecture/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-15 | ✓→✓ | = Same ✓ | 37% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 80% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 120% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 63% | 0% |
RWKV (RwaKuv) combines Transformer parallelization (training) with RNN efficiency (inference).
Installation:
bash# Install PyTorch pip install torch --upgrade --extra-index-url https://download.pytorch.org/whl/cu121 # Install dependencies pip install pytorch-lightning==1.9.5 deepspeed wandb ninja --upgrade # Install RWKV pip install rwkv
Basic usage (GPT mode + RNN mode):
pythonimport os from rwkv.model import RWKV os.environ["RWKV_JIT_ON"] = '1' os.environ["RWKV_CUDA_ON"] = '1' # Use CUDA kernel for speed # Load model model = RWKV( model='/path/to/RWKV-4-Pile-1B5-20220903-8040', strategy='cuda fp16' ) # GPT mode (parallel processing) out, state = model.forward([187, 510, 1563, 310, 247], None) print(out.detach().cpu().numpy()) # Logits # RNN mode (sequential processing, same result) out, state = model.forward([187, 510], None) # First 2 tokens out, state = model.forward([1563], state) # Next token out, state = model.forward([310, 247], state) # Last tokens print(out.detach().cpu().numpy()) # Same logits as above!
Efficient token-by-token generation:
pythonfrom rwkv.model import RWKV from rwkv.utils import PIPELINE model = RWKV(model='RWKV-4-Pile-14B-20230313-ctx8192-test1050', strategy='cuda fp16') pipeline = PIPELINE(model, "20B_tokenizer.json") # Initial prompt prompt = "The future of AI is" state = None # Generate token by token for token in prompt: out, state = pipeline.model.forward(pipeline.encode(token), state) # Continue generation for _ in range(100): out, state = pipeline.model.forward(None, state) token = pipeline.sample_logits(out) print(pipeline.decode(token), end='', flush=True)
Key advantage: Constant memory per token (no growing KV cache)
Process million-token sequences:
pythonmodel = RWKV(model='RWKV-4-Pile-14B', strategy='cuda fp16') # Process very long document state = None long_document = load_document() # e.g., 1M tokens # Stream through entire document for chunk in chunks(long_document, chunk_size=1024): out, state = model.forward(chunk, state) # State now contains information from entire 1M token document # Memory usage: O(1) (constant, not O(n)!)
Standard fine-tuning workflow:
python# Training script import pytorch_lightning as pl from rwkv.model import RWKV from rwkv.trainer import RWKVTrainer # Configure model config = { 'n_layer': 24, 'n_embd': 1024, 'vocab_size': 50277, 'ctx_len': 1024 } # Setup trainer trainer = pl.Trainer( accelerator='gpu', devices=8, precision='bf16', strategy='deepspeed_stage_2', max_epochs=1 ) # Train model = RWKV(config) trainer.fit(model, train_dataloader)
Memory comparison (1M token sequence):
python# Transformer (GPT) # Memory: O(n²) for attention # KV cache: 1M × hidden_dim × n_layers × 2 (keys + values) # Example: 1M × 4096 × 24 × 2 = ~400GB (impractical!) # RWKV # Memory: O(1) per token # State: hidden_dim × n_layers = 4096 × 24 = ~400KB # 1,000,000× more efficient!
Speed comparison (inference):
python# Transformer: O(n) per token (quadratic overall) # First token: 1 computation # Second token: 2 computations # ... # 1000th token: 1000 computations # RWKV: O(1) per token (linear overall) # Every token: 1 computation # 1000th token: 1 computation (same as first!)
Use RWKV when:
Key advantages:
Use alternatives instead:
Issue: Out of memory during training
Use gradient checkpointing and DeepSpeed:
pythontrainer = pl.Trainer( strategy='deepspeed_stage_3', # Full ZeRO-3 precision='bf16' )
Issue: Slow inference
Enable CUDA kernel:
pythonos.environ["RWKV_CUDA_ON"] = '1'
Issue: Model not loading
Check model path and strategy:
pythonmodel = RWKV( model='/absolute/path/to/model.pth', strategy='cuda fp16' # Or 'cpu fp32' for CPU )
Issue: State management in RNN mode
Always pass state between forward calls:
python# WRONG: State lost out1, _ = model.forward(tokens1, None) out2, _ = model.forward(tokens2, None) # No context from tokens1! # CORRECT: State preserved out1, state = model.forward(tokens1, None) out2, state = model.forward(tokens2, state) # Has context from tokens1
Time-mixing and channel-mixing: See references/architecture-details.md for WKV operation, time-decay mechanism, and receptance gates.
State management: See references/state-management.md for att_x_prev, att_kv, ffn_x_prev states, and numerical stability considerations.
RWKV-7 improvements: See references/rwkv7.md for latest architectural improvements (March 2025) and multimodal capabilities.
Performance (vs Transformers):
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | pass→pass | 18,491 | 10,603 | -43% | 1 | 1 | 0% | 2,877 | 3,938 | +37% | 0 | 0 | — |
case-01 | fail→fail | 19,899 | 13,295 | -33% | 1 | 1 | 0% | 4,134 | 4,972 | +20% | 0 | 0 | — |
case-02 | pass→pass | 8,362 | 3,415 | -59% | 1 | 1 | 0% | 1,561 | 2,806 | +80% | 0 | 0 | — |
case-03 | pass→pass | 8,080 | 5,256 | -35% | 1 | 1 | 0% | 1,422 | 3,122 | +120% | 0 | 0 | — |
case-04 | pass→pass | 10,577 | 5,677 | -46% | 1 | 1 | 0% | 2,016 | 3,279 | +63% | 0 | 0 | — |
case-05 | pass→pass | 11,953 | 9,261 | -23% | 1 | 1 | 0% | 2,348 | 3,959 | +69% | 0 | 0 | — |
case-06 | pass→pass | 6,072 | 3,600 | -41% | 1 | 1 | 0% | 1,155 | 2,774 | +140% | 0 | 0 | — |
case-07 | pass→pass | 12,928 | 5,073 | -61% | 1 | 1 | 0% | 2,274 | 3,056 | +34% | 0 | 0 | — |
case-08 | pass→pass | 9,107 | 1,592 | -83% | 1 | 1 | 0% | 1,676 | 2,380 | +42% | 0 | 0 | — |
case-09 | pass→pass | 9,967 | 4,999 | -50% | 1 | 1 | 0% | 2,021 | 3,045 | +51% | 0 | 0 | — |
case-10 | pass→pass | 11,518 | 6,456 | -44% | 1 | 1 | 0% | 2,104 | 3,252 | +55% | 0 | 0 | — |
case-11 | pass→pass | 14,035 | 12,222 | -13% | 1 | 1 | 0% | 2,396 | 4,542 | +90% | 0 | 0 | — |
case-12 | pass→pass | 9,316 | 2,118 | -77% | 1 | 1 | 0% | 1,612 | 2,430 | +51% | 0 | 0 | — |
case-13 | pass→pass | 3,325 | 4,204 | +26% | 1 | 1 | 0% | 504 | 2,789 | +453% | 0 | 0 | — |
case-14 | pass→pass | 8,568 | 2,719 | -68% | 1 | 1 | 0% | 1,732 | 2,606 | +50% | 0 | 0 | — |
case-16 | pass→pass | 9,271 | 5,708 | -38% | 1 | 1 | 0% | 1,808 | 3,172 | +75% | 0 | 0 | — |
case-17 | fail→pass | 9,598 | 1,600 | -83% | 1 | 1 | 0% | 1,760 | 2,385 | +36% | 0 | 0 | — |
case-18 | pass→pass | 14,050 | 10,320 | -27% | 1 | 1 | 0% | 2,309 | 3,831 | +66% | 0 | 0 | — |
case-19 | pass→pass | 12,964 | 8,737 | -33% | 1 | 1 | 0% | 2,121 | 3,625 | +71% | 0 | 0 | — |
case-20 | pass→pass | 12,111 | 8,586 | -29% | 1 | 1 | 0% | 1,999 | 3,513 | +76% | 0 | 0 | — |
case-21 | pass→pass | 6,781 | 1,635 | -76% | 1 | 1 | 0% | 1,319 | 2,370 | +80% | 0 | 0 | — |
case-22 | pass→pass | 5,078 | 1,262 | -75% | 1 | 1 | 0% | 969 | 2,315 | +139% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.