Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL algorithms for LLM post-training at scale with flexible infrastructure backends.
.claude/skills/openlair-verl-rl-training/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 253% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 39% | 0% |
verl is a flexible, efficient, and production-ready RL training library for large language models from ByteDance's Seed team. It implements the HybridFlow framework (EuroSys 2025) and powers models like Doubao-1.5-pro achieving O1-level performance on math benchmarks.
Choose verl when you need:
Consider alternatives when:
bash# Option 1: pip install pip install verl[vllm] # or verl[sglang] for SGLang backend # Option 2: Docker (recommended for production) docker pull verlai/verl:vllm011.latest # Option 3: From source git clone https://github.com/volcengine/verl.git cd verl && pip install -e .[vllm,math]
bashpython3 -m verl.trainer.main_ppo \ algorithm.adv_estimator=grpo \ data.train_files=~/data/gsm8k/train.parquet \ actor_rollout_ref.model.path=Qwen/Qwen2.5-7B \ actor_rollout_ref.rollout.n=8 \ actor_rollout_ref.actor.use_kl_loss=True \ trainer.n_gpus_per_node=8
verl uses a HybridFlow programming model separating control flow from computation:
┌─────────────────────────────────────────────────────────┐
│ Single-Process Controller (Ray) │
│ - Orchestrates: rollout → reward → train → sync │
└─────────────────────┬───────────────────────────────────┘
│
┌─────────────────────▼───────────────────────────────────┐
│ Multi-Process Workers │
│ ├── ActorRolloutRefWorker (policy + generation) │
│ ├── CriticWorker (value estimation, PPO only) │
│ └── RewardManager (model-based or rule-based rewards) │
└─────────────────────────────────────────────────────────┘Use this workflow for training reasoning models on math tasks like GSM8K or MATH.
prompt and reward_model columnspythonimport pandas as pd data = [ { "prompt": [{"role": "user", "content": "What is 15 + 27?"}], "reward_model": {"ground_truth": "42"} }, # ... more examples ] df = pd.DataFrame(data) df.to_parquet("train.parquet")
python# reward_function.py import re def compute_reward(responses, ground_truths): rewards = [] for response, gt in zip(responses, ground_truths): # Extract answer from response match = re.search(r'\\boxed{([^}]+)}', response) if match and match.group(1).strip() == gt.strip(): rewards.append(1.0) else: rewards.append(0.0) return rewards
yaml# config/grpo_math.yaml algorithm: adv_estimator: grpo gamma: 1.0 lam: 1.0 data: train_files: /path/to/train.parquet val_files: /path/to/val.parquet train_batch_size: 256 max_prompt_length: 512 max_response_length: 2048 actor_rollout_ref: model: path: Qwen/Qwen2.5-7B-Instruct actor: use_kl_loss: true kl_loss_coef: 0.001 ppo_mini_batch_size: 64 rollout: name: vllm n: 8 # samples per prompt temperature: 0.7 top_p: 0.95 trainer: total_epochs: 3 n_gpus_per_node: 8 save_freq: 100
bashpython3 -m verl.trainer.main_ppo \ --config-path config \ --config-name grpo_math \ trainer.experiment_name=grpo_math_qwen7b
Use this workflow when you need value-based advantage estimation (GAE).
yamlalgorithm: adv_estimator: gae # Use GAE instead of GRPO gamma: 0.99 lam: 0.95 critic: model: path: Qwen/Qwen2.5-7B-Instruct # Can be same or different from actor ppo_mini_batch_size: 64 actor_rollout_ref: actor: use_kl_loss: true kl_loss_coef: 0.02 clip_ratio: 0.2 # PPO clipping
bashpython3 -m verl.trainer.main_ppo \ algorithm.adv_estimator=gae \ critic.model.path=Qwen/Qwen2.5-7B-Instruct \ trainer.n_gpus_per_node=8
Use this workflow for models >70B parameters or when you need expert parallelism.
pip install mbridgeyamlactor_rollout_ref: model: path: /path/to/megatron/checkpoint backend: megatron actor: strategy: megatron tensor_model_parallel_size: 8 pipeline_model_parallel_size: 2 rollout: name: vllm tensor_parallel_size: 8
bash# On head node ray start --head --port=6379 # On worker nodes ray start --address='head_ip:6379' # Launch training python3 -m verl.trainer.main_ppo \ trainer.nnodes=4 \ trainer.n_gpus_per_node=8
| Algorithm | adv_estimator | Use Case | |-----------|-----------------|----------| | GRPO | grpo | Critic-free, math/reasoning | | PPO/GAE | gae | Dense rewards, value estimation | | REINFORCE++ | reinforce_plus_plus | Variance reduction | | RLOO | rloo | Leave-one-out baseline | | ReMax | remax | Maximum reward baseline | | OPO | opo | Optimal policy optimization |
yaml# Rollout parameters actor_rollout_ref.rollout.n: 8 # Samples per prompt actor_rollout_ref.rollout.temperature: 0.7 # Sampling temperature actor_rollout_ref.rollout.top_p: 0.95 # Nucleus sampling # Training parameters actor_rollout_ref.actor.lr: 1e-6 # Learning rate actor_rollout_ref.actor.ppo_mini_batch_size: 64 actor_rollout_ref.actor.clip_ratio: 0.2 # PPO clip range # KL control actor_rollout_ref.actor.use_kl_loss: true actor_rollout_ref.actor.kl_loss_coef: 0.001 algorithm.kl_ctrl.target_kl: 0.1 # For adaptive KL control
Symptoms: CUDA out of memory during generation phase
Solutions:
yaml# Reduce batch size actor_rollout_ref.rollout.log_prob_micro_batch_size: 4 # Enable gradient checkpointing actor_rollout_ref.model.enable_gradient_checkpointing: true # Use FSDP2 with CPU offloading actor_rollout_ref.actor.strategy: fsdp2 actor_rollout_ref.actor.fsdp_config.offload_policy: true
Symptoms: Loss spikes, reward collapse
Solutions:
yaml# Reduce learning rate actor_rollout_ref.actor.lr: 5e-7 # Increase KL penalty actor_rollout_ref.actor.kl_loss_coef: 0.01 # Enable gradient clipping actor_rollout_ref.actor.max_grad_norm: 1.0
Symptoms: Long pauses between rollout and training
Solutions:
bash# Use FSDP2 for faster resharding actor_rollout_ref.actor.strategy=fsdp2 # Enable async weight transfer trainer.async_weight_update=true
Symptoms: Import errors or generation failures
Solution: Use compatible versions:
bashpip install vllm>=0.8.5,<=0.12.0 # Avoid vLLM 0.7.x (known bugs)
See references/multi-turn.md for agentic workflows with tool use.
yamlactor_rollout_ref: model: path: Qwen/Qwen2.5-VL-7B-Instruct rollout: name: vllm enable_vision: true
yamlactor_rollout_ref: actor: lora: enabled: true r: 16 alpha: 32 target_modules: ["q_proj", "v_proj"]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,901 | 8,916 | -47% | 1 | 1 | 0% | 3,234 | 5,017 | +55% | 0 | 0 | — |
case-02 | fail→fail | 17,385 | 14,386 | -17% | 1 | 1 | 0% | 3,175 | 5,661 | +78% | 0 | 0 | — |
case-03 | fail→fail | 20,081 | 11,074 | -45% | 1 | 1 | 0% | 3,771 | 5,502 | +46% | 0 | 0 | — |
case-04 | pass→pass | 15,585 | 9,688 | -38% | 1 | 1 | 0% | 2,806 | 4,897 | +75% | 0 | 0 | — |
case-05 | fail→pass | 11,982 | 3,333 | -72% | 1 | 1 | 0% | 2,104 | 3,582 | +70% | 0 | 0 | — |
case-06 | fail→pass | 18,316 | 5,602 | -69% | 1 | 1 | 0% | 3,054 | 3,881 | +27% | 0 | 0 | — |
case-07 | pass→pass | 7,950 | 2,639 | -67% | 1 | 1 | 0% | 1,609 | 3,556 | +121% | 0 | 0 | — |
case-08 | fail→pass | 5,796 | 2,705 | -53% | 1 | 1 | 0% | 990 | 3,496 | +253% | 0 | 0 | — |
case-09 | fail→pass | 13,519 | 2,226 | -84% | 1 | 1 | 0% | 2,401 | 3,340 | +39% | 0 | 0 | — |
case-10 | fail→fail | 13,744 | 9,859 | -28% | 1 | 1 | 0% | 2,776 | 5,095 | +84% | 0 | 0 | — |
case-11 | fail→pass | 11,675 | 2,399 | -79% | 1 | 1 | 0% | 2,002 | 3,380 | +69% | 0 | 0 | — |
case-12 | fail→fail | 11,859 | 10,050 | -15% | 1 | 1 | 0% | 2,007 | 4,981 | +148% | 0 | 0 | — |
case-13 | fail→pass | 15,214 | 3,700 | -76% | 1 | 1 | 0% | 2,431 | 3,675 | +51% | 0 | 0 | — |
case-14 | pass→pass | 5,201 | 2,534 | -51% | 1 | 1 | 0% | 910 | 3,458 | +280% | 0 | 0 | — |
case-15 | pass→pass | 5,699 | 1,747 | -69% | 1 | 1 | 0% | 982 | 3,216 | +227% | 0 | 0 | — |
case-16 | fail→pass | 7,375 | 2,768 | -62% | 1 | 1 | 0% | 1,207 | 3,479 | +188% | 0 | 0 | — |
case-17 | fail→pass | 11,192 | 1,775 | -84% | 1 | 1 | 0% | 2,002 | 3,246 | +62% | 0 | 0 | — |
case-18 | fail→pass | 10,594 | 1,996 | -81% | 1 | 1 | 0% | 1,784 | 3,264 | +83% | 0 | 0 | — |
case-19 | fail→pass | 14,177 | 7,533 | -47% | 1 | 1 | 0% | 2,647 | 4,423 | +67% | 0 | 0 | — |
case-20 | fail→fail | 17,551 | 11,276 | -36% | 1 | 1 | 0% | 3,289 | 5,128 | +56% | 0 | 0 | — |
case-21 | pass→pass | 6,263 | 1,770 | -72% | 1 | 1 | 0% | 1,197 | 3,246 | +171% | 0 | 0 | — |
case-22 | pass→pass | 7,097 | 2,766 | -61% | 1 | 1 | 0% | 1,457 | 3,507 | +141% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.