Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Provides guidance for LLM post-training with RL using slime, a Megatron+SGLang framework. Use when training GLM models, implementing custom data generation workflows, or needing tight Megatron-LM integration for RL scaling.
.claude/skills/openlair-slime-rl-training/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 122% | 0% |
slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation.
Choose slime when you need:
Consider alternatives when:
┌─────────────────────────────────────────────────────────┐
│ Data Buffer │
│ - Prompt initialization and management │
│ - Custom data generation and filtering │
│ - Rollout sample storage │
└─────────────┬───────────────────────────┬───────────────┘
│ │
┌─────────────▼───────────┐ ┌─────────────▼───────────────┐
│ Training (Megatron-LM) │ │ Rollout (SGLang + Router) │
│ - Actor model training │ │ - Response generation │
│ - Critic (optional) │ │ - Reward/verifier output │
│ - Weight sync to rollout│ │ - Multi-turn support │
└─────────────────────────┘ └─────────────────────────────┘bash# Recommended: Docker docker pull slimerl/slime:latest docker run --rm --gpus all --ipc=host --shm-size=16g \ -it slimerl/slime:latest /bin/bash # Inside container cd /root/slime && pip install -e . --no-deps
bashgit clone https://github.com/THUDM/slime.git cd slime pip install -r requirements.txt pip install -e .
bash# Source model configuration source scripts/models/qwen3-4B.sh # Launch training python train.py \ --actor-num-nodes 1 \ --actor-num-gpus-per-node 4 \ --rollout-num-gpus 4 \ --advantage-estimator grpo \ --use-kl-loss --kl-loss-coef 0.001 \ --rollout-batch-size 32 \ --n-samples-per-prompt 8 \ --global-batch-size 256 \ --num-rollout 3000 \ --prompt-data /path/to/data.jsonl \ ${MODEL_ARGS[@]} ${CKPT_ARGS[@]}
Use this workflow for training reasoning models with group-relative advantages.
python# data.jsonl format {"prompt": "What is 2 + 2?", "label": "4"} {"prompt": "Solve: 3x = 12", "label": "x = 4"}
Or with chat format:
python{ "prompt": [ {"role": "system", "content": "You are a math tutor."}, {"role": "user", "content": "What is 15 + 27?"} ], "label": "42" }
Choose a pre-configured model script:
bash# List available models ls scripts/models/ # glm4-9B.sh, qwen3-4B.sh, qwen3-30B-A3B.sh, deepseek-v3.sh, llama3-8B.sh, ... # Source your model source scripts/models/qwen3-4B.sh
bashpython train.py \ --actor-num-nodes 1 \ --actor-num-gpus-per-node 8 \ --rollout-num-gpus 8 \ --advantage-estimator grpo \ --use-kl-loss \ --kl-loss-coef 0.001 \ --prompt-data /path/to/train.jsonl \ --input-key prompt \ --label-key label \ --apply-chat-template \ --rollout-batch-size 32 \ --n-samples-per-prompt 8 \ --global-batch-size 256 \ --num-rollout 3000 \ --save-interval 100 \ --eval-interval 50 \ ${MODEL_ARGS[@]}
tensorboard --logdir outputs/Use async mode for higher throughput by overlapping rollout and training.
bashpython train_async.py \ --actor-num-nodes 1 \ --actor-num-gpus-per-node 8 \ --rollout-num-gpus 8 \ --advantage-estimator grpo \ --async-buffer-size 4 \ --prompt-data /path/to/train.jsonl \ ${MODEL_ARGS[@]}
bash--async-buffer-size 4 # Number of rollouts to buffer --update-weights-interval 2 # Sync weights every N rollouts
Use this workflow for training agents with tool use or multi-step reasoning.
python# custom_generate.py async def custom_generate(args, samples, evaluation=False): """Multi-turn generation with tool calling.""" for sample in samples: conversation = sample.prompt for turn in range(args.max_turns): # Generate response response = await generate_single(conversation) # Check for tool call tool_call = extract_tool_call(response) if tool_call: tool_result = execute_tool(tool_call) conversation.append({"role": "assistant", "content": response}) conversation.append({"role": "tool", "content": tool_result}) else: break sample.response = response sample.reward = compute_reward(sample) return samples
bashpython train.py \ --custom-generate-function-path custom_generate.py \ --max-turns 5 \ --prompt-data /path/to/agent_data.jsonl \ ${MODEL_ARGS[@]}
See examples/search-r1/ for a complete multi-turn search example.
slime uses three types of arguments:
1. Megatron Arguments (passed directly):
bash--tensor-model-parallel-size 2 --pipeline-model-parallel-size 1 --num-layers 32 --hidden-size 4096
2. SGLang Arguments (prefixed with --sglang-):
bash--sglang-mem-fraction-static 0.8 --sglang-context-length 8192 --sglang-log-level INFO
3. slime Arguments:
bash# Resource allocation --actor-num-nodes 1 --actor-num-gpus-per-node 8 --rollout-num-gpus 8 --colocate # Share GPUs between training/inference # Data --prompt-data /path/to/data.jsonl --input-key prompt --label-key label # Training loop --num-rollout 3000 --rollout-batch-size 32 --n-samples-per-prompt 8 --global-batch-size 256 # Algorithm --advantage-estimator grpo # or: gspo, ppo, reinforce_plus_plus --use-kl-loss --kl-loss-coef 0.001
rollout_batch_size × n_samples_per_prompt = global_batch_size × num_steps_per_rolloutExample: 32 × 8 = 256 × 1
slime's data buffer enables flexible data management:
pythonclass RolloutDataSource: def get_samples(self, num_samples): """Fetch prompts from dataset.""" return self.dataset.sample(num_samples) def add_samples(self, samples): """Called after generation (no-op by default).""" pass
pythonclass RolloutDataSourceWithBuffer(RolloutDataSource): def __init__(self): self.buffer = [] def add_samples(self, samples): """Store generated samples for reuse.""" self.buffer.extend(samples) def buffer_filter(self, args, buffer, num_samples): """Custom selection logic (prioritized, stratified, etc.).""" return select_best(buffer, num_samples)
Symptoms: Inference engine dies mid-training
Solutions:
bash# Enable fault tolerance --use-fault-tolerance # Increase memory allocation --sglang-mem-fraction-static 0.85 # Reduce batch size --rollout-batch-size 16
Symptoms: Training hangs after rollout
Solutions:
bash# Increase sync interval --update-weights-interval 5 # Use colocated mode (no network transfer) --colocate
Symptoms: CUDA OOM in backward pass
Solutions:
bash# Enable gradient checkpointing --recompute-activations # Reduce micro-batch size --micro-batch-size 1 # Enable sequence parallelism --sequence-parallel
Symptoms: GPU idle during data fetch
Solutions:
bash# Increase data workers --num-data-workers 4 # Use streaming dataset --streaming-data
| Model Family | Configurations | |--------------|----------------| | GLM | GLM-4.5, GLM-4.6, GLM-4.7, GLM-Z1-9B | | Qwen | Qwen3 (4B, 8B, 30B-A3B), Qwen3-MoE, Qwen2.5 | | DeepSeek | V3, V3.1, R1 | | Llama | Llama 3 (8B, 70B) | | Others | Kimi K2, Moonlight-16B |
Each model has pre-configured scripts in scripts/models/.
Share GPUs between training and inference to reduce memory:
bashpython train.py \ --colocate \ --actor-num-gpus-per-node 8 \ --sglang-mem-fraction-static 0.4 \ ${MODEL_ARGS[@]}
python# custom_rm.py class CustomRewardModel: def __init__(self, model_path): self.model = load_model(model_path) def compute_reward(self, prompts, responses): inputs = self.tokenize(prompts, responses) scores = self.model(inputs) return scores.tolist()
bash--custom-rm-path custom_rm.py
bash--eval-prompt-data aime /path/to/aime.jsonl \ --eval-prompt-data gsm8k /path/to/gsm8k.jsonl \ --n-samples-per-eval-prompt 16
examples/ directory for 14+ worked examples| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | pass→pass | 13,704 | 8,045 | -41% | 1 | 1 | 0% | 2,563 | 4,596 | +79% | 0 | 0 | — |
case-06 | fail→pass | 14,441 | 7,008 | -51% | 1 | 1 | 0% | 2,570 | 4,469 | +74% | 0 | 0 | — |
case-01 | fail→fail | 23,499 | 14,563 | -38% | 1 | 1 | 0% | 4,762 | 6,116 | +28% | 0 | 0 | — |
case-02 | fail→pass | 17,094 | 14,242 | -17% | 1 | 1 | 0% | 3,527 | 6,499 | +84% | 0 | 0 | — |
case-03 | fail→pass | 17,241 | 12,572 | -27% | 1 | 1 | 0% | 3,244 | 5,570 | +72% | 0 | 0 | — |
case-04 | fail→pass | 14,468 | 4,222 | -71% | 1 | 1 | 0% | 2,549 | 3,947 | +55% | 0 | 0 | — |
case-07 | fail→pass | 9,172 | 4,080 | -56% | 1 | 1 | 0% | 1,859 | 4,121 | +122% | 0 | 0 | — |
case-08 | fail→pass | 6,556 | 2,307 | -65% | 1 | 1 | 0% | 1,295 | 3,561 | +175% | 0 | 0 | — |
case-09 | fail→pass | 17,464 | 4,276 | -76% | 1 | 1 | 0% | 1,649 | 4,123 | +150% | 0 | 0 | — |
case-10 | pass→pass | 9,114 | 3,024 | -67% | 1 | 1 | 0% | 1,896 | 3,830 | +102% | 0 | 0 | — |
case-11 | pass→pass | 9,015 | 5,474 | -39% | 1 | 1 | 0% | 1,981 | 4,428 | +124% | 0 | 0 | — |
case-12 | fail→pass | 16,892 | 3,060 | -82% | 1 | 1 | 0% | 3,224 | 3,797 | +18% | 0 | 0 | — |
case-13 | fail→pass | 9,599 | 1,750 | -82% | 1 | 1 | 0% | 1,715 | 3,504 | +104% | 0 | 0 | — |
case-14 | fail→pass | 10,964 | 3,345 | -69% | 1 | 1 | 0% | 2,052 | 3,854 | +88% | 0 | 0 | — |
case-15 | fail→pass | 7,368 | 1,882 | -74% | 1 | 1 | 0% | 1,281 | 3,572 | +179% | 0 | 0 | — |
case-16 | fail→pass | 20,286 | 1,908 | -91% | 1 | 1 | 0% | 2,454 | 3,564 | +45% | 0 | 0 | — |
case-17 | pass→pass | 2,362 | 3,522 | +49% | 1 | 1 | 0% | 503 | 3,904 | +676% | 0 | 0 | — |
case-18 | fail→pass | 9,829 | 2,725 | -72% | 1 | 1 | 0% | 1,766 | 3,718 | +111% | 0 | 0 | — |
case-19 | fail→pass | 4,159 | 2,123 | -49% | 1 | 1 | 0% | 747 | 3,621 | +385% | 0 | 0 | — |
case-20 | fail→pass | 11,132 | 2,891 | -74% | 1 | 1 | 0% | 2,005 | 3,745 | +87% | 0 | 0 | — |
case-21 | fail→pass | 7,994 | 2,299 | -71% | 1 | 1 | 0% | 1,342 | 3,609 | +169% | 0 | 0 | — |
case-22 | fail→pass | 12,618 | 2,611 | -79% | 1 | 1 | 0% | 2,264 | 3,617 | +60% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +77 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.