Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.
.claude/skills/openlair-miles-rl-training/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 170% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 4% | 0% |
miles is a high-performance, enterprise-ready RL framework optimized for large-scale model post-training. Built as a production fork of slime, it addresses critical challenges in MoE training stability, low-precision training, and train-inference alignment.
Choose miles when you need:
Consider alternatives when:
bash# Recommended: Docker docker pull radixark/miles:latest docker run --rm --gpus all --ipc=host --shm-size=16g \ -it radixark/miles:latest /bin/bash # From source git clone https://github.com/radixark/miles.git cd miles pip install -r requirements.txt pip install -e .
miles inherits slime's configuration system. Basic training:
bashpython train.py \ --advantage-estimator grpo \ --model-name qwen3-30b-a3b \ --hf-checkpoint /path/to/qwen3-30b-a3b-hf \ --rollout-batch-size 512 \ --n-samples-per-prompt 8
Use this workflow for training large MoE models like DeepSeek V3 or Qwen3-MoE.
bash# FP8 block scaling (recommended for stability) export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1 export CUDA_DEVICE_MAX_CONNECTIONS=1
bashpython train.py \ --actor-num-gpus-per-node 8 \ --rollout-num-gpus 8 \ --hf-checkpoint /path/to/deepseek-v3 \ --advantage-estimator grpo \ --tensor-model-parallel-size 8 \ --expert-model-parallel-size 4 \ --prompt-data /path/to/data.jsonl \ --num-rollout 3000
Use this workflow for maximum rollout throughput with EAGLE speculative decoding.
miles supports EAGLE speculative decoding via SGLang:
bashpython train.py \ --actor-num-gpus-per-node 8 \ --hf-checkpoint /path/to/target-model \ --sglang-speculative-algorithm EAGLE \ --sglang-speculative-num-steps 3 \ --sglang-speculative-eagle-topk 1 \ --sglang-speculative-num-draft-tokens 4 \ --sglang-speculative-draft-model-path /path/to/draft-model \ --advantage-estimator grpo \ --prompt-data /path/to/data.jsonl
For online SFT of draft model during training:
bash--mtp-num-layers 1 \ --enable-mtp-training \ --mtp-loss-scaling-factor 0.2
Note: Online MTP training requires a torch dist checkpoint with MTP weights. Add --mtp-num-layers 1 during checkpoint conversion from HuggingFace.
miles inherits all slime arguments. See slime API Reference for the complete list.
bash--actor-num-nodes 1 --actor-num-gpus-per-node 8 --rollout-num-gpus 8 --rollout-num-gpus-per-engine 2 --colocate
bash--tensor-model-parallel-size 8 --pipeline-model-parallel-size 2 --expert-model-parallel-size 4 # MoE expert parallelism
bash--sglang-speculative-algorithm EAGLE --sglang-speculative-num-steps 3 --sglang-speculative-eagle-topk 1 --sglang-speculative-num-draft-tokens 4 --sglang-enable-draft-weights-cpu-backup --sglang-speculative-draft-model-path /your/draft/model/path
bash--mtp-num-layers 1 --enable-mtp-training --mtp-loss-scaling-factor 0.2
The following features are documented in miles but specific CLI flags may vary. Consult the miles repository for latest configuration.
End-to-end FP8 sampling and training that eliminates quantization-induced discrepancy causing RL collapse in MoE models.
Records expert routing decisions during SGLang inference and replays them during Megatron training for bit-wise expert alignment.
How R3 Works:
sample.rollout_routed_expertsEnables single-machine deployment of 1TB+ models (e.g., on H200).
Memory Savings with INT4:
| Model Size | BF16 VRAM | INT4 VRAM | Reduction | |------------|-----------|-----------|-----------| | 70B | 140GB | 45GB | 3.1x | | 235B | 470GB | 150GB | 3.1x | | 671B | 1.3TB | 420GB | 3.1x |
miles achieves "exactly 0 KL divergence" between training and inference through:
torch.compile integrationmiles uses the same Sample dataclass as slime with the rollout_routed_experts field for MoE routing replay:
python@dataclass class Sample: prompt: str | list[dict] tokens: list[int] response: str reward: float | dict loss_mask: list[int] status: Status metadata: dict rollout_log_probs: list[float] rollout_routed_experts: list[list[int]] # MoE routing for R3
See slime API Reference for the complete Sample definition.
Symptoms: Loss explodes, NaN values
Solutions:
export NVTE_FP8_BLOCK_SCALING_FP32_SCALES=1--lr 5e-7Symptoms: Low acceptance rate over time
Solutions:
--sglang-speculative-num-steps 2--sglang-enable-draft-weights-cpu-backupSymptoms: Policy divergence, reward collapse
Solutions:
--use-tis --tis-threshold 0.9| Family | Models | MoE Support | |--------|--------|-------------| | DeepSeek | R1, V3, V3.2 | Full | | Qwen | 2, 2.5, 3 (including MoE) | Full | | Llama | 3, 3.1, 3.3, 4 | Dense only | | Gemma | 2, 3, 3N | Dense only | | GLM | 4.5, 4.6, 4.7 | Dense only | | MiniMax | M2, M2.1 | Full |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,941 | 5,434 | -66% | 1 | 1 | 0% | 3,079 | 3,771 | +22% | 0 | 0 | — |
case-07 | pass→pass | 3,840 | 1,700 | -56% | 1 | 1 | 0% | 742 | 2,850 | +284% | 0 | 0 | — |
case-02 | fail→pass | 17,787 | 5,542 | -69% | 1 | 1 | 0% | 1,427 | 3,847 | +170% | 0 | 0 | — |
case-03 | fail→pass | 20,887 | 9,823 | -53% | 1 | 1 | 0% | 3,802 | 4,646 | +22% | 0 | 0 | — |
case-04 | pass→pass | 10,562 | 4,973 | -53% | 1 | 1 | 0% | 1,712 | 3,419 | +100% | 0 | 0 | — |
case-05 | fail→pass | 20,408 | 9,995 | -51% | 1 | 1 | 0% | 3,281 | 4,389 | +34% | 0 | 0 | — |
case-06 | pass→pass | 9,898 | 6,104 | -38% | 1 | 1 | 0% | 1,817 | 3,852 | +112% | 0 | 0 | — |
case-08 | pass→pass | 10,385 | 2,106 | -80% | 1 | 1 | 0% | 1,727 | 2,935 | +70% | 0 | 0 | — |
case-09 | fail→pass | 20,047 | 4,263 | -79% | 1 | 1 | 0% | 3,212 | 3,339 | +4% | 0 | 0 | — |
case-10 | fail→pass | 20,854 | 3,004 | -86% | 1 | 1 | 0% | 4,249 | 3,110 | -27% | 0 | 0 | — |
case-11 | fail→pass | 9,915 | 2,708 | -73% | 1 | 1 | 0% | 2,294 | 3,087 | +35% | 0 | 0 | — |
case-12 | fail→pass | 11,758 | 1,731 | -85% | 1 | 1 | 0% | 2,298 | 2,880 | +25% | 0 | 0 | — |
case-13 | pass→pass | 9,907 | 2,039 | -79% | 1 | 1 | 0% | 1,815 | 2,860 | +58% | 0 | 0 | — |
case-14 | fail→pass | 12,424 | 1,416 | -89% | 1 | 1 | 0% | 2,347 | 2,756 | +17% | 0 | 0 | — |
case-15 | fail→pass | 13,813 | 2,137 | -85% | 1 | 1 | 0% | 2,365 | 2,867 | +21% | 0 | 0 | — |
case-16 | fail→pass | 14,215 | 2,942 | -79% | 1 | 1 | 0% | 2,326 | 3,006 | +29% | 0 | 0 | — |
case-17 | pass→pass | 11,555 | 5,815 | -50% | 1 | 1 | 0% | 2,662 | 3,772 | +42% | 0 | 0 | — |
case-18 | fail→pass | 9,458 | 1,417 | -85% | 1 | 1 | 0% | 2,101 | 2,760 | +31% | 0 | 0 | — |
case-19 | fail→pass | 15,989 | 3,837 | -76% | 1 | 1 | 0% | 2,510 | 3,263 | +30% | 0 | 0 | — |
case-20 | pass→pass | 13,503 | 1,825 | -86% | 1 | 1 | 0% | 2,291 | 2,842 | +24% | 0 | 0 | — |
case-21 | fail→pass | 10,253 | 3,117 | -70% | 1 | 1 | 0% | 2,039 | 3,183 | +56% | 0 | 0 | — |
case-22 | fail→pass | 9,291 | 2,211 | -76% | 1 | 1 | 0% | 1,660 | 3,017 | +82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +68 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.