Install any skill in seconds. Free to start, no credit card required.
Get Started Free →MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-01 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
For the higher-level overview, see:
Use MoE communication overlap when:
EP > 1Avoid turning it on as an early bring-up step. It is easier to validate after the dispatcher, routing mode, and recompute plan are already stable.
pythoncfg.comm_overlap.overlap_moe_expert_parallel_comm = True # Optional: delayed wgrad for additional overlap cfg.comm_overlap.delay_wgrad_compute = True # IMPORTANT: disable shared expert overlap when using dispatch overlap cfg.model.moe_shared_expert_overlap = False
expert_model_parallel_size > 1num_moe_experts > 1moe_token_dispatcher_type must be "alltoall" or "flex"virtual_pipeline_model_parallel_size) must be set (non-None)Setting moe_flex_dispatcher_backend alone does not activate flex dispatch. You must also set moe_token_dispatcher_type = "flex".
delay_wgrad_compute adds further constraints if CUDA-graph scopes includeattention or MoE-router work.
A 2026-05-18 current-main H100 x16 smoke on Qwen3 30B-A3B mock pretraining used EP=16, alltoall, global batch size 1024, CUDA graphs disabled, and moe_permute_fusion=false because the PyTorch 25.11 / TE / Triton stack failed in Transformer Engine fused permutation in prior bring-up.
Results were directional rather than release-grade:
delay_wgrad_compute: 31.20s steady-state mean overiterations 3-8
Treat this as evidence that EP overlap can help an inter-node alltoall MoE shape when communication is exposed. It is not proof that delayed wgrad is a separate win, and it does not validate the fused permutation path. An earlier 2026-05-16 short smoke on the same shape showed the same pattern.
src/megatron/bridge/training/comm_overlap.pysrc/megatron/bridge/training/flex_dispatcher_backend.pysrc/megatron/bridge/training/config.pytests/unit_tests/training/test_comm_overlap.pytests/unit_tests/training/test_deepep.pymoe_shared_expert_overlap andoverlap_moe_expert_parallel_comm can conflict. Disable shared expert overlap when using the dispatch overlap path.
active. Without it, the overlap scheduling cannot interleave correctly.
moe_flex_dispatcher_backend="deepep" alonedoes nothing if moe_token_dispatcher_type is still "alltoall".
disabled. You need to explicitly enable it via overrides.
communication is already a visible slice of step time. It is not guaranteed to help every small or lightly loaded EP run.
Look for overlap-related log messages during initialization. The comm overlap validation in comm_overlap.py will raise if prerequisites are not met, so a clean startup confirms the feature is active.
For a short performance-harness smoke, keep the command shape explicit and vary only one overlap knob at a time:
bashuv run python scripts/performance/run_script.py \ -m qwen \ -mr qwen3_30b_a3b \ --task pretrain \ -g h100 \ -c bf16 \ -ng 16 \ -gn 8 \ --max_steps 8 \ --cuda_graph_impl none \ --moe_flex_dispatcher_backend None \ --moe_a2a_overlap false \ --tokenizer_type NullTokenizer \ comm_overlap.overlap_moe_expert_parallel_comm=true \ comm_overlap.delay_wgrad_compute=false \ model.moe_shared_expert_overlap=false
If fused MoE permutation fails during bring-up, add model.moe_permute_fusion=false to separate overlap timing from runtime-stack validation, then retest with the matched production container.
Other measured skills in the registry, with their headline benchmark lift.