Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 156% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 256% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 186% | 0% |
NeMo AutoModel uses PyTorch-native distributed training. All parallelism is orchestrated through a single MeshContext object that holds device meshes, strategy configs, and axis names. <!-- NVSkills catalog signing requested after PR #2937 (2026-07-31). -->
For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question.
Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks.
Recommend strategy: fsdp2. Mention tp_size, pp_size, cp_size, ep_size, and the pipeline sub-config. State that dp_size is inferred from world_size / (tp_size * pp_size * cp_size).
yamldistributed: strategy: fsdp2 tp_size: 8 pp_size: 4 cp_size: 1 ep_size: 1 pipeline: pp_schedule: interleaved1f1b pp_microbatch_size: 1
Recommend strategy: fsdp2 with ep_size > 1. Say this creates a separate moe_mesh; include the moe sub-config when relevant; state that ep_size must divide dp_size * cp_size. Do not recommend megatron_fsdp or ddp.
yamldistributed: strategy: fsdp2 ep_size: 8 moe: reshard_after_forward: false
Say no for pipeline parallelism, expert parallelism, and sequence_parallel. Recommend fsdp2 for PP, EP, or sequence_parallel; mention that DDP is only simple data parallelism.
Three strategies are available, selected via the distributed.strategy YAML key:
| Strategy | YAML value | Best for | |---|---|---| | FSDP2 | fsdp2 | General use, recommended default. Supports TP, PP, CP, EP, HSDP. | | MegatronFSDP | megatron_fsdp | NVIDIA Megatron-style FSDP. No PP, no EP, no sequence_parallel. | | DDP | ddp | Simple data parallelism only. No TP, PP, CP, or EP. |
Decision tree:
fsdp2 (default). Use ddp only if you need the simplest possible setup.fsdp2 with appropriate TP/PP sizing.fsdp2 with ep_size > 1 (creates a separate moe_mesh).fsdp2 with PP + TP.cp_size > 1).When answering strategy-selection questions, state the chosen distributed.strategy first, then enumerate the YAML fields the user must set.
Quick TP + PP answer:
strategy: fsdp2; do not use megatron_fsdp when pipeline parallelism is required.tp_size for tensor parallelism and pp_size for pipeline parallelism.pipeline: sub-config with pp_schedule and pp_microbatch_size.dp_size unset or none; it is inferred as world_size / (tp_size * pp_size * cp_size).Quick MoE expert-parallel answer:
strategy: fsdp2 and ep_size > 1.moe: sub-config only when ep_size > 1; it maps to MoEParallelizerConfig.moe_mesh for expert parallelism in addition to the main device_mesh.megatron_fsdp or ddp for expert parallelism; megatron_fsdp has no EP support.ep_size must divide dp_size * cp_size and that megatron_fsdp does not support EP, PP, or sequence_parallel.The distributed section in the recipe YAML maps directly to parse_distributed_section() in recipes/_dist_utils.py:
yamldistributed: strategy: fsdp2 # fsdp2 | megatron_fsdp | ddp dp_size: none # auto-calculated from world_size / (tp * pp * cp) dp_replicate_size: none # FSDP2-only, for HSDP tp_size: 1 pp_size: 1 cp_size: 1 ep_size: 1 # Strategy-specific flags (forwarded to the strategy dataclass): sequence_parallel: false activation_checkpointing: false defer_fsdp_grad_sync: true # FSDP2 only # Sub-configs (optional): pipeline: pp_schedule: 1f1b pp_microbatch_size: 1 # ... see PipelineConfig fields moe: reshard_after_forward: false # ... see MoEParallelizerConfig fields
The dp_size is always inferred:
dp_size = world_size / (tp_size * pp_size * cp_size)initialize_distributed() [components/distributed/init_utils.py]
-> initializes torch.distributed process group and returns DistInfo
YAML distributed section + DistInfo.world_size
-> parse_distributed_section() [recipes/_dist_utils.py]
-> create_distributed_setup_from_config() [recipes/_dist_utils.py]
-> DistributedSetup.build() [components/distributed/config.py]
-> instantiate_infrastructure() [_transformers/infrastructure.py]
-> _instantiate_distributed() -> FSDP2Manager / MegatronFSDPManager / DDPManager
-> _instantiate_pipeline() -> AutoPipeline (if pp_size > 1)
-> parallelize_fn -> MoE parallelizer (if ep_size > 1) or PP wrapper
-> apply_model_infrastructure() [_transformers/infrastructure.py]
-> _shard_pp() or _shard_ep_fsdp() (applies sharding to the model)yamldistributed: strategy: fsdp2 tp_size: 1 cp_size: 1
This auto-calculates dp_size = world_size and applies fully_shard() per transformer block via DTensor-based sharding.
Keep TP within a single NVLink domain (typically one node):
yamldistributed: strategy: fsdp2 tp_size: 4 # 2, 4, or 8 -- must divide GPUs per node sequence_parallel: true
The TP plan is auto-selected based on the model type. Pass a custom plan via the Python API if needed:
pythonconfig = FSDP2Config(sequence_parallel=True, tp_plan=my_custom_plan)
yamldistributed: strategy: fsdp2 pp_size: 2 pipeline: pp_schedule: interleaved1f1b # 1f1b, gpipe, interleaved_1f1b, etc. pp_microbatch_size: 4 scale_grads_in_schedule: false
The model must have a _pp_plan attribute (set on the HF model class) for AutoPipeline to know how to split layers across stages. Models without _pp_plan are not compatible with PP.
Intra-node full sharding + inter-node replication via a 2D DeviceMesh:
yamldistributed: strategy: fsdp2 dp_replicate_size: 2 # must divide dp_size
Constraint: dp_replicate_size < dp_size (pure replication with no sharding is not supported by FSDP2).
Trades compute for memory by recomputing activations during backward:
yamldistributed: activation_checkpointing: true
This is a model-build/training behavior flag, not mesh topology. Dense strategies read it from the strategy config; EP/MoE paths pass the recipe-level flag directly into model infrastructure.
FSDP2 defers gradient sync to the final micro-batch by default for communication overlap:
yamldistributed: defer_fsdp_grad_sync: true # default
FSDP2Config defaults to bfloat16 for all three precision knobs via MixedPrecisionPolicy(param_dtype=bf16, reduce_dtype=bf16, output_dtype=bf16, cast_forward_inputs=True). Override via the Python API:
pythonfrom torch.distributed.fsdp import MixedPrecisionPolicy config = FSDP2Config( mp_policy=MixedPrecisionPolicy(param_dtype=torch.float16, reduce_dtype=torch.float32), )
_pp_plan (a dict mapping module FQNs to stages).pp_size > 1 in the distributed section.pipeline sub-config with schedule and microbatch size.Defined in PipelineConfig.pp_schedule:
1f1b (one-forward-one-backward, default)gpipeinterleaved_1f1b / interleaved1f1blooped_bfsdfsv_schedulezero_bubbleyamldistributed: strategy: fsdp2 pp_size: 2 pipeline: pp_schedule: interleaved1f1b pp_microbatch_size: 4 scale_grads_in_schedule: false checkpoint: model_save_format: safetensors save_consolidated: final
AutoPipeline.build() calls pipeline_model() which splits the model into stages using the model's _pp_plan, creates PipelineStage objects, and builds the schedule. During training, schedule.step() drives forward and backward through the pipeline.
Use CP for long sequences (8K+). CP shards Q/K/V on the sequence dimension as DTensors.
yamldistributed: strategy: fsdp2 cp_size: 2 # or 4, 8
attention. SDPBackend.MATH is not compatible with DTensor.
is_causal=True is set viaforward pre-hooks registered by attach_context_parallel_hooks().
apply_model_infrastructure() callsattach_context_parallel_hooks() on each model part (for non-TE models).
make_cp_batch_and_ctx() creates a CP contextmanager that shards the batch along the sequence dimension and sets up context_parallel() from torch.distributed.tensor.experimental.
make_cp_batch_for_te() uses THD format andTE's thd_get_partitioned_indices for sharding.
CP works with packed sequences. The packed_sequence_size must be divisible by cp_size. When using TE, chunks are sharded per-chunk via _shard_thd_chunk_for_te().
Packing multiple sequences into a single training sample for efficiency.
yamlpacked_sequence: packed_sequence_size: 4096 # 0 = disabled step_scheduler: local_batch_size: 1 # must be 1 for packed sequences
When packed_sequence_size > 0, the dataset collator packs sequences up to that length. local_batch_size must be 1 because each "sample" is already a packed batch.
Set ep_size > 1 to distribute experts across GPUs. This creates a separate moe_mesh alongside the main device_mesh:
yamldistributed: strategy: fsdp2 ep_size: 8 activation_checkpointing: true
The moe_mesh shape is (pp_size, ep_shard_size, ep_size) with dimension names ("pp", "ep_shard", "ep").
Constraint: dp_cp_size (= dp_size * cp_size) must be divisible by ep_size.
yamldistributed: strategy: fsdp2 ep_size: 8 activation_checkpointing: true moe: reshard_after_forward: false ignore_router_for_ac: false wrap_outer_model: true
The moe sub-section maps to MoEParallelizerConfig and is only instantiated when ep_size > 1.
yamldistributed: strategy: fsdp2 tp_size: 1 cp_size: 1 pp_size: 1 ep_size: 8 sequence_parallel: false activation_checkpointing: true
Despite its name, megatron_fsdp does not support expert parallelism (ep_size > 1), pipeline parallelism (pp_size > 1), or sequence_parallel. Use fsdp2 for these features.
| Model size | TP | PP | CP | Strategy | |---|---|---|---|---| | < 3B | 1 | 1 | 1 | FSDP2 (DP only) | | 3-13B | 2-4 | 1 | 1 | FSDP2 + TP | | 13-70B | 4-8 | 2-4 | 1 | FSDP2 + TP + PP | | 70B+ | 8 | 4-8 | 1 | FSDP2 + TP + PP | | Any + long seq (8K+) | as above | as above | 2-8 | add CP |
MoE models need less TP than dense models of similar total parameter count because only a fraction of parameters are active per token. EP is the primary scaling dimension:
| Model | TP | PP | EP | Notes | |---|---|---|---|---| | Small MoE (<10B total) | 1 | 1 | 8 | EP only | | Medium MoE (10-30B total) | 1-2 | 1 | 8 | small TP for shared layers | | Large MoE (100B+ total) | 1-2 | 4+ | 8-64 | PP for depth, EP for experts |
When not using YAML recipes, configure distributed training via Python:
pythonfrom nemo_automodel.components.distributed import ( DistributedSetup, FSDP2Config, ParallelismSizes, initialize_distributed, ) dist_env = initialize_distributed("nccl") distributed_setup = DistributedSetup.build( strategy=FSDP2Config(sequence_parallel=True), parallelism_sizes=ParallelismSizes(tp_size=2), activation_checkpointing=True, world_size=dist_env.world_size, )
Or pass directly to from_pretrained:
pythonfrom nemo_automodel import NeMoAutoModelForCausalLM model = NeMoAutoModelForCausalLM.from_pretrained( "meta-llama/Llama-3.2-1B", distributed_setup=distributed_setup, )
Strategy config dataclasses:
components/distributed/config.py
FSDP2Config -- sequence_parallel, tp_plan, mp_policy, offload_policy,
activation_checkpointing, defer_fsdp_grad_sync
MegatronFSDPConfig -- zero_dp_strategy, overlap_grad_reduce, overlap_param_gather, etc.
DDPConfig -- activation_checkpointing onlyMeshContext (single source of truth for parallelism):
components/distributed/mesh.py
MeshContext -- device_mesh, moe_mesh
Properties: tp_size, pp_size, cp_size, ep_size, dp_size, dp_replicate_size
MeshAxisName -- PP, DP, DP_REPLICATE, DP_SHARD, DP_SHARD_CP, DP_CP, CP, TP, EP, EP_SHARDMesh context and raw mesh creation:
components/distributed/config.py
DistributedSetup.build() -- builds MeshContext from strategy + parallelism
components/distributed/mesh_utils.py
_create_device_meshes() -- routes to FSDP2/MegatronFSDP/DDP raw mesh creation
_create_fsdp2_device_mesh() -- shape (pp, dp_replicate, dp_shard, cp, tp) + flattened submeshes
_create_megatron_fsdp_device_mesh() -- shape (dp, cp, tp)Distributed managers:
components/distributed/fsdp2.py -- FSDP2Manager.parallelize()
components/distributed/megatron_fsdp.py -- MegatronFSDPManager.parallelize()
components/distributed/ddp.py -- DDPManagerPipeline parallelism:
components/distributed/pipelining/config.py -- PipelineConfig dataclass
components/distributed/pipelining/autopipeline.py -- AutoPipeline orchestrator
components/distributed/pipelining/functional.py -- pipeline_model(), schedule creation
components/distributed/pipelining/hf_utils.py -- HF model validation for PPContext parallelism:
components/distributed/context_parallel/utils.py
make_cp_batch_and_ctx() -- creates CP context manager + shards batch
create_context_parallel_ctx() -- wraps torch.distributed.tensor.experimental.context_parallel
attach_context_parallel_hooks() -- strips attention_mask, sets is_causal=True
make_cp_batch_for_te() -- TE-specific CP batch sharding (THD format)Infrastructure orchestration:
_transformers/infrastructure.py
instantiate_infrastructure() -- config objects -> runtime objects
apply_model_infrastructure() -- applies sharding, PEFT, checkpoints to model
_shard_pp() -- pipeline parallel path
_shard_ep_fsdp() -- EP + FSDP path (non-PP)YAML parsing:
recipes/_dist_utils.py
parse_distributed_section() -- YAML dict -> typed configs + sizes
create_distributed_setup_from_config() -- recipe adapter: parse + create DistributedSetup; does not init process groupMoE config:
components/distributed/config.py
MoEParallelizerConfig -- reshard_after_forward, ignore_router_for_ac, wrap_outer_model, etc.
components/moe/config.py
MoEConfig -- n_routed_experts, n_activated_experts, score_func, etc.NVLink domain. Use PP or DP for cross-node scaling.
_pp_plan on the model class. Not all HF models have this.Check validate_hf_model_for_pipeline_support() before enabling PP.
(interleaved_1f1b) and smaller microbatches to reduce bubble time.
safetensors withsave_consolidated: final for final HF export, or save_consolidated: false plus the generated model/consolidate.sh helper for offline export.
Attention) or TE attention only. SDPBackend.MATH is not compatible with DTensor.
dp_size * cp_size. The device meshcreation asserts dp_cp_size % ep_size == 0.
(pp_size > 1), EP (ep_size > 1), or sequence_parallel. The MeshContext validation raises on these combinations.
HSDP. Validation raises on any of these.
recomputing activations during backward, but adds ~30% compute overhead.
bfloat16 policy works for most models. FP16 models may need a custom MixedPrecisionPolicy.
packed_sequence_size must be divisible by cp_size when using CPwith packed sequences.
dp_replicate_size is FSDP2-only. Passing it with megatron_fsdpor ddp raises a ValueError.
Run the smallest recipe that exercises the requested strategy. Success means exit code 0, finite loss, no NCCL timeout, and log output matching the expected TP/PP/CP/EP sizes.
Other measured skills in the registry, with their headline benchmark lift.