Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek.
Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek.
Megatron-Core trains LLMs from 2B to 462B parameters with up to 47% Model FLOP Utilization on H100 GPUs through advanced parallelism strategies.
Installation:
# Docker (recommended)
docker run --gpus all -it --rm nvcr.io/nvidia/pytorch:25.04-py3
# Or pip
pip install megatron-core
Simple distributed training:
# Train with 2 GPUs using data parallelism
torchrun --nproc_per_node=2 examples/run_simple_mcore_train_loop.py
# Or LLaMA-3 8B training
./examples/llama/train_llama3_8b_fp8.sh
Common workflows
Workflow 1: Train LLaMA-style model with 3D parallelism
# Start with 1, increase until OOMfor MBS in 1 2 4 8; doecho"Testing micro-batch-size=$MBS"
torchrun ... --micro-batch-size $MBSdone
Typical values:
7B model: 4-8
70B model: 1-2
405B model: 1
Step 4: Tune parallelism degrees
Rules of thumb:
Tensor Parallel: Use ≤8 (limited by NVLink within node)
Pipeline Parallel: Use for >70B models
Context Parallel: Use for sequences >8K tokens
Data Parallel: Fill remaining GPUs
Pipeline bubbles: Use interleaved pipeline schedule
--num-layers-per-virtual-pipeline-stage 2
Data loading: Use fast data loader
--dataloader-type cyclic
Issue: Diverging loss
Stabilize training:
--lr-warmup-iters 2000 # Longer warmup
--clip-grad 1.0 # Gradient clipping
--init-method-std 0.006 # Smaller init
--attention-dropout 0.0 # No dropout in attention
--hidden-dropout 0.0 # No dropout in FFN
Advanced topics
Parallelism strategies: See references/parallelism-guide.md for detailed comparison of TP/PP/DP/CP/EP with performance analysis and when to use each.
Performance benchmarks: See references/benchmarks.md for MFU numbers across different model sizes and GPU configurations.
Production configurations: See references/production-examples.md for real-world setups from LLaMA 3 405B, Nemotron-4 340B, and DeepSeek-V3 671B.
Training recipes: See references/training-recipes.md for complete hyperparameter configurations for GPT/LLaMA/Mixtral architectures.
Hardware requirements
GPU: NVIDIA Ampere+ (A100, H100, B200)
Turing works but slower
FP8 requires Hopper/Ada/Blackwell
Network: InfiniBand or 400Gb+ Ethernet for multi-node
Memory per GPU:
7B model: 40GB+
70B model: 80GB (with TP=4)
405B model: 80GB (with TP=8, PP=8)
Storage: Fast NVMe for checkpoints (1TB+ for 70B+ models)
Known Conflicts
Do not install alongside deepspeed in the same environment. Both depend on Apex and specific PyTorch builds. Version mismatches cause CUDA compilation errors. Use separate environments.