Bringing PyTorch Monarch to AMD GPUs

PyTorch Monarch has been ported to AMD Instinct GPUs, enabling more reliable and fault-tolerant distributed training for large language models. This update allows training jobs to recover from node failures dynamically without requiring a full restart from checkpoints.
Why it matters
Improving fault tolerance in AI training reduces wasted computational resources and accelerates the development of large-scale AI models.
Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed training across hundreds or thousands of GPUs. At this scale, hardware failures are not exceptional events—they are expected. A single GPU memory error, network partition, or node crash can bring down an entire training run that has been progressing for days or weeks. While our previous work demonstrated near-linear scaling of FP8 training at scale (achieving 96.16% scaling efficiency on a 1024-GPU MI325 cluster with DeepSeekV3-671B), the key challenge remains: reliability at scale.
To address these challenges, we have brought PyTorch Monarch to AMD Instinct GPUs with ROCm, expanding the single-controller model beyond CUDA environments and bringing this emerging runtime to a broader hardware ecosystem.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in