GPU Communication Library in Meta-Scale AI Clusters.pdf
1、GPU Communication Libraryin Meta-Scale AI ClustersJames Hongyi ZengMetaAI CLUSTERSMetas AI Clusters Getting Big 1GW,Announced 2026Lebanon,IndianaUp to 5GW,ETA 2028HyperionRichland Parish,Louisiana1GWPrometheusNew Albany,OhioDiverse Communication NeedsLLM StagesExample ChallengesCritical CommsExample
2、 OSS FrameworksExample OSS Comms LibrariesPre-TrainingMaximize massive throughput across thousands of GPUsFSDP(AllGather/ReduceScatter)torchtitan,Megatron-LMNCCL,RCCLPost-Training(RLHF)Rapid weight shipping from Learners to ActorsP2P(Send/Recv)torchforge,verlRay RPCInferenceKV Cache shipping for Pre
3、fill-Decode disaggregationP2P(Send/Recv)Expert Parallelism(AlltoAll)vLLM,SGLangDeepEP,Mooncake TEPre-training:Fault Tolerance Frequent failures are inherent risk for large scale synchronous pre-training18 min at 100K GPUs10 min restart time8 min effective training!Diverse set of interruptionUnavoida
4、ble Hardware FailuresPre-training:Fault Tolerance Solution:Parallelism Aware Fault ToleranceDivide GPUs into multiple replicasSame rank in each replica communicates with AllReduceDynamically scale up and downMeta Collective Communication Library(MCCL)Fast recoveryOne replica down,the rest continuesF





点击查看更多