Patel identifies disaggregated PD and wide EP as networking techniques where performance matters significantly.
“When someone has really good networking with these leading edge techniques to run a model such as KIMI or DeepSeq called disaggregated PD or wide EP, these are two different techniques, network performance is a humongous factor.”
Patel identifies distributed training across data centers with lower bandwidth as a key unsolved problem that would unlock massive scaling.
“Everything that we've seen so far is that large scale training has to happen in an individual data center with very high speed networking.”