Patel explains that AdamW optimizer requires four bytes per parameter, creating 400 gigabytes of data for 100B parameter models.
“when you look at LLMs, people use AdamW. And AdamW as an optimizer is like four bytes per parameter roughly, if I recall correctly the optimizer state.”
Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data every two seconds during training.
“They're doing it for like multi trillion. Right? So let's call it 10,000,000,000,000 parameters, four bytes parameter, that's 40 terabytes of data you need to transmit in two seconds.”
Patel says NVIDIA GPU bandwidth increased less than 10x while FLOPS increased 100x from 2016 to 2023.
“The bandwidth has not even gone up one order of magnitude. Right? Whereas flops have gone up two orders of magnitude. So so less than one order of magnitude versus two orders of magnitude increase.”