Patel reports Microsoft aims to 10x training capacity every 18-24 months, representing a 10x increase from GPT-5 training.
“We try to 10x the training capacity every eighteen to twenty four months. And so this would be effectively a 10x increase. 10x from what GPD five was trained with.”
Patel states $10 billion data centers target automated software engineering, not chat models.
“So no one is trying to make with these $10,000,000,000 data centers, they're not trying to make chat models. Right?”
Patel says current models use 100,000 GPUs while next generation will require hundreds of thousands or millions.
“And next generation models that are trained on hundreds of thousands or even millions GPUs, right?”
Patel calculates next-generation clusters deliver 15x more compute through five times more GPUs and 3x performance gains.
“So you got you have five x the GPUs, and you have three x the performance per GPU roughly. So then you're at, like, 15 x more compute.”
Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data every two seconds during training.
“They're doing it for like multi trillion. Right? So let's call it 10,000,000,000,000 parameters, four bytes parameter, that's 40 terabytes of data you need to transmit in two seconds.”