Patel argues OpenAI models are optimized for GPUs while Anthropic and Google models are optimized for TPUs due to co-design.
“And the way that Anthropic and Google's models are headed, it's actually a terrible decision potentially for them to train with GPUs.”
Patel argues three years of improvement came from model layer advances plus hardware-software co-design, citing smaller Qwen models outperforming GPT-4.
“If you look back three years, it's GPT-four. Now it's maybe like quen, one of the smaller quen models that's like 27B parameters total and 2,000,000,000 active is way better.”
Patel argues co-optimizing hardware, software, and models produces 100x gains instead of 8x from multiplicative 2x improvements per layer.
“The real breakthrough innovation is when you leapfrog a few layers, you co optimize and co design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x, because you've optimized it across all three layers.”