Patel claims smaller Quen models with 27B total parameters and 2B active outperform GPT-4 from three years ago.
“If you look back three years, it's GPT-four. Now it's like maybe like quen one of the smaller quen models that's like 27 b parameters total and like 2,000,000,000 active is like way better.”
Patel argues three years of improvement came from model layer advances plus hardware-software co-design, citing smaller Qwen models outperforming GPT-4.
“If you look back three years, it's GPT-four. Now it's maybe like quen, one of the smaller quen models that's like 27B parameters total and 2,000,000,000 active is way better.”
Patel argues co-optimizing hardware, software, and models produces 100x gains instead of 8x from multiplicative 2x improvements per layer.
“The real breakthrough innovation is when you leapfrog a few layers, you co optimize and co design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x, because you've optimized it across all three layers.”
post Patel says less compute efficiency means fewer tokens and intelligence for everyone
https://x.com/dylan522p/status/2086333447556227503
post Patel asks if there is a single good airport in Europe, claiming zero efficiency.
https://x.com/dylan522p/status/2093962660492546469