Dylan Patel on

scaling

5 quotes · Dec 2023 – Aug 2026

Saidverbatim, newest first

  1. Patel reports Microsoft aims to 10x training capacity every 18-24 months, representing a 10x increase from GPT-5 training.

    “We try to 10x the training capacity every eighteen to twenty four months. And so this would be effectively a 10x increase. 10x from what GPD five was trained with.”

    1:19 · Dwarkesh Patel · 12 Nov 2025 · permalink
  2. Patel states $10 billion data centers target automated software engineering, not chat models.

    “So no one is trying to make with these $10,000,000,000 data centers, they're not trying to make chat models. Right?”

    34:22 · Alex Kantrowitz · 23 Apr 2025 · permalink
  3. Patel says current models use 100,000 GPUs while next generation will require hundreds of thousands or millions.

    “And next generation models that are trained on hundreds of thousands or even millions GPUs, right?”

    19:49 · Special Competitive Studies Project · 13 Mar 2025 · permalink
  4. Patel calculates next-generation clusters deliver 15x more compute through five times more GPUs and 3x performance gains.

    “So you got you have five x the GPUs, and you have three x the performance per GPU roughly. So then you're at, like, 15 x more compute.”

    26:01 · Unsupervised Learning: With Jacob Effron · 21 Jan 2025 · permalink
  5. Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data every two seconds during training.

    “They're doing it for like multi trillion. Right? So let's call it 10,000,000,000,000 parameters, four bytes parameter, that's 40 terabytes of data you need to transmit in two seconds.”

    23:38 · Scaling Intelligence · 12 Nov 2024 · permalink

Everything Dylan Patel is on record saying · RSS