Dylan Patel on

gpu

15 quotes · Feb 2023 – Aug 2026

Saidverbatim, newest first

  1. Patel argues silicon supply can be optimized for either high throughput or high interactivity use cases.

    “Supply of silicon can go many ways. You can either leverage it to high throughput things or high interactivity things.”

    26:41 · SemiAnalysis · 17 Aug 2026 · permalink
  2. Patel reveals SemiAnalysis operates over $80 million in compute for daily automated benchmarking across all major AI chips.

    “Every day it runs on an automated CI, and we run it on all the latest Chinese models from GLM, Zebu, Moonshot, Kimi, Alibaba, all these models we run.”

    16:08 · RAISE Summit · 16 Jul 2026 · permalink
  3. Supermicro

    “It's 72 GPUs. It's got more memory than Vera Rubin. It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later.”

    2:04 · Supermicro · 15 Jul 2026 · permalink
  4. Patel says AMD MI450X has 72 GPUs, more memory and flops than Vera Rubin, launching three to six months later.

    “It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later. I mean, three to six months after,”

    2:09 · Supermicro · 15 Jul 2026 · permalink
  5. Patel reports Trainium rents for under $10 billion per gigawatt while GPUs cost $12-13 billion.

    “Tranium sells at sub $10,000,000,000 per gigawatt rental rate to Anthropic and to OpenAI. GPUs, at least before the craziness of the last six months, usually went around 12 to $13,000,000,000 per gigawatt.”

    58:22 · Sequoia Capital · 30 Jun 2026 · permalink
  6. Swole as a Service

    “We try and track the entire supply chain from tools that manufacture chips, fabs, data centers, energy, industrials, and then AI models and who's using them, how much, and where.”

    1:51 · Swole as a Service · 13 May 2026 · permalink
  7. Patel reports 10-15% of NVIDIA GPUs fail and require RMA within first two weeks of cluster deployment.

    “When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine. Like you have to receipt them, whatever.”

    4:35 · TBPN · 3 Feb 2026 · permalink
  8. Patel says 10-15% of NVIDIA GPUs fail and need RMA in the first two weeks after cluster deployment.

    “When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine.”

    4:35 · TBPN · 3 Feb 2026 · permalink
  9. Patel says OpenAI and Meta run NVIDIA GPUs at lower power to fit 10% more chips despite worse TCO.

    “Even though it's terrible on a TCO basis, they were able to get, you know, 10% more GPUs in, and it's great. And Meta has done similar.”

    15:23 · Open Compute Project · 23 Oct 2025 · permalink
  10. Patel reveals DeepSeek inference implementation requires 160 GPUs worth over $10 million of hardware per replica.

    “That's over $10,000,000 of hardware, and then that's just one replica, then you'll have a lot of replicas and you share the caching servers between them.”

    6:11 · No Priors: AI, Machine Learning, Tech, & Startups · 14 Aug 2025 · permalink
  11. Patel says NVIDIA made 4 million GPUs last year and will produce 7 million this year.

    “Nvidia made over 4,000,000 GPUs last year, they're making over 7,000,000 this year, right? High end data center GPUs.”

    19:58 · Special Competitive Studies Project · 13 Mar 2025 · permalink
  12. Patel reports China imported one million H20 GPUs in recent quarters, enough for largest cluster.

    “Just in Q3, Q4 and the early part of Q1 this year, they imported a million H20s, Right?”

    20:18 · Special Competitive Studies Project · 13 Mar 2025 · permalink
  13. Patel says GPU orders take four to six months from placement to data center installation.

    “call it four or five five, six months between, you know, when an order is placed and you can actually have it installed in your data center if it got worked on immediately.”

    3:39 · The Inside View · 9 Aug 2023 · permalink
  14. Patel says NVIDIA GPU bandwidth increased less than 10x while FLOPS increased 100x from 2016 to 2023.

    “The bandwidth has not even gone up one order of magnitude. Right? Whereas flops have gone up two orders of magnitude. So so less than one order of magnitude versus two orders of magnitude increase.”

    7:28 · Gradient Flow · 2 Feb 2023 · permalink
  15. Patel states memory cost has quadrupled in NVIDIA GPUs while per-gigabyte pricing remained flat since 2016.

    “And now with the h 100, they have 96 or 80 gigabytes of memory, but the cost per gigabyte is the same.”

    9:54 · Gradient Flow · 2 Feb 2023 · permalink

Everything Dylan Patel is on record saying · RSS