Dylan Patel on

gb200

5 quotes · October 2025

Saidverbatim, newest first

  1. Patel quantifies GB200 as 10x more power efficient than H200 at certain interactivity rates.

    “at certain interactivity rates, I. E. Tokens per second per user, it's 10x more efficient per watt, right, compared to H200.”

    15:20 · Open Compute Project · 23 Oct 2025 · permalink
  2. Patel notes GB200 expanded NVLink from 8 to 72 GPUs and rack power jumped from 10 to 140 kilowatts.

    “Now you have 72 GPUs. And if you go look at the rack, right? It's completely different, right? It's liquid cooled. It's an entire rack that consumes 140 kilowatts, whereas H100 servers consumed 10 kilowatts.”

    11:27 · Together AI · 3 Oct 2025 · permalink
  3. Patel explains GB200's 72-GPU configuration creates reliability challenges compared to 8-GPU systems due to higher failure probability.

    “If the reliability of each GPU is the same, then when a single GPU fails in 72 GPUs, you have a much higher chance of something failing, right?”

    12:40 · Together AI · 3 Oct 2025 · permalink
  4. Patel confirms OpenAI runs production inference on GB200 despite reliability requiring workloads handle 64 of 72 GPUs.

    “OpenAI has said they're running production inference on GV200 a couple of months ago, in fact. Right?”

    13:42 · Together AI · 3 Oct 2025 · permalink
  5. Patel says B200 is better for training while GB200 is better for inference, reversing expected use.

    “And so you've sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.”

    29:16 · Together AI · 3 Oct 2025 · permalink

Everything Dylan Patel is on record saying · RSS