Dylan Patel on

benchmarking

19 quotes · Oct 2025 – Jul 2026

Saidverbatim, newest first

  1. Patel says InferenceX now has over $80 million of compute across multiple vendors.

    “we have over $80,000,000 of compute GPUs from NVIDIA AMD, TPUs from, Google, Tranium from Amazon, and we run this benchmark constantly on the newest inference engine, newest drivers, newest, PyTorch version,”

    15:53 · RAISE Summit · 16 Jul 2026 · permalink
  2. Patel reveals SemiAnalysis operates over $80 million in compute for daily automated benchmarking across all major AI chips.

    “Every day it runs on an automated CI, and we run it on all the latest Chinese models from GLM, Zebu, Moonshot, Kimi, Alibaba, all these models we run.”

    16:08 · RAISE Summit · 16 Jul 2026 · permalink
  3. Patel says SemiAnalysis analyzed over $5 million worth of Claude production traces to benchmark agentic workloads.

    “Initially, when we were benchmarking the difference between these chips and different engines, different schemes for parallelism, we were just running it, you know, fixed context length.”

    16:21 · RAISE Summit · 16 Jul 2026 · permalink
  4. Patel says InferenceX analyzed over $5 million worth of Claude Code production traces.

    “we've analyzed over $5,000,000 worth of Claude code traces. Right? So this is real production traffic that people have donated to us,”

    16:31 · RAISE Summit · 16 Jul 2026 · permalink
  5. Patel's benchmarking found Blackwell is 30x faster than Hopper on DeepSeek v3, exceeding Jensen's 25x claim.

    “In DeepSeek v three, Blackwell is 30 x faster than Hopper on on somewhere on the continuum.”

    12:17 · WisdomTree in Europe · 9 Jul 2026 · permalink
  6. Patel secured over $50 million in donated hardware for InferenceX, expanding to $100 million with TPUs and training.

    “We've got over $50,000,000 of hardware donated to us. Once we launch TPUs and training, would actually be over $100,000,000 of hardware.”

    16:43 · Sequoia Capital · 30 Jun 2026 · permalink
  7. Patel says Nvidia and AMD both lie about peak flops specs which are impossible to achieve.

    “All their quoted specs are lies impossible to achieve Whether it's Nvidia or AMD, neither of them you can ever hit their peak flops.”

    3:57 · TensorWave · 30 Apr 2026 · permalink
  8. Patel claims Nvidia achieves 50% sustained performance on common 8k gemm while AMD gets only 30%.

    “if you just take a eight k by eight k by eight k gem, very common shape and you run it on the map mode unit of AMD and Nvidia, you get like 50% sustained performance on Nvidia and you get like 30% sustained performance on AMD.”

    4:23 · TensorWave · 30 Apr 2026 · permalink
  9. Patel asserts vendors cannot achieve their advertised flops specifications in real-world conditions.

    “There there is no functional way to get anywhere close to their flops that they advertise.”

    4:37 · TensorWave · 30 Apr 2026 · permalink
  10. Patel explains Inference X was created because vendor-claimed performance is unattainable with any available framework.

    “There's no way to get the performance that vendors like to claim. And so our whole thing there was, well, how do we actually have a benchmark that represents what people actually get on performance?”

    11:41 · TensorWave · 30 Apr 2026 · permalink
  11. Patel notes software dependencies change daily or multiple times weekly across the entire stack.

    “And furthermore, with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”

    11:55 · TensorWave · 30 Apr 2026 · permalink
  12. Patel explains AI software stacks update multiple times per week making performance measurement a moving target.

    “with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”

    11:55 · TensorWave · 30 Apr 2026 · permalink
  13. Patel notes AI software stacks update constantly with nightly builds making performance measurement time-sensitive.

    “with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly.”

    11:55 · TensorWave · 30 Apr 2026 · permalink
  14. Patel observes VLLM and SGLANG framework developers compete publicly on Twitter using InferenceX as a leaderboard.

    “VLLM and SGLANG love competing with each other. They used to post on Twitter all the time about how they had beat the other one in certain some kind of performance”

    21:11 · TensorWave · 30 Apr 2026 · permalink
  15. Patel reports AMD and Nvidia engineers treat Inference X as a competitive leaderboard.

    “AMD and Nvidia. They love to compete with each other. And AMD engineers and Nvidia engineers look at Inference X as a leader board.”

    21:21 · TensorWave · 30 Apr 2026 · permalink
  16. Patel says InferenceMax runs daily benchmarks to track hardware performance improvements as software optimizes.

    “And why this is important is you can see the progress of hardware over time as the software stack gets more optimized.”

    4:58 · Open Compute Project · 23 Oct 2025 · permalink
  17. Patel quantifies InferenceMax uses tens of millions of dollars in GPU hardware with daily software updates.

    “we have tens of millions of dollars of hardware of the current and last generation GPUs. As I mentioned before, the software updates every single day.”

    6:43 · Open Compute Project · 23 Oct 2025 · permalink
  18. Patel quantifies GB200 as 10x more power efficient than H200 at certain interactivity rates.

    “at certain interactivity rates, I. E. Tokens per second per user, it's 10x more efficient per watt, right, compared to H200.”

    15:20 · Open Compute Project · 23 Oct 2025 · permalink
  19. Patel reports AMD MI355 beats NVIDIA B200 in some document processing scenarios with open source software.

    “in some cases, AMD actually does have a publicly usable open source implementation that beats NVIDIA even, right, with the MI355 versus V200, which is a surprise, right?”

    16:28 · Open Compute Project · 23 Oct 2025 · permalink

Everything Dylan Patel is on record saying · RSS