On the record about
4 people · 21 quotes · 2 Feb 2023 to 17 Aug 2026
3 of 4 lanes rest on fewer than 5 quotes and are marked thin. Offsets are days from the middle first-quote date, 21 May 2024 — a date, and nothing else. It is not a claim about who reached a view first.
Patel says NVIDIA GPU bandwidth increased less than 10x while FLOPS increased 100x from 2016 to 2023.
“The bandwidth has not even gone up one order of magnitude. Right? Whereas flops have gone up two orders of magnitude. So so less than one order of magnitude versus two orders of magnitude increase.”
Patel states memory cost has quadrupled in NVIDIA GPUs while per-gigabyte pricing remained flat since 2016.
“And now with the h 100, they have 96 or 80 gigabytes of memory, but the cost per gigabyte is the same.”
Patel says GPU orders take four to six months from placement to data center installation.
“call it four or five five, six months between, you know, when an order is placed and you can actually have it installed in your data center if it got worked on immediately.”
Huang explains CUDA crushed NVIDIA's gross margins because it added cost without applications to justify premium pricing.
“And if people aren't willing to pay you for it but your cost went up, then your gross margins get crushed”
Baker says GPUs got 50x faster while data centers only 4-5x faster, creating investment opportunity in infrastructure.
“GPUs have gotten 50 times faster, and the rest of the data center has only gotten four to five times faster.”
Baker explains GPUs got 50x faster while data center infrastructure only got 4-5x faster, causing low utilization.
“GPUs have gotten 50 times faster, and the rest of the data center has only gotten four to five times faster. And that is why MFU is so low.”
Patel argues export controls aim to limit AI usage scale in China, not prevent AGI development.
“And I think that is a much easier goal to achieve than trying to debate on what AGI is.”
Patel says NVIDIA made 4 million GPUs last year and will produce 7 million this year.
“Nvidia made over 4,000,000 GPUs last year, they're making over 7,000,000 this year, right? High end data center GPUs.”
Patel reports China imported one million H20 GPUs in recent quarters, enough for largest cluster.
“Just in Q3, Q4 and the early part of Q1 this year, they imported a million H20s, Right?”
Patel reveals DeepSeek inference implementation requires 160 GPUs worth over $10 million of hardware per replica.
“That's over $10,000,000 of hardware, and then that's just one replica, then you'll have a lot of replicas and you share the caching servers between them.”
Gurley flags NVIDIA's unusual promise to buy unsold CoreWeave capacity as revealed in regulatory filings.
“One of the more peculiar of all the deals is, and this was disclosed in a CoreWeave filing, was NVIDIA has promised to buy any of CoreWeave's service availability that they can't sell to anyone else. That is very unusual. That's not the same as making an investment.”
Patel says OpenAI and Meta run NVIDIA GPUs at lower power to fit 10% more chips despite worse TCO.
“Even though it's terrible on a TCO basis, they were able to get, you know, 10% more GPUs in, and it's great. And Meta has done similar.”
Patel reports 10-15% of NVIDIA GPUs fail and require RMA within first two weeks of cluster deployment.
“When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine. Like you have to receipt them, whatever.”
Patel says 10-15% of NVIDIA GPUs fail and need RMA in the first two weeks after cluster deployment.
“When you first turn on the cluster, about ten to fifteen percent of them fail RMA in the first two weeks. Wow. And then that's fine.”
“We try and track the entire supply chain from tools that manufacture chips, fabs, data centers, energy, industrials, and then AI models and who's using them, how much, and where.”
Baker predicts disaggregation extends GPU lives to 10-15 years, lowering financing costs and saving private credit.
“This is going to be really good for the whole private credit industry. It's going to help finance the AI build out.”
Patel reports Trainium rents for under $10 billion per gigawatt while GPUs cost $12-13 billion.
“Tranium sells at sub $10,000,000,000 per gigawatt rental rate to Anthropic and to OpenAI. GPUs, at least before the craziness of the last six months, usually went around 12 to $13,000,000,000 per gigawatt.”
“It's 72 GPUs. It's got more memory than Vera Rubin. It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later.”
Patel says AMD MI450X has 72 GPUs, more memory and flops than Vera Rubin, launching three to six months later.
“It's got more more, memory, flops. It's got it's it's better than Vera Rubin in most every way. It's a little bit later. I mean, three to six months after,”
Patel reveals SemiAnalysis operates over $80 million in compute for daily automated benchmarking across all major AI chips.
“Every day it runs on an automated CI, and we run it on all the latest Chinese models from GLM, Zebu, Moonshot, Kimi, Alibaba, all these models we run.”
Patel argues silicon supply can be optimized for either high throughput or high interactivity use cases.
“Supply of silicon can go many ways. You can either leverage it to high throughput things or high interactivity things.”