On the record about
5 people · 75 quotes · 26 Nov 2019 to 25 Aug 2026
2 of 5 lanes rest on fewer than 5 quotes and are marked thin. Offsets are days from the middle first-quote date, 13 Oct 2024 — a date, and nothing else. It is not a claim about who reached a view first.
Baker argues AI revolution stems from cloud computing power and mobile-generated data, not algorithmic advances.
“The only thing that has enabled the AI revolution that we're living through, which I think we're at the bottom of the first inning in, is one, we had the ability to do cloud computing, so just apply significantly more computational power to old algorithms, and then b, we had dramatically more data.”
Baker states data quantity is the single most predictive element of AI quality, not algorithms or infrastructure.
“The single most predictive element of knowledge about AI quality is the quantity of data used to train the algorithm.”
Baker cites research showing every 10x increase in training data doubles AI quality.
“and it's been very well established in multiple papers from both Google and Microsoft research that for every order of magnitude increase in the data you use to train an algorithm, the quality of the AI doubles.”
Baker predicts Tesla FSD will achieve 100x improvement quickly as compute scales to GPT-4.5 levels.
“I think they're going to go really fast to GPT-4.5 compute, which means you're going to get, using these orders of magnitude, you're going get a 100x improvement really fast.”
Gurley explains Strawberry requires 10X processing for linear improvement, meaning 10-100X inference costs for better solutions.
“in order to get linear improvement, you have to do maybe 10X the amount of processing. And this is all inference. So what are the implications of that?”
Gerstner reports Jensen predicts inference will scale 100x to billion-x with 40% of NVIDIA revenue already from inference
“he said as a consequence of that, inference is going to a 100 x, thousand x, a million x, maybe even a billion x.”
Patel says o1 generates 40k sequence lengths versus 4k for standard models, requiring lower batching and higher prices.
“If you ask it to, like, generate a web scraper, in the standard one it'll just start outputting code, and it'll be maybe like a four k sequence length.”
Patel explains o1's thinking time creates memory bandwidth issues that prevent batching users at high levels.
“But if you batch higher, k b cache is not just a memory capacity issue, it's also a memory bandwidth issue.”
Gurley notes shifting narrative suggesting inference scaling is preferable to training CapEx.
“There was a podcast recently where they kind of flipped everything on their head and they said, well, if we're not doing that anymore, it's way better because we can just move on to inference, which is getting cheaper and you won't have to spend all this CapEx.”
Patel reveals o1's reasoning process sometimes switches between Chinese and English during hidden thinking phase.
“It generates tons of things. It's like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It's thinking. Right?”
Patel describes reasoning models generating thousands of thinking tokens, sometimes switching languages mid-reasoning.
“It generates tons of things. It's like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It's thinking. Right? It's churning.”
Patel explains reasoning models increase cost ten times by outputting 11,000 tokens versus 1,000 for same query
“I outputted a thousand tokens to I outputted 11,000 tokens. I've 10x'd my spend to generate no. Not the same thing. Right? It's higher quality.”
Patel calculates reasoning models cost fifty times more per query due to batch size and token generation combined
“Cost increase for a single token to be generated is four to five x, but then I'm generating 10 x as many tokens.”
“inference compute is gonna be the the kind of derivative winner of that. We're gonna run out of GPUs, accelerators, compute in 2025 the same way we did in '23.”
“inference compute is gonna be the the kind of derivative winner of that. We're gonna run out of GPUs, accelerators, compute in 2025 the same way we did in '23.”
Baker identifies three axes of AI scaling: pretraining, inference time compute, and now reasoning as the third multiplicative dimension.
“And then we started scaling around inference time compute. And it's very clear that we have now added a third axis of scaling performance, and that is reasoning.”
Patel argues export controls primarily limit AI inference deployment in China, not frontier model training capabilities.
“A large part of export controls, if they work, is just that the amount of AI that can be run-in China is going to be much lower.”
Gerstner says Google's monthly token generation exploded 100x in a year, from 9 trillion to 980 trillion tokens.
“Today, it's 980,000,000,000,000 tokens. So from 9,000,000,000,000 to nine eighty, it's a 100 x increase in a year.”
Patel predicts AWS revenue growth will reaccelerate after consistent deceleration due to Anthropic and new data centers.
“AWS has been decelerating revenue. Year on year revenue has been falling consistently. And and our big call is that it's actually going to start reaccelerating.”
Patel confirms OpenAI runs production inference on GB200 despite reliability requiring workloads handle 64 of 72 GPUs.
“OpenAI has said they're running production inference on GV200 a couple of months ago, in fact. Right?”
Patel says B200 is better for training while GB200 is better for inference, reversing expected use.
“And so you've sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.”
Patel reveals NVIDIA's next generation will split inference into separate context processing and decode workloads, not training versus inference.
“NVIDIA's next generation actually has something very different. They're not saying, Hey, there's a training GPU and an inference GPU, right? Because either is fine.”
Patel explains NVIDIA is splitting inference into decode and prefill workloads, but provisioning for unknown future ratios is challenging.
“what sort of NVIDIA's pitching is like a split of inference into two workloads, and we'll see if they're successful. There's a lot of challenges on the infrastructure side.”
Patel cites DeepSeek's open-sourced inference system requiring 140 GPUs communicating over RDMA networks for single replica.
“One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.”
Patel explains users will pay 10x more for 10x faster inference completion, justifying Cerberus economics.
“for a lot of people, I'm fine to spend 10x the price on something that completes 10x faster. So Cerberus sort of just makes a ton of sense there.”
Huang says AI creates abundance of intelligence by orders of magnitude, compressing year-long work into hours or real-time.
“AI reduces the cost of intelligence or create the abundance of intelligence by orders of magnitude.”
Huang describes PhysicsNEMO as a physics-aware AI framework combining principled simulation with AI prediction.
“PhysicsNEMO is essentially a physics aware AI model simulation system and AI framework that allows us to create these AI models that are either trained by principled simulators or work alongside principled simulators”
Huang claims PhysicsNEMO can predict physics simulations ten thousand times faster than traditional methods.
“so it's grounded in the laws of physics, but able to predict 10,000 times faster.”
Patel claims code agent revenue grew from a couple billion to over $10 billion in a very short time.
“code agent revenue has gone from a couple billion to north of 10,000,000,000 in like a very short amount of time. And these, the horizon of these has also increased dramatically,”
Patel identifies disaggregated PD and wide EP as networking techniques where performance matters significantly.
“When someone has really good networking with these leading edge techniques to run a model such as KIMI or DeepSeq called disaggregated PD or wide EP, these are two different techniques, network performance is a humongous factor.”
Patel notes software dependencies change daily or multiple times weekly across the entire stack.
“And furthermore, with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”
Patel explains AI software stacks update multiple times per week making performance measurement a moving target.
“with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.”
Patel distinguishes OpenAI hiding reasoning chains from Anthropic showing them in code generation models.
“In the case of OpenAI, they don't show you the whole reasoning chain. In the case of Anthropic, they do.”
Patel predicts reasoning chains will close as AI response times extend from immediate to hours or days.
“And then as the horizon of AI models gets longer, right, rather than a question answer immediately, the question answer becomes ten minutes, hours, days, the value of all the reasoning is gone.”
Baker says AI models shifting to usage-based pricing with overage reveals no ceiling on spending yet.
“We're just moving from these all you can eat plans to usage based plans with overage, where those usage tokens cost a lot more, and we're finding out that there's we're nowhere near the amount of, you know, people ceiling price for how much they'll spend.”
Baker emphasizes only 0.1% of the world uses AI models properly yet there's massive shortage despite trillions spent.
“And we're in an insane shortage despite spending cumulatively trillions of dollars. What happens when 5% of the world's population is using these models the way the cutting edge 10 basis points are?”
“What happens when 5% of the world's population is using these models the way the cutting edge 10 basis points are? Like, it's just it's unimaginable. This is why orbital compute is a necessity.”
“Tranium, by far. Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.”
Baker predicts Trainium will dominate 2026 like TPUs did in 2025, with Trainium 3 ramping in second half.
“Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.”
Baker argues Trainium is most underestimated because frontier mixture-of-expert models require switched scale-up networks for inference.
“And so Trainium is for sure the most underestimated, not only because of those design choices, but because the all of these frontier models are what are called mixture of expert models.”
Baker states only NVIDIA and Amazon Trainium have functioning switched scale-up networks for inference today.
“And the only two functioning switched scale up networks in the world today are the ones that power NVIDIA GPUs and Amazon's Trainiums.”
Baker reveals Atreides could have invested over $50 million in CoreWeave at $1.1 billion valuation but was conflicted out.
“I could've Atreides could've invested over $50,000,000 in the round at 1,100,000,000, And I was conflicted out by Crusoe,”
Gerstner argues that token production and consumption is the fundamental basis of all AI intelligence.
“There is no intelligence. There's no consumer chat GPT. There's no enterprise intelligence. There's no clawed code without the production of tokens.”
Baker estimates Anthropic would be doing $100-150B ARR if not compute-constrained, versus current $50B.
“And I think maybe a true statement is that Infantropic could just wave a magic wand and get all the compute they wanted. They'd probably be doing well north of $100,000,000,000 today, maybe 150.”
Baker believes Anthropic is likely already generating cash or will start this year.
“I think Anthropic probably starts generating cash this year if they are not already generating cash, which I think is probably the case.”
Baker argues America will consume all available compute, reducing edge AI bear case concerns.
“And I just think the same is true of compute. It's why I'm probably less worried about like an edge AI bear case than I was.”
Baker argues America will consume all available compute, making him less worried about edge AI bear cases.
“It's why I'm probably less worried about like an edge AI bear case than I was. We're going to consume as much compute as we can.”
Baker explains harness engineering matters significantly and harnesses are increasingly co-developed with models.
“And it turns out that harness engineering is not as important as the model, but it really matters. And these harnesses in these models are increasingly being co developed.”
Baker says understanding frontier AI now requires enterprise usage-based plans, not $250 monthly subscriptions which are rate-limited.
“To understand what Frontier AI is capable of today, even for a non coding use case, need to have Cloud Code or Codex five point Codex.”
Baker says understanding frontier AI now requires enterprise usage-based plans, not consumer subscriptions, due to rate limiting.
“To understand what Frontier AI is capable of today, even for a non coding use case, need to have Cloud Code or Codex five point Codex. And you need to be on an enterprise plan.”
Baker says AI shifting from flat pricing to usage-based is extremely bullish as people consume more AI.
“AI is just shifting from all you can eat to pay by the drink. Then it turns out people really like to talk to their friends long distance.”
Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to shift to usage-based pricing.
“So I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.”
Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to usage-based pricing shift.
“I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.”
Baker explains continual learning as models dynamically updating weights in real-time, unlike current reinforcement learning approaches.
“Continual learning is a model that dynamically adjusts its weights or adjusts in some way in real time.”
Baker predicts GPUs will have 10-15 year useful lives due to prefill-inference disaggregation, extending older chips' value.
“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives.”
Baker predicts GPUs will have 10-15 year useful lives due to prefill-inference disaggregation, extending older chips' value.
“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives.”
Baker argues GPU useful lives will extend to 10-15 years due to inference disaggregation, contradicting AI skeptics.
“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives. The AI skeptics are like, oh, these companies are all cooking their books.”
Baker argues GPU useful lives will extend to 10-15 years due to inference disaggregation, contradicting AI skeptics.
“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives. The AI skeptics are like, oh, these companies are all cooking their books.”
Baker argues GPU useful lives will extend to 10-15 years due to prefill/decode disaggregation, not 1-2 years.
“The useful life of GPU is only a year or two. The useful life of CPU is only four years because the rapid technological change.”
Patel predicts OpenAI and Anthropic will have over 100 gigawatts combined by 2030, terawatts by 2040.
“I think by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you'll add Meta and Google and so on and so forth.”
Patel predicts OpenAI and Anthropic alone will deploy over 100 gigawatts of compute by 2030.
“by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you'll add Meta and Google and so on and so forth.”
Baker explains AI recomputes answers probabilistically each time, enabling superhuman capabilities but requiring extreme computational expense.
“AI, even if you put a harness on it, even if you do the chain of thought, even if you have multiple agents, it's probabilistic that it is recomputing the answer each time.”
Baker calculates rendering Monopoly Go with VEO3 would cost over 100x the game's revenue.
“Monopoly Go is a game where we have the revenue and the hours played and it is to render it using list prices for something like VEO3 is more than two orders of magnitude greater than its revenue. So what is the role for human creativity, man? I don't know.”
Patel describes agentic workflows using 30,000-100,000 input tokens but generating only 1,000 output tokens.
“And and that makes, you know, initially, that'd be like, okay, well now I need a ton ton of compute to calculate all the prefilled tokens.”
Patel explains KV cache storage and reuse enables massive cost decreases in inference.
“you calculate that once, you store it off in memory, whether it be system memory or storage, and then you pull it back in when you run the turn.”
Patel reports cache hit rates above 95% for many agentic workflows in production.
“we're seeing cache hit rates above 95% for many AgenTeq workflows, which which means the cost for a cache hit is it's not free,”
Baker argues contracted compute trades at massive discount to spot, repricing will accelerate cash flows and answer ROI questions.
“And so, essentially, you have the contracted base of installed compute trading at a massive discount to the current spot market.”
Baker reports GPU rental prices doubled from mid-$2 to nearly $4 per hour over seven months for identical clusters.
“And they had rented a cluster of several thousand black wells, and we'll just call it somewhere in the mid $2 per GPU hour.”
Baker reports inference cloud expects to pay 100% more for Blackwells when contracts expire, showing hyperscalers are underearning.
“They went on a podcast, and they essentially said, we are planning to pay 100% more for Blackwell's when our contract expires. And that just means that essentially all the hyperscalers are under earning.”
Baker cites analysis showing compute margins, quantity, and inference margins all rising simultaneously, driving lab acceleration.
“The amount of compute is going up and inference margins going up. And if you multiply those three, that's how you're getting this crazy acceleration into some of the labs plus open source,”
Baker cites analysis showing compute margins, quantity, and inference margins all rising simultaneously, driving lab acceleration.
“The amount of compute is going up and inference margins going up. And if you multiply those three, that's how you're getting this crazy acceleration into some of the labs plus open source,”
Baker estimates SpaceX monetizes compute at $50B per gigawatt versus $73B consensus, with Grok and Cursor hitting $10B ARR quickly.
“And they're monetizing at something like 50,000,000,000 a gig and consensus estimates for next year are 73,000,000,000. So forget Starlink v three, forget Starlink direct to sell, Grok 4.5 and Cursor.”
Patel argues that if model progress pauses while compute supply grows, demand growth will slow and prices will collapse.
“If model progress at the labs pause, then more compute comes online. It has to slow down. Right? Sort of right now we have supply demand, right?”
Patel says anyone can profitably run inference by renting GB300 racks and deploying open models.
“Go download the Kimi weights. Go download VLM or SGLANG. Set it up. You know, Codecs and Fable can actually help you do this.”
Patel argues labs will allocate less compute to inference over time, contrary to consensus belief.
“So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?”