Gavin Baker on

inference

28 quotes · 10 posts · Nov 2019 – Aug 2026

Saidverbatim, newest first

  1. Baker argues contracted compute trades at massive discount to spot, repricing will accelerate cash flows and answer ROI questions.

    “And so, essentially, you have the contracted base of installed compute trading at a massive discount to the current spot market.”

    4:03 · Invest Like the Best · 4 Aug 2026 · permalink
  2. Baker reports GPU rental prices doubled from mid-$2 to nearly $4 per hour over seven months for identical clusters.

    “And they had rented a cluster of several thousand black wells, and we'll just call it somewhere in the mid $2 per GPU hour.”

    14:58 · Invest Like the Best · 4 Aug 2026 · permalink
  3. Baker reports inference cloud expects to pay 100% more for Blackwells when contracts expire, showing hyperscalers are underearning.

    “They went on a podcast, and they essentially said, we are planning to pay 100% more for Blackwell's when our contract expires. And that just means that essentially all the hyperscalers are under earning.”

    15:40 · Invest Like the Best · 4 Aug 2026 · permalink
  4. Baker cites analysis showing compute margins, quantity, and inference margins all rising simultaneously, driving lab acceleration.

    “The amount of compute is going up and inference margins going up. And if you multiply those three, that's how you're getting this crazy acceleration into some of the labs plus open source,”

    40:55 · Invest Like the Best · 4 Aug 2026 · permalink
  5. Baker estimates SpaceX monetizes compute at $50B per gigawatt versus $73B consensus, with Grok and Cursor hitting $10B ARR quickly.

    “And they're monetizing at something like 50,000,000,000 a gig and consensus estimates for next year are 73,000,000,000. So forget Starlink v three, forget Starlink direct to sell, Grok 4.5 and Cursor.”

    57:37 · Invest Like the Best · 4 Aug 2026 · permalink
  6. Baker believes Anthropic is likely already generating cash or will start this year.

    “I think Anthropic probably starts generating cash this year if they are not already generating cash, which I think is probably the case.”

    12:12 · listen · Invest Like the Best · 20 May 2026 · permalink
  7. Baker argues America will consume all available compute, reducing edge AI bear case concerns.

    “And I just think the same is true of compute. It's why I'm probably less worried about like an edge AI bear case than I was.”

    21:24 · listen · Invest Like the Best · 20 May 2026 · permalink
  8. Baker explains harness engineering matters significantly and harnesses are increasingly co-developed with models.

    “And it turns out that harness engineering is not as important as the model, but it really matters. And these harnesses in these models are increasingly being co developed.”

    36:04 · listen · Invest Like the Best · 20 May 2026 · permalink
  9. Baker says understanding frontier AI now requires enterprise usage-based plans, not consumer subscriptions, due to rate limiting.

    “To understand what Frontier AI is capable of today, even for a non coding use case, need to have Cloud Code or Codex five point Codex. And you need to be on an enterprise plan.”

    36:49 · listen · Invest Like the Best · 20 May 2026 · permalink
  10. Baker says AI shifting from flat pricing to usage-based is extremely bullish as people consume more AI.

    “AI is just shifting from all you can eat to pay by the drink. Then it turns out people really like to talk to their friends long distance.”

    38:00 · listen · Invest Like the Best · 20 May 2026 · permalink
  11. Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to shift to usage-based pricing.

    “So I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.”

    38:16 · listen · Invest Like the Best · 20 May 2026 · permalink
  12. Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to usage-based pricing shift.

    “I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.”

    38:16 · listen · Invest Like the Best · 20 May 2026 · permalink
  13. Baker argues GPU useful lives will extend to 10-15 years due to inference disaggregation, contradicting AI skeptics.

    “The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives. The AI skeptics are like, oh, these companies are all cooking their books.”

    47:11 · listen · Invest Like the Best · 20 May 2026 · permalink
  14. Baker predicts GPUs will have 10-15 year useful lives due to prefill-inference disaggregation, extending older chips' value.

    “The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives.”

    47:11 · listen · Invest Like the Best · 20 May 2026 · permalink
  15. Baker argues GPU useful lives will extend to 10-15 years due to prefill/decode disaggregation, not 1-2 years.

    “The useful life of GPU is only a year or two. The useful life of CPU is only four years because the rapid technological change.”

    47:22 · listen · Invest Like the Best · 20 May 2026 · permalink
  16. Baker says AI models shifting to usage-based pricing with overage reveals no ceiling on spending yet.

    “We're just moving from these all you can eat plans to usage based plans with overage, where those usage tokens cost a lot more, and we're finding out that there's we're nowhere near the amount of, you know, people ceiling price for how much they'll spend.”

    14:59 · Sohn Conference Foundation · 15 May 2026 · permalink
  17. Baker emphasizes only 0.1% of the world uses AI models properly yet there's massive shortage despite trillions spent.

    “And we're in an insane shortage despite spending cumulatively trillions of dollars. What happens when 5% of the world's population is using these models the way the cutting edge 10 basis points are?”

    15:32 · Sohn Conference Foundation · 15 May 2026 · permalink
  18. Sohn Conference Foundation

    “What happens when 5% of the world's population is using these models the way the cutting edge 10 basis points are? Like, it's just it's unimaginable. This is why orbital compute is a necessity.”

    15:38 · Sohn Conference Foundation · 15 May 2026 · permalink
  19. Sohn Conference Foundation

    “Tranium, by far. Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.”

    17:03 · Sohn Conference Foundation · 15 May 2026 · permalink
  20. Baker predicts Trainium will dominate 2026 like TPUs did in 2025, with Trainium 3 ramping in second half.

    “Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.”

    17:05 · Sohn Conference Foundation · 15 May 2026 · permalink
  21. Baker argues Trainium is most underestimated because frontier mixture-of-expert models require switched scale-up networks for inference.

    “And so Trainium is for sure the most underestimated, not only because of those design choices, but because the all of these frontier models are what are called mixture of expert models.”

    17:44 · Sohn Conference Foundation · 15 May 2026 · permalink
  22. Baker states only NVIDIA and Amazon Trainium have functioning switched scale-up networks for inference today.

    “And the only two functioning switched scale up networks in the world today are the ones that power NVIDIA GPUs and Amazon's Trainiums.”

    18:04 · Sohn Conference Foundation · 15 May 2026 · permalink
  23. Baker reveals Atreides could have invested over $50 million in CoreWeave at $1.1 billion valuation but was conflicted out.

    “I could've Atreides could've invested over $50,000,000 in the round at 1,100,000,000, And I was conflicted out by Crusoe,”

    19:31 · Sohn Conference Foundation · 15 May 2026 · permalink
  24. Baker identifies three axes of AI scaling: pretraining, inference time compute, and now reasoning as the third multiplicative dimension.

    “And then we started scaling around inference time compute. And it's very clear that we have now added a third axis of scaling performance, and that is reasoning.”

    1:34:40 · listen · All-In Podcast · 4 Jan 2025 · permalink
  25. Baker predicts Tesla FSD will achieve 100x improvement quickly as compute scales to GPT-4.5 levels.

    “I think they're going to go really fast to GPT-4.5 compute, which means you're going to get, using these orders of magnitude, you're going get a 100x improvement really fast.”

    1:02:16 · listen · Invest Like the Best · 27 Aug 2024 · permalink
  26. Baker argues AI revolution stems from cloud computing power and mobile-generated data, not algorithmic advances.

    “The only thing that has enabled the AI revolution that we're living through, which I think we're at the bottom of the first inning in, is one, we had the ability to do cloud computing, so just apply significantly more computational power to old algorithms, and then b, we had dramatically more data.”

    12:55 · listen · Invest Like the Best · 26 Nov 2019 · permalink
  27. Baker states data quantity is the single most predictive element of AI quality, not algorithms or infrastructure.

    “The single most predictive element of knowledge about AI quality is the quantity of data used to train the algorithm.”

    38:57 · listen · Invest Like the Best · 26 Nov 2019 · permalink
  28. Baker cites research showing every 10x increase in training data doubles AI quality.

    “and it's been very well established in multiple papers from both Google and Microsoft research that for every order of magnitude increase in the data you use to train an algorithm, the quality of the AI doubles.”

    39:18 · listen · Invest Like the Best · 26 Nov 2019 · permalink

Postedtheir own words, on X

  1. post Baker says Atreides internal AI spend will be 100x higher in August 2026 versus March 2026 and still doubling monthly

  2. post Baker estimates his personal AI usage is up 100x and calls Grok Bot another Claude Code moment

  3. post Baker says Grok 4.6 matches Fable 5 Max at 85% discount, 80% cheaper input and 88% cheaper output, Pareto dominant

  4. post Baker would not take under on 250 billion for Anthropic unless they stumble or fail to secure compute

  5. post Baker says future is multi-model with specialized open-source models on customer data working with frontier models behind custom router

  6. post Baker credits Jalapeño as first good ASIC outside TPU and Trainium but says disaggregated GPU plus SRAM will outperform

  7. post Baker says Stripe is growing net revenue and FCF almost 2x faster than Adyen with declining share count

  8. post Baker suggests Anthropic may have shifted from gross to net ARR reporting making 65 billion more conservative than 47 billion

Everything Gavin Baker is on record saying · RSS