<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>inference: everyone on the record — The Minutes</title>
    <link>https://minutesof.com/on/inference/</link>
    <description>Everyone on the record about inference: 75 verbatim quotes from 5 people, November 2019 to August 2026, in one chronology, each with a timestamp and a…</description>
    <language>en</language>
    <lastBuildDate>Sun, 30 Aug 2026 22:33:01 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/on/inference/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Dylan Patel: Patel argues labs will allocate less compute to inference over time, contrary to consensus belief</title>
      <link>https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</link>
      <guid isPermaLink="true">https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</guid>
      <description>“So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?” — Dylan Patel, Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says anyone can profitably run inference by renting GB300 racks and deploying open models.</title>
      <link>https://minutesof.com/q/fefa3320-9068-4e30-84da-c70370affef0/</link>
      <guid isPermaLink="true">https://minutesof.com/q/fefa3320-9068-4e30-84da-c70370affef0/</guid>
      <description>“Go download the Kimi weights. Go download VLM or SGLANG. Set it up. You know, Codecs and Fable can actually help you do this.” — Dylan Patel, Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel argues that if model progress pauses while compute supply grows, demand growth will slow an</title>
      <link>https://minutesof.com/q/ccbeb5bf-7c6b-4af0-ae7d-f4c729f341ed/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ccbeb5bf-7c6b-4af0-ae7d-f4c729f341ed/</guid>
      <description>“If model progress at the labs pause, then more compute comes online. It has to slow down. Right? Sort of right now we have supply demand, right?” — Dylan Patel, SemiAnalysis</description>
      <pubDate>Mon, 17 Aug 2026 15:00:06 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker estimates SpaceX monetizes compute at $50B per gigawatt versus $73B consensus, with Grok an</title>
      <link>https://minutesof.com/q/898427d9-5d8c-40af-9feb-9256d6c3c56e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/898427d9-5d8c-40af-9feb-9256d6c3c56e/</guid>
      <description>“And they&#x27;re monetizing at something like 50,000,000,000 a gig and consensus estimates for next year are 73,000,000,000. So forget Starlink v three, forget Starlink direct to sell, Grok 4.5 and Cursor.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker cites analysis showing compute margins, quantity, and inference margins all rising simultan</title>
      <link>https://minutesof.com/q/273fa3f1-13d3-4f0d-9774-ac10ff1c4250/</link>
      <guid isPermaLink="true">https://minutesof.com/q/273fa3f1-13d3-4f0d-9774-ac10ff1c4250/</guid>
      <description>“The amount of compute is going up and inference margins going up. And if you multiply those three, that&#x27;s how you&#x27;re getting this crazy acceleration into some of the labs plus open source,” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker cites analysis showing compute margins, quantity, and inference margins all rising simultan</title>
      <link>https://minutesof.com/q/273fa3f1-13d3-4f0d-9774-ac10ff1c4250/</link>
      <guid isPermaLink="true">https://minutesof.com/q/273fa3f1-13d3-4f0d-9774-ac10ff1c4250/</guid>
      <description>“The amount of compute is going up and inference margins going up. And if you multiply those three, that&#x27;s how you&#x27;re getting this crazy acceleration into some of the labs plus open source,” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker reports inference cloud expects to pay 100% more for Blackwells when contracts expire, show</title>
      <link>https://minutesof.com/q/443a6457-4940-4dff-9022-a9659a14fe37/</link>
      <guid isPermaLink="true">https://minutesof.com/q/443a6457-4940-4dff-9022-a9659a14fe37/</guid>
      <description>“They went on a podcast, and they essentially said, we are planning to pay 100% more for Blackwell&#x27;s when our contract expires. And that just means that essentially all the hyperscalers are under earning.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker reports GPU rental prices doubled from mid-$2 to nearly $4 per hour over seven months for i</title>
      <link>https://minutesof.com/q/2f00fe00-eb15-4fb6-83b8-ad9eee2d466d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2f00fe00-eb15-4fb6-83b8-ad9eee2d466d/</guid>
      <description>“And they had rented a cluster of several thousand black wells, and we&#x27;ll just call it somewhere in the mid $2 per GPU hour.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues contracted compute trades at massive discount to spot, repricing will accelerate cas</title>
      <link>https://minutesof.com/q/c85f4400-5d03-4327-a817-cfa56c88e0a0/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c85f4400-5d03-4327-a817-cfa56c88e0a0/</guid>
      <description>“And so, essentially, you have the contracted base of installed compute trading at a massive discount to the current spot market.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 04 Aug 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel reports cache hit rates above 95% for many agentic workflows in production.</title>
      <link>https://minutesof.com/q/12c60899-9cef-4341-915a-6afdeefc7d8d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/12c60899-9cef-4341-915a-6afdeefc7d8d/</guid>
      <description>“we&#x27;re seeing cache hit rates above 95% for many AgenTeq workflows, which which means the cost for a cache hit is it&#x27;s not free,” — Dylan Patel, RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains KV cache storage and reuse enables massive cost decreases in inference.</title>
      <link>https://minutesof.com/q/3bea5703-d869-4e93-bec6-70e3ecd20e14/</link>
      <guid isPermaLink="true">https://minutesof.com/q/3bea5703-d869-4e93-bec6-70e3ecd20e14/</guid>
      <description>“you calculate that once, you store it off in memory, whether it be system memory or storage, and then you pull it back in when you run the turn.” — Dylan Patel, RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel describes agentic workflows using 30,000-100,000 input tokens but generating only 1,000 out</title>
      <link>https://minutesof.com/q/a3014961-4f96-48f2-8af1-0b0ad506d194/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a3014961-4f96-48f2-8af1-0b0ad506d194/</guid>
      <description>“And and that makes, you know, initially, that&#x27;d be like, okay, well now I need a ton ton of compute to calculate all the prefilled tokens.” — Dylan Patel, RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker calculates rendering Monopoly Go with VEO3 would cost over 100x the game&#x27;s revenue.</title>
      <link>https://minutesof.com/q/92c8f11c-3e80-4ba8-9041-57c9ce86cb86/</link>
      <guid isPermaLink="true">https://minutesof.com/q/92c8f11c-3e80-4ba8-9041-57c9ce86cb86/</guid>
      <description>“Monopoly Go is a game where we have the revenue and the hours played and it is to render it using list prices for something like VEO3 is more than two orders of magnitude greater than its revenue. So what is the role for human creativity, man? I don&#x27;t know.” — Gavin Baker, Generating Alpha Podcast</description>
      <pubDate>Wed, 08 Jul 2026 15:54:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker explains AI recomputes answers probabilistically each time, enabling superhuman capabilitie</title>
      <link>https://minutesof.com/q/f9d8be14-a3e0-4d46-93f0-a74b86319592/</link>
      <guid isPermaLink="true">https://minutesof.com/q/f9d8be14-a3e0-4d46-93f0-a74b86319592/</guid>
      <description>“AI, even if you put a harness on it, even if you do the chain of thought, even if you have multiple agents, it&#x27;s probabilistic that it is recomputing the answer each time.” — Gavin Baker, Generating Alpha Podcast</description>
      <pubDate>Wed, 08 Jul 2026 15:54:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel predicts OpenAI and Anthropic alone will deploy over 100 gigawatts of compute by 2030.</title>
      <link>https://minutesof.com/q/e9a8fba3-1f79-4310-9c95-0d513f20374b/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e9a8fba3-1f79-4310-9c95-0d513f20374b/</guid>
      <description>“by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you&#x27;ll add Meta and Google and so on and so forth.” — Dylan Patel, Sequoia Capital</description>
      <pubDate>Tue, 30 Jun 2026 12:00:25 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel predicts OpenAI and Anthropic will have over 100 gigawatts combined by 2030, terawatts by 2</title>
      <link>https://minutesof.com/q/dd349bfc-a119-4399-bf3c-9a26efbef20a/</link>
      <guid isPermaLink="true">https://minutesof.com/q/dd349bfc-a119-4399-bf3c-9a26efbef20a/</guid>
      <description>“I think by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you&#x27;ll add Meta and Google and so on and so forth.” — Dylan Patel, Sequoia Capital</description>
      <pubDate>Tue, 30 Jun 2026 12:00:25 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues GPU useful lives will extend to 10-15 years due to prefill/decode disaggregation, no</title>
      <link>https://minutesof.com/q/d97e5674-e5aa-447d-9755-f683cbf66ab1/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d97e5674-e5aa-447d-9755-f683cbf66ab1/</guid>
      <description>“The useful life of GPU is only a year or two. The useful life of CPU is only four years because the rapid technological change.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues GPU useful lives will extend to 10-15 years due to inference disaggregation, contrad</title>
      <link>https://minutesof.com/q/d42a1b86-60fc-431a-8a6b-892fc8405b51/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d42a1b86-60fc-431a-8a6b-892fc8405b51/</guid>
      <description>“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives. The AI skeptics are like, oh, these companies are all cooking their books.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues GPU useful lives will extend to 10-15 years due to inference disaggregation, contrad</title>
      <link>https://minutesof.com/q/d42a1b86-60fc-431a-8a6b-892fc8405b51/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d42a1b86-60fc-431a-8a6b-892fc8405b51/</guid>
      <description>“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives. The AI skeptics are like, oh, these companies are all cooking their books.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts GPUs will have 10-15 year useful lives due to prefill-inference disaggregation, ex</title>
      <link>https://minutesof.com/q/ca1f41f5-e05d-4029-a1db-504f8b4e433d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ca1f41f5-e05d-4029-a1db-504f8b4e433d/</guid>
      <description>“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts GPUs will have 10-15 year useful lives due to prefill-inference disaggregation, ex</title>
      <link>https://minutesof.com/q/ca1f41f5-e05d-4029-a1db-504f8b4e433d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ca1f41f5-e05d-4029-a1db-504f8b4e433d/</guid>
      <description>“The disaggregation of inference means that I think these GPUs are going to have ten or fifteen year lives.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker explains continual learning as models dynamically updating weights in real-time, unlike cur</title>
      <link>https://minutesof.com/q/b32b5e45-ab6f-4748-bcad-62727d94ff2f/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b32b5e45-ab6f-4748-bcad-62727d94ff2f/</guid>
      <description>“Continual learning is a model that dynamically adjusts its weights or adjusts in some way in real time.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to usage-based pricing sh</title>
      <link>https://minutesof.com/q/e54157aa-cef1-4bd4-9d22-a8b4df8c0d13/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e54157aa-cef1-4bd4-9d22-a8b4df8c0d13/</guid>
      <description>“I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts OpenAI and Anthropic will exceed $200B ARR this year due to shift to usage-based p</title>
      <link>https://minutesof.com/q/60e611b0-5da1-46dd-88cc-3277d267722f/</link>
      <guid isPermaLink="true">https://minutesof.com/q/60e611b0-5da1-46dd-88cc-3277d267722f/</guid>
      <description>“So I think the shift to usage based pricing is probably why you will see OpenAI and Anthropic exceed well over $200,000,000,000 in ARR this year.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker says AI shifting from flat pricing to usage-based is extremely bullish as people consume mo</title>
      <link>https://minutesof.com/q/4b52026f-d927-4a34-90b9-14245a9cc937/</link>
      <guid isPermaLink="true">https://minutesof.com/q/4b52026f-d927-4a34-90b9-14245a9cc937/</guid>
      <description>“AI is just shifting from all you can eat to pay by the drink. Then it turns out people really like to talk to their friends long distance.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker says understanding frontier AI now requires enterprise usage-based plans, not consumer subs</title>
      <link>https://minutesof.com/q/7b49081d-b3f9-445d-aba4-d290d5374ebd/</link>
      <guid isPermaLink="true">https://minutesof.com/q/7b49081d-b3f9-445d-aba4-d290d5374ebd/</guid>
      <description>“To understand what Frontier AI is capable of today, even for a non coding use case, need to have Cloud Code or Codex five point Codex. And you need to be on an enterprise plan.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker says understanding frontier AI now requires enterprise usage-based plans, not $250 monthly</title>
      <link>https://minutesof.com/q/42fa040e-9f02-4acf-911f-dfbc9a85a96b/</link>
      <guid isPermaLink="true">https://minutesof.com/q/42fa040e-9f02-4acf-911f-dfbc9a85a96b/</guid>
      <description>“To understand what Frontier AI is capable of today, even for a non coding use case, need to have Cloud Code or Codex five point Codex.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker explains harness engineering matters significantly and harnesses are increasingly co-develo</title>
      <link>https://minutesof.com/q/7991d993-a149-4684-8692-12d7dc932f1e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/7991d993-a149-4684-8692-12d7dc932f1e/</guid>
      <description>“And it turns out that harness engineering is not as important as the model, but it really matters. And these harnesses in these models are increasingly being co developed.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues America will consume all available compute, making him less worried about edge AI be</title>
      <link>https://minutesof.com/q/c36e87b2-cd1e-4c19-b84d-b929d673346d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c36e87b2-cd1e-4c19-b84d-b929d673346d/</guid>
      <description>“It&#x27;s why I&#x27;m probably less worried about like an edge AI bear case than I was. We&#x27;re going to consume as much compute as we can.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues America will consume all available compute, reducing edge AI bear case concerns.</title>
      <link>https://minutesof.com/q/c36d1b7b-a557-4ecb-b855-6610596f9fe3/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c36d1b7b-a557-4ecb-b855-6610596f9fe3/</guid>
      <description>“And I just think the same is true of compute. It&#x27;s why I&#x27;m probably less worried about like an edge AI bear case than I was.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker believes Anthropic is likely already generating cash or will start this year.</title>
      <link>https://minutesof.com/q/ec5a9c7e-99d9-4536-8535-fbba88740844/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ec5a9c7e-99d9-4536-8535-fbba88740844/</guid>
      <description>“I think Anthropic probably starts generating cash this year if they are not already generating cash, which I think is probably the case.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker estimates Anthropic would be doing $100-150B ARR if not compute-constrained, versus current</title>
      <link>https://minutesof.com/q/52541184-dd18-482e-b266-0236fe4216b1/</link>
      <guid isPermaLink="true">https://minutesof.com/q/52541184-dd18-482e-b266-0236fe4216b1/</guid>
      <description>“And I think maybe a true statement is that Infantropic could just wave a magic wand and get all the compute they wanted. They&#x27;d probably be doing well north of $100,000,000,000 today, maybe 150.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Wed, 20 May 2026 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Brad Gerstner: Gerstner argues that token production and consumption is the fundamental basis of all AI intell</title>
      <link>https://minutesof.com/q/108fa4fc-95f6-46b6-b83f-019d063ae52f/</link>
      <guid isPermaLink="true">https://minutesof.com/q/108fa4fc-95f6-46b6-b83f-019d063ae52f/</guid>
      <description>“There is no intelligence. There&#x27;s no consumer chat GPT. There&#x27;s no enterprise intelligence. There&#x27;s no clawed code without the production of tokens.” — Brad Gerstner, Altimeter Capital</description>
      <pubDate>Fri, 15 May 2026 20:31:59 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker reveals Atreides could have invested over $50 million in CoreWeave at $1.1 billion valuatio</title>
      <link>https://minutesof.com/q/4bbf9d84-bae2-460e-a752-a60a13663534/</link>
      <guid isPermaLink="true">https://minutesof.com/q/4bbf9d84-bae2-460e-a752-a60a13663534/</guid>
      <description>“I could&#x27;ve Atreides could&#x27;ve invested over $50,000,000 in the round at 1,100,000,000, And I was conflicted out by Crusoe,” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker states only NVIDIA and Amazon Trainium have functioning switched scale-up networks for infe</title>
      <link>https://minutesof.com/q/05badf52-27d8-4742-8532-91f442850ac7/</link>
      <guid isPermaLink="true">https://minutesof.com/q/05badf52-27d8-4742-8532-91f442850ac7/</guid>
      <description>“And the only two functioning switched scale up networks in the world today are the ones that power NVIDIA GPUs and Amazon&#x27;s Trainiums.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues Trainium is most underestimated because frontier mixture-of-expert models require sw</title>
      <link>https://minutesof.com/q/9f8eb50e-f98e-49c0-8051-f5a028e01a39/</link>
      <guid isPermaLink="true">https://minutesof.com/q/9f8eb50e-f98e-49c0-8051-f5a028e01a39/</guid>
      <description>“And so Trainium is for sure the most underestimated, not only because of those design choices, but because the all of these frontier models are what are called mixture of expert models.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts Trainium will dominate 2026 like TPUs did in 2025, with Trainium 3 ramping in seco</title>
      <link>https://minutesof.com/q/77488a06-75d9-432d-b0cd-0afa0e21f41f/</link>
      <guid isPermaLink="true">https://minutesof.com/q/77488a06-75d9-432d-b0cd-0afa0e21f41f/</guid>
      <description>“Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Tranium, by far. Tranium is going to be to 2026, especially in the second half of this year when</title>
      <link>https://minutesof.com/q/a91f962e-6711-454f-b1db-5c6d4437360d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a91f962e-6711-454f-b1db-5c6d4437360d/</guid>
      <description>“Tranium, by far. Tranium is going to be to 2026, especially in the second half of this year when Tranium three really ramps, as TPUs were to twenty twenty five.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: What happens when 5% of the world&#x27;s population is using these models the way the cutting edge 10</title>
      <link>https://minutesof.com/q/7c0eb714-2699-491f-b983-09b25778e0f7/</link>
      <guid isPermaLink="true">https://minutesof.com/q/7c0eb714-2699-491f-b983-09b25778e0f7/</guid>
      <description>“What happens when 5% of the world&#x27;s population is using these models the way the cutting edge 10 basis points are? Like, it&#x27;s just it&#x27;s unimaginable. This is why orbital compute is a necessity.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker emphasizes only 0.1% of the world uses AI models properly yet there&#x27;s massive shortage desp</title>
      <link>https://minutesof.com/q/20c46b63-3663-4f9f-9324-17a82821dfcb/</link>
      <guid isPermaLink="true">https://minutesof.com/q/20c46b63-3663-4f9f-9324-17a82821dfcb/</guid>
      <description>“And we&#x27;re in an insane shortage despite spending cumulatively trillions of dollars. What happens when 5% of the world&#x27;s population is using these models the way the cutting edge 10 basis points are?” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker says AI models shifting to usage-based pricing with overage reveals no ceiling on spending</title>
      <link>https://minutesof.com/q/eb16479a-ee2b-4d2c-9779-1e7eb5762367/</link>
      <guid isPermaLink="true">https://minutesof.com/q/eb16479a-ee2b-4d2c-9779-1e7eb5762367/</guid>
      <description>“We&#x27;re just moving from these all you can eat plans to usage based plans with overage, where those usage tokens cost a lot more, and we&#x27;re finding out that there&#x27;s we&#x27;re nowhere near the amount of, you know, people ceiling price for how much they&#x27;ll spend.” — Gavin Baker, Sohn Conference Foundation</description>
      <pubDate>Fri, 15 May 2026 18:17:41 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel predicts reasoning chains will close as AI response times extend from immediate to hours or</title>
      <link>https://minutesof.com/q/44c831c5-138e-45c3-8b54-9ad0278248ff/</link>
      <guid isPermaLink="true">https://minutesof.com/q/44c831c5-138e-45c3-8b54-9ad0278248ff/</guid>
      <description>“And then as the horizon of AI models gets longer, right, rather than a question answer immediately, the question answer becomes ten minutes, hours, days, the value of all the reasoning is gone.” — Dylan Patel, Swole as a Service</description>
      <pubDate>Wed, 13 May 2026 20:00:36 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel distinguishes OpenAI hiding reasoning chains from Anthropic showing them in code generation</title>
      <link>https://minutesof.com/q/a9706075-6705-4eb4-ad0e-013b2f735206/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a9706075-6705-4eb4-ad0e-013b2f735206/</guid>
      <description>“In the case of OpenAI, they don&#x27;t show you the whole reasoning chain. In the case of Anthropic, they do.” — Dylan Patel, Swole as a Service</description>
      <pubDate>Wed, 13 May 2026 20:00:36 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains AI software stacks update multiple times per week making performance measurement a</title>
      <link>https://minutesof.com/q/bf34d090-4f22-4978-bd08-307b1566b305/</link>
      <guid isPermaLink="true">https://minutesof.com/q/bf34d090-4f22-4978-bd08-307b1566b305/</guid>
      <description>“with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.” — Dylan Patel, TensorWave</description>
      <pubDate>Thu, 30 Apr 2026 18:18:14 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel notes software dependencies change daily or multiple times weekly across the entire stack.</title>
      <link>https://minutesof.com/q/7d67261b-caa3-4beb-8c61-2b0c52ec53da/</link>
      <guid isPermaLink="true">https://minutesof.com/q/7d67261b-caa3-4beb-8c61-2b0c52ec53da/</guid>
      <description>“And furthermore, with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.” — Dylan Patel, TensorWave</description>
      <pubDate>Thu, 30 Apr 2026 18:18:14 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel identifies disaggregated PD and wide EP as networking techniques where performance matters</title>
      <link>https://minutesof.com/q/a8d7ee01-6bf0-4ca4-828e-77222cf7a242/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a8d7ee01-6bf0-4ca4-828e-77222cf7a242/</guid>
      <description>“When someone has really good networking with these leading edge techniques to run a model such as KIMI or DeepSeq called disaggregated PD or wide EP, these are two different techniques, network performance is a humongous factor.” — Dylan Patel, Aria Networks</description>
      <pubDate>Thu, 16 Apr 2026 02:27:54 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel claims code agent revenue grew from a couple billion to over $10 billion in a very short ti</title>
      <link>https://minutesof.com/q/eae8cac0-9d21-4f54-97aa-ef5eae6fa378/</link>
      <guid isPermaLink="true">https://minutesof.com/q/eae8cac0-9d21-4f54-97aa-ef5eae6fa378/</guid>
      <description>“code agent revenue has gone from a couple billion to north of 10,000,000,000 in like a very short amount of time. And these, the horizon of these has also increased dramatically,” — Dylan Patel, Daytona</description>
      <pubDate>Tue, 07 Apr 2026 17:12:37 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Jensen Huang: Huang claims PhysicsNEMO can predict physics simulations ten thousand times faster than traditio</title>
      <link>https://minutesof.com/q/efa4c687-8b33-4fdc-88bc-eb9d63667ddd/</link>
      <guid isPermaLink="true">https://minutesof.com/q/efa4c687-8b33-4fdc-88bc-eb9d63667ddd/</guid>
      <description>“so it&#x27;s grounded in the laws of physics, but able to predict 10,000 times faster.” — Jensen Huang, NVIDIA</description>
      <pubDate>Thu, 05 Feb 2026 02:53:33 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Jensen Huang: Huang describes PhysicsNEMO as a physics-aware AI framework combining principled simulation with</title>
      <link>https://minutesof.com/q/eb1ea83d-8c91-4ef7-8dde-aedbbff7c2b4/</link>
      <guid isPermaLink="true">https://minutesof.com/q/eb1ea83d-8c91-4ef7-8dde-aedbbff7c2b4/</guid>
      <description>“PhysicsNEMO is essentially a physics aware AI model simulation system and AI framework that allows us to create these AI models that are either trained by principled simulators or work alongside principled simulators” — Jensen Huang, NVIDIA</description>
      <pubDate>Thu, 05 Feb 2026 02:53:33 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Jensen Huang: Huang says AI creates abundance of intelligence by orders of magnitude, compressing year-long wo</title>
      <link>https://minutesof.com/q/4f81d78a-5e9e-4884-9e86-06863a537aee/</link>
      <guid isPermaLink="true">https://minutesof.com/q/4f81d78a-5e9e-4884-9e86-06863a537aee/</guid>
      <description>“AI reduces the cost of intelligence or create the abundance of intelligence by orders of magnitude.” — Jensen Huang, Cisco</description>
      <pubDate>Wed, 04 Feb 2026 04:34:47 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains users will pay 10x more for 10x faster inference completion, justifying Cerberus e</title>
      <link>https://minutesof.com/q/d63342d7-794d-43b2-8529-800ca943aa5a/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d63342d7-794d-43b2-8529-800ca943aa5a/</guid>
      <description>“for a lot of people, I&#x27;m fine to spend 10x the price on something that completes 10x faster. So Cerberus sort of just makes a ton of sense there.” — Dylan Patel, TBPN</description>
      <pubDate>Tue, 03 Feb 2026 23:08:03 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel cites DeepSeek&#x27;s open-sourced inference system requiring 140 GPUs communicating over RDMA n</title>
      <link>https://minutesof.com/q/b5b15eca-a046-4a24-8309-0cfa9d92ceec/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b5b15eca-a046-4a24-8309-0cfa9d92ceec/</guid>
      <description>“One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.” — Dylan Patel, Clockwork</description>
      <pubDate>Fri, 21 Nov 2025 17:18:51 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains NVIDIA is splitting inference into decode and prefill workloads, but provisioning</title>
      <link>https://minutesof.com/q/e37df648-a474-40f3-9e9c-b531f05a918e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e37df648-a474-40f3-9e9c-b531f05a918e/</guid>
      <description>“what sort of NVIDIA&#x27;s pitching is like a split of inference into two workloads, and we&#x27;ll see if they&#x27;re successful. There&#x27;s a lot of challenges on the infrastructure side.” — Dylan Patel, Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel reveals NVIDIA&#x27;s next generation will split inference into separate context processing and</title>
      <link>https://minutesof.com/q/8896a5c5-5fee-4e45-86fb-6379c1c3eec2/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8896a5c5-5fee-4e45-86fb-6379c1c3eec2/</guid>
      <description>“NVIDIA&#x27;s next generation actually has something very different. They&#x27;re not saying, Hey, there&#x27;s a training GPU and an inference GPU, right? Because either is fine.” — Dylan Patel, Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says B200 is better for training while GB200 is better for inference, reversing expected us</title>
      <link>https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</guid>
      <description>“And so you&#x27;ve sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.” — Dylan Patel, Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel confirms OpenAI runs production inference on GB200 despite reliability requiring workloads</title>
      <link>https://minutesof.com/q/63137c12-d683-47eb-9bf7-16977addf85c/</link>
      <guid isPermaLink="true">https://minutesof.com/q/63137c12-d683-47eb-9bf7-16977addf85c/</guid>
      <description>“OpenAI has said they&#x27;re running production inference on GV200 a couple of months ago, in fact. Right?” — Dylan Patel, Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel predicts AWS revenue growth will reaccelerate after consistent deceleration due to Anthropi</title>
      <link>https://minutesof.com/q/62b39bc8-1e44-4f9d-b9b5-d25bb37cda66/</link>
      <guid isPermaLink="true">https://minutesof.com/q/62b39bc8-1e44-4f9d-b9b5-d25bb37cda66/</guid>
      <description>“AWS has been decelerating revenue. Year on year revenue has been falling consistently. And and our big call is that it&#x27;s actually going to start reaccelerating.” — Dylan Patel, a16z</description>
      <pubDate>Mon, 22 Sep 2025 13:01:51 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Brad Gerstner: Gerstner says Google&#x27;s monthly token generation exploded 100x in a year, from 9 trillion to 980</title>
      <link>https://minutesof.com/q/c5f66780-d7e1-4b21-9d13-444df206cd01/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c5f66780-d7e1-4b21-9d13-444df206cd01/</guid>
      <description>“Today, it&#x27;s 980,000,000,000,000 tokens. So from 9,000,000,000,000 to nine eighty, it&#x27;s a 100 x increase in a year.” — Brad Gerstner, CNBC Television</description>
      <pubDate>Thu, 28 Aug 2025 17:12:24 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel argues export controls primarily limit AI inference deployment in China, not frontier model</title>
      <link>https://minutesof.com/q/9e102380-dbd9-4dbb-acd8-a43420fa4c7f/</link>
      <guid isPermaLink="true">https://minutesof.com/q/9e102380-dbd9-4dbb-acd8-a43420fa4c7f/</guid>
      <description>“A large part of export controls, if they work, is just that the amount of AI that can be run-in China is going to be much lower.” — Dylan Patel, Lex Fridman</description>
      <pubDate>Mon, 03 Feb 2025 00:12:13 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker identifies three axes of AI scaling: pretraining, inference time compute, and now reasoning</title>
      <link>https://minutesof.com/q/5f490d92-e1c8-4877-b369-6b470be69d65/</link>
      <guid isPermaLink="true">https://minutesof.com/q/5f490d92-e1c8-4877-b369-6b470be69d65/</guid>
      <description>“And then we started scaling around inference time compute. And it&#x27;s very clear that we have now added a third axis of scaling performance, and that is reasoning.” — Gavin Baker, All-In Podcast</description>
      <pubDate>Sat, 04 Jan 2025 00:23:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: inference compute is gonna be the the kind of derivative winner of that. We&#x27;re gonna run out of G</title>
      <link>https://minutesof.com/q/b30df38d-a975-4d8d-9f97-4b9eeeed3b27/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b30df38d-a975-4d8d-9f97-4b9eeeed3b27/</guid>
      <description>“inference compute is gonna be the the kind of derivative winner of that. We&#x27;re gonna run out of GPUs, accelerators, compute in 2025 the same way we did in &#x27;23.” — Gavin Baker, All-In Podcast</description>
      <pubDate>Sat, 04 Jan 2025 00:23:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: inference compute is gonna be the the kind of derivative winner of that. We&#x27;re gonna run out of G</title>
      <link>https://minutesof.com/q/b30df38d-a975-4d8d-9f97-4b9eeeed3b27/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b30df38d-a975-4d8d-9f97-4b9eeeed3b27/</guid>
      <description>“inference compute is gonna be the the kind of derivative winner of that. We&#x27;re gonna run out of GPUs, accelerators, compute in 2025 the same way we did in &#x27;23.” — Gavin Baker, All-In Podcast</description>
      <pubDate>Sat, 04 Jan 2025 00:23:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Bill Gurley: Patel calculates reasoning models cost fifty times more per query due to batch size and token gen</title>
      <link>https://minutesof.com/q/3af95131-05a1-45d4-ae2d-9578bd46190d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/3af95131-05a1-45d4-ae2d-9578bd46190d/</guid>
      <description>“Cost increase for a single token to be generated is four to five x, but then I&#x27;m generating 10 x as many tokens.” — Bill Gurley, Bg2 Pod</description>
      <pubDate>Mon, 23 Dec 2024 20:25:20 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Bill Gurley: Patel explains reasoning models increase cost ten times by outputting 11,000 tokens versus 1,000</title>
      <link>https://minutesof.com/q/9aed22e0-c898-4e8e-9a74-966679c754b1/</link>
      <guid isPermaLink="true">https://minutesof.com/q/9aed22e0-c898-4e8e-9a74-966679c754b1/</guid>
      <description>“I outputted a thousand tokens to I outputted 11,000 tokens. I&#x27;ve 10x&#x27;d my spend to generate no. Not the same thing. Right? It&#x27;s higher quality.” — Bill Gurley, Bg2 Pod</description>
      <pubDate>Mon, 23 Dec 2024 20:25:20 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Bill Gurley: Patel describes reasoning models generating thousands of thinking tokens, sometimes switching lan</title>
      <link>https://minutesof.com/q/476ea39f-bfb6-495d-9a59-81ea69699b0d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/476ea39f-bfb6-495d-9a59-81ea69699b0d/</guid>
      <description>“It generates tons of things. It&#x27;s like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It&#x27;s thinking. Right? It&#x27;s churning.” — Bill Gurley, Bg2 Pod</description>
      <pubDate>Mon, 23 Dec 2024 20:25:20 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel reveals o1&#x27;s reasoning process sometimes switches between Chinese and English during hidden</title>
      <link>https://minutesof.com/q/1524c47b-e0e1-402a-b743-111da2d0e959/</link>
      <guid isPermaLink="true">https://minutesof.com/q/1524c47b-e0e1-402a-b743-111da2d0e959/</guid>
      <description>“It generates tons of things. It&#x27;s like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It&#x27;s thinking. Right?” — Dylan Patel, BG2 Pod</description>
      <pubDate>Mon, 23 Dec 2024 20:12:56 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Bill Gurley: Gurley notes shifting narrative suggesting inference scaling is preferable to training CapEx.</title>
      <link>https://minutesof.com/q/168ba19f-f3d8-48fd-bbbd-632f2cb46217/</link>
      <guid isPermaLink="true">https://minutesof.com/q/168ba19f-f3d8-48fd-bbbd-632f2cb46217/</guid>
      <description>“There was a podcast recently where they kind of flipped everything on their head and they said, well, if we&#x27;re not doing that anymore, it&#x27;s way better because we can just move on to inference, which is getting cheaper and you won&#x27;t have to spend all this CapEx.” — Bill Gurley, Bg2 Pod</description>
      <pubDate>Thu, 12 Dec 2024 18:20:19 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains o1&#x27;s thinking time creates memory bandwidth issues that prevent batching users at</title>
      <link>https://minutesof.com/q/b7e7fd8f-2e86-4705-b53e-c249725aab84/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b7e7fd8f-2e86-4705-b53e-c249725aab84/</guid>
      <description>“But if you batch higher, k b cache is not just a memory capacity issue, it&#x27;s also a memory bandwidth issue.” — Dylan Patel, Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says o1 generates 40k sequence lengths versus 4k for standard models, requiring lower batch</title>
      <link>https://minutesof.com/q/2548d082-f15b-4fca-b6d4-4cfc177ee7e5/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2548d082-f15b-4fca-b6d4-4cfc177ee7e5/</guid>
      <description>“If you ask it to, like, generate a web scraper, in the standard one it&#x27;ll just start outputting code, and it&#x27;ll be maybe like a four k sequence length.” — Dylan Patel, Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Brad Gerstner: Gerstner reports Jensen predicts inference will scale 100x to billion-x with 40% of NVIDIA reve</title>
      <link>https://minutesof.com/q/50e70e24-1d30-4bef-bf58-58b1d0263aa5/</link>
      <guid isPermaLink="true">https://minutesof.com/q/50e70e24-1d30-4bef-bf58-58b1d0263aa5/</guid>
      <description>“he said as a consequence of that, inference is going to a 100 x, thousand x, a million x, maybe even a billion x.” — Brad Gerstner, BG2 Pod</description>
      <pubDate>Sun, 13 Oct 2024 10:35:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Bill Gurley: Gurley explains Strawberry requires 10X processing for linear improvement, meaning 10-100X infere</title>
      <link>https://minutesof.com/q/27a00324-83ae-472e-8e7b-d4900e14122e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/27a00324-83ae-472e-8e7b-d4900e14122e/</guid>
      <description>“in order to get linear improvement, you have to do maybe 10X the amount of processing. And this is all inference. So what are the implications of that?” — Bill Gurley, BG2 Pod</description>
      <pubDate>Wed, 25 Sep 2024 16:21:38 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker predicts Tesla FSD will achieve 100x improvement quickly as compute scales to GPT-4.5 level</title>
      <link>https://minutesof.com/q/a913c721-311e-4f5d-a524-73f6cda49ae1/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a913c721-311e-4f5d-a524-73f6cda49ae1/</guid>
      <description>“I think they&#x27;re going to go really fast to GPT-4.5 compute, which means you&#x27;re going to get, using these orders of magnitude, you&#x27;re going get a 100x improvement really fast.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 27 Aug 2024 08:00:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker cites research showing every 10x increase in training data doubles AI quality.</title>
      <link>https://minutesof.com/q/1c159533-7002-43df-a5db-be3d8cc9d824/</link>
      <guid isPermaLink="true">https://minutesof.com/q/1c159533-7002-43df-a5db-be3d8cc9d824/</guid>
      <description>“and it&#x27;s been very well established in multiple papers from both Google and Microsoft research that for every order of magnitude increase in the data you use to train an algorithm, the quality of the AI doubles.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 26 Nov 2019 10:30:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker states data quantity is the single most predictive element of AI quality, not algorithms or</title>
      <link>https://minutesof.com/q/32e64b43-db5d-48ee-815c-272156cf2971/</link>
      <guid isPermaLink="true">https://minutesof.com/q/32e64b43-db5d-48ee-815c-272156cf2971/</guid>
      <description>“The single most predictive element of knowledge about AI quality is the quantity of data used to train the algorithm.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 26 Nov 2019 10:30:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Gavin Baker: Baker argues AI revolution stems from cloud computing power and mobile-generated data, not algori</title>
      <link>https://minutesof.com/q/c6581532-cbbe-425f-bb38-995439fb7bb4/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c6581532-cbbe-425f-bb38-995439fb7bb4/</guid>
      <description>“The only thing that has enabled the AI revolution that we&#x27;re living through, which I think we&#x27;re at the bottom of the first inning in, is one, we had the ability to do cloud computing, so just apply significantly more computational power to old algorithms, and then b, we had dramatically more data.” — Gavin Baker, Invest Like the Best</description>
      <pubDate>Tue, 26 Nov 2019 10:30:00 +0000</pubDate>
      <category>inference</category>
    </item>
  </channel>
</rss>
