<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The Minutes of Dylan Patel on inference</title>
    <link>https://minutesof.com/dylan-patel/on/inference/</link>
    <description>Everything Dylan Patel has said on inference: 21 verbatim quotes between November 2024 and August 2026, each with a timestamp and a link to the recording…</description>
    <language>en</language>
    <lastBuildDate>Sun, 30 Aug 2026 16:22:16 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/dylan-patel/on/inference/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Patel says anyone can profitably run inference by renting GB300 racks and deploying open models.</title>
      <link>https://minutesof.com/q/fefa3320-9068-4e30-84da-c70370affef0/</link>
      <guid isPermaLink="true">https://minutesof.com/q/fefa3320-9068-4e30-84da-c70370affef0/</guid>
      <description>“Go download the Kimi weights. Go download VLM or SGLANG. Set it up. You know, Codecs and Fable can actually help you do this.” — Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel argues labs will allocate less compute to inference over time, contrary to consensus belief.</title>
      <link>https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</link>
      <guid isPermaLink="true">https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</guid>
      <description>“So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?” — Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel argues that if model progress pauses while compute supply grows, demand growth will slow and prices will</title>
      <link>https://minutesof.com/q/ccbeb5bf-7c6b-4af0-ae7d-f4c729f341ed/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ccbeb5bf-7c6b-4af0-ae7d-f4c729f341ed/</guid>
      <description>“If model progress at the labs pause, then more compute comes online. It has to slow down. Right? Sort of right now we have supply demand, right?” — SemiAnalysis</description>
      <pubDate>Mon, 17 Aug 2026 15:00:06 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel describes agentic workflows using 30,000-100,000 input tokens but generating only 1,000 output tokens.</title>
      <link>https://minutesof.com/q/a3014961-4f96-48f2-8af1-0b0ad506d194/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a3014961-4f96-48f2-8af1-0b0ad506d194/</guid>
      <description>“And and that makes, you know, initially, that&#x27;d be like, okay, well now I need a ton ton of compute to calculate all the prefilled tokens.” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel explains KV cache storage and reuse enables massive cost decreases in inference.</title>
      <link>https://minutesof.com/q/3bea5703-d869-4e93-bec6-70e3ecd20e14/</link>
      <guid isPermaLink="true">https://minutesof.com/q/3bea5703-d869-4e93-bec6-70e3ecd20e14/</guid>
      <description>“you calculate that once, you store it off in memory, whether it be system memory or storage, and then you pull it back in when you run the turn.” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel reports cache hit rates above 95% for many agentic workflows in production.</title>
      <link>https://minutesof.com/q/12c60899-9cef-4341-915a-6afdeefc7d8d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/12c60899-9cef-4341-915a-6afdeefc7d8d/</guid>
      <description>“we&#x27;re seeing cache hit rates above 95% for many AgenTeq workflows, which which means the cost for a cache hit is it&#x27;s not free,” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel predicts OpenAI and Anthropic will have over 100 gigawatts combined by 2030, terawatts by 2040.</title>
      <link>https://minutesof.com/q/dd349bfc-a119-4399-bf3c-9a26efbef20a/</link>
      <guid isPermaLink="true">https://minutesof.com/q/dd349bfc-a119-4399-bf3c-9a26efbef20a/</guid>
      <description>“I think by 2030, just OpenAI and Anthropic will have over 100 gigawatts combined, and then you&#x27;ll add Meta and Google and so on and so forth.” — Sequoia Capital</description>
      <pubDate>Tue, 30 Jun 2026 12:00:25 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel distinguishes OpenAI hiding reasoning chains from Anthropic showing them in code generation models.</title>
      <link>https://minutesof.com/q/a9706075-6705-4eb4-ad0e-013b2f735206/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a9706075-6705-4eb4-ad0e-013b2f735206/</guid>
      <description>“In the case of OpenAI, they don&#x27;t show you the whole reasoning chain. In the case of Anthropic, they do.” — Swole as a Service</description>
      <pubDate>Wed, 13 May 2026 20:00:36 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel predicts reasoning chains will close as AI response times extend from immediate to hours or days.</title>
      <link>https://minutesof.com/q/44c831c5-138e-45c3-8b54-9ad0278248ff/</link>
      <guid isPermaLink="true">https://minutesof.com/q/44c831c5-138e-45c3-8b54-9ad0278248ff/</guid>
      <description>“And then as the horizon of AI models gets longer, right, rather than a question answer immediately, the question answer becomes ten minutes, hours, days, the value of all the reasoning is gone.” — Swole as a Service</description>
      <pubDate>Wed, 13 May 2026 20:00:36 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel notes software dependencies change daily or multiple times weekly across the entire stack.</title>
      <link>https://minutesof.com/q/7d67261b-caa3-4beb-8c61-2b0c52ec53da/</link>
      <guid isPermaLink="true">https://minutesof.com/q/7d67261b-caa3-4beb-8c61-2b0c52ec53da/</guid>
      <description>“And furthermore, with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.” — TensorWave</description>
      <pubDate>Thu, 30 Apr 2026 18:18:14 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel explains AI software stacks update multiple times per week making performance measurement a moving targe</title>
      <link>https://minutesof.com/q/bf34d090-4f22-4978-bd08-307b1566b305/</link>
      <guid isPermaLink="true">https://minutesof.com/q/bf34d090-4f22-4978-bd08-307b1566b305/</guid>
      <description>“with software changing literally multiple times a week, right, PyTorch has nightlies, VLM has nightlies, Asteeling has nightlies, CUDA drivers update constantly. You just go You go through the whole list.” — TensorWave</description>
      <pubDate>Thu, 30 Apr 2026 18:18:14 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel identifies disaggregated PD and wide EP as networking techniques where performance matters significantly</title>
      <link>https://minutesof.com/q/a8d7ee01-6bf0-4ca4-828e-77222cf7a242/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a8d7ee01-6bf0-4ca4-828e-77222cf7a242/</guid>
      <description>“When someone has really good networking with these leading edge techniques to run a model such as KIMI or DeepSeq called disaggregated PD or wide EP, these are two different techniques, network performance is a humongous factor.” — Aria Networks</description>
      <pubDate>Thu, 16 Apr 2026 02:27:54 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel claims code agent revenue grew from a couple billion to over $10 billion in a very short time.</title>
      <link>https://minutesof.com/q/eae8cac0-9d21-4f54-97aa-ef5eae6fa378/</link>
      <guid isPermaLink="true">https://minutesof.com/q/eae8cac0-9d21-4f54-97aa-ef5eae6fa378/</guid>
      <description>“code agent revenue has gone from a couple billion to north of 10,000,000,000 in like a very short amount of time. And these, the horizon of these has also increased dramatically,” — Daytona</description>
      <pubDate>Tue, 07 Apr 2026 17:12:37 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel cites DeepSeek&#x27;s open-sourced inference system requiring 140 GPUs communicating over RDMA networks for s</title>
      <link>https://minutesof.com/q/b5b15eca-a046-4a24-8309-0cfa9d92ceec/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b5b15eca-a046-4a24-8309-0cfa9d92ceec/</guid>
      <description>“One example is one that DeepSeek open sourced over December of last yearJanuary, February of this year, where a single replica of inference for a single model is going to be like 140 GPUs.” — Clockwork</description>
      <pubDate>Fri, 21 Nov 2025 17:18:51 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel confirms OpenAI runs production inference on GB200 despite reliability requiring workloads handle 64 of</title>
      <link>https://minutesof.com/q/63137c12-d683-47eb-9bf7-16977addf85c/</link>
      <guid isPermaLink="true">https://minutesof.com/q/63137c12-d683-47eb-9bf7-16977addf85c/</guid>
      <description>“OpenAI has said they&#x27;re running production inference on GV200 a couple of months ago, in fact. Right?” — Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel says B200 is better for training while GB200 is better for inference, reversing expected use.</title>
      <link>https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</guid>
      <description>“And so you&#x27;ve sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.” — Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel reveals NVIDIA&#x27;s next generation will split inference into separate context processing and decode worklo</title>
      <link>https://minutesof.com/q/8896a5c5-5fee-4e45-86fb-6379c1c3eec2/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8896a5c5-5fee-4e45-86fb-6379c1c3eec2/</guid>
      <description>“NVIDIA&#x27;s next generation actually has something very different. They&#x27;re not saying, Hey, there&#x27;s a training GPU and an inference GPU, right? Because either is fine.” — Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel explains NVIDIA is splitting inference into decode and prefill workloads, but provisioning for unknown f</title>
      <link>https://minutesof.com/q/e37df648-a474-40f3-9e9c-b531f05a918e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e37df648-a474-40f3-9e9c-b531f05a918e/</guid>
      <description>“what sort of NVIDIA&#x27;s pitching is like a split of inference into two workloads, and we&#x27;ll see if they&#x27;re successful. There&#x27;s a lot of challenges on the infrastructure side.” — Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel reveals o1&#x27;s reasoning process sometimes switches between Chinese and English during hidden thinking pha</title>
      <link>https://minutesof.com/q/1524c47b-e0e1-402a-b743-111da2d0e959/</link>
      <guid isPermaLink="true">https://minutesof.com/q/1524c47b-e0e1-402a-b743-111da2d0e959/</guid>
      <description>“It generates tons of things. It&#x27;s like it it sometimes switches between Chinese and English. Right? Like, whatever it is. It&#x27;s thinking. Right?” — BG2 Pod</description>
      <pubDate>Mon, 23 Dec 2024 20:12:56 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel says o1 generates 40k sequence lengths versus 4k for standard models, requiring lower batching and highe</title>
      <link>https://minutesof.com/q/2548d082-f15b-4fca-b6d4-4cfc177ee7e5/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2548d082-f15b-4fca-b6d4-4cfc177ee7e5/</guid>
      <description>“If you ask it to, like, generate a web scraper, in the standard one it&#x27;ll just start outputting code, and it&#x27;ll be maybe like a four k sequence length.” — Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>inference</category>
    </item>
    <item>
      <title>Patel explains o1&#x27;s thinking time creates memory bandwidth issues that prevent batching users at high levels.</title>
      <link>https://minutesof.com/q/b7e7fd8f-2e86-4705-b53e-c249725aab84/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b7e7fd8f-2e86-4705-b53e-c249725aab84/</guid>
      <description>“But if you batch higher, k b cache is not just a memory capacity issue, it&#x27;s also a memory bandwidth issue.” — Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>inference</category>
    </item>
  </channel>
</rss>
