<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The Minutes of Dylan Patel on inference economics</title>
    <link>https://minutesof.com/dylan-patel/on/inference-economics/</link>
    <description>Everything Dylan Patel has said on inference economics: 30 verbatim quotes between December 2023 and August 2026, each with a timestamp and a link to the…</description>
    <language>en</language>
    <lastBuildDate>Sun, 30 Aug 2026 16:22:16 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/dylan-patel/on/inference-economics/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Patel says Anthropic generates $50 million per megawatt revenue, enabling 5x return on inference spending.</title>
      <link>https://minutesof.com/q/6e3f2d04-6524-4dca-bc2d-340a73646ca2/</link>
      <guid isPermaLink="true">https://minutesof.com/q/6e3f2d04-6524-4dca-bc2d-340a73646ca2/</guid>
      <description>“In the case of Anthropic, the the revenue has gone as high as $50,000,000 per megawatt.” — Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel reveals Anthropic turned first profit in Q2 and will hit billion-dollar operating profit in Q3 before IP</title>
      <link>https://minutesof.com/q/ab25a556-53cd-4e2e-ba7e-2336d5d78b64/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ab25a556-53cd-4e2e-ba7e-2336d5d78b64/</guid>
      <description>“So Anthropic turned their first gross profit in Q2 in June. And then in Q3, they will be turning a billion dollars of operating profit, slightly over.” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel reports OpenAI gross margins rose from 30% to 55% overall, 50% to 65% excluding free users.</title>
      <link>https://minutesof.com/q/552929f1-c8ae-43c3-8501-0ab7970989ce/</link>
      <guid isPermaLink="true">https://minutesof.com/q/552929f1-c8ae-43c3-8501-0ab7970989ce/</guid>
      <description>“You look at OpenAI late last year, their margins had were roughly 30% gross margin, but if you stripped away the free users, they were at 50%.” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel says their AI spending equals over a third of employee compensation for their 90-person firm.</title>
      <link>https://minutesof.com/q/6f7b1cde-3481-4cd5-89a6-c38b8939b845/</link>
      <guid isPermaLink="true">https://minutesof.com/q/6f7b1cde-3481-4cd5-89a6-c38b8939b845/</guid>
      <description>“We&#x27;re spending, you know, more like, you know, more than a third of the spend, you know, employee spend, a third of it is also on top of that is AI.” — WisdomTree in Europe</description>
      <pubDate>Thu, 09 Jul 2026 06:16:39 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel argues Mythos fast mode is probably cheaper than Claude 4.6 fast mode for most tasks due to token effici</title>
      <link>https://minutesof.com/q/9d03e1ee-9f3e-449f-bbd7-ef96c09864e7/</link>
      <guid isPermaLink="true">https://minutesof.com/q/9d03e1ee-9f3e-449f-bbd7-ef96c09864e7/</guid>
      <description>“the flip side is, is, is mythos is more token efficient. Methos fast mode is probably cheaper than like, four, six fast mode for most tasks.” — SemiAnalysis Weekly</description>
      <pubDate>Wed, 06 May 2026 00:55:06 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel cites Anthropic adding $67 billion ARR per month as evidence of expanding AI revenue beyond hyperscalers</title>
      <link>https://minutesof.com/q/8d242968-225b-41c7-86f6-4f27678a4907/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8d242968-225b-41c7-86f6-4f27678a4907/</guid>
      <description>“You&#x27;re starting to see it with Anthropix revenue adding $67,000,000,000 of ARR a month, but there&#x27;s so many more firms coming online with revenue streaming in,” — CNBC Television</description>
      <pubDate>Mon, 16 Mar 2026 21:58:55 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel states Anthropic now adds $23 billion revenue monthly versus few hundred million earlier.</title>
      <link>https://minutesof.com/q/a9df3580-18d1-4ac3-9f9d-afcb758ea6d7/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a9df3580-18d1-4ac3-9f9d-afcb758ea6d7/</guid>
      <description>“Anthropics, you know, adding $23,000,000,000 of revenue a month now Mhmm. Versus they were just adding a few 100,000,000 of revenue a month earlier. So clearly, we&#x27;re in the take off period.” — Latent Space</description>
      <pubDate>Thu, 26 Feb 2026 21:15:18 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel reports Claude Code doubled to 4% of GitHub commits in January alone, with overall AI coding likely at 1</title>
      <link>https://minutesof.com/q/d7ab6afb-29cc-4e7c-9868-7f4eb011caa2/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d7ab6afb-29cc-4e7c-9868-7f4eb011caa2/</guid>
      <description>“Just in this month just in January, it went from 4% of or 2% of commits on GitHub to 4% of GitHub commits were done by Cloud Code.” — Latent Space</description>
      <pubDate>Thu, 26 Feb 2026 21:15:18 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel explains prefill costs one-fourth of decode per token, making 30,000 token inputs costlier than 2-4,000</title>
      <link>https://minutesof.com/q/187d2021-25f1-4e4f-8559-763189933ea6/</link>
      <guid isPermaLink="true">https://minutesof.com/q/187d2021-25f1-4e4f-8559-763189933ea6/</guid>
      <description>“And that&#x27;s a very common ratio. Right? And then when you think about, okay, the cost of running pre fill is roughly one fourth of running decode.” — Clockwork</description>
      <pubDate>Fri, 21 Nov 2025 17:18:51 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel states avoiding redundant prefill through KV cache can cut inference costs to one-fourth of previous lev</title>
      <link>https://minutesof.com/q/c93c6e47-fe3d-405d-b28b-2d99a4919dee/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c93c6e47-fe3d-405d-b28b-2d99a4919dee/</guid>
      <description>“So you can cut your cost to one fourth of what it was previously if you just don&#x27;t do the pre fill. Right?” — Clockwork</description>
      <pubDate>Fri, 21 Nov 2025 17:18:51 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel says inference providers sell public endpoints at flat or negative margins, compensating through private</title>
      <link>https://minutesof.com/q/a5a36fd6-ef64-421d-871a-a04382e1453c/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a5a36fd6-ef64-421d-871a-a04382e1453c/</guid>
      <description>“Most of the inference providers are selling at flat margins or even negative for their public endpoints. And they then make it up when people do private deployments.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:45 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel calculates 20% power efficiency difference translates to only 4% TCO difference on NVIDIA deployments.</title>
      <link>https://minutesof.com/q/e54db0c9-c69e-4b1e-a270-5018b18ad51e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e54db0c9-c69e-4b1e-a270-5018b18ad51e/</guid>
      <description>“Because if you have enough power, a 20% difference in performance per watt only ends up being a 4% difference in TCO.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:45 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel says the standard inference deployment unit has shifted from single nodes to hundreds of GPUs.</title>
      <link>https://minutesof.com/q/ee180f12-a142-4bbb-8e53-e032a9458215/</link>
      <guid isPermaLink="true">https://minutesof.com/q/ee180f12-a142-4bbb-8e53-e032a9458215/</guid>
      <description>“the standard unit for an inference deployment being hundreds of GPUs instead of a single node. And then there&#x27;s all these different things about traffic.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel notes NVIDIA Blackwell performance improved dramatically from launch to present, as did AMD hardware.</title>
      <link>https://minutesof.com/q/8a048f38-2c8d-4690-a1d2-66e094aaa6fb/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8a048f38-2c8d-4690-a1d2-66e094aaa6fb/</guid>
      <description>“If you tried to use NVIDIA&#x27;s Blackwell six months ago, the numbers were not amazing, right? But now they&#x27;re actually amazing. So how did that progress over time? Same with AMD, right?” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel breaks down inference TCO as 20% electricity and data center, 80% hardware on NVIDIA deployments.</title>
      <link>https://minutesof.com/q/3358ce5d-fe0f-496b-a2f4-13edb43515e9/</link>
      <guid isPermaLink="true">https://minutesof.com/q/3358ce5d-fe0f-496b-a2f4-13edb43515e9/</guid>
      <description>“20% of your cost is your electricity, your data center real estate, roughly. And then the rest of the cost is that hardware, at least on a standard NVIDIA deployment your GPUs, your networking, etcetera.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel identifies GB200&#x27;s power efficiency advantage but notes deployment challenges with backplane and liquid</title>
      <link>https://minutesof.com/q/d7dccc2a-7013-4972-b5c9-372d97cae003/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d7dccc2a-7013-4972-b5c9-372d97cae003/</guid>
      <description>“GB200 has a huge power efficiency advantage, right? Everyone here understands the challenges of running and deploying GB200. There&#x27;s a lot of challenges with the backplane.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel reports GB200 is 10x more power efficient than H200 at certain interactivity rates.</title>
      <link>https://minutesof.com/q/70c03321-ad13-4824-91b8-9ec2ebd3b7be/</link>
      <guid isPermaLink="true">https://minutesof.com/q/70c03321-ad13-4824-91b8-9ec2ebd3b7be/</guid>
      <description>“But it turns out at certain interactivity rates, I. E. Tokens per second per user, it&#x27;s 10x more efficient per watt, right, compared to H200.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel finds AMD MI355 beats NVIDIA B200 on performance TCO in certain publicly usable configurations.</title>
      <link>https://minutesof.com/q/b9346a92-e689-4ff1-a8aa-a519c318c01e/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b9346a92-e689-4ff1-a8aa-a519c318c01e/</guid>
      <description>“So we do different scenarios. We do document processing, which is 8,000 context in, 1,000 out. We do chat, which is 1,000 in, 1,000 out.” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel shows B200 has 15x raw performance advantage over H100 but only 10x performance per TCO.</title>
      <link>https://minutesof.com/q/101225bc-685d-4ce2-a64b-7e2df33b1f98/</link>
      <guid isPermaLink="true">https://minutesof.com/q/101225bc-685d-4ce2-a64b-7e2df33b1f98/</guid>
      <description>“If we don&#x27;t divide by TCO, then it looks like the performance of B200 is actually 15x that of H100, versus the performance TCO is only 10x,” — Open Compute Project</description>
      <pubDate>Thu, 23 Oct 2025 06:01:19 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel states Supermicro&#x27;s liquid cooling reduces Hopper server power from 10 kilowatts to 7-8 kilowatts versus</title>
      <link>https://minutesof.com/q/98989f27-8eb3-4a61-b86b-0dd09799deee/</link>
      <guid isPermaLink="true">https://minutesof.com/q/98989f27-8eb3-4a61-b86b-0dd09799deee/</guid>
      <description>“Everyone else&#x27;s hopper servers h 100 air cooled. And so each server consumes, you know, 10 kilowatts almost. Right?” — Supermicro</description>
      <pubDate>Thu, 16 Oct 2025 17:00:47 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel calculates that without Supermicro&#x27;s liquid cooling solution, servers would consume 30% more power.</title>
      <link>https://minutesof.com/q/10a69170-989c-42e0-9bd5-da11329eee97/</link>
      <guid isPermaLink="true">https://minutesof.com/q/10a69170-989c-42e0-9bd5-da11329eee97/</guid>
      <description>“So that&#x27;s why Shibbol Micro dedicated deep cooling so aggressively. And if the liquid cooling solution from Super Micro didn&#x27;t exist, then those servers would be consuming 30% more power.” — Supermicro</description>
      <pubDate>Thu, 16 Oct 2025 17:00:47 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel describes diagnostic challenges in AI data centers where failures occur across multiple infrastructure l</title>
      <link>https://minutesof.com/q/b7a7f145-e89d-4c61-b6f7-3bea9c8248c6/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b7a7f145-e89d-4c61-b6f7-3bea9c8248c6/</guid>
      <description>“And people have had problems where their data center wasn&#x27;t ready and something stopped working. They don&#x27;t know, is it the chip? No.” — Supermicro</description>
      <pubDate>Thu, 16 Oct 2025 17:00:47 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel explains OpenAI&#x27;s router will send low-value queries to cheaper models but use expensive compute for hig</title>
      <link>https://minutesof.com/q/0502aa27-d821-4087-8f4b-18a954c32235/</link>
      <guid isPermaLink="true">https://minutesof.com/q/0502aa27-d821-4087-8f4b-18a954c32235/</guid>
      <description>“if the user asks a low value query like, hey, why is the sky blue? Just route them to mini. The model can answer perfectly fine, and that is a chunk of queries.” — a16z</description>
      <pubDate>Mon, 18 Aug 2025 18:26:43 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel says AI model costs dropped 1200x from GPT-3&#x27;s $60 per million outputs to current pricing.</title>
      <link>https://minutesof.com/q/41d2ba76-f8de-4943-bca0-4e696568c0ea/</link>
      <guid isPermaLink="true">https://minutesof.com/q/41d2ba76-f8de-4943-bca0-4e696568c0ea/</guid>
      <description>“GPT three was it cost, you know, it cost $60 for the million output. And over time, that kept reducing. Right? OpenAI released new models.” — MedBricks Webcast</description>
      <pubDate>Thu, 27 Mar 2025 13:55:08 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel explains reasoning models dramatically increase memory usage and reduce batch size, multiplying serving</title>
      <link>https://minutesof.com/q/fb69a564-812e-47e6-9f3d-e6f80b0878ca/</link>
      <guid isPermaLink="true">https://minutesof.com/q/fb69a564-812e-47e6-9f3d-e6f80b0878ca/</guid>
      <description>“So your your memory usage is going way up with these reasoning models, and you still have a lot of users.” — Lex Fridman</description>
      <pubDate>Mon, 03 Feb 2025 00:12:13 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel explains reasoning models dramatically increase serving costs due to long output context and memory cons</title>
      <link>https://minutesof.com/q/8ea24247-feff-4e7f-beec-571f6109d01d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8ea24247-feff-4e7f-beec-571f6109d01d/</guid>
      <description>“your memory usage is going way up with these reasoning models, and you still have a lot of users. So effectively, the cost to serve multiplies by a ton.” — Lex Fridman</description>
      <pubDate>Mon, 03 Feb 2025 00:12:13 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel estimates Blackwell delivers 10-15x cost improvement for inference despite NVIDIA claiming 30x at GTC.</title>
      <link>https://minutesof.com/q/879abc84-3fdc-4dd7-9b63-baf6f6541b7c/</link>
      <guid isPermaLink="true">https://minutesof.com/q/879abc84-3fdc-4dd7-9b63-baf6f6541b7c/</guid>
      <description>“But now, like, Blackwell, NVIDIA&#x27;s pitching 10 to 15 x improvement in cost. It&#x27;s like, well, you know, they&#x27;re massaging the numbers marketing.” — Unsupervised Learning: With Jacob Effron</description>
      <pubDate>Tue, 21 Jan 2025 14:00:12 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel explains OpenAI&#x27;s o1 model thinks for 5-20 seconds before outputting, creating inference throughput chal</title>
      <link>https://minutesof.com/q/179104b9-35d2-44c2-9d8a-85d97e3bd8f8/</link>
      <guid isPermaLink="true">https://minutesof.com/q/179104b9-35d2-44c2-9d8a-85d97e3bd8f8/</guid>
      <description>“When you use OpenAI&#x27;s o one, it thinks for ten seconds, twenty seconds, five seconds. It varies a lot, but thinks for a while and then it sends you tokens.” — Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel reports H100 rental prices have dropped from $3-4 per hour to $2.15 or less, approaching natural cost of</title>
      <link>https://minutesof.com/q/e377715e-4aba-45ba-9b19-f37de850e568/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e377715e-4aba-45ba-9b19-f37de850e568/</guid>
      <description>“An hour. Right? For shorter term or midterm deals. Right now, it&#x27;s like, if you want a six month deal, you could get, like, $2.15 or less.” — Dwarkesh Patel</description>
      <pubDate>Wed, 02 Oct 2024 14:59:36 +0000</pubDate>
      <category>inference economics</category>
    </item>
    <item>
      <title>Patel argues AI software will have lower R&amp;D costs but much higher cost of goods sold from operating services.</title>
      <link>https://minutesof.com/q/cdca4ea2-53cc-42c9-a237-c2927fb44bb3/</link>
      <guid isPermaLink="true">https://minutesof.com/q/cdca4ea2-53cc-42c9-a237-c2927fb44bb3/</guid>
      <description>“Like the R and D cost is much lower in terms of people, but the cost of goods sold in terms of actually operating the service, I think will be much higher.” — Latent Space</description>
      <pubDate>Tue, 05 Dec 2023 18:16:28 +0000</pubDate>
      <category>inference economics</category>
    </item>
  </channel>
</rss>
