<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The Minutes of Dylan Patel on gpu utilization</title>
    <link>https://minutesof.com/dylan-patel/on/gpu-utilization/</link>
    <description>Everything Dylan Patel has said on gpu utilization: 2 verbatim quotes between December 2023 and July 2026, each with a timestamp and a link to the…</description>
    <language>en</language>
    <lastBuildDate>Mon, 31 Aug 2026 13:10:13 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/dylan-patel/on/gpu-utilization/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Patel argues prefill caching drastically raises GPU utilization by eliminating repeated context recalculation</title>
      <link>https://minutesof.com/q/1c63392d-c6b8-44b1-8b80-f8d850c63669/</link>
      <guid isPermaLink="true">https://minutesof.com/q/1c63392d-c6b8-44b1-8b80-f8d850c63669/</guid>
      <description>“Now that means your GPU utilization rises drastically, and the amount of time that the GPUs are generating new tokens is far, far higher than the amount of times that GPUs are generating or recalculating the context, which is not necessarily providing any value to your operations.” — RAISE Summit</description>
      <pubDate>Thu, 16 Jul 2026 16:44:57 +0000</pubDate>
      <category>gpu utilization</category>
    </item>
    <item>
      <title>Patel explains model bandwidth utilization is the critical metric for inference, unlike training where MFU mat</title>
      <link>https://minutesof.com/q/815c0b87-b8a9-42b6-bfc0-c6e3c82677fb/</link>
      <guid isPermaLink="true">https://minutesof.com/q/815c0b87-b8a9-42b6-bfc0-c6e3c82677fb/</guid>
      <description>“But on inference, it&#x27;s not being talked about much, but model MBU, model bandwidth utilization is the important factor.” — Latent Space</description>
      <pubDate>Tue, 05 Dec 2023 18:16:28 +0000</pubDate>
      <category>gpu utilization</category>
    </item>
  </channel>
</rss>
