<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>training: everyone on the record — The Minutes</title>
    <link>https://minutesof.com/on/training/</link>
    <description>Everyone on the record about training: 15 verbatim quotes from 2 people, June 2020 to August 2026, in one chronology, each with a timestamp and a link to…</description>
    <language>en</language>
    <lastBuildDate>Sun, 30 Aug 2026 22:33:18 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/on/training/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Dylan Patel: Patel argues labs will allocate less compute to inference over time, contrary to consensus belief</title>
      <link>https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</link>
      <guid isPermaLink="true">https://minutesof.com/q/469a18e6-8605-46bd-bea0-fa21bd5502b3/</guid>
      <description>“So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?” — Dylan Patel, Dwarkesh Patel</description>
      <pubDate>Tue, 25 Aug 2026 15:57:53 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains agentic training only needs to sync tokens every few minutes versus weights every</title>
      <link>https://minutesof.com/q/86748cbb-055a-4578-92be-361a647d16ac/</link>
      <guid isPermaLink="true">https://minutesof.com/q/86748cbb-055a-4578-92be-361a647d16ac/</guid>
      <description>“When you&#x27;re doing these rollouts and especially as things get more and more agentic and training, you might not only need to send not the entire weights but just the tokens that are relevant.” — Dylan Patel, TBPN</description>
      <pubDate>Tue, 03 Feb 2026 23:08:03 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says B200 is better for training while GB200 is better for inference, reversing expected us</title>
      <link>https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</link>
      <guid isPermaLink="true">https://minutesof.com/q/b34c2dfe-463a-4923-8ba7-41d8b89b38a6/</guid>
      <description>“And so you&#x27;ve sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.” — Dylan Patel, Together AI</description>
      <pubDate>Fri, 03 Oct 2025 20:23:30 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>David Friedberg: Friedberg attended a 24-day intensive Jay Robinson wrestling camp at University of Minnesota</title>
      <link>https://minutesof.com/q/a52409fd-1020-4c4e-98c8-069c694790f2/</link>
      <guid isPermaLink="true">https://minutesof.com/q/a52409fd-1020-4c4e-98c8-069c694790f2/</guid>
      <description>“It was like a twenty four day intensive wrestling camp in Minnesota. I forgot how many days specifically, but it was a Jay Robinson wrestling camp at the University of Minnesota.” — David Friedberg, NPTE Final Frontier</description>
      <pubDate>Fri, 03 Oct 2025 11:51:33 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>David Friedberg: Friedberg credits a 24-day intensive Jay Robinson wrestling camp at University of Minnesota f</title>
      <link>https://minutesof.com/q/6877f97a-1da9-45a3-8085-2cb2668ed870/</link>
      <guid isPermaLink="true">https://minutesof.com/q/6877f97a-1da9-45a3-8085-2cb2668ed870/</guid>
      <description>“It was, like, a twenty four day intensive wrestling camp in Minnesota. I forgot how many days specifically, but it was a Jay Robinson wrestling camp at the University of Minnesota.” — David Friedberg, NPTE Final Frontier</description>
      <pubDate>Thu, 02 Oct 2025 02:28:47 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says GPT-5 will combine massive pre-training like GPT-4.5 with massive post-training like o</title>
      <link>https://minutesof.com/q/808fcaf5-9216-420e-b11b-f8a91aeadae8/</link>
      <guid isPermaLink="true">https://minutesof.com/q/808fcaf5-9216-420e-b11b-f8a91aeadae8/</guid>
      <description>“GPT five, as Sam calls it, is is gonna be a model that has huge pre training scale, right, like GPT 4.5, but also huge post training scale like o one and o three and continuing to scale that up.” — Dylan Patel, Alex Kantrowitz</description>
      <pubDate>Wed, 23 Apr 2025 16:30:06 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel reports that OpenAI&#x27;s Orion training run did not improve enough over GPT-4.5 to qualify as</title>
      <link>https://minutesof.com/q/8373ef31-935a-4e8d-90e3-7c6565db225c/</link>
      <guid isPermaLink="true">https://minutesof.com/q/8373ef31-935a-4e8d-90e3-7c6565db225c/</guid>
      <description>“There were hopes that Orion could be used for for GPT five, but its improvement was, like, not enough to be, like, really a GPT five.” — Dylan Patel, Alex Kantrowitz</description>
      <pubDate>Wed, 23 Apr 2025 16:30:06 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains training a 70 billion parameter model requires passing four times that data every</title>
      <link>https://minutesof.com/q/2bc44e6d-a415-482b-bcb0-f9c59a1b4198/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2bc44e6d-a415-482b-bcb0-f9c59a1b4198/</guid>
      <description>“So if you&#x27;re training like a 70,000,000,000 parameter model, you need to pass four x that number every single training step. And when you&#x27;re passing that, you can&#x27;t really overlap communications and compute.” — Dylan Patel, Prime Intellect AI</description>
      <pubDate>Tue, 25 Mar 2025 18:21:45 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says GPUs in a 16,000 GPU cluster show up to 10% speed variation from manufacturing differe</title>
      <link>https://minutesof.com/q/e4cb87fa-869e-4b13-8b85-bd4d3fb6cb18/</link>
      <guid isPermaLink="true">https://minutesof.com/q/e4cb87fa-869e-4b13-8b85-bd4d3fb6cb18/</guid>
      <description>“But even within a cluster of like 16 k GPUs, GPUs will be slower by up to 10%. Right? So there there&#x27;s a lot of variation even in the manufacturing.” — Dylan Patel, Prime Intellect AI</description>
      <pubDate>Tue, 25 Mar 2025 18:21:45 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel says current models use 100,000 GPUs while next generation will require hundreds of thousan</title>
      <link>https://minutesof.com/q/561bd369-2c88-4ea1-8fa4-3d02389be50d/</link>
      <guid isPermaLink="true">https://minutesof.com/q/561bd369-2c88-4ea1-8fa4-3d02389be50d/</guid>
      <description>“And next generation models that are trained on hundreds of thousands or even millions GPUs, right?” — Dylan Patel, Special Competitive Studies Project</description>
      <pubDate>Thu, 13 Mar 2025 16:42:15 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat ver</title>
      <link>https://minutesof.com/q/02fd4a40-e301-43f7-be83-33801b8010a9/</link>
      <guid isPermaLink="true">https://minutesof.com/q/02fd4a40-e301-43f7-be83-33801b8010a9/</guid>
      <description>“This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it&#x27;s confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.” — Dylan Patel, Lex Fridman</description>
      <pubDate>Mon, 03 Feb 2025 00:12:13 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data e</title>
      <link>https://minutesof.com/q/2b1dd77d-4b29-490c-a1f9-fd32caff427b/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2b1dd77d-4b29-490c-a1f9-fd32caff427b/</guid>
      <description>“They&#x27;re doing it for like multi trillion. Right? So let&#x27;s call it 10,000,000,000,000 parameters, four bytes parameter, that&#x27;s 40 terabytes of data you need to transmit in two seconds.” — Dylan Patel, Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel explains that AdamW optimizer requires four bytes per parameter, creating 400 gigabytes of</title>
      <link>https://minutesof.com/q/9b0a4eda-79e5-4213-b0df-76f45b715526/</link>
      <guid isPermaLink="true">https://minutesof.com/q/9b0a4eda-79e5-4213-b0df-76f45b715526/</guid>
      <description>“when you look at LLMs, people use AdamW. And AdamW as an optimizer is like four bytes per parameter roughly, if I recall correctly the optimizer state.” — Dylan Patel, Scaling Intelligence</description>
      <pubDate>Tue, 12 Nov 2024 04:17:00 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>Dylan Patel: Patel predicts five to seven companies will train GPT-4 scale models in the next year.</title>
      <link>https://minutesof.com/q/fa2d7b01-051f-42f8-ab72-c441507a20aa/</link>
      <guid isPermaLink="true">https://minutesof.com/q/fa2d7b01-051f-42f8-ab72-c441507a20aa/</guid>
      <description>“I do believe that there&#x27;s going to be five to seven companies that will have a GPT-four size model, at least, right, in terms of total flops.” — Dylan Patel, The Inside View</description>
      <pubDate>Wed, 09 Aug 2023 15:00:23 +0000</pubDate>
      <category>training</category>
    </item>
    <item>
      <title>David Friedberg: Friedberg advocates ending qualified immunity and retraining police as a service rather than</title>
      <link>https://minutesof.com/q/c100e8e1-efd4-420f-aa63-79335d6d7771/</link>
      <guid isPermaLink="true">https://minutesof.com/q/c100e8e1-efd4-420f-aa63-79335d6d7771/</guid>
      <description>“I I&#x27;m a huge fan of ending qualified immunity. I think that doesn&#x27;t make any sense. I think we have to stop arming our police like their military.” — David Friedberg, All-In Podcast</description>
      <pubDate>Sat, 20 Jun 2020 01:15:25 +0000</pubDate>
      <category>training</category>
    </item>
  </channel>
</rss>
