<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>The Minutes of Dylan Patel on model architecture</title>
    <link>https://minutesof.com/dylan-patel/on/model-architecture/</link>
    <description>Everything Dylan Patel has said on model architecture: 4 verbatim quotes between February 2025 and June 2026, each with a timestamp and a link to the…</description>
    <language>en</language>
    <lastBuildDate>Sun, 30 Aug 2026 16:22:19 +0000</lastBuildDate>
    <atom:link href="https://minutesof.com/dylan-patel/on/model-architecture/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Patel explains how MOE routing works and how some of Meta&#x27;s experts weren&#x27;t being used at all.</title>
      <link>https://minutesof.com/q/d8c531cb-5cc2-4ab8-b1ae-b23a834b6b94/</link>
      <guid isPermaLink="true">https://minutesof.com/q/d8c531cb-5cc2-4ab8-b1ae-b23a834b6b94/</guid>
      <description>“And and each expert learns its own independent things and it&#x27;s like really not something observable by people. But what you can see is tokens when which experts do they route to?” — Matthew Berman</description>
      <pubDate>Mon, 30 Jun 2025 17:27:19 +0000</pubDate>
      <category>model architecture</category>
    </item>
    <item>
      <title>Patel says GPT-5 will be first model to massively scale both pre-training and post-training simultaneously.</title>
      <link>https://minutesof.com/q/00545820-383f-45e6-ac01-d8e99c981b60/</link>
      <guid isPermaLink="true">https://minutesof.com/q/00545820-383f-45e6-ac01-d8e99c981b60/</guid>
      <description>“Right? And this would be the first time we see a model that was was a step up in both at the same time.” — Alex Kantrowitz</description>
      <pubDate>Wed, 23 Apr 2025 16:30:06 +0000</pubDate>
      <category>model architecture</category>
    </item>
    <item>
      <title>Patel explains training a 70 billion parameter model requires passing four times that data every training step</title>
      <link>https://minutesof.com/q/2bc44e6d-a415-482b-bcb0-f9c59a1b4198/</link>
      <guid isPermaLink="true">https://minutesof.com/q/2bc44e6d-a415-482b-bcb0-f9c59a1b4198/</guid>
      <description>“So if you&#x27;re training like a 70,000,000,000 parameter model, you need to pass four x that number every single training step. And when you&#x27;re passing that, you can&#x27;t really overlap communications and compute.” — Prime Intellect AI</description>
      <pubDate>Tue, 25 Mar 2025 18:21:45 +0000</pubDate>
      <category>model architecture</category>
    </item>
    <item>
      <title>Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning</title>
      <link>https://minutesof.com/q/02fd4a40-e301-43f7-be83-33801b8010a9/</link>
      <guid isPermaLink="true">https://minutesof.com/q/02fd4a40-e301-43f7-be83-33801b8010a9/</guid>
      <description>“This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it&#x27;s confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.” — Lex Fridman</description>
      <pubDate>Mon, 03 Feb 2025 00:12:13 +0000</pubDate>
      <category>model architecture</category>
    </item>
  </channel>
</rss>
