Dylan Patel on

model architecture

4 quotes · Feb 2025 – Jun 2026

Saidverbatim, newest first

  1. Patel explains how MOE routing works and how some of Meta's experts weren't being used at all.

    “And and each expert learns its own independent things and it's like really not something observable by people. But what you can see is tokens when which experts do they route to?”

    2:45 · Matthew Berman · 30 Jun 2025 · permalink
  2. Patel says GPT-5 will be first model to massively scale both pre-training and post-training simultaneously.

    “Right? And this would be the first time we see a model that was was a step up in both at the same time.”

    37:03 · Alex Kantrowitz · 23 Apr 2025 · permalink
  3. Patel explains training a 70 billion parameter model requires passing four times that data every training step without overlapping communication.

    “So if you're training like a 70,000,000,000 parameter model, you need to pass four x that number every single training step. And when you're passing that, you can't really overlap communications and compute.”

    6:12 · Prime Intellect AI · 25 Mar 2025 · permalink
  4. Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning models.

    “This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it's confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.”

    4:38 · Lex Fridman · 3 Feb 2025 · permalink

Everything Dylan Patel is on record saying · RSS