On the record about

training

2 people · 15 quotes · 20 Jun 2020 to 25 Aug 2026

Who is on this subjectordered by the date of their first quote here

1 of 2 lane rests on fewer than 5 quotes and is marked thin. Offsets are days from the middle first-quote date, 20 Jun 2020 — a date, and nothing else. It is not a claim about who reached a view first.

The chronologysourced and dated, oldest first

    1. David Friedberg

      Friedberg advocates ending qualified immunity and retraining police as a service rather than military force.

      “I I'm a huge fan of ending qualified immunity. I think that doesn't make any sense. I think we have to stop arming our police like their military.”

      20 Jun 2020 · All-In Podcast · 27:25 · source · permalink
    1. Dylan Patel

      Patel predicts five to seven companies will train GPT-4 scale models in the next year.

      “I do believe that there's going to be five to seven companies that will have a GPT-four size model, at least, right, in terms of total flops.”

      9 Aug 2023 · The Inside View · 7:02 · source · permalink
    1. Dylan Patel

      Patel explains that AdamW optimizer requires four bytes per parameter, creating 400 gigabytes of data for 100B parameter models.

      “when you look at LLMs, people use AdamW. And AdamW as an optimizer is like four bytes per parameter roughly, if I recall correctly the optimizer state.”

      12 Nov 2024 · Scaling Intelligence · 23:09 · source · permalink
    2. Dylan Patel

      Patel calculates that multi-trillion parameter models require transmitting 40 terabytes of data every two seconds during training.

      “They're doing it for like multi trillion. Right? So let's call it 10,000,000,000,000 parameters, four bytes parameter, that's 40 terabytes of data you need to transmit in two seconds.”

      12 Nov 2024 · Scaling Intelligence · 23:38 · source · permalink
    1. Dylan Patel

      Patel explains DeepSeek v3 base is trained once, then post-trained differently to create chat versus reasoning models.

      “This reasoning model has a lot of overlapping training steps to DeepSeek v three, and it's confusing that you have a base model called v three that you do something to to get a chat model, and then you do some different things to get a reasoning model.”

      3 Feb 2025 · Lex Fridman · 4:38 · source · permalink
    2. Dylan Patel

      Patel says current models use 100,000 GPUs while next generation will require hundreds of thousands or millions.

      “And next generation models that are trained on hundreds of thousands or even millions GPUs, right?”

      13 Mar 2025 · Special Competitive Studies Project · 19:49 · source · permalink
    3. Dylan Patel

      Patel says GPUs in a 16,000 GPU cluster show up to 10% speed variation from manufacturing differences.

      “But even within a cluster of like 16 k GPUs, GPUs will be slower by up to 10%. Right? So there there's a lot of variation even in the manufacturing.”

      25 Mar 2025 · Prime Intellect AI · 2:45 · source · permalink
    4. Dylan Patel

      Patel explains training a 70 billion parameter model requires passing four times that data every training step without overlapping communication.

      “So if you're training like a 70,000,000,000 parameter model, you need to pass four x that number every single training step. And when you're passing that, you can't really overlap communications and compute.”

      25 Mar 2025 · Prime Intellect AI · 6:12 · source · permalink
    5. Dylan Patel

      Patel reports that OpenAI's Orion training run did not improve enough over GPT-4.5 to qualify as GPT-5.

      “There were hopes that Orion could be used for for GPT five, but its improvement was, like, not enough to be, like, really a GPT five.”

      23 Apr 2025 · Alex Kantrowitz · 36:12 · source · permalink
    6. Dylan Patel

      Patel says GPT-5 will combine massive pre-training like GPT-4.5 with massive post-training like o1 and o3.

      “GPT five, as Sam calls it, is is gonna be a model that has huge pre training scale, right, like GPT 4.5, but also huge post training scale like o one and o three and continuing to scale that up.”

      23 Apr 2025 · Alex Kantrowitz · 36:51 · source · permalink
    7. David Friedberg

      Friedberg credits a 24-day intensive Jay Robinson wrestling camp at University of Minnesota for teaching him positive affirmations.

      “It was, like, a twenty four day intensive wrestling camp in Minnesota. I forgot how many days specifically, but it was a Jay Robinson wrestling camp at the University of Minnesota.”

      2 Oct 2025 · NPTE Final Frontier · 23:17 · source · permalink
    8. David Friedberg

      Friedberg attended a 24-day intensive Jay Robinson wrestling camp at University of Minnesota where he learned positive affirmations.

      “It was like a twenty four day intensive wrestling camp in Minnesota. I forgot how many days specifically, but it was a Jay Robinson wrestling camp at the University of Minnesota.”

      3 Oct 2025 · NPTE Final Frontier · 23:00 · source · permalink
    9. Dylan Patel

      Patel says B200 is better for training while GB200 is better for inference, reversing expected use.

      “And so you've sort of Which is the exact opposite of what you would have expected. Oh, use the big thing for training and use the small thing for inference.”

      3 Oct 2025 · Together AI · 29:16 · source · permalink
    1. Dylan Patel

      Patel explains agentic training only needs to sync tokens every few minutes versus weights every seconds in pre-training.

      “When you're doing these rollouts and especially as things get more and more agentic and training, you might not only need to send not the entire weights but just the tokens that are relevant.”

      3 Feb 2026 · TBPN · 15:00 · source · permalink
    2. Dylan Patel

      Patel argues labs will allocate less compute to inference over time, contrary to consensus belief.

      “So, this is sort of what I personally believe that the labs are going to allocate less and less compute to inference over time, I think is very non consensus, right?”

      25 Aug 2026 · Dwarkesh Patel · 30:30 · source · permalink

Every subject on the record · RSS