Dylan Patel on

reinforcement learning

2 entries, 3 Feb 2025 to 15 Jan 2026

On the recordsourced and dated, oldest first

    1. spoken

      Patel says DeepSeek r one zero shows reasoning behaviors emerge naturally from RL on verifiable rewards without human data.

      “So it's the remarkable thing about these reasoning results, and especially the DeepSeek r one paper, is this result that they call DeepSeek r one zero, which is they took one of these pretrained models.”

      3 Feb 2025 · Lex Fridman · 2:43:33 · source · permalink
    1. spoken

      Patel says last year at NeurIPS was first time he focused on RL and verifiable test-time scaling research.

      “last year was the first time I, like, stopped at a bunch of, like, RL stuff and verifiable, you know, test times, you know, sort of scaling stuff that, you know, verifiable RL stuff, all that kind of stuff.”

      15 Jan 2026 · SAIL Media · 1:04 · source · permalink

Dylan Patel ontheir other subjects

Everything Dylan Patel is on record saying · RSS