Patel says DeepSeek r one zero shows reasoning behaviors emerge naturally from RL on verifiable rewards without human data.
“So it's the remarkable thing about these reasoning results, and especially the DeepSeek r one paper, is this result that they call DeepSeek r one zero, which is they took one of these pretrained models.”
Patel says last year at NeurIPS was first time he focused on RL and verifiable test-time scaling research.
“last year was the first time I, like, stopped at a bunch of, like, RL stuff and verifiable, you know, test times, you know, sort of scaling stuff that, you know, verifiable RL stuff, all that kind of stuff.”