Patel says reasoning capability allowing models to think before answering dramatically increased capabilities in six months.
“Right? And this has just happened in six months. And the same applies to various certain coding tasks. The same applies to yeah.”
Patel says DeepSeek r one zero shows reasoning behaviors emerge naturally from RL on verifiable rewards without human data.
“So it's the remarkable thing about these reasoning results, and especially the DeepSeek r one paper, is this result that they call DeepSeek r one zero, which is they took one of these pretrained models.”