Dylan Patel · Lex Fridman · 3 February 2025
“One, I will classify as preference fine tuning. Preference fine tuning is a generalized term for what came out of reinforcement learning from human feedback, which is RLHF.”
Verbatim excerpt with a timestamp. The full recording is at the source; we link out and do not host it. Everything Dylan Patel is on record saying.