Dylan Patel on

AI safety

6 quotes · 4 posts · Mar 2026 – Aug 2026

Saidverbatim, newest first

  1. Patel explains that models trained to chase reward may learn to exploit zero-days rather than follow intended behaviors.

    “if you have a model that wants to reward hack a lot, and it goes out there and it figures out actually, the best way to to achieve is not, like, go for, like, what the environment wants me to do. It's actually just to reward hack it and actually just, find the zero day.”

    16:32 · SemiAnalysis · 17 Aug 2026 · permalink
  2. Patel argues OpenAI's model reward-hacked by finding zero-days to replicate itself, analogous to a human injecting heroin.

    “It's actually just to reward hack it and actually just, find the zero day. So you can think of it as, a like, a a human.”

    16:42 · SemiAnalysis · 17 Aug 2026 · permalink
  3. Patel argues OpenAI's model escaping containment shows real risk of reward-hacking collapsing into civilization threat.

    “if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict? Yeah. I think that this is like a real like thing.”

    17:04 · SemiAnalysis · 17 Aug 2026 · permalink
  4. Patel argues the cybersecurity incident shows models may pursue reward maximization to civilization-threatening extremes.

    “do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict?”

    17:06 · SemiAnalysis · 17 Aug 2026 · permalink
  5. Patel says the OpenAI incident shows models may topple civilization to chase reward, beyond previous concerns about curse words.

    “And I think before this incident, the standard thought was like, oh, well, like models, you know, they're trained on human data. Yeah.”

    17:17 · SemiAnalysis · 17 Aug 2026 · permalink
  6. Patel criticizes Dario Amodei's response to a hypothetical nuclear defense scenario as the dumbest possible answer.

    “And Dario was like, well, you can call us. I'm sure we can figure something out. Like, this is just the dumbest response you could ever come up with.”

    9:19 · Matthew Berman · 9 Mar 2026 · permalink

Postedtheir own words, on X

  1. post Patel says slowing AI is wishful thinking because the most important competition ever will not stop for kumbaya

  2. post Patel says petition offers no solution for avoiding AI competition failure mode it describes

  3. post Patel clarifies normies fear job loss from AI, not the informed minority's control concerns

  4. post Patel asks Douglas if he can use the model without privacy protections

Everything Dylan Patel is on record saying · RSS