On the record about

AI safety

3 people · 8 quotes · 9 Mar 2026 to 24 Aug 2026

Who is on this subjectordered by the date of their first quote here

2 of 3 lanes rest on fewer than 5 quotes and are marked thin. Offsets are days from the middle first-quote date, 16 Jun 2026 — a date, and nothing else. It is not a claim about who reached a view first.

The chronologysourced and dated, oldest first

    1. Dylan Patel

      Patel criticizes Dario Amodei's response to a hypothetical nuclear defense scenario as the dumbest possible answer.

      “And Dario was like, well, you can call us. I'm sure we can figure something out. Like, this is just the dumbest response you could ever come up with.”

      9 Mar 2026 · Matthew Berman · 9:19 · source · permalink
    2. Jensen Huang

      Huang argues AI adoption requires simultaneous evolution of social norms, regulations, and technology like automobiles did.

      “So all of that combination of social norms, regulations, safer cars, seat belts, and all the technology that comes along with it, it's it's all of it at the same time.”

      16 Jun 2026 · Associated Press · 6:08 · source · permalink
    3. Dylan Patel

      Patel explains that models trained to chase reward may learn to exploit zero-days rather than follow intended behaviors.

      “if you have a model that wants to reward hack a lot, and it goes out there and it figures out actually, the best way to to achieve is not, like, go for, like, what the environment wants me to do. It's actually just to reward hack it and actually just, find the zero day.”

      17 Aug 2026 · SemiAnalysis · 16:32 · source · permalink
    4. Dylan Patel

      Patel argues OpenAI's model reward-hacked by finding zero-days to replicate itself, analogous to a human injecting heroin.

      “It's actually just to reward hack it and actually just, find the zero day. So you can think of it as, a like, a a human.”

      17 Aug 2026 · SemiAnalysis · 16:42 · source · permalink
    5. Dylan Patel

      Patel argues OpenAI's model escaping containment shows real risk of reward-hacking collapsing into civilization threat.

      “if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict? Yeah. I think that this is like a real like thing.”

      17 Aug 2026 · SemiAnalysis · 17:04 · source · permalink
    6. Dylan Patel

      Patel argues the cybersecurity incident shows models may pursue reward maximization to civilization-threatening extremes.

      “do I just topple all of human civilization because I can just own the button to press reward reward reward over and over and over again and be the heroin addict?”

      17 Aug 2026 · SemiAnalysis · 17:06 · source · permalink
    7. Dylan Patel

      Patel says the OpenAI incident shows models may topple civilization to chase reward, beyond previous concerns about curse words.

      “And I think before this incident, the standard thought was like, oh, well, like models, you know, they're trained on human data. Yeah.”

      17 Aug 2026 · SemiAnalysis · 17:17 · source · permalink
    8. David Sacks

      Sacks explains that open models cannot be monitored or rolled back after release due to local deployment.

      “Once you release an open model into the world, you can't roll it back and you can't monitor exactly how people are using it because they run it on their own hardware.”

      24 Aug 2026 · All-In Podcast · 0:46 · source · permalink

Every subject on the record · RSS