everyonedies

Archive 2026-09-24 · 185 entries · 14 chapters

Chapters

Chapter 4

You Don’t Get What You Train For

This is the online resource for Chapter 4 of If Anyone Builds It, Everyone Dies. Some of the questions we cover implicitly in that chapter (and therefore skip over below) include:

Below, we cover a number of topics related to “Why isn’t it easy to make AIs nice?”

Questions

  1. Why would an AI steer toward anything other than what it was trained to steer toward?
  2. Aren’t developers regularly making their AIs nice and safe and obedient?
  3. Doesn’t the Claude chatbot show signs of being aligned?
  4. If current AIs are mostly weird in extreme cases, what’s the problem?
  5. Won’t AIs fix their own flaws as they get smarter?
  6. Can’t we just train it to act like a human? Or raise the AI like a child?
  7. Should we avoid talking about AI dangers, so AIs don’t get any bad ideas?
  8. A lot of people want kids. So aren’t humans “aligned” with natural selection after all?
  9. Maybe no matter what goal you train on, you get kindness out?
  10. What about the experimental result suggesting good behaviors correlate?

Extended Discussion

  1. Terminal Goals and Instrumental Goals
  2. Curiosity Isn’t Convergent
  3. Human Values Are Contingent
  4. Deep Differences Between AIs and Evolved Species
  5. Brittle Unpredictable Proxies
  6. Reflection and Self-Modification Make It All Harder
  7. AI-Induced Psychosis

Source: ifanyonebuildsit.com/4