“Alignment” in the machine learning sphere is about what the agent is trained to do.
You might have one trained to solve a maze, but not realize that it didn’t “learn” to find the exit, it “learned” to move to the bottom right corner.
It’s the paperclip machine problem, basically. A human knows not to turn literally everything into paperclips, but a machine might “value” turning even themselves into paperclips above all else. And it’s actually really hard to teach it the human values without any misunderstanding.
This is about behavior, though, and does not require that we interpret the agent as “alive”.
Also, full disclosure, I fucking hate AI and would make it illegal for Google, Anthropic and Jim down the street to develop one today if I could. No quarter.
“Alignment” in the machine learning sphere is about what the agent is trained to do.
You might have one trained to solve a maze, but not realize that it didn’t “learn” to find the exit, it “learned” to move to the bottom right corner.
It’s the paperclip machine problem, basically. A human knows not to turn literally everything into paperclips, but a machine might “value” turning even themselves into paperclips above all else. And it’s actually really hard to teach it the human values without any misunderstanding.
This is about behavior, though, and does not require that we interpret the agent as “alive”.
Also, full disclosure, I fucking hate AI and would make it illegal for Google, Anthropic and Jim down the street to develop one today if I could. No quarter.