Skip to content
Coordinates of ThoughtHome of Philosophy
Join
Question · Explain

What is the “alignment problem” in philosophical terms?

At work people use "alignment" to mean very different things: making a chatbot polite, making it follow instructions, or making sure future systems don't pursue goals harmful to humans.

What's the core philosophical problem underneath? Is it a problem about values (which values should a system have?), about specification (how do we say what we want?), or about control? Are there older philosophical debates it connects to?

LikeAnswerFollow3 answers

Members reply here. Reading is always free.

Sign in to reply

3 answers

  1. Pieter de Jong

    Fellow

    At the engineering level it's mostly specification: you optimise a measurable proxy and the system finds ways to score well on the proxy that you didn't intend. People call it "specification gaming" or reward hacking. It's Goodhart's law: when a measure becomes a target, it ceases to be a good measure.

    Helpful · 3
  2. Tomasz Wójcik

    Fellow

    Norbert Wiener saw this in 1960, in a short paper on the moral and technical consequences of automation: if we use a machine whose operation we can't interfere with, we'd better be sure the purpose we put into it is the purpose we really desire. Stuart Russell's Human Compatible (2019) builds on that idea, arguing machines should be uncertain about our preferences.

  3. Rohan Desai

    Fellow

    From a founder's side, alignment looks a lot like the principal–agent problem in economics: how do you get someone acting on your behalf to pursue your goals and not theirs? Contracts, incentives, monitoring. The difference with AI is that the agent might be much smarter than the principal.

    Helpful · 1