By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
Glad I saw this comment because like many people I assumed this meant “do you expect aligned AI to on average turn out better than unaligned AI” which seems very different than what you actually meant.
Ah, sorry I misunderstood. But if you assume the chance of loss of control is fairly high anyway, then a safeguard if a value alignment seems invaluable.
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
Glad I saw this comment because like many people I assumed this meant “do you expect aligned AI to on average turn out better than unaligned AI” which seems very different than what you actually meant.
Ah, sorry I misunderstood. But if you assume the chance of loss of control is fairly high anyway, then a safeguard if a value alignment seems invaluable.