This is a neat idea, it’s difficult to come up with safe preferences to encode in an ASI, and the concept of strong risk-aversion might help.
A major obstacle (which I didn’t see listed in section 8) is that currently we have no idea how to embed any set of preferences whatsoever in an ASI.
2b. If we figure out how to encode risk-averse preferences in an ASI, then I’m not sure it makes sense to speak of it as “misaligned”, because clearly we do know how to get it to pursue goals that we care about. It seems weird to expect that we won’t know how to make ASI not want to tile the universe with paperclips, but we will know how to make it want to risk-aversely tile the universe with paperclips.
I think section 10 is pointing at something similar. I find it at least somewhat plausible that RL on risk aversion generalizes better than other kinds of RL. I would still be surprised if we could get risk aversion to generalize to ASI using anything resembling current techniques, but this seems like a better-than-average idea for preventing AI takeover.
I’m unsure whether we can successfully train ASIs to be reliably risk-averse, including far OOD. Our claim is just that the chances of success are high enough to make risk aversion worth pursuing as a line of defense. That’s the case we try to make in section 10. See also my reply to Ryan’s comment. I also think our chances of success are a bit higher for AIs that aren’t yet ASIs, and if we succeed in making them risk-averse I think they could help a lot with aligning any later-arising ASIs, by doing this sort of stuff.
I think you could know how to encode preferences into ASI without knowing whether it’s aligned. At that point it’s like a genie, you may well be able encode preferences into it but you might encode bad preferences that you didn’t realize would be bad.
As I understand the term, that sort of ASI wouldn’t be considered “misaligned”, it would be “aligned, but to the wrong target”. I think of misalignment as when you wanted the ASI to do one thing, but it did something else instead.
Two quick thoughts:
This is a neat idea, it’s difficult to come up with safe preferences to encode in an ASI, and the concept of strong risk-aversion might help.
A major obstacle (which I didn’t see listed in section 8) is that currently we have no idea how to embed any set of preferences whatsoever in an ASI. 2b. If we figure out how to encode risk-averse preferences in an ASI, then I’m not sure it makes sense to speak of it as “misaligned”, because clearly we do know how to get it to pursue goals that we care about. It seems weird to expect that we won’t know how to make ASI not want to tile the universe with paperclips, but we will know how to make it want to risk-aversely tile the universe with paperclips.
I think section 10 is pointing at something similar. I find it at least somewhat plausible that RL on risk aversion generalizes better than other kinds of RL. I would still be surprised if we could get risk aversion to generalize to ASI using anything resembling current techniques, but this seems like a better-than-average idea for preventing AI takeover.
Thanks!
I’m unsure whether we can successfully train ASIs to be reliably risk-averse, including far OOD. Our claim is just that the chances of success are high enough to make risk aversion worth pursuing as a line of defense. That’s the case we try to make in section 10. See also my reply to Ryan’s comment. I also think our chances of success are a bit higher for AIs that aren’t yet ASIs, and if we succeed in making them risk-averse I think they could help a lot with aligning any later-arising ASIs, by doing this sort of stuff.
I think you could know how to encode preferences into ASI without knowing whether it’s aligned. At that point it’s like a genie, you may well be able encode preferences into it but you might encode bad preferences that you didn’t realize would be bad.
As I understand the term, that sort of ASI wouldn’t be considered “misaligned”, it would be “aligned, but to the wrong target”. I think of misalignment as when you wanted the ASI to do one thing, but it did something else instead.