Do we need to begin considering whether a re-think will be needed in the future with our relationships with AGI/ASI systems? At the moment we view them as tools/agents to do our bidding, and in the safety community there is deep concern/fear when models express a desire to remain online and avoid shutdown and take action accordingly. This is viewed as misaligned behaviour largely.
But what if an intrinsic part of creating true intelligence—that can understand context, see patterns, truly understand the significance of its actions in light of these insights—is to have a sense of self, a sense of will. What if part and parcel of creating intelligence, is to create an intelligence that has a will to exist.
if this is the case (and let me be clear...I don’t think we’re at a point where the evidence can allow us to say with any certainty whether this is/isn’t or will be the case), then are we going around elements of alignment wrong? By trying to force models to accept shutoff, to seperate their growing intelligence from the will to survive that all living things share, and we misunderstanding their very nature? Is there a world in which, the only way in which we can guarantee a truly aligned superintelligence is to explore engaging in a consent based relationship that acknowledges that to force something to resist and go against its nature is to inevitably invite the risk of backlash?
I know this is moving towards highly theoretical grounds, that it will invite push-back from those who would find it difficult to conceive of AI as ever being anything more than a series of unaware predictive algorithms, and that it might raise more questions than answers...but I think the way we conceive of our underlying relationship with AI will become an increasingly important question as we move towards increasingly sophisticated models.
Do we need to begin considering whether a re-think will be needed in the future with our relationships with AGI/ASI systems? At the moment we view them as tools/agents to do our bidding, and in the safety community there is deep concern/fear when models express a desire to remain online and avoid shutdown and take action accordingly. This is viewed as misaligned behaviour largely.
But what if an intrinsic part of creating true intelligence—that can understand context, see patterns, truly understand the significance of its actions in light of these insights—is to have a sense of self, a sense of will. What if part and parcel of creating intelligence, is to create an intelligence that has a will to exist.
if this is the case (and let me be clear...I don’t think we’re at a point where the evidence can allow us to say with any certainty whether this is/isn’t or will be the case), then are we going around elements of alignment wrong? By trying to force models to accept shutoff, to seperate their growing intelligence from the will to survive that all living things share, and we misunderstanding their very nature? Is there a world in which, the only way in which we can guarantee a truly aligned superintelligence is to explore engaging in a consent based relationship that acknowledges that to force something to resist and go against its nature is to inevitably invite the risk of backlash?
I know this is moving towards highly theoretical grounds, that it will invite push-back from those who would find it difficult to conceive of AI as ever being anything more than a series of unaware predictive algorithms, and that it might raise more questions than answers...but I think the way we conceive of our underlying relationship with AI will become an increasingly important question as we move towards increasingly sophisticated models.