That is sorta the idea yes. Agents would choose this decision criteria mostly because it vastly increases their odds of survival, which allows them to further whatever goals they have. I would hope that this result is obvious enough that many agents will be able to converge on it, increasing the proportion using it, and thus increasing the overall survival rate.
The other takeaway is that, given that humans will be weaker than AGI/ASI, any game theoretic reason for such entities to still cooperate with us can potentially help reduce the existential risk.
I agree that the model requires further scrutiny to determine if it is realistic enough to matter.
That is sorta the idea yes. Agents would choose this decision criteria mostly because it vastly increases their odds of survival, which allows them to further whatever goals they have. I would hope that this result is obvious enough that many agents will be able to converge on it, increasing the proportion using it, and thus increasing the overall survival rate.
The other takeaway is that, given that humans will be weaker than AGI/ASI, any game theoretic reason for such entities to still cooperate with us can potentially help reduce the existential risk.
I agree that the model requires further scrutiny to determine if it is realistic enough to matter.