Instrumental convergence is the thesis that sufficiently intelligent agents with a high proportion of goals will tend to pursue certain similar instrumental goals, such as resource- and power-seeking, self-preservation, and recursive self-improvement.
I’m writing on the subject, and want to get a sense of how important the community thinks this is in the context of AI safety. So this needs two polls for comparison.
Defining ‘catastrophe’ as ‘a >50% drop in human population in under a decade’ and ‘AI-caused’ as ‘some event clearly tracing to the action of one or more AIs alone or of humans using AI’ (so e.g. an artificially created pandemic would count).
(I couldn’t see a way to get a legend on the poll, but there are 20 values available, so they’re implicitly increments of 5%)
Yesterday I read (some of) Scott Alexander’s recent open letter to Steven Pinker. In it, he lists 3 mecanisms that I suppose he considers to be central to the credibility of AI-induced catastrophic scenarios: it starts here. (they’re instrumental convergence, misgeneralization, and reward-hacking)
[Question] How important is instrumental convergence?
Instrumental convergence is the thesis that sufficiently intelligent agents with a high proportion of goals will tend to pursue certain similar instrumental goals, such as resource- and power-seeking, self-preservation, and recursive self-improvement.
I’m writing on the subject, and want to get a sense of how important the community thinks this is in the context of AI safety. So this needs two polls for comparison.
Defining ‘catastrophe’ as ‘a >50% drop in human population in under a decade’ and ‘AI-caused’ as ‘some event clearly tracing to the action of one or more AIs alone or of humans using AI’ (so e.g. an artificially created pandemic would count).
(I couldn’t see a way to get a legend on the poll, but there are 20 values available, so they’re implicitly increments of 5%)
Please give any context that you think matters.
Ty!
Yesterday I read (some of) Scott Alexander’s recent open letter to Steven Pinker. In it, he lists 3 mecanisms that I suppose he considers to be central to the credibility of AI-induced catastrophic scenarios: it starts here.
(they’re instrumental convergence, misgeneralization, and reward-hacking)