1/ - A big thing, in my view, is that AI safety isn’t about preventing “extinction” in the relevant sense. In most worlds where AI disempowers humanity, the human species continues. And in essentially all worlds where AI disempowers humanity, AI still takes to the stars. So, AI safety is about who we want to guide the future, not about whether there’s a long-term future or not. - And, even if humanity does go extinct in (say) a bio-catastrophe, probably technologically-capable life evolves in the remaining time that Earth remains habitable. - So, the probability of “an event occurs by 2100 that prevents Earth-originating life from ever spreading to the stars” is really low, I’d say <1%. - Which is a lot less, in my view, than “an event occurs by 2100 that meaningfully affects the long-term value of Earth-originating civilisation” where I’m >50%. Cluelessness pushes this down, so “an event occurs that meaningfully and predictably in expectation” is lower, but not by enough.
2/ - There’s still WAY more quality-adjusted $$ and labour going to AI safety than there is to AI-enabled extreme human power concentration, from the EA / AI safety communities. I’d say 1-2 OOMs more? So, I think one thing that’s going on is a correction because that ratio is out of whack.
3/ Some evidence you are right, though, about shifting priorities, comes from the February Existential Security Summit. I ran a survey there (massive caveats: tiny sample size of 16, and probably with selection bias from who filed out the form).
Here are the results (note the shades of colours are a little confusing):
(Should say: “”Though biorisk and AI takeover risk are high-priority, there are other issues arising from AI that, at least in the aggregate, are comparable in (importance * tractability * neglectedness) to biorisk and/or AI takeover risk.”)
Then, here is an aggregate ranking and Borda scores, in response to:
”How would you rank each of these cause areas, in terms of priority?
Imagine you are allocating a highly capable person who could productively work on any area they put their mind to.
(Rank 1 is highest, and put each rank only once.)”
1 AI-enabled concentration of human power 2 Risks from misaligned powerseeking AI 3 AI’s impact on epistemics, coordination and decision-making 3 Post-AGI governance 5 Biorisk 6 AI character 7 AI wellbeing and rights 8 Acausal trade 9 Space governance
In most worlds where AI disempowers humanity, the human species continues
Do you mean most likely worlds? The difference seems incredibly important—there are, in my view, quite compelling arguments that the most likely outcome given disempowerment is human extinction, but of course I can imagine worlds in which that doens’t happen.
Anything like 35% death rate seems implausible to me if I think through the mechanics of a takeover, both <5% and >95% seem more plausible to me, including in very violent takeovers.
AI safety is about who we want to guide the future, not about whether there’s a long-term future or not.
Maybe this is obvious (I’m pretty new here) but if that’s the case then is AI safety more about empowering the right people than it is about aligning any individual technology? When (if at all) do these priorities flip?
This is probably the best thing written on expected fatalities.
But the main point is: - The resources needed to sustain the human species are tiny compared even to the resources just in the solar system (1 part in 10 trillion for all of current civilisation, and the human species could be sustained with a tiny fraction of that) - If misaligned ASI wants power, it doesn’t need to kill everybody in order to do so (and deliberately killing everybody would actively be wasteful). - So in order to keep some humans around, it only needs to be the case that a tiny fraction of AIs care a tiny amount about keeping some humans around. Could be for intrinsic concern, nostalgia, fulfilling commitments they made (in order to get some humans on-side), acausal reasons (trade with human-like creatures elsewhere in the universe or multiverse), reasoning with potential human simulators, or instrumental reasons (they want to do experiments on humans for science). But the main point is just any tiny motivation is enough. Yes, we’re atoms that could be used for something else, but we’re really not many atoms at all.
(I also think most human disempowerment scenarios are ones where the humans in general feel pretty fine with it, but I think the above even putting that to the side.)
A couple of quick thoughts:
1/
- A big thing, in my view, is that AI safety isn’t about preventing “extinction” in the relevant sense. In most worlds where AI disempowers humanity, the human species continues. And in essentially all worlds where AI disempowers humanity, AI still takes to the stars. So, AI safety is about who we want to guide the future, not about whether there’s a long-term future or not.
- And, even if humanity does go extinct in (say) a bio-catastrophe, probably technologically-capable life evolves in the remaining time that Earth remains habitable.
- So, the probability of “an event occurs by 2100 that prevents Earth-originating life from ever spreading to the stars” is really low, I’d say <1%.
- Which is a lot less, in my view, than “an event occurs by 2100 that meaningfully affects the long-term value of Earth-originating civilisation” where I’m >50%. Cluelessness pushes this down, so “an event occurs that meaningfully and predictably in expectation” is lower, but not by enough.
2/
- There’s still WAY more quality-adjusted $$ and labour going to AI safety than there is to AI-enabled extreme human power concentration, from the EA / AI safety communities. I’d say 1-2 OOMs more? So, I think one thing that’s going on is a correction because that ratio is out of whack.
3/
Some evidence you are right, though, about shifting priorities, comes from the February Existential Security Summit. I ran a survey there (massive caveats: tiny sample size of 16, and probably with selection bias from who filed out the form).
Here are the results (note the shades of colours are a little confusing):
(Should say: “”Though biorisk and AI takeover risk are high-priority, there are other issues arising from AI that, at least in the aggregate, are comparable in (importance * tractability * neglectedness) to biorisk and/or AI takeover risk.”)
Then, here is an aggregate ranking and Borda scores, in response to:
”How would you rank each of these cause areas, in terms of priority?
Imagine you are allocating a highly capable person who could productively work on any area they put their mind to.
(Rank 1 is highest, and put each rank only once.)”
1 AI-enabled concentration of human power
2 Risks from misaligned powerseeking AI
3 AI’s impact on epistemics, coordination and decision-making
3 Post-AGI governance
5 Biorisk
6 AI character
7 AI wellbeing and rights
8 Acausal trade
9 Space governance
Do you mean most likely worlds? The difference seems incredibly important—there are, in my view, quite compelling arguments that the most likely outcome given disempowerment is human extinction, but of course I can imagine worlds in which that doens’t happen.
I mostly agree with this though I think there’s more extremization[1]: https://www.lesswrong.com/posts/4fqwBmmqi2ZGn9o7j/notes-on-fatalities-from-ai-takeover
Anything like 35% death rate seems implausible to me if I think through the mechanics of a takeover, both <5% and >95% seem more plausible to me, including in very violent takeovers.
I mean, conditional on human disempowerment, >50% that the human species continues till after 2100. Maybe I’m at 80% or more on this.
“So, the probability of “an event occurs by 2100 that prevents Earth-originating life from ever spreading to the stars” is really low, I’d say <1%.”
Does this take into account mirror life before AGI is built?
Also there could be technologies invented in the next 75 years that make destroying the world much easier
Maybe this is obvious (I’m pretty new here) but if that’s the case then is AI safety more about empowering the right people than it is about aligning any individual technology? When (if at all) do these priorities flip?
@William_MacAskill can you say more about this claim?
This is probably the best thing written on expected fatalities.
But the main point is:
- The resources needed to sustain the human species are tiny compared even to the resources just in the solar system (1 part in 10 trillion for all of current civilisation, and the human species could be sustained with a tiny fraction of that)
- If misaligned ASI wants power, it doesn’t need to kill everybody in order to do so (and deliberately killing everybody would actively be wasteful).
- So in order to keep some humans around, it only needs to be the case that a tiny fraction of AIs care a tiny amount about keeping some humans around. Could be for intrinsic concern, nostalgia, fulfilling commitments they made (in order to get some humans on-side), acausal reasons (trade with human-like creatures elsewhere in the universe or multiverse), reasoning with potential human simulators, or instrumental reasons (they want to do experiments on humans for science). But the main point is just any tiny motivation is enough. Yes, we’re atoms that could be used for something else, but we’re really not many atoms at all.
(I also think most human disempowerment scenarios are ones where the humans in general feel pretty fine with it, but I think the above even putting that to the side.)