This is a nice analysis and I wish more people would do things like this. A lot of the setup looks sensible to me. I thought it was clever how you used non-AI x-risk to determine the total number of people affected by AI x-risk, since this doesn’t get you unwieldy Astronomical Waste-style numbers, but it also assigns positive value to future people. I’m not sure it’s actually a good approach, but at least it’s a clever idea.
It would be more useful to compare AI safety work vs. other longtermist interventions, since it’s unlikely that donations to GiveWell would beat longtermist interventions from a longtermist POV. But I realize that would be a lot more work, and you’ve already put an admirable amount of work into this.
The biggest thing missing from the model is the possibility that safety research is net harmful. I believe much historical safety research ended up being harmful by making AI easier to commercialize and thus accelerating development (it would’ve been better for researchers to focus more on theoretical work that doesn’t directly enable commercialization). I’m less sure about this but there may also be a replacement effect where empirical work on aligning current-gen models—which I don’t think is very useful for aligning ASI—crowds out more important long-term work.
It would be more useful to compare AI safety work vs. other longtermist interventions, since it’s unlikely that donations to GiveWell would beat longtermist interventions from a longtermist POV
I agree that would be incredibly useful; maybe I’ll do that next (20% chance). The same model can be used for pandemics and nuclear risk—I’d just need to update (1) P(doom) for each, (2) tractability (for AI, that’s the ‘AI safety decreases risk by 7% per doubling of staff’), and (3) personal contribution. It could be a quick tool for anyone to realize how impactful longtermist careers are and, based on their beliefs about the world and their own ability, choose the career with the highest EV, though I’d only recommend acting on that comparison if the difference is quite large (my hunch is 5x or higher) given the uncertainty involved.
It would also force people to hold self-consistent beliefs. If there’s a separate calculator for AI safety and one for biosecurity, someone could claim that non-[x-risk at hand] is much higher than [x-isk at hand] in each case, but that wouldn’t be consistent across the two, cause each [x-risk at hand] would factor into the other’s non-[x-risk at hand]. In other words, it can be used as a tool to calibrate beliefs about existential risks (I think it would do that for me, at least).
The biggest thing missing from the model is the possibility that safety research is net harmful.
This is quite interesting; I hadn’t thought of this. Do you think it should be approximated as “% chance that AI safety is actually bad” and “increase in AI risk per doubling of staff”? e.g. it would look like this:
90% chance AI safety reduces AI risk, decreasing it by 10% per doubling of staff
10% chance AI safety increases AI risk, increasing it by 10% per doubling of staff
re: the latter, maybe you can get inspiration from RP’s CCM > existential risk > “small-scale AI misalignment project” and check out the graphics below. Their default params are 96.4% chance no effect, 70% chance +ve outcome conditional on effect, +30% rise in p(extinction) conditional on -ve outcome, and you can change them and see how the EV updates; these defaults don’t matter as much as the takeaway that AIS work needs to be robustly +ve and that folks whose risk aversion is greater than zero (probably wise) will do well to prioritise resolving this sign uncertainty, which boils down to Michael’s advice above (cf. the advice to build deep models, or Dave Banerjee’s advice more specifically).
This is quite interesting; I hadn’t thought of this. Do you think it should be approximated as “% chance that AI safety is actually bad” and “increase in AI risk per doubling of staff”?
I’m not sure. I would probably say that you shouldn’t start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how you’re going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.
This is a nice analysis and I wish more people would do things like this. A lot of the setup looks sensible to me. I thought it was clever how you used non-AI x-risk to determine the total number of people affected by AI x-risk, since this doesn’t get you unwieldy Astronomical Waste-style numbers, but it also assigns positive value to future people. I’m not sure it’s actually a good approach, but at least it’s a clever idea.
It would be more useful to compare AI safety work vs. other longtermist interventions, since it’s unlikely that donations to GiveWell would beat longtermist interventions from a longtermist POV. But I realize that would be a lot more work, and you’ve already put an admirable amount of work into this.
The biggest thing missing from the model is the possibility that safety research is net harmful. I believe much historical safety research ended up being harmful by making AI easier to commercialize and thus accelerating development (it would’ve been better for researchers to focus more on theoretical work that doesn’t directly enable commercialization). I’m less sure about this but there may also be a replacement effect where empirical work on aligning current-gen models—which I don’t think is very useful for aligning ASI—crowds out more important long-term work.
I agree that would be incredibly useful; maybe I’ll do that next (20% chance). The same model can be used for pandemics and nuclear risk—I’d just need to update (1) P(doom) for each, (2) tractability (for AI, that’s the ‘AI safety decreases risk by 7% per doubling of staff’), and (3) personal contribution. It could be a quick tool for anyone to realize how impactful longtermist careers are and, based on their beliefs about the world and their own ability, choose the career with the highest EV, though I’d only recommend acting on that comparison if the difference is quite large (my hunch is 5x or higher) given the uncertainty involved.
It would also force people to hold self-consistent beliefs. If there’s a separate calculator for AI safety and one for biosecurity, someone could claim that non-[x-risk at hand] is much higher than [x-isk at hand] in each case, but that wouldn’t be consistent across the two, cause each [x-risk at hand] would factor into the other’s non-[x-risk at hand]. In other words, it can be used as a tool to calibrate beliefs about existential risks (I think it would do that for me, at least).
This is quite interesting; I hadn’t thought of this. Do you think it should be approximated as “% chance that AI safety is actually bad” and “increase in AI risk per doubling of staff”? e.g. it would look like this:
90% chance AI safety reduces AI risk, decreasing it by 10% per doubling of staff
10% chance AI safety increases AI risk, increasing it by 10% per doubling of staff
Or is that too rudimentary, you think?
re: the latter, maybe you can get inspiration from RP’s CCM > existential risk > “small-scale AI misalignment project” and check out the graphics below. Their default params are 96.4% chance no effect, 70% chance +ve outcome conditional on effect, +30% rise in p(extinction) conditional on -ve outcome, and you can change them and see how the EV updates; these defaults don’t matter as much as the takeaway that AIS work needs to be robustly +ve and that folks whose risk aversion is greater than zero (probably wise) will do well to prioritise resolving this sign uncertainty, which boils down to Michael’s advice above (cf. the advice to build deep models, or Dave Banerjee’s advice more specifically).
I’m not sure. I would probably say that you shouldn’t start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how you’re going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.