It would be more useful to compare AI safety work vs. other longtermist interventions, since itâs unlikely that donations to GiveWell would beat longtermist interventions from a longtermist POV
I agree that would be incredibly useful; maybe Iâll do that next (20% chance). The same model can be used for pandemics and nuclear riskâIâd just need to update (1) P(doom) for each, (2) tractability (for AI, thatâs the âAI safety decreases risk by 7% per doubling of staffâ), and (3) personal contribution. It could be a quick tool for anyone to realize how impactful longtermist careers are and, based on their beliefs about the world and their own ability, choose the career with the highest EV, though Iâd only recommend acting on that comparison if the difference is quite large (my hunch is 5x or higher) given the uncertainty involved.
It would also force people to hold self-consistent beliefs. If thereâs a separate calculator for AI safety and one for biosecurity, someone could claim that non-[x-risk at hand] is much higher than [x-isk at hand] in each case, but that wouldnât be consistent across the two, cause each [x-risk at hand] would factor into the otherâs non-[x-risk at hand]. In other words, it can be used as a tool to calibrate beliefs about existential risks (I think it would do that for me, at least).
The biggest thing missing from the model is the possibility that safety research is net harmful.
This is quite interesting; I hadnât thought of this. Do you think it should be approximated as â% chance that AI safety is actually badâ and âincrease in AI risk per doubling of staffâ? e.g. it would look like this:
90% chance AI safety reduces AI risk, decreasing it by 10% per doubling of staff
10% chance AI safety increases AI risk, increasing it by 10% per doubling of staff
re: the latter, maybe you can get inspiration from RPâs CCM > existential risk > âsmall-scale AI misalignment projectâ and check out the graphics below. Their default params are 96.4% chance no effect, 70% chance +ve outcome conditional on effect, +30% rise in p(extinction) conditional on -ve outcome, and you can change them and see how the EV updates; these defaults donât matter as much as the takeaway that AIS work needs to be robustly +ve and that folks whose risk aversion is greater than zero (probably wise) will do well to prioritise resolving this sign uncertainty, which boils down to Michaelâs advice above (cf. the advice to build deep models, or Dave Banerjeeâs advice more specifically).
This is quite interesting; I hadnât thought of this. Do you think it should be approximated as â% chance that AI safety is actually badâ and âincrease in AI risk per doubling of staffâ?
Iâm not sure. I would probably say that you shouldnât start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how youâre going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.
I agree that would be incredibly useful; maybe Iâll do that next (20% chance). The same model can be used for pandemics and nuclear riskâIâd just need to update (1) P(doom) for each, (2) tractability (for AI, thatâs the âAI safety decreases risk by 7% per doubling of staffâ), and (3) personal contribution. It could be a quick tool for anyone to realize how impactful longtermist careers are and, based on their beliefs about the world and their own ability, choose the career with the highest EV, though Iâd only recommend acting on that comparison if the difference is quite large (my hunch is 5x or higher) given the uncertainty involved.
It would also force people to hold self-consistent beliefs. If thereâs a separate calculator for AI safety and one for biosecurity, someone could claim that non-[x-risk at hand] is much higher than [x-isk at hand] in each case, but that wouldnât be consistent across the two, cause each [x-risk at hand] would factor into the otherâs non-[x-risk at hand]. In other words, it can be used as a tool to calibrate beliefs about existential risks (I think it would do that for me, at least).
This is quite interesting; I hadnât thought of this. Do you think it should be approximated as â% chance that AI safety is actually badâ and âincrease in AI risk per doubling of staffâ? e.g. it would look like this:
90% chance AI safety reduces AI risk, decreasing it by 10% per doubling of staff
10% chance AI safety increases AI risk, increasing it by 10% per doubling of staff
Or is that too rudimentary, you think?
re: the latter, maybe you can get inspiration from RPâs CCM > existential risk > âsmall-scale AI misalignment projectâ and check out the graphics below. Their default params are 96.4% chance no effect, 70% chance +ve outcome conditional on effect, +30% rise in p(extinction) conditional on -ve outcome, and you can change them and see how the EV updates; these defaults donât matter as much as the takeaway that AIS work needs to be robustly +ve and that folks whose risk aversion is greater than zero (probably wise) will do well to prioritise resolving this sign uncertainty, which boils down to Michaelâs advice above (cf. the advice to build deep models, or Dave Banerjeeâs advice more specifically).
Iâm not sure. I would probably say that you shouldnât start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how youâre going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.