This is quite interesting; I hadn’t thought of this. Do you think it should be approximated as “% chance that AI safety is actually bad” and “increase in AI risk per doubling of staff”?
I’m not sure. I would probably say that you shouldn’t start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how you’re going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.
I’m not sure. I would probably say that you shouldn’t start a career in AI safety unless you can articulate a theory of why safety work has been harmful in the past, and how you’re going to avoid more of the same. Building that theory is more important than adjusting the model inputs on a cost-effectiveness model.