So, the cases seem like they are being driven, not by the fact that they are realistic or unrealistic, but by the fact that they specify the long-term causal consequences.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don’t know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don’t need (full) specification but a situation where the uncertainty is manageable.
Perhaps. But the case is considerably less tight here. It works, to the extent it does, because we think this is plausibly a case where we have a grip on long-term EV (though, maybe we don’t). And, if that is what is doing the work, then it seems like there is a reason for thinking that the problem is more closely tied to “promoting the total good” because this is what makes course awareness of total consequences in real world conditions likely.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don’t know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don’t need (full) specification but a situation where the uncertainty is manageable.
Perhaps. But the case is considerably less tight here. It works, to the extent it does, because we think this is plausibly a case where we have a grip on long-term EV (though, maybe we don’t). And, if that is what is doing the work, then it seems like there is a reason for thinking that the problem is more closely tied to “promoting the total good” because this is what makes course awareness of total consequences in real world conditions likely.
Just a thought.