I’ve been thinking a bit on this reasoning. You write: “consider, e.g., an unrealistic scenario where you could press a button with guaranteed overall catastrophic consequences or one that would get the universe rid of all unwanted suffering forever while keeping everything else equal.” I agree that those seem like clear cases, but they are also cases in which the action is specified in terms of its causal consequences. This confounds the discussion a bit because it gives you the long-term value of the act upfront. For instance, consider if you took the case to be “push the button that reads ′ guaranteed overall catastrophic consequences’”, but didn’t assume that the label was veridical. Now the case falls apart. So, the cases seem like they are being driven, not by the fact that they are realistic or unrealistic, but by the fact that they specify the long-term causal consequences. And any action so-described can then be compared against any other so-described; they are no longer incommensurable.
So, I am wondering if your case can be made without building in the long-term consequences upfront. At the moment, I don’t see how. But perhaps I am missing something.
So, the cases seem like they are being driven, not by the fact that they are realistic or unrealistic, but by the fact that they specify the long-term causal consequences.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don’t know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don’t need (full) specification but a situation where the uncertainty is manageable.
Perhaps. But the case is considerably less tight here. It works, to the extent it does, because we think this is plausibly a case where we have a grip on long-term EV (though, maybe we don’t). And, if that is what is doing the work, then it seems like there is a reason for thinking that the problem is more closely tied to “promoting the total good” because this is what makes course awareness of total consequences in real world conditions likely.
I’ve been thinking a bit on this reasoning. You write: “consider, e.g., an unrealistic scenario where you could press a button with guaranteed overall catastrophic consequences or one that would get the universe rid of all unwanted suffering forever while keeping everything else equal.” I agree that those seem like clear cases, but they are also cases in which the action is specified in terms of its causal consequences. This confounds the discussion a bit because it gives you the long-term value of the act upfront. For instance, consider if you took the case to be “push the button that reads ′ guaranteed overall catastrophic consequences’”, but didn’t assume that the label was veridical. Now the case falls apart. So, the cases seem like they are being driven, not by the fact that they are realistic or unrealistic, but by the fact that they specify the long-term causal consequences. And any action so-described can then be compared against any other so-described; they are no longer incommensurable.
So, I am wondering if your case can be made without building in the long-term consequences upfront. At the moment, I don’t see how. But perhaps I am missing something.
What about a button that stops a superintelligent AI from trying to eliminate all humans? (assuming it would be bad if it succeeds). The outcome is not specified because we don’t know for sure that the superintelligence will succeed. However, that seems likely enough for us to be warranted in pushing that button. We don’t need (full) specification but a situation where the uncertainty is manageable.
Perhaps. But the case is considerably less tight here. It works, to the extent it does, because we think this is plausibly a case where we have a grip on long-term EV (though, maybe we don’t). And, if that is what is doing the work, then it seems like there is a reason for thinking that the problem is more closely tied to “promoting the total good” because this is what makes course awareness of total consequences in real world conditions likely.
Just a thought.