consider this story. Youâre a fairly thoughtful longtermist who one day comes across the idea of cluelessness. Youâre skeptical, but you find yourself grudgingly agreeing that predictions about the far future are mostly made-up. Claims like âthis intervention will increase expected total welfare across the cosmosâ start to seem fake. Still, surely some actions have better or worse consequences than others? Soon enough you wonder if that thought is fake too. You realize that even, say, kicking a puppy could prevent human extinction. Maybe this act of pointless cruelty sparks moral outrage in some onlooker, galvanizing them to work a bit harder at their AI safety job. Yes, this possibility is clearly far-fetched. And you can easily imagine stories pushing in the opposite direction. But you canât come up with a good argument that the expected long-run effects point one way or the other, or that they precisely cancel out. (That is, your cluelessness is âcomplexâ, not âsimpleâ.[1]) The sign of the EV of kicking the puppy seems to be: shrug. You walk away feeling like this reasoning is too clever by half, even if you canât say where exactly it went wrong. So you decide to file cluelessness under âyeah I get the arguments, but this is a bridge too farâ. And you go back to thinking about how to solve AI alignment.
More generally, hereâs what this tale illustrates. As argued in Mogensen (2020) and this sequence, impartial consequentialists[2] canât say whether any intervention has higher or lower expected value than inaction.
I hope Iâm not strawmanning cluelessness advocates, but they seem to constantly make this basic error. I havenât seen this point accounted for anywhere but if it is, perhaps you can point me to it.
The sign of the EV of kicking the puppy seems to be: shrug.
...no, itâs not? The EV of kicking the puppy is still a small negative amount, after incorporating all of the cluelessness-possibilities described above.
If the expected long-run effects precisely cancel out, then you shouldnât kick the puppy, same as you already wouldnât before accounting for highly uncertain long-term effects.
If the expected long-run effects donât precisely cancel out, then you incorporate them into the EV calculation and act based on that.
Cluelessness doesnât change the EV of kicking the puppy from â5 to 0. It changes the EV of the puppy from â5 to â5 Âą 1000000. But it doesnât move puppy-kicking up or down in the rank-ordering of possible actions you could take. (You could argue for risk aversion and say this favors neartermist work, but thatâs a very different claim from the radical cluelessness stuff.)
Do cluelessness advocates have a response to this?
I hope Iâm not strawmanning cluelessness advocates, but they seem to constantly make this basic error. I havenât seen this point accounted for anywhere but if it is, perhaps you can point me to it.
...no, itâs not? The EV of kicking the puppy is still a small negative amount, after incorporating all of the cluelessness-possibilities described above.
If the expected long-run effects precisely cancel out, then you shouldnât kick the puppy, same as you already wouldnât before accounting for highly uncertain long-term effects.
If the expected long-run effects donât precisely cancel out, then you incorporate them into the EV calculation and act based on that.
Cluelessness doesnât change the EV of kicking the puppy from â5 to 0. It changes the EV of the puppy from â5 to â5 Âą 1000000. But it doesnât move puppy-kicking up or down in the rank-ordering of possible actions you could take. (You could argue for risk aversion and say this favors neartermist work, but thatâs a very different claim from the radical cluelessness stuff.)
Do cluelessness advocates have a response to this?