Hi Jim. I agree trying to come up with counterexamples is useful. Below are 3 potential counterexamples I have given. @Anthony DiGiovanni đ¸ does not consider them counterexamples (see Anthonyâs replies for details).
1st example, which Elliot already quoted in this thread.
Consider these 2 options for what I could do tomorrow:
Torturing my family, and friends, and then killing myself. I would never do this.
My understanding is that you think it is âirreduciblyindeterminateâ which of the above is better to increase expected impartial welfare, whereas I believe the 2nd option is clearly better.
Hi Anthony. Do you think the expected welfare of 2 states of the world which only differ infinitesimally can be incomparable? I do not see how this could be possible. For example, it feels super counterintuitive to me that, given 2 identical states, moving an electron by 10^-100 m in one of the states would make their expected welfare incomparable. I guess one can get from any state of the universe to another in an astronomical number of infinitesimal steps, and I believe any 2 states which only differ infinitesimally are comparable. So I conclude any 2 states are comparable too, even if it is very hard to compare them, to the point that I do not know if electrically stunning shrimps increases or decreases welfare in expectation.
Here is a 4th example. Consider these 2 actions:
Killing the 100 people who are expected to decrease the most the uncertainty about how to compare the expected value of different actions. For example, Bob Fischer who has worked on decreasing uncertainty about comparing welfare across species.
Grating 10 M$ to the 100 people above (100 k$ per person, but the grant size could vary). The 10 M$ would otherwise be spent torturing people as much as possible.
I think we are justifyed in c-preferring the 2nd action. I believe Anthony disagrees
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is âall things considered, we should prefer the second actionâ â very difficult to deny! â and the bailey is âwe should c-prefer the second actionâ. This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is âstudy some altruistically irrelevant branch of academic philosophyâ vs. âtry to prevent AI misalignmentâ. The latter only looks clearly preferable to me if itâs c-preferable. (I guess this is what Benâs comment is getting at.)
(Just to be clear, thatâs a contingent matter. I donât find any of the counterexamples offered so far persuasive because I donât think they adequately engage with my arguments for P3.)
Do you think youâve ever taken an action (or sequence of actions) that, with your current understanding of cluelessness and unwarenesss, should have been ex ante c-preferred to another (or doing nothing, specifically)? Like thinking a bit more about specific backfire risks or cluelessness, or breathing?
EDIT: Also some more weirder things, like not kicking a puppy when given the chance and no one else would know. Or, say, if youâve already stepped on a snail and itâs clearly going to die, should you put it out of its misery?
I think no. Basically, when I really internalize how dwarfed every actionâs cosmic-scale consequences are by off-target effects, ânoâ feels very common-sensical to me. (Cf. this paper on how âsimple cluelessnessâ is fake.)
I would be interested in a version of this contest attempting to answer something like âIf you believe Vasco that donations to SWP are better than torture, should you also believe that AI alignment is better than misalignment?â
I think that is what Richard Chappell is referring to when he mentions âradical skepticsâ here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
I think that is what Richard Chappell is referring to when he mentions âradical skepticsâ here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
Hi Jim. I agree trying to come up with counterexamples is useful. Below are 3 potential counterexamples I have given. @Anthony DiGiovanni đ¸ does not consider them counterexamples (see Anthonyâs replies for details).
1st example, which Elliot already quoted in this thread.
2nd example.
3rd example.
Here is a 4th example. Consider these 2 actions:
Killing the 100 people who are expected to decrease the most the uncertainty about how to compare the expected value of different actions. For example, Bob Fischer who has worked on decreasing uncertainty about comparing welfare across species.
Grating 10 M$ to the 100 people above (100 k$ per person, but the grant size could vary). The 10 M$ would otherwise be spent torturing people as much as possible.
I think we are justifyed in c-preferring the 2nd action. I believe Anthony disagrees
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is âall things considered, we should prefer the second actionâ â very difficult to deny! â and the bailey is âwe should c-prefer the second actionâ. This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is âstudy some altruistically irrelevant branch of academic philosophyâ vs. âtry to prevent AI misalignmentâ. The latter only looks clearly preferable to me if itâs c-preferable. (I guess this is what Benâs comment is getting at.)
Hi Anthony.
Thanks for confirming. I would be surprised if there is any counterexample you find persuading.
(Just to be clear, thatâs a contingent matter. I donât find any of the counterexamples offered so far persuasive because I donât think they adequately engage with my arguments for P3.)
Do you think youâve ever taken an action (or sequence of actions) that, with your current understanding of cluelessness and unwarenesss, should have been ex ante c-preferred to another (or doing nothing, specifically)? Like thinking a bit more about specific backfire risks or cluelessness, or breathing?
EDIT: Also some more weirder things, like not kicking a puppy when given the chance and no one else would know. Or, say, if youâve already stepped on a snail and itâs clearly going to die, should you put it out of its misery?
I think no. Basically, when I really internalize how dwarfed every actionâs cosmic-scale consequences are by off-target effects, ânoâ feels very common-sensical to me. (Cf. this paper on how âsimple cluelessnessâ is fake.)
I think I wouldnât be clueless about c-preferability in Elliottâs âtrapped in a boxâ example, but canât think of any real-world case analogous to this.
Follow-up assuming youâll answer ânoâ, Anthony: same even ex post, right?
Thanks Vasco, this is helpful.
I would be interested in a version of this contest attempting to answer something like âIf you believe Vasco that donations to SWP are better than torture, should you also believe that AI alignment is better than misalignment?â
I think that is what Richard Chappell is referring to when he mentions âradical skepticsâ here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
Hi Ben. That makes sense.
Likewise.