Cluelessness Critiques Week: Discussion Thread
This week is Cluelessness Critiques Week on the EA Forum.
Weâll be publishing eligible entries[1] to the Cluelessness Critiques Essay Competition on the Forum for discussion.
Read the entriesThe comment competition
While the essays are being judged, weâd love for this event to kick off more valuable discussion of Anthonyâs Sequence, and the problems for impartial altruists it raises (for a refresher, read this summary).
To help with this goal, weâre hosting a comment competition throughout the week. The best comments, according to me, @Anthony DiGiovanni đž and @Will Aldred, will be awarded in multiples of $100, up to a total of $2000[2].
Weâre looking in particular for comments that:
Show a deep understanding of Anthonyâs sequence and the problem of unawareness,
And advance the conversation, by introducing further considerations, clarifying key points, making a new critique.
We may also award comments simply because they are very helpful for the discussionâi.e. comments that clear up a persistent confusion.
Awards will look like this:
How to use this discussion thread:
This is a place to post any questions or comments you have after reading the sequence, or the competition entries. This could be in the form of:
Clarification questions. If something confused you, it probably confused someone else.
Statements that youâd like to debate with others. Consider adding a poll to your comment if you are making a bold statement.
Quick arguments, even scrappy ones, that another commenter could disagree with or take further.
- ^
I.e., those that engage with Anthonyâs sequence, and offer a critique or solution. Iâll also refrain from publishing some of the more unreadable (because AI-generated) pieces, since they wonât be interesting to the Forum audience.
- ^
Same caveats here as listed in the disclaimer section here. Additionally, we may award less than the total $2000 if we donât find enough comments that we consider worth awarding. Note also that we will endeavour to award comments throughout the week, but some awards may be given the week after.
Why complete cluelessness is counterintuitive to me
I mean, specifically, this kind of pervasive cluelessness, where you canât justifiably decide between any two actions. It seems that such pervasive cluelessness is in tension with the idea of instrumental convergence. Iâm personally sympathetic to imprecise probabilities, but intuitively Iâm not convinced that cluelessness is that bad that we canât justify even basic learning and other (supposedly convergent) instrumental strategies, such as (at least) those in the low-footprint capacity building category.
And I think, if we have to show that cluelessness is at least not that pervasive, it seems useful to look at the most seemingly absurd cases. Being clueless about whether epistemic improvement is worthwhile is one of them. If we (or other aligned agents that we create) can in principle be non-clueless about some specific class of strategies, for instance, why wouldnât we at least sometimes be justified in taking actions that would make us non-clueless about them?[1] (hence, being additionally non-clueless about whether this transition is better than doing nothing or pursuing some other alternative with a small opportunity cost) And if we are indeed justified in doing so, why canât we derive some broader system of instrumental strategies from this fact?[2] (hence, potentially, having even more non-clueless states) It sounds strange that even strategies aimed at becoming non-clueless wouldnât be better justified than doing something random.
I have to admit, the situation is much less obvious than I first thought. Interestingly, some imprecise probabilists may have to pay to avoid free knowledge (unlike precise ones). And there is a thing called dilation: new information may make your intervals wider than before. But most importantly, knowledge is not actually free. And itâs yet unclear to me how we should account for long-term changes in the structure of our entire decision trees in general. This seems to involve the topic of sequential rationality.
Anyway, why is impartial altruism so different from other value systems in this respect?[3] DiGiovanni writes:
In principle, we can come up with examples where learning goes wrong, but itâs not yet clear why we should treat them as severely undermining the whole idea of epistemic improvement.
But we donât have to think that our future selves are such perfectly coherent extensions. They are indeed different, but they may differ to a degree small enough for us to make a justified decision.
So far, it is not clear why impartiality is special. Learning seems to have broadly similar structures across many domains. Impartial altruism doesnât seem to be an exceptional domain such that some magical demons suddenly appear from nowhere to make you go mad if you know too much.[4] Then why arenât there some learning strategies available to us that are more justified than doing nothing?[5]
This seems to be relevant when deciding between some direct interventions (such as donations), but how does the associated cluelessness infect learning itselfâwhich is supposed to help us compare them? Cluelessness may emerge here if we donât know how to compare learning with its alternative, which is possibly higher-EV. But what if the alternative we consider is something which has low direct impact on your environment, such as resting? Then by choosing to learn, you may miss something valuable if it lies further down your decision tree and resting makes that opportunity more accessible. This is certainly normal, but isnât unique to impartial altruism.
We shouldnât always prefer learning to doing nothingâitâs even possible that we should do it less often than some other things. This is because successful approaches to epistemic improvement should mix learning with other things, such as resting when tired, eating when hungry, not dying in the process, etc. But isnât it possible for us to identify some combination of such low-impact actions that is better than, for instance, indefinitely pursuing just one of them? Empirically, this seems possible in many domains, and we may be able to identify some similarities across them. So again, it seems unclear, what is it about impartial altruism that makes things more clueless hereâto the degree that we canât even decide whether we better be alive and learn things.
It will be helpful to have greater clarity about the potential sources and mechanisms of cluelessness (and their relative strength) in such specific most-absurdly-looking cases.[6]
Alternatively, we may also consider the possibility that all realistic agents with sufficiently ambitious value systems should be clueless. Like, should realistic impartial paperclip maximizers be clueless too? This doesnât necessarily imply unconditional cluelessness, but it would also be very pessimistic.
For example, maybe we could look at strategies that increase the probability of successfully creating aligned successors while reducing the probability of failure.
Alternatively, maybe we should be clueless about some small-scale areas too, not only about impartial altruism? If so, it might be useful to study such areas to identify where exactly cluelessness starts to emerge.
In principle, we may consider a possibility that there is such a hypothetical level of self-improvement, after which something strange happensâas in the case where the simulation hypothesis is true and the masters of simulation decide to stop you from being too successful :) Or perhaps there are some cognitohazards waiting out there in philosophy that will make you abandon your impartial altruism for some reason. But do these possibilities seem severe enough to make us really clueless about learning?
To be clear, there is a huge difference between learning some random stuff and deliberately trying to learn more about possible crucial considerations, for instance. Iâm mainly concerned with the decision-relevant knowledge here. Many things may be helpful when deciding between interventions directly aimed at helping our moral patients, and many other things are also helpful even if they are less directly relatedâsuch as learning how to live longer, or how to acquire more resources to have more impact. And âdoing nothingâ here is, of course, doing somethingâbut with a (relatively) small direct impact on what happens around us.
By the way, maybe we should use less-demanding decision rules than maximalityâsuch as in the graded approach to comparison.
DiGiovanni states:
A single counterexample of a pair of actions where one is c-preferred over the other would suffice to disprove his claim.
It seems noteworthy that none of the solutions attempted this.
I donât think thatâs so surprising. There are obvious candidate counterexamples, e.g.:
If you think youâre justified in c-preferring the SWP donation, then Anthonyâs claim is disproved. But Anthony has said he doesnât consider this a counterexample, so it seems unlikely that offering more candidate counterexamples would move the debate forward.
Thatâs especially so since I think the scope of Anthonyâs claim is intended to rule out other candidate counterexamples, e.g.:
That seems true to me, but I think this pair of actions falls outside the scope of Anthonyâs claim. Heâs talking about actions with effects that arenât so tightly limited in space and time.
So the debate calls for something more than just a bare counterexample. As Anthony says in another comment, I try to give that âsomething moreâ in my post. Donating $5 to MAWF is justifiably c-dispreferred to some mixed action, as is any other action that fails the âflanking variantsâ test. That likely rules out almost all actions.
Yeah, I think itâs easy to find a counterexample to âunconditional cluelessnessâ (i.e., even in the box situation), but much harder to find one relevant to what you and I should do in our present non-simplified situations (which is presumably what Anthony meant for us to discuss).
Agreed. I actually wish people would try to find one counterexample instead of trying to prove too much. This might allow us to specify some cruxes.
Hi Jim. I agree trying to come up with counterexamples is useful. Below are 3 potential counterexamples I have given. @Anthony DiGiovanni đž does not consider them counterexamples (see Anthonyâs replies for details).
1st example, which Elliot already quoted in this thread.
2nd example.
3rd example.
Here is a 4th example. Consider these 2 actions:
Killing the 100 people who are expected to decrease the most the uncertainty about how to compare the expected value of different actions. For example, Bob Fischer who has worked on decreasing uncertainty about comparing welfare across species.
Grating 10 M$ to the 100 people above (100 k$ per person, but the grant size could vary). The 10 M$ would otherwise be spent torturing people as much as possible.
I think we are justifyed in c-preferring the 2nd action. I believe Anthony disagrees
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is âall things considered, we should prefer the second actionâ â very difficult to deny! â and the bailey is âwe should c-prefer the second actionâ. This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is âstudy some altruistically irrelevant branch of academic philosophyâ vs. âtry to prevent AI misalignmentâ. The latter only looks clearly preferable to me if itâs c-preferable. (I guess this is what Benâs comment is getting at.)
Hi Anthony.
Thanks for confirming. I would be surprised if there is any counterexample you find persuading.
(Just to be clear, thatâs a contingent matter. I donât find any of the counterexamples offered so far persuasive because I donât think they adequately engage with my arguments for P3.)
Do you think youâve ever taken an action (or sequence of actions) that, with your current understanding of cluelessness and unwarenesss, should have been ex ante c-preferred to another (or doing nothing, specifically)? Like thinking a bit more about specific backfire risks or cluelessness, or breathing?
EDIT: Also some more weirder things, like not kicking a puppy when given the chance and no one else would know. Or, say, if youâve already stepped on a snail and itâs clearly going to die, should you put it out of its misery?
I think no. Basically, when I really internalize how dwarfed every actionâs cosmic-scale consequences are by off-target effects, ânoâ feels very common-sensical to me. (Cf. this paper on how âsimple cluelessnessâ is fake.)
I think I wouldnât be clueless about c-preferability in Elliottâs âtrapped in a boxâ example, but canât think of any real-world case analogous to this.
Follow-up assuming youâll answer ânoâ, Anthony: same even ex post, right?
Thanks Vasco, this is helpful.
I would be interested in a version of this contest attempting to answer something like âIf you believe Vasco that donations to SWP are better than torture, should you also believe that AI alignment is better than misalignment?â
I think that is what Richard Chappell is referring to when he mentions âradical skepticsâ here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).
Hi Ben. That makes sense.
Likewise.
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer âlow-footprint capacity-buildingâ over âdoing nothingâ. And this post more generally argues that âdonate $5 to Make-A-Wish Foundationâ is c-dispreferable to some mixed action (maybe thatâs not specific enough for what you have in mind). (I donât yet buy either of these arguments, though.)
Not sure whether new posts count for the commentary competition, but this seemed too long to paste here as a literal comment!
The maximality rule is too demanding. Under maximality, we canât even say that [1, 10 000] is better than [-10 000, 1.0001]. Are there any good reasons to avoid more permissive rules even in cases like this?
Of course, if your UEV intervals overlap and you choose one of them, you may make a mistake if your idealized version would choose the other one. But choosing at random is no better in this regard.
What seems important is how costly such mistakes are, i.e., how the overall performance of your choice rule compares with that of other rulesâsuch as choosing at random.
Can these performance measures be defined without additional assumptions about how values are distributed within UEV intervals?
Thatâs not quite right. Maximality says an action A is impermissible when some alternative B has higher EV on every probability function in your representor. And that can be true even when Aâs and Bâs EV ranges overlap.
Example:
If A had the same range but sloped the other way, then it would be permissible by maximality:
So to figure out whatâs permissible under maximality, we canât just look at ranges. We need to look at the representor.
Yes, but this means that you know something additional about the structure of the representor, not just its range. Iâm asking whether we can do better than maximality even for intervals alone, without adding any further details.
(By the way, pictures are broken.)
UPD: Though, you probably mean that we do know some additional structure for EV if we look at how it is constructed from probabilities and utilities, for which we have just intervals without structure.
Thanks.
Oops, pictures should be fixed now.
And yes I think often we know more than just intervals of EVs. For example, we know whether the EV of some action increases or decreases with the probability of some proposition X.
Just sharing my quick takes on how to navigate cluelessness with a rather long-term strategy here. It doesnât get to predictability land but to enough insight to move forward carefully despite vast uncertainties.
Reading through all the essays and comments this week, Iâve become much more convinced that cluelessness is pervasive and standard ways of handling it fail. I agree with the normative and conceptual premises.
But the empirical premise is untrue. In particular, it is untrue that for any pair of actions, the actionsâ consequences are too coarse to be compared.
Consider the following counterexample: I could either donate 10% of my income to GiveDirectly or take a nice vacation. Is it true that our knowledge of the consequences of these actions is so coarse-grained that the comparison between them is indeterminate? I donât think so, even though in principle the same concerns about the catch-all and unawareness apply.
The âbest guessâ justification for donating is simple. My lifetime income is very likely to be in the top 5% globally; under all plausible theories of utility, marginal utility declines with income; thus, the marginal transfer from me to someone in extreme poverty improves aggregate well-being.
How might this be incorrect or imprecise?
The recipient of the donation could use it for something bad. (This is analogous to the meat-eater problem, but given my own consumption of animal products the meat-eater problem itself doesnât apply.)
Catch-up growth in recipient countries could lead to an unstable multi-polar world order, which eventually causes a catastrophic WWIII.
Because I donât go on vacation, I donât accidentally step on a butterfly I otherwise would have. The butterfly flaps its wings and causes a tornado that kills 50 people. Also, those 50 people were all wild animal suffering researchers.
Everyone else is a p-zombie; only my consumption matters, and I should take more vacations.
In case this seems too glib, I promise I tried to come up with better examples (and asked AI). I do not think there is any plausible probability distribution under which the above are likely enough that taking the vacation is better than donating. I would welcome suggestions.
âPlausibleâ is doing a lot of work here. I donât have a good justification for why no probability distribution where these events are likely enough to change the outcome of the comparison is plausible to me. All I can say is that if we consider these type of events to be equally plausible to the âbest guessâ, then I agree with Richard Y Chappell that cluelessness considerations approach radical skepticism.
More formally, I would claim that the set S of all possible actions contains a non-empty proper subset C of actions over which an ideal impartially altruistic agent has complete preferences. The comparison between each action in C and each action in S\C is indeterminate. This still offers a lot of guidance on impartial altruistic action, because I would claim we all have more than one action within C available to us.
There is still a lot of debate to be had and research to be done to figure out which actions are in C and which arenât (see Jim Buhlerâs âThe Train to Arbitrary Landâ for a good elaboration of this problem.)
Still, this offers a meaningful change from previous cause prioritization discussions, which generally assumed that all actions are in C. Practically, these considerations push me towards near-term interventions, whose sign is clear and comparisons of which with alternative uses of resources I believe to be determinate.
Here are two potential reasons I find far more plausible than yours, fwiw:
- the GiveWell donation increases farmed animal suffering more than it increases human welfare.
- it slightly increases the total human population, which slightly accelerates human progress and hence increases AI x-risks. The difference is small but the long-term future may overwhelm the short-term benefits of GW so much that this is enough.
The thesis that we are clueless about the overall sign of donating to effective global health and development charities is actually one of the most consensual in the cluelessness literature (see Mogensen 2021; Kollin et al. 2025, §1; Greaves 2016).
Longtermist causes are those that generate significant disagreement (see, e.g., this overview and refs therein).
Thank you for the response, and in particular for the references for further reading!
Iâm curious to hear more about your reasoning for those examples, or why they seem especially likely under some plausible distribution. (For what itâs worth, Iâll clarify the hypothetical donation goes to GiveDirectly, not GiveWell. Also, I eat meat, and my meat consumption is sensitive to my income.)
I agree that it is possible for them to occur, but there is no probability distribution I would include in my representor under which they are sufficiently likely to occur that the comparison between actions is indeterminate.
Iâm not sure how helpful it is to dwell on specific examples, but this is the disagreement I have with many arguments for cluelessness about neartermist causes.
The story given in Mogensen 2021 of the priest saving Hitler from drowning as a child is illustrative. This was the right thing to do. There have been billions and billions of children; a handful have had both the opportunity and desire to commit horrific genocides. I donât believe there is a plausible probability distribution that make it so likely an anonymous child will commit genocide that it is better to let them drown.
Oops sorry for the GW-GD confusion, but yeah, this changes nothing to my point I think.
This seems hardly defensible.
- The average human being (including in poor countries) contributes to the farming of so many animals throughout their lifetime (at least in expectation). You would have to be astonishingly confident that the welfare of the animals we eat do not significantly matter for your above conclusion to follow.
- The number of far-future lives we indirectly influence might be astronomical such that long-term effects (almost) always dominate. I donât see on what basis you can exclude this possibility from your probability distributions, given the arguments given here and refs therein.
(I donât recall this specific example from Mogensen and donât have time to dive back into it, so I wonât comment on that, sorry.)
The problem of Cluelessness has been grappled with for millennia; two relevant examples:
1. The story of Khidr in the Quran, who, through a series of nonsensical or seemingly evil actions, demonstrates to Moses how we are Clueless about consequences.
2. The Bhagavad Gita,which advises disregarding consequences altogether:
To action alone hast thou a right and never at all to its fruits; let not the fruits of action be thy motive; neither let there be in thee any attachment to inaction.â
The solutions end up looking like Deontology or Virtue Ethics, which Anthony (inadvertently?) references in his summary: âOther values and moral norms still matter to us, for example, rules like avoiding dishonesty or virtues like compassion.â
Anyone grappling with the possibility of Cluelessness would do well to consider prior art.
For the sake of provocation: How is managing Cluelessness in Philanthropy any different from managing Cluelessness in, say, running a neighborhood coffee shop?
There is an unknowable amount of cause-effect relationships involved in running a business. Is there anything unique about Philanthropic modeling? Modeling dynamical systems is, in general, notoriously hard.
I wrote a few paragraphs fleshing out my question here if anybody wants to respond to me more in full.
This is briefly addressed by DiGiovanni here and there (his cluelessness X near-term AW post is also relevant), and I discuss some aspect of his point in the first ref a bit further here.
Awesomeâthanks Jim. Reading through some of these now