Researcher at the Center on Long-Term Risk. All opinions my own.
Anthony DiGiovanni đ¸
I think no. Basically, when I really internalize how dwarfed every actionâs cosmic-scale consequences are by off-target effects, ânoâ feels very common-sensical to me. (Cf. this paper on how âsimple cluelessnessâ is fake.)
I think I wouldnât be clueless about c-preferability in Elliottâs âtrapped in a boxâ example, but canât think of any real-world case analogous to this.
Yeah, spacetime bracketing or some more formally nebulous form of bracketing along the lines here.
Thanks!
I donât think the arguments against the capacity-building strategies I discuss are as strong as those in their favor
My core objection to capacity-building in the sequence is: For any concrete capacity-building strategy, our understanding of that strategyâs full range of possible consequences is extremely coarse. And it seems very plausible to me that these consequences will include large off-target effects on, e.g., lock-in events â in which case, capacity-building strategies inherit the non-robustness of strategies aimed at influencing lock-in events. This is for pretty similar reasons to how the off-target effects of AMF donations seem to dominate. I donât yet see why you think otherwise.
(So in particular, I donât think your responses in your appendix to specific backfire risks I mentioned in the post address this core objection.)
Moreover, much of the value of building capacity is the value of being able to act on considerations we arenât yet aware of. So, in my view, unawareness bears asymmetrically on these strategies rather than neutrally (I realize this latter point is stated very briefly and needs further development).
Yeah, Iâd be interested in seeing this spelled out a lot more sometime. Per the above, even if a strategy might enable us to act on considerations we arenât yet aware of, this doesnât help us with cluelessness if the strategyâs impact is still very plausibly dominated by off-target effects.
Iâm one of the judges of the competition. My comments shouldnât be taken as a full review of a post. And, unfortunately, I wonât have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered!
Iâm confused by your responses to the vignettes and thought experiments you quote in this post.
The insensitivity to mild sweetening example isnât meant to be an argument for imprecision. Itâs in the section of the post thatâs focused on spelling out implications of imprecision.
Re: the pause AI example: I donât think your response engages with the core intuition the example is getting at, which is that giving a determinate answer seems arbitrary. I guess you donât share that intuition â fair enough â but it seems misleading to present the example as if itâs intended to pump a brute intuition (of incompleteness) rather than the underlying intuition that justifies it.
Re: the vignette in unawareness post #1, you say: âThat quote suggests that he doesnât consider this vignette to have persuasive force.â But I was saying that the vignette alone doesnât establish the whole conclusion. And why should it? (By that point in the sequence, I havenât gotten into all the rest of the argument.) The role of vignettes generally is to put concrete color on an otherwise abstract argument, not to substitute for the argument.
(Just to be clear, thatâs a contingent matter. I donât find any of the counterexamples offered so far persuasive because I donât think they adequately engage with my arguments for P3.)
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is âall things considered, we should prefer the second actionâ â very difficult to deny! â and the bailey is âwe should c-prefer the second actionâ. This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is âstudy some altruistically irrelevant branch of academic philosophyâ vs. âtry to prevent AI misalignmentâ. The latter only looks clearly preferable to me if itâs c-preferable. (I guess this is what Benâs comment is getting at.)
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer âlow-footprint capacity-buildingâ over âdoing nothingâ. And this post more generally argues that âdonate $5 to Make-A-Wish Foundationâ is c-dispreferable to some mixed action (maybe thatâs not specific enough for what you have in mind). (I donât yet buy either of these arguments, though.)
Alternatively, if you posit incommensurable values then you should probably reject P1.
Itâs a fair point that deference principles come into conflict with prospective reasons in Hareâs case, and the prospective reasons argument seems really plausible there. I donât feel very confident, but hereâs how Iâm thinking about this:
Deference to the idealized self isnât actually necessary for my argument.
I wrote P1 in those terms as one way of making precise the idea that we should account for hypotheses weâre unaware of, when weighing up prospective reasons.[1]
But we could instead just directly say, as I do in the sequence: Insofar as weâre impartial altruists, (1) we want to in some sense weigh up all possible outcomes by their value and plausibility; and (2) our values are defined over metaphysically possible outcomes, not over the extremely coarse-grained versions of such outcomes we conceive of. So when we (perhaps roughly, informally) compare actions as impartial altruists, we should do so in a way that tries to account for fine-grainings of possible outcomes that we havenât explicitly conceived of.
Still, Iâm pretty sympathetic to the âdeference to the idealized selfâ idea, so how do I reconcile this with Hareâs case? Well, itâs not clear to me that this particular deference principle is substantively analogous to the deference principle invoked in Hareâs case, even though theyâre structurally similar.
The deference principle Iâm invoking is, roughly: (D1) My current evidence doesnât warrant considering A c-preferable to B, from the perspective of an agent who assesses that evidence the way Iâd want to if I didnât have my computational (etc.) limits. So I, in my current decision situation, shouldnât consider A c-preferable.
The principle in Hareâs case is, roughly: (D2) If I had more evidence (thereby putting me in a different decision situation), I wouldnât be warranted in considering A c-preferable to B no matter what that evidence is. So I shouldnât consider A c-preferable.
It seems plausible to me that we should endorse D1 but not D2, because:
If you ask me why I donât c-prefer A over B, and I invoke D1, my answer only makes reference to features of the decision problem Iâm actually in â itâs just that those features are assessed from a less computationally limited perspective.
Whereas if you asked me why I donât c-prefer taking the sugar, and I invoked D2, I wouldnât be telling you why my actual decision problem fails to recommend taking the sugar. Iâd be telling you that versions of myself in different problems donât c-prefer taking the sugar. Theyâre more informed versions of myself, yes. But those versions of myself have different reasons!
(Compare to how itâs coherent for an agent who endorses causal decision theory to two-box when âdropped intoâ Newcombâs problem, yet to commit to one-box in the future if the prediction of their decision hasnât yet been made. The former agent has different reasons than the latter.)
(Anyway, again, not confident in this, and I think the more important point is that the cluelessness argument as such doesnât depend on deference principles.)
- ^
In particular, this framing doesnât require that humans even have well-defined âhypothesesâ in our epistemic state. I donât think the alternative framing I use in the sequence itself requires us to have unrealistically precisely defined hypotheses, either. But this was a bit of a sticking point when discussing the problem of unawareness with one thoughtful interlocutor â which is (AIUI) what led to Jesse writing the post on deference to the idealized self that Iâm drawing on.
Hi Richard, thanks for the reply! Iâll reply to your critiques of two premises in separate comments.
DiGiovanniâs case for imprecise credences rests heavily on the intuition that it would be objectionably arbitrary to settle on any precise number given the breadth of our uncertainty in the face of conflicting considerations. I get the intuition: from the inside, Iâm not sure how I would go about determining a precise credence in non-arbitrary fashion. But I think this confuses a procedural question (by what process should we settle on a credence?) with the substantive epistemic question of what credence is in fact most warranted. The latter question might have a precise answer even if we canât easily articulate a general high-level process that will reliably yield this correct answer.
I donât understand your critique yet, can you explain this more? The claim âPrecise credences are arbitraryâ doesnât assume we need to âeasily articulate a general high-level processâ. Iâm making the substantive epistemic claim: in basically any particular real-world case, when I compare one precise credence to another, it doesnât seem to me that the net balance of reasons favors one over the other. These reasons can and do include hard-to-articulate intuitions and such â my argument is that even when we account for all those, choosing one precise credence seems totally arbitrary. Denying that seems like an extraordinary claim, to me, see e.g. the intuition pump from this post.
Thatâs all about P2a, separate from whether judgments like âA is all-things-considered c-better than B, independent of precise credences per seâ are arbitrary. I think itâs important to separate these things, as I do with the split of P2a vs. P2b.
Iâm not sure how to make progress on our disagreement about P2b, other than to say: I think youâre too hastily jumping from âwe need to take some things as foundational (not requiring further justification), on pain of radical skepticismâ to âthese particular gestalt judgments about c-preferability can reasonably be taken as foundational (+ not undermined by defeaters)â. The latter seems very implausible to me, for the reason I give in the table here (response to âSome actions are obviously c-preferable to others...â).
thereâs some reason, that does not come to mind rn, why bracketing over those would fail to provide action guidance?
The worry is very similar to the worry about bracketing on persons: identity of person-moments is super fragile.
Letâs say I donate to AMF. There are some far-future person-moments for which I am clueful that donating to AMF seriously harms them, because these person-moments only exist in the possible worlds corresponding to [insert far-future backfire risk here]. And we can say the same substituting in âbenefitsâ for âharmsâ. We then end up a maximal bracket-set on which the donation is good, and one on which the donation is bad.
(If this needs more unpacking, Iâm happy to do that!)
(I endorse this comment. Not exactly sure Iâd put the worry in these terms, and I wonder if Titotalâs crux is normative neartermism â but if so, Iâd just say that thatâs out scope because itâs not impartial.)
Iâm one of the judges of the competition. My comments shouldnât be taken as a full review of a post. And, unfortunately, I wonât have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered!
I think this post contains a few important misunderstandings of the argument. Sorry this wasnât super clear from the summary; the argument is very nuanced, and itâs important to look at the full sequence to understand how the epistemology Iâm arguing for works.
P1 says we need to expect our idealized selfâs (perhaps imprecise) EV(A) to be higher than EV(B), not that we need to know this.
In your coin flip example, the idealized self might know the laws of physics in such detail that theyâre certain the coin will land heads. But they might also know these laws in such detail that theyâre certain it will land tails. From your perspective, these possibilities are just as plausible, so thatâs why youâre justified in a credence of (roughly) 0.5.
So the âprefer A only if you expect your idealized self to prefer Aâ reasoning generalizes regular old EV reasoning. I might say more on this in another post, but basically, the reason I used the former framing was to (1) avoid P1 requiring us to assign literal EVs ourselves, and (2) have P1 require that we shouldnât just ignore hypotheses that we havenât conceived of (but our idealized self has).
(I recommend reading Jesseâs post on Ideal Reflection in full, for more context.)
Several of your proposed counterexamples to the premises, if I understand them correctly, are claims that we can still compare two options with imprecise âEVâ in principle. Which I donât deny.
Your âq<p and Q<Pâ example is a case where we can compare imprecise âEVsâ.
Same with âI can have beliefs about me being able to lift less than a really buff guy even if I am uncertain about how much each of us could liftâ. This is not a counterexample to P3 â itâs a case where the quantity in question is not so coarse-grained that the consequent of P2 (âAâs and Bâs âEVsâ are incomparableâ) applies.
When I said, âC-preferability is (arguably) not something we can directly perceiveâ: in the context of the rest of the counterpoint, my point was that c-preferability needs to be justified by analysis of the possible consequences themselves (or an appeal to an intuition that we have reason to believe implicitly tracks such an analysis). Iâm not saying we need to directly perceive consequences in order to have beliefs about them.
On your âmartingaleâ comment: I think youâre implicitly assuming the only options are âupdate towards higher âEVââ, âupdate towards lower âEVââ, or âno updateâ. The point of the imprecise framework argued for in post #2 is that thereâs a fourth option, âupdate of ambiguous directionâ. Thatâs consistent with each precisification of the âEVâ being a martingale.
I donât see what difference the stochastic dominance point makes to the conclusion. Under severe imprecision (P3), it seems at least as hard to compare actions with the (imprecise) stochastic dominance relation, as with (imprecise) EV-broadly-construed.
Iâm one of the judges of the competition. My comments shouldnât be taken as a full review of a post. And, unfortunately, I wonât have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered!
Thanks for this. I have some sympathy to the critique of maximality as far as it goes. (Though personally Iâm not that compelled by the symmetry arguments for your proposed alternative, the midpoint rule.)
My worry is that the threat of indeterminacy recurs, at the level of vague degrees of ex ante betterness. When I try to weigh up the reasons in favor of any A vs. B w.r.t. their total consequences, fully accounting for unawareness, I really donât see why I should consider the âupper endpointâ larger or smaller in absolute value than the âlower endpointâ. I think the arguments in posts #3 and #4 of the sequence support this worry, though I didnât explicitly frame them in such terms and used the precise-endpoints formalism. You acknowledge that this is possible here:
Just as we can admit vagueness in the endpoints of our imprecise credences, we can admit corresponding vagueness in the value yielded by D. This makes it difficult to say whether a given option is better than the null action when D is close to 0, which seems quite plausible: cases where D â 0 are exactly the kinds of cases where betterness or comparative justification may be too unclear or too vague to discriminate either way.
But I would like to see more of an argument that this isnât our situation, perhaps building on what you say in your appendix about low-footprint capacity-building.
(If we are indeed in a situation where even the âendpointsâ are hopelessly vague, then I donât think your defense of first-order wagering goes through either. You can have insensitivity to mild sweetening under vague degrees of betterness; it doesnât assume maximality.)
Hi Impartial, thanks for this reply! Canât speak for Jesse, but I definitely agree with this:
It is important not to lose sight of the fact that overall indeterminacy is not to be avoided tout court. We should avoid it where, and only where, it does not actually obtain from an impartial altruistic point of view.
So, if bracketing (either TD or BU) is entirely non-action-guiding for the reasons Jesse mentions, that doesnât bear on whether bracketing is a good normative theory. But it does bear on your claim at the start of your post: âBottom-up, person-centered bracketing can thus preserve principled impartial action guidanceâ.
(As another judge of the competition: Seconded!)
Thanks Vasco. I think it would help a lot if you spelled out the premises more, because theyâre quite opaque to me as written. E.g. I donât know what âthe norm reveals a choice functionâ means. (I think this kind of use of jargon without giving context is a common failure mode of current LLM summaries.)
Also, if I understand correctly, âbehaviour maximizes a family of preference orderings...â means that the result only shows that we can represent an agentâs behavior as satisfying completeness. But my unawareness argument isnât about what our behavior can be represented as. The question is: When weâre comparing our options when making decisions in the first place, should we have complete preferences? Cf âWinning isnât enoughâ:
But what these arguments really show is that you are disposed to playing a dominated strategy if we cannot model your behavior as if you were a Bayesian with a certain prior and utility function. They donât say anything about the procedure by which you need to make your decisions. I.e., they donât say that you have to write down precise probabilities, utilities, and make decisions by solving for the Bayes-optimal policy for those.
(But let me know if Iâve misunderstood the result.)
The standard isnât either (a) or (b) exactly. I think no one, precise Bayesian or otherwise, has a complete standard for how to set credences. Iâd gesture at something like the example in this comment: try setting credences in a similar way to precise Bayesians, but whenever you find that it seems arbitrary which distribution you pick among many, include them all.
So insofar as I understand what is meant by âa distribution needs a sponsorâ, Iâm not committed to (b) by virtue of agreeing that we should have P(grey) = 1â2 in Joyceâs case.
In particular:
I think whatâs going on in your last paragraph is an equivocation between:
âIf you donât assign a precisely symmetric distribution over some set of hypotheses H, it must be because there is some respect in which the hypotheses are not symmetric.â
âIf you donât assign a precisely symmetric distribution over H, it must be because youâre explicitly aware of a pair of hypotheses in H that are not symmetric.â
(1) seems very plausible. But (2) isnât. It can be the case that Iâm not aware of the hypotheses in H, yet I have reasons (based on the arguments given in sections 3.2.1 and 4.1.1) to consider them not symmetric.
I think that from an impartial POV:
currently, probably[1] weâre clueless about the comparison of any two actions
itâs plausible that, if we had a lot more evidence and more developed conceptual models of the cosmos, weâd be non-clueless about the comparisons of some actions. I canât say âhow oftenâ this would occur, in the sense of how easy it would be to get such evidence and models. Seems really hard; see here.
- ^
Referring to my uncertainty about the logical implications of our evidence for whether or not P3 is true.
The sequence gives general arguments for this, especially sections 2.3, 3.2, and 4.1. Iâm not exactly sure what you find uncompelling about them. The worry is that severe coarseness makes the degree of justification so severely vague that we have all the same qualitative problems as under the maximality analysis.
Of course, the claim is defeasible by arguments to the contrary for some specific A vs. B. But the burden of proof seems quite high to me.