Researcher at the Center on Long-Term Risk. All opinions my own.
Anthony DiGiovanni đ¸
Good question! I think âother theoretically possible aggregations of all or most of the possible consequences of A and Bâ would also suffice, yeah. (Of course, if we ourselves canât specify what this alternative is, we have our work cut out for us if weâre gonna argue that we should expect our idealized self to prefer A over B on this basis.)
Interesting, thatâs helpful to know.
Not a comprehensive reply, but: I think many of the examples youâre talking about are arguably cases of coarse awareness. People were coarsely aware of the potential backfire risks earlier on, but (arguably) the reason they didnât give these risks enough weight was that they didnât have a more fine-grained awareness of the specific causal pathways. I think such cases count as evidence for the pessimistic induction.
Thanks Vasco. Iâve summarized my reply on LessWrong here (figured that this might be of (more?) interest to LW readers).
Iâd guess weâre getting slightly better, yep. I might put less weight on the evidence you mention, than on: âWeâre living in a period of really unprecedented AI progress, seems like that puts in a better position to reason about the mechanisms governing the far future than ever before.â
I think itâs mostly (1), but Iâm open to something like (2) or (3) as well.
(Following (1):) There is in principle some (a) amount of information that non-ideal agents could attain about the cosmos with non-Pascalian probability,[1] + (b) a priori modeling and induction we could apply to that information, such that we wouldnât be clueless. So I donât think we need to observe the target variable, or empirically âvalidateâ the theory, to be non-clueless.
But the bar to achieve such an (a)+(b) seems very high, because:
If we do try to empirically validate the theory by appealing to calibration on near-term proxies:
I indeed donât see why we should expect such calibration to transfer, up to the degree of precision we need to escape cluelessness (sec. 2.3.1.1). This bites even if, say, we use AI to get much more calibrated on ~years-long time horizons.
If we donât, and instead try to argue conceptually that the theory captures enough of the relevant considerations in fine-grained enough detail:
So my suspicion is that yeah, weâd still be clueless given the kind of theory you mention. But I find it hard to say, because I canât imagine exactly what âcomparably goodâ looks like, concretely. I appreciate that thatâs hard to spell out on your end.
Maybe sufficiently advanced AI could get around this. Maybe not, e.g. if âthe universal priorâ is irreducibly imprecise, or if (following (3)) information about simulators or causally disconnected worlds is fundamentally inaccessible.
(Iâm happy to unpack any of this more if useful, not sure if I answered your question properly!)
Hi Ben, I like the spirit of this question, though Iâm not sure itâs the most relevant formulation. Thoughts on that, before I answer your literal question:
To get clear on terms:
Iâm assuming by âCIâ here, you mean something like your precise 90% (or whatever) confidence interval for your idealized selfâs EV of the intervention (as per Premise 1).
By ârobustâ, I donât mean a narrow /â strictly positive CI. I mean that the verdict âthis intervention is positive âin expectationââ isnât sensitive to arbitrary choices about how to factor in the considerations weâre unaware of.
I think itâs fair to say most people involved in early AI risk advocacy considered their work robustly positive in that sense.
This definition of ârobustâ is what matters for Premise 3. Because P3 says, our verdicts about interventions having positive âEVâ are sensitive to such arbitrary choices â even if we admit weâre very uncertain (i.e. we have wide CIs).
So, if weâre asking whether a given âsign flipâ counts as evidence for P3, I donât see why the bar should be ânarrow CIs with opposite-sign center points before and after the considerationâ.
If weâve discovered a consideration that flipped us from âwide CI centered at a positive EVâ to âwide CI centered at a negative EVâ, isnât that some evidence that our initial âpositive in expectationâ verdict wasnât robust, in the sense above? (And hence inductive evidence that our current âpositive in expectationâ verdicts arenât robust (i.e., evidence for P3), as argued here.)
Anyway, Iâd agree that the clearest evidence for P3 would come from sign flips that meet your bar. Maybe the small animal replacement problem? Iâd guess lots of people who care about animal welfare thought that getting people to eat less beef was clearly good before being aware of SARP, and think itâs clearly net-bad after being aware of SARP. (Itâs harder to come by examples of sign flips by your def for longtermist causes, because non-clueless longtermists typically agree that we should be very uncertain about the far future. But per the above, this is to be expected if weâre clueless.)
Hi Dan/âFable, thanks for the critique! The three most important problems I see:
1. Claiming that the unawareness argument only has practical implications if thereâs a privileged âdefaultâ
The sequence argues carefully for incomparability. It cannot argue that incomparability favors inaction, because if A and the status quo are incomparable, the status quo is not better. Yet the practical gloss everyone puts on the conclusion resolves every incomparability toward the default
As discussed here and in the introduction of the sequence, my claim was never âwe should default to inactionâ. Itâs that we have no impartial altruistic reason to favor any intervention over any other option, including âinactionâ.
Your response to this is that if we donât favor a default under incomparability, âthe argument has almost no practical biteâ. But we can have reasons for choices other than impartial altruistic ones â see here.
(This point is upstream of one of your replies to objection #2: âso everything becomes incomparable with everything, the argument again supplies no reason to resolve toward a defaultâ.)
2. Conflating two kinds of awareness growth
In your response to the objection âThe considerations that matter most may be ones no engagement revealsâ, you say that the evidence for P3 is âWeâve become aware of new considerations through active engagement.â But âbecoming aware of new considerationsâ is a much lower bar than âbecoming aware of all the considerations that the sign of an actionâs âEVâ is sensitive toâ. You need the latter for your critique via Model 3 to work. This is important for the following:
3. Neglecting unawareness about/âimprecision in the time horizon
Your argument in Model 3 seems to be: Suppose you have T time steps to (1) actively grow your awareness by repeated exploration and then (2) exploit the strategy that does best w.r.t. your credences over this fleshed-out awareness set â before you die. Then (1)+(2) would beat âstick to the familiar domainâ. Iâm happy to grant here[1] that this is true for some T.
But we donât know T.[2] If our beliefs about T are imprecise enough, we canât say whether (a) the benefits of eventually cashing in on our grown awareness outweigh (b) the potential backfire effects of actions we take to grow our awareness. To meet the bar of awareness growth noted in (2) above, T might need to be very large indeed.
- ^
This is just for the sake of argument. I think the model is importantly unrealistic in some ways I donât cover here for lack of time.
- ^
Suppose we instead say âConditional on living forever, the upsides are unbounded, so the utility from this case swamps all the finite-T cases.â One problem is that if weâre going to allow for unbounded upsides from an infinite T, we should also consider unbounded downsides. (You acknowledge this possibility in objection #2, but what matters is whether your critique via Model 3 actually works, not what the sequence as written says.)
- ^
Hmm yeah maybe I shouldnât have fully endorsed Benâs summary. I think many forms of bracketing are impartial in the sense that they donât arbitrarily favor some moral patients over others. But the forms of bracketing Iâm aware of are either (1) not âimpartialâ in the sense that some moral patients/âconsequences are bracketed out for not very well-motivated reasons, or (2) not action-guiding.
That latter definition of impartial might be confusing though, so in general Iâd just list my specific dissatisfactions with different forms of bracketing, âimpartialityâ aside.
AnÂnouncÂing the Safe Pareto ImÂproveÂments (SPI) FunÂdaÂmenÂtals Program
Hey Ben â your understanding is correct. Option 2 is meant to allow for the possibility that the inference from P1-P3 to Conclusion, as stated in my summary post, is logically invalid. (Of course I think thatâs unlikely, but philosophy can be subtle.) Does that clear things up?
AMA: AnÂthony DiGioÂvanni, auÂthor of the âChallenge of UnawareÂnessâ sequence
Hi Toby, glad it was helpful!
Fair question â the idea is:
P1 claims that c-preferences can only be normatively justified by arguing that one option has better âexpectedâ consequences.
The objections youâre referring to are all implicitly of the form, âWe should always c-prefer either A or B (or be indifferent) because of [some reason other than a comparison of the âexpectedâ consequences].â
E.g. in the first row, [some reason other than a comparison of the âexpectedâ consequences] = âwe always have to choose somethingâ.
Iâll add a note on this to the post itself.
(Due to time constraints I expect I can only give brief replies/âclarifications, going forward. I hope a full read of the sequence will suffice, though I realize itâs quite long, sorry!)
But you donât need to restrict yourself to concerns about the whole future lightcone to run into this problemâat the foundational level this is true of every statement. ⌠I donât see why we should do so with credences (which are of course usually non-foundational statements)
(See my last para for the âfuture lightconeâ thing.)
I donât understand your Munchausen trilemma argument yet. You say credences are âof course usually non-foundationalâ. Agreed! Thatâs exactly why I think our choices of credences require deeper justification. (Whereas foundational things, like Huemerâs âseemingsâ, donât.[1])
forecasting 0.1234567% chance of rain if the extra precision was actually decision-relevant
The extra precision might be âdecision-relevantâ in the sense that: if you were justified in a credence of 0.1234567% + 0.0000001%, you should choose A, and if you were justified in a credence of 0.1234567% â 0.0000001%, you should choose B. But the whole question is why weâd be justified in the former vs. the latter, epistemically. (âI need to make a choiceâ isnât a justification for any particular option you choose.)
Your counterpoint seems to be that in some cases that feel sort-of- equal (and about which, in the cases you describe we actually have a lot of information), we might be inclined to give equal credence.
Thatâs not what Iâm saying, sorry â Iâm denying we should give equal credence. Please see my reply to a similar comment here, and section 3.2.1 and 4.1.1 of the sequence (you might need to CTRL+F some terms defined earlier in the sequence). If itâs still unclear, Iâm happy to try to explain further if you could point to particular passages that need clarification.
The precise EV approach is well evidenced in short-term decision-making
I donât know what exactly this means. If you mean âwe seem to be justified in using precise EVs in short term decision makingâ:
I think our beliefs shouldnât be precise in basically any real-world case, not just beliefs about the far future. (Sec 2.2)
So I think whatâs going on is simply that short term decisions arenât sensitive to the imprecision in the beliefs weâre actually justified in having. The principled difference from the far future case is that in the latter, our decisions are sensitive to the imprecision.
- ^
That is, they donât require deeper justification prima facie. Theyâre still defeasible.
Hi Arepo, thanks for sharing your cruxes here.
The argument I give against assigning precise credences is that itâs arbitrary â literally, you pick one precise credence over many others for no reason. To me, âyou have no reason to do this thingâ is a pretty darn strong argument. :) (ETA: I like the intuition pump in this very short post, if it helps.)
doesnât say why this means we shouldnât/âcanât pick credences according to our best effort.
Why does âour best effortâ need to be precise? Can you say more what exactly you mean? (If the intuition is that more precision = more information, I address that in the post.)
It also doesnât say why, if we can measure short term value, we shouldnât use that as a justification for our decisionmaking process and assume EV from events that we donât think we can assess is 0.
I address this in the unawareness sequence. I recommend reading the table in my summary post â the row with âEven if our impact is dominated by consequences weâre unaware of...â â for the high-level idea, and the links therein for details.
especially when weâre not given an alternative
Isnât this privileging the hypothesis? My claim is that we donât have a positive argument in favor of doing what the precise EV approach recommends (or fuzzier âbest guessesâ, either). If our best defense of that approach is âwhat else is there?â, that seems rather damning.
Ah, sorry, I thought you were making the first-order wager argument (Q3 here), but IIUC youâre making a metanormative wager argument as Toby suggested. I discuss why Iâm unconvinced of that here. (And as another commenter pointed out, this is supplemented by âWhy cluelessness mattersâ in the OP.)
taking your probability-weighted expectation of that range
I think youâre misunderstanding the framework. The whole problem is that we canât assign a (non-arbitrary) âprobability-weighted expectationâ. Thatâs the motivation for representing with a range rather than a single expectation.
(ETA: By default I plan not to reply further.)
I address this objection here (Q3), if I understand what youâre saying correctly. (Iâd recommend first reading sec. 2.1 of the post for crucial background on the epistemology, though, as I noted in another comment.)
(In general, I think you should not expect this post to be a self-contained explanation of the argument by any means. Itâs a high-level summary.)
Iâm not saying we have information that updates us in a particular direction about the bias. Iâm saying we have information suggesting various different directions, and itâs ambiguous what the update should be overall â which is fundamentally different from âno updateâ. I strongly recommend reading sections 2.1 and 2.4 of this post, as well as 3.2.1 of this post, to understand the epistemology thatâs at play here.
(ETA: The final paragraphs of sec 4.1.1, linked right after the part you quoted, also discuss this point.)
(Edited for tone)
Sorry, I donât understand. The snippet I quoted â about acausal stuff and simulations â is whatâs at issue in this discussion.
Regardless, Iâm still interested in where you object to my response to Extrapolation. Could you please say more on that?
I mean that we have what I call âcoarse awarenessâ here: we conceive of crude groups of possible worlds, rather than possible worlds specified in fine-grained enough detail to assign them precise values (wrt impartial altruist axiologies). See also here for some examples. Happy to unpack more if those sections donât answer things!