For example, compare (1) a researcher spends some time thinking about what happens if a cosmic ray flips a bit (or a programmer makes a sign error, like in the famous GPT-2 incident), versus (2) nobody spends any time thinking about that. (1) is clearly better, right? We can always be concerned that the person wonât do a great job, or that it will be counterproductive because theyâll happen across very dangerous information and then publish it, etc. But still, the expected value here is clearly positive, right?
If you arenât publishing anything, then sure, research into what to do seems mostly harmless (other than opportunity costs) in expectation, but it doesnât actually follow that it would necessarily be good in expectation, if you have enough deep uncertainty (or complex cluelessness); I think this example illustrates this well, and is basically the kind of thing Iâm worried about all of the time now. In the particular case of sign flip errors, I do think it was useful for me to know about this consideration and similar ones, and I act differently than I would have otherwise as a result, but one of the main effects since learning about these kinds of s-risks is that Iâm (more) clueless about basically every intervention now, and am looking to portfolios and hedging.
If you are publishing, and your ethical or empirical views are sufficiently different from others working on the problem so that you make very different tradeoffs, then that could be good, bad or ambiguous. For example, if you didnât really care about s-risks, then publishing a useful considerations for those who are concerned about s-risks might take attention away from your own priorities, or it might increase cooperation, and the default position to me should be deep uncertainty/âcluelessness here, not that itâs good in expectation or bad in expectation or 0 in expectation.
Maybe you can eliminate this ambiguity or at least constrain its range to something relatively insignificant by building a model, doing a sensitivity analysis, etc., but a lot of things donât work out, and the ambiguity could be so bad that it infects everything else. This is roughly where I am now: I have considerations that result in complex cluelessness about AI-related interventions and I want to know how people work through this.
Re: cost-effectiveness analyses always turning up positive, perhaps especially in longtermism. FWIW that hasnât been my experience. Instead, my experience is that every time I investigate the case for some AI-related intervention being worth funding under longtermism, I conclude that itâs nearly as likely to be net-negative as net-positive given our great uncertainty and therefore I end up stuck doing almost entirely âmetaâ things like creating knowledge and talent pipelines.
Of course, that doesnât mean he never finds good âdirect workâ, or that the âdirect workâ already being funded isnât better than nothing in expectation overall, and I would guess he thinks it is.
Hmm, it seems to me (and you can correct me) that we should be able to agree that there are SOME technical AGI safety research publications that are positive under some plausible beliefs/âvalues and harmless under all plausible beliefs/âvalues, and then we donât have to talk about cluelessness and tradeoffs, we can just publish them.
And we both agree that there are OTHER technical AGI safety research publications that are positive under some plausible beliefs/âvalues and negative under others. And then we should talk about your portfolios etc. Or more simply, on a case-by-case basis, we can go looking for narrowly-tailored approaches to modifying the publication in order to remove the downside risks while maintaining the upside.
I feel like weâre arguing past each other: I keep saying the first category exists, and you keep saying the second category exists. We should just agree that both categories exist! :-)
Perhaps the more substantive disagreement is what fraction of the work is in which category. I see most but not all ongoing technical work as being in the first category, and I think you see almost all ongoing technical work as being in the second category. (I think you agreed that âpublishing an analysis about what happens if a cosmic ray flips a bitâ goes in the first category.)
(Luke says âAI-relatedâ but my impression is that he mostly works on AGI governance not technical, and the link is definitely about governance not technical. I would not be at all surprised if proposed governance-related projects were much more heavily weighted towards the second category, and am only saying that technical safety research is mostly first-category.)
For example, if you didnât really care about s-risks, then publishing a useful considerations for those who are concerned about s-risks might take attention away from your own priorities, or it might increase cooperation, and the default position to me should be deep uncertainty/âcluelessness here, not that itâs good in expectation or bad in expectation or 0 in expectation.
This points to another (possible?) disagreement. I think maybe you have the attitude where (to caricature somewhat) if thereâs any downside risk whatsoever, no matter how minor or far-fetched, you immediately jump to âIâm clueless!â. Whereas Iâm much more willing to say: OK, I mean, if you do anything at all thereâs a âdownside riskâ in a sense, just because life is uncertain, who knows what will happen, but thatâs not a good reason to let just sit on the sidelines and let nature take its course and hope for the best. If I have a project whose first-order effect is a clear and specific and strong upside opportunity, I donât want to throw that project out unless thereâs a comparably clear and specific and strong downside risk. (And of course we are obligated to try hard to brainstorm what such a risk might be.) Like if a firefighter is trying to put out a fire, and they aim their hose at the burning interior wall, they donât stop and think, âWell I donât know what will happen if the wall gets wet, anything could happen, so Iâll just not pour water on the fire, yâknow, donât want to mess things up.â
The âcluelessnessâ intuition gets its force from having a strong and compelling upside story weighed against a strong and compelling downside story, I think.
If the first-order effect of a project is âdirectly mitigating an important known s-riskâ, and the second-order effects of the same project are âI dunno, itâs a complicated world, anything could happenâ, then I say we should absolutely do that project.
Perhaps the more substantive disagreement is what fraction of the work is in which category. I see most but not all ongoing technical work as being in the first category, and I think you see almost all ongoing technical work as being in the second category. (I think you agreed that âpublishing an analysis about what happens if a cosmic ray flips a bitâ goes in the first category.)
Ya, I think this is the crux. Also, considerations like the cosmic ray flips a bit tend to force a lot of things into the second category when they otherwise wouldnât have been, although Iâm not specifically worried about cosmic ray bit flips, since they seems sufficiently unlikely and easy to avoid.
(Luke says âAI-relatedâ but my impression is that he mostly works on AGI governance not technical, and the link is definitely about governance not technical. I would not be at all surprised if proposed governance-related projects were much more heavily weighted towards the second category, and am only saying that technical safety research is mostly first-category.)
(Fair.)
The âcluelessnessâ intuition gets its force from having a strong and compelling upside story weighed against a strong and compelling downside story, I think.
This is actually what Iâm thinking is happening, though (not like the firefighter example), but we arenât really talking much about the specifics. There might indeed be specific cases where I agree that we shouldnât be clueless if we worked through them, but I think there are important potential tradeoffs between incidental and agential s-risks, between s-risks and other existential risks, even between the same kinds of s-risks, etc., and there is a ton of uncertainty in the expected harm from these risks, so much that itâs inappropriate to use a single distribution (without sensitivity analysis to âreasonableâ distributions, and with this sensitivity analysis, things look ambiguous), similar to this example, and weâre talking about âsweeteningâ one side or the other i, but thatâs totally swamped by our uncertainty.
If the first-order effect of a project is âdirectly mitigating an important known s-riskâ, and the second-order effects of the same project are âI dunno, itâs a complicated world, anything could happenâ, then I say we should absolutely do that project.
What I have in mind is more symmetric in upsides and downsides (or at least, Iâm interested in hearing why people think it isnât in practice), and I donât really distinguish between effects by order*. My post points out potential reasons that I actually think could dominate. The standard Iâm aiming for is âCould a reasonable person disagree?â, and I default to believing a reasonable person could disagree when I point out such tradeoffs until we actually carefully work through them in detail and it turns out itâs pretty unreasonable to disagree.
*Although thinking more about it now, I suppose longer chains are more fragile and likely to have unaccounted for effects going in the opposite direction, so maybe we ought to give them less weight, and maybe this solves the issue if we did this formally? I think ignoring higher-order effects is formally irrational using vNM rationality or stochastic dominance, although maybe fine in practice, if what weâre actually doing is just an approximation of giving them far less weight with a skeptical prior and then they actually just get dominated completely by more direct effects.
I donât really distinguish between effects by order*
I agree that direct and indirect effects of an action are fundamentally equally important (in this kind of outcome-focused context) and I hadnât intended to imply otherwise.
If you arenât publishing anything, then sure, research into what to do seems mostly harmless (other than opportunity costs) in expectation, but it doesnât actually follow that it would necessarily be good in expectation, if you have enough deep uncertainty (or complex cluelessness); I think this example illustrates this well, and is basically the kind of thing Iâm worried about all of the time now. In the particular case of sign flip errors, I do think it was useful for me to know about this consideration and similar ones, and I act differently than I would have otherwise as a result, but one of the main effects since learning about these kinds of s-risks is that Iâm (more) clueless about basically every intervention now, and am looking to portfolios and hedging.
If you are publishing, and your ethical or empirical views are sufficiently different from others working on the problem so that you make very different tradeoffs, then that could be good, bad or ambiguous. For example, if you didnât really care about s-risks, then publishing a useful considerations for those who are concerned about s-risks might take attention away from your own priorities, or it might increase cooperation, and the default position to me should be deep uncertainty/âcluelessness here, not that itâs good in expectation or bad in expectation or 0 in expectation.
Maybe you can eliminate this ambiguity or at least constrain its range to something relatively insignificant by building a model, doing a sensitivity analysis, etc., but a lot of things donât work out, and the ambiguity could be so bad that it infects everything else. This is roughly where I am now: I have considerations that result in complex cluelessness about AI-related interventions and I want to know how people work through this.
For another source of pessimism, Luke Muehlhauser from Open Phil wrote:
Of course, that doesnât mean he never finds good âdirect workâ, or that the âdirect workâ already being funded isnât better than nothing in expectation overall, and I would guess he thinks it is.
Hmm, it seems to me (and you can correct me) that we should be able to agree that there are SOME technical AGI safety research publications that are positive under some plausible beliefs/âvalues and harmless under all plausible beliefs/âvalues, and then we donât have to talk about cluelessness and tradeoffs, we can just publish them.
And we both agree that there are OTHER technical AGI safety research publications that are positive under some plausible beliefs/âvalues and negative under others. And then we should talk about your portfolios etc. Or more simply, on a case-by-case basis, we can go looking for narrowly-tailored approaches to modifying the publication in order to remove the downside risks while maintaining the upside.
I feel like weâre arguing past each other: I keep saying the first category exists, and you keep saying the second category exists. We should just agree that both categories exist! :-)
Perhaps the more substantive disagreement is what fraction of the work is in which category. I see most but not all ongoing technical work as being in the first category, and I think you see almost all ongoing technical work as being in the second category. (I think you agreed that âpublishing an analysis about what happens if a cosmic ray flips a bitâ goes in the first category.)
(Luke says âAI-relatedâ but my impression is that he mostly works on AGI governance not technical, and the link is definitely about governance not technical. I would not be at all surprised if proposed governance-related projects were much more heavily weighted towards the second category, and am only saying that technical safety research is mostly first-category.)
This points to another (possible?) disagreement. I think maybe you have the attitude where (to caricature somewhat) if thereâs any downside risk whatsoever, no matter how minor or far-fetched, you immediately jump to âIâm clueless!â. Whereas Iâm much more willing to say: OK, I mean, if you do anything at all thereâs a âdownside riskâ in a sense, just because life is uncertain, who knows what will happen, but thatâs not a good reason to let just sit on the sidelines and let nature take its course and hope for the best. If I have a project whose first-order effect is a clear and specific and strong upside opportunity, I donât want to throw that project out unless thereâs a comparably clear and specific and strong downside risk. (And of course we are obligated to try hard to brainstorm what such a risk might be.) Like if a firefighter is trying to put out a fire, and they aim their hose at the burning interior wall, they donât stop and think, âWell I donât know what will happen if the wall gets wet, anything could happen, so Iâll just not pour water on the fire, yâknow, donât want to mess things up.â
The âcluelessnessâ intuition gets its force from having a strong and compelling upside story weighed against a strong and compelling downside story, I think.
If the first-order effect of a project is âdirectly mitigating an important known s-riskâ, and the second-order effects of the same project are âI dunno, itâs a complicated world, anything could happenâ, then I say we should absolutely do that project.
Ya, I think this is the crux. Also, considerations like the cosmic ray flips a bit tend to force a lot of things into the second category when they otherwise wouldnât have been, although Iâm not specifically worried about cosmic ray bit flips, since they seems sufficiently unlikely and easy to avoid.
(Fair.)
This is actually what Iâm thinking is happening, though (not like the firefighter example), but we arenât really talking much about the specifics. There might indeed be specific cases where I agree that we shouldnât be clueless if we worked through them, but I think there are important potential tradeoffs between incidental and agential s-risks, between s-risks and other existential risks, even between the same kinds of s-risks, etc., and there is a ton of uncertainty in the expected harm from these risks, so much that itâs inappropriate to use a single distribution (without sensitivity analysis to âreasonableâ distributions, and with this sensitivity analysis, things look ambiguous), similar to this example, and weâre talking about âsweeteningâ one side or the other i, but thatâs totally swamped by our uncertainty.
What I have in mind is more symmetric in upsides and downsides (or at least, Iâm interested in hearing why people think it isnât in practice), and I donât really distinguish between effects by order*. My post points out potential reasons that I actually think could dominate. The standard Iâm aiming for is âCould a reasonable person disagree?â, and I default to believing a reasonable person could disagree when I point out such tradeoffs until we actually carefully work through them in detail and it turns out itâs pretty unreasonable to disagree.
*Although thinking more about it now, I suppose longer chains are more fragile and likely to have unaccounted for effects going in the opposite direction, so maybe we ought to give them less weight, and maybe this solves the issue if we did this formally? I think ignoring higher-order effects is formally irrational using vNM rationality or stochastic dominance, although maybe fine in practice, if what weâre actually doing is just an approximation of giving them far less weight with a skeptical prior and then they actually just get dominated completely by more direct effects.
I agree that direct and indirect effects of an action are fundamentally equally important (in this kind of outcome-focused context) and I hadnât intended to imply otherwise.