Given the optimizer’s cause, how do I optimally pick a cause area? Two observations:
In the standard model above, I cannot improve upon picking the cause with the highest estimate. While this cause will likely be overestimated by several OOMs and will likely not be the top x-risk cause, it is still the optimal cause to choose in expectation.
But in the extended model with grounded and speculative causes, the logic changes. Here, it can be optimal to pick the cause with the highest x-risk estimate in the grounded class, even if there are causes with higher x-risk estimates in the speculative class.
I think this has very interesting implications.
2) implies that working in a more grounded cause area (like global health?) can be better than working on speculative x-risk. This is a powerful implication and I think EA should take this very seriously.
1) implies that even if everyone makes an individually optimal cause-prioritization decision, some people’s top causes will still look highly implausible to others.
Good question, but I implore you not to take this post too seriously. It’s a real phenomenon, but it’s a real stretch to claim that this applies in the implied way to cause areas like AI safety.
The model in the article is just a toy model of a world where existential threats are randomly distributed according to a power law, where genuinely high-probability threats are, by assumption, basically absent from the space of possible threats, and where there’s no process of updating based on evidence.
A more narrow claim like: “single narrowly defined x-risk estimates of genuinely speculative, unresearched causes are likely to be inflated” might be valid, but that doesn’t seem to be what titotal is implying. He’s making a claim way beyond anything implied by the model, even if the model were a valid representation of the phenomenon. He seems to believe that almost all of today’s concern for AI risk is all downstream of a belief cultivated within a narrow subcommunity subject to the optimiser’s curse—an extraordinary claim that requires a lot more evidence than that supplied in the article.
The claim being made is something like:
Some time in the 2000s, Eliezer Yudkowsky and friends made up some numbers for AI risk
Community dynamics (rather than the merits of the arguments) spread these numbers from Bostrom to Tegmark to Sam Harris to Elon Musk to the EA community etc.
These social dynamics had such an effect that multiple seemingly independent and unrelated people/experts from Geoffrey Hinton to Yoshua Bengio to Chinese academics and tech people from completely different intellectual lineages have absorbed this false belief from the cultural milieu (again, not at all based on the merits of the arguments), leading many of these people, as well as specialists across unrelated fields, to make estimates of AI x-risk that remain orders of magnitude too high
Note that this requires some pretty wild, difficult-to-justify assumptions on how this belief has spread.
The opposing narrative (which I would advocate for) is that:
AI x-risk is something that has independently been identified as a risk by multiple people—often far before quantifying or ranking risks.
The spread of these beliefs was obviously affected by community dynamics, but people largely adopted somewhat independent beliefs based on the merits of rational arguments
Quantitative predictions were inherently imprecise because they’re so dependent on messy world models, but they were grounded enough in reason and evidence, and composed of enough independent estimates, that any optimiser’s curse is massively weakened
As certain predictions came to pass, and current LLMs approach AGI in non-ideal geopolitical and competitive circumstances, this is increasingly being seen by a wider range of thinkers as a >1% x-risk
Even if there were “optimiser curse” risks in initial prioritisation of AI, it’s now increasingly recognised that AI will be a massive deal. AI-assisted engineered pandemic uplift work, observed cyber-capabilities, and signs of misalignment/scheming etc. are building on the strong theoretical evidence base that AI-generated catastrophe is possible.
And on your particular question of how to act, even given the optimiser’s curse as stated in the toy model, working on the speculative thing could still be optimal. If a highly uncertain intervention seems exceptionally promising or x-risky, the value of information becomes incredibly high, because accurate or well-reasoned research will lead to this intervention being prioritised or not by far more people. If you have a research focus, it’s therefore probably more recommended to focus on a more uncertain, “high-risk-high-reward” area. You could also draw up a toy model for explore/exploit based on the optimiser’s curse.
Finally, you don’t necessarily have to “pick” from a narrow set of pre-defined cause areas. You can also divide or merge risks, cause areas, skill-sets etc. to be more robust, precise, or coherent with your own world model (e.g. focusing on engineered pandemics because you realise this could interact with AI-related x-risk, GCBRs, and global health).
Cause prioritisation is a function of the marginal impact on the outcome per marginal dollar/hour spent or similar.
This adds another layer of complexity because you can’t just eg “integrate out” the existential risk first and then reason about the impact, you kind of need to do it jointly.
This means you also need to include sources of impact uncertainty. To me, this ~ removes many longtermist cause areas because it drags most of the “impact mass” too close to zero—but people have different opinions on this of course.
Point 1 is interesting—you can do better if you have and incorporate a prior, but the question is where that comes from. I think often it’s easier to have priors over intervention success than existential outcomes per se.
though a corollary of it is “don’t assume that just because you’ve picked direct work that your career choice is maximally good and stuff like donations and helping others is just a distraction”. This is arguably true for speculative career choices even if the optimal cause is the correct one (i.e. even if AI x-risk really does dominate everything, lots of the promising approaches to resolving it that people might choose will have no impact)
I like diversification as a reaction to this type of uncertainty, but it does not trivially follow? I might be missing something—do you have a favourite minimal set of assumptions that rigorously yield diversification as a function of this?
the x-risk coming from each cause decreases with each additional person working on this cause
Whether diversification is better (in expectation) depends on how a cause’s x-risk decreases as additional people work on this cause. If x-risk decreases linearly (the 1000th person makes the same marginal contribution as the 1st), then diversification is not better in expectation. But if the contribution to x-risk prevention is marginally decreasing in people, diversification is better.
(By diversification I mean each person choosing their top estimated x-risk cause individually. But it can also mean that some people deliberately do not work on the cause with the highest aggregated risk estimate.)
I was more referring to the diversification as implied by “don’t focus the vast majority of efforts on one cause”, which to me meant more “if you’re a decision maker over some amount of resources, you should diversify the allocation across cause areas”. Which I agree with, but it’s quite hard to really justify.
Yes, via nonlinearity you can get to diversification, but this means making additional assumptions beyond sampling error/publication bias/optimiser curse type effects. The nonlinearity you’re describing matters on a movement level, but not on a individual decision makers level. ”Impact risk aversion” is another mechanism to get diversification which I think can be reasonable in cases where eg low impact reduces the probability of future donations or similar.
One channel I think is under explored and might work well as a justification for diversification in practice is something like this (I haven’t thought about this rigorously though): if I predictably optimise and my objective function is known to others, they will (in the worst case, possibly thru misaligned incentives) feed me biased information to influence my decision, and optimisation is very sensitive to noise, therefore I subject myself to adverse selection. So basically by not optimising but diversifying across good options you reduce the negative impact of this type of adverse selection. In this case, the “errors” are not iid. Hard to say how much diversification that yields.
I think the conclusion that diversification is a good strategy follows trivially from the optimizers’ curse: if you focus all your efforts on the apparent biggest threat, you’ve probably just focused on the cause with the largest risk assessment error and entirely neglected the actual biggest threat. A more diverse allocation is more likely to address the actual biggest threat. If there are diminishing returns to resources allocated to mitigate particular risk areas that makes diversification look better (complex nonlinear returns complicate it). As does the possibility that larger errors in risk assessment for a particular type of risk are inversely correlated with ability to invest in the best mitigation strategy for that type of risk.[1]
But your point about adverse selection is a good one too. Metrics are gameable, and there are stronger incentives to do so when funding is “winner takes all” rather than “we disburse funds to a wide selection of causes and value rigour and disclosure of uncertainties”
I think there are probably exceptions to this, but I think it’s generally true. Good understanding of celestial mechanics and early warning systems, for example, are absolutely essential to potentially preventing hypothetical large space rocks colliding with earth, but also mean that we are less likely to overestimate the imminence of destruction by a rogue asteroid than we are for more unpredictable phenomenon.
While this sounds intuitively right, I think in the simplest utility maximising setting (iid additive errors with mean zero) your first claim does not seem true? The best looking noisy option is still most likely to be the best?
(I need to think more about the maths, but at least you need some kind of shrinkage to a prior that can change the ranking, which you’re unlikely to get, and if you’re maximising utility the solution is always fully concentrated?)
I’m not sure naive total utility maximization [in a static framework] is the best framework to be thinking about dealing with existential risk over time.[1]
Assuming the number of risks and error bars are not trivially small, the universal outcome of concentrating all your risk mitigations on one is that most risks continue to be a high as they could possibly be. The modal outcome is that the risks ignored includes at least one risk greater than the one all efforts are concentrated on mitigating. Some reasonable assumptions in the article above show this can hold even where the actual biggest risk is orders of magnitude greater than the one targeted. In the diversified approach, less money are devoted to reducing the perceived biggest risk, but the rest is apportioned to reducing other risks. This seems more robust to conventional assumptions like uncertainty and some risks being easier to mitigate than others.
And tbh I’m not even seeing an average utility boost from concentrating on the single largest risk as opposed to mitigating lots of risks without ancillary assumptions like increasing returns to risk reduction expenditure or the actual value of many risks under consideration being 0.
Yeah I agree—expected utility maximisation really starts to fall apart in this existential risk regime, even over trajectories rather than applied statically, and it only makes sense “locally” and at the margin.
Personally I’m very happy to bite the bullet and not be rigorously utilitarian, but I’m also a global health focussed “old school EA” thinking about how much to diversify donations across charities ;)
Interesting. Who might these people be who deliberately feed you biased information? How do they benefit from you focusing on cause area y instead of cause area z?
I think 1) implies that you should give up some substantial optimisation for the sake of greater versatility (which seems approx titotal’s view with reference to overcommitting)
2) feels correct and important to me, also since I’ve been arguing in the post op linked and elsewhere that treating extinction as special is a heuristic that was useful for initial cause prioritisation but isn’t a valid reason for focusing on it two decades later.
Given the optimizer’s cause, how do I optimally pick a cause area? Two observations:
In the standard model above, I cannot improve upon picking the cause with the highest estimate. While this cause will likely be overestimated by several OOMs and will likely not be the top x-risk cause, it is still the optimal cause to choose in expectation.
But in the extended model with grounded and speculative causes, the logic changes. Here, it can be optimal to pick the cause with the highest x-risk estimate in the grounded class, even if there are causes with higher x-risk estimates in the speculative class.
I think this has very interesting implications.
2) implies that working in a more grounded cause area (like global health?) can be better than working on speculative x-risk. This is a powerful implication and I think EA should take this very seriously.
1) implies that even if everyone makes an individually optimal cause-prioritization decision, some people’s top causes will still look highly implausible to others.
Good question, but I implore you not to take this post too seriously. It’s a real phenomenon, but it’s a real stretch to claim that this applies in the implied way to cause areas like AI safety.
The model in the article is just a toy model of a world where existential threats are randomly distributed according to a power law, where genuinely high-probability threats are, by assumption, basically absent from the space of possible threats, and where there’s no process of updating based on evidence.
A more narrow claim like: “single narrowly defined x-risk estimates of genuinely speculative, unresearched causes are likely to be inflated” might be valid, but that doesn’t seem to be what titotal is implying. He’s making a claim way beyond anything implied by the model, even if the model were a valid representation of the phenomenon. He seems to believe that almost all of today’s concern for AI risk is all downstream of a belief cultivated within a narrow subcommunity subject to the optimiser’s curse—an extraordinary claim that requires a lot more evidence than that supplied in the article.
The claim being made is something like:
Some time in the 2000s, Eliezer Yudkowsky and friends made up some numbers for AI risk
Community dynamics (rather than the merits of the arguments) spread these numbers from Bostrom to Tegmark to Sam Harris to Elon Musk to the EA community etc.
These social dynamics had such an effect that multiple seemingly independent and unrelated people/experts from Geoffrey Hinton to Yoshua Bengio to Chinese academics and tech people from completely different intellectual lineages have absorbed this false belief from the cultural milieu (again, not at all based on the merits of the arguments), leading many of these people, as well as specialists across unrelated fields, to make estimates of AI x-risk that remain orders of magnitude too high
Note that this requires some pretty wild, difficult-to-justify assumptions on how this belief has spread.
The opposing narrative (which I would advocate for) is that:
AI x-risk is something that has independently been identified as a risk by multiple people—often far before quantifying or ranking risks.
The spread of these beliefs was obviously affected by community dynamics, but people largely adopted somewhat independent beliefs based on the merits of rational arguments
Quantitative predictions were inherently imprecise because they’re so dependent on messy world models, but they were grounded enough in reason and evidence, and composed of enough independent estimates, that any optimiser’s curse is massively weakened
As certain predictions came to pass, and current LLMs approach AGI in non-ideal geopolitical and competitive circumstances, this is increasingly being seen by a wider range of thinkers as a >1% x-risk
Even if there were “optimiser curse” risks in initial prioritisation of AI, it’s now increasingly recognised that AI will be a massive deal. AI-assisted engineered pandemic uplift work, observed cyber-capabilities, and signs of misalignment/scheming etc. are building on the strong theoretical evidence base that AI-generated catastrophe is possible.
And on your particular question of how to act, even given the optimiser’s curse as stated in the toy model, working on the speculative thing could still be optimal. If a highly uncertain intervention seems exceptionally promising or x-risky, the value of information becomes incredibly high, because accurate or well-reasoned research will lead to this intervention being prioritised or not by far more people. If you have a research focus, it’s therefore probably more recommended to focus on a more uncertain, “high-risk-high-reward” area. You could also draw up a toy model for explore/exploit based on the optimiser’s curse.
Finally, you don’t necessarily have to “pick” from a narrow set of pre-defined cause areas. You can also divide or merge risks, cause areas, skill-sets etc. to be more robust, precise, or coherent with your own world model (e.g. focusing on engineered pandemics because you realise this could interact with AI-related x-risk, GCBRs, and global health).
Cause prioritisation is a function of the marginal impact on the outcome per marginal dollar/hour spent or similar.
This adds another layer of complexity because you can’t just eg “integrate out” the existential risk first and then reason about the impact, you kind of need to do it jointly.
This means you also need to include sources of impact uncertainty.
To me, this ~ removes many longtermist cause areas because it drags most of the “impact mass” too close to zero—but people have different opinions on this of course.
Point 1 is interesting—you can do better if you have and incorporate a prior, but the question is where that comes from. I think often it’s easier to have priors over intervention success than existential outcomes per se.
Above all it implies don’t focus the vast majority of efforts on one cause.
That might not be practical for career choices,[1] but it’s certainly possible for a funder or movement
though a corollary of it is “don’t assume that just because you’ve picked direct work that your career choice is maximally good and stuff like donations and helping others is just a distraction”. This is arguably true for speculative career choices even if the optimal cause is the correct one (i.e. even if AI x-risk really does dominate everything, lots of the promising approaches to resolving it that people might choose will have no impact)
I like diversification as a reaction to this type of uncertainty, but it does not trivially follow? I might be missing something—do you have a favourite minimal set of assumptions that rigorously yield diversification as a function of this?
One simple model is:
each person can choose only one cause area
errors are iid across people
the x-risk coming from each cause decreases with each additional person working on this cause
Whether diversification is better (in expectation) depends on how a cause’s x-risk decreases as additional people work on this cause. If x-risk decreases linearly (the 1000th person makes the same marginal contribution as the 1st), then diversification is not better in expectation. But if the contribution to x-risk prevention is marginally decreasing in people, diversification is better.
(By diversification I mean each person choosing their top estimated x-risk cause individually. But it can also mean that some people deliberately do not work on the cause with the highest aggregated risk estimate.)
I was more referring to the diversification as implied by “don’t focus the vast majority of efforts on one cause”, which to me meant more “if you’re a decision maker over some amount of resources, you should diversify the allocation across cause areas”. Which I agree with, but it’s quite hard to really justify.
Yes, via nonlinearity you can get to diversification, but this means making additional assumptions beyond sampling error/publication bias/optimiser curse type effects.
The nonlinearity you’re describing matters on a movement level, but not on a individual decision makers level.
”Impact risk aversion” is another mechanism to get diversification which I think can be reasonable in cases where eg low impact reduces the probability of future donations or similar.
One channel I think is under explored and might work well as a justification for diversification in practice is something like this (I haven’t thought about this rigorously though): if I predictably optimise and my objective function is known to others, they will (in the worst case, possibly thru misaligned incentives) feed me biased information to influence my decision, and optimisation is very sensitive to noise, therefore I subject myself to adverse selection. So basically by not optimising but diversifying across good options you reduce the negative impact of this type of adverse selection. In this case, the “errors” are not iid. Hard to say how much diversification that yields.
I think the conclusion that diversification is a good strategy follows trivially from the optimizers’ curse: if you focus all your efforts on the apparent biggest threat, you’ve probably just focused on the cause with the largest risk assessment error and entirely neglected the actual biggest threat. A more diverse allocation is more likely to address the actual biggest threat. If there are diminishing returns to resources allocated to mitigate particular risk areas that makes diversification look better (complex nonlinear returns complicate it). As does the possibility that larger errors in risk assessment for a particular type of risk are inversely correlated with ability to invest in the best mitigation strategy for that type of risk.[1]
But your point about adverse selection is a good one too. Metrics are gameable, and there are stronger incentives to do so when funding is “winner takes all” rather than “we disburse funds to a wide selection of causes and value rigour and disclosure of uncertainties”
I think there are probably exceptions to this, but I think it’s generally true. Good understanding of celestial mechanics and early warning systems, for example, are absolutely essential to potentially preventing hypothetical large space rocks colliding with earth, but also mean that we are less likely to overestimate the imminence of destruction by a rogue asteroid than we are for more unpredictable phenomenon.
While this sounds intuitively right, I think in the simplest utility maximising setting (iid additive errors with mean zero) your first claim does not seem true? The best looking noisy option is still most likely to be the best?
(I need to think more about the maths, but at least you need some kind of shrinkage to a prior that can change the ranking, which you’re unlikely to get, and if you’re maximising utility the solution is always fully concentrated?)
I’m not sure naive total utility maximization [in a static framework] is the best framework to be thinking about dealing with existential risk over time.[1]
Assuming the number of risks and error bars are not trivially small, the universal outcome of concentrating all your risk mitigations on one is that most risks continue to be a high as they could possibly be. The modal outcome is that the risks ignored includes at least one risk greater than the one all efforts are concentrated on mitigating. Some reasonable assumptions in the article above show this can hold even where the actual biggest risk is orders of magnitude greater than the one targeted. In the diversified approach, less money are devoted to reducing the perceived biggest risk, but the rest is apportioned to reducing other risks. This seems more robust to conventional assumptions like uncertainty and some risks being easier to mitigate than others.
And tbh I’m not even seeing an average utility boost from concentrating on the single largest risk as opposed to mitigating lots of risks without ancillary assumptions like increasing returns to risk reduction expenditure or the actual value of many risks under consideration being 0.
Yeah I agree—expected utility maximisation really starts to fall apart in this existential risk regime, even over trajectories rather than applied statically, and it only makes sense “locally” and at the margin.
Personally I’m very happy to bite the bullet and not be rigorously utilitarian, but I’m also a global health focussed “old school EA” thinking about how much to diversify donations across charities ;)
Interesting. Who might these people be who deliberately feed you biased information? How do they benefit from you focusing on cause area y instead of cause area z?
To be clear, it need not be deliberate and they need not benefit personally!
I think 1) implies that you should give up some substantial optimisation for the sake of greater versatility (which seems approx titotal’s view with reference to overcommitting)
2) feels correct and important to me, also since I’ve been arguing in the post op linked and elsewhere that treating extinction as special is a heuristic that was useful for initial cause prioritisation but isn’t a valid reason for focusing on it two decades later.