Non-EA interests include chess and TikTok (@benthamite). Formerly @ CEA, METR + a couple now-acquired startups.
Ben_Westđ¸
DiGiovanni states:
We arenât justified in c-preferring any action over any other
A single counterexample of a pair of actions where one is c-preferred over the other would suffice to disprove his claim.
It seems noteworthy that none of the solutions attempted this.
Interesting post, thanks Richard.
Am I correct in understanding you to be suggesting âstarting a nuclear warâ as an example of something so obviously c-dispreferable that only a âradical skepticâ would disagree?
If not, do you have an example of a pair of actions such that only a radical skeptic could doubt the c-preferability of one of them?
The stochastic dominance point is helpful, ty
Thanks! Not sure I agree with âheuristicâ: Stockfishâs stopping rule is a heuristic, but itâs backed by a theorem (Cantelli) that holds with no distributional assumptions. What is distribution-dependent is the estimate of
that feeds into it, and I agree that some skepticism is warranted there.The chess analogy is meant to show the framework is implementable, not merely theoretical. That moves the question from âare expected values well-defined?â to âwhat is
and how large is relative to it?â I consider that progress because the answer isnât uniformly âsuspend judgmentâ: some parameter values license acting, others clearly donât.
Thanks! The Cantelli bound is a basic result of probability theory, so I think itâs hard to change chess such that it would no longer apply at all.
But an experiment I would be interested to do is to run an engine with varying parameters and see how much this changes performance. If the âdeliberate until you donât see many sign flipsâ approach only works because we have a precisely-tuned definition of âmanyâ, then I do think this weakens the analogy.
UnawareÂness as a comÂpuÂtaÂtional constraint
There are even AI-safety-pun-based sports teams!
Thanks! Tbc, I agree that there are persistence effects; I just donât think they are so strong that we can reasonably expect that slight changes in starting conditions result in a completely different world.
if we were to travel back in time and alter starting conditions just slightly, it seems reasonable to expect that the world today would be completely different.
This is possible, but seems pretty unclear to me. cf The Gods of Straight Lines:
a few short years after the bombs stopped falling in 1945, the world economy returned to trend as if nothing had happened.
In 1776, America rebelled in the name of freedom and democracy: the origin myth of the modern world order. And yet, somehow, unrebellious Canada ended up just as free and democratic. An unrebellious America likely would have too.
For two decades, North Vietnam battled under the banner of communism, and won against all odds. And yet, somehow, Vietnam is now the most pro-capitalist country in the world.
$35/âattendee is an extremely impressive cost. Thanks for doing this and writing it up!
Thanks for writing this Sarah and best wishes for the new position!
I have been pretty pleased with the 2026 Forum output (particularly the unawareness event caused me to think more about my own work than most other things, maybe more than any other online event in 2026). Kudos to the rest of the team, and to your leadership for enabling that.
Median european seed stage pre-money valuation is âŹ5.0 million. People often raise the seed round within 1-2 years of starting. So âŹ1M is donating 20-40% of your company.[1]
- ^
Of course a large fraction of companies never make it even so far as a seed round. Itâs very hard to figure out what reference class someone should be in but it is worth noting that there are a good number of people who read this forum who have achieved valuations greater than that average within a couple years of trying.
- ^
Cool idea! I like the move of taking a selfish ambitious drive and trying to transform it into something more altruistic.
To see this, letâs be generous and say the expected number of future well-off people added by preventing an existential catastrophe is only 10^40
I think your argumentation supports the 10^40 number more than the âwell-offâ claim. Iâm not sure for the best canonical source expressing skepticism of the âwell-offâ bit, but The Future Might Not Be So Great contains a long list of arguments about whether future people will be well off; I would be interested in you responding to some of those.
Anthony cites Greaves and MacAskill giving an example similar to your gunpowder one:
Consider, for example, would-be longtermists in the Middle Ages. It is plausible that the considerations most relevant to their decision â such as the benefits of science, and therefore the enormous value of efforts to help make the scientific and industrial revolutions happen sooner â would not have been on their radar. Rather, they might instead have backed attempts to spread Christianity, perhaps by violence: a putative route to value that, by our more enlightened lights today, looks wildly off the mark. The suggestion, then, is that our current predicament is relevantly similar to that of our medieval would-be longtermists.
I personally think these examples are less compelling than they first appear (e.g. the persistence literature generally finds weaker effects than what you might imagine), but I agree that a failure of EAs to find examples of sign flips doesnât mean that future ones wonât exist.
Thanks! I canât tell if this is cruxy, but for what itâs worth your âpessimal inductionâ vignettes donât resonate with me in a way which makes me less motivated by the unawareness concerns.
For example, Bostrom coined the phrase âattention hazardâ in 2011. I remember someone telling me that MIRI was net-negative for this reason at EAG 2015, and I would be surprised if e.g. Habryka hadnât considered this risk before starting Lightcone. So I disagree with citing him/âthis as a good example of unawareness; itâs more that they mis-estimated a known risk factor.
Similarly, I remember talking about SARP at one of my first EAGs. I think I came across it in Brian Tomasikâs 2007 post, maybe even before I had encountered EA. Perhaps Iâve mis-estimated those concerns, but it doesnât seem like unawareness.
My overall experience is kind of the opposite of yours: when I got involved in EA people talked a lot about âCause Xâ and âCrucial Considerationsâ and now theyâve mostly just⌠stopped? Like people tried to find other considerations, and thereâs some new stuff around s-risks and weird decision theories etc., but if you look at what people talk about at EAGs today it feels mostly like more precise versions of what was discussed in 2016, rather than a large and unpredictable jump from the older understanding. Or, more technically: it feels like weâve had updates in evidence-space, but not as many updates in hypothesis-space, and I understand the latter to be motivating imprecision.
Obviously, this could be because EAs suck at cause prio research, or we just havenât been hit yet with the big update, etc., but the âpessimal inductionâ seems less pessimal to me.
Itâs great that youâre thinking about this!
Iâm confused why you are denominating options in robotics-startup-days saved. This feels like a narrow definition of âimpactâ. Iâd encourage you to consider other ways to benefit the world; parts 4-6 of the 80k career guide might be helpful. Specifically, under the assumption that the thing you terminally value is more like âreducing sufferingâ than ârobotics progressâ, I would encourage you to first consider which causes advance those values, and only then drill into job options. (The 80k career guide will walk you through this.)
Also, even to the extent you do just terminally value robotics progress, you might want to consider whether robotics will advance too quickly for your estimation to be accurate.
You respond to Richard Ngo here:
> do you think that, if we had a theory of sociopolitics that was about as good as 20th-century economics, then we wouldnât be clueless about how to do sociopolitical interventions (like founding AI safety movements) effectively?
No, because I think âfounding AI safety movements that succeed at making the far future go betterâ is a pretty out-of-distribution kind of sociopolitical intervention.
Suppose instead we had a comparably good theory of the right reference class, e.g. âmovements trying to shape transformative technologies.â Would we still be clueless about AI safety movement-building?
More generally: you list various considerations across your posts and I have a hard time understanding which is load-bearing for your answer here. Some possibilities:
Weâre clueless because we havenât yet developed the relevant theory (Richardâs reading IIUC, on which cluelessness is contingent and reducible)
No such theory could be validated even in principle, because we never observe the target variable (far-future value) and calibration on near-term proxies doesnât transfer
Even a validated theory wouldnât help, because impact is dominated by considerations inaccessible to any theory (e.g. unconceived hypothesis classes)
Whatâs the clearest example of a complex cluelessness sign flip youâre aware of?
(By âclearâ I mean âhad a very narrow confidence interval before encountering some consideration and a narrow interval after encountering that consideration but the CIs now center points with opposite signsâ.[1])
The clearest examples I know of (e.g. rescuing Hitler as a child) seem to me like examples of simple cluelessness. You list some examples here, but they donât seem that clear to me, e.g. I disagree that âEarly awareness-raising about AGI x-risk presumably seemed robustly goodâ and would guess most people involved in that had CIs which comfortably straddled zero.
- ^
Or alternatively: there are two representors with narrow but non-overlapping CIs.
- ^
Thanks Vasco, this is helpful.
I would be interested in a version of this contest attempting to answer something like âIf you believe Vasco that donations to SWP are better than torture, should you also believe that AI alignment is better than misalignment?â
I think that is what Richard Chappell is referring to when he mentions âradical skepticsâ here. But I would be interested in a version of his post which defends the claim that EAs are reasonably justified in making the trade-offs we currently make (even if a radical skeptic would not be convinced of this defense).