Non-EA interests include chess and TikTok (@benthamite). Formerly @ CEA, METR + a couple now-acquired startups.
Ben_Westđ¸
Thanks! Tbc, I agree that there are persistence effects; I just donât think they are so strong that we can reasonably expect that slight changes in starting conditions result in a completely different world.
if we were to travel back in time and alter starting conditions just slightly, it seems reasonable to expect that the world today would be completely different.
This is possible, but seems pretty unclear to me. cf The Gods of Straight Lines:
a few short years after the bombs stopped falling in 1945, the world economy returned to trend as if nothing had happened.
In 1776, America rebelled in the name of freedom and democracy: the origin myth of the modern world order. And yet, somehow, unrebellious Canada ended up just as free and democratic. An unrebellious America likely would have too.
For two decades, North Vietnam battled under the banner of communism, and won against all odds. And yet, somehow, Vietnam is now the most pro-capitalist country in the world.
$35/âattendee is an extremely impressive cost. Thanks for doing this and writing it up!
Thanks for writing this Sarah and best wishes for the new position!
I have been pretty pleased with the 2026 Forum output (particularly the unawareness event caused me to think more about my own work than most other things, maybe more than any other online event in 2026). Kudos to the rest of the team, and to your leadership for enabling that.
Median european seed stage pre-money valuation is âŹ5.0 million. People often raise the seed round within 1-2 years of starting. So âŹ1M is donating 20-40% of your company.[1]
- ^
Of course a large fraction of companies never make it even so far as a seed round. Itâs very hard to figure out what reference class someone should be in but it is worth noting that there are a good number of people who read this forum who have achieved valuations greater than that average within a couple years of trying.
- ^
Cool idea! I like the move of taking a selfish ambitious drive and trying to transform it into something more altruistic.
To see this, letâs be generous and say the expected number of future well-off people added by preventing an existential catastrophe is only 10^40
I think your argumentation supports the 10^40 number more than the âwell-offâ claim. Iâm not sure for the best canonical source expressing skepticism of the âwell-offâ bit, but The Future Might Not Be So Great contains a long list of arguments about whether future people will be well off; I would be interested in you responding to some of those.
Anthony cites Greaves and MacAskill giving an example similar to your gunpowder one:
Consider, for example, would-be longtermists in the Middle Ages. It is plausible that the considerations most relevant to their decision â such as the benefits of science, and therefore the enormous value of efforts to help make the scientific and industrial revolutions happen sooner â would not have been on their radar. Rather, they might instead have backed attempts to spread Christianity, perhaps by violence: a putative route to value that, by our more enlightened lights today, looks wildly off the mark. The suggestion, then, is that our current predicament is relevantly similar to that of our medieval would-be longtermists.
I personally think these examples are less compelling than they first appear (e.g. the persistence literature generally finds weaker effects than what you might imagine), but I agree that a failure of EAs to find examples of sign flips doesnât mean that future ones wonât exist.
Thanks! I canât tell if this is cruxy, but for what itâs worth your âpessimal inductionâ vignettes donât resonate with me in a way which makes me less motivated by the unawareness concerns.
For example, Bostrom coined the phrase âattention hazardâ in 2011. I remember someone telling me that MIRI was net-negative for this reason at EAG 2015, and I would be surprised if e.g. Habryka hadnât considered this risk before starting Lightcone. So I disagree with citing him/âthis as a good example of unawareness; itâs more that they mis-estimated a known risk factor.
Similarly, I remember talking about SARP at one of my first EAGs. I think I came across it in Brian Tomasikâs 2007 post, maybe even before I had encountered EA. Perhaps Iâve mis-estimated those concerns, but it doesnât seem like unawareness.
My overall experience is kind of the opposite of yours: when I got involved in EA people talked a lot about âCause Xâ and âCrucial Considerationsâ and now theyâve mostly just⌠stopped? Like people tried to find other considerations, and thereâs some new stuff around s-risks and weird decision theories etc., but if you look at what people talk about at EAGs today it feels mostly like more precise versions of what was discussed in 2016, rather than a large and unpredictable jump from the older understanding. Or, more technically: it feels like weâve had updates in evidence-space, but not as many updates in hypothesis-space, and I understand the latter to be motivating imprecision.
Obviously, this could be because EAs suck at cause prio research, or we just havenât been hit yet with the big update, etc., but the âpessimal inductionâ seems less pessimal to me.
Itâs great that youâre thinking about this!
Iâm confused why you are denominating options in robotics-startup-days saved. This feels like a narrow definition of âimpactâ. Iâd encourage you to consider other ways to benefit the world; parts 4-6 of the 80k career guide might be helpful. Specifically, under the assumption that the thing you terminally value is more like âreducing sufferingâ than ârobotics progressâ, I would encourage you to first consider which causes advance those values, and only then drill into job options. (The 80k career guide will walk you through this.)
Also, even to the extent you do just terminally value robotics progress, you might want to consider whether robotics will advance too quickly for your estimation to be accurate.
You respond to Richard Ngo here:
> do you think that, if we had a theory of sociopolitics that was about as good as 20th-century economics, then we wouldnât be clueless about how to do sociopolitical interventions (like founding AI safety movements) effectively?
No, because I think âfounding AI safety movements that succeed at making the far future go betterâ is a pretty out-of-distribution kind of sociopolitical intervention.
Suppose instead we had a comparably good theory of the right reference class, e.g. âmovements trying to shape transformative technologies.â Would we still be clueless about AI safety movement-building?
More generally: you list various considerations across your posts and I have a hard time understanding which is load-bearing for your answer here. Some possibilities:
Weâre clueless because we havenât yet developed the relevant theory (Richardâs reading IIUC, on which cluelessness is contingent and reducible)
No such theory could be validated even in principle, because we never observe the target variable (far-future value) and calibration on near-term proxies doesnât transfer
Even a validated theory wouldnât help, because impact is dominated by considerations inaccessible to any theory (e.g. unconceived hypothesis classes)
Whatâs the clearest example of a complex cluelessness sign flip youâre aware of?
(By âclearâ I mean âhad a very narrow confidence interval before encountering some consideration and a narrow interval after encountering that consideration but the CIs now center points with opposite signsâ.[1])
The clearest examples I know of (e.g. rescuing Hitler as a child) seem to me like examples of simple cluelessness. You list some examples here, but they donât seem that clear to me, e.g. I disagree that âEarly awareness-raising about AGI x-risk presumably seemed robustly goodâ and would guess most people involved in that had CIs which comfortably straddled zero.
- ^
Or alternatively: there are two representors with narrow but non-overlapping CIs.
- ^
Thanks! âDonât arbitrarily favor some moral patients over othersâ was the most intuitive definition of âimpartialâ in my mind, so I was confused about how it was being used here.
I feel somewhat confused about what exactly the challenge here is. You say:
Grant all the premises but argue that the conclusionâthat we have no impartial-altruistic reason to prefer any action over anotherâdoesnât follow.
My understanding is that Anthony agrees we can justifiably prefer actions, and even do so on altruistic consequence-based grounds if something like bracketing works â but heâd deny that either deontological reasons or bracketed reasons are âimpartial altruistic reasonsâ in the sense his conclusion targets: the first arenât consequence-comparisons at all, and the second are consequence-comparisons that give up full impartiality to stay determinate.
Is my understanding correct? If so, I would find a more precise statement of the inference people are supposed to challenge helpful.
My understanding is that Anthony agrees that there are still reasons to do things:
First, the unawareness argument doesnât imply that ânothing we do mattersâ all things considered. It only implies that impartial altruism, or any very far-reaching value system, isnât action-guiding. Other values and moral norms still matter to us, for example, rules like avoiding dishonesty or virtues like compassion. These can be action-guiding even if weâre clueless about total consequences.
I think your justification âbecause all of my attempts to do good actually end up being a net positive for me in terms of my own self-interestâ doesnât disagree with his conclusion?
Surprising that your MP was willing to meet with you for so long. Thanks for doing this (and writing it up)!
Sure, not the metaphor I would use but I broadly agree that applicants who are willing to plug the metaphorical USB stick into their computer (e.g. by following people they want to work with and applying to the jobs that they post) have a much lower rejection rate.
Sure, âhiring managers being bad at marketing is the bottleneck, not fundingâ is at least partially true. It still implies that if you happen to stumble across a poorly advertised position, you shouldnât expect the acceptance rate to be low!
My version of Mattâs critique that you quoted is something like:
Imagine youâre running a mining company, and you want to start mining Venus. You could either embark on a massive terraforming project to make Venus habitable by biological humans who can work in your mining colony, or you could just build a bunch of robots who can naturally withstand Venusâ climate, think faster than a human, make better decisions than a human, etc. etc. What do you do?
Obviously you are going to choose to send the robots, and the robots arenât going to want to eat meat, so you donât need to worry about factory farming on Venus.I donât think this argument is bulletproof. For example, ports in the U.S. are required to pay human dock workers to sit around and do nothing after their jobs had been automated. I could imagine some sort of analogous regulatory capture in the future which would require mining companies to send humans to other planets even when robots would be more efficient. Preventing this kind of lock-in is one of the few interventions targeting a post-singularity world that I feel positive about.
There are even AI-safety-pun-based sports teams!