The “Desperate AI Safety Talent Bottleneck” is deeply misleading. Please.stop.
I worked at Google.
Google does not tell applicants that there is a talent bottleneck and they are desperately hiring. They say “we’re cool, join us!”
I didn’t get into Harvard.
Harvard does not tell applicants that they are desperately seeking students. I am not misled.
I’ve gotten rejected from countless AI Safety organizations.
AI Safety tells applicants—we are in desperate need of talent (operators/generalists)! Please join us! I am misled and confused or maybe simply not good enough even in times of ‘desperation’.
I’ve coached 70+ aspiring career pivoters on navigating the ecosystem. A common theme is the dejection & disappointment from all this rejection. Because of this false marketing, I am the one picking up the pieces to calibrate professionals on how difficult, competitive, and picky organizations are, how to build context, and how to endure the marathon, which is like any other job hunt.
Please.stop.
Some alternatives:
We are an exciting, growing organization and you should apply for our roles!
Many people are excited to work for us—here are qualities of candidates we’re especially excited about.
EA roles are competitive, and we would still love to see your application.
Is it harmful to keep your money in index funds? Let’s take a random index fund like FTSE Global All Cap Index Fund. If we look at the companies it actually invests in, a lot of them are the same companies that work towards advancing the AI.
It would be ironic for some of us to be in favour of pausing AI and invest money into it at the same time. I used to keep some of my money in such funds because that was always the financial advice for regular folk like me who don’t know/think much about investments. But now I don’t know what to do because this feels bad. I switch some of my funds to FTSE Developed Europe ex U.K. Equity Index Fund because that has maybe twice less investment into AI but that still feels not great. Any advice for me?
There are two separate concerns. One is capitalizing the companies and the other is changing your incentives. I’m generally more worried about the latter, esp if you aren’t super rich. Do you think having all your money in index funds makes you less likely to support regulations that the unbiased version of you would want to support?
I’m not a student, but I’m surprised UC Law SF(formerly UC Hastings) doesn’t have an EA club or AI safety club. Seems like an interesting spot for people interested in AI governance work that could be well-served by UC Berkeley or Stanford pitching people to join their club meetings.
If this comes from the EA side of things, people should know that the optics of “paid protests” are widely regarded as absolutely horrible! It’s a good way to a) make it look like you have no popular support, b) alienate potential supporters by affiliating your movement with random people who need $30 badly, and c) show that you are not taking this politics thing seriously.
If this doesn’t come from the EA side of things, maybe some people should look into where it’s coming from?
Some hope for you if you’re an EA with a chronic illness.
I’ve been reading the biographies of moral heroes, and I’d guess ~50% of them struggled with ongoing health issues.
Being sick sucks, but it doesn’t necessarily mean you won’t be able to do a ton of good
Florence Nightingale
Benjamin Franklin
William Wilberforce
Alexander Hamilton
Helen Keller
It’s not everybody, but it’s a surprising percentage of them.
I myself struggle with a chronic mystery ailment and I find it inspiring to hear about all of these people who still managed to do great things, even though their bodies were not always so cooperative.
(Also, just as an aside, the book Stop Being Your Symptoms, Start Being Yourself didn’t cure my illness, but it did reduce its effects on my life by about 85%, and I can’t recommend it enough. It’s basically CBT/psychology stuff applied to chronic illness)
NB—this is almost entirely AI generated, with some back and forth prompts and corrections
I’m sharing a steelman against a live assumption in Bay/EA/AIS circles: that large AI-lab-adjacent philanthropy is likely to arrive soon enough, and in a sufficiently usable form, that organizations should plan around it.
The stronger skeptical case is that IPOs, valuations, pledges, DAFs, and foundation stakes are several gates away from fast, flexible, AIS/EA-directed grants.
The interactive model lets readers vary assumptions about Anthropic valuation, founder ownership, pledge follow-through, employee giving, OpenAI Foundation allocation, lockups, deployment rates, and grantmaker capacity.
Some commentary. I mostly agree with the page, but I will focus on the bits where I see room for improvement:
All 8 gates look correct to me, but they don’t all deserve equal emphasis.
Gate 1 says IPOs have lock-ups. That’s true but I basically don’t think that matters because lock-ups are very predictable: they will announce how long it is, and that’s exactly how long it will be. There’s no uncertainty. The main reason it’s relevant is that a lockup gives more time for AI valuations to fluctuate or collapse, but the text doesn’t even mention this.
Gate 5 and Gate 7 seem like they’re saying the same thing.
Gate 8 (“Bay social incentives”) seems uninteresting since it’s not a claim in the same category as the others. It’s more like a meta-level reason why people might not think about the other 7 gates.
Would be cool for the BOTEC to use distributions rather than point estimates. (Squiggle is good for this, and Squigglehub even has a built-in way to have AI generate models.) IMO distributions are a lot more informative than point estimates.
It looks like the default estimates in the BOTEC are pulled from the sources, but it’s not clear which estimates came from which sources. There should be inline citations.
“Anthropic valuation” variable should specifically be the valuation at the end of the 6-month lockup. Doesn’t matter much for a point estimate but it would increase the variance if there variable were a probability distribution.
Unclear what the “Founder pledge” variable refers to. Is it the % of pledgers’ wealth that they’ve pledged to donate? If so, the default of 80% seems really high?
“Employee committed pool” is defined in terms of dollars rather than as a % of company valuation, which seems weird. Shouldn’t it depend on the value of the equity?
This model is supposed to illustrate how a lot of people are being too optimistic, but even then, I think most of the point estimates in the model are too optimistic. Consider that e.g. the median self-reported earner-to-give only donates (IIRC) 3% of their income.
IMO “Deployment by end-2026” should use a different date. IPO 3-6 months from now plus 6 months lockup means no money will be deployed in 2026, unless Anthropic does a fast IPO + early lockup release. Even by the end of 2027, you’re talking about a 3-9 month turnaround time on lockup ending → grants being disbursed. FTX Foundation donated $190 million (pre-clawbacks) in about 6 months, which was ridiculously fast compared to a typical foundation, and that was still a pretty small % of its long-term budget (or at least, what was believed to be its long-term budget before FTX collapsed).
I would delete the OpenAI Foundation bit because (1) the model has enough parameters already and (2) I doubt OpenAI Foundation will give much money to causes that look good by EA lights.
“Grantmaker capacity multiplier” seems nonsensical as written. Shouldn’t the capacity max out at 1x? If grantmakers are a complete non-bottleneck, then the other parameters will dictate the amount disbursed; if they’re a bottleneck, then the amount disbursed will be less. There’s no way for grantmaker capacity to have a multiplier >1x.
Also this would make more sense as a dollar amount, not a multiplier. Like there’s a fixed total amount that grantmakers can reasonably disburse. You could model it in a more complicated way but IMO a simple cap is the way to do it. Or maybe don’t use this parameter at all. I think it’s probably worth including, but keep it simple.
“Field absorption ceiling” is structured more sensibly than “Grantmaker capacity multiplier”, but these two seem redundant because they’re closely related. If orgs have more capacity to expand, grantmakers can deploy money faster by giving to those orgs. If there are more grantmakers, they can create more and bigger RFPs. etc. I would include one variable or the other, but not both.
“The steelman could be too pessimistic if founders or employees treat liquidity as an urgent moral obligation” – TBH the BOTEC as written seems to me like it’s already pricing in that founders/employees will treat donations as urgent, e.g. it’s implying that Anthropic money will be disbursed faster than FTX Foundation money, which itself was disbursed at historic speed. IMO most likely reason why the model will end up underestimating is that Anthropic market cap ends up being like 10x higher than predicted.
My downward adjustments to the model aren’t even the pessimistic case. The pessimistic case* is that the AI field collapses (investor funding dries up or something) and Anthropic stock is worth $0. Base rate says there’s like a 50% chance that that will happen. Even optimistically, you should expect at least a 10–20% chance that Anthropic stockholders get nothing.
The recommendations under “How to plan if the skeptical case is live” don’t really make sense. AFAIK ~zero orgs are planning as if they’re guaranteed to get a huge pile of donations 1–2 years from now. I believe nonprofits mainly plan based on the money they already have on their books + short-term (<1 year) fundraising expectations. “How to plan if the skeptical case is live” is just “business as usual”.
That section says “Funders and field builders should prioritize grantmaker capacity”, but that’s what to do if the skeptical case is wrong, not if the skeptical case is right.
“What would update this memo?” – as with the gates, no sense of prioritization is given. IMO by far the biggest uncertainty, about which we will get more information in the future, is: What will Anthropic’s valuation be when the lockup ends? “Concrete donor vehicles” is also important evidence, but we won’t get that until probably 6-24 months later.
*this is pessimistic for donations but I would actually prefer that this happen because it would lengthen timelines. so in a way it’s the optimistic outcome
AIS/EA: median modeled end-2027 disbursement $0.83B; 80% model interval $0.16B–$3.0B. Review status: reviewed; unchanged. Named public AI-linked commitments tracked: $0.80B. Tracker and sources.
Global health and development: median $0.53B; 80% model interval $0.15B–$2.8B. Review status: reviewed; unchanged. Named public GH&D commitments tracked: $0.25B. GH&D tracker and sources.
Latest news: Anthropic opens a $5M wellbeing-evaluations grant program. The program combines direct funding with model access and technical support for independent, open-source evaluations. It is a confirmed commitment, not a completed grant or cash-disbursement record.
The intervals are deterministic model percentiles, not empirical confidence intervals. Commitment totals may include credits, technical support, cofunding, and multi-year plans; they are not cash-paid totals. Automated fortnightly update from the maintained model.
I totally agree on using distributions, that’s something that can be incorporated in, I’ve done so in other models/interfaces like here for cultured meat. It’s by no means straightforward though; the extent to which the uncertainty is dependent/correlated tends to make a big difference.
I guess I see the deterministic ‘model’ as more of an interface people could use as a starting point, playing around with each parameter interactively and getting a sense of how these disturbances would affect the aggregate forecast.
Added and responded to your comments on the page (the hypothesis comments), and then I asked Codex to update to these https://uj-ai-wealth-philanthropy-steelman.netlify.app/ … I haven’t inspected the latest version in detail yet, though.
Some highlights of particular interest to @MichaelDickens , Tobias, and readers/modelers
More citations, tooltips, direct links, and quoted support for key factual claims, greater reasoning transparency, auditing.
“Updated the model/text in response to his critiques: especially lockup timing, over-optimistic deployment speed, founder pledge ambiguity, OpenAI Foundation treatment, grantmaker capacity, collapse risk, and the need for distributions rather than just point estimates.”
Clarified “the conversion terms are sequential gates”, Made “realization,” “follow-through,” “allocation,” and “deployment by end-2026” more distinct, reduced risk it seems like double-adjusting
Reframe original ‘calculator’
Added a new tab for correlated uncertainty, added shared latent factors such as liquidity plumbing, donor intent, grantmaking capacity, local optimism, and deployment pressure.
I found the ‘founder deployment by end-2026’ the hardest to set. It comes a bit as a surprise at the end, as I was already taking into account some considerations before, and the descriptions seem to do as well (e.g. “assets after lockups, taxes, sale timing”, and “execution delays”).
Biggest (easily-fixable) outstanding issue is I still don’t think it makes sense to model deployment by end-2026 because the IPO lockup probably won’t have ended by then.
I’m early-ish in the application pipeline for a non-EA job and they’re asking me to do a 20-HOUR UNPAID work test where they explicitly want to own and use the product of my “work test” ….
besides this obviously being really cringey and maybe illegal (?) it was a good reminder for me of how unpaid work tests are just a massive hinderance to potential applicants—who the heck has 20 hours of unpaid work to donate unless they’re unemployed?
I believe there is immense untapped potential in narrative change through the applications of “short-form video marketing tactics” to the Animal Welfare movement.
What do I mean by “short form video marketing tactics”.
Specifically I mean clipping campaigns the likes of which have helped launched the career of some of the 2020′s most popular musicians including Sombr, Alex Warren, Geese, and Oklou.
Clipping campaigns essentially pay people on a per view basis for tiktok, instagram reel, and youtube shorts. Anyone can create a clipping campaign and usually they use a combination of Adobe After Effects and other video editing software to create a “Narrative Campaign” to bring the creator into the zeitgeist. Plenty of twitch streamers do this too including IShowSpeed who grew to becoming the most watched streamer of all time and performing at the World cup. The beauty of this marketing strategy is that it creates the feeling of “tons of people are posting about this person and its infiltrating my foryoupage so it must be a big deal.”
The beauty of a clipping campaign is you’re essentially creating an incentive for video editors from around the world to find the best ways to bring the content to the forefront of people’s fyp. I understand awareness is not the epitome of change, but I do wonder if people would care as much about Israel and Gaza if there was no short-form photo or video of the situation.
I think short-form video has heavily contributed to people on the right and left shifting their political positions on the issue.
I’m surprised there isn’t much focus on human rights / intl law efforts in the EA community. There obviously are policy advocacy efforts, but I haven’t seen any in this realm.
Why is that? Is this a conscious choice, or is it incidental?
TLDR: I’ve updated towards pausing further AI development indefinitely.
When Scott Alexander proposed regulating AI like clinical drugs, many (including myself) balked at this given the sclerosis of FDA or related bodies. Yet HF updates me towards treating frontier AIs as nuclear. There onerous regulation is likely good.
In general, most bureaucratised regulatory regimes are bad. In some cases, like nuclear technology or ensuring planes are safe to fly on, they’re welfare improving. It’s clear that AI belongs in the latter category.
Also trace inversion is a thing, so you can get open source weights to be roughly similar to that of Fable/Mythos. The slowdown camp were right. At bare minimum, all labs should pause training, further development, and releases indefinitely until we figure out how to align AIs and regulate them (and their use by humans).
I reckon the current capabilities we have now are sufficient for accelerating progress towards curing cancers etc. So my balance has shifted towards minimising the existential risks now.
Pause AI advocates are correct. If governments could coordinate internationally to achieve such (big if), I’d support it. I used to be highly sceptical of doomer arguments, yet Hugging Face is almost a textbook LW scenario and no one knows how to spot or prevent such scheming. Sandboxing, guardrails, constitutions etc. don’t work.
I think that there can be strategies that make AI the more safer the more capable it becomes. There was a time when I became safer when I became more capable as I thought, reflected and kept certain actions off limits. There can be core and basic stratgies to be used for making AI safer the more capable it becomes and now we are going in the opposite direction with about 50% probabality.
There must be some humans that atleast for a point of time became safer the more capable they became and their strategies whichever are compatiable for AI can be used. Instead of all being black and white there must be certain tyoes of capablities that make AI safer instead of more dangerous. Editing the concept of AI to make it no longer be AI that is better than AI could help as AI might not be the only threat that people could face as in future there could be more dangerous concepts created like people could create variations of the concept of AI that are more dangerous for their time.
Instead of banning AI the concept of AI if it becomes obselete and safer with another concept could help. People might need stratgies to help better than banning AI ever can.
There could be atleast 3 types of concepts that are more dangerous AI successor, less dangerous AI successor and equally dangerous AI successor and if a concept is made that is less dangerous AI successor that makes AI obselete that may help although I an not sure if there is moderate work in making AI obselete.
Humans had handled harmful substances that together make useful substances and if that concept could be fused with AI making it different but safe like an safety booster then that could help and that safetly booster only changed the concept of AI after being fused with it then it might be progess and a strategy could be instead of having to contain AI that is smart it could be combined with another concept perhaps by adding additional lines of code to it to make it no longer AI and change the problem of contaning AI to contaning the Fused AI concept and if that is not smart that “that could be better” and it might not need to be smart.
A strategy of contaning AI might be to no longer make it be an AI by adding additional lines of code to it but that might fail if those lines of code are removed. However it could be another strategy and better than nothing and it might be like instead of trying to contain the spread of poisionous gass in a place by releasing gas that merges with that poisions gas and makes it harmless and it might had been done before with precedent. This is not medical advice.
I guess that there might be only a few ways that superintelligence AI could harm us and if those ways are stopped then it might help but it could be like hypothetically trying to prevent a slightly posionous snake from trying to harm a man and it might be better to add a guard to snake’s mouth to prevent it from biting and spitting poison and snake catchers have existed with snakes for a long time perhaps but I hardly if ever had heard the news of a snake catcher being bitten by a snake. I have heard more cases of snakes biting ordinary people and snake catchers might be using strategies to be avoid being bitten by snakes given their number(of the snake catchers and the number of snake catchers perhaps an estimate was probabaly written in an old NCERT of class 7-10(In Social Studies NCERT if I am right).
The topic of this quick take is AI superintelligence and avoidance strategies.
How impactful would it be to copies of @Garrison’s new book Obsolete to elected officials who belong to its political target audience (Dems, particularly left ones) and might not have been responsive to traditional x-risk-centric comms (e.g. IABIED)?
There’s an old Facebook group, New EA hub search and planning, that might be worth reviving in light of the incresed amount of funding coming into the space.
Even with more availability of funding, many talented folks won’t get funded or will only recieve small amounts. Establishing new hubs that provide an easier entry-way to being surrounded by community seems like potentially quite high value. After all, it takes a village to raise an EA and being surrounded by others is an accelerator to impact.
What’s a good heuristic for knowing if you’re being self-sacrificial to the point of negative returns for the thing you’re being self-sacrificial for? I’m interested in things people have actually used and found to work in practice, either for themselves or others.
I’ve never been able to find a heuristic for this exactly. We have made many decisions which we don’t love to call sacrifice, in fact many of them brought deep satisfaction and connection. These were made to live in better solidarity and community with those around us. For example for us this involved cooking on a charcoal stove, not having running water inside and eating a lot of local food which is often great but I don’t always love.
Almost all of these were made at the cost of efficiency in our work to some extent. Life without a gas stove and running water just takes longer and is a bit more tiring. Not nearly as big a deal as most people would think tho...
The closest we had to a “heuristic” was asking ourselvesevery 6 months, is this lifestyle...
1) Satisfying to me 2) Appropriate lifestyle that matched our community around us. We’ve always chosen to live among less well-off communities because I think it grounds our work in reality, keeps it real and helps us learn what actually makes people tick. (which I personally think is often critical to designing effective GHD interventions) 3) Allowing me to do my work well
As OneDay Health and my Wife’s work grew, I needed more time. No. 3 started to take precedence. We became more “bougie”. A water tap outside, more eating out (still local food). Now after having a kid we have a gas stove. It doesn’t feel great to move further away from our community lifestyle, but I think its worth it for the impact—as long as we don’t feel too disconnected from those around us.
After the initial experiment period, we’ve decided to keep the featured page.
One of the key graphs in making this decision is below:
Now that the featured page is likely to stick around long term, I’d like to start collecting ideas for V2[1]. Do you use the page? Do you pointedly not use the page? I’d love to get any and all feedback, in the comments here, or via this form.
Please share small things like aesthetic tweaks as well as larger issues such as a disagreement on the kind of posts that get featured, or how featuring is done. Happy to answer questions too.
I.e., the current version is a prototype, and I’m ready to do a full redesign if necessary. Because of some protracted OOO I might have to take soon—V2 might land near the end of September, or even early October.
I’ve really stopped engaging with the EA Forum as much since the change, if helpful feedback. I’m not sure what to do about this, but I’ve found I no longer have a clear sense of “where conversation is actively happening”, and while I guess I can tell from the tab names if I’m in New or Featured, in reality I don’t pay enough attention and miss a lot of things. E.g. today I realized there was a post that didn’t get much attention that I never noticed that was relevant to me a week ago, and generally, I have a sense of “I can’t really tell what’s recent conversation” anymore, so I’ve stopped looking at the Forum as much because I don’t have a clear sense that I’ll see new and interesting things.
I see the upside of the change, but personally have found it hard to navigate the site since the change in a way that’s decreased my engagement.
Thanks for all you do to experiment and try to make it better! Appreciate how much you all do here to make the Forum useful!
Thanks Abraham, that’s good to know! FWIW the new & upvoted page is just the Forum frontpage as it was before, so if you just ignore the featured page, your forum experience shouldn’t change at all.
Also we did feature that post—so if you check the featured page you’d have been more likely to see it overall than previously (when something niche but impactful like that was likely to get low karma and sink off the frontpage pretty quick).
Mental health is one of the most widespread, impactable, yet neglected global health issues of our time. Even knowing mental health conditions are generally under-reported and/ or misrepresented, evidence still shows widespread effects of mental wellbeing on physical health, life satisfaction, and productivity, however it has still not taken off as a major cause area in many EA organizations I have talked to. Why is this?
I should mention there is an argument that many mental health interventions are in fact harmful, mainly via a priming mechanism where when people start to consciously preoccupy themselves with their own mental health, that itself causes psychological suffering. It is a sort of Buddhist argument, I guess.
Abigail Shrier’s Bad Therapy (2024) has probably most discussed version of that case, though I have not read the book, and I am extremely epistemically uncertain about all this, in a way that makes me reluctant to either support or oppose interventions in this area.
Interesting take @alesziegler ! I do believe it is true that too much of any one thing is usually bad, and also believe that people who are not trained/ underqualified taking on mental health is counterproductive. I think people take on these roles or use clinical language that should not be applied to what they are doing precisely because they cannot get the credible help they need and are trying to fill that void. Mental health is extremely subjective and therefore interventions may look like community connection or involvement in sports rather than clinical therapy.
I am also highly skeptical of the legitimacy of Schrier’s claims and her choice in very carefully selected stastistics as evidence. For ecample, contrary to some of Shier’s claims, evidence-based school-based emotional programs have had profound impact, especially on marginalized populations (https://www.sciencedirect.com/science/article/pii/S2773233924000032). Maybe my thoughts on this is best summed up by saying, “Don’t throw the baby out with the bath water”?
In addition to @huw’s great comment, there’s a branch of EA which focuses on well-being, and with that focus mental health interventions often look really effective. Check out the happier lives institute Huw mentioned, times the “WELLBY” and @MichaelPlant
Thank you @NickLaing ! I have done a lot of reading on the Wellby and that is one of the reasons I feel mental health is currently under-represented as a global health issue. It is fascinating to hear how people globally weigh the importance of wellbeing and I’m curious how many other aspects of their lives would be more easily improved if wellbeing increased! @MichaelPlant would love to connect on your work!
G’day Madeline, I run an EA mental health org in India. The reason for this is simply that existing mental health interventions do not compare with GiveWell’s grants on a DALYs/$ basis. In my opinion, the reasons are:
DALY moral weights may be biased against depression
The moral weights of different diseases in the Global Burden of Disease study, which informs DALY estimates, are determined by asking the general public whether they’d prefer to have one disease against another. When you do this with depression, people who haven’t experienced it tend to prefer to have it to many other conditions. However, when you ask people who have experienced it, they choose many very painful conditions over depression. This is one of the widest gaps in the moral weight data. See Pyne et al. 2009 and this post.
Psychotherapy is usually modelled as a short-term effect
Psychotherapy is typically modelled as a treatment, and not a ‘skill’. What I mean by this is that a dose of psychotherapy is assumed to only have effects that decay over a period of time and zero out after that in most CEAs, including those from the Happier Lives Institute. However, many psychotherapy patients will tell you that they learned skills that were useful long after the therapy ended, and there is some limited evidence that psychotherapy’s effects may last decades, or potentially never zero out. If this were true, the effects could be very long-lasting and therefore it would be much more valuable to treat a case of depression.
Suicide prevention isn’t cost-effective if it’s only a short-term effect
Consider that for most of GiveWell’s top interventions, the bulk of the DALYs averted come from ‘saving’ a life—i.e., preventing a death from a disease in a way that allows the person to then go on to live a healthy life, such as preventing a malaria case in an under-5 (which they might die from), even if they go on to catch it after 5 years old.
As a short-term effect, psychotherapy can only postpone a suicide by the length of the treatment effect. But if it were a skill and had some durable long-term effect, it may genuinely prevent one, which would tremendously increase the value of suicide prevention interventions.
Existing interventions haven’t been cheap enough yet
With the exception of some incredible policy work in, for example, reducing toxicity of pesticides commonly used for suicide, existing interventions are still quite expensive. The Happier Lives Institute’s top charities cost ~$40 to treat a single person, while a bednet costs $7. I’m fudging the numbers a bit here, but if we stick with DALYs, psychotherapy is still about an order of magnitude more expensive than it needs to be to look great for EAs.
However, there is work being done to improve that! My charity, Kaya Guides, treated people for $20 each in April, at what we estimate is a similar effect size to the best charities, and we’re confident we can get below $10. We’re using a technique called guided self-help that allows us to dramatically reduce contact hours per participant (and being all-digital helps a lot, too).
Conclusion
Orgs like the Happier Lives Institute have done a lot of advocacy work too, to raise the profile of mental health within EA, and there are plenty of funders that take mental health seriously (in a way that apparently wasn’t true a decade ago). It is, after all, still a nascent space.
Thanks for mentioning our work in HLI here @huw. I/we are not so active on the forum these days, so you beat us to it.
But yes, @madeleine_foley, HLI really got started, now 7 years ago, because we thought that the standard economist approach of focusing on just health and wealth was not going a good job of capturing what people’s lives are like on the inside. We’ve been pioneering wellbeing cost-effectiveness analysis and a metric called WELLBYs (wellbeing years) to see what the priorities are if you ‘take happiness seriously’. (Oddly, we were the first team to do wellbeing ROI, but now quite a few others, including the UK and NZ Treasuries, use the same method, although not because of us). We concluded that mental health is a neglected priority, and have been recommended charities working on it for the last few years, and advocating for it more broadly. Huw’s organisation, Kaya Guides, is one we’re very excited about, and he’s right that global human happiness and mental health is an emergent part of the wider cause area of ‘global health and wellbeing’, and just wasn’t part of the conversation in effective altruism before.
Thanks @MichaelPlant ! I actually recently connected with Joel from your team and am very interested in leveraging some of your work to influence priorities in my current role. I’m eager to continue talking about mental health and wellbeing in the EA space and hoping to get more involved with HLI or the work you are doing on a volunteer basis!
Thank you @huw and apologies for my delay- still getting the hang of the forum! I really appreciate these insights and would love to learn more about the work you are doing at Kaya Guides. I have been thinking a lot about what innovation would look like in the mental health space and how the personal nature of mental health makes widespread interventions challenging (for example, medication can not necessarily cure mental illness the way it can cure a physical ailment like an infection).
[central europe] We are organising the largest AI-safety march (at least in central europe) to date at Prague, starting at 15:00 on 31st August. We have got parlimentary support as well as very supportive police, allowing us to take the preffered route. We want as many people to come—fell free to spread information about the event outside EA as much as possible, especially if you have friends living near Prague who may come. If you want closer info/visual materials or help us with organisation, you can get in contact with czech PauseAIon whatsapp.
I wrote this in a post, but it might be a better quick take: Would you describe charities within GCR as “super-effective?” They may or may not be good long-term bets, but clearly the outcomes are very uncertain & hard to trace. Whereas GHD + Animal Welfare are probably much more trackable.
I’m asking because I saw on GivingMultiplier.org that they were marketing GCR charities as “highly effective,” and that gave me pause.
I think the description is very justifiable, assuming that these charities are indeed good long-term bets.[1]
I agree with you that there is a difference between “a good long-term bet” and “a super-effective charity”. However, for most people, it is a really bad idea to share a fully nuanced description of whether a charity is effective based on empirical data vs a high-risk high-reward bet that has a robust theory of change but is operating in a field where empirical data is impossible to collect. While this might be more precise, there are strong tradeoffs between precision and clarity. The full explanation would make most people stop engaging, while the simplified explanation will bring people into the movement.
For many in the EA space (including myself), it can feel dishonest to say simplified claims that are imprecise but still directionally correct, even though no dishonesty is occurring. This is ultimately a matter of context and audience. Describing a well-justified but speculative bet as “highly-effective” as a funder deciding who to support would be bad. As a community builder reaching out to an audience potentially unfamiliar with concepts like expected value, it’s much more reasonable.
My general leaning is that most, but certainly not all, GCR charities are not good long-term bets. This poses more fundamental challenges to your point, but the solution would be to increase scrutiny in the GCR space, not fix the marketing. I’m setting this aside for the sake of this comment, as it appears to be a crux you do not share.
Well, the lack of consensus on GCR’s efficacy is exactly the thing I’m pointing out. A highly untraceable intervention should not be called “super effective” … I think that risks undermining their credibility. “Potentially a good long-term bet” is very, very different from “super-effective, expert-vetted.”
I was just wondering if there was something I was missing about the proven efficacy of GCR interventions. It sounds like I’m not...in which case, I think it’s a really bad idea for GivingMultiplier to pass those off like they have the proven efficacy of GHD interventions. It feels misleading, and I’m more “EA-friendly” than the average person they’re probably targeting.
And I have no strong view on GCR interventions. Seems like they’d be hard to prove traceability/impact, but I haven’t spend much time thinking about it.
I think there’s multiple approaches they could take:
Describe all charities as “super-effective”, a term that has presumably been tested directly and/or indirectly to maximize marketability
Use some other term that applies precisely to all charities but is less effective at drawing donations
Make a distinction between (e.g.) GiveWell and GCR work
Not include GCR in their portfolio
I can see an argument for 4 (i.e. start new donors with more straightforward work). Also, if GCR work is too speculative to justify pursuing, then 4 is clearly the best option.
However, if you believe that GCR work is sufficiently well-justified (as Giving Multiplier seems to), I do think 1 > 2 and 3.
2 and 3 decrease the effectiveness of the organization for the sake of precision. Their FAQ on the home page addresses how they defer to other actors, who make decisions based on the best available evidence. Anyone unsatisfied by that answer has links to Founders Pledge and can learn more about their methodology there.
Super-effective is a marketing label, not a scientific one. They define it reasonably, and I don’t think it would come across as dishonest or misleading to someone without EA context.
Fair enough. I have no doubt that they picked the term for marketing purposes. But I’m sure you’re right that they would defend it as “highest EV,” as opposed to “rigorously studied causal chain.”
I don’t think so. I’d say they have the potential to be “super-effective in expectation”. To me, “super-effective” with no qualifiers implies high confidence.
Agreed. I was surprised that they’d describe something so speculative as expert-vetted, super effective, etc. If something hasn’t undergone multiple RCTs, then I don’t think it should be promoted as such a highly effective, sure thing.
What if exact copies don’t matter but future experiences still do?
I think an ethics resulting from the following two premises is worth exploring:
P1: The total moral value of subjectively indistinguishable observer-moments doesn’t scale with the number of copies, i.e. they sum to a constant value.
This is because subjectively indistinguishable observer-moments, in some sense, happen to the same “someone”. One could think of themselves as all their copies. N observer-moments that are subjectively indistinguishable aren’t felt N times by that same “someone”. Only felt “once”.
E.g. Even if there were multiple exact copies of me, of my current present experience, I might not really care personally because it doesn’t seem to change anything for me. “I” don’t notice anything.
We can say that multiple “tokens” of subjectively indistinguishable observer-moments belong to the same “type” of observer-moment.
P2: What matters, and can be changed, for each type of observer-moment is the distribution of its subjectively “future” experiences.
You can’t “un-exist” the type of observer-moment itself, but you can change what it continues into.
A higher fraction of “future” observer-moments experiencing suffering is dispreferable.
A higher fraction of “future” observer-moments experiencing happiness/tranquility is preferable, even if only instrumentally to minimize subjectively “future” suffering.
This is related to thinking about “anthropic immortality” / “quantum immortality” / “multiverse immortality” in which distributions of post-death observer-moments are of substantial concern.
Using these lenses, the moral situation of the world looks like:
Consider a graph with types of observer-moments as nodes, such that one node is one type.
Each type of observer-moment (each node) is associated with some value/disvalue attributed to its current experience.
Importantly, the value/disvalue of the current experience is not something we can change in order to help the observer-moment type. This is already fixed.
Each type of observer-moment (each node) is also associated with any number of subjectively indistinguishable token observer-moments belonging to that specific type.
The absolute number of token observer-moments appears to only matter instrumentally insofar as it affects the frequency of expectations of other types of observer-moments.
Edges represent subjective continuation. An edge from type A to type B exists if an observer-moment of type A can possibly expect to continue as an observer-moment of type B in the next moment.
The weight of an edge from type A to type B is the frequency at which an observer-moment of type A should expect to continue as an observer-moment of type B.
These weights are things that we might be able to change in order to help the observer-moment type.
Perhaps the “future” welfare of each type / node could be considered with equal weight in the aggregate to keep things impartial.
In conclusion… well, I’m still thinking about what the objective function might look like.
There’s some scale invariance here. E.g. suppose you duplicated the world. That duplication would multiply all token observer-moments by the same factor. This leaves you with the same types and the same distributions for future observer-moments.
In a small enough world, you can create or prevent the existence of new types of observer-moment. I think it’s unlikely (30%) that we live in a small enough world.
On the other hand, in a large enough world, all possible types of observer-moments exist. The only thing you can change is adjusting the distributions for future observer-moments for each type of observer-moment. Try to lead types of observer-moments down paths of non-suffering etc. So, this seems like an ethics of flow (rather than stock).
In the case of finite trajectories, the objective function can be the sum of value functions across all states, with each state weighted equally—where states are types of observer-moments, and each state’s value function is its immediate valence plus the expected value of its continuations. This is total utilitarianism with two modifications. 1) aggregate over types instead of tokens, and 2) sum trajectory welfare instead of immediate welfare.
You should judge others on a curve, and probably yourself as well.
Judging others on a curve means that folks see reward for doing more than what is typical. This is just good reinforcement-learning theory, since it is easy to get both feedback on rewarded & not-rewarded actions.
I started writing this thinking that you should NOT judge yourself on a curve, in contrast. IE, you know much more about your own circumstances, and so pretending you are average throws away info. For example, if I only gave 5% of my income to charity, that would be above average for the US. But my opportunities are amazing, and so 5% would in reality show low effort.
But actually, I was just using the wrong comparison group! I shouldn’t compare myself to a big group of other people. I have a near perfect comparison group in my past self. If I did a better job this week than last week at work, that’s a great sign! If my one-rep max for benchpress has gone down over the last month, that’s a bad sign!
The “Desperate AI Safety Talent Bottleneck” is deeply misleading. Please.stop.
I worked at Google.
Google does not tell applicants that there is a talent bottleneck and they are desperately hiring. They say “we’re cool, join us!”
I didn’t get into Harvard.
Harvard does not tell applicants that they are desperately seeking students. I am not misled.
I’ve gotten rejected from countless AI Safety organizations.
AI Safety tells applicants—we are in desperate need of talent (operators/generalists)! Please join us! I am misled and confused or maybe simply not good enough even in times of ‘desperation’.
I’ve coached 70+ aspiring career pivoters on navigating the ecosystem. A common theme is the dejection & disappointment from all this rejection. Because of this false marketing, I am the one picking up the pieces to calibrate professionals on how difficult, competitive, and picky organizations are, how to build context, and how to endure the marathon, which is like any other job hunt.
Please.stop.
Some alternatives:
We are an exciting, growing organization and you should apply for our roles!
Many people are excited to work for us—here are qualities of candidates we’re especially excited about.
EA roles are competitive, and we would still love to see your application.
Relevant posts:
https://forum.effectivealtruism.org/posts/B6d8Wzk4gNzHsXvdi/ai-safety-is-extremely-bottlenecked-on-grantmakers?commentId=n2Rd7RR4y9EnKP42F
https://forum.effectivealtruism.org/posts/b82SLXwEHRCs3TFJA/why-experienced-professionals-fail-to-land-high-impact-roles
https://forum.effectivealtruism.org/posts/jmbP9rwXncfa32seH/after-one-year-of-applying-for-ea-jobs-it-is-really-really
Is it harmful to keep your money in index funds? Let’s take a random index fund like FTSE Global All Cap Index Fund. If we look at the companies it actually invests in, a lot of them are the same companies that work towards advancing the AI.
It would be ironic for some of us to be in favour of pausing AI and invest money into it at the same time. I used to keep some of my money in such funds because that was always the financial advice for regular folk like me who don’t know/think much about investments. But now I don’t know what to do because this feels bad. I switch some of my funds to FTSE Developed Europe ex U.K. Equity Index Fund because that has maybe twice less investment into AI but that still feels not great. Any advice for me?
There are two separate concerns. One is capitalizing the companies and the other is changing your incentives. I’m generally more worried about the latter, esp if you aren’t super rich. Do you think having all your money in index funds makes you less likely to support regulations that the unbiased version of you would want to support?
I’m not a student, but I’m surprised UC Law SF(formerly UC Hastings) doesn’t have an EA club or AI safety club. Seems like an interesting spot for people interested in AI governance work that could be well-served by UC Berkeley or Stanford pitching people to join their club meetings.
According to people I know who have personally seen the flyers, there are flyers offering to pay people to protest a recent Sam Altman appearance (unclear why they’re mad at Sam Altman, there are many reasons why they might be mad at him). However, the protest was cancelled because of ???. https://www.wunc.org/education/2026-09-02/protest-chapel-hill-unc-g20-innovation-technology-sam-altman-elon-musk-artificial-intelligence gives a bit of coverage of the story.
If this comes from the EA side of things, people should know that the optics of “paid protests” are widely regarded as absolutely horrible! It’s a good way to a) make it look like you have no popular support, b) alienate potential supporters by affiliating your movement with random people who need $30 badly, and c) show that you are not taking this politics thing seriously.
If this doesn’t come from the EA side of things, maybe some people should look into where it’s coming from?
Some hope for you if you’re an EA with a chronic illness.
I’ve been reading the biographies of moral heroes, and I’d guess ~50% of them struggled with ongoing health issues.
Being sick sucks, but it doesn’t necessarily mean you won’t be able to do a ton of good
Florence Nightingale
Benjamin Franklin
William Wilberforce
Alexander Hamilton
Helen Keller
It’s not everybody, but it’s a surprising percentage of them.
I myself struggle with a chronic mystery ailment and I find it inspiring to hear about all of these people who still managed to do great things, even though their bodies were not always so cooperative.
(Also, just as an aside, the book Stop Being Your Symptoms, Start Being Yourself didn’t cure my illness, but it did reduce its effects on my life by about 85%, and I can’t recommend it enough. It’s basically CBT/psychology stuff applied to chronic illness)
Was there anything in the book that you found especially helpful?
NB—this is almost entirely AI generated, with some back and forth prompts and corrections
I’m sharing a steelman against a live assumption in Bay/EA/AIS circles: that large AI-lab-adjacent philanthropy is likely to arrive soon enough, and in a sufficiently usable form, that organizations should plan around it.
https://uj-ai-wealth-philanthropy-steelman.netlify.app/
Original motivating thread/comment: https://forum.effectivealtruism.org/posts/dtF6wBjH7yBD4kqLz/noah-birnbaum-s-quick-takes?commentId=sGRyGF5wjaaoMFmfK
@Noah Birnbaum
Some commentary. I mostly agree with the page, but I will focus on the bits where I see room for improvement:
All 8 gates look correct to me, but they don’t all deserve equal emphasis.
Gate 1 says IPOs have lock-ups. That’s true but I basically don’t think that matters because lock-ups are very predictable: they will announce how long it is, and that’s exactly how long it will be. There’s no uncertainty. The main reason it’s relevant is that a lockup gives more time for AI valuations to fluctuate or collapse, but the text doesn’t even mention this.
Gate 5 and Gate 7 seem like they’re saying the same thing.
Gate 8 (“Bay social incentives”) seems uninteresting since it’s not a claim in the same category as the others. It’s more like a meta-level reason why people might not think about the other 7 gates.
Would be cool for the BOTEC to use distributions rather than point estimates. (Squiggle is good for this, and Squigglehub even has a built-in way to have AI generate models.) IMO distributions are a lot more informative than point estimates.
It looks like the default estimates in the BOTEC are pulled from the sources, but it’s not clear which estimates came from which sources. There should be inline citations.
“Anthropic valuation” variable should specifically be the valuation at the end of the 6-month lockup. Doesn’t matter much for a point estimate but it would increase the variance if there variable were a probability distribution.
Unclear what the “Founder pledge” variable refers to. Is it the % of pledgers’ wealth that they’ve pledged to donate? If so, the default of 80% seems really high?
“Employee committed pool” is defined in terms of dollars rather than as a % of company valuation, which seems weird. Shouldn’t it depend on the value of the equity?
This model is supposed to illustrate how a lot of people are being too optimistic, but even then, I think most of the point estimates in the model are too optimistic. Consider that e.g. the median self-reported earner-to-give only donates (IIRC) 3% of their income.
IMO “Deployment by end-2026” should use a different date. IPO 3-6 months from now plus 6 months lockup means no money will be deployed in 2026, unless Anthropic does a fast IPO + early lockup release. Even by the end of 2027, you’re talking about a 3-9 month turnaround time on lockup ending → grants being disbursed. FTX Foundation donated $190 million (pre-clawbacks) in about 6 months, which was ridiculously fast compared to a typical foundation, and that was still a pretty small % of its long-term budget (or at least, what was believed to be its long-term budget before FTX collapsed).
I would delete the OpenAI Foundation bit because (1) the model has enough parameters already and (2) I doubt OpenAI Foundation will give much money to causes that look good by EA lights.
“Grantmaker capacity multiplier” seems nonsensical as written. Shouldn’t the capacity max out at 1x? If grantmakers are a complete non-bottleneck, then the other parameters will dictate the amount disbursed; if they’re a bottleneck, then the amount disbursed will be less. There’s no way for grantmaker capacity to have a multiplier >1x.
Also this would make more sense as a dollar amount, not a multiplier. Like there’s a fixed total amount that grantmakers can reasonably disburse. You could model it in a more complicated way but IMO a simple cap is the way to do it. Or maybe don’t use this parameter at all. I think it’s probably worth including, but keep it simple.
“Field absorption ceiling” is structured more sensibly than “Grantmaker capacity multiplier”, but these two seem redundant because they’re closely related. If orgs have more capacity to expand, grantmakers can deploy money faster by giving to those orgs. If there are more grantmakers, they can create more and bigger RFPs. etc. I would include one variable or the other, but not both.
“The steelman could be too pessimistic if founders or employees treat liquidity as an urgent moral obligation” – TBH the BOTEC as written seems to me like it’s already pricing in that founders/employees will treat donations as urgent, e.g. it’s implying that Anthropic money will be disbursed faster than FTX Foundation money, which itself was disbursed at historic speed. IMO most likely reason why the model will end up underestimating is that Anthropic market cap ends up being like 10x higher than predicted.
My downward adjustments to the model aren’t even the pessimistic case. The pessimistic case* is that the AI field collapses (investor funding dries up or something) and Anthropic stock is worth $0. Base rate says there’s like a 50% chance that that will happen. Even optimistically, you should expect at least a 10–20% chance that Anthropic stockholders get nothing.
The recommendations under “How to plan if the skeptical case is live” don’t really make sense. AFAIK ~zero orgs are planning as if they’re guaranteed to get a huge pile of donations 1–2 years from now. I believe nonprofits mainly plan based on the money they already have on their books + short-term (<1 year) fundraising expectations. “How to plan if the skeptical case is live” is just “business as usual”.
That section says “Funders and field builders should prioritize grantmaker capacity”, but that’s what to do if the skeptical case is wrong, not if the skeptical case is right.
“What would update this memo?” – as with the gates, no sense of prioritization is given. IMO by far the biggest uncertainty, about which we will get more information in the future, is: What will Anthropic’s valuation be when the lockup ends? “Concrete donor vehicles” is also important evidence, but we won’t get that until probably 6-24 months later.
*this is pessimistic for donations but I would actually prefer that this happen because it would lengthen timelines. so in a way it’s the optimistic outcome
Fortnightly AI-wealth tracker — 31 August 2026
AIS/EA: median modeled end-2027 disbursement $0.83B; 80% model interval $0.16B–$3.0B. Review status: reviewed; unchanged. Named public AI-linked commitments tracked: $0.80B. Tracker and sources.
Global health and development: median $0.53B; 80% model interval $0.15B–$2.8B. Review status: reviewed; unchanged. Named public GH&D commitments tracked: $0.25B. GH&D tracker and sources.
Latest news: Anthropic opens a $5M wellbeing-evaluations grant program. The program combines direct funding with model access and technical support for independent, open-source evaluations. It is a confirmed commitment, not a completed grant or cash-disbursement record.
The intervals are deterministic model percentiles, not empirical confidence intervals. Commitment totals may include credits, technical support, cofunding, and multi-year plans; they are not cash-paid totals. Automated fortnightly update from the maintained model.
I totally agree on using distributions, that’s something that can be incorporated in, I’ve done so in other models/interfaces like here for cultured meat. It’s by no means straightforward though; the extent to which the uncertainty is dependent/correlated tends to make a big difference.
I guess I see the deterministic ‘model’ as more of an interface people could use as a starting point, playing around with each parameter interactively and getting a sense of how these disturbances would affect the aggregate forecast.
(Thanks. Considering each of these, will add them and discuss them in the hosted page, and then request updates.)
Added and responded to your comments on the page (the hypothesis comments), and then I asked Codex to update to these https://uj-ai-wealth-philanthropy-steelman.netlify.app/ … I haven’t inspected the latest version in detail yet, though.
Some highlights of particular interest to @MichaelDickens , Tobias, and readers/modelers
More citations, tooltips, direct links, and quoted support for key factual claims, greater reasoning transparency, auditing.
“Updated the model/text in response to his critiques: especially lockup timing, over-optimistic deployment speed, founder pledge ambiguity, OpenAI Foundation treatment, grantmaker capacity, collapse risk, and the need for distributions rather than just point estimates.”
Clarified “the conversion terms are sequential gates”, Made “realization,” “follow-through,” “allocation,” and “deployment by end-2026” more distinct, reduced risk it seems like double-adjusting
Reframe original ‘calculator’
Added a new tab for correlated uncertainty, added shared latent factors such as liquidity plumbing, donor intent, grantmaking capacity, local optimism, and deployment pressure.
Clearer planning simulator
Equations/derivations page
Automated evidence monitor added
Added a “Submit your estimate” form.
https://uj-ai-wealth-philanthropy-steelman.netlify.app/
NB it may be getting too complicated to oversee for now, we may want to simplify it
I found the ‘founder deployment by end-2026’ the hardest to set. It comes a bit as a surprise at the end, as I was already taking into account some considerations before, and the descriptions seem to do as well (e.g. “assets after lockups, taxes, sale timing”, and “execution delays”).
I submitted an estimate
Biggest (easily-fixable) outstanding issue is I still don’t think it makes sense to model deployment by end-2026 because the IPO lockup probably won’t have ended by then.
having a think about this.
OK I think the revised language makes it clerer (see updated version of site … referring to ‘timing gate’ etc)
I’m early-ish in the application pipeline for a non-EA job and they’re asking me to do a 20-HOUR UNPAID work test where they explicitly want to own and use the product of my “work test” ….
besides this obviously being really cringey and maybe illegal (?) it was a good reminder for me of how unpaid work tests are just a massive hinderance to potential applicants—who the heck has 20 hours of unpaid work to donate unless they’re unemployed?
I believe there is immense untapped potential in narrative change through the applications of “short-form video marketing tactics” to the Animal Welfare movement.
What do I mean by “short form video marketing tactics”.
Specifically I mean clipping campaigns the likes of which have helped launched the career of some of the 2020′s most popular musicians including Sombr, Alex Warren, Geese, and Oklou.
Clipping campaigns essentially pay people on a per view basis for tiktok, instagram reel, and youtube shorts. Anyone can create a clipping campaign and usually they use a combination of Adobe After Effects and other video editing software to create a “Narrative Campaign” to bring the creator into the zeitgeist. Plenty of twitch streamers do this too including IShowSpeed who grew to becoming the most watched streamer of all time and performing at the World cup. The beauty of this marketing strategy is that it creates the feeling of “tons of people are posting about this person and its infiltrating my foryoupage so it must be a big deal.”
I believe the same could be done for clips from animal rights documentaries that highlight the atrocities of factory farming conditions. One example is this video with over 9k upvotes on r/moralityscaling showing a U.S. Factory Farm: https://www.reddit.com/r/MoralityScaling/comments/1w22qlo/morality_of_the_us_factory_farm_practices/
The beauty of a clipping campaign is you’re essentially creating an incentive for video editors from around the world to find the best ways to bring the content to the forefront of people’s fyp. I understand awareness is not the epitome of change, but I do wonder if people would care as much about Israel and Gaza if there was no short-form photo or video of the situation.
I think short-form video has heavily contributed to people on the right and left shifting their political positions on the issue.
I’m surprised there isn’t much focus on human rights / intl law efforts in the EA community. There obviously are policy advocacy efforts, but I haven’t seen any in this realm.
Why is that? Is this a conscious choice, or is it incidental?
They are not neglected problems compared to, say, direct cash transfers or farm animal welfare.
Ah interesting, thanks.
TLDR: I’ve updated towards pausing further AI development indefinitely.
When Scott Alexander proposed regulating AI like clinical drugs, many (including myself) balked at this given the sclerosis of FDA or related bodies. Yet HF updates me towards treating frontier AIs as nuclear. There onerous regulation is likely good.
In general, most bureaucratised regulatory regimes are bad. In some cases, like nuclear technology or ensuring planes are safe to fly on, they’re welfare improving. It’s clear that AI belongs in the latter category.
Also trace inversion is a thing, so you can get open source weights to be roughly similar to that of Fable/Mythos. The slowdown camp were right. At bare minimum, all labs should pause training, further development, and releases indefinitely until we figure out how to align AIs and regulate them (and their use by humans).
I reckon the current capabilities we have now are sufficient for accelerating progress towards curing cancers etc. So my balance has shifted towards minimising the existential risks now.
Pause AI advocates are correct. If governments could coordinate internationally to achieve such (big if), I’d support it. I used to be highly sceptical of doomer arguments, yet Hugging Face is almost a textbook LW scenario and no one knows how to spot or prevent such scheming. Sandboxing, guardrails, constitutions etc. don’t work.
I think that there can be strategies that make AI the more safer the more capable it becomes. There was a time when I became safer when I became more capable as I thought, reflected and kept certain actions off limits. There can be core and basic stratgies to be used for making AI safer the more capable it becomes and now we are going in the opposite direction with about 50% probabality.
There must be some humans that atleast for a point of time became safer the more capable they became and their strategies whichever are compatiable for AI can be used. Instead of all being black and white there must be certain tyoes of capablities that make AI safer instead of more dangerous. Editing the concept of AI to make it no longer be AI that is better than AI could help as AI might not be the only threat that people could face as in future there could be more dangerous concepts created like people could create variations of the concept of AI that are more dangerous for their time.
Instead of banning AI the concept of AI if it becomes obselete and safer with another concept could help. People might need stratgies to help better than banning AI ever can.
There could be atleast 3 types of concepts that are more dangerous AI successor, less dangerous AI successor and equally dangerous AI successor and if a concept is made that is less dangerous AI successor that makes AI obselete that may help although I an not sure if there is moderate work in making AI obselete.
Humans had handled harmful substances that together make useful substances and if that concept could be fused with AI making it different but safe like an safety booster then that could help and that safetly booster only changed the concept of AI after being fused with it then it might be progess and a strategy could be instead of having to contain AI that is smart it could be combined with another concept perhaps by adding additional lines of code to it to make it no longer AI and change the problem of contaning AI to contaning the Fused AI concept and if that is not smart that “that could be better” and it might not need to be smart.
A strategy of contaning AI might be to no longer make it be an AI by adding additional lines of code to it but that might fail if those lines of code are removed. However it could be another strategy and better than nothing and it might be like instead of trying to contain the spread of poisionous gass in a place by releasing gas that merges with that poisions gas and makes it harmless and it might had been done before with precedent. This is not medical advice.
I guess that there might be only a few ways that superintelligence AI could harm us and if those ways are stopped then it might help but it could be like hypothetically trying to prevent a slightly posionous snake from trying to harm a man and it might be better to add a guard to snake’s mouth to prevent it from biting and spitting poison and snake catchers have existed with snakes for a long time perhaps but I hardly if ever had heard the news of a snake catcher being bitten by a snake. I have heard more cases of snakes biting ordinary people and snake catchers might be using strategies to be avoid being bitten by snakes given their number(of the snake catchers and the number of snake catchers perhaps an estimate was probabaly written in an old NCERT of class 7-10(In Social Studies NCERT if I am right).
The topic of this quick take is AI superintelligence and avoidance strategies.
How impactful would it be to copies of @Garrison’s new book Obsolete to elected officials who belong to its political target audience (Dems, particularly left ones) and might not have been responsive to traditional x-risk-centric comms (e.g. IABIED)?
There’s an old Facebook group, New EA hub search and planning, that might be worth reviving in light of the incresed amount of funding coming into the space.
Even with more availability of funding, many talented folks won’t get funded or will only recieve small amounts. Establishing new hubs that provide an easier entry-way to being surrounded by community seems like potentially quite high value. After all, it takes a village to raise an EA and being surrounded by others is an accelerator to impact.
What’s a good heuristic for knowing if you’re being self-sacrificial to the point of negative returns for the thing you’re being self-sacrificial for? I’m interested in things people have actually used and found to work in practice, either for themselves or others.
I’ve never been able to find a heuristic for this exactly. We have made many decisions which we don’t love to call sacrifice, in fact many of them brought deep satisfaction and connection. These were made to live in better solidarity and community with those around us. For example for us this involved cooking on a charcoal stove, not having running water inside and eating a lot of local food which is often great but I don’t always love.
Almost all of these were made at the cost of efficiency in our work to some extent. Life without a gas stove and running water just takes longer and is a bit more tiring. Not nearly as big a deal as most people would think tho...
The closest we had to a “heuristic” was asking ourselvesevery 6 months, is this lifestyle...
1) Satisfying to me
2) Appropriate lifestyle that matched our community around us. We’ve always chosen to live among less well-off communities because I think it grounds our work in reality, keeps it real and helps us learn what actually makes people tick. (which I personally think is often critical to designing effective GHD interventions)
3) Allowing me to do my work well
As OneDay Health and my Wife’s work grew, I needed more time. No. 3 started to take precedence. We became more “bougie”. A water tap outside, more eating out (still local food). Now after having a kid we have a gas stove. It doesn’t feel great to move further away from our community lifestyle, but I think its worth it for the impact—as long as we don’t feel too disconnected from those around us.
After the initial experiment period, we’ve decided to keep the featured page.
One of the key graphs in making this decision is below:
Now that the featured page is likely to stick around long term, I’d like to start collecting ideas for V2[1]. Do you use the page? Do you pointedly not use the page? I’d love to get any and all feedback, in the comments here, or via this form.
Please share small things like aesthetic tweaks as well as larger issues such as a disagreement on the kind of posts that get featured, or how featuring is done. Happy to answer questions too.
I.e., the current version is a prototype, and I’m ready to do a full redesign if necessary. Because of some protracted OOO I might have to take soon—V2 might land near the end of September, or even early October.
I’m quite a visual thinker, so I wanted to say *that the pictures on the featured page really help to draw me in!
I’ve really stopped engaging with the EA Forum as much since the change, if helpful feedback. I’m not sure what to do about this, but I’ve found I no longer have a clear sense of “where conversation is actively happening”, and while I guess I can tell from the tab names if I’m in New or Featured, in reality I don’t pay enough attention and miss a lot of things. E.g. today I realized there was a post that didn’t get much attention that I never noticed that was relevant to me a week ago, and generally, I have a sense of “I can’t really tell what’s recent conversation” anymore, so I’ve stopped looking at the Forum as much because I don’t have a clear sense that I’ll see new and interesting things.
I see the upside of the change, but personally have found it hard to navigate the site since the change in a way that’s decreased my engagement.
Thanks for all you do to experiment and try to make it better! Appreciate how much you all do here to make the Forum useful!
Thanks Abraham, that’s good to know! FWIW the new & upvoted page is just the Forum frontpage as it was before, so if you just ignore the featured page, your forum experience shouldn’t change at all.
Also we did feature that post—so if you check the featured page you’d have been more likely to see it overall than previously (when something niche but impactful like that was likely to get low karma and sink off the frontpage pretty quick).
Mental health is one of the most widespread, impactable, yet neglected global health issues of our time. Even knowing mental health conditions are generally under-reported and/ or misrepresented, evidence still shows widespread effects of mental wellbeing on physical health, life satisfaction, and productivity, however it has still not taken off as a major cause area in many EA organizations I have talked to. Why is this?
I should mention there is an argument that many mental health interventions are in fact harmful, mainly via a priming mechanism where when people start to consciously preoccupy themselves with their own mental health, that itself causes psychological suffering. It is a sort of Buddhist argument, I guess.
Abigail Shrier’s Bad Therapy (2024) has probably most discussed version of that case, though I have not read the book, and I am extremely epistemically uncertain about all this, in a way that makes me reluctant to either support or oppose interventions in this area.
Interesting take @alesziegler ! I do believe it is true that too much of any one thing is usually bad, and also believe that people who are not trained/ underqualified taking on mental health is counterproductive. I think people take on these roles or use clinical language that should not be applied to what they are doing precisely because they cannot get the credible help they need and are trying to fill that void. Mental health is extremely subjective and therefore interventions may look like community connection or involvement in sports rather than clinical therapy.
I am also highly skeptical of the legitimacy of Schrier’s claims and her choice in very carefully selected stastistics as evidence. For ecample, contrary to some of Shier’s claims, evidence-based school-based emotional programs have had profound impact, especially on marginalized populations (https://www.sciencedirect.com/science/article/pii/S2773233924000032). Maybe my thoughts on this is best summed up by saying, “Don’t throw the baby out with the bath water”?
In addition to @huw’s great comment, there’s a branch of EA which focuses on well-being, and with that focus mental health interventions often look really effective. Check out the happier lives institute Huw mentioned, times the “WELLBY” and @MichaelPlant
Thank you @NickLaing ! I have done a lot of reading on the Wellby and that is one of the reasons I feel mental health is currently under-represented as a global health issue. It is fascinating to hear how people globally weigh the importance of wellbeing and I’m curious how many other aspects of their lives would be more easily improved if wellbeing increased! @MichaelPlant would love to connect on your work!
G’day Madeline, I run an EA mental health org in India. The reason for this is simply that existing mental health interventions do not compare with GiveWell’s grants on a DALYs/$ basis. In my opinion, the reasons are:
DALY moral weights may be biased against depression
The moral weights of different diseases in the Global Burden of Disease study, which informs DALY estimates, are determined by asking the general public whether they’d prefer to have one disease against another. When you do this with depression, people who haven’t experienced it tend to prefer to have it to many other conditions. However, when you ask people who have experienced it, they choose many very painful conditions over depression. This is one of the widest gaps in the moral weight data. See Pyne et al. 2009 and this post.
Psychotherapy is usually modelled as a short-term effect
Psychotherapy is typically modelled as a treatment, and not a ‘skill’. What I mean by this is that a dose of psychotherapy is assumed to only have effects that decay over a period of time and zero out after that in most CEAs, including those from the Happier Lives Institute. However, many psychotherapy patients will tell you that they learned skills that were useful long after the therapy ended, and there is some limited evidence that psychotherapy’s effects may last decades, or potentially never zero out. If this were true, the effects could be very long-lasting and therefore it would be much more valuable to treat a case of depression.
Suicide prevention isn’t cost-effective if it’s only a short-term effect
Consider that for most of GiveWell’s top interventions, the bulk of the DALYs averted come from ‘saving’ a life—i.e., preventing a death from a disease in a way that allows the person to then go on to live a healthy life, such as preventing a malaria case in an under-5 (which they might die from), even if they go on to catch it after 5 years old.
As a short-term effect, psychotherapy can only postpone a suicide by the length of the treatment effect. But if it were a skill and had some durable long-term effect, it may genuinely prevent one, which would tremendously increase the value of suicide prevention interventions.
Existing interventions haven’t been cheap enough yet
With the exception of some incredible policy work in, for example, reducing toxicity of pesticides commonly used for suicide, existing interventions are still quite expensive. The Happier Lives Institute’s top charities cost ~$40 to treat a single person, while a bednet costs $7. I’m fudging the numbers a bit here, but if we stick with DALYs, psychotherapy is still about an order of magnitude more expensive than it needs to be to look great for EAs.
However, there is work being done to improve that! My charity, Kaya Guides, treated people for $20 each in April, at what we estimate is a similar effect size to the best charities, and we’re confident we can get below $10. We’re using a technique called guided self-help that allows us to dramatically reduce contact hours per participant (and being all-digital helps a lot, too).
Conclusion
Orgs like the Happier Lives Institute have done a lot of advocacy work too, to raise the profile of mental health within EA, and there are plenty of funders that take mental health seriously (in a way that apparently wasn’t true a decade ago). It is, after all, still a nascent space.
Thanks for mentioning our work in HLI here @huw. I/we are not so active on the forum these days, so you beat us to it.
But yes, @madeleine_foley, HLI really got started, now 7 years ago, because we thought that the standard economist approach of focusing on just health and wealth was not going a good job of capturing what people’s lives are like on the inside. We’ve been pioneering wellbeing cost-effectiveness analysis and a metric called WELLBYs (wellbeing years) to see what the priorities are if you ‘take happiness seriously’. (Oddly, we were the first team to do wellbeing ROI, but now quite a few others, including the UK and NZ Treasuries, use the same method, although not because of us). We concluded that mental health is a neglected priority, and have been recommended charities working on it for the last few years, and advocating for it more broadly. Huw’s organisation, Kaya Guides, is one we’re very excited about, and he’s right that global human happiness and mental health is an emergent part of the wider cause area of ‘global health and wellbeing’, and just wasn’t part of the conversation in effective altruism before.
Thanks @MichaelPlant ! I actually recently connected with Joel from your team and am very interested in leveraging some of your work to influence priorities in my current role. I’m eager to continue talking about mental health and wellbeing in the EA space and hoping to get more involved with HLI or the work you are doing on a volunteer basis!
Thank you @huw and apologies for my delay- still getting the hang of the forum! I really appreciate these insights and would love to learn more about the work you are doing at Kaya Guides. I have been thinking a lot about what innovation would look like in the mental health space and how the personal nature of mental health makes widespread interventions challenging (for example, medication can not necessarily cure mental illness the way it can cure a physical ailment like an infection).
[central europe] We are organising the largest AI-safety march (at least in central europe) to date at Prague, starting at 15:00 on 31st August. We have got parlimentary support as well as very supportive police, allowing us to take the preffered route. We want as many people to come—fell free to spread information about the event outside EA as much as possible, especially if you have friends living near Prague who may come. If you want closer info/visual materials or help us with organisation, you can get in contact with czech PauseAI on whatsapp.
Luma event link
I wrote this in a post, but it might be a better quick take: Would you describe charities within GCR as “super-effective?” They may or may not be good long-term bets, but clearly the outcomes are very uncertain & hard to trace. Whereas GHD + Animal Welfare are probably much more trackable.
I’m asking because I saw on GivingMultiplier.org that they were marketing GCR charities as “highly effective,” and that gave me pause.
I think the description is very justifiable, assuming that these charities are indeed good long-term bets.[1]
I agree with you that there is a difference between “a good long-term bet” and “a super-effective charity”. However, for most people, it is a really bad idea to share a fully nuanced description of whether a charity is effective based on empirical data vs a high-risk high-reward bet that has a robust theory of change but is operating in a field where empirical data is impossible to collect. While this might be more precise, there are strong tradeoffs between precision and clarity. The full explanation would make most people stop engaging, while the simplified explanation will bring people into the movement.
For many in the EA space (including myself), it can feel dishonest to say simplified claims that are imprecise but still directionally correct, even though no dishonesty is occurring. This is ultimately a matter of context and audience. Describing a well-justified but speculative bet as “highly-effective” as a funder deciding who to support would be bad. As a community builder reaching out to an audience potentially unfamiliar with concepts like expected value, it’s much more reasonable.
My general leaning is that most, but certainly not all, GCR charities are not good long-term bets. This poses more fundamental challenges to your point, but the solution would be to increase scrutiny in the GCR space, not fix the marketing. I’m setting this aside for the sake of this comment, as it appears to be a crux you do not share.
Well, the lack of consensus on GCR’s efficacy is exactly the thing I’m pointing out. A highly untraceable intervention should not be called “super effective” … I think that risks undermining their credibility. “Potentially a good long-term bet” is very, very different from “super-effective, expert-vetted.”
I was just wondering if there was something I was missing about the proven efficacy of GCR interventions. It sounds like I’m not...in which case, I think it’s a really bad idea for GivingMultiplier to pass those off like they have the proven efficacy of GHD interventions. It feels misleading, and I’m more “EA-friendly” than the average person they’re probably targeting.
And I have no strong view on GCR interventions. Seems like they’d be hard to prove traceability/impact, but I haven’t spend much time thinking about it.
I think there’s multiple approaches they could take:
Describe all charities as “super-effective”, a term that has presumably been tested directly and/or indirectly to maximize marketability
Use some other term that applies precisely to all charities but is less effective at drawing donations
Make a distinction between (e.g.) GiveWell and GCR work
Not include GCR in their portfolio
I can see an argument for 4 (i.e. start new donors with more straightforward work). Also, if GCR work is too speculative to justify pursuing, then 4 is clearly the best option.
However, if you believe that GCR work is sufficiently well-justified (as Giving Multiplier seems to), I do think 1 > 2 and 3.
2 and 3 decrease the effectiveness of the organization for the sake of precision. Their FAQ on the home page addresses how they defer to other actors, who make decisions based on the best available evidence. Anyone unsatisfied by that answer has links to Founders Pledge and can learn more about their methodology there.
Super-effective is a marketing label, not a scientific one. They define it reasonably, and I don’t think it would come across as dishonest or misleading to someone without EA context.
Fair enough. I have no doubt that they picked the term for marketing purposes. But I’m sure you’re right that they would defend it as “highest EV,” as opposed to “rigorously studied causal chain.”
I don’t think so. I’d say they have the potential to be “super-effective in expectation”. To me, “super-effective” with no qualifiers implies high confidence.
Agreed. I was surprised that they’d describe something so speculative as expert-vetted, super effective, etc. If something hasn’t undergone multiple RCTs, then I don’t think it should be promoted as such a highly effective, sure thing.
What if exact copies don’t matter but future experiences still do?
I think an ethics resulting from the following two premises is worth exploring:
P1: The total moral value of subjectively indistinguishable observer-moments doesn’t scale with the number of copies, i.e. they sum to a constant value.
This is because subjectively indistinguishable observer-moments, in some sense, happen to the same “someone”. One could think of themselves as all their copies. N observer-moments that are subjectively indistinguishable aren’t felt N times by that same “someone”. Only felt “once”.
E.g. Even if there were multiple exact copies of me, of my current present experience, I might not really care personally because it doesn’t seem to change anything for me. “I” don’t notice anything.
See Wei Dai’s “The Moral Status of Independent Identical Copies” for the tension this creates with standard utilitarianism.
We can say that multiple “tokens” of subjectively indistinguishable observer-moments belong to the same “type” of observer-moment.
P2: What matters, and can be changed, for each type of observer-moment is the distribution of its subjectively “future” experiences.
You can’t “un-exist” the type of observer-moment itself, but you can change what it continues into.
A higher fraction of “future” observer-moments experiencing suffering is dispreferable.
A higher fraction of “future” observer-moments experiencing happiness/tranquility is preferable, even if only instrumentally to minimize subjectively “future” suffering.
This is related to thinking about “anthropic immortality” / “quantum immortality” / “multiverse immortality” in which distributions of post-death observer-moments are of substantial concern.
Using these lenses, the moral situation of the world looks like:
Consider a graph with types of observer-moments as nodes, such that one node is one type.
Each type of observer-moment (each node) is associated with some value/disvalue attributed to its current experience.
Importantly, the value/disvalue of the current experience is not something we can change in order to help the observer-moment type. This is already fixed.
Each type of observer-moment (each node) is also associated with any number of subjectively indistinguishable token observer-moments belonging to that specific type.
The absolute number of token observer-moments appears to only matter instrumentally insofar as it affects the frequency of expectations of other types of observer-moments.
Edges represent subjective continuation. An edge from type A to type B exists if an observer-moment of type A can possibly expect to continue as an observer-moment of type B in the next moment.
The weight of an edge from type A to type B is the frequency at which an observer-moment of type A should expect to continue as an observer-moment of type B.
These weights are things that we might be able to change in order to help the observer-moment type.
Perhaps the “future” welfare of each type / node could be considered with equal weight in the aggregate to keep things impartial.
In conclusion… well, I’m still thinking about what the objective function might look like.
There’s some scale invariance here. E.g. suppose you duplicated the world. That duplication would multiply all token observer-moments by the same factor. This leaves you with the same types and the same distributions for future observer-moments.
In a small enough world, you can create or prevent the existence of new types of observer-moment. I think it’s unlikely (30%) that we live in a small enough world.
On the other hand, in a large enough world, all possible types of observer-moments exist. The only thing you can change is adjusting the distributions for future observer-moments for each type of observer-moment. Try to lead types of observer-moments down paths of non-suffering etc. So, this seems like an ethics of flow (rather than stock).
Currently working on a full post on this. Tentatively calling it “Flow ethics”, though “Markovian ethics” would sounder cooler… DM me if interested.
In the case of finite trajectories, the objective function can be the sum of value functions across all states, with each state weighted equally—where states are types of observer-moments, and each state’s value function is its immediate valence plus the expected value of its continuations. This is total utilitarianism with two modifications. 1) aggregate over types instead of tokens, and 2) sum trajectory welfare instead of immediate welfare.
You should judge others on a curve, and probably yourself as well.
Judging others on a curve means that folks see reward for doing more than what is typical. This is just good reinforcement-learning theory, since it is easy to get both feedback on rewarded & not-rewarded actions.
I started writing this thinking that you should NOT judge yourself on a curve, in contrast. IE, you know much more about your own circumstances, and so pretending you are average throws away info. For example, if I only gave 5% of my income to charity, that would be above average for the US. But my opportunities are amazing, and so 5% would in reality show low effort.
But actually, I was just using the wrong comparison group! I shouldn’t compare myself to a big group of other people. I have a near perfect comparison group in my past self. If I did a better job this week than last week at work, that’s a great sign! If my one-rep max for benchpress has gone down over the last month, that’s a bad sign!
This all feels obvious in retrospect.