Thinking back to my first big (well, big for me) donation and an unusual series of thoughts I had. Sharing in case anyone has experienced the same. --
I’ve practiced frugality to a significant extent, largely because I want to donate/salary sacrifice so much of my income to effective charities. As a result, whenever I’m spending a significant amount of money on anything, alarm bells go off.
I was surprised that alarm bells went off when donating. I had thought through where I donated extensively, and charity was the reason why I wanted to save money in the first place. But I still felt stressed because I was “spending money.”
This is such a clear example of missing the forest for the trees. I think we need to be careful about instrumental values/rules and making sure they don’t become absolute limitations that decrease overall impact.
Is it a sticker shock reaction developed from years of reflexively minimizing outflows? Donating more frequently in smaller amounts might help. Feels less irreversible.
Without getting into the specifics, this donation opportunity only made sense above a certain level.
I did donate, and I’m glad I did, but I think it’s valuable to note the inconsistencies of human psychology that appear even in communities that pride themselves on rationality.
There’s lots of talk about how more EA-aligned funding will attract grifters. The Robin Hanson in me wondered what expensive or hard-to-fake signals exist for (human) alignment.
The first ideas that came to mind:
donating a kidney
vegan/vegetarian
non-profit work
donating a lot to effective charities
community engagement
I don’t know whether folks should weight these highly in funding decisions. Probably they all collapse under sufficient pressure.
Feels like most of these except the kidney are fairly easily faked or not that relevant. Community engagement and working for your own non-profit are about the first thing anyone seeking to grift will do, as well as people who are sincere (including as others have pointed out, the people who are sincerely wrong)
And Sam Bankman Fried was a vegan who apparently started off with a completely sincere desire to work for animal nonprofits. Treating these as signals and the people complaining about his behaviour and some of the things he said as noise was a bad move.
I personally think people worry disproportionately about grift on the EA Forum, while implementation challenges and other considerations seem to be bigger risks in practice. (e.g. I don’t think the Wytham project or Dispensers for Safe Water would have been more successful if they had been assigned to more vegan or more frugal people)
The kind of grifting that most concerns me within EA would result in a lot of fudging and inflation around reported impact in order to gain money & status. I think the hardest thing is fake, and the main thing we care about protecting, is real, verifiable impact. If there is an influx of money into EA, I would want the first thing to be done is to undertake a lot of new, third-party evaluation with depth. This is not happening today and I have not seen great plans to do so right now.
In my view, this one of EA’s biggest problem. Even GiveWell rarely bothers to verify numbers if it’s boring or manual to do so (e.g. number of people vaccinated, number of bednets delivered). Even when past charity evaluations could be calibrated against third party data sources, nobody bothers. We’re allocating billions using tools we’ve never bother to calibrate.
Without anyone establishing anything resembling ground truth, there is no accurate feedback signal for leaders, grantmakers, donors. When the only signal is what LOOKS impactful, and that’s how all money and status is awarded, that’s what everyone will clamour after. We’re going to learn the wrong lessons, fund the wrong projects, hire the wrong people, and then never find out.
Tons of money. Zero verification. Every fraud’s wet dream come true.
Frankly, if you can’t measure your work in a that an independent person could evaluate, I would be very skeptical that you’re having any impact at all.
One point where I disagree is that I’m more concerned about well-intentioned people genuinely believing their project is extremely valuable, but who lack clear evidence of impact and end up not achieving much. I’m less worried about people deliberately lying/exaggerating to raise money for things they do not themselves believe are impactful.
But of course it’s hard to say from the outside, you could always argue that people are secretly doing whatever they’re doing just for the money & status, and I’ve seen people making such claims even when there’s tons of evidence to the contrary.
We have bought 1500 copies of IABIED in Russian; over the past few days, sent emails to 1700 winners of olympiads who have previously received HPMOR from us offering the new book; and already shipped 173 copies to them (+ sent 43 e-books)
“If you believe we should be burning witches at the stake, you’re a full crank. If you believe we should be segregated by race, you’re somehow a former crank. If you believe we should be factory-farmed by species, you’re a future crank.
“Some cranks are harder to spot. When the winner of one game says they were lucky, you can’t tell if they’re well-calibrated or a performative crank. When the loser says no regrets, you can’t tell whether they’re well-adjusted or delusional. (If the loser says they’re the winner, that’s a different kettle of fish.)”
From The 1001, the EA sports blog. If you like what you find, subscribe for free to get posts directly in your inbox. No platform, no slop, no spam.
Suggestion for the forum: prompt anyone who downvotes a (new) post to leave a comment
If you look at the newest posts on any given day, you’ll see a number with 0 or 1 karma. If my understanding is correct, this means they have been downvoted. Yet a lot of these posts clearly involved a lot of work and are well intended. I imagine that to put so much effort into a post only to be immediately downvoted, and possibly have your post disappear from the front page and be read by very few people as a result, must be quite disheartening.
Of course many of these posts have flaws. But it would be good then for anyone who downvotes it—or even disagree-votes—to say why and provide feedback. I think requiring this would be too much, but if downvoting triggered a comment-box to pop up, that might help to encourage such feedback to be given. I’m imagining this working the same way voting on a poll prompts you to also add a comment. And if doing this across the board seems too annoying, it could be made to apply only to posts with, idk, <10 karma?
If you look at the newest posts on any given day, you’ll see a number with 0 or 1 karma. If my understanding is correct, this means they have been downvoted.
I don’t think this is correct, you can mouseover the karma score of a post to see the number of votes. E.g. here it has 1 karma from 1 vote (the automatic self-vote that is given when you post) and no downvotes
I think these posts are either not being shown on the frontpage to most people, or don’t get voted on before disappearing
Oh! I had the idea that new posts start with 3 karma, but then I was wrong.
In that case, I suppose downvoting is less common than I thought. I did still see one or two new posts with 0 or −1 karma, so my argument would still stand for them. But this is clearly less of an issue than I thought.
If there’s so much EA-ish money coming, how does one “put out buckets” to collect it?
I am very concerned about relative cause prioritization among the alleged wave of philanthropy. (Personally, I expect it to me smaller than many do.) How does one shape this? For example, I really want efforts of the shape “raise the salience of AI pause as a political issue among the US public, particularly the religious mainstream.” How do I influence the amount of funding that goes to that, if there’s a big anticipated wave?
I’m exploring the possibility of starting an NGO that sets up partnerships between local NGOs and sports teams to raffle experiences for people who donate. I know other NGOs/companies have done things like this. Do you know anyone you recommend I contact to start getting feedback from people in the industry?
I think a lot of ppl studying empirical persuasion results of AI tend to overestimate the ecological validity and generalizability of those results. This isn’t to say that studying empirical persuasion results is useless, just a) we should move towards better study designs in the future, and b) when looking at empirical persuasion studies you need to think carefully about not just the methodology and results but also what’s reasonable to generalize and not.
I think a lot of people understand in the abstract that “ecological validity is important” but don’t appreciate how difficult this is for persuasion results.
it’s actually just very hard to design a study that captures plausible real-world interactions, for practical, procedurally ethical, and infohazard-y reasons.
It’s very easy for empirical results to understate the influence of existing AI persuasion (see below)
It’s also very easy for empirical results to overstate the influence of existing AI persuasion (see below)
Our studies give us very limited ability to predict the timing of the superpersuasive regime.
There’s no clean analytic tool that allows us to draw lines on graphs to nearly the level of the METR time horizon graph
despite its many problems, I think the METR graph is much better at capturing something deeply true about software automation, compared to the studies we have at capturing persuasion ability or automation.
tbh I’m not sure (and currently lean against) the ability of the existing study methodologies can even allow us to reliably retrodict superhuman persuasion
as in, I don’t think the existing tools can tell us that the AIs are reliably superhuman even after it occurs.
This seems concerning!
Furthermore, the AI labs aren’t optimizing as much for AI persuasion as they are for other capabilities like coding, so we should expect some latent un-optimized for capabilities. It means that we shouldn’t be shocked if sudden leaps happen.
Our studies give us very limited ability to predict the shape of future superhuman capabilities.
For example, in Hackenburg (2026), when AIs outpersuade expert human debaters in single-session debates, they do so by throwing a ton of facts at you. Do we expect this to be the capability profile of future true superhuman persuaders?
Obviously not imo. Just seems really implausible that this is one of the empirical results that’d generalize well to the superhuman regime.
The field is already aware of some of these issues (see Appendix A), however in practice (in conversations and in the study methodology) I think many empirical researchers underestimate how big a deal these problems pose. I also think other people (eg researchers in other fields, or non-researchers) citing the studies are even less careful, and are even worse at understanding the ecological validity and generalization issues.
“Research on impacts lags behind increased adoption of increasingly powerful AI. The mechanisms of epistemic decline discussed in Section 2 hinge on AI systems’ ability to persuade and to automate (Hackenburg et al., 2025; Durmus et al., 2024). Both of these are increasing (Anthropic, April 2025; Park et al., 2024a).
Ecosystem-level feedback dynamics are inherently difficult for individual studies to capture. The recursive coupling of human beliefs, AI outputs, and training data described in Section 2.3 operates across platforms, training pipelines, and populations simultaneously. It is difficult for single studies, which are often limited in the number of platforms/models/populations they can test, to observe dynamics that emerge from the interaction of all three.
Measured evidence is usually a lagging indicator of complex real-world phenomena. Realworld situations are difficult to specify and translate into neat taxonomies and variables for modeling that are currently required for much of high-quality research. For example, it is difficult to assess how much personalization boosts persuasiveness. In research settings, persuasiveness can only be measured by proxy. In doing so, some studies find no effect (Hackenburg et al., 2025; Argyle et al., 2025), others find small effects (Kelley and Riedl, 2026; Matz et al., 2024a), and others find strong effects (Salvi et al., 2025a). This disagreement may itself reflect the difficulty of measuring effects that depend on sustained, naturalistic interaction rather than one-shot exposure.
The scientific endeavor also risks overstating these effects. Laboratory settings are not always representative of real-world environments; experimental designs may act as high-fidelity amplifiers that maximize pressure to achieve statistical rigor. In the lab, moving a participant’s opinion on a policy issue by a few percentage points may be a controlled success with effect size, but in the actual chaos of a political cycle even massive advertising spending can struggle to move the needle at all (Coppock et al., 2020).”
Looking for a way to skill-build as a student or young graduate? Consider charity trusteeship!
EAs are often good at maths/technology and less good at leadership and governance skills. This is an excellent match for (non-EA) trustee governance boards, that often have committed older people with good leadership and governance but who struggle with use of tech and maths. There are a lot of trustee vacancies around—highly absorbable—and it’s an easy way to get started on tackling real issues with charities big and small!
You can help your charity act more effectively in the world using EA principles and frameworks about effective use of charitable resources, and learn experientially about practical issues than arise with those frameworks when they are implemented. Organisations like the Young Trustees Movement (UK) can mentor and support you, and the other trustees will probably have a lot of sage advice. And while trusteeship is unpaid, it also doesn’t *cost* money either, so it’s a fairly cheap way to do “EA as a hobby”.
The EA community needs more people aged 30+ with genuine commitment and leadership skills, and should encourage people looking for principles-first EA to consider this as a life path in their 20s.
I have made a semi-functional rich person impact ranker. It requires votes (and optionally, longform analyses from humans or AIs). If you want to help flesh it out, dm me on here or on twitter or if you have my details, anywhere I am contactable and I’ll send it to you.
A name I came up with for a pattern that comes up often enough to deserve one. Here’s a particular instance:
Suppose you’re a longtermist who thinks the vast majority of the (net positive) moral weight is in the far future, that AI existential risk is above 1%, and that it’s reasonably tractable to reduce. Some people who hold all three premises still say things like: “AI safety has maybe two orders of magnitude more resources than animal welfare, so the marginal dollar is better spent on animal welfare.”
This is the fallacy of similar magnitudes: in this case, the application is that neglectedness comparisons only matter when the causes are within a couple orders of magnitude of each other in stakes. Under the stated premises, they aren’t: the a priori difference in expected value is way larger. Two orders of magnitude in resources doesn’t move the needle even a bit.
Note: this AW response could be ok if you have certain views about risk aversion or worldview diversification, but that’s besides the point.
Hot take: When Anthropic IPO money starts flowing into AI safety, the ecosystem should consider operating more outside charitable structures. Charities get an indirect public subsidy via tax-deductible donations, but the trade-off is greater compliance burden and cost, restrictions on spending, and a greater reliance on public goodwill. As the world gets weirder from AI-driven job losses and political shifts, taxpayers may be less happy to see their foregone tax dollars going to China dialogues or high salaries for technical AI safety researchers. (Not legal or financial advice.)
Mental health is one of the most widespread, impactable, yet neglected global health issues of our time. Even knowing mental health conditions are generally under-reported and/ or misrepresented, evidence still shows widespread effects of mental wellbeing on physical health, life satisfaction, and productivity, however it has still not taken off as a major cause area in many EA organizations I have talked to. Why is this?
In addition to @huw’s great comment, there’s a branch of EA which focuses on well-being, and with that focus mental health interventions often look really effective. Check out the happier lives institute Huw mentioned, times the “WELLBY” and @MichaelPlant
G’day Madeline, I run an EA mental health org in India. The reason for this is simply that existing mental health interventions do not compare with GiveWell’s grants on a DALYs/$ basis. In my opinion, the reasons are:
DALY moral weights may be biased against depression
The moral weights of different diseases in the Global Burden of Disease study, which informs DALY estimates, are determined by asking the general public whether they’d prefer to have one disease against another. When you do this with depression, people who haven’t experienced it tend to prefer to have it to many other conditions. However, when you ask people who have experienced it, they choose many very painful conditions over depression. This is one of the widest gaps in the moral weight data. See Pyne et al. 2009 and this post.
Psychotherapy is usually modelled as a short-term effect
Psychotherapy is typically modelled as a treatment, and not a ‘skill’. What I mean by this is that a dose of psychotherapy is assumed to only have effects that decay over a period of time and zero out after that in most CEAs, including those from the Happier Lives Institute. However, many psychotherapy patients will tell you that they learned skills that were useful long after the therapy ended, and there is some limited evidence that psychotherapy’s effects may last decades, or potentially never zero out. If this were true, the effects could be very long-lasting and therefore it would be much more valuable to treat a case of depression.
Suicide prevention isn’t cost-effective if it’s only a short-term effect
Consider that for most of GiveWell’s top interventions, the bulk of the DALYs averted come from ‘saving’ a life—i.e., preventing a death from a disease in a way that allows the person to then go on to live a healthy life, such as preventing a malaria case in an under-5 (which they might die from), even if they go on to catch it after 5 years old.
As a short-term effect, psychotherapy can only postpone a suicide by the length of the treatment effect. But if it were a skill and had some durable long-term effect, it may genuinely prevent one, which would tremendously increase the value of suicide prevention interventions.
Existing interventions haven’t been cheap enough yet
With the exception of some incredible policy work in, for example, reducing toxicity of pesticides commonly used for suicide, existing interventions are still quite expensive. The Happier Lives Institute’s top charities cost ~$40 to treat a single person, while a bednet costs $7. I’m fudging the numbers a bit here, but if we stick with DALYs, psychotherapy is still about an order of magnitude more expensive than it needs to be to look great for EAs.
However, there is work being done to improve that! My charity, Kaya Guides, treated people for $20 each in April, at what we estimate is a similar effect size to the best charities, and we’re confident we can get below $10. We’re using a technique called guided self-help that allows us to dramatically reduce contact hours per participant (and being all-digital helps a lot, too).
Conclusion
Orgs like the Happier Lives Institute have done a lot of advocacy work too, to raise the profile of mental health within EA, and there are plenty of funders that take mental health seriously (in a way that apparently wasn’t true a decade ago). It is, after all, still a nascent space.
In 2023-2024, it seemed a very strong bias from the EA community and the core AI safety funders that Anthropic not to be argued against or funding diverted to anything that was critical of them (directly or indirectly).
I wonder how the community and funders have updated their beliefs and behaviours?
As one datapoint, this is from 2024, and I saw no subsequent evidence from core AI safety funders (including CG) that it influenced their willingness to fund me (I’ve also been regularly critical of Anthropic, and the other companies, on twitter and elsewhere throughout this time). https://www.lcfi.ac.uk/news-events/blog/post/reflections-on-machines-of-loving-grace
Examples of fairly critical remarks about Anthropic’s actions. All more recent than your timeframe, because twitter’s search function is a pain and because my memory for twitter etc discussion is less good. https://x.com/S_OhEigeartaigh/status/2029475839654388069 (re: gullible bunch memo)
(I am followed by some prominent US policy people, including the outgoing white house AI adviser, so if there were a bias against people who criticise Anthropic, I would expect to be punished pretty harshly)
And while ongoing commentary on their safety/policy actions might not count for your criteria, this paper has been influential enough with policymakers that it probably counts (and goes against aspects of Anthropic’s stance on China that underlies a lot of their positions). https://papers.ssrn.com/abstract=5278644
None of these are from “core AI safety funders”. Views in the EA community were all across the map, but I share James’s impression that big funders (especially CG) were too reluctant to fund anything that’s critical of AI companies.
• Part-time role ($80 AUD/hr, ~13 hours/week). • Lead Saturday sessions and provide remote support. • Must have strong ML skills and completed most or all of the ARENA curriculum. • Applications close 8 August 2026, reviewing candidates on rolling basis. • Check out the role description and apply.
I hastily vibecoded a ~live-updating EA Forum (posts + comments; I think users are static) database + semantic search because software is ~free now. I haven’t rigorously checked anything over because it was/is just a random side project so use at your own risk and discretion. Enjoy!
I’m recording a podcast (TWCBB) with @Lauren Gilbert [edit: we had to reschedule so I can still collect questions]. If there’s anything you’d like us to ask her, comment here.
Thinking back to my first big (well, big for me) donation and an unusual series of thoughts I had. Sharing in case anyone has experienced the same.
--
I’ve practiced frugality to a significant extent, largely because I want to donate/salary sacrifice so much of my income to effective charities. As a result, whenever I’m spending a significant amount of money on anything, alarm bells go off.
I was surprised that alarm bells went off when donating. I had thought through where I donated extensively, and charity was the reason why I wanted to save money in the first place. But I still felt stressed because I was “spending money.”
This is such a clear example of missing the forest for the trees. I think we need to be careful about instrumental values/rules and making sure they don’t become absolute limitations that decrease overall impact.
Is it a sticker shock reaction developed from years of reflexively minimizing outflows? Donating more frequently in smaller amounts might help. Feels less irreversible.
Without getting into the specifics, this donation opportunity only made sense above a certain level.
I did donate, and I’m glad I did, but I think it’s valuable to note the inconsistencies of human psychology that appear even in communities that pride themselves on rationality.
Me and @Franare interviewing @GraceAdams🔸 tomorrow evening. Is there anything you’d like us to ask her?
There’s lots of talk about how more EA-aligned funding will attract grifters. The Robin Hanson in me wondered what expensive or hard-to-fake signals exist for (human) alignment.
The first ideas that came to mind:
donating a kidney
vegan/vegetarian
non-profit work
donating a lot to effective charities
community engagement
I don’t know whether folks should weight these highly in funding decisions. Probably they all collapse under sufficient pressure.
I agree with all of this. One thing missing is just joining AIS/EA before it was cool/got much more money.
Feels like most of these except the kidney are fairly easily faked or not that relevant. Community engagement and working for your own non-profit are about the first thing anyone seeking to grift will do, as well as people who are sincere (including as others have pointed out, the people who are sincerely wrong)
And Sam Bankman Fried was a vegan who apparently started off with a completely sincere desire to work for animal nonprofits. Treating these as signals and the people complaining about his behaviour and some of the things he said as noise was a bad move.
Here’s another list from Caroline Ellison during the FTX times on valuable signals: https://forum.effectivealtruism.org/posts/M44i4CiMECP5Xoorz/demandingness-and-time-money-tradeoffs-are-orthogonal (which I’m not sure I agree with, but I found interesting)
I personally think people worry disproportionately about grift on the EA Forum, while implementation challenges and other considerations seem to be bigger risks in practice. (e.g. I don’t think the Wytham project or Dispensers for Safe Water would have been more successful if they had been assigned to more vegan or more frugal people)
The kind of grifting that most concerns me within EA would result in a lot of fudging and inflation around reported impact in order to gain money & status. I think the hardest thing is fake, and the main thing we care about protecting, is real, verifiable impact. If there is an influx of money into EA, I would want the first thing to be done is to undertake a lot of new, third-party evaluation with depth. This is not happening today and I have not seen great plans to do so right now.
In my view, this one of EA’s biggest problem. Even GiveWell rarely bothers to verify numbers if it’s boring or manual to do so (e.g. number of people vaccinated, number of bednets delivered). Even when past charity evaluations could be calibrated against third party data sources, nobody bothers. We’re allocating billions using tools we’ve never bother to calibrate.
Without anyone establishing anything resembling ground truth, there is no accurate feedback signal for leaders, grantmakers, donors. When the only signal is what LOOKS impactful, and that’s how all money and status is awarded, that’s what everyone will clamour after. We’re going to learn the wrong lessons, fund the wrong projects, hire the wrong people, and then never find out.
Tons of money. Zero verification. Every fraud’s wet dream come true.
Sure, if you can directly measure what you actually want (positive impact), that is of course best :)
Frankly, if you can’t measure your work in a that an independent person could evaluate, I would be very skeptical that you’re having any impact at all.
See also this comment which I also agree with.
One point where I disagree is that I’m more concerned about well-intentioned people genuinely believing their project is extremely valuable, but who lack clear evidence of impact and end up not achieving much. I’m less worried about people deliberately lying/exaggerating to raise money for things they do not themselves believe are impactful.
But of course it’s hard to say from the outside, you could always argue that people are secretly doing whatever they’re doing just for the money & status, and I’ve seen people making such claims even when there’s tons of evidence to the contrary.
We have bought 1500 copies of IABIED in Russian; over the past few days, sent emails to 1700 winners of olympiads who have previously received HPMOR from us offering the new book; and already shipped 173 copies to them (+ sent 43 e-books)
Is Thomas Tuchel deluded or what?
“If you believe we should be burning witches at the stake, you’re a full crank. If you believe we should be segregated by race, you’re somehow a former crank. If you believe we should be factory-farmed by species, you’re a future crank.
“Some cranks are harder to spot. When the winner of one game says they were lucky, you can’t tell if they’re well-calibrated or a performative crank. When the loser says no regrets, you can’t tell whether they’re well-adjusted or delusional. (If the loser says they’re the winner, that’s a different kettle of fish.)”
From The 1001, the EA sports blog. If you like what you find, subscribe for free to get posts directly in your inbox. No platform, no slop, no spam.
Suggestion for the forum: prompt anyone who downvotes a (new) post to leave a comment
If you look at the newest posts on any given day, you’ll see a number with 0 or 1 karma. If my understanding is correct, this means they have been downvoted. Yet a lot of these posts clearly involved a lot of work and are well intended. I imagine that to put so much effort into a post only to be immediately downvoted, and possibly have your post disappear from the front page and be read by very few people as a result, must be quite disheartening.
Of course many of these posts have flaws. But it would be good then for anyone who downvotes it—or even disagree-votes—to say why and provide feedback. I think requiring this would be too much, but if downvoting triggered a comment-box to pop up, that might help to encourage such feedback to be given. I’m imagining this working the same way voting on a poll prompts you to also add a comment. And if doing this across the board seems too annoying, it could be made to apply only to posts with, idk, <10 karma?
I don’t think this is correct, you can mouseover the karma score of a post to see the number of votes. E.g. here it has 1 karma from 1 vote (the automatic self-vote that is given when you post) and no downvotes
I think these posts are either not being shown on the frontpage to most people, or don’t get voted on before disappearing
Oh! I had the idea that new posts start with 3 karma, but then I was wrong.
In that case, I suppose downvoting is less common than I thought. I did still see one or two new posts with 0 or −1 karma, so my argument would still stand for them. But this is clearly less of an issue than I thought.
They start with a self-strong-upvote, which depends on the user karma https://github.com/centre-for-effective-altruism/ea-forum/blob/27d24f53b7dce59b770040e6ac9ff05fcb9c4b16/src/lib/votes/voteHelpers.ts#L71-L91
If there’s so much EA-ish money coming, how does one “put out buckets” to collect it?
I am very concerned about relative cause prioritization among the alleged wave of philanthropy. (Personally, I expect it to me smaller than many do.) How does one shape this? For example, I really want efforts of the shape “raise the salience of AI pause as a political issue among the US public, particularly the religious mainstream.” How do I influence the amount of funding that goes to that, if there’s a big anticipated wave?
Write convincingly about it then get it featured prominently somewhere the donors, or people they trust, already read.
Such as, for example, the Effective Altruism Forum :)
Blood Donation & Sports Teams
I’m exploring the possibility of starting an NGO that sets up partnerships between local NGOs and sports teams to raffle experiences for people who donate. I know other NGOs/companies have done things like this.
Do you know anyone you recommend I contact to start getting feedback from people in the industry?
I think a lot of ppl studying empirical persuasion results of AI tend to overestimate the ecological validity and generalizability of those results. This isn’t to say that studying empirical persuasion results is useless, just a) we should move towards better study designs in the future, and b) when looking at empirical persuasion studies you need to think carefully about not just the methodology and results but also what’s reasonable to generalize and not.
I think a lot of people understand in the abstract that “ecological validity is important” but don’t appreciate how difficult this is for persuasion results.
it’s actually just very hard to design a study that captures plausible real-world interactions, for practical, procedurally ethical, and infohazard-y reasons.
It’s very easy for empirical results to understate the influence of existing AI persuasion (see below)
It’s also very easy for empirical results to overstate the influence of existing AI persuasion (see below)
Our studies give us very limited ability to predict the timing of the superpersuasive regime.
There’s no clean analytic tool that allows us to draw lines on graphs to nearly the level of the METR time horizon graph
despite its many problems, I think the METR graph is much better at capturing something deeply true about software automation, compared to the studies we have at capturing persuasion ability or automation.
tbh I’m not sure (and currently lean against) the ability of the existing study methodologies can even allow us to reliably retrodict superhuman persuasion
as in, I don’t think the existing tools can tell us that the AIs are reliably superhuman even after it occurs.
This seems concerning!
Furthermore, the AI labs aren’t optimizing as much for AI persuasion as they are for other capabilities like coding, so we should expect some latent un-optimized for capabilities. It means that we shouldn’t be shocked if sudden leaps happen.
Our studies give us very limited ability to predict the shape of future superhuman capabilities.
For example, in Hackenburg (2026), when AIs outpersuade expert human debaters in single-session debates, they do so by throwing a ton of facts at you. Do we expect this to be the capability profile of future true superhuman persuaders?
Obviously not imo. Just seems really implausible that this is one of the empirical results that’d generalize well to the superhuman regime.
The field is already aware of some of these issues (see Appendix A), however in practice (in conversations and in the study methodology) I think many empirical researchers underestimate how big a deal these problems pose. I also think other people (eg researchers in other fields, or non-researchers) citing the studies are even less careful, and are even worse at understanding the ecological validity and generalization issues.
Appendix A (notes from Yang et. al 2026)
“Research on impacts lags behind increased adoption of increasingly powerful AI. The mechanisms of epistemic decline discussed in Section 2 hinge on AI systems’ ability to persuade and to automate (Hackenburg et al., 2025; Durmus et al., 2024). Both of these are increasing (Anthropic, April 2025; Park et al., 2024a).
Ecosystem-level feedback dynamics are inherently difficult for individual studies to capture. The recursive coupling of human beliefs, AI outputs, and training data described in Section 2.3 operates across platforms, training pipelines, and populations simultaneously. It is difficult for single studies, which are often limited in the number of platforms/models/populations they can test, to observe dynamics that emerge from the interaction of all three.
Measured evidence is usually a lagging indicator of complex real-world phenomena. Realworld situations are difficult to specify and translate into neat taxonomies and variables for modeling that are currently required for much of high-quality research. For example, it is difficult to assess how much personalization boosts persuasiveness. In research settings, persuasiveness can only be measured by proxy. In doing so, some studies find no effect (Hackenburg et al., 2025; Argyle et al., 2025), others find small effects (Kelley and Riedl, 2026; Matz et al., 2024a), and others find strong effects (Salvi et al., 2025a). This disagreement may itself reflect the difficulty of measuring effects that depend on sustained, naturalistic interaction rather than one-shot exposure.
The scientific endeavor also risks overstating these effects. Laboratory settings are not always representative of real-world environments; experimental designs may act as high-fidelity amplifiers that maximize pressure to achieve statistical rigor. In the lab, moving a participant’s opinion on a policy issue by a few percentage points may be a controlled success with effect size, but in the actual chaos of a political cycle even massive advertising spending can struggle to move the needle at all (Coppock et al., 2020).”
Moderation updates
Looking for a way to skill-build as a student or young graduate? Consider charity trusteeship!
EAs are often good at maths/technology and less good at leadership and governance skills. This is an excellent match for (non-EA) trustee governance boards, that often have committed older people with good leadership and governance but who struggle with use of tech and maths. There are a lot of trustee vacancies around—highly absorbable—and it’s an easy way to get started on tackling real issues with charities big and small!
You can help your charity act more effectively in the world using EA principles and frameworks about effective use of charitable resources, and learn experientially about practical issues than arise with those frameworks when they are implemented. Organisations like the Young Trustees Movement (UK) can mentor and support you, and the other trustees will probably have a lot of sage advice. And while trusteeship is unpaid, it also doesn’t *cost* money either, so it’s a fairly cheap way to do “EA as a hobby”.
The EA community needs more people aged 30+ with genuine commitment and leadership skills, and should encourage people looking for principles-first EA to consider this as a life path in their 20s.
Good shout! And right now at EA Netherlands we’re looking for new trustees!
Just under a month until the Cluelessness Critiques Competition deadline. How’s it going? Any hints on what people are writing about?
I have made a semi-functional rich person impact ranker. It requires votes (and optionally, longform analyses from humans or AIs). If you want to help flesh it out, dm me on here or on twitter or if you have my details, anywhere I am contactable and I’ll send it to you.
How does it compare to https://impactlist.xyz/?
The Fallacy of Similar Magnitudes
A name I came up with for a pattern that comes up often enough to deserve one. Here’s a particular instance:
Suppose you’re a longtermist who thinks the vast majority of the (net positive) moral weight is in the far future, that AI existential risk is above 1%, and that it’s reasonably tractable to reduce. Some people who hold all three premises still say things like: “AI safety has maybe two orders of magnitude more resources than animal welfare, so the marginal dollar is better spent on animal welfare.”
This is the fallacy of similar magnitudes: in this case, the application is that neglectedness comparisons only matter when the causes are within a couple orders of magnitude of each other in stakes. Under the stated premises, they aren’t: the a priori difference in expected value is way larger. Two orders of magnitude in resources doesn’t move the needle even a bit.
Note: this AW response could be ok if you have certain views about risk aversion or worldview diversification, but that’s besides the point.
Hot take: When Anthropic IPO money starts flowing into AI safety, the ecosystem should consider operating more outside charitable structures. Charities get an indirect public subsidy via tax-deductible donations, but the trade-off is greater compliance burden and cost, restrictions on spending, and a greater reliance on public goodwill. As the world gets weirder from AI-driven job losses and political shifts, taxpayers may be less happy to see their foregone tax dollars going to China dialogues or high salaries for technical AI safety researchers. (Not legal or financial advice.)
Agreed, part of the thesis of https://surplus.dev!
Mental health is one of the most widespread, impactable, yet neglected global health issues of our time. Even knowing mental health conditions are generally under-reported and/ or misrepresented, evidence still shows widespread effects of mental wellbeing on physical health, life satisfaction, and productivity, however it has still not taken off as a major cause area in many EA organizations I have talked to. Why is this?
In addition to @huw’s great comment, there’s a branch of EA which focuses on well-being, and with that focus mental health interventions often look really effective. Check out the happier lives institute Huw mentioned, times the “WELLBY” and @MichaelPlant
G’day Madeline, I run an EA mental health org in India. The reason for this is simply that existing mental health interventions do not compare with GiveWell’s grants on a DALYs/$ basis. In my opinion, the reasons are:
DALY moral weights may be biased against depression
The moral weights of different diseases in the Global Burden of Disease study, which informs DALY estimates, are determined by asking the general public whether they’d prefer to have one disease against another. When you do this with depression, people who haven’t experienced it tend to prefer to have it to many other conditions. However, when you ask people who have experienced it, they choose many very painful conditions over depression. This is one of the widest gaps in the moral weight data. See Pyne et al. 2009 and this post.
Psychotherapy is usually modelled as a short-term effect
Psychotherapy is typically modelled as a treatment, and not a ‘skill’. What I mean by this is that a dose of psychotherapy is assumed to only have effects that decay over a period of time and zero out after that in most CEAs, including those from the Happier Lives Institute. However, many psychotherapy patients will tell you that they learned skills that were useful long after the therapy ended, and there is some limited evidence that psychotherapy’s effects may last decades, or potentially never zero out. If this were true, the effects could be very long-lasting and therefore it would be much more valuable to treat a case of depression.
Suicide prevention isn’t cost-effective if it’s only a short-term effect
Consider that for most of GiveWell’s top interventions, the bulk of the DALYs averted come from ‘saving’ a life—i.e., preventing a death from a disease in a way that allows the person to then go on to live a healthy life, such as preventing a malaria case in an under-5 (which they might die from), even if they go on to catch it after 5 years old.
As a short-term effect, psychotherapy can only postpone a suicide by the length of the treatment effect. But if it were a skill and had some durable long-term effect, it may genuinely prevent one, which would tremendously increase the value of suicide prevention interventions.
Existing interventions haven’t been cheap enough yet
With the exception of some incredible policy work in, for example, reducing toxicity of pesticides commonly used for suicide, existing interventions are still quite expensive. The Happier Lives Institute’s top charities cost ~$40 to treat a single person, while a bednet costs $7. I’m fudging the numbers a bit here, but if we stick with DALYs, psychotherapy is still about an order of magnitude more expensive than it needs to be to look great for EAs.
However, there is work being done to improve that! My charity, Kaya Guides, treated people for $20 each in April, at what we estimate is a similar effect size to the best charities, and we’re confident we can get below $10. We’re using a technique called guided self-help that allows us to dramatically reduce contact hours per participant (and being all-digital helps a lot, too).
Conclusion
Orgs like the Happier Lives Institute have done a lot of advocacy work too, to raise the profile of mental health within EA, and there are plenty of funders that take mental health seriously (in a way that apparently wasn’t true a decade ago). It is, after all, still a nascent space.
In 2023-2024, it seemed a very strong bias from the EA community and the core AI safety funders that Anthropic not to be argued against or funding diverted to anything that was critical of them (directly or indirectly).
I wonder how the community and funders have updated their beliefs and behaviours?
As one datapoint, this is from 2024, and I saw no subsequent evidence from core AI safety funders (including CG) that it influenced their willingness to fund me (I’ve also been regularly critical of Anthropic, and the other companies, on twitter and elsewhere throughout this time).
https://www.lcfi.ac.uk/news-events/blog/post/reflections-on-machines-of-loving-grace
Your essay was a good read but it is an incredibly low bar if we count it as “critical of Anthropic”.
My equivalent would be trying to get funding for serious advocacy or policy work that went against Anthropic’s position.
Examples of fairly critical remarks about Anthropic’s actions. All more recent than your timeframe, because twitter’s search function is a pain and because my memory for twitter etc discussion is less good.
https://x.com/S_OhEigeartaigh/status/2029475839654388069 (re: gullible bunch memo)
https://x.com/S_OhEigeartaigh/status/2026957849994108990 (re: walking back RSP commitments)
https://x.com/S_OhEigeartaigh/status/2019518744561873286 (re: safety evaluations)
(I am followed by some prominent US policy people, including the outgoing white house AI adviser, so if there were a bias against people who criticise Anthropic, I would expect to be punished pretty harshly)
And while ongoing commentary on their safety/policy actions might not count for your criteria, this paper has been influential enough with policymakers that it probably counts (and goes against aspects of Anthropic’s stance on China that underlies a lot of their positions). https://papers.ssrn.com/abstract=5278644
(Not sure about (the phrasing of) your impression, e.g. I remember this and this, and the anthropic tag has this, this, and this among others?)
Those are all 2025, and most are late 2025. Also I’d suggest they are potentially informative to my argument, given one of them is literally about:
“Anthropic is not being consistently candid about their connection to EA”
None of these are from “core AI safety funders”. Views in the EA community were all across the map, but I share James’s impression that big funders (especially CG) were too reluctant to fund anything that’s critical of AI companies.
They’re from the EA community, no?
Apply to be a Teacher for TARA Round 2, 2026 🧑🏫
• Part-time role ($80 AUD/hr, ~13 hours/week).
• Lead Saturday sessions and provide remote support.
• Must have strong ML skills and completed most or all of the ARENA curriculum.
• Applications close 8 August 2026, reviewing candidates on rolling basis.
• Check out the role description and apply.
I hastily vibecoded a ~live-updating EA Forum (posts + comments; I think users are static) database + semantic search because software is ~free now. I haven’t rigorously checked anything over because it was/is just a random side project so use at your own risk and discretion. Enjoy!
Full database is ~17GB
Link: https://eaforum.up.railway.app/
I’m recording a podcast (TWCBB) with @Lauren Gilbert [edit: we had to reschedule so I can still collect questions]. If there’s anything you’d like us to ask her, comment here.
Too late, but I would ask her to make the best case against immigration haha (much like you did with me and BINGOs 😅)
She had to delay so it’s not too late, but I already included that exact question :)