I’m a doctor working towards the dream that every human will have access to high quality healthcare. I’m a medic and director of OneDay Health, which has launched 53 simple but comprehensive nurse-led health centers in remote rural Ugandan Villages. A huge thanks to the EA Cambridge student community in 2018 for helping me realise that I could do more good by focusing on providing healthcare in remote places.
NickLaing
Yep that’s the one!
I don’t think it would really help vs other interventions much (unless they are also doing rigorous research, which I doubt).
Because you’ll get very specifid information about the differences, It might help you understand in which ways the program helps and doesn’t as well so you could adjust programming as well. Like you might find it helps people get more jobs but not better quality jobs than those who don’t do the program. Or you might find less people enter capabilities rather than safety etc.
I think you misunderstood me here. I wasn’t suggesting you research your selection criteria, I was suggesting a basic Randomised controlled trial design possibliity for one intake to assess the impact of the program. An RCT woudln’t help you much (if at all) with selection criteria). The candidates who were selected and not selected would then be followed over time to see if the MATS program made a difference to their trajectories.
”every MATS cohort counts” might be true, but if the difference in fellows isn’t that big why not randomise one round at least? Not sure how big your cohorts are though. If they are too small this wouldn’t work at all.
Gotcha. Have no idea how good cG are at impact analysis—I’m not sure I’ve seen one publicised by them, I think they have decided against publicising things in general. I don’t think that AI safety people are necessarily that good at impact analysis though—very few would have trained at all in that field. Being good at impact analysis and understanding what is best for AI safety are completely different and underated things.
I get that people have different ideas about what is a “good” outcome, but CoGi at least must have a criteria if they are doing an assessment at all. The criteria could be very complex and nuanced, that’s fine—as long the criteria is clear before you start the study.
I would argue if you don’t know what is a good outcome, there’s probably no point in doing any work at all.
“I don’t think that an independent analysis is necessarily better than the default.” What do you mean by this exactly?
@Ryan Kidd this is great feedback you are getting and the surveys seem high quality, but just surveying participants will always be among the lowest forms of evidence. Surveys are useful (I use them a lot) but in the global health field we would want to move up the evidence ladder to get stronger evidence before rolling million plus dollar interventions out. People’s Opinions and perceptions are not necesarily true (we’ve seen this many times in global health) while good observational and experimental studies can answer questions more conclusively and discover things we don’t expect.
There’s no real counterfactual or experimentation if you are doing a survey. To get better data you could at least do a form of Case-control study (with rejected participants) or a prospective cohort study with similar groups of people. Doing an RCT with applicants seems harsh on the best participants and is a bit unfair for sure (see thread with @simon below), but maybe running one rigorous RCT could be a good idea especially if people know what they are signing up for before they applied.
Running one MATS application round where you randomised entry among the best participants would get you insanely good data not only about how much your program helps, but also the specific ways it helps. This can help you iterate your program in future.
If you’re looking at accelerated career progress you might need to follow people for 3-5 years after finishing the course to measure the effects. This length of time isn’t unusual in global health studies.
I like the author’s design too, but I have an innate dislike of retrospective studies. There’s just so much room for fiddling results and decisions that favour the org being researched, even publishing bias. Tiny data manipulations before data handovers can rig the study really easily. For example MATS could just copy paste @Laura Thomas-Walters paragraph into Claude Fable, feed in their applicant data and just run the study themselves right now. IF the results were favourable they could hand you the data, if not they could just ignore.
1.
I generally agree with you that there’s a lot of candidate noise in rank ordering, but many EAs seem to think there are huge differences between first and second tier candidates in candidate selection in these situatoins. I don’t think a lottery process is implausible here, but it still does feel unfair that you could be the top person selected for a course and not get in because of randomisation.
This can be ethically a bit savage and unfair, especially for a flagship program at a criticall point in someone’s career. You could do this, but it has downsides.
From my perspective, I would be happy with them at least comparing with the applicants closest to the bar. This comparison favours the course for sure, but at least it gives you clear info. I would want them to state the study protocol and make sure they really compared vs. the best rejections though,
If this is even mostly true, from a Global health perspective at least this is borderline neglectful lack of Monitoring and Evaluation. Why do these orgs almost assume their program works, or state vague non-counterfactual data as evidence?
I find the BlueDot example almost hard to believe. BlueDot is such a flagship AI safety program, I would have thought they would have had much more rigorous impact evidence and follow up published publicly. As much to improve their program as well as demonstrate impact.
”BlueDot Impact has raised $35 million and publishes “25% of our graduates land impactful roles within six months” with no definition of impactful roles, no denominator and no method.” doesn’t tell you much at all especially without a counterfactual
Sometimes we give orgs like GiveWell a hard time for poor follow up of orgs they fund, which has seemed a bit rough and a bit of a stretch for me. They copped a bit of flack after the Evidence Action Water Dispenser fail. And that’s not even the org itself, but the funder getting a hard time This here seems like another level.
IT would be so easy to look at a counterfactual here—its not perfect as they are rejected applicants but at least you need to show that your people are doing far better than the next teir down of rejected applicants who didn’t even do the program. Surely at least you could pay the rejected people all $50 to fill in a survey and say what they are doing now and get most of the info you need. Basic counterfactual comparison like this seems like it should be the baseline, with some more rigorous studies done as well.
I know programs might hate this, but randomly assigning applicants to one program or another then comparing their progress after the course might be a great way to compare courses as well.
I hope we’re missing something here because this seems pretty neglectful on a first pass...
I kind of like the idea, but I must confess still don’t entirely understand how an impact market would work exactly? Lets say someone independently assessed we at OneDay Health at achieving the equivalent of 1 life saved for every 2,500 patients that we treat. How would an impact market work from there?
Well done you just won the “not scared of calling out givewell prize!” ;)
I couldn’t agree with this more. I try and keep this vibe as much as possible, even though I do now run a bigger, more “professionalised” institution.
“Sometimes I’d fuck it up and say something like “Shit, I fucked up the pitch.” It humanizes you. After all, the person isn’t joining some abstract thing called Binghamton EA. At least at the very very beginning, they’re joining something that is mostly just us and the people around us. You are not an institution”
Love this solidarity poem, thank you!
Personally I’m very happy that long termists have a special bias towards humans. I think your statement here probably only holds if your are a pure hedonistic utilitarian, which many of us are not.
“II’m highly uncertain, but if I had to guess, humanity has had a net negative impact to date due to the immense scale of suffering on factory farms.”
I don’t think the bias is very hidden either. My not-, very-considered take would be I don’t think we necessarily need to “change course, but I think a truly impartial school of philosophy is useful and important too.. I would maybe favor 2 separate schools of thought and research about the future.
A school that explicitly favors human and serves human interests.
A school that tries to be completely impartial like you discuss.
I’ve never been able to find a heuristic for this exactly. We have made many decisions which we don’t love to call sacrifice, in fact many of them brought deep satisfaction and connection. These were made to live in better solidarity and community with those around us. For example for us this involved cooking on a charcoal stove, not having running water inside and eating a lot of local food which is often great but I don’t always love.
Almost all of these were made at the cost of efficiency in our work to some extent. Life without a gas stove and running water just takes longer and is a bit more tiring. Not nearly as big a deal as most people would think tho...
The closest we had to a “heuristic” was asking ourselvesevery 6 months, is this lifestyle...
1) Satisfying to me
2) Appropriate lifestyle that matched our community around us. We’ve always chosen to live among less well-off communities because I think it grounds our work in reality, keeps it real and helps us learn what actually makes people tick. (which I personally think is often critical to designing effective GHD interventions)
3) Allowing me to do my work well
As OneDay Health and my Wife’s work grew, I needed more time. No. 3 started to take precedence. We became more “bougie”. A water tap outside, more eating out (still local food). Now after having a kid we have a gas stove. It doesn’t feel great to move further away from our community lifestyle, but I think its worth it for the impact—as long as we don’t feel too disconnected from those around us.
Love this great work! I think we need more careful, thorough but straightforward explanations like this. Understanding why some charities are better than others can often be harder than we think to people outside our bubble.
For me this is the crux and the key line.”Climate’s SROI becomes competitive with GHD when accounting for high-risk, lower-certainty effects like tipping points and endogenous growth”
If you’re the kind of person who both has a high risk appetite, and believes we have some plausible control now over low certainty events in the far future then I would agree with the thesis that we should put more money into climate stuff.
Personally I think the medium-far climate future might be as difficult to predict as the trajectory of AI. We have political uncertainty (e.g. Trump scrapping wind), economic uncertainty (e.g. solar became unexpectadly cheap), potential AI tech progress, potential rogue actors (someone could release sulfur). We could be at net zero by 2040 or never.
Anyway I think the uncertainty with climate is so large that you need to be a certain kind of giver for that to make sense—especially with the “existential risk” component perhaps not being there.
I don’t find much solace and meaning in error bars that wide—but many others do.
Thanks for this great write up, and it all sounds very reasonable. I really respect that the orgs here didn’t go crazy claiming they could spend ludicrous amounts of money straight away. I agree with @MichaelDickens that is crazy how little funding there is in this area, and that’s from someone who might rate animal welfare stuff 1000x less important than him haha.
As a rule of thumb in the global health world at least, I think there are few orgs it’s where it makes sense to increase spending than 2x − 3x funding year on year, scaling up orgs and operations isn’t easy.
I think I said this on another post, but in the wild animal welfare space I would love to see a couple of clear practical wins, even if really small to show that change is possible here, and learn how to make change. Even if it’s more token than meaningful. Most of these orgs seem heavily focused on research which is helpful, but I think we can learn a lot as well from doing things and seeing what happens.
Kudos for the economist for my favourite mainstream news article on the Anthropalypse yet
https://www.economist.com/international/2026/08/13/silicon-valleys-ai-boom-is-remaking-american-charity(Behind a paywall)
I think it outlines the issues very well, and you won’t see a mainstream news article start as smart as this very often.
”EACH DOLLAR given to charity may soon do less good. In the coming years the marginal cost of saving a child’s life from disease or starvation could jump from about $5,000 to $15,000, or more. This sounds worrying. In fact it is good news, argues Alexander Berger of Coefficient Giving, one of Silicon Valley’s most influential grantmakers. Non-profits like his spend on cheap, scalable interventions first—say, by buying malaria nets before malaria vaccines. If an influx of donations pays for all the inexpensive ways of doing good, then the rest flows into costlier acts of altruism that can save yet more lives.”
I read this just because of this into “every 5 minutes while writing this post, I had to stop to do 13 push-ups, and when I could no longer complete my required number, I had to upload it.
And I can’t believe you did 468 pushups that’s crazy. I would have published after maybe 10 minutes under these rules?
Good reminder to be a bit more focused with my podcasts/reading. I like to have broad understanding, but its probably going too far that way.
@Laura Thomas-Walters I think he’s arguing against my idea not yours. I agree 100% I think there’s no reason at all not to do the retrospective design (AI could do most of it anyway), I’m just saying its still a pretty low level of research really. I agree a big improvenet though.