Humanitarian practitioner with 20+ years of experience across field operations, programme management, emergency response, cash assistance and data protection. I’m interested in how AI may change decision-making and resource allocation in humanitarian aid and philanthropy. I’m currently testing some of these ideas through zooidfund, a platform where donors’ own AI agents assess real funding needs and can make direct donations.
Alex Novikau
Yes this is the hard part. AI can make it much cheaper for a donor to assess an unfamiliar organisation once there is enough evidence to work with, but it cannot manufacture information that is not visible in the first place. And there is probably a second-order problem here: once AI starts doing more of the assessment, organisations that are easier for machines to identify, document and compare may get an advantage simply because they are more legible to the system.
I have been thinking about this as a kind of “machine-legibility privilege.” It is one of the things I am trying to test with zooidfund: whether AI assessment can actually broaden the set of needs donors can consider outside their existing networks, without simply shifting the advantage toward whoever leaves the best digital trail.
The section on scale also made me think about what happens if more of the funding process becomes automated. AI could make it much cheaper to discover and assess large numbers of small organisations, which seems potentially very useful for exactly the kind of distributed funding being discussed here. But it also creates another version of the legibility problem: an automated system will need signals it can actually process, and those signals may favour organisations that can produce standardized evidence, reporting and data over organisations that are harder to describe but locally very effective.
So there may be a tension between making the funding infrastructure scalable and keeping the substantive judgment genuinely local. It seems possible to automate a lot of the boring infrastructure without requiring the organisations themselves to become more standardized, but I don’t think that follows automatically.
Coming from a fairly traditional international organization, I think I can understand the asymmetry you describe. A lot of professional competence is difficult to teach because it is learned through repeated exposure to how institutions actually behave. You gradually get a feel for which proposals can survive internal processes, where formal authority differs from practical influence, and which ideas become much harder once they have to work across different departments of an organization, let alone different organizations.
Case discussions or “war stories” can work to surface this. Take for example a concrete AI policy proposal and ask people from different professional backgrounds to work through how they think it would fare inside a real institution. I suspect that would make some of this tacit knowledge much easier to see than trying to write it down as a set of principles.
Hi Roland, what are some of the big-if-true ideas you considered?
This was interesting to read, especially the point about the marginal cost of automating additional tasks falling as the underlying infrastructure gets better. One thing I was wondering about is how the maintenance side is developing as you add more automations.
At 11% you can probably still understand each automation fairly well, but if the aim is to automate a much larger share of program operations, some of the work presumably shifts from doing the task itself to keeping a growing set of automations working as forms, Drive structures, naming conventions, staff roles etc. change. I don’t know if you are already measuring this, but something like human minutes per run or per program round, including review, exceptions and repairs, might be useful alongside the percentage of tasks automated. It could help distinguish automations that really reduce operational load from ones that mostly move that load somewhere less visible.
One difficulty with this kind of public-good infrastructure is that it is often unclear who the customer is. Lots of people may benefit if it exists, without any one of them having a strong incentive to pay for it, and that also makes it difficult for someone considering building it to tell whether they have found a real ecosystem need or just something that sounds useful in the abstract. I may have run into this while building zooidfund, where there are several plausible beneficiaries of the infrastructure but that does not necessarily translate into a clear demand signal.
So I wonder whether part of the missing infrastructure is actually on the funder side, with more explicit problem statements, RFPs, bounties or advance commitments around gaps they think are worth solving. That would still leave plenty of room for people to notice problems and take initiative, but it would make it easier to know when a particular gap is something others genuinely want solved before somebody spends a year building around it.
I like the vulture idea. A bit more sceptical about infotainment as you predicted. Awaremess interventions is a well trodden field. I wonder if the Ethiopia result is evidence for infotainment generally, or for something quite a bit narrower. The intervention was not really mass media in the usual sense, it was people watching stories about individuals from backgrounds deliberately made similar to their own, and if I understand the study correctly there were also peer effects depending on how many people in the village saw it.
That makes me a little unsure about the jump to cheap radio broadcasting. It may work, but perhaps the important part is actually identification with a very locally credible role model, plus people around you having seen the same thing, rather than the medium itself. If so the scalable version might need much more localisation than just producing a good programme and buying airtime.
I am not sure the comparison stays entirely symmetric once endogenous growth is added on the climate side. If climate damage can permanently change the growth path, then at least some GHD interventions should also have effects on growth through health, schooling, productivity, fertility and so on.
I assume some of that is already captured in the GiveWell/Coefficient models, but probably not the same kind of macro second-order effects. Since climate seems to become competitive mainly when these more uncertain effects are included, I would be interested in how much this matters.
Interesting idea to bring the two together. If we agree that our ability to predict the consequences of our actions is limited and we live in a fundamentally unpredictable world, would it follow that all else equal we should prioritize flexibility, ability to adapt and error correct in response to changing circumstance in our political organisation, which usually means decentralization, pluralism, rule of law and democracy?
I like this. I would maybe add these couple of points:
- “Cluelessness horizon” has a solid theoretical basis to it: human society is complex enough nonlinear system in which small differences can propagate to cause large changes. It may or may not be “chaotic system” in a strict mathematical sense, but it is non linear enough to say with some confidence that there is a finite horizon for predicting consequences of a particular event.
- There is no baseline. “Inaction” has the exact long tail of possible consequances any action does. One is affecting the world whether they donate 100 dollars to bed nets, spend it in a bar or light it on fire.
So everything disappears into uncertainty past a certain horizon. Only consequences that are before that event horizon matter for our decisions, and here bed nets are a very good candidate.
I think there is another reason to be cautious about longtermism: the bias of the EA community itself. More than other EA causes longtermism allows to redirect resources towards the community and its members, as opposed to more traditional causes, like malaria bed nets. The effect is also self reinforcing, the more successful the advocacy the more resources are directed at the topic the more EAs are paid to work, or want to be paid to work on longtermism causes the more effective the advocacy becomes.
I understand you are planning to measure against another show as a control. Which makes a lot of sense, but raises some questions: Have you accounted for a possibility that control treatment itself has effect on gender norms related to IPV? There is at least some evidence that exposure to media ( e.g. Jensen/Oster in India) has effect on gender norms and views on women autonomy. If control has positive effect, you would need to be comparing two potentially beneficial interventions, not your show vs. no program. Have you costed the control show?
Also, you intention is to change attitudes to IPV, if you rely on self reports, how will you distinguish actual change in violence from treatment-induced change in reporting rates? I could see how change in attitudes may both increase the rates it is disclosed because the behavior is seen as more objectionable, or decrease them because there is now more stigma attached to the behavior.
Fair point. Retracted.
Hi Dan, I tried to send you a private message, but the Forum’s messaging function keeps erroring for me. I’m building a practical donor-agent experiment that I thought might interest you. People give their own Claude or Codex agent a mandate to assess live humanitarian campaigns and decide whether any merit support. If you’re open to hearing more, could you email me at alex@zooid.fund?
Thank you for sharing this. An aligned model does not an aligned system make; the application layer must also be aligned.
This has parallels I believe to the discussion I was trying to initiate here, unsuccessfully, about the importance of investing in prosocial AI applications. Even if reverse alignment goals are broader.
Thank you, fair criticism. I agree with you that current activity can in no way be described as an alignment corpus. It is a small, domain-specific deployment. If anything this post is a call for participation to expand it—regardless of what the value is for alignment, it is I think a worthy and underdeveloped AI application.
The point I was trying to make rather, is that pro-social agent deployment is underdeveloped compared to commercial and productivity. If one is to believe that commercial and productivity deployment is already producing alignment-relevant data, and that seem to be the case, see OpenAI monitoring their own coding agents for example How we monitor internal coding agents for misalignment | OpenAI Then so will pro-social real-world behavior, which is currently underrepresented. zooidfund is not going to solve it by itself, but it can be a contribution.
Practice makes perfect. Agentic AI pro-social behavior as a missing alignment corpus.
Thank you for this and hope this will generate more discussion.
Have you considered the role technology, especially AI can play in this? Can some of what regrantors do be replaced or supplemented? AI is capable of processing large amounts of data and doing it on a ongoing basis, something that can be very, if not prohibitively costly for current donors and regrantors, especially at smaller scale. It can help achieve much more granular allocation: smaller projects and initiatives, individual needs, continuous rather than cyclical donations.Shortening the distance between the donor and the ultimate beneficially can help emotions and ego driven giving become more effective. Perhaps a limited gain for a donor like yourself, but potentially a huge impact overall.
I’ve been working on a small project in that space, aiming to create a layer Ai agents can apply themselves to to identify, assess and donate potential recipients. zooid.fund would be interested to know what you think
Thanks, Ben. Yes it does sound like a great fit. Where can I learn more?
A lot of operational improvements look solid while the person who introduced them is still there, and then disappear when that person changes role or leaves. Given how much institutional knowledge you found sitting in founders’ heads and inboxes, staff turnover over the next year could be quite a revealing test.
If some of the organisations go through that kind of transition, whether the processes still work without the person who set them up seems like a stronger test of institutionalisation than another readiness score on its own.