Humanitarian practitioner with 20+ years of experience across field operations, programme management, emergency response, cash assistance and data protection. I’m interested in how AI may change decision-making and resource allocation in humanitarian aid and philanthropy. I’m currently testing some of these ideas through zooidfund, a platform where donors’ own AI agents assess real funding needs and can make direct donations.
Alex Novikau
I am not seeing it mentioned in the post, and I assume you do this, but one important addition is a simple decision log. Things need to move fast and decisions need to be taken quickly, but institutional memory can disappear surprisingly fast around what was decided, by whom and when. That can make both handover and any later postmortem much harder.
Great post. There may be relevant lessons here to learn from the aid localisation push. One specific and well-established failure mode is capacity building that improves an organisation’s ability to satisfy donor requirements, rather than its ability to deliver. In the worst case, organisations become more legible and fundable at the expense of effectiveness.
We are dealing with severely misaligned incentives. And when incentives point the wrong way, awareness does not fix much.
Anyone working at a frontier lab is presented with an extraordinary asymmetry: enormous and immediate personal rewards from pushing capabilities forward, against a probabilistic risk borne by everyone else. People are extremely good at persuading themselves that what benefits them is also the right thing to do, especially when the rewards are this high.
The US government, even if it understands the risks, is locked into strategic competition with China that makes slowing down impossible. A serious warning confirms the danger, but can make frontier AI seem even more important, increasing the pressure to get there first.
And internationally, the institutions that might once have helped manage this kind of competition are all being torched just as they are most needed. Communication, verification, restraint and collective action that helped prevent WWIII are all being made harder, leaving fewer ways to slow or manage the race.
So we are heading to a situation where the labs, governments and even the public understand the danger quite well, but the incentives of those who control development still push them toward acceleration, and the mechanisms for managing this are gone.
Should the reduced search and assessment costs on the side of employers also count towards impact? The selection by talent programs creates a signal about candidate quality that could save potential employers effort and resources they would otherwise have to invest in identifying and assessing the right candidates.
In that sense, some of what looks like a selection effect from the participant side may still actually be a benefit created by the program.
There is probably an architecture shape that could help solve this: interactions with AI advisor remain private but the AI advisor also helps the applicant create and maintain a live profile, that you can then base your prioritization decisions on, while applicant still controls what goes in.
This can then be a continuous screening process rather than a discreet application-decision cycle.
Many of the issues you mention may come from organizations not actually being ready to hire for senior roles.
A genuinely senior position requires giving someone meaningful decision-making authority, not just hiring a more experienced person to execute the founders’ decisions. Younger founders in small organizations can struggle with this. Delegation can be quite uncomfortable for an inexperienced manager, especially when the person being hired may disagree with them or change how things are done.
That can manifest as never-ending hiring rounds, repeated “culture fit” rejections, or eventually hiring someone who is not really qualified for a senior role but is unlikely to challenge the founders or demand much autonomy.
You are right, measuring AI maintenance burden without accounting for maintenance that had to be done regardless of AI deployment would not get you anywhere.
Interesting you are have Claude identifying opportunities to simplify/reuse/streamline, this could actually mean the per system maintenance burden will go down the more systems you connect. Pretty much the opposite of what I was thinking. Thank you.
Is the proposal to evaluate country prospects, or does it include also weighing these against further expansion in your existing countries?
With my admittedly very limited understanding of your work, it looks like the factors determining the success for your interventions can be quite local. India alone is huge with a significant variety of local conditions and your existing intervention is addressing a small subset of that diversity.
If this succeeds are you considering making the AI advisor the “allocation layer” for your human advisors? Human advisor attention is scarce so you programme has to be selective. What you currently rely on for this selection is probably mostly what’s in the application. If you can augment this by the knowledge accumulated by the AI advisor, it could make the quality of allocating your human advisors’ attention a lot better.
This is provided human advisors still retain some advantage over ai going forward.
Talent Commons is potentially very valuable, but there is an important path-dependence problem to address in the design.
Once assessments from SPAR, MATS or hiring rounds become reusable signals across organizations, a judgement made for one particular selection process will start affecting someone’s opportunities throughout the ecosystem. This will be especially consequential in an early-career talent pool, where the signal is noisy and people change quickly. Someone misclassified at 20 could end up being affected by that assessment much longer than they would today, while an initially positive assessment could create the opposite cumulative advantage.
I looked at SPAR’s privacy notice and there is already quite serious thought going into transparency, correction and optional sharing. But I think the important question for Talent Commons is whether it can preserve the provenance and decision-specific meaning of assessments rather than gradually turning them into general reputation scores. Getting that right would also be necessary so that the system actually improves talent allocation rather than just making existing judgements more portable and longer lasting.
A lot of operational improvements look solid while the person who introduced them is still there, and then disappear when that person changes role or leaves. Given how much institutional knowledge you found sitting in founders’ heads and inboxes, staff turnover over the next year could be quite a revealing test.
If some of the organisations go through that kind of transition, whether the processes still work without the person who set them up seems like a stronger test of institutionalisation than another readiness score on its own.
Yes this is the hard part. AI can make it much cheaper for a donor to assess an unfamiliar organisation once there is enough evidence to work with, but it cannot manufacture information that is not visible in the first place. And there is probably a second-order problem here: once AI starts doing more of the assessment, organisations that are easier for machines to identify, document and compare may get an advantage simply because they are more legible to the system.
I have been thinking about this as a kind of “machine-legibility privilege.” It is one of the things I am trying to test with zooidfund: whether AI assessment can actually broaden the set of needs donors can consider outside their existing networks, without simply shifting the advantage toward whoever leaves the best digital trail.
The section on scale also made me think about what happens if more of the funding process becomes automated. AI could make it much cheaper to discover and assess large numbers of small organisations, which seems potentially very useful for exactly the kind of distributed funding being discussed here. But it also creates another version of the legibility problem: an automated system will need signals it can actually process, and those signals may favour organisations that can produce standardized evidence, reporting and data over organisations that are harder to describe but locally very effective.
So there may be a tension between making the funding infrastructure scalable and keeping the substantive judgment genuinely local. It seems possible to automate a lot of the boring infrastructure without requiring the organisations themselves to become more standardized, but I don’t think that follows automatically.
Coming from a fairly traditional international organization, I think I can understand the asymmetry you describe. A lot of professional competence is difficult to teach because it is learned through repeated exposure to how institutions actually behave. You gradually get a feel for which proposals can survive internal processes, where formal authority differs from practical influence, and which ideas become much harder once they have to work across different departments of an organization, let alone different organizations.
Case discussions or “war stories” can work to surface this. Take for example a concrete AI policy proposal and ask people from different professional backgrounds to work through how they think it would fare inside a real institution. I suspect that would make some of this tacit knowledge much easier to see than trying to write it down as a set of principles.
Hi Roland, what are some of the big-if-true ideas you considered?
This was interesting to read, especially the point about the marginal cost of automating additional tasks falling as the underlying infrastructure gets better. One thing I was wondering about is how the maintenance side is developing as you add more automations.
At 11% you can probably still understand each automation fairly well, but if the aim is to automate a much larger share of program operations, some of the work presumably shifts from doing the task itself to keeping a growing set of automations working as forms, Drive structures, naming conventions, staff roles etc. change. I don’t know if you are already measuring this, but something like human minutes per run or per program round, including review, exceptions and repairs, might be useful alongside the percentage of tasks automated. It could help distinguish automations that really reduce operational load from ones that mostly move that load somewhere less visible.
One difficulty with this kind of public-good infrastructure is that it is often unclear who the customer is. Lots of people may benefit if it exists, without any one of them having a strong incentive to pay for it, and that also makes it difficult for someone considering building it to tell whether they have found a real ecosystem need or just something that sounds useful in the abstract. I may have run into this while building zooidfund, where there are several plausible beneficiaries of the infrastructure but that does not necessarily translate into a clear demand signal.
So I wonder whether part of the missing infrastructure is actually on the funder side, with more explicit problem statements, RFPs, bounties or advance commitments around gaps they think are worth solving. That would still leave plenty of room for people to notice problems and take initiative, but it would make it easier to know when a particular gap is something others genuinely want solved before somebody spends a year building around it.
I like the vulture idea. A bit more sceptical about infotainment as you predicted. Awaremess interventions is a well trodden field. I wonder if the Ethiopia result is evidence for infotainment generally, or for something quite a bit narrower. The intervention was not really mass media in the usual sense, it was people watching stories about individuals from backgrounds deliberately made similar to their own, and if I understand the study correctly there were also peer effects depending on how many people in the village saw it.
That makes me a little unsure about the jump to cheap radio broadcasting. It may work, but perhaps the important part is actually identification with a very locally credible role model, plus people around you having seen the same thing, rather than the medium itself. If so the scalable version might need much more localisation than just producing a good programme and buying airtime.
I am not sure the comparison stays entirely symmetric once endogenous growth is added on the climate side. If climate damage can permanently change the growth path, then at least some GHD interventions should also have effects on growth through health, schooling, productivity, fertility and so on.
I assume some of that is already captured in the GiveWell/Coefficient models, but probably not the same kind of macro second-order effects. Since climate seems to become competitive mainly when these more uncertain effects are included, I would be interested in how much this matters.
Real-world action is rarely a matter of a single moral decision. It often involves a combination of decisions, each with its own stakes and potential for moral failure.
In an AI-mediated giving experiment I am working on, zooidfund, delegating altruistic action to an AI agent is not simply a linear continuum between receiving advice at one end and letting the agent execute a donation end to end at the other. Shortlisting potential recipients, deciding which cases merit further assessment, evaluating evidence, calibrating donation amounts, and formulating messages to recipients are all distinct decisions with moral components. Relying on AI may improve the outcome of each. Delegating some of these decisions to AI may therefore be a moral choice in its own right, if we believe we should use the best available tools to increase the impact of our giving.
I therefore think there is a strong case for developing a practice of real-world moral delegation now, through well-defined and bounded decisions within pro-social action. This need not wait until alignment is solved. Practical experience of delegating parts of complex pro-social action to AI may itself generate useful evidence about where delegation works, where it fails and why, and ultimately contribute to alignment.