This is a really beautifully designed interactive tool and I appreciate how transparent you’ve been with the code.
But looking at this from the perspective of someone who works on the ground in global health and development, I worry that analysing a $26 billion portfolio built for systemic equity through a narrow, health only calculator could mislead philanthropists. It’s a bit like looking at a massive civil rights or education budget and finding it lacking because it didn’t buy malaria nets.
If you talk to practitioners who actually execute this work, a few big real world missing pieces stand out:
The charities this model uses as a benchmark focus on narrow, medical interventions with quick, easily measured metrics. MacKenzie Scott’s giving is intentionally structural. She’s investing in areas such as civic engagement, local leadership and systemic justice. Discounting the value of structural work just because it doesn’t yield a clinical trial isn’t scientific rigour. It’s just a measurement blind spot for things that cannot be neatly randomised.
You cannot scale single disease programs linearly. If you suddenly dump billions into narrow medical tracks, you hit a hard wall on the ground. You trigger local wage inflation for scarce medical staff, overwhelm local supply chains and completely pull local nurses away from general primary care just to service a single subsidised metric.
Since the baseline assumptions (the credibility tiers and realisation numbers) were drafted by AI, the model is essentially running complex math on top of text approximations. It takes what is essentially an AI’s best guess and wraps it in a bunch of equations to make it look definitive, rather than building the model on actual data from the field.
This is an incredibly interesting data exercise, but by leaving out field realities, the model relies on flat, linear scaling assumptions that completely overlook the law of diminishing returns. If the goal is to make this genuinely informative for philanthropists trying to drive long-term change, I’d love to see it evolve by bringing in some of those real world constraints.
A QALY is a health metric. Most of Scott’s giving targets economic mobility, education, and equity, whose value is largely non-health; a WELLBY or consumption frame would credit those buckets far more. The model is deliberately scoped to one question: how much health does the money buy? It is not a verdict on her choices.
and with your 2nd bullet:
The model also prices the same dollars at the global-health frontier. … At the skeptical defaults, the frontier delivers roughly 1,500× more health per marginal dollar.
That multiple is a marginal comparison — the next dollar, not the whole portfolio. Frontier-priced opportunities are scarce… And the implied frontier counterfactual — ~105 million QALYs at the default settings — would mean averting ~4.2 million child deaths, most of a full year of the world’s under-5 deaths; no amount of money buys that at bed-net prices. Redeploying the full $30 billion would ride up the marginal-cost curve — to a few hundred times rather than ~1,500×, at my guess — softer in magnitude, same in direction. The floor is direct cash, the one option with effectively unbounded capacity, which GiveWell now scores at 3–4× its own historic benchmark: in health-only terms, somewhere under 50–100×.
Thanks for the context but putting these caveats in the disclaimers doesn’t change how the tool works.
The underlying studies are solid. The problem is the layer on top of them. The scores and weights are AI set guesses on a slider, not field data. Again, that’s most true for the civic engagement, equity and justice buckets, which is where most of Scott’s money goes. The author says himself those buckets get “wide skeptical priors” because no health pathway has credible evidence. That isn’t a real effect adjusted for uncertainty. It’s a number filling in for missing data. Since those are the biggest buckets, that number is doing a lot of the work behind the headline frontier multiple. The comparison is weakest exactly where most of her money sits.
That’s my real concern. The tool takes her hardest to measure work and makes it look like it buys almost nothing, when that low number is a measurement gap, not a fact about the work. A model is only as strong as its weakest layer. For most of this portfolio, that layer is a guess standing in for how funding actually works on the ground, like whether a local coalition can absorb a sudden influx of cash, or the fact that structural change doesn’t run in a straight line to a health outcome.
Flagging that in the fine print is good practice, but it doesn’t fix it. The calculator is still running precise maths on top of guesses.
Hm, I empathise with your underlying concern, but don’t the sliders help with that?
(Tangentially, “precise maths on top of guesses” is basically how I’d describe all CEAs, including GiveWell’s which move hundreds of millions annually. The guesses get better in all kinds of ways, but at least when I last checked the AMF CEA it wasn’t using field data like you’d like. So in a sense you seem to be asking for more than whole teams of specialists can deliver, let alone one person and a bunch of subscriptions can do in their limited spare time?)
That’s a fair point on how these models work in general and I really don’t want to diminish Max’s efforts here. Building a tool like this takes serious effort. I certainly could not do it and I genuinely appreciate the craft that went into it!
I see where you’re coming from on the sliders, but my worry is more fundamental. Adjusting a slider lets us tweak the numbers, but it doesn’t change what the model is built to look for. When the baseline assumes that hard-to-measure community work has almost no value unless it’s backed by a trial, a slider just lets us pick a different guess. I’ll say it again that it can’t capture how real-world, long-term change actually happens on the ground.
But technical debate aside, the most disheartening part of all this has been seeing how the post is being used online. It’s been really disappointing to watch the wave of harsh backlash and inflammatory comments directed at MacKenzie Scott off the back of this.
She is giving away huge amounts of money with real trust, flexibility, and respect for local leaders doing good, complex work. When that kind of thoughtful giving gets turned into a single headline about “inefficiency,” it opens the door for a lot of unfair criticism directed at someone who is genuinely trying to tackle important issues.
That’s really why those like me on the field push back. Numbers don’t live in a vacuum and when complex human work gets reduced to a simple health calculator, the real-world conversation takes a hit.
I hear you on the mis-aimed backlash. I also interpret what she’s doing as fundamentally capabilitarian, which is underappreciated in these parts.
Thoughts on how the calculator might more fairly represent what MacKenzie Scott’s giving aims to achieve? I can ask Claude to create a mockup based on your bullets :)
This is a really beautifully designed interactive tool and I appreciate how transparent you’ve been with the code.
But looking at this from the perspective of someone who works on the ground in global health and development, I worry that analysing a $26 billion portfolio built for systemic equity through a narrow, health only calculator could mislead philanthropists. It’s a bit like looking at a massive civil rights or education budget and finding it lacking because it didn’t buy malaria nets.
If you talk to practitioners who actually execute this work, a few big real world missing pieces stand out:
The charities this model uses as a benchmark focus on narrow, medical interventions with quick, easily measured metrics. MacKenzie Scott’s giving is intentionally structural. She’s investing in areas such as civic engagement, local leadership and systemic justice. Discounting the value of structural work just because it doesn’t yield a clinical trial isn’t scientific rigour. It’s just a measurement blind spot for things that cannot be neatly randomised.
You cannot scale single disease programs linearly. If you suddenly dump billions into narrow medical tracks, you hit a hard wall on the ground. You trigger local wage inflation for scarce medical staff, overwhelm local supply chains and completely pull local nurses away from general primary care just to service a single subsidised metric.
Since the baseline assumptions (the credibility tiers and realisation numbers) were drafted by AI, the model is essentially running complex math on top of text approximations. It takes what is essentially an AI’s best guess and wraps it in a bunch of equations to make it look definitive, rather than building the model on actual data from the field.
This is an incredibly interesting data exercise, but by leaving out field realities, the model relies on flat, linear scaling assumptions that completely overlook the law of diminishing returns. If the goal is to make this genuinely informative for philanthropists trying to drive long-term change, I’d love to see it evolve by bringing in some of those real world constraints.
FWIW the app agrees with your 1st bullet:
and with your 2nd bullet:
Thanks for the context but putting these caveats in the disclaimers doesn’t change how the tool works.
The underlying studies are solid. The problem is the layer on top of them. The scores and weights are AI set guesses on a slider, not field data. Again, that’s most true for the civic engagement, equity and justice buckets, which is where most of Scott’s money goes. The author says himself those buckets get “wide skeptical priors” because no health pathway has credible evidence. That isn’t a real effect adjusted for uncertainty. It’s a number filling in for missing data. Since those are the biggest buckets, that number is doing a lot of the work behind the headline frontier multiple. The comparison is weakest exactly where most of her money sits.
That’s my real concern. The tool takes her hardest to measure work and makes it look like it buys almost nothing, when that low number is a measurement gap, not a fact about the work. A model is only as strong as its weakest layer. For most of this portfolio, that layer is a guess standing in for how funding actually works on the ground, like whether a local coalition can absorb a sudden influx of cash, or the fact that structural change doesn’t run in a straight line to a health outcome.
Flagging that in the fine print is good practice, but it doesn’t fix it. The calculator is still running precise maths on top of guesses.
Hm, I empathise with your underlying concern, but don’t the sliders help with that?
(Tangentially, “precise maths on top of guesses” is basically how I’d describe all CEAs, including GiveWell’s which move hundreds of millions annually. The guesses get better in all kinds of ways, but at least when I last checked the AMF CEA it wasn’t using field data like you’d like. So in a sense you seem to be asking for more than whole teams of specialists can deliver, let alone one person and a bunch of subscriptions can do in their limited spare time?)
That’s a fair point on how these models work in general and I really don’t want to diminish Max’s efforts here. Building a tool like this takes serious effort. I certainly could not do it and I genuinely appreciate the craft that went into it!
I see where you’re coming from on the sliders, but my worry is more fundamental. Adjusting a slider lets us tweak the numbers, but it doesn’t change what the model is built to look for. When the baseline assumes that hard-to-measure community work has almost no value unless it’s backed by a trial, a slider just lets us pick a different guess. I’ll say it again that it can’t capture how real-world, long-term change actually happens on the ground.
But technical debate aside, the most disheartening part of all this has been seeing how the post is being used online. It’s been really disappointing to watch the wave of harsh backlash and inflammatory comments directed at MacKenzie Scott off the back of this.
She is giving away huge amounts of money with real trust, flexibility, and respect for local leaders doing good, complex work. When that kind of thoughtful giving gets turned into a single headline about “inefficiency,” it opens the door for a lot of unfair criticism directed at someone who is genuinely trying to tackle important issues.
That’s really why those like me on the field push back. Numbers don’t live in a vacuum and when complex human work gets reduced to a simple health calculator, the real-world conversation takes a hit.
I hear you on the mis-aimed backlash. I also interpret what she’s doing as fundamentally capabilitarian, which is underappreciated in these parts.
Thoughts on how the calculator might more fairly represent what MacKenzie Scott’s giving aims to achieve? I can ask Claude to create a mockup based on your bullets :)