It seems plausible to me that the existence of higher degrees of random error could inflate a more error-tolerant evaluator’s CEAs for funded grants as a class. Someone could probably quantify that intuition a whole lot better, but here’s one thought experiment:
Suppose ResourceHeavy and QuickMover [which are not intended to be GiveWell and FP!] are evaluating a pool of 100 grant opportunities and have room to fund 16 of them. Each has a policy of selecting the grants that score highest on cost-effectiveness. ResourceHeavy spends a ton of resources and determines the precise cost-effectiveness of each grant opportunity. To keep the hypo simple, let’s suppose that all 100 have a true cost effectiveness of 10.00-10.09 Units, and ResourceHeavy nails it on each candidate. QuickMover’s results, in contrast, include a normally-distributed error with a mean of 0 and a standard deviation of 3.
In this hypothetical, QuickMover is the more efficient operator because the underlying opportunities were ~indistinguishable anyway. However, QuickMover will erroneously claim that its selected projects have a cost-effectiveness of ~13+ Units because it unknowingly selected the 16 projects with the highest positive error terms (i.e., those with an error of +1 SD or above). Moreover, the random distribution of error determined which grants got funded and which did not—which is OK here since all candidates were ~indistinguishable but will be problematic in real-world situations.
While the hypo is unrealistic in some ways, it seems that given a significant error term, which grants clear a 10-Unit bar may be strongly influenced by random error, and that might undermine confidence in QuickMover’s selections. Moreover, significant error could result in inflated CEAs on funded grants as a class (as opposed to all evaluated grants as a class) because the error is a in some ways a one-way rachet—grants with significant negative error terms generally don’t get funded.
I’m sure someone with better quant skills than I could emulate a grant pool with variable cost-effectiveness in addition to a variable error term. And maybe these kinds of issues, even if they exist outside of thought experiments, could be too small in practice to matter much?
It’s definitely true that all else equal, uncertainty inflates CEAs of funded grants, for the reasons you identify. (This is an example of the optimizer’s curse.) However:
This risk is lower when the variance in true CE is large, especially if its larger than the variance due to measurement error. To the extent we think this is true in the opportunities we evaluate, this reduces the quantitative contribution of measurement error to CE inflation. More elaboration in this comment.
Good CEAs are conservative in their choices of inputs for exactly this reason. The goal should be to establish the minimal conditions for a grant to be worth making, as opposed to providing precise point estimates of CE.
[highly speculative]
It seems plausible to me that the existence of higher degrees of random error could inflate a more error-tolerant evaluator’s CEAs for funded grants as a class. Someone could probably quantify that intuition a whole lot better, but here’s one thought experiment:
Suppose ResourceHeavy and QuickMover [which are not intended to be GiveWell and FP!] are evaluating a pool of 100 grant opportunities and have room to fund 16 of them. Each has a policy of selecting the grants that score highest on cost-effectiveness. ResourceHeavy spends a ton of resources and determines the precise cost-effectiveness of each grant opportunity. To keep the hypo simple, let’s suppose that all 100 have a true cost effectiveness of 10.00-10.09 Units, and ResourceHeavy nails it on each candidate. QuickMover’s results, in contrast, include a normally-distributed error with a mean of 0 and a standard deviation of 3.
In this hypothetical, QuickMover is the more efficient operator because the underlying opportunities were ~indistinguishable anyway. However, QuickMover will erroneously claim that its selected projects have a cost-effectiveness of ~13+ Units because it unknowingly selected the 16 projects with the highest positive error terms (i.e., those with an error of +1 SD or above). Moreover, the random distribution of error determined which grants got funded and which did not—which is OK here since all candidates were ~indistinguishable but will be problematic in real-world situations.
While the hypo is unrealistic in some ways, it seems that given a significant error term, which grants clear a 10-Unit bar may be strongly influenced by random error, and that might undermine confidence in QuickMover’s selections. Moreover, significant error could result in inflated CEAs on funded grants as a class (as opposed to all evaluated grants as a class) because the error is a in some ways a one-way rachet—grants with significant negative error terms generally don’t get funded.
I’m sure someone with better quant skills than I could emulate a grant pool with variable cost-effectiveness in addition to a variable error term. And maybe these kinds of issues, even if they exist outside of thought experiments, could be too small in practice to matter much?
It’s definitely true that all else equal, uncertainty inflates CEAs of funded grants, for the reasons you identify. (This is an example of the optimizer’s curse.) However:
This risk is lower when the variance in true CE is large, especially if its larger than the variance due to measurement error. To the extent we think this is true in the opportunities we evaluate, this reduces the quantitative contribution of measurement error to CE inflation. More elaboration in this comment.
Good CEAs are conservative in their choices of inputs for exactly this reason. The goal should be to establish the minimal conditions for a grant to be worth making, as opposed to providing precise point estimates of CE.