Thank you for spending time analyzing our methods. We appreciate those who are willing to engage with our work and help us improve the accuracy of our recommendations and reduce animal suffering as much as possible.
Based on previously received feedback and internal reflection, we have significantly updated our evaluation methods in the past year and will be publishing the details next Tuesday when we release our charity recommendations for 2024. From what we can tell from a quick skim, we think that our changes largely address Vetted Causes’ concerns here, as well as the detailed feedback we received last year from Giving What We Can (see also our response at the time) as part of their program that evaluates evaluators. Our cost-effectiveness analyses no longer use achievement or intervention scores, but rather directly calculate cost-effectiveness by dividing impact by cost, as you suggest. That being said, our work will never be perfect so we invite anyone reading this with the expertise to improve the rigor of our work to reach out, now or in the future.
Although your comments are related to methods that we no longer use, we’d like to spend more time understanding and engaging with them, learning from them, and potentially correcting any misconceptions. Unfortunately, we won’t have the opportunity to do so until after our charity recommendations are released next week. Additionally, it might be a comfort to know that for the past few months, Giving What We Can has been assessing ACE’s new evaluation methods along with a panel of other experts and that they intend to publish the results later this month.
we have significantly updated our evaluation methods in the past year and will be publishing the details next Tuesday when we release our charity recommendations for 2024. From what we can tell from a quick skim, we think that our changes largely address Vetted Causes’ concerns here, as well as the detailed feedback we received last year from Giving What We Can (see also our response at the time) as part of their program that evaluates evaluators. Our cost-effectiveness analyses no longer use achievement or intervention scores, but rather directly calculate cost-effectiveness by dividing impact by cost, as you suggest.
We are glad to hear that ACE has changed their evaluation methods, and we hope that the changes effectively address the concerns listed in our review.
We look forward to seeing ACE’s new charity recommendations when they are released next week.
Hi Isaac! Now that we’ve announced our 2024 Recommended Charities, we’ve had more time to process your feedback. Thanks again for engaging with our work.
As mentioned before, we’ve substantively updated our evaluation methods this year. This was informed in part by detailed feedback we received as part of Giving What We Can’s 2023 ‘Evaluating the Evaluators’ project, some of which aligns with your feedback.
One of these changes is that we now seek to conduct more direct cost-effectiveness analyses, rather than the 1-7 scoring method that we used last year. This more direct approach is possible in part thanks to Ambitious Impact’s recent work to allow quantification of animal suffering averted per dollar. Of course, these kinds of calculations are still extremely challenging, limited, and subject to significant uncertainties; we describe our methods and their limitations on our website. For example, while cost-effectiveness = impact divided by cost, it can be difficult to measure impact meaningfully in a way that is also quantifiable, so we rely on other criteria to help us make our assessments.
Another major change was introducing a formal Theory of Change assessment to understand the reasoning, evidence base, and limitations around each charity’s main programs. In our 2023 Evaluations, we discussed these considerations in our Recommendations Decisions meetings but did not systematically incorporate them into our public reviews. Together, we think these changes allow for a more nuanced assessment of charities’ work and (we hope) more informative and accessible reviews.
Regarding the impact of our recommendations, this year, we conducted an assessment of ACE’s programs and our counterfactual influence on funding. As part of this work, we surveyed donors to our Recommended Charity Fund (RCF) and asked them where they’d donate if ACE didn’t exist. This indicated that over 60% of our RCF donors would donate less to animal charities if ACE were not to exist, of whom around 12% would not donate to animal charities at all. We aim to publish these influenced-giving reports on November 29th. We hope this reassures you that animals are not worse off because of ACE’s charity recommendations.
In terms of your specific feedback on last year’s methodology:
‘Charities can receive a worse Cost-Effectiveness Score by spending less money to achieve the exact same results’ / ‘Charities can receive a better Cost-Effectiveness Score by spending more money to achieve the exact same results’ / ‘Charities can rearrange their budget and achieve the exact same results (with the exact same total expenditures), but their Cost-Effectiveness Score can significantly change.’
Your findings here are correct. Because the weighted averages in this model depended on the percentage of expenditure for each factor, they sometimes produced unintended and unhelpful results. In part due to this, we interrogated the outputs of our models in our Recommendation Decisions meetings at the time and considered cost-effectiveness scores alongside other decision-relevant factors (such as their Impact Potential and Room For More Funding), rather than taking cost-effectiveness as the only relevant factor to consider when evaluating charities or prioritizing giving opportunities. This was informed in part by each charity’s uncertainty scores, which helped inform how much weight to assign to cost-effectiveness and other criteria in our final recommendations decisions. As you would expect given their work, Legal Impact for Chickens’ uncertainty scores were among the highest of our 2023 evaluated charities. Of all our 2023 Recommended Charities, our Recommendations Decisions discussions played the biggest role for Legal Impact for Chickens given that our models were not as well-suited to their work compared to those for other charities.
How we addressed this in our 2024 Evaluations: As noted above, we now do direct cost-effectiveness analysis rather than using a weighted factor model. We think the role of our Theory of Change-focused discussions in our 2023 Recommendation Decisions meetings should have been more systematic and more clearly communicated in our 2023 reviews, which is one reason why we introduced the new Theory of Change assessment this year.
‘Charities can have 1,000,000 times the impact at the exact same price, and their Normalized Achievement Scores and Cost-Effectiveness Score can remain the same.’
This isn’t the case, but we didn’t publish the full details about our method for assessing the impact of books, podcasts, and other interventions, so we see why this wasn’t clear. Essentially for each intervention in our Menu of Interventions we identified proxies for its likely impact. For books, we had intended to include sales/views as well as a rating of the overall audience response/reviews. In practice, this wasn’t possible for various reasons given the wide variation in types of publication (e.g., some publications had not been released yet, or had been provided directly to the audience with no feedback collected), so we had to factor in such considerations on a more case-by-case basis in our Recommendations Decisions discussions. Issues such as this highlighted to us the inherent limitations of seeking to distill a charity’s work in a weighted factor model, given e.g. the large variation in tactics used by animal advocacy charities and the challenges involved in obtaining the necessary data to score their achievements based on pre-set criteria.
We used additive rather than multiplicative scoring because our objective was to create weighted factor models that reflect the quality of achievements (rather than, e.g., estimating the number of animals helped by books being written). Since we’ve transitioned to directly estimating cost-effectiveness, we now use a straightforward multiplication for factors like “likelihood of implementation.”
How we addressed this in our 2024 Evaluations: Same as the point above: we have now updated to a more direct cost-effectiveness analysis rather than using this weighted factor model, and have also introduced a new Theory of Change assessment.
‘Charities can increase their Normalized Achievement Scores and Cost-Effectiveness Score by breaking down actions into smaller steps, even if the overall results remain unchanged.’
This actually isn’t the case (sorry if this wasn’t clear). Breaking down an achievement into smaller steps would drive up the ‘Achievement quantity’ score, but would be offset by lower ‘Achievement quality’ scores for each achievement. However, there was still a risk of this introducing inconsistency into the model, which is another reason why we updated our methods this year.
‘The most important factor in determining the Normalized Achievement Score of an intervention (Impact Potential Score) is decided before the intervention even begins. This makes the maximum Normalized Achievement Score for certain interventions relatively low, even if they have extremely high impact.’
We developed this model because past evaluations have shown that the intervention type drives much of the impact of a charity’s achievements. Starting with a baseline intervention score and adjusting it still allows for particularly strong implementations to at least partially make up for a lower intervention score. That said, we agree with you on this model’s shortcomings. As with the cost-effectiveness model, we interrogated the model’s outputs in our Recommendation Decisions meetings and had a mechanism to weight Impact Potential lower in our decision-making when we were less certain about its relevance.
How we addressed this in our 2024 Evaluations: We have updated this model and now only use it in a very limited way to supplement a qualitative assessment of charities’ work during the charity selection phase, rather than during the Evaluations themselves.
‘Legal Impact for Chickens did not achieve any favorable legal outcomes, yet ACE rated them a Recommended Charity.’
When ACE considers impact for animals, we consider all the ways that animal suffering might be reduced when interventions are implemented. While they did not secure a litigation win, Legal Impact for Chickens’ Costco lawsuit garnered significant media attention that put pressure on the companies being litigated. While their work is more ‘hits-based’ than some of our other Recommended Charities, we think the considerable impact of any future legal wins means high expected value for this work overall, especially now that funding from ACE, Open Philanthropy, and the EA Animal Welfare Fund has allowed them to hire more litigators. Check out Alene Anello’s recent EA Forum post for an update on Legal Impact for Chickens’ latest achievements.
Thanks again for your engagement with our evaluations. We hope you get in touch with us directly if you come across new evidence-based methods to meaningfully capture cost-effectiveness or to improve the evaluation of animal charities. We might also reach out to you via email in the coming weeks as we go through retrospectives and plan for next year’s evaluation. Because of the complexity of the animal welfare cause area, the many uncertainties and knowledge gaps in the field of charity evaluation, and the urgency and scope of suffering, we embrace productive collaboration.
Thank you for taking the time to read our review and for responding to each of our points. We really appreciate ACE’s willingness to engage with feedback and acknowledge problems.
Regarding your clarifications related to the calculation of Normalized Achievement Scores:
‘Charities can have 1,000,000 times the impact at the exact same price, and their Normalized Achievement Scores and Cost-Effectiveness Score can remain the same.’
This isn’t the case, but we didn’t publish the full details about our method for assessing the impact of books, podcasts, and other interventions, so we see why this wasn’t clear. Essentially for each intervention in our Menu of Interventions we identified proxies for its likely impact. For books, we had intended to include sales/views as well as a rating of the overall audience response/reviews. In practice, this wasn’t possible for various reasons given the wide variation in types of publication (e.g., some publications had not been released yet, or had been provided directly to the audience with no feedback collected), so we had to factor in such considerations on a more case-by-case basis in our Recommendations Decisions discussions.
We are glad to hear that ACE was accounting for these factors behind the scenes.
‘Charities can increase their Normalized Achievement Scores and Cost-Effectiveness Score by breaking down actions into smaller steps, even if the overall results remain unchanged.’
This actually isn’t the case (sorry if this wasn’t clear). Breaking down an achievement into smaller steps would drive up the ‘Achievement quantity’ score, but would be offset by lower ‘Achievement quality’ scores for each achievement. However, there was still a risk of this introducing inconsistency into the model, which is another reason why we updated our methods this year.
Thank you for clarifying this. From the publicly available rubrics for calculating Achievement Quality Scores, it did not seem like breaking down an achievement into smaller steps would decrease the Achievement Quality Score at all. However, given that ACE was accounting for factors outside of the publicly available rubrics, it makes sense that this decrease could occur.
That being said, we believe it is important for ACE to fully disclose its methodology to the public and avoid relying on hidden evaluation criteria. This transparency would allow people from outside the organization to understand how ACE’s charity evaluation metrics (i.e. Normalized Achievement Scores) were calculated.
We might also reach out to you via email in the coming weeks as we go through retrospectives and plan for next year’s evaluation. Because of the complexity of the animal welfare cause area, the many uncertainties and knowledge gaps in the field of charity evaluation, and the urgency and scope of suffering, we embrace productive collaboration.
We appreciate your openness to collaboration. Feel free to reach out to us at any time at hello@vettedcauses.com
Thank you for spending time analyzing our methods. We appreciate those who are willing to engage with our work and help us improve the accuracy of our recommendations and reduce animal suffering as much as possible.
Based on previously received feedback and internal reflection, we have significantly updated our evaluation methods in the past year and will be publishing the details next Tuesday when we release our charity recommendations for 2024. From what we can tell from a quick skim, we think that our changes largely address Vetted Causes’ concerns here, as well as the detailed feedback we received last year from Giving What We Can (see also our response at the time) as part of their program that evaluates evaluators. Our cost-effectiveness analyses no longer use achievement or intervention scores, but rather directly calculate cost-effectiveness by dividing impact by cost, as you suggest. That being said, our work will never be perfect so we invite anyone reading this with the expertise to improve the rigor of our work to reach out, now or in the future.
Although your comments are related to methods that we no longer use, we’d like to spend more time understanding and engaging with them, learning from them, and potentially correcting any misconceptions. Unfortunately, we won’t have the opportunity to do so until after our charity recommendations are released next week. Additionally, it might be a comfort to know that for the past few months, Giving What We Can has been assessing ACE’s new evaluation methods along with a panel of other experts and that they intend to publish the results later this month.
Thank you.
- The ACE team
Hi,
Thank you for your response!
We are glad to hear that ACE has changed their evaluation methods, and we hope that the changes effectively address the concerns listed in our review.
We look forward to seeing ACE’s new charity recommendations when they are released next week.
Hi Isaac! Now that we’ve announced our 2024 Recommended Charities, we’ve had more time to process your feedback. Thanks again for engaging with our work.
As mentioned before, we’ve substantively updated our evaluation methods this year. This was informed in part by detailed feedback we received as part of Giving What We Can’s 2023 ‘Evaluating the Evaluators’ project, some of which aligns with your feedback.
One of these changes is that we now seek to conduct more direct cost-effectiveness analyses, rather than the 1-7 scoring method that we used last year. This more direct approach is possible in part thanks to Ambitious Impact’s recent work to allow quantification of animal suffering averted per dollar. Of course, these kinds of calculations are still extremely challenging, limited, and subject to significant uncertainties; we describe our methods and their limitations on our website. For example, while cost-effectiveness = impact divided by cost, it can be difficult to measure impact meaningfully in a way that is also quantifiable, so we rely on other criteria to help us make our assessments.
Another major change was introducing a formal Theory of Change assessment to understand the reasoning, evidence base, and limitations around each charity’s main programs. In our 2023 Evaluations, we discussed these considerations in our Recommendations Decisions meetings but did not systematically incorporate them into our public reviews. Together, we think these changes allow for a more nuanced assessment of charities’ work and (we hope) more informative and accessible reviews.
Regarding the impact of our recommendations, this year, we conducted an assessment of ACE’s programs and our counterfactual influence on funding. As part of this work, we surveyed donors to our Recommended Charity Fund (RCF) and asked them where they’d donate if ACE didn’t exist. This indicated that over 60% of our RCF donors would donate less to animal charities if ACE were not to exist, of whom around 12% would not donate to animal charities at all. We aim to publish these influenced-giving reports on November 29th. We hope this reassures you that animals are not worse off because of ACE’s charity recommendations.
In terms of your specific feedback on last year’s methodology:
‘Charities can receive a worse Cost-Effectiveness Score by spending less money to achieve the exact same results’ / ‘Charities can receive a better Cost-Effectiveness Score by spending more money to achieve the exact same results’ / ‘Charities can rearrange their budget and achieve the exact same results (with the exact same total expenditures), but their Cost-Effectiveness Score can significantly change.’
Your findings here are correct. Because the weighted averages in this model depended on the percentage of expenditure for each factor, they sometimes produced unintended and unhelpful results. In part due to this, we interrogated the outputs of our models in our Recommendation Decisions meetings at the time and considered cost-effectiveness scores alongside other decision-relevant factors (such as their Impact Potential and Room For More Funding), rather than taking cost-effectiveness as the only relevant factor to consider when evaluating charities or prioritizing giving opportunities. This was informed in part by each charity’s uncertainty scores, which helped inform how much weight to assign to cost-effectiveness and other criteria in our final recommendations decisions. As you would expect given their work, Legal Impact for Chickens’ uncertainty scores were among the highest of our 2023 evaluated charities. Of all our 2023 Recommended Charities, our Recommendations Decisions discussions played the biggest role for Legal Impact for Chickens given that our models were not as well-suited to their work compared to those for other charities.
How we addressed this in our 2024 Evaluations: As noted above, we now do direct cost-effectiveness analysis rather than using a weighted factor model. We think the role of our Theory of Change-focused discussions in our 2023 Recommendation Decisions meetings should have been more systematic and more clearly communicated in our 2023 reviews, which is one reason why we introduced the new Theory of Change assessment this year.
‘Charities can have 1,000,000 times the impact at the exact same price, and their Normalized Achievement Scores and Cost-Effectiveness Score can remain the same.’
This isn’t the case, but we didn’t publish the full details about our method for assessing the impact of books, podcasts, and other interventions, so we see why this wasn’t clear. Essentially for each intervention in our Menu of Interventions we identified proxies for its likely impact. For books, we had intended to include sales/views as well as a rating of the overall audience response/reviews. In practice, this wasn’t possible for various reasons given the wide variation in types of publication (e.g., some publications had not been released yet, or had been provided directly to the audience with no feedback collected), so we had to factor in such considerations on a more case-by-case basis in our Recommendations Decisions discussions. Issues such as this highlighted to us the inherent limitations of seeking to distill a charity’s work in a weighted factor model, given e.g. the large variation in tactics used by animal advocacy charities and the challenges involved in obtaining the necessary data to score their achievements based on pre-set criteria.
We used additive rather than multiplicative scoring because our objective was to create weighted factor models that reflect the quality of achievements (rather than, e.g., estimating the number of animals helped by books being written). Since we’ve transitioned to directly estimating cost-effectiveness, we now use a straightforward multiplication for factors like “likelihood of implementation.”
How we addressed this in our 2024 Evaluations: Same as the point above: we have now updated to a more direct cost-effectiveness analysis rather than using this weighted factor model, and have also introduced a new Theory of Change assessment.
‘Charities can increase their Normalized Achievement Scores and Cost-Effectiveness Score by breaking down actions into smaller steps, even if the overall results remain unchanged.’
This actually isn’t the case (sorry if this wasn’t clear). Breaking down an achievement into smaller steps would drive up the ‘Achievement quantity’ score, but would be offset by lower ‘Achievement quality’ scores for each achievement. However, there was still a risk of this introducing inconsistency into the model, which is another reason why we updated our methods this year.
‘The most important factor in determining the Normalized Achievement Score of an intervention (Impact Potential Score) is decided before the intervention even begins. This makes the maximum Normalized Achievement Score for certain interventions relatively low, even if they have extremely high impact.’
We developed this model because past evaluations have shown that the intervention type drives much of the impact of a charity’s achievements. Starting with a baseline intervention score and adjusting it still allows for particularly strong implementations to at least partially make up for a lower intervention score. That said, we agree with you on this model’s shortcomings. As with the cost-effectiveness model, we interrogated the model’s outputs in our Recommendation Decisions meetings and had a mechanism to weight Impact Potential lower in our decision-making when we were less certain about its relevance.
How we addressed this in our 2024 Evaluations: We have updated this model and now only use it in a very limited way to supplement a qualitative assessment of charities’ work during the charity selection phase, rather than during the Evaluations themselves.
‘Legal Impact for Chickens did not achieve any favorable legal outcomes, yet ACE rated them a Recommended Charity.’
When ACE considers impact for animals, we consider all the ways that animal suffering might be reduced when interventions are implemented. While they did not secure a litigation win, Legal Impact for Chickens’ Costco lawsuit garnered significant media attention that put pressure on the companies being litigated. While their work is more ‘hits-based’ than some of our other Recommended Charities, we think the considerable impact of any future legal wins means high expected value for this work overall, especially now that funding from ACE, Open Philanthropy, and the EA Animal Welfare Fund has allowed them to hire more litigators. Check out Alene Anello’s recent EA Forum post for an update on Legal Impact for Chickens’ latest achievements.
Thanks again for your engagement with our evaluations. We hope you get in touch with us directly if you come across new evidence-based methods to meaningfully capture cost-effectiveness or to improve the evaluation of animal charities. We might also reach out to you via email in the coming weeks as we go through retrospectives and plan for next year’s evaluation. Because of the complexity of the animal welfare cause area, the many uncertainties and knowledge gaps in the field of charity evaluation, and the urgency and scope of suffering, we embrace productive collaboration.
Thank you.
- The ACE team
Hi,
Thank you for taking the time to read our review and for responding to each of our points. We really appreciate ACE’s willingness to engage with feedback and acknowledge problems.
Regarding your clarifications related to the calculation of Normalized Achievement Scores:
We are glad to hear that ACE was accounting for these factors behind the scenes.
Thank you for clarifying this. From the publicly available rubrics for calculating Achievement Quality Scores, it did not seem like breaking down an achievement into smaller steps would decrease the Achievement Quality Score at all. However, given that ACE was accounting for factors outside of the publicly available rubrics, it makes sense that this decrease could occur.
That being said, we believe it is important for ACE to fully disclose its methodology to the public and avoid relying on hidden evaluation criteria. This transparency would allow people from outside the organization to understand how ACE’s charity evaluation metrics (i.e. Normalized Achievement Scores) were calculated.
We appreciate your openness to collaboration. Feel free to reach out to us at any time at hello@vettedcauses.com