How Can We Solve Animal Welfare’s Evidence Problem? (Post 1)

TL;DR:
The movement has an evidence problem. To solve it, for each intervention we need to know:

  • What types of evidence that intervention needs

  • What types of evidence different sources can give us

In this post, we outline a framework for evidencing an intervention through five stages:

Figure 1: the five stages of the evidence pipeline

Figure 1: The five stages of the evidence pipeline.

Each stage poses a different question, needing different types of evidence, and thus different methods (and potentially different organisations and experts). A follow-up post will cover practical next steps.

1. Introduction

Two months ago, Matthes argued that animal welfare has an evidence problem. Discussion mostly focused on the specific examples she gave. Her wider claim received less attention: that the movement lacks enough well-evidenced interventions, and that we need to take ownership of the entire evidence pipeline to change that.

Discussion so far has centred on interventions that are highly invested in, and whether their evidence is strong or weak (the top left of Figure 2). However, Most interventions are neither invested in nor well-evidenced. We want to add a second discussion point: how do we get more interventions to be well-evidenced (ideally, the top right of Figure 2)?

Figure 2 · movement investment against strength of evidence

Figure 2: Interventions by movement investment and strength of evidence.

We agree with Matthes that there is an evidence problem. To solve it, we need to rethink how evidence generation works within the animal movement. We need to consider:

  1. What evidence do we need?

  2. What would a movement capable of generating it look like?

This post is part one of a two-part response to Matthes’ post. Here, we set out to answer the first question: To define what types of evidence we need and what types of evidence different sources can give us. Part two covers the practical next steps.

2. What Do You Mean We Have An Evidence Problem?

Responding to Matthes, Krzysztof Wojtas argued that evidence-building is not “a linear process”, with research finishing before implementation starts. Field implementation is itself part of the learning, and gives different types of evidence that are equally needed.

We agree, and the point exposes a useful distinction: Evidence is not a monolith. Each intervention raises many different questions. Each question calls for its own kind of evidence, and each kind of evidence has its own best source.

Implementation is one of those sources, and it answers some questions well and others poorly. Let’s imagine an intervention helping fish in Brazil through improved biosecurity. Here are four questions it might need to answer, and whether its own monitoring data can answer them:

The questionWhat would answer it?Can monitoring data answer it?
Can farmers implement biosecurity practices in their context?Records from the farms in the programmeYes.
How common are disease outbreaks across Brazil?A representative survey of farmsNo. The farms in the programme are not a representative sample.
Does the intervention reduce disease outbreaks?A comparison of disease rates between fish that received the intervention (a treatment group) and those that didn’t (a control group)Potentially. Depending on whether a control group was also being monitored.
How much suffering does the intervention avert?A comparison of welfare indicators (e.g., cortisol levels) between a treatment group and a control group.No. Precisely measuring welfare typically requires sophisticated methods executed by experts, ideally within a controlled setting.

So, monitoring data is an important part of an intervention’s evidence sources, but it cannot be expected to fill all evidence gaps. Nor can implementers be expected to resolve all evidence gaps in an intervention. To solve the evidence gap, we will need diverse types of evidence and a diverse set of skills.

The movement’s evidence problem, then, is actually composed of many smaller evidence problems, each a kind of question the movement needs to be able to answer about its interventions. To solve it, we need to understand:

  • 1a) What types of evidence an intervention needs

  • 1b) What types of evidence different sources can give us

3. The Evidence Pipeline

So what are the questions the movement typically needs to answer about its interventions? There are many ways to break that down. Here is our first attempt at a framework based on systems from tech development, clinical health, and global development:[1]

Figure 3: the evidence pipeline, five stages with cross-dependencies

Figure 3: The evidence pipeline. Each stage answers a different question about an intervention. As we invest in answering these questions for an intervention (y-axis), the strength of evidence (x-axis) increases.[2] The stages do not run in a strict order: findings at any stage can reshape the others.

These five questions trace an intervention from idea to a well-evidenced scaled programme. The steps build on each other, and there is a rough logical flow from 1 to 5. But in practice the stages typically overlap, and later stages can be valuable in informing earlier stage work.

The unit moving through the pipeline is an intervention, not an organisation. An organisation may enter at any stage, including the last, so long as the earlier questions have been answered by someone.

Stage 1: Relevance

Is this the right problem to solve?

Relevance asks what welfare problems exist, and which problem is the right target in the chosen context. Left unanswered, we risk solving the wrong problem.

For direct farm animal welfare work, relevance often includes reviewing the welfare conditions for animals and the constraints of the producer. For advocacy or behaviour change, it often includes understanding what motivates and constrains the target stakeholders. Relevance is diagnostic—the aim is to characterise a situation, not to test a solution.

Suitable methods include needs assessments, burden estimates, and political or market feasibility analyses. Each method may draw on primary research (structured interviews, focus groups, and surveys) and on secondary research (reviews of existing studies and data).

Skill sets: social and behavioural scientists, ethnographers, market and political analysts, and researchers who know the local language and context.

Stage 2: Efficacy

Can the intervention work in ideal conditions?

Efficacy tests whether an intervention works under controlled conditions (e.g., a tight protocol with expert delivery). A good efficacy trial establishes the ceiling of an intervention’s impact. Left unanswered, we risk spending years on interventions that cannot work.

Suitable methods are controlled experiments. For direct farm animal welfare work, that means research stations, model farms, and laboratories. For advocacy or behaviour change, it means message and framing tests, online randomised experiments, lab studies, or experiments that isolate specific mechanisms.

Skill sets: animal welfare scientists, veterinary researchers, engineers, experimental psychologists, and statisticians (who help design and analyse trials).

Stage 3: Effectiveness

Does the intervention work in the real world?

Effectiveness again asks whether an intervention works, but this time under the real-world conditions it will run in. Left unanswered, we risk spending years on interventions that do not work.

For direct farm animal welfare work, effectiveness may mean implementing on real farms with ordinary staff who only half-remember the welfare protocols. For behavioural and social change work, it may mean people’s real choices many weeks later and without a researcher watching them. Effectiveness is tested through implementation, and so this is where actual impact for animals may start.

The major challenge for effectiveness is establishing the counterfactual. To say something created change, you need to know what would have happened if it had not been implemented. Thus, a programme’s monitoring data (which generally only records what happened where the intervention was being run) can struggle to establish effectiveness.

Where a control group is possible, suitable methods compare a group (of farms, consumers, companies) receiving the intervention against a similar group that does not. A control group is not always possible, for practical or ethical reasons, and advocacy work is most often evaluated with non-experimental designs. Those designs may still be rigorous when built carefully, and choosing among them is an evaluator’s job.[3]

A process evaluation often runs before or alongside effectiveness research. It establishes whether the intervention was delivered and adopted as intended. This is important because bad implementation of a process can make a good intervention look ineffective. Also, more qualitative observations and a wider set of outcome variables may surface side effects of an intervention.

Skill sets: impact evaluators, development economists, implementation researchers, and applied social scientists, working in collaboration with implementers.

Stage 4: Initial Scaling

Does the intervention keep working at scale?

Initial scaling asks whether an intervention still works, and at what cost, once it is being implemented and is growing. Left unanswered, we risk letting a promising pilot’s impact fade away at scale.

Does impact hold after the founders leave, or you don’t have a relationship with every farmer, when the most amenable targets are exhausted, or once you cannot spend a week training every volunteer? Development economists find that effects reliably shrink as programmes grow, sometimes called the voltage effect.[4] Something that worked on five farms has to be checked on five hundred. This stage is where implementation has fully taken off. It is the last barrier before an intervention counts as well-evidenced and ready for mass adoption.

Suitable methods include coverage, cost, and fidelity evaluations (fidelity being the degree to which an intervention is delivered as intended), measured against the equivalent evidence from before scaling. If a programme has changed substantially at scale, for example, an in-person humane education programme moving online, it may need reassessing for effectiveness.

Skill sets: implementation researchers, evaluators, programme designers, and cost-effectiveness analysts.

Stage 5: Mass Adoption

Can the industry/​society be moved?

Mass adoption is a different kind of question. The intervention is already well-evidenced and implemented well at (some level of) scale. The question becomes whether it can move from something a few organisations run to something that changes an entire industry or system. Left unanswered, we risk an excellent intervention running in a handful of places while the industry carries on unchanged.

For farm animal welfare work, mass adoption might mean a global industry norm for a welfare standard. The Open Wing Alliance is an example of an organisation built for this stage. For behavioural and social change, mass adoption might mean a population-wide shift in purchasing habits or beliefs.

Many organisations hope to eventually create mass adoption, which is a worthy goal. However, we would caution against prioritising mass adoption prior to establishing a programme as well-evidenced (as per the first four stages). Skipping the first four stages may mean scaling an intervention that does not work.

Matthes’ Examples & Cross-Dependencies

Framed in our terms, we would diagnose Matthes’ three main examples as such:

  1. Shrimp stunning moved into effectiveness before efficacy had been established.

  2. Cage-free campaigning moved into mass adoption before effectiveness had been established.

  3. Alternative proteins are failing to prove effective.

There was significant discourse on whether efficacy or effectiveness had been established for each. However we would emphasise (in line with Wojtas’ response) that in reality often the processes for evidencing interventions is messy, and that different stages can learn from each other. Though this does not detract from Matthes’ broader point that the animal movement lacks enough interventions that have successfully gone through the stages to become well-evidenced.

5. Some Ramifications

1. Each stage would ideally have its own specialised organisations.
Currently, much of the burden lands on implementers to evidence their interventions. However, more mature spaces will have people/​organisations that focus on different stages:

Figure 4: who carries each stage, today and in a more mature movement

Figure 4: Who carries the evidence burden at each stage today and in a more mature movement.

No single organisation can run lab tests, implement field trials, scale operations, and run mass adoption campaigns at once. We will, thus, need to figure out how interventions can be handed between organisations as they advance. Until then, implementers will likely continue to carry the burden of evidencing interventions. For as long as that is the case, we should design systems to support them (for example, access to expert support for organisations going through these stages).

2. Interventions at different stages need different evidence.
Any evaluation of a programme/​intervention should use stage-appropriate thresholds. Funders should not ask a programme working at the relevance stage for effectiveness data, and they should only fund effectiveness research if relevance and efficacy have been (somewhat) established.

3. Most ideas should exit the pipeline.
A pipeline that pushes everything through to mass adoption is a conveyor belt. The pipeline should work as a filter. Programmes closing because the evidence did not support the intervention should be expected, and celebrated as movement learning.

4. The skill sets are the bottleneck.
Specialised research needs specialists, which is why each stage above lists its skill sets. Our movement is built largely from implementers and secondary researchers. It is short of development economists, impact evaluators, implementation researchers, engineers, etc.

Part of attaining these skill sets could be outsourcing. Responding to Matthes, Reinstein wrote: “I don’t think EA/​AW people necessarily need to personally design every study, collect the data, or do all the econometrics/​biology. But I do think the community needs to take responsibility for setting the agenda.”

6. Caveats with Our Pipeline

1. Ideas enter at different points.
Relevance research on a US cage-free campaign would add little, since relevance there is established. A programme should consider what the existing evidence base looks like before starting work on the pipeline.

2. It does not cover all research.
For example, there is no clear home here for foundational research such as assessing the sentience of various animal species.

3. It is not universally applicable.
The structure borrows from technology development, clinical health, and global development, so it fits less well the less an intervention has in common with those models. Policy advocacy, systems change, and capacity building may find it fits poorly to their context.
Also, our experience comes mostly from on-farm welfare. Though we have tried to apply these findings to advocacy efforts and social and behavioural change, we very well may be missing things from these angles.
We also recognise that a movement-level framework that defaults to experimental evidence would disadvantage approaches that cannot be tested that way, which is not our intent and something that should be considered seriously as the movement develops.

4. It needs supporting roles we have not described.
The pipeline also depends on enabling actors (such as funders, platforms for sharing findings, those who set expected standards of evidence, etc.). We will get more into these in our next post.

7. Conclusion

The farmed animal movement is asking whether it generates enough well-evidenced interventions, and the consensus is that it does not. Generating that evidence is a process in itself, and one we currently do not have the infrastructure for.

This post offers a framework that we hope can be scaffolding for that infrastructure. We present that framework in five stages: relevance, efficacy, effectiveness, initial scaling, and mass adoption. The stages depend on each other, each needs its own skills, and no single organisation should be expected to hold them all.

Our suggestions here are a first attempt., Our main hope is to prompt more discussion and action on this topic. Our second post will set out the practical next steps we suggest.

Note on the Authors

Tom Billington and Henning Peters are both team members at The Centre for Animal Life and Farming (CALF), a research organisation. We build evidence-based animal welfare interventions for low- and middle-income countries. Many of the ideas above come from our work there, and it should be noted that we have a vested interest in the animal movement directing more resources towards research.

  1. ^

    The pipeline draws on staged models of evidence from clinical health, technology and global development: the translational research spectrum, NASA’s Technology Readiness Levels, Brownson, Colditz and Proctor’s Dissemination and Implementation Research in Health (3rd ed., 2024), the OECD DAC evaluation criteria, IPA’s Stage-Based Learning guide, and the standards of evidence for efficacy, effectiveness and dissemination from prevention science.

  2. ^

    The “Movement Investment” axis is crude and does not directly reflect monetary investment. We might expect some research steps to be significantly more expensive than others.

  3. ^

    See the Advocacy Toolkit companion hosted by Better Evaluation, which notes that non-experimental designs are the most common approach for evaluating advocacy and sets out ways to make them more rigorous.

  4. ^

    John List, The Voltage Effect (2022). See also this overview of the voltage effect.