One statistical/methodological point I’d add (something I always harp on). I don’t think “not statistically significant” should directly cited as evidence for a lack of a difference. If the question is whether mortality differences are small enough to be decision-irrelevant, we’d want something closer to an equivalence test or Bayesian posterior over the mortality difference, plus a welfare model translating mortality causes, morbidity, behavioral deprivation, fear/stress, and transition dynamics into aggregate welfare burden.
A forest plot or explicit meta-analytic summary could also make the cage-free evidence easier to interpret than the table of pairwise significance checks.
Related: I wouldn’t always treat small sample sizes or mixed statistical significance as automatically implying “no useful inference.” Small-N studies can be informative if underlying measurement noise is low. For example if I ask 4 people to taste a drink and they all wince deeply in pain and disgust, I’m going to be highly confident it tastes bad. If all 4 smile and praise it, I’ll be fairly confident that it’s at least tolerable.
One statistical/methodological point I’d add (something I always harp on). I don’t think “not statistically significant” should directly cited as evidence for a lack of a difference. If the question is whether mortality differences are small enough to be decision-irrelevant, we’d want something closer to an equivalence test or Bayesian posterior over the mortality difference, plus a welfare model translating mortality causes, morbidity, behavioral deprivation, fear/stress, and transition dynamics into aggregate welfare burden.
A forest plot or explicit meta-analytic summary could also make the cage-free evidence easier to interpret than the table of pairwise significance checks.
Related: I wouldn’t always treat small sample sizes or mixed statistical significance as automatically implying “no useful inference.” Small-N studies can be informative if underlying measurement noise is low. For example if I ask 4 people to taste a drink and they all wince deeply in pain and disgust, I’m going to be highly confident it tastes bad. If all 4 smile and praise it, I’ll be fairly confident that it’s at least tolerable.
McElreath’s globe-tossing example illustrates how much we can sometimes learn from small samples.
(Still, in the shrimp case, it does seem like there is some substantial underlying variation unrelated to the different slaughter methods.)