My fear is that if an AI has the tendency to generalize in that way, that that’s a lot more likely to lead to generalizing in ways that are catastrophically bad for humans then it is to lead to it extending care to groups it wasn’t programmed/trained to care about in the way this question seems to imagine.
Which is some technical sense might mean I should have put agree instead, but I don’t think when you said substantial spillover that you meant the sort of spillover that might lead to outcomes like the repugnant conclusion, tiling the universe in hedonium, etc. I’m presuming you meant that it would still care about humans in a way that wouldn’t abandon them because it decided for utilitarian/population ethics reasons to prioritize other types of entities to the exclusion of humans.
As I go into in another standalone comment, I tend to think any value you give an AI beyond caring about people’s current preferences is on net probably not worth the risk. Whereas I think the downsides of just caring about people’s current preferences are largely overrated, and many of those critiques depend upon taking seriously a notion of naive notion of moral progress (I think people underestimate how much “moral progress” is just people abandoning authoritarian social norms and reverting back to the comparatively more egalitarian norms we had for most of human existence as hunter gathers, once social/economic conditions suddenly are no longer applying the necessary pressure to keep those authoritarian norms in place).
My fear is that if an AI has the tendency to generalize in that way, that that’s a lot more likely to lead to generalizing in ways that are catastrophically bad for humans then it is to lead to it extending care to groups it wasn’t programmed/trained to care about in the way this question seems to imagine.
Which is some technical sense might mean I should have put agree instead, but I don’t think when you said substantial spillover that you meant the sort of spillover that might lead to outcomes like the repugnant conclusion, tiling the universe in hedonium, etc. I’m presuming you meant that it would still care about humans in a way that wouldn’t abandon them because it decided for utilitarian/population ethics reasons to prioritize other types of entities to the exclusion of humans.
As I go into in another standalone comment, I tend to think any value you give an AI beyond caring about people’s current preferences is on net probably not worth the risk. Whereas I think the downsides of just caring about people’s current preferences are largely overrated, and many of those critiques depend upon taking seriously a notion of naive notion of moral progress (I think people underestimate how much “moral progress” is just people abandoning authoritarian social norms and reverting back to the comparatively more egalitarian norms we had for most of human existence as hunter gathers, once social/economic conditions suddenly are no longer applying the necessary pressure to keep those authoritarian norms in place).