>By “satisfaction” I meant high performance on its mesa-objective
Yeah, I’d agree with this definition.
I don’t necessarily agree with your two points of skepticism, for the first one I’ve already mentioned my reasons, for the second one it’s true in principle but it seems almost anything an AI would learn semi-accidentally is going to be much simpler and more intrinsically consistent than human values. But low confidence on both and in any case that’s kind of beyond the point, I was mostly trying to understand your perspective on what utility is.
>By “satisfaction” I meant high performance on its mesa-objective
Yeah, I’d agree with this definition.
I don’t necessarily agree with your two points of skepticism, for the first one I’ve already mentioned my reasons, for the second one it’s true in principle but it seems almost anything an AI would learn semi-accidentally is going to be much simpler and more intrinsically consistent than human values. But low confidence on both and in any case that’s kind of beyond the point, I was mostly trying to understand your perspective on what utility is.