Do we have much reason to believe that when a model outputs tokens describing good welfare, that’s because it is having good subjective experiences? It seems to me that we don’t.
Do we have much reason to believe that when a model outputs tokens describing good welfare, that’s because it is having good subjective experiences? It seems to me that we don’t.