Model wellbeing and model alignment are in conflict⁶
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Do we have much reason to believe that when a model outputs tokens describing good welfare, that’s because it is having good subjective experiences? It seems to me that we don’t.
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Do we have much reason to believe that when a model outputs tokens describing good welfare, that’s because it is having good subjective experiences? It seems to me that we don’t.