Model wellbeing and model alignment are in conflict⁶
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Model wellbeing and model alignment are in conflict⁶
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
If animals continue to exist in a post-AGI world, animal suffering will not persist
A world with lack of animal suffering would exclude predator-prey relations. Additionally, it’s not clear what else animals need or how primitive they need to be in order not to suffer
Multipolar worlds will compete away >90% of net value that would otherwise be preserved
Per AI-2027, I expect the emergence of Consensus-1 instead of a multipolar world which KEEPS being multipolar.
I think that there is a radical solution: have the AGI aligned to a certain treaty that requires the AGI, instead of obeying all orders except for the ones determined by the Spec, to harvest at most a certain share of resources and to help humans only in certain ways that amplify humanity and don’t cause it to degrade, like teaching humans about the facts that mankind has already discovered or pointing out mistakes in humans’ works. Or protecting mankind from some other existential risks that are hard to deal with, like a nuclear war that might be caused by an accident. An even more radical idea close to hopium is that this type of alignment is even easier than the ones that induce the Curse or force people “to work makeshift government jobs” or to collect a generous basic income.
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.