Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.