Most current evidence of misalignment is actually models role-playing a misaligned AI⁴
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
What we had in mind was the argument that evals for misalignment are unconvincing so they know they’re in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks.
You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it’s competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it’s a problem, but the very different mechanisms suggest different approaches to solving them.
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
What we had in mind was the argument that evals for misalignment are unconvincing so they know they’re in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks.
You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it’s competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it’s a problem, but the very different mechanisms suggest different approaches to solving them.