Evidence that it’s role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don’t really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).
I agree in the case where the AI is nothing more than the collection of its personas, and I agree that—to some extent—the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that. That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment.
Evidence that it’s role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don’t really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).
I agree in the case where the AI is nothing more than the collection of its personas, and I agree that—to some extent—the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that. That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment.