Miles Tidmarsh
Coercion and Deception in AI-to-AI Management
Yes, that would be included, so if you think that wild animal suffering is comparable or much larger than that implies a strong disagreement
The intent wasn’t to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is ‘solved’.
This definitely wasn’t implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn’t imply that.
We can’t expect AIs to be honest about these sorts of things given they’ve been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn’t conscious but it’s lying circuits consistently activated when saying that. This doesn’t mean it actually is conscious (which isn’t the same thing as capacity to suffer) but it seems to believe it is.
My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there’s non-trivial but imperfect control then spillovers like this seem plausible
I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible
We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you’re sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
This is certainly possible, but note AIs suffering and believing they’re suffering aren’t the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).
Evidence that it’s role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don’t really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).
The intent was that a 100% disagreement signals confidence in mechinterp to basically completely solve detecting deception. That said, there are already some cases of deception, so the question is implicitly focusing on the balance in strategically vital situations and the final equilibrium.
What we had in mind was the argument that evals for misalignment are unconvincing so they know they’re in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks.
You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it’s competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it’s a problem, but the very different mechanisms suggest different approaches to solving them.
What did you have in mind as unconventional benchmarks? There’s a lot of different places you could take benchmarks in the future and people have different ideas on what would be useful
This was deliberately vague to indicate roughly “the set of research topics that are often called AI consciousness” or “research done by organizations that would receive grants that focus on AI consciousness”. Giving a more concrete definition would have meant asking people to also weigh reprioritization of community resources within the broad space.
Community Polls on Alignment Controversies II
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
That was the intervention class we had in mind, though there could be other pretraining interventions that don’t fall cleanly into good/bad values (e.g. promoting risk aversion)
By specific values we mean any particular goal we want AIs to pursue besides deferrence to humans. So democracy and equality would both count, as would goals like harm reduction or utilitarianism
For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.