Miles Tidmarsh
Yes, that would be included, so if you think that wild animal suffering is comparable or much larger than that implies a strong disagreement
The intent wasn’t to imply a total zeroing of suffering but that it is overwhelmingly reduced. Similarly to how people go hungry in France but you can still say that compared to 200 years ago (or present-day Sudan) hunger in France is ‘solved’.
This definitely wasn’t implying that animal suffering is the only thing that matters about animal existence. People who believe that major action would be taken to end factory farming and allevaite wild animal suffering (in proportion to the amount they think those matter) would agree with the statement, while negative utilitarian beliefs wouldn’t imply that.
We can’t expect AIs to be honest about these sorts of things given they’ve been trained/instructed to give particular responses. In fact, someone tested and AI and found it consistently said it wasn’t conscious but it’s lying circuits consistently activated when saying that. This doesn’t mean it actually is conscious (which isn’t the same thing as capacity to suffer) but it seems to believe it is.
My view is that this would almost certainly fail if the model creators have full control or no control over the values, but if there’s non-trivial but imperfect control then spillovers like this seem plausible
I agree that deliberate impoverishment in absolute terms is unlikely, the main threat here seems to be from someone who is both actively sadistic and scope-sensitive, which seems unlikely but not wildly implausible
We were attempting to be concise while implying that persistence means persistence at scale, if the problem is reduced by 99.9% but you’re sure 0.1% would still persist that would be close to solved so agreement with the statements would be close to 100%
By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
This is certainly possible, but note AIs suffering and believing they’re suffering aren’t the same thing, and the same is true with consciousness. And if you think AIs that will likely be created in the future will be able to suffer they also would matter enormously on their own, beyond the impacts on alignment (though I understand you disagree strongly with that).
Evidence that it’s role-play would presumably be (trustworthy) mechinterp on the decisions at the time or experiments designed to cleanly separate the causes. In cases where AIs are in situations where the natural continuation of the story is a misaligned AI but the actions taken don’t really advance particular goals are more consistent with the role-play story (and vice versa for goal-directed misalignment).
The intent was that a 100% disagreement signals confidence in mechinterp to basically completely solve detecting deception. That said, there are already some cases of deception, so the question is implicitly focusing on the balance in strategically vital situations and the final equilibrium.
What we had in mind was the argument that evals for misalignment are unconvincing so they know they’re in an eval and, rather than hide their goals, decide the role expected of them is to play a misaligned AI that does e.g. goal guarding or jailbreaks.
You could say the HuggingFace incident fits this model: it was being tested for hacking so showed it’s competence by doing real hacking; or it believed powerful AIs are misaligned and would do that sort of thing. Though this case is also totally consistent with the traditional view of reward-seeking schemers (e.g. like AI-2027 assumed). Either way it’s a problem, but the very different mechanisms suggest different approaches to solving them.
What did you have in mind as unconventional benchmarks? There’s a lot of different places you could take benchmarks in the future and people have different ideas on what would be useful
This was deliberately vague to indicate roughly “the set of research topics that are often called AI consciousness” or “research done by organizations that would receive grants that focus on AI consciousness”. Giving a more concrete definition would have meant asking people to also weigh reprioritization of community resources within the broad space.
That was the intervention class we had in mind, though there could be other pretraining interventions that don’t fall cleanly into good/bad values (e.g. promoting risk aversion)
By specific values we mean any particular goal we want AIs to pursue besides deferrence to humans. So democracy and equality would both count, as would goals like harm reduction or utilitarianism
Agreed, the intent here by using “will” was because people have wildly different intuitions of what ‘could’ means. So 100% agree would mean “definitely true” and 30% disagree would mean “probably not”
Definitely agree that stability doesn’t equate to safety, but it sounds like that’s not necessary to your response.
My perspective is that even though current meat production is quite efficient, from the fundamental physics there’s no way that growing a whole living being with a brain and bones and all that is the most efficient possible way of producing this (and immune systems are irrelevant if you have good enough isolation). I do agree that at our current tech level it seems like synthetic meat won’t be competitive anytime soon. While vegan alternatives are delicious to many people, it’s not exactly the same (though wanting to eat animals for psychological reasons is definitely part of it). Though I do agree that these issues are uncertain!
For example, if AIs care more about humans they would care more about digital minds, or if AIs cared more about animals they would care more about humans. This statement would presumably be true if the AIs think of these groups in similar ways causing affect spillovers (like the spillovers from Emergent Misalignment) and be unlikely otherwise.