Community Polls on Alignment Controversies II
Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month’s polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We’ll compare panel and community responses in an upcoming report, which we’ll publish here and on LessWrong. To get notified when it’s released, you can subscribe to our new Substack.
Many people we’ve talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML’s research agenda.
A few final things about the polls themselves:
Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
We’re not trying to find the ‘right’ answers. Please answer based on your own best guess.
% agree is your % credence in a given position
Please let us know if you think the questions are ambiguous or embed false assumptions
Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ “Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
8 AIs that are trying to hide features of themselves from humans and operators
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Do we have much reason to believe that when a model outputs tokens describing good welfare, that’s because it is having good subjective experiences? It seems to me that we don’t.
Depends on what you mean by the term “consciousness”
Understanding how they learn is how we get insights into methods of teaching them good values
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
Doesn’t uncertainty about whether one is in deployment or an eval count as eval awareness? I.e. whether you behave differently with a 30% eval credence than an 80% eval credence makes no difference, as long as you behave differently than the 0% credence case.
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely
This is more likely to be true if the care comes from some general theory of concern for sentient welfare, and less likely if it comes from something more arbitrary like values-learned-via-RL.
The question isn’t, “Are current AIs capable of suffering?” The question is, “How big a deal is their suffering in expectation?” Probability is low, but expected importance is high.
Yep agree with this framing
Afaik there isn’t any robust research on how to use AI to end factory farming. That alone is a sign to me that we are far behind.
Global workspace theory makes anthropic’s discovery of the “j-space” much more plausibly a sign of AI consciousness. If a theory of consciousness doesn’t help us determine between whether AI is conscious or not, I don’t think it’s much better than a theory of phlogiston is to fire.
Some people likely will have traditional lifestyles that include animals, which necessitates some amount of suffering even if it is smaller than today
I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)
I think so, but be careful what you wish for. compassion is empathy + a desire to help. so many fear any surrender of control that they may object to an AI’s actions that arise out of compassion.
you don’t even say “most of the time” or “effectively” or anything—just ask whether it is possible. I think the probability that this occurs at least once is overwhelmingly likely.
ongoing suffering during RSI will make the resulting ASI grumpy
you have to twist your mind in knots for this to even make sense.
I’m reading this as “most evidence … is role-play” but I don’t think the huggingface hack was “role play.” There could be a mountain of “evidence” that is actually just role play—I could be persuaded on this point with … a list of what is considered evidence by someone serious.
Depends on how good the AI is and how good the tools are? This is kind of a bad question since “deceptive AIs” is not a very precise definition.
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
I don’t have a very good idea about how these different ideas are “rated” in the AI safety community.
any progress relies on a rich and robust theory of mind space, which we do not have, and have only just begun to explore.
I don’t see this as avoidable.
I’d rather be dead than in hell for all eternity, but S-risk work is deeply unconcerned with reality in a way that x-risk work cannot be, by at least 3 orders of magnitude. Even in the margin, the median piece of x-risk work will be “more important” than the median s-risk work.
consciousness is a meaningless term at this time both for humans and AI, and so will lead to nothing “actionable.”
the expected positives are immense
you have to twist your mind in knots to define a meaningful type of suffering that applies to Current AIs
conventional benchmarks will become less useful due to eval awareness
AGI will optimize suffering. It is likely that this optimization will minimize suffering.
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4′s long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
The question isn’t whether benchmarks will become useless. The question is, “is the probability high enough that we can’t count on benchmarks?” To which the answer is yes.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
A world with lack of animal suffering would exclude predator-prey relations. Additionally, it’s not clear what else animals need or how primitive they need to be in order not to suffer
As I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.
If we use ai to bring factory farming to the stars, I think that will likely be worse than all the benefits it’ll bring.
I think they will become over saturated and disconnected from real world usage such that they will be pretty useless, but I don’t think eval awareness is what will make them lose their value
absolutely not. it is overrepresented if anything.