Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month’s polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We’ll compare panel and community responses in an upcoming report, which we’ll publish here and on LessWrong. To get notified when it’s released, you can subscribe to our new Substack.
Many people we’ve talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML’s research agenda.
A few final things about the polls themselves:
Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
We’re not trying to find the ‘right’ answers. Please answer based on your own best guess.
% agree is your % credence in a given position
Please let us know if you think the questions are ambiguous or embed false assumptions
Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ “Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
8 AIs that are trying to hide features of themselves from humans and operators
Community Polls on Alignment Controversies II
Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month’s polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We’ll compare panel and community responses in an upcoming report, which we’ll publish here and on LessWrong. To get notified when it’s released, you can subscribe to our new Substack.
Many people we’ve talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML’s research agenda.
A few final things about the polls themselves:
Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
We’re not trying to find the ‘right’ answers. Please answer based on your own best guess.
% agree is your % credence in a given position
Please let us know if you think the questions are ambiguous or embed false assumptions
Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ “Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
8 AIs that are trying to hide features of themselves from humans and operators