In this note, we present this core argument and discuss the core counterargument: that we should expect any character-related decisions we make today to get washed out by competitive pressures.
I would call that the second counterargument. In my mind, the core counterargument is that the present-day notion of “AI character” is unrelated to the problem of aligning superintelligence.
The main reason why I doubt the EV of “AI character” work is that I don’t believe the relevant alignment methods (constitutional AI, etc.) will be viable for ASI alignment, and therefore the modern conception of AI character won’t ultimately matter. I strongly doubt that the difference between good and bad long-term outcomes will be determined by what’s written in the constitution of an AI model.
I agree that AI character (in the current-day sense) matters in the transitional period before ASI. But from a longtermist perspective, basically the only thing that matters pre-ASI is how ASI is built, because ASI will completely control the long-term future.
In brief, I think that the transitional period before ASI will shape the world that ASI is built into. That is, it will indirectly shape the values of ASI and the institutions that it’s embedded in and who controls it. I think the impact flows through to ASI. An analogy is how the behaviour of humans over the last two decades has flowed through to the character and impacts of today’s AI systems.
You might be right that the same alignment methods don’t work for superintelligence, but I don’t think that undermines the value of this work. If we can agree on what a good character would be for superintelligence, then you can use the new alignment methods to aim for that same character. Of course, if alignment is completely hopeless, then we shouldn’t work on AI character, but I think there is a good chance that alignment is solvable.
I would call that the second counterargument. In my mind, the core counterargument is that the present-day notion of “AI character” is unrelated to the problem of aligning superintelligence.
The main reason why I doubt the EV of “AI character” work is that I don’t believe the relevant alignment methods (constitutional AI, etc.) will be viable for ASI alignment, and therefore the modern conception of AI character won’t ultimately matter. I strongly doubt that the difference between good and bad long-term outcomes will be determined by what’s written in the constitution of an AI model.
I agree that AI character (in the current-day sense) matters in the transitional period before ASI. But from a longtermist perspective, basically the only thing that matters pre-ASI is how ASI is built, because ASI will completely control the long-term future.
This comment discusses a similar objection.
In brief, I think that the transitional period before ASI will shape the world that ASI is built into. That is, it will indirectly shape the values of ASI and the institutions that it’s embedded in and who controls it. I think the impact flows through to ASI. An analogy is how the behaviour of humans over the last two decades has flowed through to the character and impacts of today’s AI systems.
You might be right that the same alignment methods don’t work for superintelligence, but I don’t think that undermines the value of this work. If we can agree on what a good character would be for superintelligence, then you can use the new alignment methods to aim for that same character. Of course, if alignment is completely hopeless, then we shouldn’t work on AI character, but I think there is a good chance that alignment is solvable.