I basically agree, although “dangerous approaches are tamped down” is doing most of the work here IMO. By default (i.e. no tamping-down), I expect the situation with a weakly-superhuman Scientist AI to be:
a small number of sane people ask “are we doomed if we go with the following alignment strategy”, and when it says yes, they don’t do it
a lot of people don’t bother to ask at all, they just ask the Scientist AI how to build ASI
a lot of people say “we have to build ASI before the reckless people in group 2”, they build ASI using their best-guess alignment strategy that has an 88.5% chance of failing, and we die with 88.5% probability
(I think Bengio would agree that this is a concern, and would agree that we need global coordination on AI safety to make this work.)
I guess the default for me is that Scientist AI won’t be competitive, so we live in a world with both scientist AI and non-scientist AI. Conditional upon successfully tamping down other approaches enough that Scientist AI gets to the weakly superhuman point while we’re still alive, I’m more optimistic that we can continue to coordinate on doing things safely.
I basically agree, although “dangerous approaches are tamped down” is doing most of the work here IMO. By default (i.e. no tamping-down), I expect the situation with a weakly-superhuman Scientist AI to be:
a small number of sane people ask “are we doomed if we go with the following alignment strategy”, and when it says yes, they don’t do it
a lot of people don’t bother to ask at all, they just ask the Scientist AI how to build ASI
a lot of people say “we have to build ASI before the reckless people in group 2”, they build ASI using their best-guess alignment strategy that has an 88.5% chance of failing, and we die with 88.5% probability
(I think Bengio would agree that this is a concern, and would agree that we need global coordination on AI safety to make this work.)
I guess the default for me is that Scientist AI won’t be competitive, so we live in a world with both scientist AI and non-scientist AI. Conditional upon successfully tamping down other approaches enough that Scientist AI gets to the weakly superhuman point while we’re still alive, I’m more optimistic that we can continue to coordinate on doing things safely.