Not this person, but many AI risk arguments are necessarily logical rather than empirical—there are good reasons to believe the relevant behaviors won’t appear, or be trivially easy to counter (at least re: harmful outputs), until you have very capable systems.
Like, if I can construct a deceptive response-to-training strategy (but current models can’t), that’s enough evidence to be concerned future superhuman models might do similar deceptive alignment. Other concerns like inner optimizers (e.g. humans stopped being kid-maxxers at high capability, because our proxy decoupled from evolution’s target) might not show up, or change in character, as models become less limited. And even when you can demonstrate the behavior empirically, people dismiss it as overly-induced or a toy environment—which was the whole point, just to show plausibility not prove it.
More fundamentally: If I argue that a future thing logically implies certain risks arise, responding with “there’s no empirical evidence” is silly. Logical chains and structural arguments are still valid epistemic tools.
Not this person, but many AI risk arguments are necessarily logical rather than empirical—there are good reasons to believe the relevant behaviors won’t appear, or be trivially easy to counter (at least re: harmful outputs), until you have very capable systems.
Like, if I can construct a deceptive response-to-training strategy (but current models can’t), that’s enough evidence to be concerned future superhuman models might do similar deceptive alignment. Other concerns like inner optimizers (e.g. humans stopped being kid-maxxers at high capability, because our proxy decoupled from evolution’s target) might not show up, or change in character, as models become less limited. And even when you can demonstrate the behavior empirically, people dismiss it as overly-induced or a toy environment—which was the whole point, just to show plausibility not prove it.
More fundamentally: If I argue that a future thing logically implies certain risks arise, responding with “there’s no empirical evidence” is silly. Logical chains and structural arguments are still valid epistemic tools.