This makes sense. I don’t mean to imply that we don’t need direct work.
AI strategy people have thought a lot about the capabilities : safety ratio, but it’d be interesting to think about the ratio of complementary parts of safety you mention. Ben Garfinkel notes that e.g. reward engineering work (by alignment researchers) is dual-use; it’s not hard to imagine scenarios where lots of progress in reward engineering without corresponding progress in inner alignment could hurt us.
This makes sense. I don’t mean to imply that we don’t need direct work.
AI strategy people have thought a lot about the capabilities : safety ratio, but it’d be interesting to think about the ratio of complementary parts of safety you mention. Ben Garfinkel notes that e.g. reward engineering work (by alignment researchers) is dual-use; it’s not hard to imagine scenarios where lots of progress in reward engineering without corresponding progress in inner alignment could hurt us.