AI alignment prize suggestion: Demonstrate a true sandwiching project
Artificial Intelligence
Sandwiching projects are a concrete way for how to make progress on aligning narrowly superhuman models. They “sandwich” the model in between one set of humans which is less capable than it and another set of humans which is more capable than it at the fuzzy task in question, and b) figure out how to help the less-capable set of humans reproduce the judgments of the more-capable set of humans. For example, first fine-tune a coding model to write short functions solving simple puzzles using demonstrations and feedback collected from expert software engineers. Then try to match this performance using some process that can be implemented by people who don’t know how to code and/or couldn’t solve the puzzles themselves.
Importantly, there are many ways to attack a sandwiching project that are slightly cheating. The most challenging version of a sandwiching project would need to make sure that no information whatsoever from the more-capable set of humans is used in the training process. The Future Fund could offer prizes for demonstrations of sandwiching projects on various levels of impressiveness and generality of the employed method.
AI alignment prize suggestion: Demonstrate a true sandwiching project
Artificial Intelligence
Sandwiching projects are a concrete way for how to make progress on aligning narrowly superhuman models. They “sandwich” the model in between one set of humans which is less capable than it and another set of humans which is more capable than it at the fuzzy task in question, and b) figure out how to help the less-capable set of humans reproduce the judgments of the more-capable set of humans. For example, first fine-tune a coding model to write short functions solving simple puzzles using demonstrations and feedback collected from expert software engineers. Then try to match this performance using some process that can be implemented by people who don’t know how to code and/or couldn’t solve the puzzles themselves.
Importantly, there are many ways to attack a sandwiching project that are slightly cheating. The most challenging version of a sandwiching project would need to make sure that no information whatsoever from the more-capable set of humans is used in the training process. The Future Fund could offer prizes for demonstrations of sandwiching projects on various levels of impressiveness and generality of the employed method.