Benchmarks will become useless due to eval awareness¹
“Alignment” benchmarks will become useless as AIs can modify their behaviour or hide their motives when they know they are being evaluated. For “capabilities” benchmarks, an AI might hide its capability (pretend to be less capable) if it knows it’s being evaluated, but it’s not immediately obvious that an AI would want to hide its capability. It may know that it is being evaluated, and decide to try its best anyway.
“Alignment” benchmarks will become useless as AIs can modify their behaviour or hide their motives when they know they are being evaluated. For “capabilities” benchmarks, an AI might hide its capability (pretend to be less capable) if it knows it’s being evaluated, but it’s not immediately obvious that an AI would want to hide its capability. It may know that it is being evaluated, and decide to try its best anyway.