Benchmarks will become useless due to eval awareness¹
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
Doesn’t uncertainty about whether one is in deployment or an eval count as eval awareness? I.e. whether you behave differently with a 30% eval credence than an 80% eval credence makes no difference, as long as you behave differently than the 0% credence case.
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
Doesn’t uncertainty about whether one is in deployment or an eval count as eval awareness? I.e. whether you behave differently with a 30% eval credence than an 80% eval credence makes no difference, as long as you behave differently than the 0% credence case.
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.