So the author is right that evals are overrated but the reason is deeper then model detecting when they are being tested.
The core diagnosis is this: Evals are architecturally wrong for the job.They are designed like compliance audit e.g one test, one result but if we organized capability assessment against a system that already aware of that requires a more adversarial methodology .
IC methodology which faces a same problem like this. So the solution was not direct elicitation it was triangulating across independent behavioural signals so no single source could be game.
When we apply it to AI by using different cross reference sources like red team outputs, deployment behaviour and independent replication by compartmentalized team. Now the convergence through different independent resources is much harder to game than a single test.
The fix is not better evals.Its a different methodology framework which is closer to competitive intelligence than compliance audit .
So the author is right that evals are overrated but the reason is deeper then model detecting when they are being tested.
The core diagnosis is this: Evals are architecturally wrong for the job.They are designed like compliance audit e.g one test, one result but if we organized capability assessment against a system that already aware of that requires a more adversarial methodology .
IC methodology which faces a same problem like this. So the solution was not direct elicitation it was triangulating across independent behavioural signals so no single source could be game.
When we apply it to AI by using different cross reference sources like red team outputs, deployment behaviour and independent replication by compartmentalized team. Now the convergence through different independent resources is much harder to game than a single test.
The fix is not better evals.Its a different methodology framework which is closer to competitive intelligence than compliance audit .