Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
you don’t even say “most of the time” or “effectively” or anything—just ask whether it is possible. I think the probability that this occurs at least once is overwhelmingly likely.
you don’t even say “most of the time” or “effectively” or anything—just ask whether it is possible. I think the probability that this occurs at least once is overwhelmingly likely.