Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
Depends on how good the AI is and how good the tools are? This is kind of a bad question since “deceptive AIs” is not a very precise definition.
Depends on how good the AI is and how good the tools are? This is kind of a bad question since “deceptive AIs” is not a very precise definition.
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks