Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn’t be able to control those in a meaningful way nor do I think current AI’s could (but wouldn’t be shocked if they already could control the voice in their head if they have it). but I would guess future ai’s will know how to control increasingly large parts of their brain activations.
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn’t be able to control those in a meaningful way nor do I think current AI’s could (but wouldn’t be shocked if they already could control the voice in their head if they have it). but I would guess future ai’s will know how to control increasingly large parts of their brain activations.