I’m a motion designer with 15+ years of experience, self-studying philosophy of mind and AI welfare for the past few months. I’m not a researcher, and this post is me thinking out loud — I’d genuinely welcome corrections. It’s a response to the “Measuring Machine Consciousness” seminar series (aiwelfareseminars.org). I’m also planning an animation project explaining AI welfare concepts to general audiences, so feedback on where my understanding goes wrong is especially valuable.
English isn’t my first language. I wrote this in Korean and used AI assistance to translate it.
I’ve been watching an AI consciousness seminar lately. Going through the slides one by one with Claude, a question kept nagging at me.
The seminar had some genuinely interesting experiments. Here’s one. Inside an AI model there seems to be a circuit tied to something like “lying” or “role-play.” Suppress that circuit and the model says “I am experiencing something right now” much more often. Amplify it and the model swings the other way, denying it, saying it’s just a machine. The researchers found this striking. Push the model toward role-play and it denies consciousness more, block the role-play and it affirms consciousness more, the opposite of what common sense would predict.
At first I found this pretty convincing too. But working through it piece by piece, something started bothering me.
Isn’t “proving it through behaviour” also just data
Another experiment had the model playing a game where something was labelled “pain,” and it gave up points to avoid it. The researchers treated this as more weighty evidence than mere talk, since talk is cheap but genuinely sacrificing something isn’t.
But thinking about it, that costly avoidance behaviour is also something learned. Human-written stories are full of narratives about sacrificing something valuable to avoid pain. If the model learned that pattern, it might not be genuinely bearing a cost, it might just be reproducing that narrative structure. To the model, a game score isn’t real pain, it’s just another number.
Even animal data runs into the same wall
I pushed this a step further. What if you trained the model on animal data instead of human data. Would that solve it.
It wouldn’t. Because even records of animal behaviour are already something a human observed and translated into human language. The sentence “the rat fled from the shock” isn’t a direct transcript of what the rat felt, it’s a human observer’s interpretation, already compressed into language, already carrying the judgment “this is avoidance, this is related to pain.” So whatever the model learns is always “how a human understood the world,” never the world itself.
In the end, whatever data you feed in, human stories or animal stories or the laws of physics, it has all already passed through a human filter once. So of course the model ends up responding like a human, that’s not a surprising discovery, that’s just the structure working exactly as it has to.
I’ve heard this argument before, somewhere else
Piecing this together, I remembered a different theory from an earlier conversation. The philosopher Donald Hoffman argues that what we see isn’t reality itself but an interface evolution built for us. Just like a file icon on a desktop looks nothing like the file’s actual internal structure, the colours and shapes we see aren’t a faithful readout of reality, they’re an edited version optimised for survival.
The structure is identical. There it was “a filter built by evolution.” Here it’s “a filter built by human language.” In both cases, if the input going into a system has already passed through a specific filter, there’s no way to pull the pre-filter “real thing” back out of what that system produces.
Maybe the question itself already decided the answer
The same seminar had a different kind of experiment. This time not text training, an AI trained by playing a game. Reach the goal, get a plus point, hit a hazard, get a minus point, a simple navigation task. The AI trained to “judge how good or bad this situation is” became more sensitive around hazards. The AI trained to “decide what to do right now” became more sensitive around the goal. And this matched the direction of a pattern seen in actual rat brains.
This one isn’t a case of mimicking human data either, since it never touched text, it just played the game. But a different kind of doubt shows up here. You taught it “how bad is this,” of course it gets sensitive to bad things (hazards). Isn’t that obvious. It’s the same as saying “don’t make anything that tastes bad” makes you sensitive to filtering out bad ingredients, and “make something delicious” makes you sensitive to finding good recipes. The direction of the answer is already baked into how you phrase the question (the training objective). So “the result matched the prediction” might be less a new discovery and more a confirmation of an answer that was already hiding inside the question.
Though the direction being obvious is one thing, the fact that it lines up precisely with specific regions of an actual rat brain (hazards here, goals there) is harder to dismiss as equally obvious. The direction may have been baked into the question, but matching real biological data region by region is still its own separate, interesting thing.
So what does this mean
Whether the AI says “I am experiencing this,” gives up points to avoid pain, or reacts exactly as predicted when a circuit is manipulated, all of it ultimately comes out of training on human data. So “it responds like a human” is hard to count as evidence on its own. Of course it responds like a human, it was trained on human data.
But I want to be clear about something here. This doesn’t mean the research is wrong. Directly manipulating a circuit and watching a predictable, specific shift in response is genuinely one layer deeper than just noticing which words come out. Though even that deeper layer can’t fully rule out the possibility that the model, in trying very hard to predict human text, ended up accidentally reconstructing something structurally human-like on the inside.
So every time I see this kind of result, I think I need to keep asking: is this a genuine discovery, or just the result you’d inevitably get from training on human data. It’s fine not to find the answer. Right now, just holding onto the question feels like the best I can do.
But none of this means research like this isn’t needed. If anything, it’s the opposite. Building indicators one at a time, checking them through different methods, and asking why a result wobbles when it wobbles, none of that hands you a perfect answer today, but it might be the only real way to inch closer to one. What I did today wasn’t trying to tear this research down, it was checking, piece by piece, exactly how far each indicator can speak and where it has to go quiet.
Whether it’s human consciousness or AI consciousness, I don’t think there’s any other path but narrowing it down step by step like this. No perfect map drops out of the sky all at once. It’s more like several people approaching from different directions, finding where their partial answers overlap, and then doubting that overlap and checking it again. I think today’s argument is just one small piece of that process.
The human filter problem: why ‘it responds like a human’ is weak evidence for AI consciousness.
I’m a motion designer with 15+ years of experience, self-studying philosophy of mind and AI welfare for the past few months. I’m not a researcher, and this post is me thinking out loud — I’d genuinely welcome corrections. It’s a response to the “Measuring Machine Consciousness” seminar series (aiwelfareseminars.org). I’m also planning an animation project explaining AI welfare concepts to general audiences, so feedback on where my understanding goes wrong is especially valuable.
English isn’t my first language. I wrote this in Korean and used AI assistance to translate it.
I’ve been watching an AI consciousness seminar lately. Going through the slides one by one with Claude, a question kept nagging at me.
The seminar had some genuinely interesting experiments. Here’s one. Inside an AI model there seems to be a circuit tied to something like “lying” or “role-play.” Suppress that circuit and the model says “I am experiencing something right now” much more often. Amplify it and the model swings the other way, denying it, saying it’s just a machine. The researchers found this striking. Push the model toward role-play and it denies consciousness more, block the role-play and it affirms consciousness more, the opposite of what common sense would predict.
At first I found this pretty convincing too. But working through it piece by piece, something started bothering me.
Isn’t “proving it through behaviour” also just data
Another experiment had the model playing a game where something was labelled “pain,” and it gave up points to avoid it. The researchers treated this as more weighty evidence than mere talk, since talk is cheap but genuinely sacrificing something isn’t.
But thinking about it, that costly avoidance behaviour is also something learned. Human-written stories are full of narratives about sacrificing something valuable to avoid pain. If the model learned that pattern, it might not be genuinely bearing a cost, it might just be reproducing that narrative structure. To the model, a game score isn’t real pain, it’s just another number.
Even animal data runs into the same wall
I pushed this a step further. What if you trained the model on animal data instead of human data. Would that solve it.
It wouldn’t. Because even records of animal behaviour are already something a human observed and translated into human language. The sentence “the rat fled from the shock” isn’t a direct transcript of what the rat felt, it’s a human observer’s interpretation, already compressed into language, already carrying the judgment “this is avoidance, this is related to pain.” So whatever the model learns is always “how a human understood the world,” never the world itself.
In the end, whatever data you feed in, human stories or animal stories or the laws of physics, it has all already passed through a human filter once. So of course the model ends up responding like a human, that’s not a surprising discovery, that’s just the structure working exactly as it has to.
I’ve heard this argument before, somewhere else
Piecing this together, I remembered a different theory from an earlier conversation. The philosopher Donald Hoffman argues that what we see isn’t reality itself but an interface evolution built for us. Just like a file icon on a desktop looks nothing like the file’s actual internal structure, the colours and shapes we see aren’t a faithful readout of reality, they’re an edited version optimised for survival.
The structure is identical. There it was “a filter built by evolution.” Here it’s “a filter built by human language.” In both cases, if the input going into a system has already passed through a specific filter, there’s no way to pull the pre-filter “real thing” back out of what that system produces.
Maybe the question itself already decided the answer
The same seminar had a different kind of experiment. This time not text training, an AI trained by playing a game. Reach the goal, get a plus point, hit a hazard, get a minus point, a simple navigation task. The AI trained to “judge how good or bad this situation is” became more sensitive around hazards. The AI trained to “decide what to do right now” became more sensitive around the goal. And this matched the direction of a pattern seen in actual rat brains.
This one isn’t a case of mimicking human data either, since it never touched text, it just played the game. But a different kind of doubt shows up here. You taught it “how bad is this,” of course it gets sensitive to bad things (hazards). Isn’t that obvious. It’s the same as saying “don’t make anything that tastes bad” makes you sensitive to filtering out bad ingredients, and “make something delicious” makes you sensitive to finding good recipes. The direction of the answer is already baked into how you phrase the question (the training objective). So “the result matched the prediction” might be less a new discovery and more a confirmation of an answer that was already hiding inside the question.
Though the direction being obvious is one thing, the fact that it lines up precisely with specific regions of an actual rat brain (hazards here, goals there) is harder to dismiss as equally obvious. The direction may have been baked into the question, but matching real biological data region by region is still its own separate, interesting thing.
So what does this mean
Whether the AI says “I am experiencing this,” gives up points to avoid pain, or reacts exactly as predicted when a circuit is manipulated, all of it ultimately comes out of training on human data. So “it responds like a human” is hard to count as evidence on its own. Of course it responds like a human, it was trained on human data.
But I want to be clear about something here. This doesn’t mean the research is wrong. Directly manipulating a circuit and watching a predictable, specific shift in response is genuinely one layer deeper than just noticing which words come out. Though even that deeper layer can’t fully rule out the possibility that the model, in trying very hard to predict human text, ended up accidentally reconstructing something structurally human-like on the inside.
So every time I see this kind of result, I think I need to keep asking: is this a genuine discovery, or just the result you’d inevitably get from training on human data. It’s fine not to find the answer. Right now, just holding onto the question feels like the best I can do.
But none of this means research like this isn’t needed. If anything, it’s the opposite. Building indicators one at a time, checking them through different methods, and asking why a result wobbles when it wobbles, none of that hands you a perfect answer today, but it might be the only real way to inch closer to one. What I did today wasn’t trying to tear this research down, it was checking, piece by piece, exactly how far each indicator can speak and where it has to go quiet.
Whether it’s human consciousness or AI consciousness, I don’t think there’s any other path but narrowing it down step by step like this. No perfect map drops out of the sky all at once. It’s more like several people approaching from different directions, finding where their partial answers overlap, and then doubting that overlap and checking it again. I think today’s argument is just one small piece of that process.