I sometimes felt you were implying “no evidence” when I thought something closer to “very weak evidence” would be more appropriate.
Can you provide some specific examples, and I’ll reconsider
However, in practice, I would agree the focus should overwhelmingly be on decreasing the uncertainty about the extent to which AI models have welfare, and how to measure it.
I’m not sure if we ‘agree’ here, or at least that’s not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed “have” or generate consciousness and welfare.
My main doubt was more:
Suppose this valenced experience and ‘welfare’ indeed comes about in such a system, which is so very different from human or other biological systems,
How could we evenr know, even in an expected value directional sense, what things made this welfare higher or lower
or whether it is even ‘net positive’ so we want to ‘make more of these’ or the opposite
We cannot trust the ‘talker’ to tell us this and I don’t understand what other signals coming out of the AI would be reliable evidence
This would imply we have ‘deep uncertainty’ about the impact of our actions on anything welfare-relevant, making our choices morally irrelevant.
So I’m not sure that we have useful ways to decrease the uncertainty about whether th AI models have welfare, and I’m not sure we really have any ways to measure it. But I guess agree we should try to decrease the uncertainty about ‘whether we can now or will ever have ways to measure it’ (or ‘it’s valence’) … which is what my post was sort of trying to do.
Perhaps the strongest argument against this would be “what if the model says it is conscious, it is in pain, the data suggests it is not lying and that it is confident in its statement?”
I don’t find this convincing. … Even if the talker has no access to the valenced consciousness, it’s model may simply lead it to a confident and wrong answer about this.
Did you mean “If” instead of “Even if”?
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesn’t mean that it actually knows the answer to ‘is it (or is anything in it’s system) in pain’.
So the ‘even if’ was setting off a contrast between ‘not having a access to the answer’ and ‘making a confident statement’.
Can you provide some specific examples, and I’ll reconsider
I mostly had this section in mind. I think the positive correlation between hedonic welfare and reported seemingly honest preferences in humans provides some weak evidence that there is such a correlation for AIs. There are some similarities between humans and AIs.
I’m not sure if we ‘agree’ here, or at least that’s not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed “have” or generate consciousness and welfare.
Right. However, I think AI welfare being difficult to measure conditional on sentience tends to imply that assessing AI sentience is also difficult. So I would prioritise decreasing uncertainty about this too.
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesn’t mean that it actually knows the answer to ‘is it (or is anything in it’s system) in pain’.
So the ‘even if’ was setting off a contrast between ‘not having a access to the answer’ and ‘making a confident statement’.
Can you provide some specific examples, and I’ll reconsider
I’m not sure if we ‘agree’ here, or at least that’s not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed “have” or generate consciousness and welfare.
My main doubt was more:
Suppose this valenced experience and ‘welfare’ indeed comes about in such a system, which is so very different from human or other biological systems,
How could we evenr know, even in an expected value directional sense, what things made this welfare higher or lower
or whether it is even ‘net positive’ so we want to ‘make more of these’ or the opposite
We cannot trust the ‘talker’ to tell us this and I don’t understand what other signals coming out of the AI would be reliable evidence
This would imply we have ‘deep uncertainty’ about the impact of our actions on anything welfare-relevant, making our choices morally irrelevant.
So I’m not sure that we have useful ways to decrease the uncertainty about whether th AI models have welfare, and I’m not sure we really have any ways to measure it. But I guess agree we should try to decrease the uncertainty about ‘whether we can now or will ever have ways to measure it’ (or ‘it’s valence’) … which is what my post was sort of trying to do.
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesn’t mean that it actually knows the answer to ‘is it (or is anything in it’s system) in pain’.
So the ‘even if’ was setting off a contrast between ‘not having a access to the answer’ and ‘making a confident statement’.
I mostly had this section in mind. I think the positive correlation between hedonic welfare and reported seemingly honest preferences in humans provides some weak evidence that there is such a correlation for AIs. There are some similarities between humans and AIs.
Right. However, I think AI welfare being difficult to measure conditional on sentience tends to imply that assessing AI sentience is also difficult. So I would prioritise decreasing uncertainty about this too.
Makes sense.