I sometimes felt you were implying āno evidenceā when I thought something closer to āvery weak evidenceā would be more appropriate. However, in practice, I would agree the focus should overwhelmingly be on decreasing the uncertainty about the extent to which AI models have welfare, and how to measure it.
I donāt find this convincing. (NB: Iām still working on this response.) Even if the talker has no access to the valenced consciousness, itās model may simply lead it to a confident and wrong answer about this.
I sometimes felt you were implying āno evidenceā when I thought something closer to āvery weak evidenceā would be more appropriate.
Can you provide some specific examples, and Iāll reconsider
However, in practice, I would agree the focus should overwhelmingly be on decreasing the uncertainty about the extent to which AI models have welfare, and how to measure it.
Iām not sure if we āagreeā here, or at least thatās not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed āhaveā or generate consciousness and welfare.
My main doubt was more:
Suppose this valenced experience and āwelfareā indeed comes about in such a system, which is so very different from human or other biological systems,
How could we evenr know, even in an expected value directional sense, what things made this welfare higher or lower
or whether it is even ānet positiveā so we want to āmake more of theseā or the opposite
We cannot trust the ātalkerā to tell us this and I donāt understand what other signals coming out of the AI would be reliable evidence
This would imply we have ādeep uncertaintyā about the impact of our actions on anything welfare-relevant, making our choices morally irrelevant.
So Iām not sure that we have useful ways to decrease the uncertainty about whether th AI models have welfare, and Iām not sure we really have any ways to measure it. But I guess agree we should try to decrease the uncertainty about āwhether we can now or will ever have ways to measure itā (or āitās valenceā) ⦠which is what my post was sort of trying to do.
Perhaps the strongest argument against this would be āwhat if the model says it is conscious, it is in pain, the data suggests it is not lying and that it is confident in its statement?ā
I donāt find this convincing. ⦠Even if the talker has no access to the valenced consciousness, itās model may simply lead it to a confident and wrong answer about this.
Did you mean āIfā instead of āEven ifā?
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesnāt mean that it actually knows the answer to āis it (or is anything in itās system) in painā.
So the āeven ifā was setting off a contrast between ānot having a access to the answerā and āmaking a confident statementā.
Can you provide some specific examples, and Iāll reconsider
I mostly had this section in mind. I think the positive correlation between hedonic welfare and reported seemingly honest preferences in humans provides some weak evidence that there is such a correlation for AIs. There are some similarities between humans and AIs.
Iām not sure if we āagreeā here, or at least thatās not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed āhaveā or generate consciousness and welfare.
Right. However, I think AI welfare being difficult to measure conditional on sentience tends to imply that assessing AI sentience is also difficult. So I would prioritise decreasing uncertainty about this too.
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesnāt mean that it actually knows the answer to āis it (or is anything in itās system) in painā.
So the āeven ifā was setting off a contrast between ānot having a access to the answerā and āmaking a confident statementā.
Thanks. Iāve read much of that ābullshitā paper and I think thereās some interesting overlap. Some ways I think it relates:
Their ālack of external validation of AI welfareā problem seems at the root of the problem I name that āthe signals we get from the AI may not tells us anything directional about the valence of the part of the AI that has consciousnessā (if it does). If there is no ground truth (opportunity for āfalsificationā) here I donāt see how we can credibly make that link.
The evidence they cite about the sensitivity of the measures signals of valence to seemingly irrelevant engineering choices reinforces my doubts.
Hi David. Great post. I broadly agree. Relatedly, readers may be interested in the post AI Welfare Is (Frankfurtian) Bullshit.
I sometimes felt you were implying āno evidenceā when I thought something closer to āvery weak evidenceā would be more appropriate. However, in practice, I would agree the focus should overwhelmingly be on decreasing the uncertainty about the extent to which AI models have welfare, and how to measure it.
Did you mean āIfā instead of āEven ifā?
Nitpick. āWhat mightā.
Can you provide some specific examples, and Iāll reconsider
Iām not sure if we āagreeā here, or at least thatās not the point I was making. I was accepting that AI models (either the information flow itself, or the physical electron flows or something) might indeed āhaveā or generate consciousness and welfare.
My main doubt was more:
Suppose this valenced experience and āwelfareā indeed comes about in such a system, which is so very different from human or other biological systems,
How could we evenr know, even in an expected value directional sense, what things made this welfare higher or lower
or whether it is even ānet positiveā so we want to āmake more of theseā or the opposite
We cannot trust the ātalkerā to tell us this and I donāt understand what other signals coming out of the AI would be reliable evidence
This would imply we have ādeep uncertaintyā about the impact of our actions on anything welfare-relevant, making our choices morally irrelevant.
So Iām not sure that we have useful ways to decrease the uncertainty about whether th AI models have welfare, and Iām not sure we really have any ways to measure it. But I guess agree we should try to decrease the uncertainty about āwhether we can now or will ever have ways to measure itā (or āitās valenceā) ⦠which is what my post was sort of trying to do.
What I meant was that even if the AI is making a confident statement that it is in pain, and it is not deliberately lying, it doesnāt mean that it actually knows the answer to āis it (or is anything in itās system) in painā.
So the āeven ifā was setting off a contrast between ānot having a access to the answerā and āmaking a confident statementā.
I mostly had this section in mind. I think the positive correlation between hedonic welfare and reported seemingly honest preferences in humans provides some weak evidence that there is such a correlation for AIs. There are some similarities between humans and AIs.
Right. However, I think AI welfare being difficult to measure conditional on sentience tends to imply that assessing AI sentience is also difficult. So I would prioritise decreasing uncertainty about this too.
Makes sense.
Thanks. Iāve read much of that ābullshitā paper and I think thereās some interesting overlap. Some ways I think it relates:
Their ālack of external validation of AI welfareā problem seems at the root of the problem I name that āthe signals we get from the AI may not tells us anything directional about the valence of the part of the AI that has consciousnessā (if it does). If there is no ground truth (opportunity for āfalsificationā) here I donāt see how we can credibly make that link.
The evidence they cite about the sensitivity of the measures signals of valence to seemingly irrelevant engineering choices reinforces my doubts.