Yeah, I agree this is unclear. But, staying away from the word ‘intention’ entirely, I think we can & should still ask: what is the best explanation for why this model is the one that minimizes the loss function during training? Does that explanation involve this argument about changing user preferences, or not?
One concrete experiment that could feed into this: if it were the case that feeding users extreme political content did not cause their views to become more predictable, would training select a model that didn’t feed people as much extreme political content? I’d guess training would select the same model anyway, because extreme political content gets clicks in the short-term too. (But I might be wrong.)
Yeah, I agree this is unclear. But, staying away from the word ‘intention’ entirely, I think we can & should still ask: what is the best explanation for why this model is the one that minimizes the loss function during training? Does that explanation involve this argument about changing user preferences, or not?
One concrete experiment that could feed into this: if it were the case that feeding users extreme political content did not cause their views to become more predictable, would training select a model that didn’t feed people as much extreme political content? I’d guess training would select the same model anyway, because extreme political content gets clicks in the short-term too. (But I might be wrong.)