Good idea! I just read through the study preprint last night and was puzzling over how you could get a good grasp on how much of an impact the wikipedia edits are making.
Another option would be to see if there was a model that took in ~half of the PAW edits in it’s training corpus, and test it on all the claims. Then you could see roughly how much of a difference having the info on wikipedia specifically made. Might still need to control for prominence outside wikipedia though.
Hi Jonah, Jasmine and Miles,
Cool idea! Thanks for writing it up.
I wonder if the causal effect of adding true stuff to wikipedia could be tested directly? For example:
Treatment: claims that are on Wikipedia
Control: claims that aren’t on Wikipedia
Design: Match by how likely the claim is to be true, and by how frequently the claim appears outside of Wikipedia.[1] Compute contrasts.
A Wikipedia claim might be mirrored on various other parts of the training corpus. We might want to control for that.
Good idea! I just read through the study preprint last night and was puzzling over how you could get a good grasp on how much of an impact the wikipedia edits are making.
Another option would be to see if there was a model that took in ~half of the PAW edits in it’s training corpus, and test it on all the claims. Then you could see roughly how much of a difference having the info on wikipedia specifically made. Might still need to control for prominence outside wikipedia though.