I am one of the authors of the aforementioned novice uplift randomized controlled trial (Hong et al. 2025), and I want to clarify that while the study did not detect a statistically significant difference on the full model reverse-genetics workflow, nonetheless we did find uplift on component tasks such as mammalian cell culture!
Here, a reasonable takeaway is that tacit knowledge in biology is not a strong barrier for uplift — instead, participants in the AI arm performed much better with AI assistance, even on tasks that we would consider to be tacit-knowledge heavy.
An important caveat is that the study was conducted mid-2025, using the then-frontier models of that generation but more importantly, with a participant pool that is representative of the average level of AI proficiency of that period. As time goes on, secular trends in AI skill development will likely make people better at using models.
Finally — I’m curious to hear what’s the second RCT that you mentioned in your post? I’d be grateful to receive the link to it!
To add a bit of colour to your caveat, when folks bring up uplift papers using yesteryear’s models I’m reminded of these charts from Anthropic and OpenAI respectively, and I think “pre-late 2025 frontier models gave you basically no uplift for even the most uplift-able domain (code), but Mythos/Astra-class models’ uplift is so great for coding that conclusions about them based on pre-late 2025 models don’t seem informative at all, I wish there were uplift RCTs looking at Mythos/Astra-class models for other domains”. For now the qualitative remarks in section 4.4.3 of Anthropic’s report on CB-2 evidence for Mythos Preview, Fable 5, and Mythos 5 will have to do, which concluded novices probably wouldn’t get significant uplift but experts would (I assume you’re already aware of them, given that the preceding subsection quoted your paper). And stacked on top of this would be, as you say, the secular trend of laypeople learning to use frontier AI better over time.
I am one of the authors of the aforementioned novice uplift randomized controlled trial (Hong et al. 2025), and I want to clarify that while the study did not detect a statistically significant difference on the full model reverse-genetics workflow, nonetheless we did find uplift on component tasks such as mammalian cell culture!
Here, a reasonable takeaway is that tacit knowledge in biology is not a strong barrier for uplift — instead, participants in the AI arm performed much better with AI assistance, even on tasks that we would consider to be tacit-knowledge heavy.
An important caveat is that the study was conducted mid-2025, using the then-frontier models of that generation but more importantly, with a participant pool that is representative of the average level of AI proficiency of that period. As time goes on, secular trends in AI skill development will likely make people better at using models.
Finally — I’m curious to hear what’s the second RCT that you mentioned in your post? I’d be grateful to receive the link to it!
To add a bit of colour to your caveat, when folks bring up uplift papers using yesteryear’s models I’m reminded of these charts from Anthropic and OpenAI respectively, and I think “pre-late 2025 frontier models gave you basically no uplift for even the most uplift-able domain (code), but Mythos/Astra-class models’ uplift is so great for coding that conclusions about them based on pre-late 2025 models don’t seem informative at all, I wish there were uplift RCTs looking at Mythos/Astra-class models for other domains”. For now the qualitative remarks in section 4.4.3 of Anthropic’s report on CB-2 evidence for Mythos Preview, Fable 5, and Mythos 5 will have to do, which concluded novices probably wouldn’t get significant uplift but experts would (I assume you’re already aware of them, given that the preceding subsection quoted your paper). And stacked on top of this would be, as you say, the secular trend of laypeople learning to use frontier AI better over time.