Thanks for your reply, and apologies my OP was unclear. Hopefully I do better here.
1) Your distinction between âthere is a shared shape across all link functionsâ and âwhich shape across all the link functionsâ is correct.
Re âwhich shape?â: if you switch from a logistic curve to clipped linear one (or, I expect, most/âall reasonable alternatives) you likely twiddle the exact values and rankings, but the broad order and ~linearly-increasing-over-time sweep of the data should remain much the same. Very loosely, even if the tick-marks on your ruler are all over the place, longer distances measured by multiple rulers laid end to end will still be in the right ballpark.
Re. âshared shapeâ: But you need to assert all the items/âbenchmarks have identical shape to resolve a single interval scale across them (very loosely again: that all the rulers laid end to end have the same length). Perhaps one way of understanding this goes back to the âendogeneity graphâ:
Handwaving: one way you could move from the middle panel to either of ones flanking it is to assert the y-axis should be re-parameterized: âTrue capability is really exponential/âlogarithmic in the ECI capability trait valueâ. Yet (loosely) you could bake the same transformation into the capability trait itself, by stipulating a progressive warp of link functions across benchmarks as they increase in difficulty.
For âwhy doubt shared shape?â, I think thereâs a fair case for brute incredulity: why should benchmarks for (e.g.) chess puzzles, science questions, verbal reasoning, (log) time horizons for coding tasks (etc.) all share the same relationship between score progress and general capability? Empirically, browsing some of the âbenchmark score vs. ECI plotsâ on Epochâs website shows some look more or less S-curvy than others, and the data too noisy and sparse (and likely reliant on tail behaviour) to uniquely identify an logistic curve versus others that look similar.
2) My broad criticisms draw on ~fundamental limitations to IRT or underspecification/âambiguity around âAI capabilityâ - in either case, I donât think thereâs a crisp statistical test or methodological fix. Although I think this covers the endogeneity worry in general (it seems hard for the data alone to distinguish methodological distortion from secular trend), trying to pin down particular niggles is handy in its own right, and offers a gesture at the broader complaint: if the âjust read the number straightforwardlyâ gets the right answer and my metrological quibbling just chases shadows, so much the worse for meâand vice versa.
Your suggestions could be worth doing, although I think the results would only be cruxy in one direction. On my end, Iâd guess the speeding up result is robust to link function choice, and would only mildly attenuate if you restrict to pre-model benchmarks (or somehow upweight them relative to post-model benchmarks) - i.e. I wouldnât take positive results as confirmatory proof. Yet if you found (e.g.) speeding up is sensitive to link function choice, that seems much stronger negative evidence.
Some further ideas, although Iâm not sure theyâre worth much:
I understand thereâs âquasi-parametricâ IRT which lets you twiddle the link functionsâand anyway you could play with simulation data to see how shape-sensitive the results are. By and large, the response curves of the benchmarks certainly tend to be S-shaped, so (e.g.) âBut if you only tweak the s-curves slightly between benchmarks you can generate huge de/âaccelerationsâ vs. âTo ablate results like this you need to stipulate hugely warps which we donât actually see in the dataâ could be informative.
I guess all of this will be confounded by reducing power (see fn. 17), but besides testing on earlier or later subsets, you could try and curate an equal discrimination set of benchmarks across the range, or narrow the set to those which bracket the trend break in 2024.
One could review all the benchmarks, and subjectively score them by how âreasoning-y/âmultistepâ each is, then somehow reweigh/âcompare classes to see if it lines up (i.e. do the benchmarks get more reasoning-y over time? Does this correspond to benchmark discrimination? If we somehow âcorrectâ, does the trend break vanish). But even if so (and going back to the problem where external validity is up for grabs), it doesnât seem crazy to assert that âtrue/âimportantâ capabilities do load more heavily on reasoning at you go up the range, so even if the benchmark measurement axis is rotating, this rotation keeps it parallel to the âtrue capabilityâ axis.
Hello Charlie,
Thanks for your reply, and apologies my OP was unclear. Hopefully I do better here.
1) Your distinction between âthere is a shared shape across all link functionsâ and âwhich shape across all the link functionsâ is correct.
Re âwhich shape?â: if you switch from a logistic curve to clipped linear one (or, I expect, most/âall reasonable alternatives) you likely twiddle the exact values and rankings, but the broad order and ~linearly-increasing-over-time sweep of the data should remain much the same. Very loosely, even if the tick-marks on your ruler are all over the place, longer distances measured by multiple rulers laid end to end will still be in the right ballpark.
Re. âshared shapeâ: But you need to assert all the items/âbenchmarks have identical shape to resolve a single interval scale across them (very loosely again: that all the rulers laid end to end have the same length). Perhaps one way of understanding this goes back to the âendogeneity graphâ:
Handwaving: one way you could move from the middle panel to either of ones flanking it is to assert the y-axis should be re-parameterized: âTrue capability is really exponential/âlogarithmic in the ECI capability trait valueâ. Yet (loosely) you could bake the same transformation into the capability trait itself, by stipulating a progressive warp of link functions across benchmarks as they increase in difficulty.
For âwhy doubt shared shape?â, I think thereâs a fair case for brute incredulity: why should benchmarks for (e.g.) chess puzzles, science questions, verbal reasoning, (log) time horizons for coding tasks (etc.) all share the same relationship between score progress and general capability? Empirically, browsing some of the âbenchmark score vs. ECI plotsâ on Epochâs website shows some look more or less S-curvy than others, and the data too noisy and sparse (and likely reliant on tail behaviour) to uniquely identify an logistic curve versus others that look similar.
2) My broad criticisms draw on ~fundamental limitations to IRT or underspecification/âambiguity around âAI capabilityâ - in either case, I donât think thereâs a crisp statistical test or methodological fix. Although I think this covers the endogeneity worry in general (it seems hard for the data alone to distinguish methodological distortion from secular trend), trying to pin down particular niggles is handy in its own right, and offers a gesture at the broader complaint: if the âjust read the number straightforwardlyâ gets the right answer and my metrological quibbling just chases shadows, so much the worse for meâand vice versa.
Your suggestions could be worth doing, although I think the results would only be cruxy in one direction. On my end, Iâd guess the speeding up result is robust to link function choice, and would only mildly attenuate if you restrict to pre-model benchmarks (or somehow upweight them relative to post-model benchmarks) - i.e. I wouldnât take positive results as confirmatory proof. Yet if you found (e.g.) speeding up is sensitive to link function choice, that seems much stronger negative evidence.
Some further ideas, although Iâm not sure theyâre worth much:
I understand thereâs âquasi-parametricâ IRT which lets you twiddle the link functionsâand anyway you could play with simulation data to see how shape-sensitive the results are. By and large, the response curves of the benchmarks certainly tend to be S-shaped, so (e.g.) âBut if you only tweak the s-curves slightly between benchmarks you can generate huge de/âaccelerationsâ vs. âTo ablate results like this you need to stipulate hugely warps which we donât actually see in the dataâ could be informative.
I guess all of this will be confounded by reducing power (see fn. 17), but besides testing on earlier or later subsets, you could try and curate an equal discrimination set of benchmarks across the range, or narrow the set to those which bracket the trend break in 2024.
One could review all the benchmarks, and subjectively score them by how âreasoning-y/âmultistepâ each is, then somehow reweigh/âcompare classes to see if it lines up (i.e. do the benchmarks get more reasoning-y over time? Does this correspond to benchmark discrimination? If we somehow âcorrectâ, does the trend break vanish). But even if so (and going back to the problem where external validity is up for grabs), it doesnât seem crazy to assert that âtrue/âimportantâ capabilities do load more heavily on reasoning at you go up the range, so even if the benchmark measurement axis is rotating, this rotation keeps it parallel to the âtrue capabilityâ axis.