Gregory—thank you for the interesting post. I admit I have *read most of it twice and not fully understood. May I ask two clarifying questions? Specifically on ECI.
First: you say link function choice is “relatively minor” (rank and ECI correlate ~0.97), but you also raise an endogeneity point, which “entirely relies on the link function.” I think the first claim holds link functions vary, shared-shape assumption stays fixed; the second is about dropping the shared-shape assumption itself. You say a shared-shape assumption is “dubious” but I’m unsure what your reasons are?
Second: I’m wondering what the empirical test(s) is for your critiques. In general, if you think measures of capabilities are not robust in some way, what is the sensitivity analysis? I am thinking like an economist here.
One idea—refit ECI under different (common) link functions, see if the “speeding up” trend survives. I don’t think Epoch runs this, though their appendix comes close, although I don’t think this would be a crux for you.
Two other ideas that came to mind, specifically on ‘measure endogeneity’:
Fit using only benchmarks that predate the model being scored.
Let discrimination depend on whether the benchmark came before or after the model.
Thanks for your reply, and apologies my OP was unclear. Hopefully I do better here.
1) Your distinction between “there is a shared shape across all link functions” and “which shape across all the link functions” is correct.
Re ‘which shape?’: if you switch from a logistic curve to clipped linear one (or, I expect, most/all reasonable alternatives) you likely twiddle the exact values and rankings, but the broad order and ~linearly-increasing-over-time sweep of the data should remain much the same. Very loosely, even if the tick-marks on your ruler are all over the place, longer distances measured by multiple rulers laid end to end will still be in the right ballpark.
Re. ‘shared shape’: But you need to assert all the items/benchmarks have identical shape to resolve a single interval scale across them (very loosely again: that all the rulers laid end to end have the same length). Perhaps one way of understanding this goes back to the ‘endogeneity graph’:
Handwaving: one way you could move from the middle panel to either of ones flanking it is to assert the y-axis should be re-parameterized: “True capability is really exponential/logarithmic in the ECI capability trait value”. Yet (loosely) you could bake the same transformation into the capability trait itself, by stipulating a progressive warp of link functions across benchmarks as they increase in difficulty.
For ‘why doubt shared shape?‘, I think there’s a fair case for brute incredulity: why should benchmarks for (e.g.) chess puzzles, science questions, verbal reasoning, (log) time horizons for coding tasks (etc.) all share the same relationship between score progress and general capability? Empirically, browsing some of the ‘benchmark score vs. ECI plots’ on Epoch’s website shows some look more or less S-curvy than others, and the data too noisy and sparse (and likely reliant on tail behaviour) to uniquely identify an logistic curve versus others that look similar.
2) My broad criticisms draw on ~fundamental limitations to IRT or underspecification/ambiguity around ‘AI capability’ - in either case, I don’t think there’s a crisp statistical test or methodological fix. Although I think this covers the endogeneity worry in general (it seems hard for the data alone to distinguish methodological distortion from secular trend), trying to pin down particular niggles is handy in its own right, and offers a gesture at the broader complaint: if the ‘just read the number straightforwardly’ gets the right answer and my metrological quibbling just chases shadows, so much the worse for me—and vice versa.
Your suggestions could be worth doing, although I think the results would only be cruxy in one direction. On my end, I’d guess the speeding up result is robust to link function choice, and would only mildly attenuate if you restrict to pre-model benchmarks (or somehow upweight them relative to post-model benchmarks) - i.e. I wouldn’t take positive results as confirmatory proof. Yet if you found (e.g.) speeding up is sensitive to link function choice, that seems much stronger negative evidence.
Some further ideas, although I’m not sure they’re worth much:
I understand there’s ‘quasi-parametric’ IRT which lets you twiddle the link functions—and anyway you could play with simulation data to see how shape-sensitive the results are. By and large, the response curves of the benchmarks certainly tend to be S-shaped, so (e.g.) “But if you only tweak the s-curves slightly between benchmarks you can generate huge de/accelerations” vs. “To ablate results like this you need to stipulate hugely warps which we don’t actually see in the data” could be informative.
I guess all of this will be confounded by reducing power (see fn. 17), but besides testing on earlier or later subsets, you could try and curate an equal discrimination set of benchmarks across the range, or narrow the set to those which bracket the trend break in 2024.
One could review all the benchmarks, and subjectively score them by how ‘reasoning-y/multistep’ each is, then somehow reweigh/compare classes to see if it lines up (i.e. do the benchmarks get more reasoning-y over time? Does this correspond to benchmark discrimination? If we somehow ‘correct’, does the trend break vanish). But even if so (and going back to the problem where external validity is up for grabs), it doesn’t seem crazy to assert that “true/important” capabilities do load more heavily on reasoning at you go up the range, so even if the benchmark measurement axis is rotating, this rotation keeps it parallel to the ‘true capability’ axis.
Gregory—thank you for the interesting post. I admit I have *read most of it twice and not fully understood. May I ask two clarifying questions? Specifically on ECI.
First: you say link function choice is “relatively minor” (rank and ECI correlate ~0.97), but you also raise an endogeneity point, which “entirely relies on the link function.” I think the first claim holds link functions vary, shared-shape assumption stays fixed; the second is about dropping the shared-shape assumption itself. You say a shared-shape assumption is “dubious” but I’m unsure what your reasons are?
Second: I’m wondering what the empirical test(s) is for your critiques. In general, if you think measures of capabilities are not robust in some way, what is the sensitivity analysis? I am thinking like an economist here.
One idea—refit ECI under different (common) link functions, see if the “speeding up” trend survives. I don’t think Epoch runs this, though their appendix comes close, although I don’t think this would be a crux for you.
Two other ideas that came to mind, specifically on ‘measure endogeneity’:
Fit using only benchmarks that predate the model being scored.
Let discrimination depend on whether the benchmark came before or after the model.
Would you find these interesting or cruxy?
Thanks, Charlie
Hello Charlie,
Thanks for your reply, and apologies my OP was unclear. Hopefully I do better here.
1) Your distinction between “there is a shared shape across all link functions” and “which shape across all the link functions” is correct.
Re ‘which shape?’: if you switch from a logistic curve to clipped linear one (or, I expect, most/all reasonable alternatives) you likely twiddle the exact values and rankings, but the broad order and ~linearly-increasing-over-time sweep of the data should remain much the same. Very loosely, even if the tick-marks on your ruler are all over the place, longer distances measured by multiple rulers laid end to end will still be in the right ballpark.
Re. ‘shared shape’: But you need to assert all the items/benchmarks have identical shape to resolve a single interval scale across them (very loosely again: that all the rulers laid end to end have the same length). Perhaps one way of understanding this goes back to the ‘endogeneity graph’:
Handwaving: one way you could move from the middle panel to either of ones flanking it is to assert the y-axis should be re-parameterized: “True capability is really exponential/logarithmic in the ECI capability trait value”. Yet (loosely) you could bake the same transformation into the capability trait itself, by stipulating a progressive warp of link functions across benchmarks as they increase in difficulty.
For ‘why doubt shared shape?‘, I think there’s a fair case for brute incredulity: why should benchmarks for (e.g.) chess puzzles, science questions, verbal reasoning, (log) time horizons for coding tasks (etc.) all share the same relationship between score progress and general capability? Empirically, browsing some of the ‘benchmark score vs. ECI plots’ on Epoch’s website shows some look more or less S-curvy than others, and the data too noisy and sparse (and likely reliant on tail behaviour) to uniquely identify an logistic curve versus others that look similar.
2) My broad criticisms draw on ~fundamental limitations to IRT or underspecification/ambiguity around ‘AI capability’ - in either case, I don’t think there’s a crisp statistical test or methodological fix. Although I think this covers the endogeneity worry in general (it seems hard for the data alone to distinguish methodological distortion from secular trend), trying to pin down particular niggles is handy in its own right, and offers a gesture at the broader complaint: if the ‘just read the number straightforwardly’ gets the right answer and my metrological quibbling just chases shadows, so much the worse for me—and vice versa.
Your suggestions could be worth doing, although I think the results would only be cruxy in one direction. On my end, I’d guess the speeding up result is robust to link function choice, and would only mildly attenuate if you restrict to pre-model benchmarks (or somehow upweight them relative to post-model benchmarks) - i.e. I wouldn’t take positive results as confirmatory proof. Yet if you found (e.g.) speeding up is sensitive to link function choice, that seems much stronger negative evidence.
Some further ideas, although I’m not sure they’re worth much:
I understand there’s ‘quasi-parametric’ IRT which lets you twiddle the link functions—and anyway you could play with simulation data to see how shape-sensitive the results are. By and large, the response curves of the benchmarks certainly tend to be S-shaped, so (e.g.) “But if you only tweak the s-curves slightly between benchmarks you can generate huge de/accelerations” vs. “To ablate results like this you need to stipulate hugely warps which we don’t actually see in the data” could be informative.
I guess all of this will be confounded by reducing power (see fn. 17), but besides testing on earlier or later subsets, you could try and curate an equal discrimination set of benchmarks across the range, or narrow the set to those which bracket the trend break in 2024.
One could review all the benchmarks, and subjectively score them by how ‘reasoning-y/multistep’ each is, then somehow reweigh/compare classes to see if it lines up (i.e. do the benchmarks get more reasoning-y over time? Does this correspond to benchmark discrimination? If we somehow ‘correct’, does the trend break vanish). But even if so (and going back to the problem where external validity is up for grabs), it doesn’t seem crazy to assert that “true/important” capabilities do load more heavily on reasoning at you go up the range, so even if the benchmark measurement axis is rotating, this rotation keeps it parallel to the ‘true capability’ axis.