Very provocative post. A few minor comments, not criticisms!
Maybe some of your reference cases are more complex than you imagine. Temperature, for example, is not trivial. Depending how you measure it (ideal gas vs electrical conductivity), you get two aifferent seemingly linear scales that agree only at the two points where you force agreement (e.g. Oc and 100C). But also even using Kelvin, it’s not at all obvious that this is the correct linear scale to use. Maybe the true scale is exponential or logarithmic relative to this: (e.g. one could argue that the differencebetween 10 K and 100 K is the same as that between 100 K and 1000 K … or you could argue (if you’re a chemist, that in the correct scale 300 k is 2x 290 K which is 2x 280 K …). Today, after mant years, we have a settled system, but it doesn’t mean anything—anybody working with quantum computers will know that one degree close to 0K is a lot more “difficult” than going from 285 to 284 K, while maybe cosmologists use phrases like “10′s of millions of degrees” where Celsius and Kelvin are basically the same, and only worry if the temperature changes by a million degrees …
All this to point out that the question of measurement is rarely trivial once you get past counting things!
What people do in science and engineering is to make two lists:
what are the tangible outcomes we need? (e.g. what is the cost per kJ of energy delivered to a home? …)
what are the physical, tangible measures we can make? (e.g. how many kg of fuel did we use? what temperature was reached? what was the % completeness of the reaction? …)
And then, they try to build a model in which the available data will predict the desired output. Sometimes it’s possible, but usually it requires a lot of assumptions, which need to be tested and validated continuously.
They always also keep in mind that all models are wrong, but some are useful, and constantly re-check if their model is useful and remind themselves of the ways it is wrong, especially if applied in and area it wasn’t intended for.
In this context, IMHO asking how capable an AI is is just far too difficult and too broad. To use your football analogy, Messi might be better than anyone in the world at scoring with his left foot, but there are guys on my amateur team who would be better at defending a corner than he would.
What’s most realistic is to ask more granular questions (e.g. likelihood to enable a bioterrorist to produce a new virus, likelihood to successfully complete a problem that takes a skilled human 1 week, … ), and then look at more tangible ways to test exactly that. Which of course is what AI Safety labs do. Just as IQ is meaningful but flawed, any measure of general capability will be flawed.
All this was a long-winded way of saying that yours is a great analysis and a valuable contribution, but that with the right mentality, people have found ways to measure and predict in complex situations before. It feels that AI is another level of complexity, but the principles that give us the most insight do not change.
Very provocative post. A few minor comments, not criticisms!
Maybe some of your reference cases are more complex than you imagine. Temperature, for example, is not trivial. Depending how you measure it (ideal gas vs electrical conductivity), you get two aifferent seemingly linear scales that agree only at the two points where you force agreement (e.g. Oc and 100C). But also even using Kelvin, it’s not at all obvious that this is the correct linear scale to use. Maybe the true scale is exponential or logarithmic relative to this: (e.g. one could argue that the differencebetween 10 K and 100 K is the same as that between 100 K and 1000 K … or you could argue (if you’re a chemist, that in the correct scale 300 k is 2x 290 K which is 2x 280 K …). Today, after mant years, we have a settled system, but it doesn’t mean anything—anybody working with quantum computers will know that one degree close to 0K is a lot more “difficult” than going from 285 to 284 K, while maybe cosmologists use phrases like “10′s of millions of degrees” where Celsius and Kelvin are basically the same, and only worry if the temperature changes by a million degrees …
All this to point out that the question of measurement is rarely trivial once you get past counting things!
What people do in science and engineering is to make two lists:
what are the tangible outcomes we need? (e.g. what is the cost per kJ of energy delivered to a home? …)
what are the physical, tangible measures we can make? (e.g. how many kg of fuel did we use? what temperature was reached? what was the % completeness of the reaction? …)
And then, they try to build a model in which the available data will predict the desired output. Sometimes it’s possible, but usually it requires a lot of assumptions, which need to be tested and validated continuously.
They always also keep in mind that all models are wrong, but some are useful, and constantly re-check if their model is useful and remind themselves of the ways it is wrong, especially if applied in and area it wasn’t intended for.
In this context, IMHO asking how capable an AI is is just far too difficult and too broad. To use your football analogy, Messi might be better than anyone in the world at scoring with his left foot, but there are guys on my amateur team who would be better at defending a corner than he would.
What’s most realistic is to ask more granular questions (e.g. likelihood to enable a bioterrorist to produce a new virus, likelihood to successfully complete a problem that takes a skilled human 1 week, … ), and then look at more tangible ways to test exactly that. Which of course is what AI Safety labs do. Just as IQ is meaningful but flawed, any measure of general capability will be flawed.
All this was a long-winded way of saying that yours is a great analysis and a valuable contribution, but that with the right mentality, people have found ways to measure and predict in complex situations before. It feels that AI is another level of complexity, but the principles that give us the most insight do not change.