My original argument was aimed at P1. I argued that there is a class of decisions—decisions about how a model learns and changes—that cannot straightforwardly be reduced to comparisons of downstream consequences within that model.
The discussion has made me realise that this points to a more constructive response to DiGiovanni’s challenge.
Suppose the cluelessness argument is correct. What should an impartial altruist do?
I think the answer is: we should shift part of our attention from predicting the future to preserving our capacity to correct our predictions.
This is not a rejection of expected value. EV remains useful whenever our model is sufficiently informative to support a comparison. The problem is what to do when we cannot know whether it is.
A finite intelligence faces a peculiar problem:
It can discover errors in its model, but it cannot guarantee that its mechanisms for detecting errors will detect all of its errors.
This creates a distinction between two forms of rationality.
Predictive rationality:
Given my model, which action has the best expected consequences?
Corrective rationality:
Given that my model may be wrong, what keeps me capable of discovering and correcting that wrongness?
The second is not simply epistemic humility. It concerns the architecture of the decision-making process itself.
And this matters because all models are necessarily limited. The relevant question is therefore not whether we can construct a model that is complete, but whether we can maintain an intelligence that remains responsive to what its model leaves out.
Consider two strategies. One attempts to maximise the apparent value of actions under a particular model. The other preserves the conditions through which that model can be challenged: listening to independent perspectives, genuinely exchanging information, allowing disagreement to remain visible, acknowledging errors, and changing one’s representation when it no longer accommodates what one encounters.
The second strategy does not necessarily have higher expected value. That is precisely the point. If we cannot reliably compare long-run consequences, we need to consider properties of the intelligence doing the comparing.
This suggests a progression:
Prediction → model uncertainty → correction → meta-correction.
But correction introduces a further problem.
An intelligence can fail not only because its model is wrong, but because the way it responds to being wrong is itself inadequate. It may hear another perspective without allowing that perspective to affect its model. It may encounter disagreement without investigating it. It may acknowledge an error without examining what made the error possible. It may correct a conclusion while preserving the assumptions that generated the error.
So the problem is no longer simply:
“Is my model wrong?”
It becomes:
“Is the way I respond to being wrong itself capable of correction?”
This also gives a more precise answer to the question raised in the comments: what makes a model “open to correction” if all models are flawed?
Not the expectation that correction will make it correct.
Rather:
A model is corrigible when it preserves channels through which what it encounters can change how it represents reality.
And because those channels can themselves fail, we need to examine and change the way we listen, interpret, disagree, update, and correct.
This opens a new domain of rationality: the practice of meta-thinking. Under cluelessness, the impartial altruist should ask not only whether its model is wrong, but what it must do to remain capable of discovering that it is wrong: Whose perspective must it genuinely hear? What information must it allow to challenge its current representation? When confronted with disagreement, must it defend its model, or make the disagreement itself an object of inquiry? When an error is exposed, must it merely correct the conclusion, or examine why the error was possible in the first place? What must it be willing to change, and what must it keep open, for correction to remain possible?
These are not merely questions about acquiring better information about the future. They concern the present conditions under which an impartial altruist can justifiably move from an “is” to an “ought.” The paradox is that the actions required to maintain those conditions are themselves actions for which we need an “ought.”
The impartial altruist therefore has to reason reflexively:
To determine what it ought to do, it must also determine what it ought to preserve about the conditions that make determining what it ought to do possible.
This is what I mean by reflexive intelligence: intelligence begins to reason about the conditions of its own ability to reason.
And I think this gives us a constructive response to cluelessness that is neither a rejection of consequentialism nor an attempt to restore confidence in prediction.
Under genuine cluelessness, rationality does not end. It becomes reflexive.
One more attempt: what if collective decision-making is really a learning problem?
I have noticed that I have become unusually invested in this problem.
Partly because I believe I see a solution, and partly because I seem to keep failing to articulate what I see in a way that makes it possible for other people to see it too.
So I want to try one more time.
This time I will start from my own profession.
I work as an infectious disease modeller and epidemiologist in a national public health institute. Ultimately, my job is about reducing disease burden in populations.
Some of the interventions we work with are remarkably effective. Measles vaccination is one example. Measles is also one of the diseases for which global eradication is conceivable.
And eradication creates a peculiar decision problem.
If we eradicate measles, we don’t merely reduce disease burden for the people alive today. We change the world inherited by every future generation. Eventually, vaccination against measles would no longer be necessary either.
The important point for my argument is not whether this future benefit can be assigned some particular expected value. It is that eradication changes the future decision problem itself.
To eradicate measles, we need a long sequence of connected decisions.
We need surveillance. We need models. We need vaccination. We need to observe what happens. We need to discover where our assumptions were wrong. We need to change the strategy. We need to coordinate with other countries. We need to learn from failures elsewhere.
The decision is therefore not really:
A or B?
It is a continuous process:
What do we currently believe? What should we do given that belief? What happened? What did we get wrong? What does this tell us about the world? What should we believe now? What should we do next?
This brings me back to the cluelessness problem.
We live in an evolving universe. Our models are necessarily incomplete, and the future is not simply unknown in the sense of being a hidden answer waiting to be calculated. The world itself changes.
So perhaps the central problem of decision-making is not how to make the correct decision from a fixed model.
Perhaps it is how to remain capable of correcting the model on which the next decision will depend.
And this introduces another layer.
We can learn about the problem.
But we can also learn about how we learn about the problem.
For example, better surveillance may produce better information. Better information may improve our model. A better model may lead to better decisions. The results of those decisions provide new information, which can improve the surveillance and modelling process again.
So there is a recursive possibility:
we don’t only optimise the solution; we improve the process by which we discover what the solution should be.
That is the point I was trying to get at in my previous post.
But there is a second part of the problem that I had not made sufficiently explicit.
Measles eradication is not an individual decision problem.
No individual can eradicate measles.
No single institution can eradicate measles.
It requires collective action.
And now the problem becomes considerably harder.
We need a collective that can continuously adapt its understanding of a changing world.
But nobody possesses the complete model.
Different people see different things. Different institutions have different information. Different models will be wrong in different ways.
So perhaps the objective should not be to make everybody hold the same model.
Perhaps the objective is to maintain a network in which different imperfect models can continuously correct one another.
This changes what I think “alignment” means.
We cannot necessarily align people on every proposition about the world.
We cannot even guarantee that we will initially agree on the correct solution.
But perhaps we can align on the process by which we discover that we are wrong.
This is where I think something interesting happens.
The things required for this process sound, at first, like ordinary virtues:
Be honest.
Listen.
Be curious.
Say when you are uncertain.
Correct others when you believe they are wrong.
Allow yourself to be corrected.
Distinguish what you observed from what you inferred.
Help others understand information they do not have.
Forgive mistakes sufficiently that people remain willing to participate.
Remain connected even when you disagree.
But viewed from the perspective of a distributed learning system, these are not merely virtues.
They are functional properties of the system.
Honesty protects the fidelity of the information entering the network.
Listening determines whether information actually reaches another person’s model.
Correction provides an error signal.
Curiosity drives exploration of uncertainty.
Forgiveness helps preserve the connections through which future information can travel.
And willingness to be corrected keeps the individual model open to updating.
The interesting thing is that this can become self-reinforcing.
Better relationships allow better information exchange.
Better information exchange allows better collective learning.
Better collective learning can increase trust in the relationships that made it possible.
And the reverse is also possible:
Fear produces concealment.
Concealment produces poorer information.
Poorer information produces worse models.
Worse models produce worse decisions.
Worse decisions produce distrust.
And distrust produces more fear and concealment.
So perhaps collective intelligence is not simply a property of how much information a group possesses.
Perhaps it is partly a property of whether the relationships between its members preserve the capacity for correction.
This also changes how I think about decentralisation.
If we don’t know beforehand what the correct model of the future problem will be, we cannot simply distribute a correct solution from the centre.
Instead, we may need to distribute the capacity to learn.
Not agreement on the answer.
Agreement on the practices that allow answers to be challenged, revised and improved.
This is why I have started thinking about the “verbs of learning”.
Not what everyone should believe.
What everyone should be capable of doing.
Listening.
Questioning.
Explaining.
Correcting.
Updating.
Admitting uncertainty.
Seeking disconfirming information.
Supporting someone else’s learning.
Being willing to change one’s mind.
These behaviours do not automatically produce agreement. Nor do they guarantee that a collective will choose the right goal.
But they may create something more fundamental:
a collective that remains capable of discovering that it is wrong.
And perhaps that is the deeper problem I have been trying to describe.
Decision theory asks:
What should we do?
But under radical uncertainty, and especially when the decision is collective, perhaps we also need to ask:
How do we remain capable of discovering that what we are doing is wrong?
And if that capacity itself can be improved, then we arrive at a recursive problem:
We learn about the world.
We learn how to learn about the world.
And we learn together how to become better at learning together.
I suspect that this is not a new idea in the literature. Many pieces of it almost certainly exist in different disciplines.
But I increasingly suspect that putting these pieces together points toward something important:
The fundamental requirement for collective intelligence may not be agreement. It may be collective corrigibility.
Not a collective that is permanently right.
A collective that remains capable of becoming less wrong.