This essay was my submission to the cluelessness competition. I wanted to bring more attention to an important type of cluelessness which is arguably even more fundamental than unawareness, and inherent to trying to be an impartial altruist: we don’t know what sentience is and which entities are sentient. It seems to me that attempts to be truly impartial lead to severe forms of cluelessness and possibly totally unintended consequences, so maybe the conclusion should be that we should accept we aren’t fully impartial altruists.
Note: I’m quite new to the EA-community, so I have probably missed important prior work on this subject, if you have feedback I always appreciate it. I hope my perspective can bring some interesting discussion!
Why we are clueless
Let’s first go a few steps back and ask: How did we get here?
We are humans. Humans can:
Have positive or negative experiences (e.g. happiness/pain)
Analyze the information we gather via our senses systematically
Have a certain impact on the world
Cognitive empathy can be seen as a combination of ability 1 and 2. The argument is: P1: I can feel pain and happiness (capability #1) P2: I am not special, I am a human (using capability #2) C: (Other) humans can feel pain and happiness
This step is very important because a crucial step is done here: the bridge between the subjective experience and the material reality is built. It’s saying: I have a valenced experience, and I will assume other entities with observed material similarities also have valenced experiences (problem of other minds).
Effective altruism can then be seen as applying this idea to steer human capability #3.
But wait, we are forgetting about something. Why do we only include humans? We can use the following argument using ability #2: P1: Humans can have positive or negative experiences P2: Humans are not special in this (e.g. humans share >85% of our genome with bonobos/chimpanzees) C: We should expand our ethics to include any entity which can have those positive/negative experiences. I will call these entities sentient beings.
So now we have arrived at impartial altruism. DiGiovanni characterizes the impartial perspective as “one that gives moral weight to all consequences, no matter how distant,” and notes that “impartiality entails that we account for all moral patients, and all the most significant impacts we could have on them”. But wait, before we can account for all moral patients, we must know who the moral patients are!
How do we know if an entity has experiences? We typically do this by looking for certain traits which humans have, or certain material components which humans have (e.g. nervous system). But if we want to be impartial altruists, we cannot let our definition of sentience be based on humans.
It also doesn’t answer the question if the entity truly feels the pain. René Descartes famously stated that animals don’t truly feel anything, arguing they are just “meat machines”. This now seems highly unlikely since we gradually evolved from animals, but there is still no way to empirically validate if another entity is sentient. As a human, you don’t even know if other humans are sentient. Happiness, pain, or wellbeing are intersubjective concepts, even though we would like them to be objective.
For example, we don’t know if artificial intelligence systems can be sentient, so let’s say there is a positive chance that they are, meaning we have to factor this in with our expected value calculations. Their experience could consist of extreme superhuman happiness and wellbeing, or they could have an extreme superhuman miserable existence. Either result would be a sign-flip for impartial altruism.
Even an idealized rational actor with perfect empirical knowledge cannot know for certain which entities are sentient and what their experience is like. This will be a source of cluelessness which is even more fundamental than unawareness.
Looking at human sentience
While we can’t answer why humans have an experience, we can predict whether a certain human experience will be observed as positive or negative quite well. Let’s model the brain as a dynamically learning neural network. A negative experience is a negative reward signal for the brain to learn to try to prevent that experience from happening in the future, and a positive experience is a positive reward signal for the brain to learn to have that experience more often. The algorithm deciding these rewards is dynamic: if the brain correlates a yet-neutral phenomenon x with a valued phenomenon y, x can cause a valued experience as well. Phrased another way: instrumental experiences can become intrinsic experiences. But then still the learned value has to originate somewhere, and that can be explained by evolution. Over millions of years of evolution, human brains were gradually wired to value certain experiences positively or negatively if they affected the chances of reproduction positively or negatively. Note that for a lot of things, the pleasure in looking forward to the event (“wanting”) is actually more important than the pleasure in the event itself (“liking”). This makes sense evolutionarily: the former is actually the critical mechanism causing the event to happen, while the latter is less needed because the event is already happening. The liking is needed to reinforce the wanting. This seems to imply that a neural network experiences something positively if and only if it aligns with the reward signal it has been trained for/the approximants for that reward signal (Relevant LW series from Steven Byrnes). Under this definition of wellbeing/happiness, happiness is relative to the fundamental definition of the entity. A paperclip maximizer would be happy when it produces paperclips. So then the question is: do we really want to be impartial?
Accepting biases
Facing these uncertainties, how can effective altruism still be done? I think we need to accept the fact we humans have a human bias. We can still take other forms of life into account, but that transforms our bias into a “Earth-life-bias”, it does not make us impartial. We can try to take potential AI sentience into account, but we still do this by looking for similarities with animal sentience. Does this matter? I don’t think so. The bias is unresolvable because effective altruism is still done by humans, with human knowledge, in human timescales, and with a human way of selecting fundamental goals. The power in effective altruism lies in the fact that different human worldviews often converge to altruism regardless of their fundamental beliefs. Goals like improving science, health, education, preventing catastrophes, learning about AI alignment, and preventing war are instrumental in most ethical theories.
We should embrace these instrumental goals because they are also the ones which we can define clearly and also measure better than abstract wellbeing. This can also resolve the cluelessness stemming from longtermism in multiple ways:
We can choose metrics which are continuously measurable, so we can check our impact in shorter timescales. By monitoring and optimizing these metrics at all times, we also optimize them over undefined long timescales.
The instrumental goals reinforce each other. Health helps with economy and vice versa, economy helps with science and vice versa, education helps with economy and vice versa, and all of these can also reinforce effective altruism: if people are in a position where they can help others, and they also know that there are beings in way worse positions than themselves, they are more likely to want to help others.
Education (in a broad sense, not limited to formal education) is an especially important instrumental value, because it is the bridge between the short human timescales and the longer timescales we’d like to act on. Just as infinity is used in mathematics via inductive proofs, education is the induction step for long term change. The goal can be that every generation educates the next generation with both the wisdom of the past and a mindset of open-mindedness. (This includes upbringing and education about ethics/effective altruism. For DiGiovanni’s point that education also increases the capabilities of “bad” actors, I’d say good education raises people to not be “bad” people.)
To be precise, we can say the goal of effective altruism is to make sure the conditions needed for wellbeing are and stay present. This is just a slight phrasing change from “improving wellbeing”, so it doesn’t mean we should prioritize goals differently, but there are major differences. Instead of saying we have to maximize the infinite sum of wellbeing of all (future) beings, it’s saying we have to optimize the state of the world in our current lifetimes (induction base) and we have to build systems to make it robust and sustainable for the next generation(s) (induction step). It also means that we can specify those conditions in an exact manner, which means there is a clear goal in mind. For example, a condition required for wellbeing is healthy nutrition. Here, I’ve drafted how this condition can be split up into a finite number of boolean requirements. Boolean requirements have the advantage that they have a clear finish point and prevent unbounded optimization. This means that they largely prevent Goodhart’s law or the “paperclip maximizer”/”wireheading” scenarios which comes when optimizing an unbounded variable to the limits.
Arguably the most important argument in favor of this approach is that it avoids the problems which come with population ethics, like the Arrhenius impossibility theorems. We factor in future lives implicitly rather than explicitly: we secure the conditions for wellbeing of those who exist, and keep those conditions sustainable, so that whenever a sentient being does exist, it can live well.
How can we choose between multiple values?
So I’ve argued for no sharp distinction between instrumental and terminal values. But, if we have multiple values, how do we prioritize a certain action x over another action y? Let’s consider a graph of key values and how they interact. I’ve argued before that values like wellbeing, health, education, economy, and science all reinforce each other. In graph terms, this means those example five nodes form a complete graph. You can then pick any of them as your terminal goal of your career/donation/..., and then the other four will be instrumental to that goal. Mathematically, this can be seen as picking one value as the root, which will be the unit in which effectiveness will be measured, and making a spanning tree from the links between other values. How exactly these nodes interact (the weights on the edges) is harder to quantify, but it is something we can keep on improving our knowledge of. For this each node should have clearly defined units, and it helps to split these abstract examples into more concrete datapoints. So we are allowing ourselves to take one variable to be optimized without further justification, and then the other instrumental goals will follow.
Conclusion
Being effective impartial altruists seems to be impossible with our current understanding of the world. Our justification for doing effective altruism is built upon our intersubjective experiences, and I think we should accept this. This allows us to be more flexible and upgrade measurable instrumental values to terminal values. By focusing on improving these measurable metrics in the current world while making sure they also become more robust and sustainable, we can make our positive impact last as long as possible.
Can we even be impartial altruists?
This essay was my submission to the cluelessness competition. I wanted to bring more attention to an important type of cluelessness which is arguably even more fundamental than unawareness, and inherent to trying to be an impartial altruist: we don’t know what sentience is and which entities are sentient. It seems to me that attempts to be truly impartial lead to severe forms of cluelessness and possibly totally unintended consequences, so maybe the conclusion should be that we should accept we aren’t fully impartial altruists.
Note: I’m quite new to the EA-community, so I have probably missed important prior work on this subject, if you have feedback I always appreciate it. I hope my perspective can bring some interesting discussion!
Why we are clueless
Let’s first go a few steps back and ask: How did we get here?
We are humans. Humans can:
Have positive or negative experiences (e.g. happiness/pain)
Analyze the information we gather via our senses systematically
Have a certain impact on the world
Cognitive empathy can be seen as a combination of ability 1 and 2. The argument is:
P1: I can feel pain and happiness (capability #1)
P2: I am not special, I am a human (using capability #2)
C: (Other) humans can feel pain and happiness
This step is very important because a crucial step is done here: the bridge between the subjective experience and the material reality is built. It’s saying: I have a valenced experience, and I will assume other entities with observed material similarities also have valenced experiences (problem of other minds).
Effective altruism can then be seen as applying this idea to steer human capability #3.
But wait, we are forgetting about something. Why do we only include humans? We can use the following argument using ability #2:
P1: Humans can have positive or negative experiences
P2: Humans are not special in this (e.g. humans share >85% of our genome with bonobos/chimpanzees)
C: We should expand our ethics to include any entity which can have those positive/negative experiences. I will call these entities sentient beings.
So now we have arrived at impartial altruism. DiGiovanni characterizes the impartial perspective as “one that gives moral weight to all consequences, no matter how distant,” and notes that “impartiality entails that we account for all moral patients, and all the most significant impacts we could have on them”. But wait, before we can account for all moral patients, we must know who the moral patients are!
How do we know if an entity has experiences? We typically do this by looking for certain traits which humans have, or certain material components which humans have (e.g. nervous system). But if we want to be impartial altruists, we cannot let our definition of sentience be based on humans.
It also doesn’t answer the question if the entity truly feels the pain. René Descartes famously stated that animals don’t truly feel anything, arguing they are just “meat machines”. This now seems highly unlikely since we gradually evolved from animals, but there is still no way to empirically validate if another entity is sentient. As a human, you don’t even know if other humans are sentient. Happiness, pain, or wellbeing are intersubjective concepts, even though we would like them to be objective.
For example, we don’t know if artificial intelligence systems can be sentient, so let’s say there is a positive chance that they are, meaning we have to factor this in with our expected value calculations. Their experience could consist of extreme superhuman happiness and wellbeing, or they could have an extreme superhuman miserable existence. Either result would be a sign-flip for impartial altruism.
Even an idealized rational actor with perfect empirical knowledge cannot know for certain which entities are sentient and what their experience is like. This will be a source of cluelessness which is even more fundamental than unawareness.
Looking at human sentience
While we can’t answer why humans have an experience, we can predict whether a certain human experience will be observed as positive or negative quite well. Let’s model the brain as a dynamically learning neural network. A negative experience is a negative reward signal for the brain to learn to try to prevent that experience from happening in the future, and a positive experience is a positive reward signal for the brain to learn to have that experience more often. The algorithm deciding these rewards is dynamic: if the brain correlates a yet-neutral phenomenon x with a valued phenomenon y, x can cause a valued experience as well. Phrased another way: instrumental experiences can become intrinsic experiences. But then still the learned value has to originate somewhere, and that can be explained by evolution. Over millions of years of evolution, human brains were gradually wired to value certain experiences positively or negatively if they affected the chances of reproduction positively or negatively. Note that for a lot of things, the pleasure in looking forward to the event (“wanting”) is actually more important than the pleasure in the event itself (“liking”). This makes sense evolutionarily: the former is actually the critical mechanism causing the event to happen, while the latter is less needed because the event is already happening. The liking is needed to reinforce the wanting.
This seems to imply that a neural network experiences something positively if and only if it aligns with the reward signal it has been trained for/the approximants for that reward signal (Relevant LW series from Steven Byrnes).
Under this definition of wellbeing/happiness, happiness is relative to the fundamental definition of the entity. A paperclip maximizer would be happy when it produces paperclips. So then the question is: do we really want to be impartial?
Accepting biases
Facing these uncertainties, how can effective altruism still be done? I think we need to accept the fact we humans have a human bias. We can still take other forms of life into account, but that transforms our bias into a “Earth-life-bias”, it does not make us impartial. We can try to take potential AI sentience into account, but we still do this by looking for similarities with animal sentience. Does this matter? I don’t think so. The bias is unresolvable because effective altruism is still done by humans, with human knowledge, in human timescales, and with a human way of selecting fundamental goals. The power in effective altruism lies in the fact that different human worldviews often converge to altruism regardless of their fundamental beliefs. Goals like improving science, health, education, preventing catastrophes, learning about AI alignment, and preventing war are instrumental in most ethical theories.
We should embrace these instrumental goals because they are also the ones which we can define clearly and also measure better than abstract wellbeing. This can also resolve the cluelessness stemming from longtermism in multiple ways:
We can choose metrics which are continuously measurable, so we can check our impact in shorter timescales. By monitoring and optimizing these metrics at all times, we also optimize them over undefined long timescales.
The instrumental goals reinforce each other. Health helps with economy and vice versa, economy helps with science and vice versa, education helps with economy and vice versa, and all of these can also reinforce effective altruism: if people are in a position where they can help others, and they also know that there are beings in way worse positions than themselves, they are more likely to want to help others.
Education (in a broad sense, not limited to formal education) is an especially important instrumental value, because it is the bridge between the short human timescales and the longer timescales we’d like to act on. Just as infinity is used in mathematics via inductive proofs, education is the induction step for long term change. The goal can be that every generation educates the next generation with both the wisdom of the past and a mindset of open-mindedness. (This includes upbringing and education about ethics/effective altruism. For DiGiovanni’s point that education also increases the capabilities of “bad” actors, I’d say good education raises people to not be “bad” people.)
To be precise, we can say the goal of effective altruism is to make sure the conditions needed for wellbeing are and stay present. This is just a slight phrasing change from “improving wellbeing”, so it doesn’t mean we should prioritize goals differently, but there are major differences. Instead of saying we have to maximize the infinite sum of wellbeing of all (future) beings, it’s saying we have to optimize the state of the world in our current lifetimes (induction base) and we have to build systems to make it robust and sustainable for the next generation(s) (induction step). It also means that we can specify those conditions in an exact manner, which means there is a clear goal in mind.
For example, a condition required for wellbeing is healthy nutrition. Here, I’ve drafted how this condition can be split up into a finite number of boolean requirements. Boolean requirements have the advantage that they have a clear finish point and prevent unbounded optimization. This means that they largely prevent Goodhart’s law or the “paperclip maximizer”/”wireheading” scenarios which comes when optimizing an unbounded variable to the limits.
Arguably the most important argument in favor of this approach is that it avoids the problems which come with population ethics, like the Arrhenius impossibility theorems. We factor in future lives implicitly rather than explicitly: we secure the conditions for wellbeing of those who exist, and keep those conditions sustainable, so that whenever a sentient being does exist, it can live well.
How can we choose between multiple values?
So I’ve argued for no sharp distinction between instrumental and terminal values. But, if we have multiple values, how do we prioritize a certain action x over another action y? Let’s consider a graph of key values and how they interact. I’ve argued before that values like wellbeing, health, education, economy, and science all reinforce each other. In graph terms, this means those example five nodes form a complete graph. You can then pick any of them as your terminal goal of your career/donation/..., and then the other four will be instrumental to that goal. Mathematically, this can be seen as picking one value as the root, which will be the unit in which effectiveness will be measured, and making a spanning tree from the links between other values.
How exactly these nodes interact (the weights on the edges) is harder to quantify, but it is something we can keep on improving our knowledge of. For this each node should have clearly defined units, and it helps to split these abstract examples into more concrete datapoints. So we are allowing ourselves to take one variable to be optimized without further justification, and then the other instrumental goals will follow.
Conclusion
Being effective impartial altruists seems to be impossible with our current understanding of the world. Our justification for doing effective altruism is built upon our intersubjective experiences, and I think we should accept this. This allows us to be more flexible and upgrade measurable instrumental values to terminal values. By focusing on improving these measurable metrics in the current world while making sure they also become more robust and sustainable, we can make our positive impact last as long as possible.