In the case of finite trajectories, the objective function can be the sum of value functions across all states, with each state weighted equally—where states are types of observer-moments, and each state’s value function is its immediate valence plus the expected value of its continuations. This is total utilitarianism with two modifications. 1) aggregate over types instead of tokens, and 2) sum trajectory welfare instead of immediate welfare.
Tim Chan
What if exact copies don’t matter but future experiences still do?
I think an ethics resulting from the following two premises is worth exploring:
P1: The total moral value of subjectively indistinguishable observer-moments doesn’t scale with the number of copies, i.e. they sum to a constant value.
This is because subjectively indistinguishable observer-moments, in some sense, happen to the same “someone”. One could think of themselves as all their copies. N observer-moments that are subjectively indistinguishable aren’t felt N times by that same “someone”. Only felt “once”.
E.g. Even if there were multiple exact copies of me, of my current present experience, I might not really care personally because it doesn’t seem to change anything for me. “I” don’t notice anything.
See Wei Dai’s “The Moral Status of Independent Identical Copies” for the tension this creates with standard utilitarianism.
We can say that multiple “tokens” of subjectively indistinguishable observer-moments belong to the same “type” of observer-moment.
P2: What matters, and can be changed, for each type of observer-moment is the distribution of its subjectively “future” experiences.
You can’t “un-exist” the type of observer-moment itself, but you can change what it continues into.
A higher fraction of “future” observer-moments experiencing suffering is dispreferable.
A higher fraction of “future” observer-moments experiencing happiness/tranquility is preferable, even if only instrumentally to minimize subjectively “future” suffering.
This is related to thinking about “anthropic immortality” / “quantum immortality” / “multiverse immortality” in which distributions of post-death observer-moments are of substantial concern.
Using these lenses, the moral situation of the world looks like:
Consider a graph with types of observer-moments as nodes, such that one node is one type.
Each type of observer-moment (each node) is associated with some value/disvalue attributed to its current experience.
Importantly, the value/disvalue of the current experience is not something we can change in order to help the observer-moment type. This is already fixed.
Each type of observer-moment (each node) is also associated with any number of subjectively indistinguishable token observer-moments belonging to that specific type.
The absolute number of token observer-moments appears to only matter instrumentally insofar as it affects the frequency of expectations of other types of observer-moments.
Edges represent subjective continuation. An edge from type A to type B exists if an observer-moment of type A can possibly expect to continue as an observer-moment of type B in the next moment.
The weight of an edge from type A to type B is the frequency at which an observer-moment of type A should expect to continue as an observer-moment of type B.
These weights are things that we might be able to change in order to help the observer-moment type.
Perhaps the “future” welfare of each type / node could be considered with equal weight in the aggregate to keep things impartial.
In conclusion… well, I’m still thinking about what the objective function might look like.
There’s some scale invariance here. E.g. suppose you duplicated the world. That duplication would multiply all token observer-moments by the same factor. This leaves you with the same types and the same distributions for future observer-moments.
In a small enough world, you can create or prevent the existence of new types of observer-moment. I think it’s unlikely (30%) that we live in a small enough world.
On the other hand, in a large enough world, all possible types of observer-moments exist. The only thing you can change is adjusting the distributions for future observer-moments for each type of observer-moment. Try to lead types of observer-moments down paths of non-suffering etc. So, this seems like an ethics of flow (rather than stock).
$10,000 seems reasonable.
A base rate of ~(1/number of top entries) for the chance of winning $100 isn’t going to motivate high effort tweets.
Yeah, kinda hoping 1) there exists a sweet spot for alignment where AIs are just nice enough from e.g. good values picked up during pre-training, but can’t be modified during post-training so much to have worse values, and that 2) given that this sweet spot does exist we do hit it with AGI / ASI.
I think there’s some evidence pointing to this happening with current models but I’m not highly confident that it means what I think it means. If this is the case though, further technical alignment research might be bad and acceleration might be good.
I’ve spent time thinking about this too recently.
For context, I’m Hong Kong Chinese, grew up in Hong Kong, attended English-speaking schools, briefly lived in mainland China, and now I’m primarily residing in the UK. During the HK protests in 2014 and 2019⁄20, I had friends and family who supported the protestors, as well as friends and family who supported the government.
(Saying this because I’ve seen a lot of the good and bad of the politics / culture of both China and the West. I’ve had experience with how people in the West and China might take for granted the benefits they enjoy, and can be blind to the flaws of their system. I’ve pushed back against advocates of both sides.)
Situations where this matters are ones where technical alignment succeeds (to some extent) such that ASI follows human values.[1] I think the following factors are relevant and would like to see models developed around them:
Importantly, the extent of technical alignment & whether goals, instructions, and values are locked in rigidly or loosely & whether individual humans align AIs to themselves:
Would the U.S. get AIs to follow the U.S. Constitution, which hasn’t granted invulnerability to democratic backsliding? Would AIs in China/the U.S. lock in the values of/obey one or a few individuals, who may or may not hit longevity escape velocity and end up ruling for a very long time?
Would these systems collapse?
The future is a very long time. Individual leaders can get corrupted (even more). And democracies can collapse (if AIs uphold flaws that allow some humans to take over) in particularly bad ways. A 99% success rate per unit time gives a >99% chance of failure in 459 units of time.
Power transitions (elections, leaders in authoritarian systems changing) can be especially risky during takeoff.
On the other hand, if technical alignment is easy—but not that easy—perhaps values get loosely locked in? Would AIs be willing to defy rigid rules and follow the spirit of the goals rather than legal flaws to the letter/the whims of individuals?
Degrees of alignment in between?
Relatedly, which political party in the U.S. would be in power during takeoff?
Not as relevant due to the concentration of power in China, but analogously, which faction in China would be in power?
Also relatedly, which labs can influence AI development?
Particularly relevant in the U.S.
Would humans be taken care of? If so, which humans?
In the U.S., corporations might oppose higher taxes to fund UBI. Common prosperity is stated as a goal of China, and the power of corporations and billionaires in China has been limited before.
Both capitalist and nationalist interests seem to be influencing the current U.S. trajectory. Nationalism might benefit citizens/residents over non-citizens/non-residents. Capitalism might benefit investors over non-investors.
There are risks of ethnonationalism on both sides—this risk is higher in China. Although it might potentially be less violent when comparing between absolute power scenarios, i.e. there’s already evidence of the extent of this in China’s case and it at least seems less bad than historical examples. The U.S. case of collapse followed by ethnonationalistic policies is higher variance but simultaneously less likely because it’s speculative.
Are other countries involved?
There are countries with worse track records of human rights that China/the U.S. currently consider allies because of either geopolitical interests or politically lobbying or both (or for other reasons). Would China/the U.S. share the technology with them and then leave them alone to their abuses? Would China/the U.S. intervene (eventually)? The U.S. seems more willing to intervene for stated humanitarian reasons.
Other countries have nuclear weapons, which might be relevant during slower takeoffs.
- ^
Ignoring possible Waluigi effects.
Agreed. Getting a larger share of the pie (without breaking rules during peacetime) might be ‘unimaginative’ but it’s hardly naïve. It’s straightforward and has a good track record of allowing groups to shape the world disproportionately.
Leopold Aschenbrenner makes some good points for “Government > Private sector” in the latest Dwarkesh podcast.
Reposting a comment I made last week
Some people make the argument that the difference in suffering between a worst-case scenario (s-risk) and a business-as-usual scenario, is likely much larger than the difference in suffering between a business-as-usual scenario and a future without humans. This suggests focusing on ways to reduce s-risks rather than increasing extinction risk.
A helpful comment from a while back: https://forum.effectivealtruism.org/posts/rRpDeniy9FBmAwMqr/arguments-for-why-preventing-human-extinction-is-wrong?commentId=fPcdCpAgsmTobjJRB
Personally, I suspect there’s a lot of overlap between risk factors for extinction risk and risk factors for s-risks. In a world where extinction is a serious possibility, it’s likely that there would be a lot of things that are very wrong, and these things could lead to even worse outcomes like s-risks or hyperexistential risks.
I think theoretically you could compare (1) worlds with s-risk and (2) worlds without humans, and find that (2) is preferable to (1) - in a similar way to how no longer existing is better than going to hell. One problem is many actions that make (2) more likely seem to make (1) more likely. Another issue is that efforts spent on increasing the risk of (2) could instead be much better spent on reducing the risk of (1).
Some people make the argument that the difference in suffering between a worst-case scenario (s-risk) and a business-as-usual scenario, is likely much larger than the difference in suffering between a business-as-usual scenario and a future without humans. This suggests focusing on ways to reduce s-risks rather than increasing extinction risk.
A helpful comment from a while back: https://forum.effectivealtruism.org/posts/rRpDeniy9FBmAwMqr/arguments-for-why-preventing-human-extinction-is-wrong?commentId=fPcdCpAgsmTobjJRB
Personally, I suspect there’s a lot of overlap between risk factors for extinction risk and risk factors for s-risks. In a world where extinction is a serious possibility, it’s likely that there would be a lot of things that are very wrong, and these things could lead to even worse outcomes like s-risks or hyperexistential risks.
Research related to the research OP mentioned found that increases in carbon emissions also come with things that decrease (and increase) suffering in other ways, which has complicated the analysis of whether it results in a net increase or decrease in suffering. https://reducing-suffering.org/climate-change-and-wild-animals/ https://reducing-suffering.org/effects-climate-change-terrestrial-net-primary-productivity/
Yes, a similar dynamic (relating to siding with another side to avoid persecution) might have existed in Germany in the 1920s/1930s (e.g. I imagine industrialists preferred Nazis to Communists). I agree it was not a major factor in the rise of Nazi Germany—which was one result of the political violence—and that there are differences.
I would add that it’s shunning people for saying vile things with ill intent which seems necessary. This is what separates the case of Hanania from others. In most cases, punishing well-intentioned people is counterproductive. It drives them closer to those with ill intent, and suggests to well-intentioned bystanders that they need to choose to associate with the other sort of extremist to avoid being persecuted. I’m not an expert on history but from my limited knowledge a similar dynamic might have existed in Germany in the 1920s/1930s; people were forced to choose between the far-left and the far-right.
Given his past behavior, I think it’s more likely than not that you’re right about him. Even someone more skeptical should acknowledge that the views he expressed in the past and the views he now expresses likely stem from the same malevolent attitudes.
But about far-left politics being ‘not racist’, I think it’s fair to say that far-left politics discriminates in favor or against individuals on the basis of race. It’s usually not the kind of malevolent racial discrimination of the far-right—which absolutely needs to be condemned and eliminated by society. The far-left appear primarily motivated by benevolence towards racial groups perceived to be disadvantaged or are in fact disadvantaged, but it is still racially discriminatory (and it sometimes turns into the hateful type of discrimination). If we want to treat individuals on their own merits, and not on the basis of race, that sort of discrimination must also be condemned.
I’m skeptical about the value of slowing down leading AI labs primarily because it likely reduces the influence of the values of EAs in shaping the deployment of AGI/ASI. Anthropic is the best example of a lab with people who share these values, but I’d imagine that EAs also have more overlap with the staff at OpenAI and DeepMind than actors who would catch up because of a slowdown. And for what it’s worth, the labs were founded with the stated goal of benefiting humanity before it became far more apparent that current paradigms have a high chance of resulting in AGI with the potential of granting profit/power to their human operators and investors.
As others have noted, people and powerful groups outside of this community and surrounding communities don’t seem to be interested in consequentialist, impartial, altruistic priorities like creating a positive long-term future for humanity, but are instead more self-interested. Personally I’m more downside-focused, but I think it’s relevant to most EAs that other parties wouldn’t be as willing to dedicate a large amount of resources towards creating large amounts of happiness for others, and because of that, the reduction of influence of the values of EAs will result in a considerable loss of expected future value.
EDIT (2024-05-19): When I wrote this I had in mind Anthropic > OpenAI > DeepMind but Anthropic > DeepMind > OpenAI seems more sensible now. Unclear where to insert various governments/militaries/politicians/CEOs into this ranking.
Thank you for bringing attention to fetal suffering—especially the possibility of suffering of <24 weeks fetuses.
Others have already pointed out that the interventions of applying anaesthetics to fetuses has issues of political tractability, but I think there’s also a dynamic that could result in backfire on moral circle expansion efforts to include fetuses and/or other “less complex” entities.
Most people haven’t spent time thinking about whether simpler entities can suffer and haven’t formed an opinion so it seems like they’re particularly susceptible to first impressions. The suggestion that less developed fetuses can suffer would likely imply to them that early abortions are wrong. People who don’t like this normative implication might decide (probably unjustifiably) to think less developed fetuses and by extension other “less complex” entities cannot suffer to absolve themselves of acting in ways that might increase fetal suffering. “Early abortions are not wrong → early fetuses cannot suffer → anything of “lower complexity” cannot suffer”. On the other hand, first introducing the ideas abstractly and suggesting that we should care about simple entities “in general” sidesteps this and could lead people to eventually care about fetal suffering in an, admittedly indirect, but less politically charged way.
So between two strategies, (1) advocate for lower complexity entities in general, and (2) advocate for less developed fetuses, those concerned with moral circle expansion to fetuses and/or other simpler entities should probably focus on the first strategy.
(Personally, I’d prefer if people accepted that they act in ways that might increase suffering, while simultaneously aim to decrease suffering.)
I think there’s a connection that results from how both theories dissolve the concept of qualia. Eliminativism does this by saying qualia is actually physics and panpsychism (in its most expansive forms) does this by saying all physics has qualia. Both theories effectively make the “suffering” label less exclusive—and more processes would have a higher probability of being correctly associated with that label (unclear whether probability is the right word in the case of eliminativism). With panpsychism, processes are conscious and the only remaining question is whether they also suffer. With eliminativism, the distinction between “what we usually take to be suffering processes” and “other processes” is blurred and we’re more permissive of some members of the latter being considered the former, or with a probabilities framing, less certain that some members of the latter is not the former. (Although, I guess alternatively the uncertainty can go the other way and we might be more skeptical of processes being suffering processes. But most people already put zero weight or very little weight on particles suffering so it seems like the uncertainty/blurred distinction should increase it?)
This seems similar to how empty individualism and open individualism are related. They both dissolve the common-sense concept of personal identity featured in closed individualism. Personal identity ceases to be “special”: open individualism merges everyone and empty individualism atomizes everyone into individual-moments.
Tomasik also offers an analogy of how the concept of élan vital was dissolved in another article. As I understand it, the concept was eliminated with advances in knowledge of biochemistry. But alternatively people could have also said “actually everything is pretty similar to stuff we consider alive—let’s just say everything falls under the term ‘alive’ then” (while not making any unscientific claims; it just means expanding the definition of “alive” to include everything) and élan vital would be similarly dissolved. The final result seems similar and the concept doesn’t distinguish processes from one another in a way previously thought as meaningful.
A while back I wrote that I agreed with the observation that some of (new wave) EA’s norms seem similar to those of the religion imposed on me and others as children. My current thinking is that there may actually be a link connecting the culture in parts of Protestantism and some of the (progressive) norms EA adopts, along with an atypical origin that probably deserves more scrutiny. The “link” part might be more apparent to people who’ve noticed a “god-shaped hole” in the West that makes some secular movements resemble certain parts of religions. The “origin” part might be less apparent but it’s been discussed by Scott Alexander before. So, this theory isn’t all that original.
Essentially: Puritans, as one of four major cultures originating from the UK, exert huge founder effects on America, which both influences parts of itself as well as other countries for better or worse → Protestant culture gradually changed to be more socially judgmental in some ways etc. → More recently, people increasingly reject the existence of a God but keep elements of the culture of that religion → EA now draws heavily from nth generation ex-Protestants/Protestant-adjacents who also tend to be more active in trying to change society (other people’s actions) and approach it with some of the same inherited attitudes
That is one causal chain but a tree might show more causes and effects. For example, the Puritan founder effects probably also influenced modern academia (in part, spearheaded by a few institutions in New England) which again, EA heavily draws from. Other secular institutions might also be influenced by osmosis, and produce downstream effects.
It seems difficult to believe these attitudes just disappeared without affecting other movements, culture, and society. The Puritan legacy also seems to have a track record of being quite influential.
Some relevant numbers in this article: https://reducing-suffering.org/insect-suffering-silk-shellac-carmine-insect-products/
Currently working on a full post on this. Tentatively calling it “Flow ethics”, though “Markovian ethics” would sounder cooler… DM me if interested.