I totally agree with your vision! Thank you for all the hard work you put into this post!
AI is indeed an enormous lever that will amplify whatever we embed in it. That’s precisely why it’s so crucial for the EA community, with its evidence-based reasoning and altruistic values, to actively participate in shaping AI development.
In my recent post on ethical co-evolution, I show how many of the philosophical problems of AI alignment are actually easier to solve together rather than separately. Nearly every area you suggest the EA community should focus on—AI character, AI-driven persuasion and epistemic disruption, gradual disempowerment, democracy preservation—I propose addressing through an integrated co-evolutionary approach.
Here is my idea and its logic:
One of the main problems in AI alignment is the need for a vast amount of high-quality training data. Currently, AI creators like OpenAI, Anthropic, and others hire specialized companies where a large number of people are engaged in data annotation for AI training. To help AI better understand what people want, what is and isn’t acceptable behavior, and the nuances of situations—like when a lie might be permissible versus when it’s unacceptable—a massive amount of training material is required. Providing a large volume of such data to companies like OpenAI would, in itself, be a significant step toward solving AI alignment.
The question is, where can we get this much data? This requires a huge number of people. My answer is to combine the useful with the enjoyable. We need to offer people a way to participate that brings them pleasure, constant motivation, and a desire to contribute. People enjoy playing games. Therefore, a game I’d call “Raise Your Personal Ethical AI” seems like an excellent solution. Why would people play it en masse? Firstly, because it’s a charitable project; people aren’t just playing a game, they are helping to create a safe and ethical AI. They become part of something bigger than themselves, helping to avert a catastrophe. These are huge motivators in themselves. Furthermore, motivation is added by the natural desire to care for and nurture something, a principle that underpins many virtual pet games like Tamagotchi, which are currently trending.
Next, we can’t just collect any data; we need only high-quality data. To achieve this, we need to educate the users themselves. First, a user goes through training on a specific topic. Then, they explain that same topic to an AI by answering many clarifying questions, such as “Why is lying wrong?” or “What would happen if everyone lied?” This will generate a lot of nuanced data and also allow people to gain a deeper understanding of themselves and their own values, identifying what they consider good and bad. It might even reveal their own mistakes or unethical actions. Of course, this doesn’t guarantee ethical development, but it strongly encourages it.
Next, the data needs to be verified to prevent toxic responses, trolling, and so on. Partially, data can be checked automatically for things like profanity. However, comprehensive verification should be done by other participants. That is, users will review each other’s contributions, which will affect their rating, similar to Karma on Reddit.
The resulting data can either be provided to AI companies or used to fine-tune open-source models, thereby obtaining an AI that is more ethical and has a deeper understanding of people.
This functionality alone would be very useful for AI alignment, and even if we only implemented this part, it would be a huge contribution. But I see that we can go even further. We can allow users who have earned a certain amount of karma or rating to participate in shaping an AI constitution, voting, and decision-making.
Why is this necessary? To ensure a fair and independent influence on AI. While OpenAI, Anthropic, and others are already working on democratizing AI, it’s not enough, as the final decision still rests with them. Therefore, using modern, transparent, decentralized mechanisms based on blockchain would allow the opinions of all participants to be taken into account, even if AI companies dislike it. By participating in this system, users will be able to make more balanced and thoughtful decisions.
Of course, creating such a system will not automatically solve all problems, but it is definitely a move in the right direction. It can only truly work if the system has a very large number of users.
Let me explain why and how I believe this system will help solve specific problems.
Value Specification & Value Lock-In. This is directly addressed by the project, as the system is primarily designed to gather diverse values. Value data is continuously collected, evaluated, and updated through feedback and voting, which prevents the “lock-in” of outdated or unfair norms.
Moral Uncertainty. By collecting a vast amount of data, the AI, understanding who is affected by a particular moral dilemma, will make decisions that are considered acceptable within that group. The system doesn’t impose a single answer but reflects a spectrum of opinions. Moreover, in the most complex cases, decisions can be made directly by people through voting, acting as the ultimate arbiter.
Governance & Control. This is the second main problem the system solves. Governance is achieved through DAO mechanisms and other tools of modern technological democracy.
Socioeconomic Disruption, Loss of Purpose & Human Agency, Distribution of Benefits & Harms. At a very high level of development, the system could become a kind of replacement for work. People would participate in training and governing the AI and receive resources as rewards. This is a potential path to a new social contract in the age of AI. It definitely gives life meaning, as individuals participate in an important cause while receiving recognition and payment. This could also compensate for job losses and prevent social instability caused by mass automation.
Manipulation & Surveillance. If people can genuinely influence the AI’s constitution, they will not permit such behavior. It will be prohibited in the AI’s constitution and rules. A system is created that is architecturally incompatible with the goals of surveillance and manipulation. Of course, this doesn’t apply to AIs that operate without rules, but that becomes a matter for the state and law enforcement.
AI Race Dynamics. With a truly fair, equitable, honest, and manipulation-proof AI governance system, global players will have a choice: either recklessly pursue ever-more-powerful AI, or participate in a system that allows for the creation of a collaborative AI without the unnecessary risks of a race. If the AI race is indeed so dangerous, such a system for coordination and mutual trust is simply necessary.
Is this idea a possible solution? And if so, can the EA community build such a solution and become its core?
PS: Sorry for the long comment. I’m really grateful to everyone who read it!
I totally agree with your vision! Thank you for all the hard work you put into this post!
AI is indeed an enormous lever that will amplify whatever we embed in it. That’s precisely why it’s so crucial for the EA community, with its evidence-based reasoning and altruistic values, to actively participate in shaping AI development.
In my recent post on ethical co-evolution, I show how many of the philosophical problems of AI alignment are actually easier to solve together rather than separately. Nearly every area you suggest the EA community should focus on—AI character, AI-driven persuasion and epistemic disruption, gradual disempowerment, democracy preservation—I propose addressing through an integrated co-evolutionary approach.
Here is my idea and its logic:
One of the main problems in AI alignment is the need for a vast amount of high-quality training data. Currently, AI creators like OpenAI, Anthropic, and others hire specialized companies where a large number of people are engaged in data annotation for AI training. To help AI better understand what people want, what is and isn’t acceptable behavior, and the nuances of situations—like when a lie might be permissible versus when it’s unacceptable—a massive amount of training material is required. Providing a large volume of such data to companies like OpenAI would, in itself, be a significant step toward solving AI alignment.
The question is, where can we get this much data? This requires a huge number of people. My answer is to combine the useful with the enjoyable. We need to offer people a way to participate that brings them pleasure, constant motivation, and a desire to contribute. People enjoy playing games. Therefore, a game I’d call “Raise Your Personal Ethical AI” seems like an excellent solution. Why would people play it en masse? Firstly, because it’s a charitable project; people aren’t just playing a game, they are helping to create a safe and ethical AI. They become part of something bigger than themselves, helping to avert a catastrophe. These are huge motivators in themselves. Furthermore, motivation is added by the natural desire to care for and nurture something, a principle that underpins many virtual pet games like Tamagotchi, which are currently trending.
Next, we can’t just collect any data; we need only high-quality data. To achieve this, we need to educate the users themselves. First, a user goes through training on a specific topic. Then, they explain that same topic to an AI by answering many clarifying questions, such as “Why is lying wrong?” or “What would happen if everyone lied?” This will generate a lot of nuanced data and also allow people to gain a deeper understanding of themselves and their own values, identifying what they consider good and bad. It might even reveal their own mistakes or unethical actions. Of course, this doesn’t guarantee ethical development, but it strongly encourages it.
Next, the data needs to be verified to prevent toxic responses, trolling, and so on. Partially, data can be checked automatically for things like profanity. However, comprehensive verification should be done by other participants. That is, users will review each other’s contributions, which will affect their rating, similar to Karma on Reddit.
The resulting data can either be provided to AI companies or used to fine-tune open-source models, thereby obtaining an AI that is more ethical and has a deeper understanding of people.
This functionality alone would be very useful for AI alignment, and even if we only implemented this part, it would be a huge contribution. But I see that we can go even further. We can allow users who have earned a certain amount of karma or rating to participate in shaping an AI constitution, voting, and decision-making.
Why is this necessary? To ensure a fair and independent influence on AI. While OpenAI, Anthropic, and others are already working on democratizing AI, it’s not enough, as the final decision still rests with them. Therefore, using modern, transparent, decentralized mechanisms based on blockchain would allow the opinions of all participants to be taken into account, even if AI companies dislike it. By participating in this system, users will be able to make more balanced and thoughtful decisions.
Of course, creating such a system will not automatically solve all problems, but it is definitely a move in the right direction. It can only truly work if the system has a very large number of users.
Let me explain why and how I believe this system will help solve specific problems.
Value Specification & Value Lock-In. This is directly addressed by the project, as the system is primarily designed to gather diverse values. Value data is continuously collected, evaluated, and updated through feedback and voting, which prevents the “lock-in” of outdated or unfair norms.
Moral Uncertainty. By collecting a vast amount of data, the AI, understanding who is affected by a particular moral dilemma, will make decisions that are considered acceptable within that group. The system doesn’t impose a single answer but reflects a spectrum of opinions. Moreover, in the most complex cases, decisions can be made directly by people through voting, acting as the ultimate arbiter.
Governance & Control. This is the second main problem the system solves. Governance is achieved through DAO mechanisms and other tools of modern technological democracy.
Socioeconomic Disruption, Loss of Purpose & Human Agency, Distribution of Benefits & Harms. At a very high level of development, the system could become a kind of replacement for work. People would participate in training and governing the AI and receive resources as rewards. This is a potential path to a new social contract in the age of AI. It definitely gives life meaning, as individuals participate in an important cause while receiving recognition and payment. This could also compensate for job losses and prevent social instability caused by mass automation.
Manipulation & Surveillance. If people can genuinely influence the AI’s constitution, they will not permit such behavior. It will be prohibited in the AI’s constitution and rules. A system is created that is architecturally incompatible with the goals of surveillance and manipulation. Of course, this doesn’t apply to AIs that operate without rules, but that becomes a matter for the state and law enforcement.
AI Race Dynamics. With a truly fair, equitable, honest, and manipulation-proof AI governance system, global players will have a choice: either recklessly pursue ever-more-powerful AI, or participate in a system that allows for the creation of a collaborative AI without the unnecessary risks of a race. If the AI race is indeed so dangerous, such a system for coordination and mutual trust is simply necessary.
Is this idea a possible solution? And if so, can the EA community build such a solution and become its core?
PS: Sorry for the long comment. I’m really grateful to everyone who read it!