By the way, this is a very good poll!
Max Clarke
I think it’s neither necessary nor sufficient for robust alignment. I’m uncertain as to whether it’s possible to get some kind of “fragile” alignment from pretraining. I don’t think robust alignment requires it, but neither do I think that it doesn’t. It definitely doesn’t hurt.
Yes, I agree with that statement. However, answer is related to “stability under reflection”—specifically I think you’re either in or out of an alignment basin (or, that might not be possible). I think if you’re in it, it’s not correct to say “partially aligned”—what you’ve got is something that’s aligned. And if you’re out of it (or there’s no such thing), then what you’ve got is not aligned. Partial alignment to me means preserving some value only under repeated reflection, which I think is plausibly possible but exponentially unlikely (I’d pick a 99.999% disagree option if it was there, basically)
Multipolar worlds will compete away >90% of net value that would otherwise be preserved
Strongly agree
Alignment to specific values is underrated in research relative to control
Yes, I think control is a waste of time. We need actual alignment to actual (universalized) values.
Partially aligned transformative AIs are likely to be stable under reflection
I disagree that “partially aligned” is a statement that has meaning here.
Research into digital mind suffering is sufficiently tractable to work on
I don’t know.
AI alignment to humans will in practice avoid moral catastrophes to animals
Alignment requires a mechanical understanding of good and bad, and it will be clear how to apply it to animals. Note that wild animal suffering arguments imply that the status quo is likely a moral catastrophe. I believe an aligned entity or system would attempt to change that.
AI alignment to humans will in practice avoid moral catastrophes to digital minds
Likewise, alignment requires a mechanical understanding of good and bad, and it will be clear how to apply it to digital minds.
Robust alignment requires alignment-relevant intervention during pretraining
Frankly I neither agree nor disagree with this statement. Robust alignment has nothing to do with the current pre training regime. It should work with or without it.
This is akin to expecting a single paper titled “The Solution to Geopolitical Stability” or “How to Achieve Permanent Marital Bliss.” These are not problems that are solved; they are conditions that are managed, maintained, and negotiated on an ongoing basis.
If I have an automated system which is generally competent at managing, maintaining and negotiating, then can I not say I have the solution to those things? This is the sense in which it means to solve alignment. It is a lofty goal, yes. I don’t think it’s incoherent, but I do tend to think that it means a system that both “wins” (against other systems, including humans) and is “aligned” (behaves to the greatest extent possible, while still winning, in the general best interest of <insert X>). Take that how you will—many do not think aiming for such a thing is advisable, due to the implications of “winning”.
The “to whom” is either not considered part of alignment itself, or it is assumed to be 5. on your list. 1, 2, 3 & 4 would not typically considered “solving alignment”, although 1. is sometimes advocated for by hardcore democracy believers. I personally think if it’s not 5., then it’s not really achieving anything of note, as it still leaves the world with many competing mutually-malign agents.
This is all to defend the technical (but arguably useless) meaning of “solving alignment”. I agree with everything else in your post. It is absolutely a “wicked problem”, involving competition with intelligent adversaries, an evolving landscape, and a moving target.
Just posting my reactions to reading this:
I find that rates are fairly high:
25% of signatories have been accused of financial misconduct, and 10% convicted
That’s really high?? Oh—this is not the giving what we can pledge😅
I estimate that Giving Pledgers are not less likely, and possibly more likely, to commit financial crimes than YCombinator entrepreneurs.
At what stage of YC? I guess that will be answered later. EDIT:
I previously estimated that 1-2% of YCombinator-backed companies with valuations over $100M had serious allegations of fraud.
.
Gina and I eventually decided that the data collection process was too time-consuming, and we stopped partway through. The final dataset includes 115 of the 232 signatories.
Random, alphabetical, or date ordered? Not that it will really matter—although I guess I would expect the earlier pledgers to be more altruistic, maybe more risk taking though.
I found that the punishment of the criminals in my data set correlated extremely poorly with my intuition for how immorally they had behaved. It would be funny if it weren’t sad that one of the longest prison sentences in my data set is from Kjell Inge Røkke, a Norwegian businessman who was convicted of having an illegal license for his yacht.
Ohhh ok 😂😅 Yeah that is funny and sad.
[Milken] was pardoned by Donald Trump in 2020.
😑
While not all Giving Pledge signatories are entrepreneurs, a large fraction are, which makes this a reasonable reference class. (An even better reference class would be “non-signatory billionaires”, of course.)
Agree
Despite this, I can find very little criticism referencing the fact that many of these signatories are criminals.
This is interesting, and is naturally raised by this post (v. interesting by the way). It makes me wonder about their screening practices. I’m guessing a random like me can’t sign up (they check one’s net wealth somehow?) but perhaps that’s all? If any billionaire can sign up, then maybe it’s not really the giving pledge that one should criticize?
Soaking screams food poisoning to me; especially with unclean water. Perhaps this is not a risk if done right, but this could be why it’s not done.
Definitely, for example if people are bikeshedding (vigorously discussing something that doesn’t matter very much)
Another proposal: Visibility karma remains 1 to 1, and agreement karma acts as a weak multiplier when either positive or negative.
So:
A comment with [ +100 | 0 ] would have a weight of 100
A comment with [ +100 | 0 ] but with 50✅ and 50❌ would have a weight of 100 + log10(50 + 50) = 200
A comment with [ +100 | 100✅ ] would have a weight of say 100 * log10(✓100) = 200
A comment with [+0 | 1000✅ ] would have a weight of 0.
Could also give karma on that basis.
However thinking about it, I think the result would be people would start using the visibility vote to express opinion even more...
Would you gift your karma if that option was available?
This is good for calibrating what the votes mean across the responses
A little ambiguous between “disagree karma & upvote karma should have equal weight” and “karma should have equal weight between people”
I think because the sorting is solely on karma, the line is “Everything above this is worth considering” / “Everything below this is not important” as opposed to “Everything above this is worth doing”
I half agree with this, so let me lay out my position. I basically agree with most or not all of your opinions on concrete failings of animal welfare (or IME, mostly animal rights) culture, including that the emphasis on being vegan is currently overdone in advocacy. But I don’t agree that nobody should be convinced to be vegan, nor that we should taboo the word (which simply represents the ideal point on a diet continuum, although that itself will trigger many v*g*n people) because if the arguments for animal welfare actually convince people, then a lot of people being vegan (or, you know, close to it) is just what the result looks like. So we shouldn’t taboo the word, we should stop treating it as all or nothing.
Let’s start with some basics. Factory farms exist right now and cause enormous suffering. Now imagine a world where they don’t exist. Why don’t they exist in that world? Well, either because there’s no demand for animal products from suffering animals, or there’s no supply of them. Both of those run through people being convinced of the arguments—demand drops when people choose to eat differently, and supply gets restricted when there’s enough political will to regulate it away.
Take the supply side first, because I think it’s the strongest case for the author’s view. I’m sympathetic to a welfarist position here. I can imagine a farm with genuinely good living conditions, enriched environments, animals living good lives and being euthanized as peacefully as possible, and personally I think that would be morally acceptable. Many people here would disagree, but grant it for a second. Even that world needs strong political buy in, because the cost incentives always push welfare down and only sustained public pressure holds the line. Buy in comes from people agreeing with the arguments. People agreeing with the arguments comes from someone convincing them. So the convincing has to happen either way. Not everyone has to do it, and I don’t think it’s hypocritical to pay other people to do it, but it has to get done. And it is my opinion that the person doing the convincing is more credible if they have committed themselves to eating less meat, for example by committing to the ideal of being vegan. It’s a costly signal that they believe what they’re saying.
As an aside, I suspect that even in that high welfare world, meat raised to that standard would be expensive enough that most people would basically be flexitarian by default anyway.
Advocacy that says “vegan or nothing” (or more likely implies it through many “microagressions”) is bad, and marginal changes is a better narrative. But it’s obvious to everyone that vegan is the ideal, and avoiding saying that doesn’t really make sense to me. And I think in practice many people should uphold that ideal personally, for the practical benefit of the costly signal.