Pause AI /â Veganish
Lets do a bunch of good stuff and have fun gang!
Pause AI /â Veganish
Lets do a bunch of good stuff and have fun gang!
Brilliant. âmaintaining access to the frontier of animal sufferingâ is so cursed it made me laugh out loud. So many gems here.
I believe Anthropic engages in quite aggressive anti-regulation lobbying with one side of itâs mouth even as itâs all like âoh no, our products are delivering too much value, weâre so scaredâ with the other. And ya, literally pushing forward the frontier all the time to attract investment is suspiciously similar to what an evil company would do, so thank god we know theyâre on our side.
I find the âmake money doing evil stuff, but be slightly less evil than the imagined counterfactualâ theory of change sort of mesmerizing. I can kind of imagine versions of it that work on paper maybe in a sort of trolley problem way, but it smells too clever by half.
In the actual case of AI companies:
It is not obvious how much the existence of Anthropic is zero-sum relative to the rest of the capabilities sector. Part of the case has to be that they exist more or less instead of some other actor, but they might just be increasing the total supply of digital minds. Even if their specific competitors today donât like them, they could still be contributing to the power of the industry through their contribution to eg. tool use through MCP, to the general lobbying accelerationist effort, to the talent pool /â salaries of capabilities work, creating demand for LLM products, attracting investment etc. Even if eventually there is only room for a few AI firms, they are all unprofitable money pits right now anyway, so it probably isnât easy to displace someone else and they could basically be growing the space more than displacing nefarious actors.
It is not obvious that Anthropic produces lower x-risk per dollar invested or token produced or whatever metric than their competitors. I mean, maybe. They write some thoughtful blog posts and some insane nationalist ones, I like some of their âsafetyâ research. The bar is really low, I kind of hate Google and Sam Altman is a known abuser (employees have said so, he lies constantly, and his sister credibly claims he sexually assaulted her). But, if you entertain the perspective that actually ~all of their production ready âsafetyâ work is unserious in the face of the alignment /â disempowerment problems at stake, then it could easily just be a wash. Like, buying from a factory farm with 20% more thumb twiddling, but essentially the same approach to mass torture.
And between those two objections, thereâs nothing really left to be said for the âmake money doing evil stuff, but be slightly less evil than the imagined counterfactualâ theory of change. I will caveat that I personally lack strong confidence on how this nets out per say⌠generally I think the âitâs bad to do evil stuffâ case wins on simplicity. Maybe, like, a true fast follower company that mostly just distilled and did for-profit safety and security work would be cool.
I tend to think safety-washing or moral cover arenât necessarily a huge deal for the companies because EAâs brand isnât that important to most people and AI Safety isnât that important to most people. Like, I donât know that the talent-attracting boost from not being evil is such a big deal. Nihilistic companies do evil stuff all the time in my opinion. Meta was using their LLMs to sensually chat up kids (fully automated grooming! robo-pedos!) and it only made their hiring a little bit harder. Iâm sure there are plenty of creeps who know ML or schmucks who can be made to look the other way for a check.
But I will say, I think the whole âgood guy with an ASI companyâ line of thinking and the for-profit interests promoting it may have significantly corroded the EA AI Safety scene itself and itâs ability to judge right from wrong. 80,000 Hours consultation recommended that I just get any âopsâ job at Google s long as it vaguely related to AI, like wtf. So much for careful philosophy, yâknow, just get right up to the finish line and say âya, idk, I guess just try to be a lab aid for who-ever is making the deadliest humanoid virusesâ. Look at how much hate Holly Elmore got for promoting moratorium advocacy as a cause area. And the level of conflict of interest going all the way to the very top of eg. Coefficient Giving, CEA was never âepistemically virtuousâ shall we say; seems sort of corrupt.
Just a few thoughts. Really good satire. Thanks!
Hey, you seem so sincere, enthusiastic, and well meaning. I appreciate that; one of the things I love about this forum is how many people on here are trying to âsave the worldâ in various sincere ways.
I apologize because I havenât actually read the Abundance book, but I am familiar with the ideas and I used to listen to Ezra Kleinâs NYT feed pretty often around the time it came out.
Where I agree:
better government services would be better; responsive government is good
instilling broadly positive bureaucratic norms (whatever that means)
probably zoning regulations should be more aggressively pro-housing, but I recognize there are trade offs there. I particularly like cheap prefab housing like Graceland portable buildings for example, but also urban infill generally seems broadly good to me.
pro-social scientific research is often good and maybe national governments should fund it more
the opportunity costs of delaying regulatory approval for medicine are often horrific; nations should probably have international reciprocity around drug testing; rapid turn around time for vaccines is a top priority of our time for both normal healthcare and plague prevention reasons
sometimes the state stands in the way of productive activity which would occur otherwise. For example, I did scouting as a kid and I can build a small shelter out of branches and brush in about an afternoon. It is a shame that it is illegal merely to even be homeless and that it is often illegal to improve land.
Where I disagree:
This shouldnât be a priority for EA. Supply-side progressivism via deregulation is distant from priority areas like global poverty abolition, existential risk prevention, and animal welfare. The benefits of eg. reforming zoning policy would be very indirect at best and yet itâs already backed by Open Phil and Emergent Ventures. I do not think there is a case that this should not prioritized more by scope sensitive, broad moral circle philanthropists on the margins.
Abundance is not neglected. It is an extremely mainstream idea popular within the democratic party in the US and elsewhere. It is also not terribly new if you donât mind my saying so. It is the current poster child of self-identified âcenter left reformistsâ who hold a good number of seats in the US legislature. It also already gets a lot of corporate sponsorship from tech and oil money.
There is not an infinite well of purely technocratic government cruft you can just keep cutting for more free lunches; you will hit a lot of very real tradeoffs very immediately.
Environmental health and safety is under-prioritized at least as often as it is over-prioritized. For example, there are well known contemporary stories about unacceptable lead pollution levels or entire poor black neighborhoods being suffocated by fumes at the mere whimsy of a tech oligarch. You hit real tradeoffs quickly and one could be forgiven for not agreeing that there is an infinite spool of environmental red tape left to cut.
Federal procurement requirements may or may not be overused as a policy tool. I donât have strong feelings here and it feels like a purely technocratic point. It is probably not a silver bullet, but I share your enthusiasm for excellence in government.
I strongly disagree with the assertions you are making about âeconomic growthâ as a singular target for welfare. That doesnât pass muster. Especially not to the exclusion of a focus on wealth inequality.
Raising peopleâs material standard of living is more complicated than this book and this post acknowledges. I donât really think that there are credible solutions to cost of living questions on offer here. For example, I think there are unaddressed tensions and unstated assumptions related to whether lowering prices is actually desirable. Lower housing prices could harm homeowners and lowering the CPI is considered contractionary.
Thatâs just my 2 cents. Hopefully some of that made sense :)
What could be more topical on the EAF than the theory of change and ethics of Anthropic PBC?
You were very thorough and I think the listicle format worked well. I largely agree with what you laid out here and I appreciate you doing the footwork of making so much of this more explicit /â legible.
A lot of this stuff feels shady and even cuts against certain justifications I have heard for Anthropic recently from eg. Holden Karnofsky on 80k and Joe Carlsmith in his blog post. It definitely seems worth being clear headed about.
There was a lot in here that felt insightful and well considered.
I agree that thinking about the end state and humanity in the limit is a fruitful area of philosophy with potentially quite important implications. I wrestle with this sort of thing a lot.
One perspective I would note here (I associate this line of thinking with Will McAskill) is that we ought to be immediately aiming for a wiser, more stable sort of middle-ground and then aim for the âend stateâ from there. I think that can make sense for a lot of practical reasons. I think there is enough of a complex truth to what is and isnât morally good that I am inclined to believe the âmoral error as an x-riskâ framing and, as such, I tend to place a high premium on option value. I think, given the practical uncertainties of the situation, I feel pretty comfortable aiming for /â punting to some more general ââprocess of wise deliberationâ over directly locking my current best guess into the cosmos.
That said, yâknow, we make decisions every day and it is still definitely worth tracking what my current best guess is for what ought actually be done with the physical matter and energy extant in the cosmos. I am partial to much of the substance that you put forward here.
âensuring the ongoing existence of sentienceâ
âsentienceâ is a bit tricky for me to parse, but I will put in for positively valenced subjective experience :)
âgaining total knowledge except that knowledge which requires inducing sufferingâ
I mean, sure, why not? I think that sort of thing is cool and inspiring for the most part. There are probably things that would count as âknowledgeâ to me, but which are so trivial that I wouldnât necessarily care about them much. But, yâknow, I will put in for the practical necessity of learning more about the universe as well as the aesthetic/â profound beauty of discovery the rules of the universe and the nature of nature.
âending all sufferingâ
Fuck ya dude! Iâm against evil and suffering seems like a central example of that. There may even be more aesthetic or injustice like things that I would consider evil even in the absence of negatively valenced experience per se which I might also entertain abolishing.
There is a lot to be said about the âend stateâ which you donât really mention here. Like, for example, I think it is good for people to be really, exceptionally happy if we can swing it. I donât know how to think about population ethics honestly.
One issue that really bites for me when I try to picture the end of the struggle and the steady end state is:
people often intrinsically value reproducing
I want immortality
Each person may require a minimum subsistence amount of stuff to live happily (even if we shrink everyone or make provably morally relevant simulations or something)
Finite materials /â scarcity
I have no reasonable way out of this conundrum and I hate biting the âpopulation controlâ bullet. That reeks of, like, âone child policyâ and overpopulation motivated genocides (cf. The Legacy of Indiaâs Quest to Sterilize Millions of Men /â Uttawar forced sterilizations). I think concerns in this general vein about the resources people use and the limits to growth are also pretty closely ties to the not uncommon concerns people have around over population /â climate heads not wanting to have kids.
Also, to make it less abstract, I will admit that my morals /â impulses are fundamentally quite natalist and I would quite like to be a Dad some day. Even if we grant that resource growth exceeds population growth for now, it seems hard to escape the Malthusian trap forever and I think this is a very fundamental tension in the limit.
Wow, I love that you ended your post in questions. I found your thesis compelling; it reminded me of how much value I used to get from more actively networking with and reaching out to people in online EA spaces. Also, I loved that it was short and salient.
What helps you ask for help when it feels uncomfortable?
Knowing relevant people who have signaled they are okay being asked for help on a given topic. Having a personalish connection to people. A lack of fear of stigma or social consequence for asking a dumb question that I shouldnât have needed help with. A sense of worthiness that I am even allowed to ask things of other people in this context.
When was the last time you asked for help, and what happened?
I ask for help multiple times every day. I am a working stiff and my day job is bench work as a technician in a clinical diagnostics lab (microbiology department). I ask the more senior technicians and medical directors for advice constantly, multiple times a day. That usually goes well and people either give me some kind of answer or at least tell me who to ask. The main downside is that it can take up my time and tbh sometimes they donât give me great advice.
Also I ask my wife for help with all the time and that goes great because they are an amazing partner that I am lucky to have! :) I love my wife!
Hey nice! AGI and improvements to representative democracy systems are both right up my alley!
That said, I think the AGI tie in might seem kind of superficial in that having more functional governance and societal coordination mechanisms would help with all sorts of stuff so I think it makes sense to frame this reasonably in a reasonably AGI timeline agnostic sort of way. That said, ya, I see your point that this sort of thing is made all the more dire when thrown into relief by our âtime of troublesâ and âlongtermists on the precipiceâ style thinking. Your call here, but I am sure it is not necessary to believe random âLLMs will change the worldâ predictions to believe that certain democratic reforms make sense.
In my experience, a lot of people in online EA spaces are pretty willing to talk to you if you reach out, so I think youâll have decent luck there if thatâs what youâre after. Not as confident about how to find more serious collaborators for a project like this.
A few ideas I would throw out there for the sake of brainstorming (many or all of which you may already be familiar with):
independent redistricting /â anti gerrymandering schemes
merging voting districts and proportionally allocating positions > requiring majorities (ie. Mixed-member proportional representation) to negate âwinner take allâ/â minority under representation
ranked choice /â transferable voting to diminish spoiler effect
open primaries might be a good idea to disincentivize the party system from filtering for radical candidates as hard
liquid democracy to let people vote directly on issues that matter to them instead of going through their rep at all (eg. imagine being able to disagree with your senator whenever you want and cast your individual .0000002% of a vote directly on whatever issue)
People talk about quadratic voting too which is probably worth knowing something about from a mechanism design standpoint, but in my opinion doesnât really stand out as a solution to anything on its own without a better way of defining what each actors budget of voting credits would actually need be applied to /â split between in any given round.
Also, I definitely second the idea of using a citizenâs assembly. In my opinion, the power of random sampling + time to learn about and focus on an issue is really OP and really under utilized by representative democracies. The statistical mathematics around approximating large populations with small random samples are really underutilized here and working in our favor. Honestly, there is tons of adverse selection in the electoral process (eg. this book deals with some elements of that).
If you havenât seen CGP Greyâs âPolitics in the Animal Kingdomâ series, you might love it! Also the Forward Party in the US tends to push for similar ideas /â platforms, so they might be worth checking out.
I think this kind of work is very valuable! Nation states might yet be the death of us. It has been terrible watch the democratic backsliding and corruption in my own US of A (in fact I will be one of the protestors this 10â18 No Kings Day). Plus, I agree with your sentiment that there is a lot of headroom. Personally, I think this has less to do with the rise of cyberspace and more to do with the fact that existing polities were just never particularly optimized around the sorts of ideals we are aspiring towards here. Classical Age Greece and the revolutionary United States were both slave states with a lot of backwards ideas after all.
In the case of a government locking in their own power, it seems like you are holding the motivations constant and just saying âpower lets you accumulate more powerâ or something right?
The obvious dis-analogy here that I am sure you are aware of on some level, but which I didnât really see you foreground here is that in the case of either the pause bootstrap or the constitutional deliberation bootstrap, the motivations of the actors are themselves in flux for this period. There isnât as clear of a story you can tell here necessarily about why acceleration should occur at all, but I take it the implied accelerant to our explosion is something like âadditional deliberation/â pausing is factually correct and goodâ and that âadditional deliberation/â pausing will improve epistemic conditionâ.
Also, let me just flag that the âconstitutional conventions of ever greater lengthâ example you gave illustrates a world that is gradually locked in for larger and larger stretches of time not merely one where there is an ever increasing amount of deliberation or something. Like, plausibly, that is an account of gradually sliding into lock in first for a one month interval, then for one year interval, etc.
Good stuff though. Iâve been wrestling with this kind of morality laden futurology and âwhat victory looks likeâ a lot lately, not all in the context of AI but also just against Malthusian traps and the wild state of nature. I tend to agree that viatopia, the great reflection, and any really âany scenario where wise deliberation will occur and be acted uponâ are beautiful and desirable waystations.
Ambitious stuff indeed! Thereâs a lot going on here.
I really appreciate discussions about âbig picture strategy about avoiding misalignmentâ.
For starters, in my opinion, solving technical alignment and control such that one could elicit the main benefits of having a âsuperintelligent servantâ are merely one threat model /â AGI-driven challenge. That said, ofc, getting that sort of thing right also basically means the rest of the planning is better left to someone else and if you are willing to additionally postulate a strong âdecisive strategic advantageâ is basically also a win condition for whatever else you could want.
I would point to eg.
robo powered ultra tyranny
gradual disempowerment /â full unemployment /â the intelligence curse
misinformation slop /â mass psychosis
terrorism, scammers, and a flood of many competent robo psychopaths
machines feeling pain and robot rights
accelerated R&D and needing to adapt at machine speeds
as issues that can all still bite more or less even in worlds where you get some level âalignmentâ esp. if you operationalize alignment as more ~ârobust instruction tuning++â rather than ~âoptimizing for the true moral law itselfâ.
That said, takeover by rogue models or systems of models is a super salient threat model in any world where machines are being made to âthinkâ more and better.
I found your list of competing framings which cut against AI Safety quite compelling. Safety washing is indeed all over the place. One thing I didnât see noted specifically is that a pretty significant contingency within EA /â AI Safety works pretty actively on apologetics for hyperscalers because they directly financially benefit and/âor they have kind of groomed themselves into being the kind of person who can work at a top AI Safety lab.
To draw contrast with how this might have been. You donât, for example, see many EAs working at and founding hot new synthetic virology companies in order to âdo biocontainment better than the competitorsâ. Ostensibly, there could be a similar grim logic of inevitably and a sense that âwe ought to do all of the virology experiments first and more responsiblyâ. Then, idk, once weâve built really powerful AIs or learned everything about virology, we can use this to exit the time of troubles. I donât actually know what eg. Anthropicâs plan is for this, but in the case of synthetic virology /â gain of function research you might imagine that once youâve learned the right stuff about all the potential pathogens, you would be super duper prepared to stop them with all your new medical interventions.
Like, I guess I am just noting my surprise at not seeing good old âkeeping safety at the frontierâ /â âracing through a minefieldâ Anthropic show up more in a screed about safety washing. The EA/ârat space in general is one of the few places where catastrophic risk from AI is a priority and the conflicts of interest here literally could not run deeper. This whole place is largely funded by one of the Meta cofounders and there are a lot of very influential EAs with a lot of personal connection to and complete financial exposure to existing AI companies. This place was a safety adjacent trade show before it was cool lol.
Lots of loving people here who really care on some level Iâm sure, but if we are talking about mixed signals, then I would reconsider the mote in our teamâs eye lol.
***
Beyond that, I guess there is the matter of timelines.
I do not share your confidence in short timelines and think interventions that take a while to pay off can be super worthwhile.
Also, idk, I feel like the assumption that it is all right around the corner and that any day now the singularity is about to happen is really central to the views of a lot of people into x-safety in a way that might explain part of why the worldview kind of struggles to spread outside the relatively limited pool of people who are open to that.
I donât know what youâd call marginal or just fiddling around the edges because I would agree that it is bad if we donât do enough soon enough and someone builds a lethally intelligent super mind and it does rise up and game over.
Maybe the only way to really push for x-safety is with If Anyone Builds It style âyou too should believe in and seek to stop the impending singularityâ outreach. That just feels like such a tough sell even if people would believe in the x-safety conditional on believing in the singularity. Agh. Iâm conflicted here. No idea.
I would love it if we could do more to ally with people who do not see the singularity as being particularly near without things descending into idle âsafety washingâ nor âtrust and safetyâ-level corporate bullshit.
Like the âAI is insane hypeâ contingency has some real stuff going for them too. I donât think they are all just blind. In my humble opinion, I also think Sam Altman looks like an asshole when he calls ChatGPT âPhD levelâ and talks about it doing ânew scienceâ. You know, in some sense, if weâre just being cute, then Wikipedia has been PhD level for a while now and it makes less shit up. There is a lot of hype. These people are marketing and sometimes they get excited.
Plus, it gives me bad vibes when I am trying to push for x-safety and I encounter (often quite justified) skepticism about the power levels of current LLMs and I end up basically just having to do marketing work or whatever for model providers. Idk.
Iâm pretty sure LLM providers arenât even profitable at this point and general robotics isnât obviously much more âright around the cornerâ than it wouldâve seemed to disinterested layperson over the past few decades. Iâm conflicted on this stuff; idk how much effort should go into âsingularity is nearâ vs âif singularity, then doom by defaultâ.
Red lines and RSPs are actually probably a pretty good way of unifying âsingularity nearâ x-safety people with âsingularity farâ or even âsingularity who?â x-safety allies.
***
As far as strategic takeaways:
I do think it is good sense to âbe readyâ and have good ideas âsitting aroundâ for when they are needed. I believe there was a recent UN general assembly where world leaders were literally asking around for, like, ideas for AI red lines. If this is a world where intelligent machines are rising, then there is a good chance we continue to see signs of that (until we donât). The natural tide of âoh shit guysâ and âwow this is realâ may be attenuated somewhat by frog boiling effects, but still. Also, the weirdness of AI Safety regulation and such under consideration will benefit from frog boiling.
Preparedness seems like a great idle time activity when the space isnât receiving the love/âattention it deserves :) .
âI dont think its undemocratic for Trump to be elected for a 3rd term, so long as proper procedures are followed here and he wins the election fairly.â
I can kind of see where you are coming from. I would invite you to consider that sometimes even that sort of thing could be bullshit /â tyranny cf. the Enabling Act of 1933.
Also, for resolution criteria:
âOther markets i would suggest would be on imprisonment/âmurder of political opponents and judges. I would suggest markets like âwill at least 4 of the following 10 people be imprisoned or murdered by Dec 31 2028âł, etc.â
Do you think specific targets would generally have been easy enough to call in advance for other autocracies /â self coups? That seems non-obvious to me?
Ya, I think thatâs right. I think making bad stuff more salient can make it more likely in certain contexts.
For example, I can imagine it to be naive to be constantly transmitting all sorts of detailed information, media, and discussion about specific weapons platforms. Raising awareness that you really hope the bad guys donât develop because it might make them too strong. I just read âPower to the People: How Open Technological Innovation Is Arming Tomorrowâs Terroristsâ by Audrey Kurth Cronin and I think it has a really relevant vibe here. Sometimes I worry about EAs doing unintentional advertisement for eg. bioweapons and superintelligence.
On the other hand, I think that topics like s-risk are already salient enough for other reasons. Like, I think extreme cruelty and torture have arisen independently at a lot of times throughout history and nature. And there are already ages worth of pretty unhinged torture porn stuff that people write which exist already on a lot of other parts of the internet. For example, the Christian conception of hell or horror fiction.
This seems sufficient to say we are unlikely to significantly increase the likelihood of âblind grabs from the memeplexâ leading to mass suffering. Even cruel torture is already pretty salient. And suffering is in some sense simple if it is just âthe opposite of pleasureâ or whatever. Utilitarians commonly talk in these terms already.
I will agree that I donât think itâs good to carelessly spread memes about specific bad stuff sometimes. I donât always know how to navigate the trade offs here; probably there is at least some stuff broadly related to GCRs and s-risks which is better left unsaid. But also a lot of stuff related to s-risk is there whether you acknowledge it or not. I submit to you that surely some level of âraise awareness so that more people and resources can be used on mitigationâ is necessary/âgood?
What dynamics do you have in mind specifically?
Always a strong unilateralist curse with infohazard stuff haha.
I think it is reasonably based and there is a lot to be said for hype, infohazards, and the strange futurist x-risk warning to product company pipeline. It may even be especially potent or likely to bite in exactly the EA milieu.
I find the idea of Waluigi a bit of a stretch given that âwhat if the robot became evilâ is a trope. And so is the Christian devil for example. âEvilâ seems at least adjacent to âstrong value pessimizationâ.
Maybe a literal bit flip utility minimizer is rare (outside of eg extortion) and talking about it would spread the memes and some cultist or confused billionaire would try to build it sort of thing?
Thanks for sharing, good to read. I got most excited about 3, 6, 7, and 8.
As far as 6 goes, I would add that I think it would probably be good if AI Safety had a more mature academic publishing scene in general and some more legit journals. There is a place for the Alignment Forum, arXiv, conference papers, and such but where is Nature AI Safety or equivalent.
I think there is a lot to be said for basically raising the waterline there. I know there is plenty of AI Safety stuff that has been published for decades in perfectly respectable academic journals and such. I personally like the part in âComputing Machinery and Intelligenceâ where Turing says that we may need to rise up against the machines to prevent them from taking control.
Still, it is a space I want to see grow and flourish big time. In general, big ups to more and better journals, forums, and conferences within such fields as AI Safety /â Robustly Beneficial AI Research, Emerging Technologies Studies, Pandemic Prevention, and Existential Security.
EA forum, LW, and the Alignment Forum have their place, but these ideas ofc need to germinate out past this particular clique/âbubble/âsubculture. I think more and better venues for publishing are probably very net good in that sense as well.
7 is hard to think about but sounds potentially very high impact. If any billionaires ever have a scary ChatGPT interaction or a similar come to Jesus moment and google âhow to spend 10 billion dollars to make AI safeâ (or even ask Deep Research), then you could bias/â frame the whole discussion /â investigation heavily from the outset. I am sure there is plenty of equivalent googling by staffers and congresspeople in the process of making legislation now.
8 is right there with AI tools for existential security. I mostly agree that an AI product which didnât push forward AGI, but did increase fact checking would be good. This stuff is so hard to think about. There is so much moral hazard in the water and I feel like I am âvibe capturedâ by all the Silicon Valley money in the AI x-risk subculture.
Like, for example, I am pretty sure I donât think it is ethical to be an AGI scaling/âracing company even if Anthropic has better PR and vibes than Meta. Is it okay to be a fast follower though? Compete in terms of fact checking, sure but is making agents more reliable or teaching Claude to run a vending machine âsafetyâ or is that merely equivocation.
Should I found a synthetic virology unicorn, but we will be way chiller than other synthetic virology companies. And itâs not completely dis-analogous because there are medical uses for synthetic virology and pharma companies are also huge capital intensive high tech operations who spend 100s of millions on a single product. Still, that sounds awful.
Maybe you think armed balance of power with nuclear weapons is a legitimate use case. It would still be bad to do a nuclear bomb research company that lets you scale and reduces costs etc. for nuclear weapons. But idk. What if you really could put in a better control system than the other guy? Should hippies start military tech startups now?
Should I start a competing plantation that, in order to stay profitable and competitive with other slave plantations uses slave labor and does a lot of bad stuff. And if I assume that the demands of the market are fixed and this is pretty much the only profitable way to farm at scale, then so as long as I grow my wares at a lower cruelty per bushel than the average of my competitors am I racing to the top? It gets bad. Same thing could apply to factory farming.
(edit: I reread this comment and wanted to go more out of my way to say that I donât think this represents a real argument made presently or historically for chattel slavery. It was merely an offhand insensitive example of a horrific tension b/âw deontology and simple goodness on the one hand and a slice of galaxy brained utilitarian reasoning on the other.)
Like I said, so much moral hazard in the idea of âAGI company for good stuffâ, but I think I am very much in favor of âAI for AI Safetyâ and âAI tools for existential security. I like âfact checkingâ as a paradigm example of a prosocial use case.
Hey, cool stuff! I have ideated and read a lot on similar topics and proposals. Love to see it!
Is the âThinking Toolsâ concept worth exploring further as a direction for building a more trustworthy AI core?
I am agnostic about whether you will hit technical paydirt. I donâłt really understand what you are proposing on a âgears levelâ I guess and Iâm not sure I could make a good guess even if I did. But, I will say that I think the vibe of your approach sounded pleasant and empowering. It was a little abstract to me I guess Iâm saying, but that need not be a bad thing maybe youâre just visionary.
It reminds me of the idea of using RAG or Toolformer to get LLMs to âshow their workâ and âcite their sourcesâ and stuff. There is surely a lot of room for improvement there bc Claude bullshits me with links on the regular.
This also reminds me of Conjectureâs Cognitive Emulation work and even just Max Tegmark and Steve Omohundroâs emphasis on making inscrutable LLMs to use deterministic proof checkers heavily to win back certain gaurantees.
Is the âLED Layerâ a potentially feasible and effective approach to maintain transparency within a hybrid AI system, or are there inherent limitations?
I donât have a clear enough sense of what youâre even talking about, but there are definitely at least some additional interventions you could run in addition to the thinking tools⌠eg. monitoring, faithful CoT techniques for marginally truer reasoning traces, you could run probes, Anthropic runs a classifier to help with robust jailbreaking for misuse etc. âŚ
I think that something like âdefense in depthâ is something like the current slogan of AI Safety. So, sure I can imagine all sorts of stuff you could try to run for more transparency beyond deterministic tool use, but w/âo a cleaer conception of the finer points it feels like I should say that there are quite an awful lot of inherent limitations, but plenty of options /â things to try as well.
Like, ârobustly managing interpretabilityâ is more like a holy grail than a design spec in some ways lol.
What are the biggest practical hurdles in considering the implementation of CCACS, and what potential avenues might exist to overcome them?
I think that a lot of what it is shooting for is aspirational and ambitious and correctly points out limitations in the current approaches and designs of AI. All of that is spot on and there is a lot to like here.
However, I think the problem of interpeting and building appropriate trust in complex learned algorithmic systems like LLMs is a tall order. âTransparency by designâ is truly one of the great technological mandates of our era, but without more context it can feel like a buzzword like âsecurity by designâ.
I think the biggest âbarrierâ I can see is just that this framing just isnât sticky enough to survive memetically and people keep trying to do transparency, tool use, control, reasoning, etc. under different frames.
But still, I think there is a lot of value in this space and you would get paid big bucks if you could even marginally improve current ablity to get trustworthy interpretable work out of LLMs. So, yâknow, keep up the good work!
Thanks, itâs not that original. I am sure I have heard them talk about AIs negotiating and forgetting stuff on the 80,000 Hours Podcast and David Brin has a book that touches on this a lot called âThe Transparent Societyâ. I havenât actually read it, but I heard a talk he gave.
Maybe technological surveillance and enforcement requirements will actually be really intense at technological maturity and you will need to be really powerful and really local and need to have a lot of context for whatâs going on. In that case, some value like privacy or âbeing aloneâ might be really hard to save.
Hopefully, even in that case, you could have other forms of restraint. Like, I can still imagine that if something like the orthogonality thesis is true, then you could maybe have a really really elegant, light-touch special focus anti super-weapons system that feels fundamentally limited to that goal in a reliable sense. If we understood the cognitive elements enough that it felt like physics or programming, then we could even say that the system meaningfully COULD NOT do certain things (violate the prime directive or whatever) and then it wouldnât feel as much like an omnipotent overlord as a special purpose tool deployed by local LE (because this place would be bombed or invaded if it could not prove it had established such a system).
If you are a poor peasant farmer world, then maybe nobody needs to know what your people are writing in their diaries. But if you are the head of fast prototyping and automated research at some relevant dual use technology firm, then maybe there should be much more oversight. Idk, there feels like lots of room for gradation, nuance, and context awareness here, so I guess I agree with you that the âproblem of libertyâ is interesting.
There was a lot to this that was worth responding to. Great work.
I think making God would actually be a bad way to handle this. I think you could probably stop this with superior forms of limited knowledge surveillance. I think there are likely socio-technical remedies to dampen some of the harsher liberty-related tradeoffs here considerably.
Imagine, for example a more distributed machine intelligence system. Perhaps itâs really not all that invasive to monitor that youâre not making a false vacuum or whatever. And it uses futuristic auto-secure hyper-delete technology to instantly delete everything it sees that isnât relevant.
Also the system itself isnât all that powerful, but rather can alert others /â draw attention to important things. And system implementation as well as the actual violent /â forceful enforcement that goes along with the system probably can and should also be implemented in a generally more cool, chill, and fair way than I associate with the Christian God centralized surveillance and control systems.
Also, a lot of these problems are already extremely salient for âhow to stop civilization ending superweapons from being createdâ-style problems we are already in the midst of here in 2025 Earth. It seems basically true that you do ~need to maintain some level of coordination with /â dominance over anything that could/âmight make a super weapon that could kill you if you want to stay alive indefinitely.
Ya, idk, I am just saying that the tradeoff framing feels unnatural. Or, like, maybe thatâs one lens, but I donât actually generally think in terms of tradeoffs b/âw my moral efforts.
Like, I get tired of various things ofc, but itâs not usually just cleanly fungible b/âw different ethical actions I might plausibly take like that. To the extent it really does work this way for you or people you know on this particular tradeoff, then yep; I would say power to ya for the scope sensitivity.
I agree that the quantitative aspect of donation pushes towards even marginal internal tradeoffs here mattering and I donât think I was really thinking about it as necessarily binary.
I agree with 1, but I think the framing feels forced for point #2.
I donât think itâs obvious that these actions would be strongly in tension with each other. Donating to effective animal charities would correlate quite strongly with being vegan.
Homo economicus deciding what to eat for dinner or something lol.
I actually totally agree that donations are an important part of personal ethics! Also, I am all aboard for the social ripple effects theory of change for effective donation. Hell yes to both of those points. I might have missed it, but I donât know that OP really argues against those contentions? I guess they donât frame it like that though.
I appreciate this survey and I found many of your questions to be charming probes. I would like to register that I object to the âis elitism good actually?â framing here. There is a very common way to define the term âelitismâ that is just straightforwardly negative. Like, âelitismâ implies classist, inegalitarian stuff that goes beyond just using it as an edgelord libertarian way of saying âmeritocracyâ.
I think there is a lot of conceptual tension between EA as a literal mass movement and EA as an usually talent dense clique /â professional network. Probably there is room in the world for both high skill professional networks and broad ethical movements, but yâknow âŚ
I think a real life scenarios where AI kills the most people today is governance stuff and military stuff.
I feel like I have heard the most unhinged haunted uses of LLMs in government and policy spaces. I think that certain people have just âlearned to stop worrying and love the hallucinationâ. They are living like it is the future already and getting people killed with their ignorance and spreading /âusing AI bs in bad faith.
Plus, there is already a lot of slaughter bot stuff going on eg. âRobots Firstâ war in Ukraine.
Maybe job automation is worth saying too. I believe Andrew Yangâs stance for example is that it is already largely here and most people just do have less labor power already, but I could be mischaracterizing this. I think âjobs stuffâ plausibly shades right into doom via âindustrial dehumanizationâ /â gradual disempowerment. In the mean time it hurts people too.
Thanks for everything Holly! Really cool to have people like you actively calling for international pause on ASI!
Hot take: Even if most people hear a really loud ass warning shot, it is just going to fuck with them a lot, but not drive change. What are you even expecting typical poor and middle class nobodies to do?
March in the street and become activists themselves? Donate somewhere? Post on social media? Call representatives? Buy ads (likely from Google or Meta)? Divest in risky AI projects? Boycott LLMs/âcompanies?
Ya, okay, I feel like the pathway from âworryâ to any of that if generally very windy, but sure. I still feel like that is just a long way from the kind of galvanized political will and real change you would need for eg. major AI companies with huge market cap to get nationalized or wiped off the market or whatever.
I donât even know how to picture a transition to an intelligence explosion resistant world and I am pretty knee deep in this stuff. I think the road from here to good outcome is just too blurry for much a lot of the time. It is easy to feel and be disempowered here.
Hey, good list fellow traveler.
In the spirit of collaboration, might I just add a few that came to mind. They may have occurred to you and simply not met your criteria.
Ecocide and ecological destruction is a broader issue than climate change. Perhaps, that is the worst of the bunch, but I believe that industrial scale destruction of land, air, and water is noteworthy. Look at eg. ocean acidification, poisoned water shelves, soil erosion, bio-accumulation of forever chemicals. I find that EA can be kind of a contrarian space such that when ecological issues are brought up it is mostly to emphasize how little we care, but these are public goods we all rely on and the piper will demand a certain level of payment.
Also, people should work less. They should have better conditions. A significant fraction of the world does back breaking work for shit money and thatâs not utopian. A bunch of the work is pretty much pointless anyways, itâs just that the workers lives arenât considered worthwhile because they are poor.
People should live longer lives and have less disease.
People should have more friends and better communities. A lot of people smoke cigarettes and scroll Facebook and theyâd probably rather do something better in a more perfect world.
There should be less war.
Also, a few points where I at least superficially disagree with you:
âgrowthâ as such is not an obvious target to me. I feel there is a bit of slight of hand here where what people mean to say is âit is good when more good stuff happensâ and what they end up saying is âit is good when commerce of any kind happensâ. You actually approach some of the most significant aspects of this critique in your critique of slave cotton and cruel farms, so I suspect weâre not as far apart on this as we could be.
GDP for example often comes up as the implied metric of âeconomic growthâ and I would argue that it actually includes a lot of bad things as well. Doing evil commerce is bad. Enclosing previously non-economic spheres of life such that they are now added to the index is bad. All manner of environmental destruction and poisoning the commons are bad and the long term negative impacts are not suitably captured by this index. Working people into pain and despair is bad. And depending on what is actually being produced, elements of this index can also just be irrelevant or unnecessary for the further increase of human living standards.
It feels like a case of âgraph madnessâ to simplify the whole world into just âmore stuffâ being produced and to be a complete nihilist about what that stuff actually is. Certainly, some forms of production are more or less essential to human well being. I would argue thereâs not much reason to think that this index captures all of that nuance even under relatively favorable âgraph worldâ assumptions. In the real world, especially given youâve already shown a commendable willingness to differentiate between âwho has wealth right nowâ and âmoralityâ, it seems even less likely that GDP would have much to say in the face of these obstacles.
I would argue that these problems are so deep that it is not even productive as short hand to say that it is good in terms of human well being when GDP goes up. I found your explanation really appealingly parsimonious and Iâm not being facetious when I say it resonated more than most cases for growth.
âEconomic growth means increasing the amount of goods that a society produces. This increases how resources people have, which increases their quality of life.â
But I would invite you to maybe interrogate:
If that is what growth means, a lot of economic activity is bullshit. Environmentally harmful activities could easily leave us with less. If we are doing commerce instead of something else, the pie might not have grown at all so much as become commercialized.
Even if the pie did get bigger, rising tides need not lift all ships. Consider the case of eg. wages stagnation, cost of living rises, and investors hoovering up all of the actual profit being made; arguably kind of the default case. I would argue it is in general quite relevant who is materially benefiting from commerce/âindustry and âgrowing the pieâ is just a bulwark against this kind of analysis. Real, âdonât look behind the curtainâ shit.
There are a lot of things other than commerce that make life worth living. If I work more and spend more, thatâs more commerce, but I would often enjoy my life less.
Thatâs just a cursory sketch of a critique of growth. I hope you kind of see what Iâm getting at and why âgrowthâ per say might not be exactly worth aiming for in its own right.
To be really honest, I would go as far as to say the idea of âgrowthâ /â GDP as the one true economic index is often employed as sort of just a cynical âpro-wealthy peopleâ talking point. There are certainly true believers, but also it has received a lot financial funding and laundered prestige it doesnât deserve. In my humble opinion, thatâs sort of a recurring problem within the ~âscienceâ of economics; loads of âNobelâ Prize winning economists have said things to that effect. Would you believe it, according to the kingâs top wealth scientist, itâs actually really good that the king is wealthy lol. I critique because I love; my undergrad minor was econ and I am a serial econ book enjoyer, but I also read books about con artists lol.
Slavery never ended in the United States. Chattel slavery did, but forced work of men in chains under threat of violence and torture happens every day until we stop it.
Sometimes they pay them shit money, <$1 a day, but not always. Personally, I donât really see why they bother except so that it can be brought up in conversations like this. Itâs an obviously insulting amount of money and there is no pretext of consent. I guess itâs the same logic as growth in a way, âcommerce washes all sinsâ. The real reason the convicts work is to avoid being assaulted by the guards and/âor thrown into solitary.
The 13th amendment explicitly allows this and under US jurisprudence this is justified as a legal form of slavery. In the end, the civil war just sort of nationalized the slave trade and changed how it operates a bit. They still pick cotton and everything.
Not that it matters, but 90%+ of these people arenât even alleged to have done anything violent. And they sure ainât white collar crimes either lol. For whatever reason, people mostly donât care about slavery as long as itâs happening to the right people I guess. I think most people donât really know because a prison is the perfect place to hide a bunch of people in chains and you get taught in school that slavery is over. They do work for the federal government and also private corporations like Walmart and McDonaldâs [https://ââapnews.com/ââarticle/ââprison-to-plate-inmate-labor-investigation-c6f0eb4747963283316e494eadf08c4e].
You already touched on prison reform, so forgive me for preaching to the converted.
Solid list! Good to have you as an accomplice!