pre-doc at Data Innovation & AI Lab
previously worked in options HFT and tried building a social media startup
founder of Northwestern EA club
Charlie_Guthmann
was just part of PDKU and now back in Philly for my “normal job”. I think this is important!
The problems in the bay (and maybe dc) and elsewhere seem very different. I think in the bay, there is enough critical mass that you could do e.g. sports league teams for the AI safety people. Exercise + social is good. Not sure if like constellation is already doing stuff like this, but maybe just pre pay for some different sports leagues and then distribute out flyers with easy signup to the various orgs and email some people. might punt a few k but would probably very little effort and decent upside. I played soccer with some AIS folks at a park and that was fun. I know I would always appreciate someone coming up to me and asking if I wanted to play on a sport team with like minded people for the next 10 weeks.
Outside the bay, idk man. If you can help start an IRL group for people looking to inform the public/gov or help organize existing EA/Rat group to do the same, that seems like a good way to channel the energy. But otherwise it’s a bit crazymaking to even think about AIS at all because you will find so little support or understanding from everyone around you. The best thing you can probably do without going nuts is to set a reminder to call your senator and congress person once a month or two and otherwise delete twitter and don’t think about it too much (unless you are working an actual AIS remote job, then idk).
In general I suppose I should have always felt the sadness associated with how bad the world is at triaging resources, that is sort of what EA is all about after all, but trying to help with AIS has made me feel this in such a visceral way that doesn’t make me feel good compared to the animal welfare/poverty stuff. Maybe I had already more successfully compartmentalized those.
https://forum.effectivealtruism.org/s/wmqLbtMMraAv5Gyqn
super dense and very well researched summary of some related things.
Agreed, I’d really want more like a RAND policy maker/rationalist to write out some theories of how people could do bio terrorism, and then grab the individual steps and ask viroligists. subject matter experts are usually quite myopic and can’t see the bigger picture.
https://dianzhuo-wang.github.io/ fwiw this guy seems like he might have an interesting perspective.
Re point 2,
https://forum.effectivealtruism.org/posts/CxMusuX8E5hiTXEWX/fruit-picking-as-an-existential-risk
https://forum.effectivealtruism.org/posts/Nc9fCzjBKYDaDJGiX/what-is-the-likelihood-that-civilizational-collapse-would-1
you might want to read these, the short is that it’s not obvious that if we collapse the modern economy that we can actually get back to interstellar, because we will have way less non renewables (phosphorus, oil/coal) at our disposal next time (don’t totally agree, but worth considering).
Point 1, strong agree esp w/ we don’t know if per human (or human descendant) ev is positive or negative (relative to counterfactual, which could be nothing or aliens), wish this was more mainstream dogma here, not sure why longtermists think they can just not engage with this.
Point 3, again I’ll re route to my response to point 2. Covid (the virus) wasn’t even that bad in a sense and it was still catastrophic (for how much it affected society). Imagine something slightly worse than covid + a record heat wave that causes a massive refugee crisis or + a war. I’m not so confident this wouldn’t massively collapse the global economy, and then route back to the fruit picking, we might only get 1-3 tries to go interstellar, so collapsing the global economy prob not == death but would == reduced chance of becoming grabby, which is ~= death from POV of total utilitarian.
We used to call this upper case (ea community/movement) vs lower case (the (meta) philosophy). I think meta normative framework/philsophy is the way I think about it; that is, you can apply EA to most normative frameworks, esp consequentalisty ones (though to many people in the community, the EA framework is actually just applied total utilitarianism). FWIW, and I say this as someone who has many issues with the EA community, the vast majority of EA criticisms from outside the community I see are just highly inaccurate and not worth engaging with from an intellectual POV (but maybe still worth engaging with for movement reputation).
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture) 8
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn’t be able to control those in a meaningful way nor do I think current AI’s could (but wouldn’t be shocked if they already could control the voice in their head if they have it). but I would guess future ai’s will know how to control increasingly large parts of their brain activations.
At the margin, S-risk work in AI is more important than x-risk work³
From a utilitarian pov, It’s not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I’d rather focus on all of us surviving.
Theories of consciousness will lead to actionable understanding of AI consciousness²
Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious.
Current AIs are capable of suffering
I don’t feel confident at all, but the behavior of llms rn sure do remind me of at least elements of stress, discomfort, and happiness.
If animals continue to exist in a post-AGI world, animal suffering will not persist
I don’t have a strong take on if agi or whoever is in control will be more moral than us but I’m guessing we will be a lot richer, and I think most likely whoever is in control won’t want to torture anything (though they might not care much), and if we are alot richer and advanced I’d think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals
Benchmarks will become useless due to eval awareness¹ (read this backwards originally)
Useless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it’s meaningless seems p high. If mech interp is “good enough” maybe the two can remain useful together.
I keep seeing people say pangram is a really good tool and will improve AI safety. I feel decently confident that the end state of AI writing detection is that AI’s can ~perfectly replicate human writing (when they want to) and that the wide scale deployment of this tool is analogous to spamming antibiotics on factory farms. I made a 100$ charitable bet at 1:1 odds that within 3 years AI writing detection will be ~useless. would be very eager to be pointed to any strong arguments for why I’m wrong.
Some musings on meta science here, kind of tangential and half baked. capabilities is a function of intelligence only once you pick a specific technology and you fix a utility function over the tech, otherwise it’s not well defined.
I haven’t done formal probability in long enough to have the exact words, but essentially you can think of building a technology as a markov chain of decisions. Both the path you take and the chance of success at each path are (partly) a function of your (intelligence). Capabilities is taking a utility function defined over the tech tree and multiplying it by the current probabilities/expected number of steps to different parts of the tree.
imagine you are trying to build a spear. You have 2 choices of sticks, 2 binding agents (glue or chords), 2 choices of stones, and only 1 of each choice is going to work. I guess you can think of this in p or bits but essentially you have 1⁄8. chance of a random walk working, given we fix p(success per task) at 100% for simplicity (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ also it already has its own discussion).
Ok so the maximum intelligence in any situation is getting to the absorbing node/final tech in the minimum number of steps (with 100% chance). For any specific technology there does exist some (not necc unique) pathway s.t. there does not exist any other pathway with less (difficulty weighted) steps. So capability is always bounded on a specific tech.
The divergence (ratio of difficulty weighted steps) between the random walk and the correct path is basically the amount of juice on that tech intelligence can give you. This is basically unbounded in theory. Also p(success per task) might be a function of the same underlying architecture as p(correct path taken at node n) idk.
Now multiply a utility function over the tech tree. The increase of capabilities is the delta in utility per step between the two intelligences.
It also gets much more complicated in reality. Paths can be self correcting if when the agent goes down a wrong path there is some probability > 0 that they can realize this, which depends both on the intelligence of the agent and the nodes in the tech path itself. Similarly, some paths are smooth and monotonic among all the reasonable choices and others will punish greedy algorithms. but I think the core intuition is what I said earlier.
https://x.com/anthropicai/status/2082965101083320543?s=46
Claude did some hacking.(Separately, was there some memo I missed—ea forum seems to have ceded all ai discussion to less wrong except second order ipo donation stuff?)
The all else equal logic is true but its basically fully internalized by normal people ++ to the point where I’d say you should be actively pushing the other way. Literally every kid I want to college with just brain off did like consulting after school largely to preserve option value/career capital (not to necc deploy it for impact later but still). People are extremely risk averse from the POV of a social planner. I’m not saying for any given subset of work you shouldn’t think hard about career capital but if you go job agnostic with this mindset it’s highly constraining. Like i’m pretty sure if I was the social planner god I’d be telling way more smart people to “send it”.
Also re tractability: yes and no. I think early in your career you don’t need to even think that hard about “success” unless you are in an extremely technical / research type thing. Like most jobs if you are consistent, responsible, decent, and good at networking you are gonna get most of the career capital juice there. Yes of course doing something really successfully at 23 is a huge boost but you aren’t going to be heavily punished for not having that. Once you are a couple years into your career, I think the onus to “own” some real successes goes up.
Something else that I didn’t see brought up is how an IPO will effect difficultly of pausing.
Thanks for the kind words!
Responding to point 4, I think your response is extremely fair and basically correct, it’s a bit unfair for me to criticize EA for not being something that it obviously isn’t. But I think my response to that is very relevant to you, which is that you should be thinking about what you want EA to be. Should EA be a library, that doesn’t weigh in or put it’s resources behind specific interventions? Or on the other end, it could be extremely “active”. I think right now it’s somewhere in the middle, which isn’t necc bad but always confusing to me. Like you could say “look EA isn’t going to get political that’s out of scope” but isn’t choosing who speaks at EAG political? isn’t giving resources for fellowships that are at minimum utilitarian adjacent political? Why is the line teaching people applied utilitarianism and not directly being an activist organization?
Not saying it should be one or the other but hope people keep thinking about this stuff.
I’m biking and Amtraking to Berkeley to join the plzdontkillme house july 1st. I’m gonna try to interview people along the way about ai/technology/practical philosophy.
You can follow me on insta https://www.instagram.com/charlie.guthmann/ or youtube https://www.youtube.com/channel/UCmTkQjHjs2cVgca3eC5vHcQ
Bear with me as I’m learning how to use social media and do content creation.
Here is the video I made a month ago (with the help of Diego, who is the linked channel), if you want to get a sense of what I’m trying to do. Advice welcome!
thanks, just donated.
lol exactly… good luck imposing your ethical will after the ipo