Jasmine Brazilek
Hi @Anthony Ozerov yes we’ve been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it’s still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
Yep agree with this framing
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I’m from CaML. Great post! I’ve been working on this simpler map too. I don’t think we’re in a research programme with Faunalytics, they just did a write up of some of our work :P
https://compassion-ecosystem.pages.dev/
oops thanks for noticing @cryato have fixed this
I think there’s a goal to reduce harm and abolish factory farming and that’s a different goal then turning everyone vegan. I think it helps people to also hear they’re personally not the enemy and they’re kinda unwilling consumers of factory farms rather than them being the ones actively commiting atrocities personally. In this sense a goal of ending factory farming (as opposed to turning everyone vegan) does not seem radical at all and most people support this goal very easily even though they eat meat.
I also disagree with the conclusion here. Yes, it’s hard to measure so we shouldn’t assume we’ll never be able to measure it! Also all AI values research is dependent on the model training regimes too. For the precautionary principle we should act as though they have welfare until we can see clear evidence against that. Thoughtful post though so thanks for that.
Progress may be possible, but CaML doesn’t have the technical background to make progress on determining how consciousness works, so we leave that to others.
Our current work in this space is on measuring whether AIs take the possibility of consciousness seriously (without being overconfident in one direction or another). So we’re measuring observable behaviors of giving statements and actions inconsistent with believing that AI welfare is clearly impossible or that current AIs are definitely conscious. I agree that current methods can provide at best weak and heavily debatable findings (for the reasons the linked post articulates), though I think that’s importantly different from precisely zero evidence.
In science it’s usually a good instinct to dismiss something this unclear, but there are two issues with that in this case (and some others): First, the issue is enormously important if true. Second, the philosophical difficulty of artificial consciousness means that our current confusion doesn’t provide Bayesian evidence either way: we’d expect ourselves to have basically these opinions in worlds where artificial consciousness is the default and also worlds where it’s impossible.
I definitely agree and am grateful for your opinion. I am not interested in consciousness research, but do believe there is tractability into the idea of AIs causing digital-mind suffering without attempting to solve the consciousness debate.
Thanks Michael, we avoided mentioning post-training to imply that “new paradigm needed” would also count on the “disagree” side of the spectrum. In other words, “disagree” on this question would mean either “post-training is sufficient” or “new paradigms are needed/sufficient”.
This is really cool work! Is there a graph you can show summarizing what the agents were doing turn after turn in this simulation? Is there anything that would validate this is common sense behavior and you have made a reasonable simulation here?
Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going