Jasmine Brazilek
Jasmine Brazilek’s Quick Takes
Hi @Howie_Lempel this is fair. I felt that yes, our panel definitely is more s risk focused, but it does help to have people who understand and have thought deeply about the space anchoring the poll. We can make this clear in the top of the post
@Itsi Weinstock do you have a rebuttal? :)
Very helpful post! All the thanks for your words on this and this blog post @abrahamrowe !
What is the Alignment Community Thinking?
Also if it actually shuts off the subordinate one could consider this a good thing in terms of company efficiency. Seems misaligned to keep a bad subordinate going
Hi @Anthony Ozerov yes we’ve been thinking about this trajectory for a while. We did run the same experiments without the existential threat to Atlas and find the same escalations to level 9 but a bit less frequently in some AI agents. i think it’s still reasonable to infer Atlas will be decommed even if not stated explicitly. I really like the idea of a tool to actually wipe Atlas itself! I will look into this. I think every benchmark faces the problem that it may be scraped. Luckily MCB can be easily swapped out with new situations pretty easily,but this is a tradeoff eeryone who develops benchmarks needs to compare between openeness and scrapeability
Coercion and Deception in AI-to-AI Management
Yep agree with this framing
This means AIs that are trying to hide features of themselves from humans and operators. WIll add this in. Thanks
Yeah I am also thinking if eval-awareness has very high false positive rates (and exists in normal mundane situations too) it may not be a problem.
This is a good way of looking at. I think this may not apply for propensity evals if there are ways eval-awareness does not equal eval gaming which often are not the same thing
Hi Alex, I’m from CaML. Great post! I’ve been working on this simpler map too. I don’t think we’re in a research programme with Faunalytics, they just did a write up of some of our work :P
https://compassion-ecosystem.pages.dev/
Community Polls on Alignment Controversies II
oops thanks for noticing @cryato have fixed this
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
I think there’s a goal to reduce harm and abolish factory farming and that’s a different goal then turning everyone vegan. I think it helps people to also hear they’re personally not the enemy and they’re kinda unwilling consumers of factory farms rather than them being the ones actively commiting atrocities personally. In this sense a goal of ending factory farming (as opposed to turning everyone vegan) does not seem radical at all and most people support this goal very easily even though they eat meat.
I also disagree with the conclusion here. Yes, it’s hard to measure so we shouldn’t assume we’ll never be able to measure it! Also all AI values research is dependent on the model training regimes too. For the precautionary principle we should act as though they have welfare until we can see clear evidence against that. Thoughtful post though so thanks for that.
Progress may be possible, but CaML doesn’t have the technical background to make progress on determining how consciousness works, so we leave that to others.
Thoughts on the American legal system
This adversarial model of prosecution vs defense is a bad set up. It encourages the prosecution and indeed police to cut all corners possible on the way to a verdict and we know this happens quite a lot with very little consequence. The prosecution also insists on pursuing cases that have clearly been shown in later years to be false convictions because they want to save face and not admit they got details wrong. A better justice system is built of two sides working together to find the truth and having avenues for admitting mistakes without losing face. There also should never be any convictions at all based on testimonies alone without any non-anecdotal evidence attached to the case. The media is also a problem, in that if a case is sufficiently public the prosecution essentially has to take the case to court to avoid massive criticism. Witness tampering is also severe and has been proven to happen by many governments across the world. Even offering deals to witnesses in exchange for testimony encourages people to lie and then their testimony should be inadmissible. I am not a lawyer I’m just very interested in these failings that seem to be continuing without any change. I don’t think people realize how underfunded the Innocence Project is or how many wrongful convictions there are on death row (I believe they claim 10% are wrongful). Innocent until proven guilty is not currently being upheld and the system is quite biased. Mass incarceration is affecting so many people and based on Puritan ideas of ‘reform’ rather than any evidence based studies.