Emerging technologies such as artificial intelligence could radically change the trajectory of our civilization. We are building a global community of researchers and professionals working to ensure that this technological transformation does not risk causing suffering on an unprecedented scale (s-risk).
We do research, award grants and scholarships, and host workshops. Our work focuses on advancing the safety and governance of artificial intelligence as well as understanding other long-term risks.
Center on Long-term Risk updates:
Research updates
Our research runs across two streams — the Model Personas agenda (empirical) and Safe Pareto Improvements (conceptual). Below we share progress on each from the first half of the year; for the bigger picture, see our research overview here.
Model Personas agenda
In the first half of the year, our empirical stream has continued the focus on the Model Personas Agenda. This has led to the following publications, in chronological order:
Conditionalization Confounds Inoculation Prompting Results
Concrete research ideas on AI personas
Shaping the exploration of the motivation-space matters for AI safety (led by Maxime Riché with external collaborators)
Inoculation Adapters Improve Upon Inoculation Prompting
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values (led by Jan Betley and Johannes Treutlein from Truthful AI with contributions from Niels Warncke)
The focus area of these publications is selective generalization of personality traits. We have recently decided to narrow down our focus on malevolent personas specifically and are planning to share an updated agenda soon.
In addition to these publications, we have participated as mentors in SPAR and led 3 projects with a total of 15 mentees on studying empirical patterns of out-of-distribution generalization, creating a benchmark for selective learning techniques, and improving inoculation prompting. All projects are expected to result in forthcoming publications.
In the team, we have had the following changes:
Daniel Tan has been on a sabbatical to work with Arcadia Impact’s Alignment team since May, which may be extended to a permanent position.
We are very excited that Vili Kohonen joined the team in June.
Safe Pareto Improvements agenda
In February, we hosted a workshop on proposals for how AI companies might make the implementation of safe Pareto improvements (SPIs) more likely. Attendees from outside CLR gave important feedback on these proposals — particularly on the importance of more conceptual deconfusion and evaluating models for unambiguously bad behaviors — and developed other ideas for how to make SPIs go well.
We incorporated this feedback into a new research agenda on SPIs, which we published in April. Since then, we’ve made progress on each of the agenda’s three streams (evals, conceptual deconfusion, and research automation):
Evals: We posted a brief report on our initial work evaluating models for willingness to make commitments that unwisely lock out SPIs.
Conceptual deconfusion: We’ve written resources on the general frames that we find helpful for thinking clearly about the game theory of SPIs (this post and a draft of a forthcoming post).
Research automation: We’re building a benchmark for comprehension of and reasoning about SPIs.
However, progress is bottlenecked by our very limited capacity on this agenda, so we’ve recently focused on ways to build capacity, like our SPI Fundamentals Program.
SPI Fundamentals Program
To support the new SPI agenda, we are running an SPI Fundamentals Program this August (with applications closing on July 24).
We think SPIs are one of the most robust levers against catastrophic AI conflict, and this program aims to build a pipeline of people working on them.
What it is: A 4-week online course (Aug 3–28), ~5-7 hrs/week. Weekly readings, exercises, Slack discussion, and office hours with Anthony DiGiovanni, who leads CLR’s SPI agenda. The final week splits into a conceptual and an empirical stream, and there’s an optional paid capstone project the week after (Aug 31 - Sep 4). we plan to hire for SPI research roles, so this doubles as a fit test.
Who it’s for: People considering a career in AI conflict reduction, or those in adjacent areas (cooperative AI, agent foundations) who want conflict risk to inform their existing work. Useful backgrounds include game theory, math, econ, decision theory, formal philosophy, CS, or theoretical physics, but no expertise required. No prior engagement with CLR’s research needed.
Deadline: 23:59 GMT, Friday July 24. If you suspect you might be a good fit for this program, or you know someone else who might be, please share and apply through this link.
The full announcement can be read here.
Program updates
Summer Research Fellowship
Our Summer Research Fellowship program started in June with 8 fellows.
Alejandro Wainstock, Benjamin Sturgeon, Fadi Benzaima, Jord Nguyen, June Hunter, and Maxime Cugnon de Sévricourt are working on the model personas agenda.
Ching Lam Choi is working on safe Pareto improvements with external mentor Nathaniel Sauerberg.
Cadence James is working on s-risk macrostrategy.
Research Affiliate program
In June we launched the Research Affiliates program. Research Affiliates will work autonomously on s-risk areas such as:
New research agendas that fall outside CLR’s current focus areas, and other high-upside research bets that need time to develop.
Automating or accelerating conceptual research.
Topics adjacent to our core agendas where the Affiliate has strong independent views or expertise.
We’re interested in taking on more Affiliates and have a form to express interest here.