A lot of AI safety people I know are excited about Anthropic and I don’t fully understand why. My instinct is to distrust company because they’re the one pushing the AI arms race, autonomously hacking three companies, and IPO’ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities?
Anthropic has done many expensive actions to credibly signal that it cares about AI safety.
Refusing to remove the ‘human must be in the kill chain’ and ‘no mass surveillance of US citizens’ clauses from its contract with the US military. This resulted in them being declared a supply chain risk, and Google/OpenAI signed contracts without such restrictions.
They’ve directed hundreds of millions in funding to AI safety causes (the equivalent for OpenAI is ~150 million).
They lobby in favor of AI safety regulation which is widely regarded as positive. This is in contrast to OpenAI, whose lobbying arm has generally opposed AI safety regulation (and has had some AI astro-turfing scandals).
There has been some genuinely impressive mech interp research coming out of Anthropic. I found the J-space work to be incredibly interesting, one of the most meaningful mech interp results I’d ever seen.
They seem generally pretty candid about the risks of AI development, and actively call for a coordinated slowdown of AI development (ex. https://www.pacingthefrontier.com/).
I think folks could reasonably say that they are engaged in race dynamics that makes them net negative. But it also doesn’t seem crazy to look at the actions above, and admire what they’ve done (especially compared to other actors in the space).
Anthropic has produced a lot of alignment research. On certain theories* of where AI danger comes from, that research has been useful enough that Anthropic’s overall impact is net positive.
*Those theories are wrong. In a sentence: Anthropic’s alignment research is almost entirely centered around how to produce desired behaviors in the short term, with no understanding of how to make an ASI continue to be aligned once you are no longer smart enough to detect misalignment.
This becomes much easier to explain when you allow for the possibility that people are biased or irrational or have conflicts of interest. People want to work on cool problems, they want to be close to the action, they want to get rich, and they want to think of their friends as good people; then they reason backward from that bottom line to determine that Anthropic must be the good guys.
I guess the discussion should not turn so much around shaming and blaming particular companies, but changing the structural incentives driving forward the current wave of AI development. Under current circumstances, any company with any CEO with any personnel would largely be incentivised to take the path that Anthropic, OpenAI etc. have taken.
By ‘this space’ I meant AI Safety. I at least see a lot of Anthropic roles being published on 80k and have seen AI safety clubs direct people towards Anthropic fellowships as they would MATS.
Publicly they claim to be all about safety, have the PBC and LTBT. They seem like the good guys and have had a meteoric rise, blowing past tech incumbents and OpenAI. Then when you read a model card or their responsible use policy you realize it’s a bit of a farce because ultimately if they deem not releasing a model makes them less competitive they will just release anyways.
A lot of AI safety people I know are excited about Anthropic and I don’t fully understand why. My instinct is to distrust company because they’re the one pushing the AI arms race, autonomously hacking three companies, and IPO’ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities?
Anthropic has done many expensive actions to credibly signal that it cares about AI safety.
Refusing to remove the ‘human must be in the kill chain’ and ‘no mass surveillance of US citizens’ clauses from its contract with the US military. This resulted in them being declared a supply chain risk, and Google/OpenAI signed contracts without such restrictions.
They’ve directed hundreds of millions in funding to AI safety causes (the equivalent for OpenAI is ~150 million).
They lobby in favor of AI safety regulation which is widely regarded as positive. This is in contrast to OpenAI, whose lobbying arm has generally opposed AI safety regulation (and has had some AI astro-turfing scandals).
There has been some genuinely impressive mech interp research coming out of Anthropic. I found the J-space work to be incredibly interesting, one of the most meaningful mech interp results I’d ever seen.
They seem generally pretty candid about the risks of AI development, and actively call for a coordinated slowdown of AI development (ex. https://www.pacingthefrontier.com/).
I think folks could reasonably say that they are engaged in race dynamics that makes them net negative. But it also doesn’t seem crazy to look at the actions above, and admire what they’ve done (especially compared to other actors in the space).
Anthropic has produced a lot of alignment research. On certain theories* of where AI danger comes from, that research has been useful enough that Anthropic’s overall impact is net positive.
*Those theories are wrong. In a sentence: Anthropic’s alignment research is almost entirely centered around how to produce desired behaviors in the short term, with no understanding of how to make an ASI continue to be aligned once you are no longer smart enough to detect misalignment.
This becomes much easier to explain when you allow for the possibility that people are biased or irrational or have conflicts of interest. People want to work on cool problems, they want to be close to the action, they want to get rich, and they want to think of their friends as good people; then they reason backward from that bottom line to determine that Anthropic must be the good guys.
I guess the discussion should not turn so much around shaming and blaming particular companies, but changing the structural incentives driving forward the current wave of AI development. Under current circumstances, any company with any CEO with any personnel would largely be incentivised to take the path that Anthropic, OpenAI etc. have taken.
I don’t think they’re respected within this space? But of course ‘this space’ is vague.
By ‘this space’ I meant AI Safety. I at least see a lot of Anthropic roles being published on 80k and have seen AI safety clubs direct people towards Anthropic fellowships as they would MATS.
Publicly they claim to be all about safety, have the PBC and LTBT. They seem like the good guys and have had a meteoric rise, blowing past tech incumbents and OpenAI. Then when you read a model card or their responsible use policy you realize it’s a bit of a farce because ultimately if they deem not releasing a model makes them less competitive they will just release anyways.