Nora Belrose

Karma: 261

AI Pause Will Likely Backfire

Nora Belrose16 Sep 2023 10:21 UTC

129 points

165 comments13 min readEA link

Counting arguments provide no evidence for AI doom

Nora Belrose27 Feb 2024 23:03 UTC

61 points

13 comments1 min readEA link

Deconstructing Bostrom’s Classic Argument for AI Doom

Nora Belrose11 Mar 2024 6:03 UTC

25 points

0 comments1 min readEA link

(www.youtube.com)

Nora Belrose 16 Sep 2023 16:38 UTC
16 points
15 ∶ 2
in reply to: Rafael Harth’s comment on: AI Pause Will Likely Backfire
I don’t think the terminal vs. instrumental goal dichotomy is very helpful, because it shifts the focus away from behavioral stuff we can actually measure (at least in principle). I also don’t think humans exhibit this distinction particularly strongly. I would prefer to talk about generalization, which is much more empirically testable and has a practical meaning.

Nora Belrose 16 Sep 2023 15:10 UTC
13 points
12 ∶ 8
in reply to: Rafael Harth’s comment on: AI Pause Will Likely Backfire
The opposing take is that all it’s doing is making the AI play a nicer character, but doesn’t lead it to internalize its goals, which is what alignment is actually about.
I think this is a misleading frame which makes alignment seem harder than it actually is. What does it mean to “internalize” a goal? It’s something like, “you’ll keep pursuing the goal in new situations.” In other words, goal-internalization is a generalization problem.
We know a fair bit about how neural nets generalize, although we should study it more (I’m working on a paper on the topic atm). We know they favor “simple” functions, which means something like “low frequency” in the Fourier domain. In any case, I don’t see any reason to think the neural net prior is malign, or particularly biased toward deceptive, misaligned generalization. If anything the simplicity prior seems like good news for alignment.

Nora Belrose 17 Sep 2023 22:03 UTC
11 points
6 ∶ 3
in reply to: DanielFilan’s comment on: AI Pause Will Likely Backfire
Yep I am aware of the value learning section of Chapter 12, which is why I used the “mostly” qualifier. That said he basically imagines something like Stuart Russell’s CIRL, rather than anything like LLMs or imitation learning.
If we treat the Orthogonality Thesis as the crux of the book, I also think the book has aged poorly. In fact it should have been obvious when the book was written that the Thesis is basically a motte-and-bailey where you argue for a super weak claim (any combo of intelligence and goals is logically possible), which is itself dubious IMO but easy to defend, and then pretend like you’ve proven something much stronger, like “intelligence and goals will be empirically uncorrelated in the systems we actually build” or something.

Nora Belrose 16 Sep 2023 13:25 UTC
11 points
1 ∶ 0
on: AI Pause Will Likely Backfire
Unfortunately, this post got published under the wrong username. I’m the Nora who wrote this post. I hope it can be fixed soon.

Nora Belrose 17 Sep 2023 20:07 UTC
10 points
6 ∶ 2
in reply to: Rafael Harth’s comment on: AI Pause Will Likely Backfire
You need to have some motivation for thinking that a fundamentally new kind of danger will emerge in future systems, in such a way that we won’t be able to handle it as it arises. Otherwise anyone can come up with any nonsense they like.
If you’re talking about e.g. Evan Hubinger’s arguments for deceptive alignment, I think those arguments are very bad, in light of 1) the white box argument I give in this post, 2) the incoherence of Evan’s notion of “mechanistic optimization,” and 3) his reliance on “counting arguments” where you’re supposed to assume that the “inner goals” of the AI are sampled “uniformly at random” from some uninformative prior over goals (I don’t think the LLM / deep learning prior is uninformative in this sense at all).

Nora Belrose 16 Sep 2023 16:57 UTC
10 points
8 ∶ 0
in reply to: Rafael Harth’s comment on: AI Pause Will Likely Backfire
Why does it have to be one or the other? I personally don’t put much stock in what Eliezer and Nate think, but many other people do.

Nora Belrose 16 Sep 2023 16:49 UTC
10 points
4 ∶ 4
in reply to: Zach Stein-Perlman’s comment on: AI Pause Will Likely Backfire
Where we agree:
“dangerous-capability-model-eval-based regulation” sounds good to me. I’m also in favor of Robin Hanson’s foom liability proposal. These seem like very targeted measures that would plausibly reduce the tail risk of existential catastrophe, and don’t have many negative side effects. I’m also not opposed to the US trying to slow down other states, although it’d depend on the specifics of the proposal.
Where we (partially) disagree:
I think there’s a plausible case to be made that publishing model weights reduces foom risk by making AI capabilities more broadly distributed, and also enhances security-by-transparency. Of course there are concerns about misuse— I do think that’s a real thing to be worried about— but I also think it’s generally exaggerated. I also relatively strongly favor open source on purely normative grounds. So my inclination is to be in favor of it but with reservations. Same goes for labs publishing capabilities research.

Nora Belrose 24 Sep 2023 16:50 UTC
8 points
1 ∶ 1
in reply to: RobertM’s comment on: AI Pause Will Likely Backfire
Anticipating the argument that, since we’re doing the training, we can shape the goals of the systems—this would certainly be reason for optimism if we had any idea what goals we would see emerge while training superintelligent systems, and had any way of actively steering those goals to our preferred ends. We don’t have either, right now.
What does this even mean? I’m pretty skeptical of the realist attitude toward “goals” that seems to be presupposed in this statement. Goals are just somewhat useful fictions for predicting a system’s behavior in some domains. But I think it’s a leaky abstraction that will lead you astray if you take it too seriously / apply it out of the domain in which it was designed for.
We clearly can steer AI’s behavior really well in the training environment. The question is just whether this generalizes. So it becomes a question of deep learning generalization. I think our current evidence from LLMs strongly suggests they’ll generalize pretty well to unseen domains. And as I said in the essay I don’t think the whole jailbreaking thing is any evidence for pessimism— it’s exactly what you’d expect of aligned human mind uploads in the same situation.

Nora Belrose 17 Sep 2023 21:53 UTC
7 points
2 ∶ 0
in reply to: Steven Byrnes’s comment on: AI Pause Will Likely Backfire
Yep it’s all meant to be disjunctive and yep it could have been clearer. FWIW this essay went through multiple major revisions and at one point I was trying to make the disjunctivity of it super clear but then that got de-prioritized relative to other stuff. In the future if/when I write about this I think I’ll be able to organize things significantly better

Nora Belrose 17 Sep 2023 22:54 UTC
6 points
0 ∶ 1
in reply to: Zach Stein-Perlman’s comment on: AI Pause Will Likely Backfire
It’s not obvious to me what alignment optimism has to do with the pause debate
Sorry, I thought it would be fairly obvious how it’s related. If you’re optimistic about alignment then the expected benefits you might hope to get out of a pause (whether or not you actually do get those benefits) are commensurately smaller, so the unintended consequences should have more relative weight in your EV calculation.
To be clear, I think slowing down AI in general, as opposed to the moratorium proposal in particular, is a more reasonable position that’s a bit harder to argue against. I do still think the overhang concerns apply in non-pause slowdowns but in a less acute manner.

Nora Belrose 17 Sep 2023 22:19 UTC
5 points
2 ∶ 1
in reply to: Steven Byrnes’s comment on: AI Pause Will Likely Backfire
It’s essentially no cost to run a gradient-based optimizer on a neural network, and I think this is sufficient for good-enough alignment. I view the the interpretability work I do at Eleuther as icing on the cake, allowing us to steer models even more effectively than we already can. Yes, it’s not zero cost, but it’s dramatically lower cost than it would be if we had to crack open a skull and do neurosurgery.
Also, if by “mechanistic interpretability” you mean “circuits” I’m honestly pretty pessimistic about the usefulness of that kind of research, and I think the really-useful stuff is lower cost than circuits-based interp.
What links here?
- Arguments for optimism on AI Alignment (I don’t endorse this version, will reupload a new version soon.) by Noosphere89 (LessWrong; 15 Oct 2023 14:51 UTC; 23 points)

Nora Belrose 28 Feb 2024 2:27 UTC
4 points
0 ∶ 0
in reply to: Matthew_Barnett’s comment on: Counting arguments provide no evidence for AI doom
The goal realism section was an argument in the alternative. If you just agree with us that the indifference principle is invalid, then the counting argument fails, and it doesn’t matter what you think about goal realism.
If you think that some form of indifference reasoning still works— in a way that saves the counting argument for scheming— the most plausible view on which that’s true is goal realism combined with Huemer’s restricted indifference principle. We attack goal realism to try to close off that line of reasoning.

Nora Belrose 17 Sep 2023 22:10 UTC
4 points
2 ∶ 2
in reply to: Chris Leong’s comment on: AI Pause Will Likely Backfire
That if there was a pause, alignment research would magically revert back to what it was back in the MIRI days
The claim is more like, “the MIRI days are a cautionary tale about what may happen when alignment research isn’t embedded inside a feedback loop with capabilities.” I don’t literally believe we would revert back to pure theoretical research during a pause, but I do think the research would get considerably lower quality.
However, I’m worried that your [white box] framing is confusing and will cause people to talk past each other.
Perhaps, but I think the current conventional wisdom that neural nets are “black box” is itself a confusing and bad framing and I’m trying to displace it.

Nora Belrose 24 Mar 2023 9:12 UTC
4 points
2 ∶ 0
in reply to: xuan’s comment on: There are no coherence theorems
“good reasoning” is really intersubjective rather than objective! There’s only pressure to find the right logical beliefs in a reasonable amount of time if there are others who would fleece you for not doing so.
This is a really interesting point that reminds me of arguments made by pragmatist philosophers like John Dewey and Richard Rorty. They also wanted to make “justification” an intersubjective phenomenon, of justifying your beliefs to other people. I don’t think they had money-pump arguments in mind though.

Nora Belrose 19 Sep 2023 15:36 UTC
3 points
0 ∶ 0
in reply to: Davidmanheim’s comment on: AI Pause Will Likely Backfire
I have now made a clarification at the very top of the post to make it 1000% clear that my opposition is disjunctive, because people repeatedly get confused / misunderstand me on this point.

Nora Belrose 19 Sep 2023 15:32 UTC
3 points
2 ∶ 0
in reply to: RobertM’s comment on: AI Pause Will Likely Backfire
Please stop saying that mind-space is an “enormously broad space.” What does that even mean? How have you established a measure on mind-space that isn’t totally arbitrary?
What if concepts and values are convergent when trained on similar data, just like we see convergent evolution in biology?

Nora Belrose 18 Sep 2023 3:03 UTC
3 points
1 ∶ 3
in reply to: Steven Byrnes’s comment on: AI Pause Will Likely Backfire
Differentiability is a pretty big part of the white box argument.
The terabyte compiled executable binary is still white box in a minimal sense but it’s going to take a lot of work to mould that thing into something that does what you want. You’ll have to decompile it and do a lot of static analysis, and Rice’s theorem gets in the way of the kinds of stuff you can prove about it. The code might be adversarially obfuscated, although literal black box obfuscation is provably impossible.
If instead of a terabyte of compiled code, you give me a trillion neural net weights, I can fine tune that network to do a lot of stuff. And if I’m worried about the base model being preserved underneath and doing nefarious things, I can generate synthetic data from the fine tuned model and train a fresh network from scratch on that (although to be fair that’s pretty compute-intensive).
What links here?
- Arguments for optimism on AI Alignment (I don’t endorse this version, will reupload a new version soon.) by Noosphere89 (LessWrong; 15 Oct 2023 14:51 UTC; 23 points)

Nora Belrose

AI Pause Will Likely Backfire

Count­ing ar­gu­ments provide no ev­i­dence for AI doom

De­con­struct­ing Bostrom’s Clas­sic Ar­gu­ment for AI Doom

Counting arguments provide no evidence for AI doom

Deconstructing Bostrom’s Classic Argument for AI Doom