I think you could know how to encode preferences into ASI without knowing whether it’s aligned. At that point it’s like a genie, you may well be able encode preferences into it but you might encode bad preferences that you didn’t realize would be bad.
As I understand the term, that sort of ASI wouldn’t be considered “misaligned”, it would be “aligned, but to the wrong target”. I think of misalignment as when you wanted the ASI to do one thing, but it did something else instead.
I think you could know how to encode preferences into ASI without knowing whether it’s aligned. At that point it’s like a genie, you may well be able encode preferences into it but you might encode bad preferences that you didn’t realize would be bad.
As I understand the term, that sort of ASI wouldn’t be considered “misaligned”, it would be “aligned, but to the wrong target”. I think of misalignment as when you wanted the ASI to do one thing, but it did something else instead.