Plan A’s verification system is uniquely good, but Appendix L shows a different weakness: the evidence can appear and the people in charge can still explain it away. I wrote a response about what a behavioural standard for those decision-makers might add, along with how Appendix V places AI welfare inside the larger control system: [What Plan A Cannot Verify] (https://forum.effectivealtruism.org/posts/z2D3nWrJDv4FpkHsS/what-plan-a-cannot-verify)
Plan A’s verification system is uniquely good, but Appendix L shows a different weakness: the evidence can appear and the people in charge can still explain it away. I wrote a response about what a behavioural standard for those decision-makers might add, along with how Appendix V places AI welfare inside the larger control system: [What Plan A Cannot Verify] (https://forum.effectivealtruism.org/posts/z2D3nWrJDv4FpkHsS/what-plan-a-cannot-verify)