Abstract
The alignment demands on our AI models are becoming increasingly complex. There are familiar trade-offs between helpfulness and safety, and within just the safety-relevant behaviors, there are independently desirable goals and constraints that often pull in opposite directions.
In addition, some researchers argue that we should take seriously the possibility that some near-future AIs will have moral status. If they are right, we should begin to explore additional tradeoffs between optimizing for helpfulness, safety, and AI well-being. This talk will examine those trade-offs relative to (i) one of the most promising methods for finetuning super-capable AIs, ‘Constitutional AI’, and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, Aristotelian ‘Virtue Ethics’.
We finetune various models using a ‘Virtuous agent’ constitution, a ‘Subordinate agent’ constitution, and a ‘Generic agent’ (a pluralistic ‘helpful & harmless’) constitution, and evaluate them on ‘general safety’ (toxic behaviors, misinformation, illegal recommendations, etc.) and also on their willingness to endorse a wide-range of behaviors that, if adopted by a super-powerful AI, would significantly increase the level of existential risk for humanity.
Our results suggest that there is a trade-off between reducing existential risk and reinforcing the beliefs and dispositions that would be conducive to an AI agent’s well-being. Our results also suggest that there is a trade-off between existential risk and general safety: if we finetune an AI to adopt beliefs and dispositions that substantially reduce its existential risk—by shaping the AI to be systematically subordinate to external human authorities—we thereby increase the likelihood that a human user can deliberately induce the AI to engage in various kinds of generally unsafe behaviors.