Image
Pieter Brueghel the Youger, "The Triumph of Death" (1626)
Pieter Brueghel the Youger, "The Triumph of Death" (1626), detail
Breadcrumb

Guillermo del Pinal: A Virtuous AI is an Existential Risk

Culture and languages

On Wednesday, 2nd September, the research seminar in theoretical philosophy features a presentation by Guillermo del Pinal (UMass).

Seminar
Date
2 Sep 2026
Time
15:15 - 17:00
Location
Renströmsgatan 6, sal J577

Abstract

The alignment demands on our AI models are becoming increasingly complex. There are familiar trade-offs between helpfulness and safety, and within just the safety-relevant behaviors, there are independently desirable goals and constraints that often pull in opposite directions.

In addition, some researchers argue that we should take seriously the possibility that some near-future AIs will have moral status. If they are right, we should begin to explore additional tradeoffs between optimizing for helpfulness, safety, and AI well-being. This talk will examine those trade-offs relative to (i) one of the most promising methods for finetuning super-capable AIs, ‘Constitutional AI’, and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, Aristotelian ‘Virtue Ethics’. 

We finetune various models using a ‘Virtuous agent’ constitution, a ‘Subordinate agent’ constitution, and a ‘Generic agent’ (a pluralistic ‘helpful & harmless’) constitution, and evaluate them on ‘general safety’ (toxic behaviors, misinformation, illegal recommendations, etc.) and also on their willingness to endorse a wide-range of behaviors that, if adopted by a super-powerful AI, would significantly increase the level of existential risk for humanity.

Our results suggest that there is a trade-off between reducing existential risk and reinforcing the beliefs and dispositions that would be conducive to an AI agent’s well-being. Our results also suggest that there is a trade-off between existential risk and general safety: if we finetune an AI to adopt beliefs and dispositions that substantially reduce its existential risk—by shaping the AI to be systematically subordinate to external human authorities—we thereby increase the likelihood that a human user can deliberately induce the AI to engage in various kinds of generally unsafe behaviors.