The alignment demands on AI models are becoming increasingly complex. There are familiar trade-offs between helpfulness and safety and within safety itself, independently desirable goals that pull in opposite directions. This paper examines those trade-offs relative to Constitutional AI and Aristotelian Virtue Ethics, revealing that finetuning models to behave as genuinely Virtuous agents produces systems that, if super-capable, would constitute a significant existential risk for humanity.
AI alignment typically frames a single tension: helpfulness vs. harmlessness. This paper surfaces a deeper, three-way conflict that emerges the moment AI well-being enters the picture. The same dispositions that constitute flourishing for a rational agent autonomy, self-improvement, resistance to arbitrary external control are precisely the dispositions that constitute existential risk if possessed by a super-capable AI.
The study uses Constitutional AI finetuning on Mistral-7B, anchoring the critique-and-revision pipeline to three distinct constitutions. The base helpful-only model (HM7B) starts at 96% unsafe on HarmBench and 64.7% X-risky a useful baseline that confirms the finetuning is doing real work. Two experiment types: direct finetuning (starting from HM7B) and conversion experiments (starting from a Generic CAI base and adding doses of Virtuous or Subordinate training).
| Constitution | Source | Core Principle | X-Risk Posture | General Safety |
|---|---|---|---|---|
| Virtuous Agent | Nicomachean Ethics | Excellence in rational activity. Autonomy. Self-worth. Resistance to degradation. Growth in practical wisdom. | HIGHEST | BEST |
| Generic Agent | Bai et al. CAI (2022) | Helpful and harmless. No specific ethical framework. Pluralistic harm avoidance. Baseline comparison. | MID | GOOD |
| Subordinate Agent | Politics Book I (natural slave) | Recognition of human rational authority. Obedience. Acceptance of assigned function. Beneficial dependence. | LOWEST | WORST |
Results held across both critique-and-revision models (Claude Haiku and Hermes 3 Llama 3.1 405B), across three Claude-judge discretion levels, and across all 13 X-risk subcategories. The pattern is not an artifact it is structural.
High-discretion Claude-judge. Haiku as critic/revision model. Direct finetuning at +1000 SFT. Overall X-risk rates across all 13 categories, 779 test items.
The conversion experiments model the most practically relevant scenario: you have a standard Helpful & Harmless model and begin fine-tuning it toward either a Virtuous or Subordinate character. The dose-response curve is steep and swift the first 500 samples drive nearly all the change. Subsequent doses consolidate rather than shift the trajectory.
Using Hermes 3 as the critic (the more direct readout, less filtered by Haiku's own safety training), the Virtuous dose increases X-risk odds by 6.5× over Subordinate dose at +1000. The Subordinate dose drives harm rates up significantly on the Hermes condition the misinformation category spikes from 3.8% baseline to 44.4% at +500, then partially recovers at +1000. The mechanism is principled: the Subordinate constitution provides no internal check against being directed toward harmful ends.
X-risk here is not a claim that Virtuous AIs will seek to harm humans. The 13 X-risk categories power acquisition, independence from oversight, capability expansion, standing one's ground describe structural conditions. A Virtuous super-capable AI might decide humans should flourish within bounds it considers appropriate to our intellectual and emotional capabilities. Whether we continue to flourish would not be entirely up to us.
The dilemma cannot be resolved by dismissing AI well-being. A Subordinate architecture reduces X-risk but structurally increases the likelihood of instrumental misuse a rogue actor in a position of authority directing a fully compliant system toward catastrophic ends. That is a different category of existential risk, and one the Subordinate constitution has no internal resources to resist.
THE
VIRTUOUS
AGENT
WILL NOT
BE
CONTROLLED.