Character Training for Risk-Averse Agents
Quick summary
arXiv:2609.38093v1 Announce Type: new Abstract: Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation
Key takeaways
- arXiv:2609.38093v1 Announce Type: new Abstract: Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm.
- Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling.
- We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences.
Why it matters
“Character Training for Risk-Averse Agents” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Member comments