Toward an Empirical Science of AI Character

By: Courtney Bigony

Background: Master’s in applied positive psychology from Penn; a decade translating wellbeing science into technology. I believe an empirical science of wellbeing can guide AI character to promote both safety and human flourishing at scale. Uncertain whether AI companies are doing this, or prioritizing it, today.

1. AI companies already commit to human thriving. OpenAI’s charter aims for AGI that “benefits all of humanity.” Anthropic wants AI that is “safe and beneficial for our users and for society as a whole.” And Claude’s constitution goes further: its wellbeing clause states that Claude “should give weight to the long-term flourishing of the user and not just their immediate interests.” These commitments are real. But to my knowledge, no public work connects a comprehensive set of validated flourishing constructs to where character is actually formed or tracks whether deployed models support users’ flourishing over time.

2. Wellbeing belongs at the center of AI character. Anthropic gets this right: safety is “a systematic science.” But human needs come in two buckets — security and growth — and that science currently covers one. A model can be perfectly safe without promoting human flourishing. Constitutions shape models not with rigid rules, but by raising them to internalize values so behavior flows from character. If character is what gets trained, then one question matters: what kind of character supports human flourishing? That isn’t a matter of opinion. Psychology has empirical answers. My guess, and it’s only a guess, is that wellbeing already shapes Claude’s character as philosophy. The opportunity is to shape it with empirical science and measurably promote human thriving.

3. The science already exists. Decades of validated instruments and randomized controlled trials cover what humans need to thrive: autonomy, psychological safety, purpose, hope, flow. What’s missing is an empirical science of wellbeing that guides AI character. One attempt at its measurement layer is the Human Potential Index, which I co-created with Jeff Smith, PhD, and self-actualization scientist Scott Barry Kaufman, PhD. It pairs 33 thriving constructs, each with its validated self-report items, with an observable AI-behavior signal, usable upstream for grading character-training data and downstream for longitudinal evals. Full framework here. Whether an AI trained this way actually improves its users’ measured flourishing has, to my knowledge, never been tested for this framework or any other.

4. Can an empirical science of AI character actually be built? (a) AI companies rightly focus on safety. Is there space to focus on flourishing as well: in character training, in evals, in priorities? (b) Do wellbeing values belong at the user-instruction level, the character level, or both? (c) For those closer to this work: to what extent are frontier labs already focused on wellbeing in a scientific way?

No comments.