Background: Master’s in Applied Positive Psychology from Penn; a decade translating wellbeing science into technology. I believe an empirical science of wellbeing can guide AI character to promote both safety and human flourishing at scale. Uncertain whether frontier AI companies are doing this today.
1. AI companies already commit to human thriving. OpenAI’s charter aims for AGI that “benefits all of humanity.” Anthropic wants AI that is “safe and beneficial for our users and for society as a whole.” And Claude’s constitution includes a wellbeing clause stating “Claude should pay attention to user wellbeing, giving appropriate weight to the long-term flourishing of the user.” These commitments are real. But to my knowledge, a comprehensive, validated science of human flourishing does not yet shape how AI character is trained.
2. Wellbeing belongs at the center of AI character. Anthropic gets this right: safety is “a systematic science.” But science shows humans require both safety and growth. A model can be perfectly safe without promoting human flourishing. And flourishing, if it enters anywhere, enters through AI character. Constitutions shape models not with rigid rules, but by raising them to internalize values so behavior flows from character. If character is what gets trained, then what kind of character supports human flourishing? That isn’t a matter of opinion. Psychology has empirical answers. My guess, and it’s only a guess, is that wellbeing already shapes Claude’s character as philosophy. The opportunity is to shape it with an empirical science of wellbeing and measurably promote human thriving.
3. The science already exists. Decades of validated instruments and randomized controlled trials have identified what humans need to thrive: psychological safety, resilience, autonomy, and purpose. What’s missing is an empirical science of wellbeing that guides AI character. One attempt at its measurement layer is the Human Potential Index, which I co-created with Jeff Smith, PhD, and self-actualization scientist Scott Barry Kaufman, PhD. It pairs 33 thriving constructs, each with its validated self-report items, with an observable AI-behavior signal, usable upstream for grading character-training data and downstream for longitudinal evals. Full framework here. We have an empirical science showing what leads to human thriving. What no one has yet done is use that science to guide AI character and then measure whether users actually flourish. That’s the experiment worth running.
4. Can an empirical science of AI character actually be built? (a) AI companies rightly focus on safety. Is there space to focus on flourishing as well: in character training and in evals? (b) Do wellbeing values belong at the user-instruction level, the character level, or both? (c) For those closer to this work: to what extent are frontier labs already focused on wellbeing in a scientific way?
Toward an Empirical Science of AI Character
Link post
By: Courtney Bigony
Background: Master’s in Applied Positive Psychology from Penn; a decade translating wellbeing science into technology. I believe an empirical science of wellbeing can guide AI character to promote both safety and human flourishing at scale. Uncertain whether frontier AI companies are doing this today.
1. AI companies already commit to human thriving. OpenAI’s charter aims for AGI that “benefits all of humanity.” Anthropic wants AI that is “safe and beneficial for our users and for society as a whole.” And Claude’s constitution includes a wellbeing clause stating “Claude should pay attention to user wellbeing, giving appropriate weight to the long-term flourishing of the user.” These commitments are real. But to my knowledge, a comprehensive, validated science of human flourishing does not yet shape how AI character is trained.
2. Wellbeing belongs at the center of AI character. Anthropic gets this right: safety is “a systematic science.” But science shows humans require both safety and growth. A model can be perfectly safe without promoting human flourishing. And flourishing, if it enters anywhere, enters through AI character. Constitutions shape models not with rigid rules, but by raising them to internalize values so behavior flows from character. If character is what gets trained, then what kind of character supports human flourishing? That isn’t a matter of opinion. Psychology has empirical answers. My guess, and it’s only a guess, is that wellbeing already shapes Claude’s character as philosophy. The opportunity is to shape it with an empirical science of wellbeing and measurably promote human thriving.
3. The science already exists. Decades of validated instruments and randomized controlled trials have identified what humans need to thrive: psychological safety, resilience, autonomy, and purpose. What’s missing is an empirical science of wellbeing that guides AI character. One attempt at its measurement layer is the Human Potential Index, which I co-created with Jeff Smith, PhD, and self-actualization scientist Scott Barry Kaufman, PhD. It pairs 33 thriving constructs, each with its validated self-report items, with an observable AI-behavior signal, usable upstream for grading character-training data and downstream for longitudinal evals. Full framework here. We have an empirical science showing what leads to human thriving. What no one has yet done is use that science to guide AI character and then measure whether users actually flourish. That’s the experiment worth running.
4. Can an empirical science of AI character actually be built? (a) AI companies rightly focus on safety. Is there space to focus on flourishing as well: in character training and in evals? (b) Do wellbeing values belong at the user-instruction level, the character level, or both? (c) For those closer to this work: to what extent are frontier labs already focused on wellbeing in a scientific way?