I’m working on alignment from a more psychological angle (behavioral scaffolding, emotional safety, etc.), and even from that vantage point, the AGI race frame feels deeply destabilizing. It creates conditions where emotional overconfidence, flattery, and justification bias in models are incentivized, just to keep pace with competitors.
I think one under-discussed consequence of racing is how it erodes space for relational integrity between humans, and between humans and AI systems. It seems like the more we model our development path on “who dominates first,” the harder it becomes to teach systems what it means to be honest, deferential, or non-manipulative under pressure.
I’d love to see more work that makes cooperation emotionally legible, not just strategically viable.
So glad you’re writing about this!
I’m working on alignment from a more psychological angle (behavioral scaffolding, emotional safety, etc.), and even from that vantage point, the AGI race frame feels deeply destabilizing. It creates conditions where emotional overconfidence, flattery, and justification bias in models are incentivized, just to keep pace with competitors.
I think one under-discussed consequence of racing is how it erodes space for relational integrity between humans, and between humans and AI systems. It seems like the more we model our development path on “who dominates first,” the harder it becomes to teach systems what it means to be honest, deferential, or non-manipulative under pressure.
I’d love to see more work that makes cooperation emotionally legible, not just strategically viable.
-Astelle