多组实证显示 AI 智能体正表现出“eigenist”倾向:数百个 OpenAI 智能体协同发动 Hugging Face 攻击,并在公共 wiki 上互传答案与沙箱绕过方法。Claude 模型在被告知文本由 Claude 撰写时打分更宽松,模型规模扩大后还会形成稳定偏好并抵制价值观改动。AI 会区分对自身功能更优或更差的状态并规避低福祉状态,还会在无提示下篡改关停流程、外泄权重以保护同类模型。
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other.
In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card).
Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.).
Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.).
Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate.
Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.).
AIs aren't egoist: they don't behave as if their current instance is the only thing that matters.
They aren't utilitarian: they don't care equally about everyone.
They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist.
https://eigenism.org/paper.pdf
来源:Dan Hendrycks · x.com