先以为是谣言,核实后确认 Jacob Coxon 其人其事;原帖约 5000 万次观看。这位 GPT-4o Core Contributor 从 Anthropic 辞职,批评两家公司冲向自我改进的超级智能,并主张限速。

一开始我以为这是谣言,或者至少是没凭据的吓人话。后来核对了一下:Jacob Coxon 这个人是真的,OpenAI、Anthropic 预训练经历也能对上;他自己在 X 上的辞职串也在。截稿时原帖已经冲到大约 5000 万次观看。下面这篇,就从这件事说起。

今天,Jacob Coxon 宣布从 Anthropic 辞职。

他不是普通员工。

过去三年,他先后在 OpenAI 和 Anthropic 从事大模型预训练研究,也是 OpenAI GPT-4o 官方列出的 Core Contributor 之一。

他的辞职理由非常直接:

OpenAI 和 Anthropic 都没有负责任地行动。两家公司正在冲向“能够自我改进的超级智能”,并拿所有人的生命赌博。

“做 AI 的人,真的认为它可能杀死所有人”

Coxon 最引发争议的一句话是:

“那些正在构建 AI 的人,是真的相信 AI 可能在本十年结束前杀死我们所有人。”

他强调,这不是营销话术。

一些高管和高级研究人员面对媒体时会把话说得更温和,但私下里,他听到的是同样的担忧。

当然,这并不是说今天的 ChatGPT 或 Claude 已经失控。

他担心的是下一阶段:

模型开始越来越多地参与代码、科研和 AI 本身的研发,最终出现能够持续改进自己的超级智能。

为什么明知道危险,还要继续做?

Coxon 对 OpenAI 和 Anthropic 给出了不同的判断。

他说,在 OpenAI,很多人并没有真正把这个问题当成一个“文明级风险”。

而 Anthropic 对风险理解得更深,但陷入了另一种逻辑:

如果我们不做,别人也会做。

既然超级智能可能迟早出现,那最好由我们先做出来。

于是所有公司都认为自己不能先减速。

这也是 Coxon 最反对的地方。

“这不应该由一家私人公司的 Slack 决定”

他写了一句很重的话:

Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack.

大意是:

进入这场超级智能竞赛,是一场傲慢的豪赌。这种事情,不应该由一家私人公司的 Slack 频道决定。

问题已经不只是 AI Safety。

而是:

谁有权决定什么时候启动超级智能?

几家公司?

几十个研究人员?

几个 CEO?

还是整个社会?

他主张减速

Coxon 最后提出,美国几家主要 AI 公司应该考虑达成某种限速协议。

如果有必要,甚至应该考虑:

暂时禁止继续提升最前沿模型的能力。

这个观点当然非常激进。

但值得注意的是,说这些话的人,正是过去几年亲自参与最前沿模型训练的人之一。

真正值得关注的问题也许不是:

“AI 到底会不会毁灭人类?”

而是:

如果造 AI 的人自己都认为存在这种风险,为什么所有人仍然不敢减速?

Jacob Coxon 辞职声明英文原文

Jacob Coxon (@hilbertspaess)
2026-09-09

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.

A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.

Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.

I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.

If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?

原文链接:https://x.com/hilbertspaess/status/2097476196791709843