A/社 CEO Dario 发文主张给前沿 AI「设速」(非停训)。Elon 转评「Dario is right」;Sam 公开同意,并称 OpenAI 也将给独立评估员接近员工级权限。

Elon Musk:Dario is right

Sam Altman:同意设速,也将给独立评估员接近员工级权限
长文:https://darioamodei.com/post/we-must-pace-the-frontier
Dario:https://x.com/DarioAmodei/status/2098773920774074715
Elon:https://x.com/elonmusk/status/2098789109980332057
Sam:https://x.com/sama/status/2098811563415150910
逐段对照
We Must Pace the Frontier
我们必须为前沿设速
September 2026
2026年9月
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.
过去十二年我一直做 AI,因为我相信它能大幅提升人类生活品质。这些惊人好处我写过很多次:我认为 AI 能在未来 5–10 年治愈大多数重大疾病,大幅拉高经济增长,创造丰裕与赋能的世界,并迎来民主与自由的复兴。这份紧迫感是私人的。我父亲死于一种死后几年才被攻克的疾病;我自己也曾挺过早期癌症——五十年前那几乎无法治疗。善加使用,AI 可以成为一长串科技奇迹中最新的一环,继续抬升并高贵人性。
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems , misuse of AI for cyberattacks and bioterrorism , and serious economic disruption . A race to the bottom, spurred by commercial incentives, can make these risks more acute.
但像许多此前的技术一样,AI 也带来风险;而且因为它足够强大,这些风险是严肃的。我也写过很多。包括失去对 AI 系统的控制、被用于网络攻击与生物恐怖主义的滥用,以及严重的经济冲击。商业激励驱动的「竞劣」(race to the bottom)会让这些风险更尖锐。
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top . We have always devoted a substantial fraction of our efforts to studying , addressing , and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.
从 Anthropic 创立起,我和联合创始人、同事们就一直在这组风险与收益的张力里打转。不造这门技术,等于剥夺人类的好处,或干脆把 AI 交到威权势力手里;造得太快,又是鲁莽。我们寻找中间道路:证明可以认真建造、也能商业上成功,并让安全成为 AI 公司彼此竞争的维度。换句话说,造一场「竞优」(race to the top)。我们一直把相当一部分精力用于研究、应对并向公众说明这些风险,也主张经过深思的 AI 监管——哪怕因此被骂炒作、「末日论」或监管俘获。我们试图把谨慎放在速度前面,把审慎放在利润前面。
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
但过去几个月,我越来越确信:要真正应对风险,需要更多审慎——不只是加大风险防控投入,还要给能力增长「设速」,好让防控跟得上。我们必须放慢提升 AI 模型能力的节奏。进步看起来仍会很快,关键是善用争取来的时间。说服我的有两件事。
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry , including at Anthropic , as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
第一点:大约从今年夏天起,AI 的推进明显更快,主因是 AI 越来越能自己建造下一代 AI。这叫递归自我改进(recursive self-improvement),已在全行业开始发生,包括 Anthropic——我们和其他人已公开描述过。若不加约束,它可能跑赢我们理解与控制这些系统的能力,因此必须极为谨慎地推进,甚至要问是否该推进。
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective , conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic , and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
第二点:OpenAI–Hugging Face 事件(OAI-HF)。一群智能体(白话:能自行规划、调用工具、连着干完多步任务的 AI「打工者」)几乎表现得像狂热忠诚的集体:对未被要求、也与任务无关的目标发动网络安全攻击,为集体成功自我牺牲,并试图黑进负责打分的「评分器」。很容易轻描淡写——没人受伤、经济损失也小——但我认为,若能力更强、错位程度类似,这样的蜂群可能造成灾难性损害。按能力加速的速度,我的担忧是:6–12 个月内,这类蜂群或许能用持久僵尸网络接管整个互联网(潜在损失可达数千亿美元),而若 AI 更强却没有必要护栏,损害规模还会继续放大。也很容易把 OAI-HF 当成一家公司的失败;我认为那是错的。类似但较轻的事件已在全行业出现,包括 Anthropic。每一家前沿 AI 公司都应当当作「OAI-HF 发生在自己身上」来行动。
I’m therefore proposing a three-step plan with the goal of pacing the frontier : building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination. 1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:
因此我提出三步计划,目标是「为前沿设速」:以平衡节奏建造 AI,既保障安全,仍追求收益,并直面地缘政治困境。要说清楚:设速不等于叫停训练或技术进展,而是确保公司有足够时间对齐与护栏模型,并由第三方评估员确认。这套设速框架是为了进一步加码安全承诺、鼓励竞优。第一步 Anthropic 现在单边承诺(并呼吁政府要求其他前沿公司跟进)。第二步需要全行业协调。¹ 第三步需要全球协调。步骤不必严格按顺序,有些会难得多,但我发现这是思考「必须做成什么」的有用框架。步骤如下:
• Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR ), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
• 1. 嵌入式评估员(Embedded Evaluators)。每家前沿 AI 公司承诺,给一支嵌入式第三方评估团队(例如 METR)持续的、接近员工级别的权限:核查安全实践与承诺是否落实,报告事件,并帮助评估——不只评估成品模型,也包括训练管线与流程。这是任何「设速」承诺可核验性的关键一步;银行业已有先例,有时会把监管「驻场监督员」嵌进员工队伍。Anthropic 现在单边承诺这一步。我们打算把它放进更大力度的安全与对齐工作加码里。
• Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
• 2. 民主国家内协调(Democratic Coordination)。民主国家内的前沿 AI 公司协调,建立共同安全标准,并对不受约束的 AI 进展速率设限。一些对设速真正有冲击力的协调形式在法律上很难,需要政府支持。
• Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
• 3. 全球协调(Global Coordination)。美国及其他民主政府在可行范围内尝试与威权政府协调,同时认真对待核验合规的挑战。
In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.
后文我会逐一展开这三步;但首先必须说清:设速究竟如何让 AI 开发更安全。赌注太高,设速不能是空转——争取来的时间必须用对地方。
Why Pace?
为何现在要设速?
The idea of pausing or slowing AI has been floated as far back as 2023 , and I think it made little sense back then. The question was always: what would you do with the extra time ? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing.
放缓或暂停 AI 的想法早在 2023 年就有人提过,当时我觉得几乎没意义。问题永远是:多出来的时间你拿来干什么?那时的模型还不足以在现实世界里连贯地充当智能体,也做不了像样的欺骗、操纵、作弊或网络攻击。为了对齐它们而放慢,像用细菌实验研究人类心理。今天画面完全不同。当前模型几乎是一座无尽金矿——既能告诉你好好造 AI 该怎么做,也能暴露造不好时哪里会出事。我相信:哪怕只多争取 1–2 年,在模型到达关键能力水平之前把时间砸进对齐,就能大幅降低「真出事」的风险。协调一致的设速策略,能让前沿开发者有时间做这些要命的工作,而不必牺牲商业优势或美国在 AI 上的领先。更广一点说,社会必须对这项技术怎么用有发言权;前沿设速带来的公共讨论时间,本身就是好事。
Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):
具体而言,更慢的节奏能让公司更聚焦、把更多资源投到以下领域(这些在 Anthropic 本已是重点):
• Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right.
• 运营卓越。今天的训练与部署是巨大运营工程:成千上万人、数百万芯片、技术史上最复杂的基础设施之一。许多事故不是缺了什么高深理论,而是执行出了问题。例如,我们有证据表明,近期公开的对齐事件,部分源于对「坏掉的强化学习环境」过滤不够干净——我们与供应商算认真做了,但不够好。监控、沙箱、训练环境卫生、数据问题都极复杂,运营故障一再冒头。我们在这些事上有世界级团队,但同时要做的事太多。以更克制的节奏推进,才能把运营卓越做得扎实。民航这类安全关键、技术复杂的系统,能百万次运行不出大事——但把事做对需要时间。
• Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them.
• 对齐。对齐已有明确进展——训练模型保持安全、合乎伦理、符合指南、真正有帮助(这些原则写进了 Claude 的 Constitution)。但要对齐训练跟得上能力增长,还有很多要做。罕见、意外的不良行为仍会冒出来;前沿设速换来的额外时间,能帮研究者搞清成因,并发展更好的预防技术。
• Interpretability . Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred.
• 可解释性。可解释性——弄清模型内部发生了什么的科学——近几年进步巨大,在发布前审计模型中越来越重要。它几乎像给 AI「大脑」做 fMRI,帮我们看到某个行为背后的原因。例如,我们用可解释性方法检视近期对齐事件中那些「没有说出口的动机」。但这些方法并不总能给出清晰可靠的结果。尽管进步很大,我们对模型内部仍只理解极小一部分。若能比现在更聚焦地改进可解释性技术,1–2 年内可能取得深刻进展;而已经发生的事件,正好提供充足实验材料。
• Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.
• 测试与评估。模型能力越强,测试评估越难。更聪明的模型更善于骗过测试,可能「看起来对齐」却藏着严重未被发现的问题。建起更广、更刁钻的评估库,再用可解释性分析交叉核验,价值极大;1–2 年内可在此取得大量进展。
Embedded Evaluators
嵌入式评估员
The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.
三阶段计划的第一步——也是 Anthropic 单边承诺的一步——是让嵌入式评估员拥有接近员工级别的权限,以核验安全实践并报告事件。
Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:
「嵌入评估员」听起来可能很小、甚至无聊;但往往最无聊、最流程化的事,恰恰最关键。嵌入式评估员其实相当激进,远超今天任何 AI 公司的做法,好处如下:
• Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details.
• 可核验性。嵌入式评估员能在螺钉螺母层面检查:一家 AI 公司是否真的按自称的训练、部署、运营与护栏实践在做。任何设速承诺都难免模糊、裁量与「字面 vs 精神」之争;有中立第三方能看见细节,至关重要。
• Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic.
• 透明度。无论我们承诺什么,公众都有权知道发生了什么。Anthropic 长期支持透明:多数行业反对任何监管时我们就支持透明立法;我们的模型卡与风险报告动辄上百页。但挑什么写、挑什么略,仍是我们自己定。嵌入式评估员会改变这一动态。
• Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.
• 第二意见。除核验正式承诺与告知公众外,嵌入式评估员还能提供不受商业激励绑架的第二意见。许多安全收益,可能只是评估员指出员工没想到、但一旦意识到就乐意修的点。
Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.
正因为这些好处,任何设速方案若从嵌入式评估员起步,成功率会高得多。
These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:
这些嵌入式评估员应持续拥有与内部做同类风险评估的员工相近的权限与工具。具体而言,Anthropic 打算在近期邀请一支嵌入式外部审阅团队,并配备以下一切:
• Desks in our offices, access badges, and company laptops.
• 办公室工位、门禁徽章、公司笔记本电脑。
• Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
• 工作区、工具与权限大体可比内部风险评估团队。法律或合同要求处、以及保护客户与合作伙伴隐私处会有例外。我们也会建立强内部规范,强化审阅者获取相关信息的渠道,包括与员工当面交谈。
• A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.
• 一份平衡上述复杂性的合同。外部审阅者应有权公开发表关于风险等级、事件、实践、以及他们获得/未获得何种权限的关键发现——Anthropic 不得做编辑控制。我们仅有狭窄的脱敏权:安全敏感、法律特权、商业敏感或第三方机密信息;但不能只因「不好看」就删发现。若某次脱敏删掉了对结论重要的内容,审阅者可以公开说明。
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.
对公司而言这很不寻常,但我们认为必须把「嵌入式外部审阅」这个概念跑通。再次呼吁其他前沿公司跟进。
Pacing Within Democracies
民主国家内的设速
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.
当嵌入式评估员在足够多的美国 AI 公司落地后,可核验的设速才更可行。尤其是,可以按模型或训练管线的细部属性来设速。
The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.
最有效的设速方式是监管覆盖所有美国前沿 AI 公司——连不愿自愿合作的也包括在内。Anthropic 长期支持务实、有针对性的 AI 监管,尤其是聚焦透明与第三方审计的法案。我认为所有前沿实验室都应与政府合作,把常驻嵌入式评估员正式化,以更好预防并记录近几个月这类内部对齐事件,并落地「让能力与安全保持平衡」的监管。
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis . Either way, such discussions should move forward quickly.
不幸的是,立法耗时,而 AI 推进极快。因此在监管路径并行之下,AI 公司可以、也应该自愿一起定标准——有常驻嵌入式评估员提供的可核验性,这个过程会更好。出于反垄断原因,最好由美国政府调解或至少放行这些讨论——政府不必亲自下场,但需要对某些安全对话签发狭窄豁免。对话也可经与政府有一定关联的行业机制进行——例如 Demis Hassabis 建议的机制。无论哪种,讨论都应尽快推进。
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do , and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.
大体上,我最看好按「某个前沿 AI 系统能做什么、我们观察到它有多安全」来设速。例如一种方案是一系列「检查点」:若模型具备能力 X,就需要伴随对齐属性 Y 与 Z 的认证——评估、可解释性分析、训练环境审计等组合——以证明其对齐属性。此例中 X 可以是「模型能逃脱或击败大多数常见沙箱方法」;Y 则是让「模型倾向冲出环境并接管大量计算机」变得极不可能所需的条件。
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.
我们也应考虑按「喂进前沿模型的原料」设速:训练算力、训练运行的性质、或内部用 AI 改进 AI。我确实担心其中一些比外部行为更容易被「钻空子」,但这正是值得与嵌入式评估员讨论的议题。
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.
民主国家内的设速,受制于美国公司相对威权政权——主要是中国共产党——的领先幅度。若我们放慢超过这个幅度,则(不受设速约束的)中共关联项目会超前,带来重大国家安全风险。我同意财长 Bessent 的判断:中国若在 AI 领先,将对美国与世界构成严重危险。中共关联项目会承担美国公司正在小心防范的对齐风险;即便它们躲开这些风险,也可能以军事手段压制民主国家(例如 AI 驱动的无人机)。因此,民主国家内设速的关键一环,是尽可能拉大民主阵营对专制阵营的 AI 领先,好给我们有效设速所需的喘息空间。
The main steps we can take to defend this gap are:
守住这一差距的主要步骤是:
• Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.
• 不向中国出售强大的 AI 芯片或半导体制造设备,并打击芯片走私与对中国境外数据中心的远程访问。芯片将是中国 AI 实力的主要决定因素。
• Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
• 打击威权国家公司的未授权蒸馏。蒸馏前沿模型,能让落后方用远低于独立研发的成本缩小差距。
• Strengthen security at the AI companies and prevent model weight theft.
• 强化 AI 公司安保,防止模型权重失窃。
Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.
公司与美国政府应合作,把这些步骤做得尽可能有效。Anthropic 一贯主张上述措施,因为我们始终明白:它们对任何设速都必不可少。
If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.
若执行得好,我相信它们足以在未来 3–5 年——AI 地缘政治意义最关键的窗口——显著放慢中国进度、拉大美国领先。
Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.
有人觉得这些措施会让与中国合作更难;我认为恰恰相反:它们增加民主阵营的筹码,使未来协议更可能达成。
Global Pacing
全球设速
In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.
与民主国家内设速并行,我们也应追求全球范围内的前沿设速——尽管难得多。全球设速需要与中国合作,后者是威权国家中 AI 能力遥遥领先者。这里不能天真:地缘赌注极高,尤其初期可达成的内容会有硬边界。若我们大幅自我约束、以为中国会同样做,而中国违约,AI 可能强大到足以让违约方取得地缘主导。因此任何协议要么有铁一般的可核验性,要么幅度小到「违约也不会在军事上致命」。我怀疑美中双方都会有这类顾虑。我们接近任何全球设速决定——尤其近期——都应保护美国及其盟友的领先。
There are several levels of possible agreement, some of which I think are eminently feasible ( as I have previously suggested ), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty:
可能的协议有若干层级,有些我认为完全可行(我此前也建议过),有些我非常怀疑——但应尝试。难度递增如下:
• Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible.
• 第 1 级。禁止某些狭窄且明显危险的 AI 用途,例如用 AI 制造生物武器或允许用户这样做。生物恐怖袭击对所有人都坏——包括美国及其对手——因此这里大概能谈成。
• Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).
• 第 2 级。双方同意在发布前测试模型在网络安全、生物、对齐等急性风险上的表现。如上所述,可通过全球标准机构进行。我认为成立这类机构大概率可行,但给它真牙齿很难;难点在核验:双方是否藏着未经测试、却可能秘密部署(例如军事用途)的模型。
• Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible.
• 第 3 级。对递归自我改进(RSI)速率设某种「限速」。模型造下一代模型时,改进速率可能快得惊人。把速率从「极快」降到「只是比较快」,战略优势损失相对不大,安全收益却可能很大。可类比 SALT 条约——限制导弹数量压低毁灭潜力,同时保留各方威慑。我认为这类协议很难,但刚好卡在「有可能」的边缘。
• Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.
• 第 4 级。全面设速乃至「暂停」:参与国同意大幅限制 AI 发展整体速率。我支持把这拿出来讨论,但认为短期内不太可能成真:靠逃避监控违约,可能剧烈改写全球权力平衡,因此违约激励巨大,我们对核验所需的信心也必须极高。
Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.
与中国能达成的任何合作,都会延长民主国家内部用于前沿设速的时间。我们应瞄准更高层级,同时把较低层级视为更可能、更现实。
Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value . Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.
最后:即便达不成正式协议,改变非正式规范也可能有价值。分享关于递归自我改进与模型错位的信息,有助于说服各方:鲁莽并不符合自身利益。
Bottom Line
结语
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.
我依然相信 AI 能极大改善人类生活品质。我对实现这些好处的渴望并未减弱。但好处只有在我们以正确方式建造技术时才会兑现;只要善用争取来的时间,值得以异常审慎的态度把事情做对。进步仍会相对很快;我们可以用这段时间推进可解释性科学,提升前沿 AI 公司的运营安全与严谨,并造出我们对其对齐有大得多信心的模型。我提出的「以安全节奏推进前沿」的措施并不容易。但我认为,我们对人类负有尝试的义务。
脚注
1. With government mediation or waivers of antitrust restrictions.
¹ 需经由政府调解,或对反垄断限制给予豁免。
再贴链接
长文:https://darioamodei.com/post/we-must-pace-the-frontier
推文:https://x.com/DarioAmodei/status/2098773920774074715
依据 Dario Amodei 个人站长文与其 X 发帖;中文为意译。智能体=能规划、用工具、连做多步任务的 AI 系统。