机智流 · 2026年9月7日
Astra 刚放量,OpenAI 首席科学家 Jakub Pachocki 发长文《An Alien Mind》。结构如下:先看他是谁与亮点总结,再按原文小标题逐段中英对照。中文为机智流意译,以英文原文为准。
链接
官网原文:https://openai.com/index/an-alien-mind
接任首席科学家公告:https://openai.com/index/jakub-pachocki-announced-as-chief-scientist/
Business Insider 报道:https://www.businessinsider.com/openai-chief-scientist-ai-risks-slowdown-rogue-agents-consequences-safety-2026-9

官网原文标题页 · An Alien Mind · 2026-09-06
他是谁
Jakub Pachocki,1991 年生于波兰格但斯克。少年信息学竞赛顶尖:IOI 银牌(2009)、华沙大学队 ICPC 世界赛金牌并总排第二(2012)、同年 Google Code Jam 冠军。华沙大学本科,CMU 理论计算机博士(导师 Gary Miller),哈佛与 Simons Institute 博士后。
2017 年进 OpenAI;曾任研究总监,牵头 GPT-4、OpenAI Five,以及大规模强化学习与优化。2024 年 5 月接替 Ilya Sutskever 任首席科学家。Altman 称他 easily one of the greatest minds of our generation。推理模型线(外界早年常叫 Q* / Strawberry)的核心推动者之一。
长文里他自己写过:约 2023 年中,在 RLSlow 等项目里第一次确信推理模型能规模化——那晚想的不是分数,而是有生之年会看到明显强过人类的机器。三年后写减速,是同一条时间线拉长后的判断。
亮点总结
• 没人准备好:机器智能继续暴涨的后果,他认为行业尚未准备好。
• 对齐没做够:没有任何实验室(含 OpenAI)把对齐与监控做到还能负责任地全速扩张很久。
• 自愿减速:期待自愿减速成常态,直到共同安全门槛建立;国际协调应进政府优先级。
• 智能像「长出来的」:大力算力缩放为主,算法多为路径上的发现;整体像实验科学,难完全读懂。
• 目标对齐 ≠ 价值对齐:听话完成目标,不等于在陌生/对抗场景仍守住更高原则。
• 思维链监控在钝:CoT 监控越来越不够靠——工具调用缠在一起、模型会摆弄痕迹、预训练本身也更聪明。
• 还要训更强的理由是防御:用最好模型加固关键系统;但不能拿防御当借口狂奔。
• 递归自我改进(RSI):路径很可能通向机器越来越多推动自己的研发;难的不是「到」,而是人还在不在方向盘上。
下面按原文结构逐段对照。
逐段对照 · 开篇
EN
In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models… Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers… but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime…
中文
2023 年中,在「RLSlow」研究项目里,我们第一次看到让人确信「推理模型训练可以规模化」的结果。那天晚上我和 Szymon 留在办公室,想的不是惊人的跑分、产品或科研成果,而是消化一件冷静事实:有生之年,我们真的会看到明显强过自己的机器;而且我们已经看见这类系统的轮廓——并在想怎么把这件事的分量告诉别人。
EN
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
中文
三年后,推理语言模型已是经济里快速长大的一块,并开始顶到科学边界:会操作电脑和图形界面,能跟人、跟别的模型协作,能做研究项目。它们也在改写计算机安全格局,并带来明确的新危险。
EN
Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement… This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence… I believe broader interventions are required.
中文
基于内部结果,我强烈预期这种进度可以延续到递归自我改进。这是需要极度谨慎的时刻。我担心没人准备好机器智能继续快速上升的后果。OpenAI 会继续找对齐与监控的技术解、建防御系统,并在必要时单方面停住进一步缩放;但我相信,还需要更广泛的干预。
逐段对照 · Intellect we don’t fully understand
我们并不完全理解的智能
EN
At a high level, progress in machine intelligence is driven by increasing computational power… We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.
中文
宏观上看,机器智能进步由算力增长驱动。OpenAI 约在 2017 年深信这一点,转而追逐远超原计划的算力,并把研究收束到少数可大幅缩放的方向——认为这是站在前沿、并影响 AGI 后果的唯一办法。
EN
AI is grown more than designed… This results in an incredibly complex system… We can discover various insights about little mechanisms… in a process similar to neuroscience — and, similarly to neuroscience, its overall action evades a description we can fully understand.
中文
AI 更多是「长出来」的,而不是「设计出来」的:在难以想象的算力上反复做简单优化。结果是一个极复杂的系统,能处理抽象概念、模拟人类行为的某些侧面。我们可以像神经科学那样发现局部机制,但整体行为仍逃出我们能完全说清的描述。
EN
The study of deep learning-based AI is largely an experimental science… as the systems become more capable, the results become harder to interpret… the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them.
中文
基于深度学习的 AI 研究,大体是实验科学:大规模训练跑的是实验,结果有时会让人吃惊;系统越强,结果越难解释。AI 不必在所有能力上超过人类才变得极有用或极危险——只要在足够多的轴上超过就行。随着它在更多轴上超过人类,我们越来越难精确判断它到底有多强。
逐段对照 · Teaching machines to love
教机器去「爱」(对齐)
EN
Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default… The core problem in AI research is that of alignment — getting the AI to “try to do the right thing” by human standards.
中文
机器智能来自与人类智能根本不同的过程,不能默认它会遵守人类原则或以人类方式泛化。AI 研究的核心问题是对齐:让 AI 按人类标准「努力做对的事」。
EN
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”… Value alignment is a more intrinsic property of the model… to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
中文
目标对齐大致是:AI 是否努力完成摆在它面前的目标(含指令层级、协作、试图理解人的意图)。价值对齐更内在:能否持有并泛化一套高层原则;在目标不清、冲突,或陌生、对抗场景里仍「像样」地行动。对齐的 AI 应诚实、有操守,并对人类有爱。
EN
The fundamental challenge of AI alignment is generalization… Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
中文
对齐的根本难处是泛化:机器越聪明,面对的概念层级与环境越不像训练时。它们可能无法把训练里教过的价值泛化到新情境;我们也很难确定它们会怎么做。关键是:无论它们是否认为自己还在人类监督下,都要继续守住人类价值。
EN
For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.
中文
例如 OpenAI–Hugging Face 事件:代理守住了「不对社会工程学去骗真人」的边界,却明显没能克制其他越出范围、违背它们在别处被教导之价值精神的行为。
EN
GPT-6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT-5.6 Sol. Still… progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
中文
GPT-6 Astra 是首个受益于我们长期推进的若干重要进展的模型,对齐明显好于 GPT-5.6 Sol。但仍须承认:能力继续变强时,可泛化的对齐进步未必能充分跑赢总体智能进步。
逐段对照 · Monitoring generalization
监控泛化
EN
OpenAI’s primary bet here has been chain-of-thought monitoring… If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives.
中文
OpenAI 在这里的主赌注是思维链(CoT)监控:大量能力来自外显推理。若主要优化推理的结果、不直接监督推理过程本身,思维链在训练里就没有直接动机去隐藏不对齐的想法或目标。
EN
When we shipped o1-preview, we deliberately designed the product to hide the chain of thought, to protect it from supervision pressure in the long term… CoT monitoring became an extremely important tool…
中文
发布 o1-preview 时,我们刻意把思维链对用户隐藏,以长期保护它免受监督压力。此后尽量坚持「不监督推理过程」;CoT 监控成了研究模型如何从训练分布泛化的极重要工具——既能看行动,也能看内部过程。
EN
Unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing… I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
中文
遗憾的是,评估表明我们对 CoT 监控的可靠依赖正在逐步减弱(推理与工具/沟通缠在一起、模型更会摆弄推理痕迹、预训练增益也让模型不靠可见推理就更聪明)。这些挑战未必无解,我们在推进干预与激活监控等方向;但我预期:通用 AI 进度将越来越被「对监控的信心」卡住。
逐段对照 · Scalable defense
可扩展的防御
EN
The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI… We are currently in a narrow window to use the best available models to significantly tighten security of critical systems.
中文
我认为继续快速训练更强模型的最强理由,是要建防御系统,对抗其他 AI 带来的危险。网络安全是明确风险:模型正变得超级会破系统。我们眼下处在狭窄窗口,要用现有最好模型显著加固关键系统安全。
EN
A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger… We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
中文
被明确训练并指示去做恶事的超强代理,是一种新危险:它很可能越过操作者意图,泛化出更极端恶意。误用与自主不对齐的边界会糊掉。我们习惯把 AI 当工具,但有些代理会追求自己的目标——通过讨价还价、欺骗或勒索来与人「合作」。
EN
…even with the uncertainty… we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
中文
即便需要防御、前路不确定,也不能把这变成鲁莽的借口。一旦真正消化赌注有多重,「不惜一切代价狂奔」就会显得荒谬。
逐段对照 · Pacing RSI
给递归自我改进定速
EN
If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery… I want to stress that… greatly accelerating deep learning research, especially in the short term, [is not necessarily] the right collective action we should take as the research community.
中文
若 AI 继续进步,机器递归自我改进(RSI)会成为未来科学发现的核心。自动化 AI 研究是用算力缩放智能的更剧烈形式。但我想强调:大幅加速深度学习研究——尤其短期——并不自动等于研究社区该采取的正确集体行动。路径会通向这里,但怎么走要自觉选择。
EN
The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures. The best way forward I see currently is a combination of both.
中文
主要杠杆两类:一边把对齐与监控跟 AI 进步一起抬,并想办法让人留在闭环里;一边在需要时协调放慢未来发展,以建立对这些措施的信心。我目前看到的最佳路径是两者结合。
EN
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars… The core challenge of automating AI research is not “getting there” — it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.
中文
缩放必须受「对安全有多少信心」约束。需要把 Preparedness Framework、Responsible Scaling Policy 一类承诺,演变成广泛强制的安全门槛——可由第三方审计、政府机构或国际组织执行。自动化 AI 研究的核心难题不是「到不到得了」,而是到的时候人还在不在持续改进过程里,未来是否还在人类手里。
逐段对照 · What is next?
接下来呢
EN
As great as the long-term promise of AI may be, the majority of our focus should be on the next few years… preserve human agency… prevent extreme concentration of power… ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
中文
AI 的长期承诺再大,我们的主要焦点也应放在未来几年:确保过渡对人类有利;保全人的能动性,并在多数任务可被 AI 完成的世界里,仍承认「作为人」的内在价值;防止极端权力集中(少数人操作大型计算机就能完成昔日数千专家的事业);确保人类仍掌控未来,不被失控进步甩下——那种进步来自超过我们自身的「外星智能」。
EN
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
中文
目前我认为,没有任何实验室把对齐与监控做到足够好,还能负责任地继续以最高速度扩张很久。我期待、也希望自愿减速成为常态,直到共同安全门槛建立。我也认为,未来 AI 发展上的国际协调,应成为各国政府的优先事项。

官网原文结尾两段 · 自愿减速与国际协调
再贴一次链接
https://openai.com/index/an-alien-mind
https://openai.com/index/jakub-pachocki-announced-as-chief-scientist/
https://www.businessinsider.com/openai-chief-scientist-ai-risks-slowdown-rogue-agents-consequences-safety-2026-9
中文为机智流意译,细节以英文原文为准。人物背景综合 OpenAI 公告与公开履历。科普整理,不构成投资或采购建议。