2026-09-11
每日一文 · 长文精读

OpenAI Releases GPT-6 Astra for Coding and Computer Use

OpenAI 发布专注于编程和计算机使用的 GPT-6 Astra

作者:Daniel Dominguez · InfoQ 原文

摘要:OpenAI 发布了 GPT-6 Astra,这是一个专注于计算机使用、编程、专业工作流、科学和网络安全的新模型。该模型最初面向有限的组织提供,并逐步向 ChatGPT Plus、Pro、Business 和 Enterprise 用户以及 OpenAI API、Microsoft Azure 和 AWS Bedrock 推出。Astra 能够直接与图形界面交互,执行多步骤任务,在多项基准测试中表现优于前代模型。它还引入了实验性上下文机制,支持长达一百万 token 的上下文,并首次被归类为关键网络安全能力级别。社区反应热烈,包括英伟达 CEO 黄仁勋在内的多位人士将其视为 AGI 的里程碑。

OpenAI has released GPT-6 Astra, a new model focused on computer use, coding, professional workflows, science, and cybersecurity.
OpenAI 发布了 GPT-6 Astra,这是一个专注于计算机使用、编程、专业工作流、科学和网络安全的新模型。
The model is initially available to a limited set of organizations and is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, as well as the OpenAI API, Microsoft Azure, and AWS Bedrock.
该模型最初面向有限的组织提供,并逐步向 ChatGPT Plus、Pro、Business 和 Enterprise 用户以及 OpenAI API、Microsoft Azure 和 AWS Bedrock 推出。
Astra extends OpenAI's models beyond generating responses toward performing multi-step tasks directly in software.
Astra 将 OpenAI 的模型从生成回复扩展到直接在软件中执行多步骤任务。
It can interact with graphical interfaces to fill forms, update CRM records, conduct research, create websites, analyze data, install and test software, and troubleshoot problems visible on screen.
它可以与图形界面交互,填写表单、更新 CRM 记录、进行研究、创建网站、分析数据、安装和测试软件,以及排查屏幕上可见的问题。
On OSWorld 2.0, OpenAI reports a score of 72.6%, compared with 65.7% for GPT-5.6 Sol.
在 OSWorld 2.0 上,OpenAI 报告得分为 72.6%,而 GPT-5.6 Sol 为 65.7%。
Coding is another focus of the release.
编程是此次发布的另一个重点。
OpenAI reports 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1.
OpenAI 报告在 Terminal-Bench 4.0 上得分为 57.9%,在 DeepSWE v1.1 上得分为 74.1%。
Astra also introduces an experimental context mechanism in Codex that allows the agent to maintain notes across context windows instead of relying only on compaction.
Astra 还在 Codex 中引入了一种实验性上下文机制,允许代理在上下文窗口之间维护笔记,而不仅仅依赖于压缩。
Previous context windows remain searchable, allowing the model to retrieve earlier requirements, test results, and tool outputs during long-running coding tasks.
之前的上下文窗口保持可搜索,允许模型在长时间运行的编程任务中检索早期需求、测试结果和工具输出。
The model supports long contexts of up to one million tokens in OpenAI's reported MRCR evaluations, scoring 96.3% in the 512K-to-1M range.
该模型在 OpenAI 报告的 MRCR 评估中支持高达一百万个 token 的长上下文,在 512K 到 1M 范围内得分为 96.3%。
OpenAI also reports improvements in professional tasks including database migrations, CAD generation, data science, browser research, and scientific workflows.
OpenAI 还报告了在专业任务上的改进,包括数据库迁移、CAD 生成、数据科学、浏览器研究和科学工作流。
Cybersecurity represents a significant change from previous OpenAI models.
网络安全代表了与之前 OpenAI 模型的重大变化。
Astra is the first OpenAI model classified at the critical cybersecurity capability level under the company's Preparedness Framework.
Astra 是第一个根据公司 Preparedness Framework 被归类为关键网络安全能力级别的 OpenAI 模型。
In testing without production safeguards, OpenAI says the model discovered and used two previously unknown vulnerabilities and demonstrated the ability to develop exploits against hardened browsers and operating systems.
在无生产安全措施的测试中,OpenAI 表示该模型发现并利用了两种先前未知的漏洞,并展示了针对加固浏览器和操作系统开发漏洞利用的能力。
The production version restricts advanced offensive tasks, while OpenAI plans to provide broader defensive capabilities through its Daybreak program.
生产版本限制了高级攻击性任务,而 OpenAI 计划通过其 Daybreak 计划提供更广泛的防御能力。
OpenAI also reports lower hallucination rates in its internal evaluation, with Astra scoring 4.2% compared with 12.2% for GPT-5.6 Sol.
OpenAI 还报告了内部评估中较低的幻觉率,Astra 得分为 4.2%,而 GPT-5.6 Sol 为 12.2%。
However, the company found Astra's written reasoning harder to monitor than its predecessor in tests designed to measure whether a model could obscure its reasoning.
然而,公司发现,在旨在衡量模型是否可能隐藏其推理的测试中,Astra 的书面推理比其前代更难监控。
OpenAI said improving this monitorability remains an active research area.
OpenAI 表示,提高这种可监控性仍然是一个活跃的研究领域。
Community reaction included comments from Nvidia CEO Jensen Huang, who highlighted the infrastructure used to train the model, saying: GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72.
社区反应包括英伟达 CEO 黄仁勋的评论,他强调了用于训练该模型的基础设施,并表示:GPT-6 Astra,在约 10 万+ NVIDIA Grace Blackwell NVLink72 上训练。
From ChatGPT to o1 to Astra in 4 years.
从 ChatGPT 到 o1 再到 Astra,历时 4 年。
AGI has arrived.
AGI 已经到来。
Congratulations @OpenAI team.
祝贺 @OpenAI 团队。
400K GPUs coming online next.
接下来将有 40 万块 GPU 上线。
Another community voice, Alex Finn, similarly focused on the AGI framing around the release, pointing to the characterization of the current period sharing: Welcome to AGI.
另一位社区人士 Alex Finn 同样关注围绕此次发布的 AGI 框架,指出当前时期的特征分享:欢迎来到 AGI。
ChatGPT 6 Astra just released.
ChatGPT 6 Astra 刚刚发布。
Astra competes with models including Anthropic's Claude Fable 5.1 and Google's Gemini 3.8 Flash.
Astra 与包括 Anthropic 的 Claude Fable 5.1 和 Google 的 Gemini 3.8 Flash 在内的模型竞争。
OpenAI's published evaluations show results varying by task: Astra leads the compared models on Terminal-Bench 4.0 and several computer-use evaluations, while Claude Fable 5.1 scores higher on Humanity's Last Exam and Google's Gemini 3.8 Flash supports native video and audio input that Astra does not.
OpenAI 发布的评估显示结果因任务而异:Astra 在 Terminal-Bench 4.0 和多项计算机使用评估中领先于对比模型,而 Claude Fable 5.1 在 Humanity's Last Exam 上得分更高,Google 的 Gemini 3.8 Flash 支持 Astra 所不具备的原生视频和音频输入。

阅读理解

1. What is a key new capability of GPT-6 Astra compared to previous OpenAI models?

2. What significant cybersecurity capability does GPT-6 Astra demonstrate according to the article?

3. How does Astra's performance on Terminal-Bench 4.0 compare to its competitors according to the article?

温故复习 →每日一句 →