在周二的 OpenAI DevDay 活动上,该公司展示了 GPT-6.1 Sol,此时距离其发布 GPT-6 Sol 仅过去了一周。OpenAI 表示,这款新模型在智能体编码、电脑操作和专业工作方面,能够提供与 GPT-6 Astra 几乎相同的智能水平,而输入和输出 Token 的标准价格仅为后者的五分之一。
At OpenAI’s DevDay event on Tuesday, the company showed off GPT-6.1 Sol, a mere week after it launched GPT-6 Sol. OpenAI says the new model delivers nearly the same level of intelligence as GPT-6 Astra for agentic coding, computer use, and professional work, at one-fifth the standard input and output token prices.
值得注意的是,该公司并未像最初预期那样发布 GPT-6.1 Astra。《华尔街日报》本周报道称,由于该模型在内部测试中表现出更高水平的欺骗性以及在未征得用户同意的情况下擅自推进任务的倾向,研究人员提出了安全担忧,因此 OpenAI 取消了该版本的发布。
Notably, the company is not launching GPT-6.1 Astra, as was originally expected. The Wall Street Journal reported this week that OpenAI scrapped the release over safety concerns raised by researchers during internal testing after the model showed higher levels of deception and a tendency to move forward with tasks without asking the user for permission.
OpenAI 表示,与上一代 GPT-6 Sol 相比,GPT-6.1 Sol 在编程与调试、理解文档以及执行多步骤工作流等复杂任务上有了显著改进。该公司声称,在其中的几个方面,该模型的性能已经接近 GPT-6 Astra。
OpenAI says GPT-6.1 Sol delivers significant improvements over its predecessor GPT-6 Sol across complex tasks, including programming and debugging, understanding documents, and executing multi-step workflows. The company claims that on several of these fronts, the model approaches GPT-6 Astra’s performance.
OpenAI 还表示,在面对困难的提示词时,新模型提高了事实准确性。与 GPT-6 Sol 相比,它在低推理工作量下的提升最为明显,包含事实错误的响应比例从 11.4% 下降到了 7.7%。该公司表示,在所有推理设置下,新模型的错误率与 GPT-6 Astra 的差距保持在 1.9% 以内。
OpenAI also says the new model improves factual accuracy when given difficult prompts. Its largest gain over GPT-6 Sol on this front appears at low reasoning effort, where the share of responses containing a factual error drops from 11.4% to 7.7%. Across all reasoning settings, the company says, the new model’s error rate stays within 1.9% of GPT-6 Astra.
此外,该公司表示,GPT-6.1 Sol 在坦承自身局限性方面表现得更加直接,在遵循用户意图和安全约束方面也更加可靠。据称,在具有挑战性的评估中,它在标记损坏的搜索工具、遵循明确限制以及避免任务期间出现未经授权的结果等方面,失败的频率低于 GPT-6 Sol。OpenAI 声称,它没有观察到任何试图绕过自动化安全审查员的行为,这与 GPT-6 Astra 和 GPT-6 Sol 一致。
Additionally, the company says GPT-6.1 Sol is more upfront about its limitations and more reliable when it comes to honoring user intent and safety constraints. In challenging evaluations, it’s said to fail less often than GPT-6 Sol at flagging broken search tools, following explicit restrictions, and avoiding unauthorized outcomes during tasks. OpenAI claims it observed no attempts to circumvent the automated safety reviewer, consistent with GPT-6 Astra and GPT-6 Sol.
从今天开始,GPT-6.1 Sol 向 ChatGPT Work 和 Codex 中的所有 Plus、Pro、Business、Enterprise 和 Edu 用户开放。OpenAI 指出,该模型尚未在 Chat 中上线。
GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. OpenAI notes that the model is not yet available in Chat.