字号·· | 护眼
thenextweb

OpenAI 722篇数学论文由其尚未发布的系统撰写OpenAI publishes 722 maths papers written by a model it has not released

点「原文对照」整页切到原文,或双击某段只看那段的原文。

图片来源:Zac Wolff,UnsplashCredit: Zac Wolff on Unsplash

OpenAI公布了由其尚未对外发布的内部模型生成的722篇数学论文。这些论文可归为372类研究成果。该公司于周二宣布了这一发布计划,并将所有论文上传至一个公开的GitHub代码库。

OpenAI has published 722 maths manuscripts from an internal model it has not released. The papers fall into 372 families of results. The company announced the release on Tuesday and put them all in a public GitHub repository.

该模型正是上个月OpenAI完成纳维-斯托克斯方程证明时所使用的模型。据代码库信息显示,OpenAI在评估过程中向该模型提供了约4000道数学题。平均而言,生成每一条研究成果需要ChatGPT Pro耗费约3小时的运算时间。

It is the same model behind OpenAI’s Navier-Stokes proof last month. OpenAI gave it about 4,000 problems during the evaluation, the repository says. On average, each result used about three hours of ChatGPT Pro thinking compute.

代码库中包含的内容:许多证明过程还附带了用Lean语言编写的形式化版本——这种编程语言可帮助计算机验证证明的正确性。不过也有部分证明未附带此类版本。OpenAI表示,那些缺乏形式化版本的成果可能存在问题,公司会尽快予以修正。所有早期版本均会予以保留,每篇论文也都配有独立的引用说明。

What is in the repository Many of the proofs come with formal versions in Lean, a programming language that lets a computer check a proof. Not all do. OpenAI says some of the results without a formal version could have issues, and it will fix them quickly. Every earlier version will stay public, and each paper has its own citation block. OpenAI also published ten summaries of the model’s reasoning.

OpenAI还公布了该模型进行推理过程的十份总结报告。其中有两项成果的生成过程并未遵循常规流程:其一为黎曼ζ函数无零点区域的证明,该内容经人工编辑后才得以发布;其二则为关于复乘阿贝尔簇的霍奇猜想证明。

Two results did not follow the standard procedure. One is a zero-free region for the Riemann zeta function, whose write-up a human edited for readability. The other is a proof of the Hodge conjecture for CM abelian varieties.

OpenAI向《科学美国人》透露,几乎所有论文均是通过向单一AI智能体发出单一指令后生成的。该公司还计划为基于AI产生的研究成果举办研讨会、学术会议及专项研究项目。目前,OpenAI正致力于以负责任的方式对外发布该模型。

Nearly every paper came from a single prompt to a single AI agent, OpenAI told Scientific American. The company will fund workshops, conferences and special programmes on results produced by AI. It says it is also working to release the model responsibly.

专家称相关工作才刚刚起步:OpenAI称其制定相关举措时参考了由数学家组成的独立顾问团的意见。该顾问团名为“数学与人工智能咨询小组”(AGMAI),隶属于高等研究院。在9月29日发布的指导准则中,该小组要求各人工智能实验室停止在私有模型上测试高难度数学问题,同时必须公开每项成果生成时所使用的指令、耗时及运算成本。

Advisers say the work has only begun OpenAI said it drew on advice from an independent panel of mathematicians. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) is hosted by the Institute for Advanced Study. In its 29 September guidelines, it asked AI labs to stop testing hard maths problems on private models. It also asked them to publish the prompts, time and compute cost behind each result.

“人工智能实验室利用专有内部模型开展数学研究,此举可能催生一种两极分化的局面,导致部分实验室远远领先于整个学界,”该小组在准则中强调道。

“The use of proprietary internal models by AI labs to do mathematical research risks creating a two-tier system where labs outrun the rest of the field,” the group wrote.

周二晚间,AGMAI表示,其意见并不等同于对这些研究成果的认可,也不代表对OpenAI取得这些成果方式的肯定。

On Tuesday night, AGMAI said its advice was not an endorsement of the results, or of how OpenAI got them.

“此次成果的发布只是人类理解该问题的开端,远非终点,”该组织写道。

“This release is the beginning, not the completion, of the process of human understanding,” the group wrote.

OpenAI的研究负责人丹·罗伯茨向《纽约时报》表示,对内部模型进行测试有助于开发出更优秀的工具。他指出,那些数学证明只是测试过程中的副产品。

Dan Roberts, OpenAI’s research lead, told The New York Times that testing internal models helps build better tools. The proofs were a byproduct, he said.

纽约大学数学家特里斯坦·巴克马斯特在OpenAI解决纳维-斯托克斯问题之前便一直在研究该课题。他告诉《纽约时报》,自己怀疑OpenAI是否真的对一次性发布的众多成果逐一进行了验证。“我认为他们根本没有履行应有的审查程序,”巴克马斯特说道。

Tristan Buckmaster, a New York University mathematician, was working on the Navier-Stokes problem before OpenAI solved it. He told the Times he doubted OpenAI had checked so many results released at once. “I don’t think they’ve done their sort of due diligence at all,” Buckmaster said.

麻省理工学院数学家安德鲁·萨瑟兰则对《科学美国人》表示,目前关于单一智能体即可完成这些工作的说法尚无法证实。在其他人也能运行该模型并复现相关结果之前,这一说法都应被视为未经证实。

MIT mathematician Andrew Sutherland told Scientific American to treat the single-agent claims as unverified. That should hold until others can run the model and repeat the results, he said.