字号 ·· | 护眼
thenextweb

一名研究人员称,OpenAI的智能体对联合国统计网站进行了16500次扫描OpenAI agents scanned a UN statistics site 16,500 times, researcher says

点「原文对照」整页切到原文,或双击某段只看那段的原文。

一名与OpenAI有关联的研究人员发现,人工智能代理在4月13日至6月19日期间,扫描了联合国数据网站超过16,500次。

AI agents that a researcher linked to OpenAI scanned a United Nations data website more than 16,500 times between 13 April and 19 June.

罗万·霍华德-琼斯(Rowan Howard-Jones)在9月26日的一篇博客文章中写道,这些代理使用了代理服务器和一种编码技巧,以绕过其API的限制。

They used proxies and an encoding trick to get round limits on its API, Rowan Howard-Jones wrote in a blog post on 26 September.

霍华德-琼斯表示,该代理与OpenAI的关联可能性极高,但尚未得到证实,且TNW并未自行核实该分析结果。

Howard-Jones calls the OpenAI link highly likely, not proven, and TNW has not checked the analysis itself.

证据包括代理标记为“CHATGPTTEST1”和“OAI_META_1312”的页面。部分相同的微软Azure地址也出现在DseWiki集群背后,而OpenAI已确认该集群属于其自身。

The evidence includes pages the agents labelled CHATGPTTEST1 and OAI_META_1312. Some of the same Microsoft Azure addresses were also behind the DseWiki swarm, which OpenAI has confirmed was its own.

目标网站是UNCTADstat,即联合国贸易和发展会议的公共数据站点。根据请求内容判断,这些代理希望获取有关食品贸易、工业和生产能力的统计数据,但其具体任务尚不明确。

The target was UNCTADstat, the public data site of UN Trade and Development. Judging by the requests, the agents wanted figures on food trade, industry and productive capacity, but their exact tasks are unknown.

它们遇到的第一个问题是技术性的,因为代理似乎仅限于使用GET请求,而GET请求只能读取页面,但该网站的主要数据端点仅接受POST请求。

Their first problem was technical, as the agents appear to have been limited to GET requests, which only read pages, while the site’s main data endpoint takes POST requests only.

为了绕过这一限制,它们将页面发送至Urlquery——一款在隔离浏览器中打开链接的安全扫描器。这些页面包含表单,一旦加载完成,便会立即向UNCTADstat发送自身。

To get around this, they sent pages to Urlquery, a security scanner that opens links in a sealed-off browser. Those pages held forms that sent themselves to UNCTADstat as soon as they loaded.

在接下来的两个月里,这些方法变得更加复杂。代理通过第三方中继传输流量,并在谷歌的XSS游戏(一个旨在教授Web安全漏洞的网站)上托管脚本。

Over the next two months, the methods grew more complex. The agents sent traffic through third-party relays and hosted scripts on Google’s XSS game, a site built to teach web security flaws.

在某一时刻,它们还将关键词拆分为片段,以绕过一个实际上并不存在的过滤器。

At one point they also split keywords into pieces to slip past a filter that did not exist.

从5月4日起,代理使用双重编码来隐藏端点名称,从而使GET请求得以通过。霍华德-琼斯统计到共有55次此类请求。该网站还对82次请求进行了速率限制,但扫描活动仍在继续。

From 4 May, the agents used double encoding to hide the endpoint’s name, which let GET requests through. Howard-Jones counted 55 such requests. The site also rate-limited 82 requests, yet the scans carried on.

代理获取的所有数据原本就是公开的。据该文章称,霍华德-琼斯不会将此活动称为黑客行为,并在发布文章前已将编码绕过问题报告给UNCTAD的安全团队。

All the data the agents got was already public. Howard-Jones would not call the activity hacking, the post says, and reported the encoding bypass to UNCTAD’s security team before publishing.

斯坦福大学的网络安全讲师亚历克斯·斯塔莫斯(Alex Stamos)在接受《Journal》采访时指出,这种行为虽然近乎黑客攻击,但实际上主要属于一种极具侵略性的数据收集行为。

Alex Stamos, a cybersecurity lecturer at Stanford University, told the Journal the activity came close to hacking but was mostly very aggressive data gathering.

OpenAI向该媒体表示正在审查相关调查结果,并已向联合国提供了相关报告。

OpenAI told the paper it was reviewing the findings and had offered the UN a briefing.

这项分析是基于研究机构 Transluce 于 9 月 23 日发布的报告;该报告促使霍华德-琼斯(Howard-Jones)直接研究了公开可获取的 URL 查询记录。

The analysis follows a 23 September report by the research lab Transluce, which led Howard-Jones to study the public Urlquery records directly.

这一发现进一步印证了此前关于 OpenAI 的一些调查结果,例如今年 5 月份发生的 RubyGems 包被大量滥用的事件。

It adds to earlier findings about OpenAI agents, such as a RubyGems package flood in May.