Anthropic发布了Claude Sonnet 5.5,在Terminal-Bench 4.0智能体编程测试中得分70.6%,而Sonnet 5仅为10.3%,更昂贵的Opus 5.5为66.4%。这是首个配备前沿级网络安全防护措施和阻止推理提取分类器的Sonnet模型。
Anthropic has released Claude Sonnet 5.5, which scores 70.6% on the Terminal-Bench 4.0 agentic coding test against 10.3% for Sonnet 5 and 66.4% for the more expensive Opus 5.5. It is the first Sonnet to ship with frontier-style cyber safeguards and with classifiers that block reasoning extraction.
Anthropic发布了Claude Sonnet 5.5,该公司表示其运行速度比前代快30%以上,每项任务成本最多降低30%。定价为每百万输入token 2美元,每百万输出token 10美元。在Terminal-Bench 4.0智能体编程测试中得分70.6%,而Sonnet 5仅为10.3%。这超越了旗舰模型。
Anthropic has released Claude Sonnet 5.5, which it says runs more than 30% faster and costs up to 30% less per task than its predecessor, the company said. It is priced at $2 per million input tokens and $10 per million output. It scores 70.6% on Terminal-Bench 4.0, an agentic coding test, against 10.3% for Sonnet 5. That beats the flagship.
Anthropic报告称,Opus 5.5在相同测试中以最高努力设置得分66.4%,而Opus每token成本是Sonnet 5.5的两倍。在涵盖44种职业的GDPval-AA测试中,Sonnet 5.5得分1,844,Opus 5.5为1,846,Sonnet 5为1,449。该公司表示,Opus在需要持续判断的开放式工作中仍然明显更强。防护措施下移了一个层级。
Anthropic reports Opus 5.5 at 66.4% on the same test at its highest effort setting, and Opus costs twice as much per token. On GDPval-AA, a test across 44 occupations, Sonnet 5.5 scores 1,844 against Opus 5.5’s 1,846 and Sonnet 5’s 1,449. The company says Opus remains clearly stronger at open-ended work requiring sustained judgement. The safeguards moved down a tier.
Anthropic表示,该模型的网络安全能力与Opus 5相当,因此发布时即附带此前仅为其最强模型保留的限制措施。高风险网络请求将明显回退至Sonnet 5。这是首个以这种方式发布的Sonnet,其生物学防护措施保持不变。第二道防线是蒸馏防护。
Anthropic says the model’s cybersecurity capabilities are comparable to those of Opus 5, so it launches with restrictions of the kind previously reserved for its most capable models. Higher-risk cyber requests will visibly fall back to Sonnet 5. It is the first Sonnet to ship that way, and its biology safeguards are unchanged. The second defence is distillation.
这也是首个配备阻止推理提取分类器的Sonnet,其保留思维功能将模型的推理过程与产生它的账户绑定。8月,研究人员从OpenAI、Anthropic和Google系统中6,708条公开智能体轨迹中解码了315,320个思维块。他们恢复了62个API密钥、33个密码和7个私钥。欧洲要求此类保护。
It is also the first Sonnet with classifiers that block reasoning extraction, and its preserved thinking ties a model’s reasoning to the account that produced it. In August, researchers decoded 315,320 thinking blocks from 6,708 public agent traces across OpenAI, Anthropic and Google systems. They recovered 62 API keys, 33 passwords and seven private keys. Europe requires that protection.
《人工智能法案》第55条要求具有系统性风险的通用模型提供者保护模型及其物理基础设施,同时进行对抗性测试,并及时报告严重事件。这些义务自2025年8月起适用。委员会于今年8月获得了对违规行为处以罚款的权力。能力声明是一把双刃剑。
Article 55 of the AI Act obliges providers of general-purpose models with systemic risk to protect the model and its physical infrastructure, alongside adversarial testing and reporting serious incidents without undue delay. Those duties have applied since August 2025. The Commission gained the power to fine breaches this August. The capability claim cuts both ways.
Anthropic表示,Sonnet 5.5并未推进其模型能力的前沿,这就是其对齐评估覆盖的风险范围较窄的原因。该公司还表示,网络能力有大幅提升,值得采用前沿级别的保障措施。这两种说法出现在同一份公告中。有一个价格需要核实。
Anthropic says Sonnet 5.5 does not advance the frontier of its models’ capabilities, which is why its alignment assessment covered a narrower set of risks. It also says the cyber capabilities are a large improvement warranting frontier-style safeguards. Both statements appear in the same announcement. One price needs checking.
Anthropic将2美元和10美元描述为与Sonnet 5相同的定价。我们6月的报告称,这些是截至8月31日的 introductory 价格,此后Sonnet 5将收费3美元和15美元。要么涨价从未发生,要么比较对象是 introductory 价格。
Anthropic describes $2 and $10 as the same pricing as Sonnet 5. Our report in June said those were introductory rates until 31 August, after which Sonnet 5 would cost $3 and $15. Either the increase never happened or the comparison is to the introductory price.