David Ha 是 Sakana AI 的联合创始人兼首席执行官,位于东京。
David Ha is co-founder and CEO of Sakana AI, based in Tokyo.
世界对前沿 AI 模型着迷。数十亿美元被投入到训练更大规模的单体系统,基于规模本身能保证主导地位的假设。但这种策略正变得经济上不合理。开源替代方案现在仅落后几个月,而许多前沿模型的推理成本已超过它们旨在提升的员工的时薪。“一模型统治一切”的方法正走向递减收益。
The world is obsessed with frontier AI models. Billions are poured into training ever-larger monolithic systems, on the assumption that scale alone guarantees dominance. But this strategy is becoming economically irrational. Open-source alternatives now trail by only a few months, and the inference costs of many frontier models exceed the hourly wages of the employees they are meant to augment. The "one model to rule them all" approach is a path to diminishing returns.
我于 2023 年共同创立的东京实验室 Sakana AI 并非在打造 OpenAI 的日本版本。我们正在构建一种不同的智能架构,根植于约束、集体行动和文化主权。
Sakana AI, the Tokyo-based lab I co-founded in 2023, is not building the Japanese version of OpenAI. We are building a different architecture for intelligence, rooted in constraints, collective action and cultural sovereignty.
真正的战略价值在于提升到堆栈的更高层次,从原始模型权重到协调它们的智能。下一阶段 AI 的赢家不是构建最大单体的公司,而是掌握编排的公司。
The real strategic value is moving up the stack, from the raw model weights to the intelligence that coordinates them. The winners of the next phase of AI will not be those who build the largest monolith, but those who own the orchestration.
鱼群 这是 David Ha 为 Sakana AI 早期概念化标志设计的一些方案。他说,在多次尝试后,最终确定了正确的设计。
The school of fish Some of David Ha's early conceptual logo designs for Sakana AI. He says that after many attempts, he eventually settled on the right one.
Sakana 在日语中意为鱼,名字反映了我们的核心赌注:集体智能。正如鱼群以单个鱼无法做到的方式行动,我们相信没有单一公司或模型会主导 AI。每个模型都有其自身的偏见、优势和盲点。未来属于模型编排者:一个学习如何协调模型池以解决复杂多步骤问题的系统。智能不仅仅是单个鱼,而是整个鱼群。
Sakana means fish in Japanese, and the name reflects our central bet: collective intelligence. Just as a school of fish behaves in ways no single fish can, we believe no single company or model will dominate AI. Every model carries its own biases, strengths, and blind spots. The future belongs to the model orchestrator: a system that learns how to coordinate a pool of models to solve complex, multistep problems. Intelligence is not just an individual fish.
我们的编排模型 Fugu 本身就是一个基础模型(即用于处理各种任务的底层模型)。与那些遵循固定规则、由人工编程设计的系统不同,Fugu 是通过强化学习在成千上万条多步骤推理链上被端到端训练出来的。它能够学习在每个步骤中应该调用哪个基础模型、如何处理相互矛盾的输出结果、如何在成本、性能与合规性之间找到平衡,以及何时将敏感数据路由到自托管的模型中,而不是云服务 API。
It's the school. Our orchestration model, Fugu, is itself a foundation model. Unlike a hard-coded router that follows static human rules, Fugu is trained end-to-end through reinforcement learning on thousands of multistep reasoning chains. It learns which underlying model to invoke at each step, how to synthesize conflicting outputs, how to balance cost, performance and compliance, and when to route sensitive data to a self-hosted model rather than a cloud API.
单个基础模型与这种经过训练的编排层之间的区别,就如同大型计算机与个人电脑之间的区别一样显著。
The difference between an individual foundation model and this learned orchestration layer is the difference between a mainframe and a personal computer.
在一个逐渐走向混合式(本地计算与云计算相结合)计算模式的世界里,编排系统就相当于操作系统:某些任务需要高性能的模型来处理;而另一些任务则可以通过运行在本地设备上的小型、高效模型来完成(这些设备的功耗可能仅为 150 瓦)。真正的智能在于能够判断何时、使用哪种模型,以及使用它们的原因。如果某个模型出现故障或改变了其运行策略,系统可以自动绕过该模型的影响,继续正常运行。
In a world moving toward hybrid local and cloud inference, the orchestrator becomes the operating system. Some tasks require the brute force of a frontier model. Others can be handled by a small, efficient model running on a local device at 150 watts. The intelligence lies in knowing which to use, when and why. And if one model goes offline or changes its policies, the system simply routes around the damage.
关于人工智能主权的争论常常被误解了。其核心并不在于构建一个完全孤立、完全依赖本国技术的系统;没有任何国家(包括美国或中国)能够独自掌握所有相关技术。例如,日本需要使用英伟达(Nvidia)的图形处理单元(GPU);而美国的模型则是利用全球数据来进行训练的。真正的主权体现在供应链的韧性上——即具备在国内开发、调整和维护人工智能技术的能力,同时还能在不依赖任何单一供应商的情况下,灵活地利用全球资源来运行人工智能系统。
Sovereignty is a supply chain, not a wall The debate around AI sovereignty is often misunderstood. It is not about building a perfectly isolated, purely domestic model. No country, not even the U.S. or China, owns the entire stack. Japan needs Nvidia graphics processing units (GPUs); American models are trained on global data. True sovereignty is supply-chain resilience: the know-how to develop, adapt and maintain AI domestically, combined with the ability to orchestrate across global resources without depending on any single vendor. This is why post-training matters.
这就是为什么“训练后优化”(post-training optimization)如此重要的原因。我们用于开发 Sakana Chat 和 Sakana Translate 的 Namazu 模型并非从零开始训练的;它们是基于现有的开源基础模型进行开发的,随后再使用我们专门为日语、日本文化及商业场景设计的数据集进行进一步训练,以减少不必要的错误处理和偏见。
Our Namazu models, which power Sakana Chat and Sakana Translate, are not trained from scratch. They start from open-weight base models and are then post-trained on our own datasets for Japanese language, culture and business contexts, with additional tuning to reduce unnecessary refusals and bias.
在涉及政治敏感内容的场景中,Namazu 在中立性表现上优于其基础模型;在全面的日语能力评估中,Namazu 也表现出色;同时,它在翻译质量上也具有明显优势。文化的适应性并不意味着需要牺牲系统的功能或性能。
Namazu outperforms its base models on benchmarks for neutrality in politically sensitive Japanese contexts, on comprehensive Japanese-language evaluations, and on translation. Cultural alignment does not require sacrificing capability.
然而,真正的“保险措施”在于系统的灵活性与冗余性——如果某个服务提供商关闭了 API 接口,我们的路由系统会自动切换到其他可用服务。在当今地缘政治日益分裂的时代,这种灵活性对于任何希望掌控自身数字基础设施的国家来说都至关重要。
But the ultimate insurance policy is orchestration. If one provider shuts off API access, the routing model fails over to others. Sovereignty is continuity through diversity. In an era of geopolitical fragmentation, this is not an academic concern. It is an existential requirement for any nation that wishes to control its own digital infrastructure.
由于资源限制(尤其是资金和人力),日本没有能力每个季度都从头开始训练参数量高达数万亿的模型。这种资源约束直接影响了我们的研究方向以及产品上市策略。
Building under constraint Japan does not have the capital or energy to train trillion-parameter models from scratch every quarter. That constraint shapes how we do research and how we go to market.
在研究领域,我们专注于“递归自我优化”(Recursive Self-Improvement, RSI)技术:这种技术允许 AI 系统帮助设计和训练下一代 AI 系统。虽然 RSI 有时被视作科幻概念(即 AI 系统通过自我修改代码来提升自身能力),但对我们而言,它其实是一种经济上的必然选择。我们发表在《自然》(Nature)杂志上的 AI Scientist 框架证明了:AI 系统已经能够生成科学见解、编写代码、运行实验并整理研究结果。
In research, our focus is recursive self-improvement, or RSI: AI systems that help design and train the next generation of AI systems. RSI is often framed as science fiction, an AI rewriting its own code toward god-like intelligence. For us, it is an economic imperative. Our AI Scientist framework, published in Nature, demonstrates that systems of AI agents can already generate scientific ideas, write code, run experiments and compile findings.
如果将这一技术应用于模型开发过程中,预计可以大幅降低预训练成本(降低一到两个数量级)。我们的目标并非完全取代人类的工作,而是让人类判断在关键决策中发挥关键作用(例如在投入大量资金进行计算之前进行合理性审查)。系统的准确性、错误检测能力与运行效率是相辅相成的;正是通过这种方式,资金相对有限的团队也能实现超越资金雄厚竞争对手的创新。
Applied to model development itself, we estimate this could reduce pre-training costs by one to two orders of magnitude. The goal is not to remove humans from the loop, but to use human judgment as a sanity check before committing millions of dollars to compute. Alignment, error-checking and efficiency go hand in hand. This is how a smaller player out-innovates a better-funded rival.
在市场上,硅谷的许多企业都坚持以消费者为中心的策略;而我们却选择了相反的道路。我们的首批项目是与日本最保守的机构合作的——那些大型银行、证券公司以及政府机构。如果能够在日本的大型银行中成功应用人工智能技术,那么这项技术就可以在任何地方得到推广。我们与三菱UFJ金融集团(MUFG)的合作成果是“AI贷款专家”(AI Lending Expert)这一系统:该系统能够生成详细的信贷评估报告,供人类贷款专员审核或修改。
In the market, many in Silicon Valley go consumer-first. We chose the opposite path. Our first deployments were with Japan's most conservative institutions: megabanks, securities companies and government agencies. If you can deploy AI in a Japanese megabank, you can deploy it anywhere. Our partnership with MUFG has produced the AI Lending Expert, a production system that generates detailed credit memos that a human loan officer can affirm or override.
另一个例子是用于并购业务的“文档自动化工具”(Document Agent)——它将制作并购提案文件的时间从几周缩短到了几小时,同时确保了银行对最终文件的所有权。这种“人在决策过程中的参与”设计并非技术上的限制,而是一种文化传统,它确保了决策的透明性与责任性。
A document agent for M&A cuts pitch-deck production from weeks to hours, but the banker retains ownership. This human-in-the-loop design is not a technological limitation; it is a cultural feature that ensures accountability.
这些深度合作的伙伴关系实际上就是我们的“实验室”;日本的企业文化虽然常被认为效率低下,但实际上却是一种竞争优势。这些企业要求软件具备极高的可靠性,他们不会因为未经验证的技术而裁员,而是会奖励那些能够带来实际成果的团队。这种严谨的态度迫使我们开发出真正实用、可靠的软件。
These deep co-development partnerships function as our laboratory. Japan's enterprise culture, often criticized as slow, is a competitive advantage. These companies demand production-grade reliability. They do not celebrate laying off staff for unproven technology; they reward results. That discipline forces us to build software that actually works.
日本成功的最后一个、非技术性的关键因素就是“乐观主义”。如果我有一根魔法杖,我会让整个国家的人民都充满乐观情绪。因为乐观会带来变革,而变革又会进一步激发更多的乐观情绪。
The need for optimism There is one final, nontechnical ingredient to Japan's success: hope. If I had a magic wand, I would make everyone in the country optimistic. With optimism comes change. And change leads to more optimism.
我在日本生活了多年,亲眼见证了当人们相信变革是可能的时候,日本社会能够迅速、集体地行动起来。战后经济的奇迹以及索尼、松下等企业的崛起,并不仅仅依靠充足的资本,更离不开一种共同的使命感。日本有着将资源匮乏转化为竞争优势的传统。
I have lived in Japan for years, and I have seen its capacity for rapid, collective mobilization when people believe change is possible. The post-war economic miracle and the rise of the likes of Sony and Panasonic were not driven by abundant capital alone, but by a collective sense of purpose. Japan has a history of turning scarcity into an advantage.
时机已经成熟:日本正从通货紧缩转向通货膨胀。日本的银行持有大量资金,这些资金亟需被投入使用;同时,日本的企业在生产和运营中采用人工智能技术的速度远超全球其他同行。现在所需要的,是一种乐观的态度和积极的舆论导向。从传统的、单一的商业模式转向更加灵活的集体协作模式,从依赖进口技术转向发展自主的基础设施,需要一个愿意相信“新的发展模式是可能的”、并且相信“日本是实现这一目标的理想之地”的社会。
The conditions are ripe. Japan is transitioning from deflation to inflation. Its banks hold vast capital that must be deployed. Its enterprises are adopting AI in production faster than their global peers realize. What is needed now is a narrative of optimism. Moving from monolithic models to collective systems, from imported technology to sovereign infrastructure, requires a society willing to believe that a different architecture is possible, and that Japan is the right place to build it.
日本历来擅长将资源短缺转化为竞争优势。如今,计算能力、能源和资本的短缺正在迫使我们构建更加智能、更加灵活、效率更高的系统。未来的十年里,人工智能领域的竞争将不会取决于谁拥有最多的GPU(图形处理单元),而是取决于谁拥有最先进的系统架构——那些能够自主调整、自我优化,并且能够真正理解并适应其所服务社会的文化和运营需求的系统。
Japan has a history of turning scarcity into an advantage. Today, the constraints of compute, energy and capital are forcing us to build smarter, more agile, and vastly more efficient systems. The next decade of AI will not be won by those with the most GPUs. It will be won by those with the most intelligent architecture: systems that route, adapt, improve themselves and respect the cultural and operational realities of the societies they serve.