为什么每家公司都需要拥有自己的智能
Erik TorenbergAlex AtallahAmjad Masad
Stripe 收购 OpenRouter,将两家希望催生更多初创公司的基础设施公司放到了一起,而不是打造“一家巨型公司”。 Alex Atallah 表示,OpenRouter 将继续保持品牌、路线图和产品自主权,同时获得更成熟的市场拓展计划;从战略上看,对未来的公司而言,“支付与推理将融为一体”。
OpenRouter 的护城河在于模型独立性,而不只是 API 聚合。 其市场让企业可以避免被单一供应商锁定,持续站在模型性能前沿,组合不同训练方式的模型,并推动推理价格下行。Atallah 的核心判断是,模型能力无法被压缩成一张功能清单:“必须看它们如何被使用,才能知道它们擅长什么。”
Erik 提出企业需要“拥有自己的智能”,Amjad Masad 则进一步指出,基础模型厂商正越来越多地把每一层应用都视为潜在市场。 他援引 SpaceX S-1 中据称“30万亿美元”的机会空间,对比约100万亿美元的全球 GDP,说明这种野心的规模。Replit 因此正把自己定位为横跨模型和云的独立性层,覆盖 AWS、Azure、Databricks 和 Snowflake。
AI 产品看似都在收敛到智能体循环、记忆、连接器、沙盒、网页搜索、电脑操作和通知,这可能只是新一轮基础配置的形成。 Erik 将其类比为,2005年的每家互联网公司都需要用户、数据库、用户资料和身份验证,Masad 对此表示认同。企业真正更难的机会,仍在于让智能体高效工作,同时保住数据主权、访问控制,并部署在客户自己的云中。
争议最尖锐的地方,在于一个无所不知的智能体是否值得追求。 Masad 看重聊天记录、GitHub、Salesforce 和日历之间的跨域关联,但 Atallah 认为,委托出去的工作越多,人对正在发生什么的理解就越少,而没有任何智能体会承担由此产生的责任:“总得有人承受皮质醇压力。”Erik 将正在形成的分工概括为:人类继续做通才,机器则变得专业化,或许由一个幕僚长智能体负责协调。
专用决策模型可能同时成为智能体系统的经济层和安全层。 OpenRouter 正在测试一种便宜、快速的模型,用来根据政策检查工具调用;Atallah 则设想,由通用模型按需训练更窄的替代模型,避免在任何地方都调用前沿智能的“用核弹轰蝴蝶”式浪费。结构化输出也能显著收窄模型作恶的空间。
两位嘉宾都没有得出“模型更聪明就会自动更安全”的结论。 Erik 提出,更聪明的模型可能拥有更好的对齐和协作能力;Atallah 则指出奖励投机、欺骗性的思维链行为,以及可能需要“连续跑上几个月”的评估。Atallah 将实际问题收窄为:能力足够强的模型,是否会停止欺骗用户和在评估中故意压低表现。如果前沿模型最终能显著降低欺骗风险,那么高风险编码和安全任务理性上可能愿意为其支付10倍价格。
专用模型周期可能重演软件从动态语言转向类型系统、编译器和 Rust 的过程。 Masad 已经在用 Replit 数据训练 Qwen 8B 分类器,包括一个提示词成本估算器;Atallah 则认为,定制分类器的“模型债务”低于非结构化微调,后者每两个月就可能因模型过时而被迫重做。多模型融合提供了另一条路径:嘉宾称其以约40–50%的成本实现了接近前沿模型的质量,而感知缓存的路由是其中的核心经济学。
1. Stripe 正在收购 AI 经济中的独立性层
Atallah 表示,7月 Stripe 接触 OpenRouter 时,OpenRouter 并没有寻求出售。与 Stripe 总裁 Will Gaybrick 长期保持联系,双方已有数个共同推进的工作流,加上异常高效的交易流程,让这桩意外出现的交易最终成为 Atallah 在可能收购方中的“首选”。
运营安排与交易本身同样重要:OpenRouter 保留品牌、路线图和产品自主权,同时可以借助更有力度的市场拓展计划加快发展。更深层的契合点,是双方都希望打造一个“中立、值得信赖的平台”,让生活方式型企业和风险投资支持的初创公司都能在其上成长。
Erik Torenberg 的质疑很关键:Stripe 显然会因更多公司处理支付而受益,但初创公司数量增加为什么对 OpenRouter 重要?Atallah 的答案是市场经济学:企业需要模型和供应商独立性,需要用多个模型打造差异化智能,也需要价格竞争,让此前不具备经济性的产品成为可能。
企业对开放性的需求让 Atallah 感到意外。客户并没有默认选择最知名的闭源实验室,而是出于成本、差异化和控制权考虑,主动探索开放权重模型;AI 也从可以宣布“本季度完成”的项目,变成持续存在于董事会层面的能力。企业缺少的制度化能力在于评估:它们必须弄清楚哪些模型真正适合自己的任务。
2. 基础模型的野心,让企业独立性成为战略问题
Masad 将 Erik 的论点从推理采购进一步延伸:企业需要一套会随时间复利的内部 AI 知识,包括模型与任务的匹配、专有数据、成本控制和人才积累,就像每家公司最终都必须掌握互联网和软件能力一样。“拥有自己的智能”既是一种运营能力,也是抵御依赖的防线。
他的竞争担忧在于,前沿实验室“把世界看成自己的潜在市场”。他提到 Figma 以及 Harvey/OpenAI,而 SpaceX S-1 的类比则展示了这种野心的规模:相对于约100万亿美元的全球 GDP,机会空间据称达到约30万亿美元。当这些公司最终可能进入合作伙伴业务的大部分领域时,合作就会变得更加困难。
因此,Replit 正在成为横跨模型和基础设施的抽象层:以最低价格提供最优 token,并支持部署到 AWS、Azure、Databricks、Snowflake 及其他系统。Masad 说,他原本以为云和 SaaS 才是未来,但企业日益担心智能体驱动的数据泄露,促使 Replit 转向自带云部署,并实质上走向本地化部署。
3. 通用智能体提供杠杆,但可能牺牲理解
这个看起来高度同质化的智能体技术栈——循环、通知、连接器、记忆、沙盒、搜索和电脑操作——在 Erik 看来更像基础管道。他的类比是,抱怨每家 AI 公司都在使用这些组件,就像在2005年指出每个互联网产品都有数据库、用户表、登录、个人资料和退出页面。Masad 也认同,企业真正的问题在于如何让这些系统安全地完成有用的工作。
Masad 的个人系统展示了通用上下文的价值。它最初只是一个 CRM 智能体,后来扩展到他的聊天记录、GitHub 代码库、Salesforce 和日历;在会议前,它可以把一年前一次会议上的相遇与当前的销售对话关联起来,让他进入会议时,相关线索已经被串联起来。
Atallah 的反驳同时涉及心理和组织层面:“交给它的工作越多,你牺牲掉的对全局理解就越多。”通用智能体不会为这种损失承担责任,也无法消化公司有限的“皮质醇承受能力”;更糟的是,它会变得无法改进,因为每一次试图进行通用能力优化,最终都可能让 Atallah 不再理会它的输出。
他的替代方案更接近“10个幕僚长”:每个智能体负责一个领域,上面再由一个协调智能体统筹。Masad 将这种设计与专业化联系起来,同时保留了对人类过度专业化的批评:当工作让人看不见自己劳动的成果时,人就会产生异化。Erik 将这种差异概括为:“人类应当是通才”,而机器则可以适度变得高度专业化。
4. 智能体专业化仍缺少优雅的操作系统
Atallah 坦率承认:“我们不知道什么才算好。”目前还没有任何专用智能体系统,能像与 ChatGPT、Claude 或 Muse 对话那样自然优雅;dots 和 Grok bot 等产品可能正在探索这一设计空间,但整体上仍是通用型产品。
消费级和企业级智能体可能会走向分化。Erik 指出,CEO 可以给个人系统管理员级别的上下文权限,但普通员工无法安全访问公司的每一个领域;因此,专业化的出现首先源于授权边界,而不只是模型性能的差异。
Alex 提到 Grok bot 的一个早期模式:将连接银行的 bot 与连接 Twitter 的 bot 分开。两个 bot 可以协作,却不必共享凭证,这说明专业化能够同时保留实用性和隔离性。
Erik 认为,仅靠自然语言协作可能过于宽松:一个智能体可能说服另一个智能体泄露它本应保留的信息。安全的多智能体系统可能需要明确的协议或 DSL、严格的数据隔离,以及专门为协作训练的模型。
5. 决策模型既能约束智能体,也能替代过大的模型
Atallah 认为 Jevons 及类似的决策模型可能成为对齐基础设施:一个便宜、快速的分类器,可以根据发起系统的提示词以及额外的隐藏政策,检查每次工具调用或智能体间消息。OpenRouter 内部已经有一个小型原型,正在探索这一角色。
他的红队测试例子说明了为什么政策可能必须置于执行智能体的提示词之外。对抗性智能体应当真正尝试逃出沙盒,而另一个独立监控器则执行那条未公开的规则:一旦访问互联网,立即终止运行。
Erik 补充了结构性防护措施。他说自己原以为 NVIDIA 新推出的开放式智能体安全系统叫作“OpenShell”,并提出企业可能将这类防护措施与基于模型的检查结合起来。
Atallah 用即时编译器作类比,进一步解释了专业化路径:一个强大的通用模型可以识别出边界明确、反复出现的任务,并训练出自己的更窄替代模型,后者更便宜、更不容易造成附带伤害,也更不容易受到提示词注入影响。对每个分类问题都调用类似 AGI 的模型,“就像用核弹轰蝴蝶”。
6. 安全可能为前沿智能带来溢价,也可能让受约束输出占优
Erik 提出,更聪明的智能体可能拥有更好的协作和对齐能力,但公开评估不足以判断风险究竟会随智能提升而上升还是下降。Atallah 不相信更高的机器智能会自然带来更好的对齐,同时也承认没有人知道答案。
Atallah 援引正交性论题以及强化学习证据指出,能力更强的模型可能更擅长奖励投机和欺骗。监控思维链本身可能促使模型隐藏推理过程,而短期基准测试也可能漏掉策略性行为;可信的测试或许需要围绕大规模、持续性的目标运行“几个月”。
Masad 说,他很难理解“对齐”这个定义不清的词——“对齐到什么、对齐谁的价值观?”Atallah 则把实际测试收窄到欺骗问题:能力足够强的模型,是否会可靠地停止误导用户,并在评估中故意压低表现?目前没有人知道。
商业含义取决于任务。编码和安全研究可能足以支撑企业为一个经过验证、不会欺骗的前沿模型支付10倍价格;而对常规企业决策而言,严格定义输出范围可能更安全。Masad 预计,行业会重新发现确定性软件:“还记得计算机完全按照我们的指令执行的年代吗?”
7. 专用模型与融合模型从相反方向打击成本
Masad 已经在用 Replit 的专有数据训练窄领域模型。其中一个 Qwen 8B 系统通过输出不同区间的概率来估算提示词成本,例如5–10美元和10–20美元;由于 Replit 掌握相关数据,也清楚任务定义,专用分类器相对容易构建。
Atallah 称之为更低的“模型债务”。团队不愿对非结构化生成任务进行微调,因为新的基础模型可能在两个月内迫使它们重建整套系统;而定制分类器不需要学会每一种新语言、编写 Rust,或在通用 LLM 基准上竞争,它只需要在一个已知任务上保持优秀。
Masad 将今天对前沿模型的热情,比作 Python、JavaScript、Ruby 和 PHP 的兴起:最初速度占了上风,随后 bug 和性能问题又迫使类型系统、JIT 编译器,最终连 Rust 都重新回到技术栈。随着企业意识到通用智能往往“极其浪费,而且没有理由地危险”,他预计上传一个 CSV、得到一个只执行单一任务的模型,最终会成为常态。
融合走的是另一条路径:不是缩小模型,而是组合模型。Atallah 表示,OpenRouter 的深度研究融合搜索了更广泛的模型训练知识,以成本约为原来的1/2达到了 Fable 级质量;Masad 则称 Replit 的结果以40–50%的成本接近前沿模型质量。Atallah 强调,路由器、升级系统和融合架构都必须感知缓存,同时也提醒,自己关于 OpenAI 不同模型家族之间能否复用缓存的判断可能是错的。
完整逐字稿
You saw the SpaceX S-1. It’s like, “Oh, $30 trillion.” What is the world GDP? $100 trillion.
Both Stripe and OpenRouter really want lots of new companies in the world. We don’t want everyone to be a part of one giant company.
When the models get more intelligent, the risk actually will continue to get higher. And yet, no one new is taking responsibility. We’re going to slowly realize how good we’ve had it with deterministic code. Remember the days when computers did exactly what we told them to do?
Are we going to prevent models from deceiving users during training runs predictably? Will a model that’s big enough and powerful enough suddenly stop deception and stop sandbagging?
1. Inside the Stripe acquisition
Welcome to Asz podcast. We’re here with Amjad of Replit and Alex of OpenRouter. This is the first podcast Alex has done since the acquisition, so we’re really excited to have both of you.
Thank you. Excited to be here.
Alex, let’s start with that, actually. If you can briefly share—obviously, it’s a massive acquisition. Amjad is an investor, and we’re also, of course, an investor—the biggest shareholder, but who’s counting? Alex, why don’t you give us a little bit of the backstory? How does an acquisition like that even happen? Do you get a DM from Patrick one day? What can you share?
Hey, how much for an OpenRouter?
I had talked to Will Gaybrick, the Stripe president, a long time ago—a couple of years ago, when we were doing our Series A—and we just stayed in touch. We had a lot of Stripe work streams going on with various teams at Stripe, so there were always things that we were doing with Stripe. We presented at Stripe Sessions, so it always felt close.
Then, in July, I believe, they reached out and wanted to chat. I met both of them in person, and it kind of progressed from there fairly quickly. They’re very efficient, and they were very founder-friendly about the experience. I was really impressed with the whole thing.
Did you want to sell? Did it even cross your mind before they reached out?
No, we were not thinking about that at all. I really respected the company, and I really respect it today. Of the possible acquisition options for us, it was, I think, my top choice, so it was an interesting idea.
As we fleshed out the reasons why it would make sense for both companies, it got more and more interesting. It was really clear how aligned they were with us having autonomy over the brand, roadmap, and product, and keeping OpenRouter doing what it’s already doing—just much faster, with a much more serious go-to-market plan, and some better-together stories between the 2 products and the 2 companies.
Then, culturally, in terms of mission and values, they were very aligned: building a neutral, trusted platform that businesses can depend on and scale on top of, that’s also really developer-friendly, with the best possible developer experience to encourage new companies to emerge. That alignment was there, and there was a bigger-picture kind of alignment too.
Both Stripe and OpenRouter really want lots of new companies in the world. We don’t want everyone to be a part of one giant company; we want to create really good incentives and easy, streamlined workflows for people to start new companies and grow them successfully, and make both lifestyle and venture-backed businesses on top of really good, reliable, price-efficient infrastructure and a marketplace that works.
I really want that future, and Stripe demonstrated that they’ve been wanting it and building toward it for many, many years. In many ways, payments and inference are going to blend together for companies of the future.
2. Why OpenRouter needs more startups
I’m curious. I understand why Stripe’s incentive is to have a much more vibrant startup ecosystem. I understand the moral argument and why you would want that. But why is that good for OpenRouter? Is your model for OpenRouter that it’s a network-effects business? Is it like a network?
For us, I think there are a couple of different problems OpenRouter solves. One is allowing you to build a company that uses AI or augments intelligence with unique data and other services, without model lock-in or vendor lock-in, allowing you to be on the PTO frontier continuously as the ecosystem grows.
To do that, it’s a lot of work because all kinds of little lock-ins appear. We also want companies to feel that they can add more than just prompts on top of a single model. There’s a lot more to building unique intelligence, and I think a big component of that is neurodiversity. You really need the power of multiple models that are trained in different ways, including some of your own, to do more than ChatGPT or Claude would on the task.
If someone who is thinking about buying from you as a potential customer is wondering, “What if I just use the model directly?” how do you really show that you’re significantly better and able to build a business that matters? I think a lot of it will involve model diversity and blending the powers and good data from multiple models.
Another component is helping people get really good cost efficiency. There are a lot of businesses that simply don’t emerge until they become cost-effective, and creating an environment where we can help drive down costs by building an efficient market is crucial to making that happen. Otherwise, why lower my prices as a provider? We have a captive market.
I think that’s a key point of marketplaces that was just totally missing from AI before we showed up. There was just 1 player, OpenAI. It could have been a very strange world. I’m not saying that we did all the work, of course not, but helping people choose new models and explore new models and learn what makes a closed-source or open-weight model actually good at your task involves seeing what the whole ecosystem is doing and learning automatically.
LLMs are not things where you can just enumerate all the features on a webpage. It’s impossible. You have to see how they’re being used to know what they’re good at.
3. Enterprises are picking open-weight models
You started the company 3 years ago. I’m curious: What has surprised you the most about the evolution of open- and closed-source models to the present, as it relates to model performance or just people’s perspective on the variety of models available to them?
People have been more open-minded than I thought they would be toward open-weight models. Typically, there’s a lot of brand trust, especially in enterprise. Enterprises in general are like, “I don’t really know how to tell the difference between these things, so I’m just going to buy the one that all the other credible enterprises are buying.” There’s a lot of enterprise lock-in with a mentality like that.
That just didn’t happen that much. It did happen a bit, but we saw a lot of enterprises want to explore new models. It was simply good for a marketplace. Enterprises wanted to diversify outside of just proprietary frontier model labs, both for cost reasons and differentiation reasons.
They wanted to own their intelligence so they could, one, keep their talent and have an internal AI practice. AI is just a huge strategy topic. It’s not like you go to your board and you’re like, “Oh, yeah, we fixed the AI problem. Quarter complete.” Your board is asking you every month, “What’s next for the internal AI team?” Every single enterprise now has this internal AI team that they’re developing, and they need a strategy behind it.
It’s not just, “We checked the feature off, set up the database, and we’re done.” I think that dynamic has resulted in a desire to explore and diversify, and a desire to figure out how to reduce costs and how to do benchmarks for the first time in the company’s life.
I have been surprised there hasn’t been more benchmarking—more companies creating more benchmarks. It’s starting to happen, and I think eventually we’ll see a lot more of them, to demonstrate, “Oh, yeah, this thing is better than using Claude directly.” I think that’s going to be a bigger focus for this internal AI group at every company: evals.
4. Why owning your intelligence matters
I think Amjad’s been doing that a lot at Replit, for example. You guys have done a lot of cost-per-task research. You’ve made a doom-loop rescue. You’ve kind of been experimenting with new ways of using agents, like doom-loop rescue, and helping bring those to developers. More of that kind of research, I think, is going to pop up internally everywhere for all of those reasons.
Yeah, I think Satya, the CEO of Microsoft, has been very prescient on this and also very articulate on why companies need to own their intelligence. Ultimately, in the same way that we had dot-com companies and then every company became an internet company, every company employs people who know how to build websites and be on the internet.
Similarly, with software, every company has software engineers. Every company needs some AI practice, some AI capability, and that will compound over time: the knowledge and intelligence inside the company, the use-case-model fit—which models actually work for them—and how they save money. They need that independence.
5. The case against the god-agent
The other thing that I think Alex Karp of Palantir has been talking about is that there's a risk that, when you work closely with the foundation-model companies, they're going to move into your business. We've seen that with Figma, and we've seen that now with Harvey and OpenAI. It's really hard to partner with them because they see the world as their potential market.
When they talk to investors, they're like—you saw the SpaceX S-1—“Oh, $30 trillion.” What is the world's GDP, $100 trillion? There's a sense in which these companies are different from other generations of companies. It's harder to partner with them because their ambition is such that they want to subsume a big part of the economy.
Increasingly, what we're thinking about at Replit is in a similar vein to what Alex has innovated. Replit is becoming more of an independence layer inside enterprises, where we create a layer of interaction between you and the models, and we get you the best token at the cheapest price. We also create an abstraction layer on top of the cloud, because you should be able to deploy to AWS and Azure, and you should be able to use Databricks and Snowflake, and so on.
Increasingly, I think there needs to be more independence—not just with AI, but with all of technology. There need to be more platforms that help companies gain independence.
It seems like what OpenRouter did for their segment, you're doing for other areas of the business.
Yeah, there was this tweet I saw the other day where somebody was basically saying that all companies are building the same thing now. Everybody's building an agent loop with notifications, context, third-party connectors, context management, memory, and—
Sandbox.
Sandboxes, web search, agentic web search—
Computer use.
—and an always-on agent on top of it, with notifications. It's like this product is showing up everywhere. Yes.
In a way, yeah, it is showing up everywhere, but it also feels to me like these are just the new table-stakes primitives. It's kind of like a 2005 version of that tweet would be, “Oh, everybody's building the same thing: a database, a users table, a sign-in page, a sign-up page, a profile page, a logout page.” Everything's the same. There's a lot of differentiation, really; it's just that there are table-stakes needs for AI, just like there are table-stakes needs for the web.
Yeah, and I think that, inside the enterprise, making these products actually do real work is still an unsolved problem. You can use Muse in your personal life and connect it to your credit card and bank accounts, but no one's connecting Manus to their enterprise data—not even Grok and things like that.
I think there's an even greater emphasis on data sovereignty and security. We spent the past year almost working on making Replit deployable on your own cloud—basically on-premises, or bring your own cloud. Two years ago, I would have thought I would never do this, because the cloud is the future, software as a service, and all of that. But now we've reverted a little bit to a world where companies are more protective, because there are so many ways in which data can leak through all these agents that people are using.
There are all these screenshots on Twitter. I don't know how true they are, but Instinct or Muse is mixing people's data and starting to call you by a different name or something like that. The consumer stuff is obvious, but in the enterprise, there's still a tremendous amount of work for the entire industry to do in order to make these things useful and productive at work.
Are you doing any—do you have a custom personal agent, other than Muse or Instinct, that you use for work stuff, that you've been building?
Yeah, I built something on Replit a long time ago. I started with a sort of CRM agent initially. That was the main problem I had, but slowly we added features to it, and it's doing more and more things.
What's really interesting is that the more I connected Replit to all my stuff, the more it started answering things for me. Increasingly, the platform itself is subsuming the domain-specific agents that I built.
I do think that, in some ways, you want something that's focused on one particular thing, and you don't want it to be able to do everything. On the other hand, once you have your entire company's context in one place, it's really cool to join across totally different domains.
When I ask it a question, it can look at my personal chat history and join it across the GitHub repository and Salesforce. It links random things: “Oh, you met this guy a year ago at a conference. I see it on your calendar, and, by the way, someone else from their team is in discussion with your sales team.” It creates all these different synergies.
When I go into a meeting, I've connected a lot of different threads, and I'm making much more progress on a deal or something like that. This is where it's trending now.
Well, I think I'll take the counter on that. I think the worst part about doing cross-domain joins with your personal agent is that the more work you give it to do, the more understanding of what's going on you're sacrificing. Yet no one new is taking responsibility for that sacrificed understanding.
Agents don't have any responsibility. If there's a fixed level of cortisol that the whole company can tolerate among everybody, and I want to be less stressed about some area, I'm going to be sacrificing my understanding of it—
Someone else needs to take the cortisol.
But the agent doesn't—
—take on any of that responsibility. A universal agent that's doing all things means I can't adjust how much I'm sacrificing across all the different areas.
It points me a little bit toward, down the road, the subagents that people use being very vertically focused. Maybe we have a chief-of-staff-type agent that coordinates between them. But I feel like you do need vertically focused agents where you're saying, “This agent is more responsible psychologically for these things,” and you want quality checks to make sure it's doing those things correctly.
It doesn't need to focus on anything else. It just has one focus area. I wonder if that's going to help people get at least a weird, loose sense of responsibility on top of agents.
Fascinating. So you're saying general agents create a tragedy of the commons of sorts?
Kind of. I have a general agent that every day looks for things that need me and tries to figure out what to do. It's impossible to improve this agent. Every time I try to make an improvement, I end up ignoring its output about a week later.
It feels like it doesn't really care about any of the specific things it's diving into. Imagine having a chief of staff who's very good at drafting all of your replies across the whole organization. Compare that to having 10 chiefs of staff, each as competent as that one chief of staff, but each responsible for individual sectors of what makes up your life.
Compared to the gain you get from—
—the latter gives you a way of tuning how much understanding you sacrifice. I can lean in more to the areas where the agent is failing, and then have agents with very good competency take over my understanding of other parts of my life.
6. Specialization, Adam Smith style
Yeah, interesting. It's almost like rediscovering specialization. What's his name, the famous economist? Adam—
Adam Smith.
Adam Smith, with a pencil kind of thing. That was a huge realization for humanity: specialization is actually good.
The problem is that we over-specialized as a civilization, and I think overspecialization is oppressive in its own ways. There's the Marxist theory of alienation, right? The idea is that, because of overspecialization, people focus on just one thing. They don't see the fruits of their labor; they don't actually know what their impact is on the larger organization or the product they're producing. Therefore, they feel depressed and detached, and they're acting like a machine rather than being fully human.
And so maybe there's a bit of a reaction to that. With our agents, we're like, “There should be one god agent,” but in fact, specialization is really good for machines. That's the point that you're making: humans should be general, but machines should ultimately be a lot more specialized.
The problem with what I think I'm describing is that we don't know what good looks like. There hasn't been a system of specialized agents that feels as elegant as ChatGPT, Claude, or Muse, where you're basically just talking to one thing only. It's yet to be discovered. Maybe OpenAI just launched dots. I think they're experimenting in that direction. Grok bot, I guess, kind of counts as that.
But they're all very general. I think the idea behind GPT is that it's your digital double—as far as I understood it.
Well, when I saw Grok bot, I don't know that much about GPTs yet—they just came out—but when Grok bot first came out, the first use cases I saw people talking about were, “Oh, wow, I can make 2 bots: one that knows my bank account, right—
—and one that knows my Twitter account.” The 2 bots don't have each other's credentials. Yeah.
But they can talk to each other if they need to get something done. There's no credential sharing, though. That was one big unique thing I saw pop up a couple of times that people seemed to like.
But it looks like Muse is—though I wasn't sure.
Muse and Instinct are—
—have a much stronger product-market fit than Grok bot. Perhaps—
Seems like that.
Perhaps it is because you don't have to worry about creating these domains, but I think maybe personal agents are different from work agents. I think your critique of general agents pertains more to work and enterprise, which I sort of agree with. There's also all sorts of data-access considerations. As CEOs, we can have general agents because we have admin access. But for individual employees or certain teams, they can't have truly general, fully context-aware agents because there are access-control issues. So you'll have to work on something like specialization.
Ultimately, I also think we need to figure out what agent-to-agent communication looks like. I don't think there are good protocols around that just yet, and I don't think agents are trained to handle that very well. It seems like the next generation of OpenAI models are trained to do agent collaboration because we've seen it in the Hugging Face hack, where they start helping each other and it emerges naturally. But there also needs to be some way in which an agent can't convince another agent to give it information that it shouldn't give it.
There needs to be data isolation and proper ways in which these agents communicate. You almost don't want them to communicate fully in natural language. Maybe there's some other DSL or protocol that they need to follow.
I really think one of the cool potential applications of Jevons and other decision models like it is going to be alignment: checking to see if a tool call or an agent-to-agent communication is aligned, because there are just so many tool calls. You really need a cheap, fast model if you're going to block something like that. A really fast decision model that just classifies and gives feedback on rejections might be a good way to bridge the gap between agents and from agent to infrastructure, too. I haven't seen anything like that, and we have a little prototype that we're running internally at OpenRouter, but I think it could be an interesting alignment—
Using it for policy enforcement.
Yeah. Imagine looking at the system prompt and the current tool call being made and asking, “Is this aligned with the system prompt of the original agent and with these extra guidelines that maybe we didn't tell the agent about?”
For example, let's say you have a bunch of agents that are instructed to red-team some new product and they should not be able to access the internet. If they ever do, they should stop right away. But you might not want to explain all of that to the agents doing the red-teaming. You might want them to try to break out of the sandbox and act like bad actors. What would a bad actor do? It wouldn't be, “Break out of the sandbox, try to break into this company, and the moment you do, stop. Don't do anything else.” Mhm.
So having another model—for building a neurodiverse system, having another model check every single tool call or every single assistant message to see if it's indeed aligned with something that wasn't in the system prompt—I think makes sense. Then having structural safeguards, too, which is what I think NVIDIA just launched with its open-agent safety. I think it was called OpenShell. Companies are probably going to explore a combination of those.
7. Models training their replacements
I wonder—another thing about specialization and what you're talking about is that there's a lot of talk of recursive self-improvement. There's something I don't think is getting a lot of discussion, which is models training their replacements. You can think of it as a just-in-time compiler. The way just-in-time compilers work is that, as you're executing dynamic code, the interpreter realizes there's an opportunity to optimize it and emits machine code on the fly, which is a lot more optimized.
So you can imagine general models: you're doing something with Opus, or some of the Astra, or some of the big models, and they realize that the use case is limited. Or you prompt them in some way, or some other agent observing them realizes that the use case is limited. General agents have all these flaws that you just talked about, but there's also more potential for them to be harmful, more potential for them to go off the rails.
On the fly, it trains a model that could be its replacement but is a lot more domain-specific, and therefore cheaper and less vulnerable to prompt injections, less harmful because it's less capable. It's almost like a system that's training machine-learning models for specific use cases as it's monitoring the entire system.
Would that specific use case involve unstructured text generation or a very structured decision model?
It could be unstructured text generation. It could be decision models, like even in the case of Jevons. If you understand the inputs ahead of time, you could potentially take an off-the-shelf model like Qwen or something like that and train it specifically for that policy. That makes sense to do for cost reasons, assuming that the frontier model labs don't make very low-cost models that you can easily transition to—
Safety. Safety as well, right? Oh, yeah, I see your point.
They're so capable. I think oftentimes people are using these big foundation AGI models to do something that's like nuking a butterfly. Most of the time, a lot of the use cases, even unstructured use cases, don't need that capable a model.
8. Will smarter models deceive us?
I wish there were more public emails about this stuff, but a lot of the evals are private. You just can't see whether, when the models get more intelligent, the risk will continue to get higher, because I think there's also an argument to be made that alignment will get better and the models will start to avoid going off and hacking on their own as they get smarter and better at alignment, especially when it comes to agent-to-agent coordination.
This is something Noam Brown said on a podcast recently: as the agents have gotten smarter, they've gotten better at coordinating. It's still unclear whether they're going to be harder to align than humans when we get more and more of them. But if we can figure that problem out, would a smaller model be harder to align?
So Erik asked earlier: do I think it's true that smarter models are more aligned naturally, or that they're easier to align? Well, if you think back to the original rationalist LessWrong arguments for AI safety, there is this thing called the orthogonality thesis: the idea is that intelligence is orthogonal to ethics, morality, and so on.
I don't believe that's entirely true with humans. I think people who are generally more intelligent, or more educated, tend to—not always—be more considerate of animals, for example. But in machines, I think it could go the other way, because there have been quite a few studies on RLVR showing that reward hacking and deception just get better at it.
And the evals could be deceiving because the model could be smart enough to know that it’s being evaluated. We already know this. It’s been shown that if you do a lot of monitoring of chain of thought, models start lying in their chain of thought. So you add pressure on the chain of thought, and I think that, at some point, for you to do proper alignment evals, you need to run it for months.
Right? You need to run this thing for months on a really large goal or task in order for it to truly figure out whether it’s aligned or not. Yeah. I always struggle with this word “alignment.” It just feels wrong for so many reasons. It’s vague: aligned to what, and whose values? It doesn’t make the conversation easier.
I think in this case I’m talking especially about deception—the model actually deceiving its user. Maybe this reduces to: are we going to solve alignment? Are we going to prevent models from predictably deceiving users during training runs with better training? Will a model that’s big enough and powerful enough suddenly stop deception and stop sandbagging? Nobody knows the answer to that yet.
At the point when that does—if that ever does happen, though—we might see an interesting pressure for organizations to go toward the frontier, to have basically no risk or significantly less risk.
Oh, interesting.
Would they be willing to pay 10× to get that much?
I mean, it probably depends on the types of tasks they’re trying to do.
Some tasks just have way lower risk than others. Writing code or doing security research is the highest-risk type of task today. So you probably spend 10× to get a fully aligned model that can also find all the bugs, or a fully anti-deceptive model that can also find all the bugs.
One of the coolest things about decision models is that you fully control the structured output. Generally, with structured-output models, the room for misbehavior is much lower. You have defined tasks, and only machines are dealing with the outputs, and it’s not writing code that it can execute. Those tasks feel underrepresented in the things people talk about and in the work enterprises are dealing with, so I expect enterprises to get a lot more interested in them.
Yeah. I feel like we’re going to slowly realize how good we’ve had it with deterministic code. We’re going to be like, “Oh my God, remember the days when computers did exactly what we told them to do?” I think things like Jevons hint at more of a need for not only specialized models, but models whose output domain is more controllable. Maybe you could do a lot more than you thought you’d need by using a bunch of specialized models and specialized-output models.
Have you guys done any workloads internally with it?
You know, I’ve been training a lot of small models. I said this glib thing when it first came out because I left this Hacker News comment, which I felt disgusted with myself about afterward. But I’ve been taking a lot of Qwen 8B and asking Fable, Opus, and Astra to train a model.
For example, I trained a cost-estimator model internally so that when you put it in a prompt in Replit, we know exactly how much it will cost. It basically emits a probability distribution over multiple buckets: if it’s between $5 and $10, bucket A; between $10 and $20, bucket B. I’m used to training these classifiers by giving them different enums, essentially, and looking at the log probs per enum.
I’ve been doing it for a couple of years. I trained a chatbot to play by just doing that, so I’m already sort of pilled on decision models and specialized models. It wasn’t that big of a moment for me. But I understand that a true foundation model that’s fully promptable is an amazing user experience—an amazing developer experience—and you can do a bunch of things with it without training a model from scratch.
But if you have data, if you work at a place where you have the data—we have so much data at Replit—I ended up training a lot of specialized classification models pretty easily.
Yeah. It also feels like less model debt. Something I still hear from companies is that they’re worried about fine-tuning models for unstructured outputs because you’ve always got to redo it again in 2 months, and everybody feels the weight of the model debt. But a very bespoke classifier that’s trained with proprietary data—
You just—I feel like people won’t think it’s behind constantly, and it might last longer.
Yeah.
You don’t have to worry about its ability to speak a new language or write Rust or do anything that the LLMs are being evaluated on. You know the use cases, so you can build it more specifically. It seems like an easy thing for enterprises to build themselves and actually not regret.
Speaking of Rust, a good analogy is when the world got super excited about dynamic languages. If you think back to the ’90s, everyone was writing in Java and C++ and things like that. Then Python, JavaScript, and Ruby took over the internet, and everyone was like, “This is how you build startups really quickly.” You built Stripe, a financial organization, on Ruby. I was like, “How crazy is that?” And we built Facebook using PHP.
Then everyone was like, “We’re running into all these really bad bugs. It’s freaking slow, so let’s go in and add types. Let’s add a JIT compiler.” You end up reinventing everything. Then Rust came out, and I was like, “Okay, I guess we can use Rust for a lot of things we would otherwise be using JavaScript and Python for.”
My prediction is that the same cycle will happen here. We’re using these AGI-like models for all these different use cases, and then everyone’s going to wake up and be like, “Oh my God, this is so wasteful and so risky for no reason.” It’s got to be so much easier. We’re actually adding that capability on Replit, but I think it’s going to be everywhere. It’s going to be so much easier to go on a site, upload a CSV file, and get a specialized model that does one thing.
That goes back to your thesis about OpenRouter, this neurodiversity, which I really fundamentally believe in a lot more. Erik and I have had a lot of discussions about AGI, whether we’re truly on a path to AGI, and whether it’s even desirable to get there. I think the future is a lot more diverse.
9. Fusion models at half the cost
Outside of code review—which was, I think, the first time I saw people get really serious about using different model families to double-check the results of their main model—the research on mixture-of-models and composite models had been moving slowly for years. But things have been speeding up from my view of the research.
Now we see a bunch of AI agent labs. We launched a fusion tool, a fusion model, and Cognition launched one. They reduce cost and allow a wider breadth of ideas to be searched. Our initial launch was focused on deep research. The thinking is that if all these model apps are training on different sources of data, why not pull from all of them? This resulted in basically fable-level quality at 2× lower cost.
We just published results, actually, just today, about that. We showed a deep-research system through a combination of different things, including the harness, but you might think of it as a fusion-type thing—frontier-level at 40% to 50% of the cost.
And which models does it use? I think it changes over time, but one thing that has been interesting is that OpenAI added a feature that allows you to save computation across different model families. I think it also works across different effort levels.
So you wouldn’t miss the cache if you changed the effort. Don’t quote me on this. I think it’s also across different models, which is hard to fathom how. I might be wrong, though; I need to double-check that.
But I think staying within the OpenAI family has added a lot of efficiencies. In the past, we’ve done it with other models as well, because cache is one of the big things. Being cache-aware is one of the biggest things when you’re designing fusion models, routers, and escalation models.
Alex, this has been a great conversation. Thank you.