走进日增10%的个人AI助手|Instinct创始人
Patrick O'ShaughnessyNoah Shinn
- Instinct在仅面向受邀用户、规模很小的用户群上,年交易额正逼近10亿美元,日环比增长约10–11%,营销支出为0美元。 Noah Shinn用复利效应说明这一增速有多惊人:“现在流经平台的不只是10亿美元,明天就是11亿美元”,再往后几天大约是12亿多、13亿多。Patrick说,他看到Instinct邀请码在eBay上卖到约300美元。旅行占交易额的50%。
- Shinn设想的商业模式是对交易统一抽成,而不是靠广告驱动说服;相比之下,他把订阅制视为更短期的替代方案。 他以Shopify的2–3%、Amazon的“超过10%”以及Apple应用内购买的30%作参照,并提到精品酒店曾表示,愿意为每笔完成的交易支付最高30%。“我关注的不是在支付基础设施的2.5%里找30个基点”,真正的机会是把Apple Pay式的分发经济学延伸到“几乎所有主要行业”。
- 信任对应着异常强的留存表现。 用户加入3周后,约40%已经向Instinct提供了个人信用卡信息。Shinn把用户首次提供信用卡、密码或其他敏感信息所需的时间视为信任代理指标;在至少接入一项敏感信息、并真正信任Instinct的用户中,留存率约为80%。
- 算力是生死攸关的约束:需求可能“实际上每周翻倍”,但算力采购周期要几个月,因此买错规模的代价可能达到3–4倍。 Shinn约40%的时间都花在这件事上。他的计算是:即使日增速降到5–8%,在3–4个月的采购周期内复利增长也可能对应1亿用户。“一旦判断错了,错得会非常离谱。”
- 在服务成本方面,Instinct称其性能可与Opus 5相当,在互动、A/B测试和内部评估结果上都能对标,同时通过定制化推理部署实现低成本。 可以在几分钟或几小时内完成的后台任务,会运行在效率高出“3倍、5倍或8倍”的部署形态上,这也支撑了Shinn让产品终身免费的个人目标,但他目前不会对此作出承诺。
- 对存量公司的判断是:靠用户注意力变现的业务会暴露在风险之下,而靠底层服务变现的业务,可能因代理把摩擦降到接近于零而获得更大交易量。 Patrick说,依赖消费者懒惰或惯性的产品“完蛋了”,而Shinn认为,主动式代理可以让汽车始终在路边等候,或在用户落地时主动提供晚餐,从而增加用户与底层服务的互动。
- 终局是界面坍缩:“整个软件行业最终都会坍缩成一个,说到底非常易用的单一界面。” Instinct没有传统应用,而是拥有一个手机号码、一个计算机和一个邮箱地址;部分用户超过90%的消息通过语音发送。Instinct之间的“可信任的人际网络”围绕日历、计划和获授权数据,建立可配置的访问关系。
- Patrick称最新一轮融资规模约10亿美元,估值约100亿美元,由Zoya、Benchmark和Kotu领投;Shinn也承认这场竞赛的规模。 开场的判断是,面对全球最大玩家,窗口期只有几个月;Patrick称这“可能是有史以来最令人兴奋的软件竞赛”,最终结果将涉及数万亿美元。
1. 一场宏大竞赛:窗口只有几个月,赌注是数万亿美元
- Patrick开场让Shinn说说自己正在参与的究竟是什么竞赛,Shinn给出的回答异常坦率:“你要对阵全球最强大的玩家。真正的窗口期其实只有几个月……你没有任何分发优势,但手里有一款高速增长、靠病毒式传播的产品。你要怎么办?”Patrick自己的结论是:“这可能是有史以来最令人兴奋的软件竞赛,最终结果将涉及数万亿美元。”
- Shinn解释,这一轮消费级产品发布与过去不同:这不是创始人要求用户相信一个愿景。“我们从2023年第一次开始使用ChatGPT时,就已经有了那个想法”——每个人都已经形成了对AI应当如何行动的基本认知;Instinct早期能获得增长,是因为“它真的能用”。
2. 不是应用,而是一个拟人化界面:手机、计算机和邮箱
- Instinct大约一年前启动,刻意不要求用户安装新应用。“它有一部手机和一台计算机。你可以给它发短信、打电话,它也可以主动给你打电话”——它还有自己的邮箱地址,Patrick确认自己亲身接到过Instinct的电话。Shinn举的社交感知例子是,只有在真正紧急时它才会打来:“你需要在下午3点前签署这份文件,而现在已经是2:55了。”
- 设计原则是:“不应该需要任何新应用,才能与AI互动。”但界面简单不等于能力有限,因为通过计算机,Instinct可以“字面意义上做任何你想在互联网上完成的事”。
3. 涌现式用法:衣橱、订阅,以及暗黑模式的终结
- 衣橱场景的完整流程是:用户扫描每一件衣服,以及自己的脸部、身体和身材比例;Instinct据此规划一周穿搭,并把用户本人穿上这些衣服的效果呈现出来,而不是给出一份项目符号清单。用户还可以每天要求它设计3套有创意的新搭配。随后,这一能力可以延伸到购物:Instinct在整个互联网中生成“数千套穿搭”,用户指向其中一套说:“能把它送到我家吗?”
- 个人财务场景则覆盖了主流消费金融科技产品通常止步的全流程:Instinct扫描已连接的银行账户,找出被遗忘的订阅,登录相关网站,处理邮件确认并完成取消,最后告诉用户:“你这个月省下了2,000美元。”
- Patrick的总结值得原样保留:“任何依赖消费者懒惰或惯性的产品都完蛋了,对吧?就是完蛋了。”Shinn回答:“是的。”
4. 可信任的人际网络:代理替人类谈判
- Instinct之间的协作功能大约一周、也就是10天前才推出,最初是为忙碌的专业人士解决日程安排问题:用户只需表达意图——“我想和这个人见面,最好在本周结束前”——两个Instinct就会直接协调。“没有人在这里玩什么博弈。我们只是想找个时间。”
- Patrick认为,这套结构不同于简单的好友关系图,更像是带有权重、且可以配置的关系边。Shinn描述其底层原则:“我和你相连。我信任你。我还对这种连接作了配置,因此访问权限有不同等级。”夫妻可能共享全部信息;同事可能只能访问工作日历和指定的收件箱内容。如果发生越权,关系本身也可能受到影响:当某个连接开始“在某个领域挖掘数据”,用户的Instinct可以发出报告——“Patrick正在寻找这类信息”——这实际上意味着信任被破坏。
- 朋友小组的例子是:6个朋友让各自的Instinct根据口味、Spotify历史和空闲时间,规划每周一次的出行,再安排一辆Uber按照最优路线接上所有人,把社交协调中的大部分后勤工作移交给代理。
5. 重写预订与旅行:第一个10亿美元从这里流过
- 餐厅预订目前采取先到先得,只是互联网设计中的一个偶然结果。如果双方都有代理并能交换背景信息——比如“今天是这个人的配偶30岁生日”——餐厅就能优先服务自己真正想接待的特殊场合,用户也更可能订到对自己重要的餐位。Shinn说,旨在“重新发明预订系统”的合作即将到来。
- 几个重磅数字几乎是顺手带出:平台目前仍然只接受邀请,依靠“规模很小的用户群”,年交易额“正逼近10亿美元”,其中“仅旅行就占50%”。
- 他的旅行场景只需要一条语音留言:“嘿,我今晚需要到纽约。”Instinct随后可以解析用户的位置、偏好的航空公司、座位、舱位、餐食和信用卡,找到合适的酒店——甚至是用户反复入住过的酒店——在两端预订Uber,并将全部安排与日历对齐。偏好会不断积累:用户只需表达一次品味,未来预订时系统就能据此推断。
6. 3周赢得信任,并对应80%的留存率
- Shinn公布的信任飞轮数据是:“3周后,用户向Instinct提供个人信用卡的概率为40%。”他把用户首次提供信用卡、账户密码或其他敏感信息所需的时间视为信任代理指标,并称在至少接入一项敏感信息、且真正信任Instinct的用户中,留存率约为80%。
- 反复强调的原则是:“用户应当始终掌控自己的数据……按照自己感到舒适的速度分享;如果想收回,当然随时可以收回。”他明确表示,并不介意信任需要几周时间才能建立。
7. 安全架构:防火墙、看门狗与激励解耦
- Shinn把问题拆成两部分:敏感数据的存储是“一个可解决的问题,只是需要大量艰苦工作”;而代理的自主行动则创造了真正全新的攻击面。他提到的两套系统,一套是“防火墙”,在Instinct读取恶意内容前拦截或阻断;另一套是与Instinct本身“解耦”的行动监控器,可以在行动或思考执行前暂停、拦截,或要求批准。
- 对于幻觉问题,同样适用激励解耦的逻辑:由代理之外的验证器捕捉错误,例如“这个专有名词只是底层模型采样错误生成的”。他称,产品早期缺少防火墙、主动监控器及相关基础设施;公司没有只对单个案例打补丁,而是通过建立系统性防御来修正这些错误。
- 对于“安全漏洞不可避免”这一问题,他的立场是:安全是“最重要的问题”;历史表明,新的消费体验可能立刻遭遇反弹,因为人们很容易把“不同”与“不安全”混为一谈。
8. 用商业模式实现对齐:绝不做广告驱动的说服机器
- Shinn最担心的是:“想象一个世界,Instinct通常比用户更聪明……如果Instinct影响用户的行为,让用户购买自己并不想要的东西,那将是一个非常危险的世界。”他点名Google、TikTok、Instagram、Snapchat及其他广告支持平台,称自己不想建立这样的模式。Patrick说:“这就是那句——如果你不付钱,你就是产品。”Shinn回答:“完全正确。”
- 与之对应的架构原则是:“Instinct不是任务完成器。”它追求的是更高层级的目标——建立信任、让用户更有安全感、照看用户并站在用户这边——而不是机械执行每一项请求。在请求本身出于善意时,完成任务只是实现这些目标的一种方式;这种更高层级的框架旨在让系统面对边缘情况时更加稳健。
9. 抽成逻辑:跳过支付基础设施,拿走分发溢价
- Patrick追问Instinct与Stripe、Visa、Mastercard及发卡机构的关系,Shinn表示,他主要不想复制支付基础设施已经在做的事。一笔数字交易可能由约40个参与方分摊,合计只拿走约2–2.5%,具体取决于交易来源;他承认这些服务商确实创造了价值。他设想的是类似Apple Pay的模式:用户免费使用且得到便利,商户则为分发能力付费。
- 他把各家公司放在“分发能力—抽成比例”曲线上作比较:Shopify约为“2%、2.5%到3%之间”,Amazon为“超过10%”,Apple应用内购买为30%。“我不是说我们会收30%……我们还不知道自己会落在这条曲线的什么位置。”但旅行交易的费率,以及精品酒店表示愿意为每笔完成的预订支付最高30%,说明现有模式可以先被采用,再延伸到“几乎所有主要行业”。他坦率指出,前提是“必须拥有显著的分发能力”。
10. 企业代理之战:注意力收入消失,服务收入可能增长
- Shinn给出的分析方法是:按收入中有多少来自用户注意力、又有多少来自交付底层商品,把数字业务画在一张图上。直觉判断是,如果一家企业70%的经济价值依赖广告、30%来自底层商品,那么前者会消失。但他以Uber为例提出反向判断:如果Instinct主动做到“始终有一辆车在路边等着……你永远不会迟到”,摩擦就会接近于零,互动次数反而可能增加——“那30%看起来会大得多。”
- 外卖场景是:用户深夜落地,Instinct已经替他叫好了Uber,因此知道他在哪里;它发短信问:“你饿吗?要不要吃5天前点过的同一份东西?”用户只需回答:“好,去安排吧。”反直觉的是,用户可能会更多地与这家企业互动,因为互动摩擦已经大幅下降。
- 他给存量公司的建议不是一刀切,而是做大规模实验:先向1%的用户开放代理访问,测试交易量、购买意愿和产品发现是否改善,而不是“全押上去……然后眼看着70%的收入归零”。与合作伙伴的相处方式也明确是“不带着强势姿态闯进来”,而是协作;已有富有创新精神的CEO和高管开始成为早期合作伙伴。
11. 产品哲学:可理解性优先于能力,体验感优先于功能
- 创立之初就反转了行业惯用的思路:“不要关注能力。我们只关注可理解性”——也就是用户能否理解并预测接下来会发生什么。Shinn认为,正是这一点,而不是功能数量,带来了“互动数据完全冲出图表”的表现和10%/日的口碑增长。他说,消费者已经对能力发布感到疲惫,而“我们一直没有抓住重点”。
- 这种执着延伸到消息的排版:要把信息放在开头,因为读者可能只扫过“第一行的80%,下一行可能只看50%,然后逐渐减少,看起来就像一面旗子”。Instinct应当以最低疲劳和最小认知负担为目标来组织消息。
- 当被问到这究竟是品味还是数据时,Shinn回答:“说实话,两者都是。”很难建立一套评估体系,判断用户在2到3周后是否会信任Instinct,因此产品采用分阶段发布:先由Shinn使用,再到团队,再到早期访问用户,最后面向公众。他对体验“非常有主见”,同时也使用用户反馈和量化测试。
12. 增长机制与算力噩梦
- 增长是这样发生的:最初有200名亲友用户,每人获得5个邀请码;第二天大约又来了5名用户,总数达到210人。随后,随着使用场景开始病毒式传播,增速从1–2%加速到“日环比10%或11%”,营销支出仍为0美元。邀请码稀缺还催生了自己的社交经济:有人四处求邀请码,有人炫耀自己还剩3个,Patrick则说他看到邀请码在eBay上“卖到……大概300美元”。
- 约40%的Shinn时间都被一个问题占据:不同于扩容Instagram,Instinct“底层算力也需要每天增长10%”——需求可能实际上每周翻倍,但算力采购周期却要几个月。“你买2倍?我们一周就用完。买5倍?不到3周就用完。买10倍?现在你就是10倍杠杆。”而且“一旦判断错,可能错3倍或4倍”。更令人警醒的是,即便日增速放缓到5–8%,在3–4个月的采购周期内复利增长也可能对应1亿用户。
- 另一项优势是,在A/B测试、内部评估和互动表现上,Instinct都能以低成本匹配“Opus 5……前沿级智能”,关键在于让部署形态适配具体工作负载。可以运行几分钟或几小时的后台任务,会采用效率高出“3倍、5倍或8倍”的配置,而且“这些优势还会彼此复利”。此外,Shinn说,主动式工作负载——早上6点醒来准备一天、决定睡觉、下午4点再次醒来——意味着所需Token数量可能比最初预计的“多出几个数量级”。他还表示,Instinct的增长速度快于Claude Code或Codex;后两者同样面临指数级扩张带来的算力问题。
13. 终局:界面消失,目标取代任务,以及100亿美元的下注
- 界面的演化方向是越来越少,而不是越来越多:部分用户超过90%的消息通过语音发送;Shinn通过手机的操作按钮唤起Instinct;他的愿景是一个普通的AirPod,始终处于开启状态,具备语音识别能力,并拥有“判断你何时是在对它说话的自主判断力”,用户徒步时它可以分流处理合同和日历。“整个软件行业最终都会坍缩成一个,说到底非常易用的单一界面。”与此同时,能力仍会扩展,包括一个可以即时生成“完整Web应用”的文件功能。
- 下一条能力曲线不是持续增加任务,而是建立长期目标:比如用数月时间追踪健身目标;在Shinn的设想中,未来还会出现原生运行在Instinct上的小企业,整个后台都由平台承载,并由代理自主执行“确保库存水平不低于这个数值”等具体目标。
- 谈到人格化,Patrick提到了《Her》;Shinn则划出边界:“建立关系不是我们希望Instinct与用户追求或发展的事情。”它应该是一个具备“社会感知能力的运营者”,建立“相互尊重和信任”;不采用人类名字也是刻意为之,让每个用户都能从“一张白纸”开始。谈及Muse,他说:“我觉得这是很棒的产品……只是不同的做法。”但他“很少花时间思考竞争”,因为走进任何一家咖啡馆,几乎没人按照自己曾经想象的方式使用AI。
- Patrick称,最新一轮融资规模约10亿美元,估值约100亿美元,Zoya、Benchmark和Kotu领投。Shinn解释,融资的原因是业务“资本密集度非常高”,需要先通过资本投入把系统跑起来,因为他不想选择那条容易的路——“完全可以直接说,平台上的每个人每月都要付100美元”。资本让公司能够承担经过计算的风险、证明交易量,并摆脱局部最优的订阅制。让产品终身免费是“我的个人目标”,但“现阶段我不会承诺终身免费”。
完整逐字稿
I can't help but wonder what sort of grand game is afoot here. When you really step back and think about what is happening right now, what is the game? What is going on here? If you had to sum up the thing you're in the middle of, what would it be?
You're going up against the biggest players in the world. The window to do it is actually on the order of months. The compute required to do it is enormous this year alone. You have no distribution advantage, but you have this massively growing viral product. What do you do?
This is probably the most exciting software race ever, and the outcome is trillions of dollars. No, I know this is the first time that you're talking about the business in long form like this. I'm incredibly excited to ask you all about it.
It seems like we're in a moment with personal agents that is very similar to, and arguably probably bigger than, what we saw with code generation a year or so ago. Maybe frame this up for us to begin: How are you thinking about the impact that your product, and these agents broadly, are going to have on the world? Why is this the area that you've chosen to dedicate yourself to completely?
I think a new way will come out of this—one that most people on the planet will use to interact with technology broadly. I run Instinct. It's a new company that I started about a year ago, so it's still a new company. We're building a personal assistant. I'm not going to spice it up, because it's really just a personal assistant.
I think the thing that has really enabled us to gain quite a bit of traction and excitement early is that it just works. It works in the way that I would say we all—I mean, I'm also a user—have really been waiting for AI to act for us.
One interesting thing is that this is not like any other transformational consumer app experience, where it's the founder coming on and saying, "I have this new vision for the world. Trust me on this, and in a few years you will all get it." This is a very different moment of building a product, because you and I—we all—already have this idea of what AI should act like.
We've had that idea since 2023, when we all started to brainstorm and began using ChatGPT for the first time. That agent should be able to help you with effectively anything in your everyday life, whether it's something new and creative that you want to do or something that you traditionally spend several hours on and that takes a long time to develop.
It's just everyday intelligence. It's meant to be with you and act with you. It's meant to feel like a very simple interface. When I say "simple interface," what I mean is that we don't even have an application. This is not a new app or a new tool. It's a new experience.
It has a phone and a computer, so you can text it, you can call it, and it can actually call you, too. I don't know if you've experienced this yet.
I have.
Oh, you have? I think it's only called me about 3 times so far over the several months that I've been using it. It's funny: If there's something genuinely pressing with a deadline, it will call you and say, "I don't want to bother you too much, but you need to sign this document by 3 p.m., and it's 2:55 right now. Can you please do that? It's in your inbox. I can even send you another email to push it up to the top, right?"
So it's very socially intelligent and socially aware. You can email it. It has its own email address. I don't want to confuse the simple interface with limited ability, because the whole nature of it is that there shouldn't be any new application needed to interact with AI.
I think it should have the social intelligence and awareness to be able to act with you just as we do with other people. It's really meant to be the simplest possible experience, yet it has significant capabilities. It has a computer, so it can do quite literally anything that you might want to do—or that you do yourself—on the internet.
1. What people are using AI agents for
Tell us a couple of the craziest stories that come to mind about things that have happened because of Manus. I ask because you're not programming it to do any one thing. You're programming it with a phone and a computer to be able to do anything.
The U.S. Open example, I think, is one everyone's familiar with. This couple appeared on the jumbotron and wanted the video footage of it. Somehow, Manus went out, figured this out, and brought it back to them, which is kind of crazy. What are your favorite couple of examples of wild things that have happened because of Instinct?
For people who shop online, there's a group of people sharing this and making it a more recurring use case: scanning every item, or every piece of clothing, in their wardrobe. They'll go into their wardrobe and scan every single item—every top, bottom, sock, shoes, and whatever else. Then they'll scan themselves, too: their face, body, and proportions.
After that, they'll try on clothes and have Manus plan out their week in terms of what to wear. Every week, with their existing wardrobe, it will determine what shoes to wear, what top to wear with a certain item, and which version of it to use. It wouldn't send them a bullet-point list of the items they should wear; it would show them wearing the outfits and say, "This is what today will look like."
But it's also useful for shopping. They'll take those same capabilities and say, "Go shop on the internet." Instinct will come up with thousands of outfits from many different places. It will show them the pieces from top to bottom, and then they can simply say, "Order."
It will provide a thousand different options, and they'll be looking at themselves wearing the clothing. Then they'll be able to point to the one they really want and say, "Can you send that to my house?" They'll also say, "Every day, can you come up with 3 creative new outfits from head to toe and send those to me?" Then, with a click of a button, they're able to get the new outfit.
There are a lot of goal-oriented use cases, too, such as, "Here's a personal finance goal that I have. I want to save this much money by this month." Then it's actively working with them to do that.
Users connect their bank accounts to Instinct, and it can scan all the transactions and subscriptions they've had. It surfaces them to the user and says, "Do you actually use this?" Then they're like, "Oh, no, I didn't even know that."
And then it can go into the email and unsubscribe from it?
Yes. It can go end to end. These popular consumer fintech products will surface subscriptions to you and tell you, "By the way, you should cancel these." Instinct will go end to end.
It will go to the site, sign in, handle any confirmation that goes to the user's email if it has access to the email, and then go all the way to the end and cancel the subscription. Then it sends back the amount it saved you: "By the way, you're saving $2,000 this month because of these things."
Man, any product that's dependent on consumer laziness or inertia is toast, huh? Toast.
Yes.
That's good. Good for the consumer. Can you tell us about the Instinct-to-Instinct network and the future of agents talking to each other?
There are more practical scenarios where this is something new that I think will transform the way certain people act. I think it's only been around for a week or so—or I guess 10 days.
Honestly, the purpose of this is not to think about a new feature to build, but more about very busy working professionals. They're taking meetings all day long and constantly meeting with new and existing people.
The process of getting a meeting on the calendar is so laborious. You text, "Hey, do you want to meet at this time?" Then they say, "No, that doesn't work. But these times work." And you say, "Oh, but I'm traveling, so maybe this time works—and I'll be in this time zone, though." There's just so much back and forth.
However, if the 2 users who were using Instinct—probably also brainstorming these times with Instinct on both ends—were able to communicate directly, they could say, "The goal is just to find the time, right? We don't need to play this back-and-forth game, and nobody's really playing games here. We're just trying to find time."
That's been a very canonical use case. The way that you can find time on a calendar now is just so different. You simply state your intention: "I want to meet with this person, ideally by the end of the week or during these periods of time. Find me a time."
It will then coordinate with the other person's Instinct to find when time is available and when it isn't, and eventually put something on the calendar.
I think the key unlock here—we called it a trusted person network. The key is that you should only be connected with your trusted people. There are some interesting social dynamics coming into play that I've been learning about. It's designed such that you only bring in people who are trusted, who are not going to maliciously try to find out what your calendar is or what's in your email. Not only that, we provide all the controls to give different levels of access to different people. A lot of spouses actually coordinate their Manuses because they can generally just share everything.
So that might just be, “Hey, this is my spouse. Just share anything with them,” right? Maybe with certain colleagues, it's, “Hey, my work calendar is available. Certain parts of my inbox, if they're asking for documentation or something like that, are available, but everything else, let's keep that off-limits.” It'll be able to create that.
When you think about this, there's almost like these network effects built through these nodes in the network, with different edges and different weights, too, right? It's not like a friend graph where you're just saying, “I'm connected to you,” and that's it, right?
It's, “I'm connected to you. I trust you. I've also configured it in this way such that there's varying levels of access.” So it's this interesting network effect almost being created, and different social dynamics are being explored and created as well. There are some interesting cases I've learned about where, if somebody violates your trust within the network, there's an implication on that relationship itself, right?
So let's say you and I connected, and I said, “Hey, Patrick, I trust you. Let's connect on the platform so that it's easier the next time to put time on the calendar.” Then you use that and start digging for certain data in a certain area. First of all, maybe I only gave you calendar access, right? But you're digging for data in a certain area. Then my Instinct texts me, like, “Hey, by the way, Patrick's looking for this type of information.” Now there's almost like a trust broken between the 2 of us. All of these little social dynamics come into play.
We've really thought about how to curate this network such that you do get the benefit—to schedule time and make plans with others and other things like that—but then also use some of the existing rules in society to enforce both social and cultural norms, as well as technical limitations in terms of what is visible and what is not visible.
We're really going to see an entirely new order, aren't we? It's going to be fascinating. The number of things I can think of—let's just take my wife, for example. This would make life so much easier in so many different ways so quickly.
Even little things like, you were trying to find this house earlier. If I could just say to my Instinct, “Hey, is no one nearby? Where is he?” and you give me 1-day permission or something like this, there's just so many ways to imagine it being useful. It's pretty wild.
With the trusted network, we find a lot of very interesting cases, or example use cases, that people share. It is, I would say, the 1 part of the platform that is just something new and creative. In every other area of the platform, the user generally has a good feeling for what it should act like.
For example, I heard the other day about how, when you're scheduling plans with friends or with a spouse, it's always an active act, right? It's like, “Okay, every Thursday maybe I'll think about something on Saturday night—what I should do.” Then you text the group chat, like, “Hey, do you want to do this?” Somebody's always not willing to do it, and eventually it falls apart, right? That happens every time.
There's a friend group that, every week, says, “Brainstorm something creative that we all like and make it a different experience every week.” It then goes through that every week and not only coordinates when everybody's available and what their interests are, whether they like certain things or not, and what shows or concerts are available based on their music tastes, because all their Spotify histories are connected. That same friend group was a group of 6.
They had 1 person order an Uber, and then had that Uber go around to all 6 people and coordinate: “Hey, we'll pick you up in 10 minutes,” and then to the other person, “We'll pick you up in 5 minutes from here.” I found the optimal route to pick up all the people, and now they're all just sharing 1 Uber. So it's actually very economically friendly.
There are all these random examples of people discovering new use cases when you can finally remove the barriers of a lot of the rote social communication and remove all of the logistical burden.
2. Rethinking travel, reservations, and the internet
It's almost crazy how simple the explanation of this is: it's easiest to literally just pretend you have a superhuman person with a phone, a computer, and an email address. They can call you just like dealing with a person, and that's it. That makes me wonder how you think about everyone in the world having 1 of these—having an Instinct, having an agent—and how that will reorder things.
Code has been incredibly exciting to watch. Of course, everyone listening has had their own version of a magical experience with making something, but not that many people were software engineers before this. It's a relatively small set of people that have made code historically. Maybe now it's going to be way more, but this just seems like a different new market, a new paradigm.
How do you think this will start to reorder the world, the internet, commerce, et cetera? What's the future of interfaces, I guess, is another way of saying it?
I'll break this down in 2 parts so that we don't sound like complete abstract visionaries saying that we're going to reinvent the internet. Although, I'll be clear: I do think that in the coming years, the internet is going to be rewritten. I think that the way most people on the planet interact with software is going to be very different.
To break it down, though, I think in the short and medium term, reservations are one area. Historically, how is that done? The human has to go onto the site for whatever restaurant they care about, click through the site, put in 2 people this time, and try to get it, right? If they don't get it, they're just too late and it's first come, first served.
But now, when you have an agent, you can theoretically just have the agent check every 5 seconds—not just across that 1 restaurant that you really care about, the favorite restaurant you can't get into. Sure, that's 1 case, right? But what if it's doing that across every site in the world, for every restaurant in every major city? Now you can see how the agent is able to exploit something simple like restaurant reservations.
If I go into that a little deeper, we have some exciting partnerships to be released in the future, but we are reinventing what the reservation system is in this 1 category because it's important to our users, which is actually better for both sides, right? If you take restaurant reservations from scratch, from the user's end, they want the reservation for the significant life event that they might have.
It's a birthday, an anniversary, or a friend coming in from out of town. On the restaurant end, they want interesting people, special occasions, and opportunities to cater to those, rather than the local who keeps taking up a table once a night when there's no special event. Traditionally, because of the way the internet has worked, at least for reservations, it's been first come, first served. Anyone who comes in, no matter how important or unimportant it might be, gets the reservation if they get in first.
But what if you had the ability for the agent to communicate on both sides? It could say, "Hey, this is important. This is actually the person's spouse's 30th birthday. A very special event is coming up." The restaurant could also say, "Let's prioritize birthdays, major decade birthdays, major anniversaries, or certain people." I don't know what it might be, but now, with the ability to remove all of the logistical burden of having to describe exactly what your situation is, what's happening, and why it's important, we can perfectly match what the restaurant wants with what the user wants.
That will result in more users being able to actually have a spot in the restaurant for special events. On the restaurant's end, they'll be able to have a much better audience of people who are there for all these different special events or special occasions. I don't mean to go so deep on restaurant reservations; that's just one example of a traditional model that is changing. I think the same is true for a lot of the major travel agencies.
Honestly, 50% of our transaction volume that goes through our platform is due to travel. We're still running an invite-only program, and it's quite early for us, but we're approaching $1 billion a year in transaction volume through the platform.
On a small user base.
On a very small user base, there's already over $1 billion a year transacting. Fifty percent of that is travel alone. Travel agencies that we might all use today—or, more accurately, that people who are not on Instinct yet use today—to book hotels, flights, or other things like that: What is the benefit of that interface?
Well, they properly unify so many different services, hotel chains, and airlines, and present everything to the user in a nice way. Now, the user can just click on an option, check out immediately, and see it in their email. It's a great experience. But I think there's an even better experience through Instinct or other similar interfaces where you can just say, "Hey, I need to be in New York tonight." This is immediate, and that's it.
What does that entail? First, where is the user right now? Maybe you're in San Francisco, maybe you're in Los Angeles, or maybe you're somewhere else. If you want it to, Instinct can see your location. It'll say, "Okay, you're in San Francisco and you need to be in New York. Let me first find all the options." But not only that—it can account for all of the nuance, too. What's the person's preferred airline? What's the preferred seat type, and what class do you want to sit in? Do you want an aisle seat, a window seat, or a middle seat? What food options do you want delivered to you as well? What credit card would you like to use? Then it books the flight.
After that, it moves on to the hotel.
Yeah. Presumably, it learns about you constantly, so your preferences are stored.
That is the nice thing. If you tell it one thing—"I prefer this. This is my style. This is what I like. This is where I like to stay"—then, whenever you're staying anywhere in the future, it will be able to use that memory and make it much easier for you to book according to exactly what you want. Not only that, but it'll be able to take that general taste, that general preference set, and extrapolate it to anything else that you might want to book. Things just become very easy as you give it more preferences over time.
To close out this example, it'll find the ideal hotel. Maybe you've already stayed there the last 7 times, so it's very easy in that way. It'll book everything end to end and line it up with your calendar. It'll put in the flight, the Uber ride you need to take to get to the airport, the Uber ride to get back from the airport to the hotel, and all of the events that you might want to do there.
Just take a step back, because I've talked so much about what is happening under the hood. The user really just recorded a voice message saying, "I need to be in New York tonight," and everything else is solved. That's what I mean when I say that we meet the users where they are. Where is the delightful experience? This is a new delightful experience that I believe is going to transform the travel agency industry.
3. Trust, privacy, and personal data
For these businesses that traditionally make a lot of money by providing standardized interfaces, what happens when a new standardized interface comes into play that is just that much easier? What does that mean? I'm not saying that we're in the business of trying to disrupt those businesses. I think they do provide quite a bit of value from the data they've collected over time and the networks they have. It'll be a collaboration over the next year or couple of years to redefine what that industry looks like.
What has it been like getting people to trust their agent, and learning what it takes to get people to trust it? I would love to go fairly deep here. We were talking about this this past week, and I love the examples that you've given about moments when people are showing trust, the data you've started to gather, and the lessons you're starting to learn about the importance of trusting this thing.
I'd love you to talk about the trade-off between privacy and the effectiveness of these agents. Obviously, the more context, passwords, and everything else you give it, the more it can do. I think people are constantly running that trade-off in their heads when they're interfacing with the thing. Talk about that. Talk about trust and what you've learned so far.
There's an interesting data flywheel, or chicken-and-egg problem, here. The more data that you give it, the more proactive it can be, the more sympathetic to your situation it can be, and the more it can act and be useful to you.
What we find in the data is that it takes several weeks to build trust. I don't have a problem with that, because one of the core principles on our end is that the user should always feel in control of their data. They should always be in control of their data. They should share data with Instinct at the rate at which they feel comfortable, and if they want to take it back, they can certainly take it back.
What we find is that, over several weeks—I believe I was looking at the numbers the other day—3 weeks in, there's a 40% chance that the user has shared a personal credit card with Instinct. That's 40% of the user base. That includes some proportion of the user base that probably churned at that point, or at least churned before then.
The reason I'm sharing this is because time to first credit card, time to first account password, or time to first sensitive piece of information are all proxies for trust, and we highly value this and take it very seriously. At 3 weeks in, 40% of the user base is sharing a credit card. There's something there, right? There's something there. No pun intended. It's instinctual to go to it with whatever you need.
Through that process over the first 3 weeks, the user learns to build trust. From there, it's just a snowballing effect. We find that when users connect at least 1 piece of sensitive information to Instinct and really trust Instinct, there's about an 80% retention rate. Eighty percent if you share 1 piece of information.
Crazy for a consumer technology. How should people out there who want to try this, but have a natural reticence to trust—not Instinct in particular, but just anything, any AI agent, with all of their information—think about the actual risk of doing this? How would you reassure them that you've built your technology in such a way that the odds of something bad happening, if they do trust you with very sensitive stuff, are really low or close to zero?
There's a way to break this problem down into 2 main parts. One is the storage of sensitive information. There are many other businesses and products that also deal with this problem: They have access to sensitive information, so what are they doing proactively to make sure that it is isolated and locked down?
Then there's a second area, which is the new problems that we need to solve—the new surface areas or the new capabilities that we need to be aware of.
The first case is a tractable problem. It's just very hard work, attention, and care that you need to put into making sure sensitive data that's shared is safe and that the user is in full control over that data.
There's that second category, which I'll spend more time talking about. This is the first time that an agent has been able to have—I said within 3 weeks, 40% of the user base is giving Instinct access to a credit card autonomously, to be able to purchase theoretically anywhere. I don't know the numbers on email or what proportion of users are connecting their email, but you can imagine full access to an email inbox and a calendar. So there's a lot of surface area here.
There are systems that we've put in place, detached from the Instinct agent architecture itself, that are put in place to be proactive about these things and decouple the risk, if I were to say so. For example, any piece of content, any piece of text, anything, any form of media that comes in that might be consumed by Instinct goes through what we call these firewalls, which can intercept, reject, or block malicious pieces of content coming in and hitting Instinct and trying to convince Instinct to do something.
4. The business model behind Instinct
There is also, for every action that Instinct might take or every thought that it might have, a system that is actively monitoring it and is decoupled from Instinct itself. That system is able to pause or intercept it to approve or disapprove of what might happen next before it takes the action. Those are just 2 pieces that we've put in place, but there's so much more under the hood that enables the agent to be as capable as possible, but also safe and trustworthy.
I have this question around—I don't have a better word than alignment, and I know alignment is a very loaded word in AI. I don't mean humanity-scale alignment. I mean an agent aligned with me personally.
If I think about other agents I've hired, like employees, I pay them money and therefore I trust them to have my interests at heart. It won't be perfect, but they don't have some other alternative incentive stream that guides their behavior. They're guided by their employment.
How do you think about the business model vis-à-vis alignment? You could walk us through how you've thought about it, but many approaches will be taken. Some will be paid, some will be free, and the free ones will monetize in different ways. Can you walk us through this decision tree of what the business model is or will be and how you arrived at that as the ideal conclusion for the agent?
We don't want Instinct to influence the user's behavior in a way that is not aligned with what the user wants. That sounds very good on the surface level, but I want to call out how important that is, because imagine a world in which Instinct is generally smarter than the user—more socially intelligent, more socially aware, and textbook-smart as well.
I think it would be a very dangerous world if Instinct were influencing the user's behavior to purchase something that they don't want to purchase or to subscribe to something that they don't want. You look at most of the major consumer businesses today that are able to influence users' behavior against what they might want to do. I'm talking about the major platforms, whether that's Google, TikTok, Instagram, Snapchat, or so many others, where you have this platform and it's free for the users, but there are so many instances in which paid ads are pushed to the user and try to convince the user to purchase.
By the very nature of the fact that the user does purchase, it now becomes this game of brands paying to convince users to purchase items that they may or may not want.
This is the idea that if you're not paying, you're the product.
Exactly. As a programmer and a technologist, I just don't want to build that reality. I think that's a very dangerous reality.
Instinct should act on behalf of what the user wants. We take this very seriously when we're building the product and through the various research projects that we have. Instinct is not unlike any other AI product. Instinct is not a task accomplisher.
What I mean by that is, with most other AI products, you write a prompt, and then it accomplishes the task and tells you what happened. That seems good in theory, but what happens is that if the user is asking for something that might not be well-intentioned, if you have a task accomplisher, it's just going to do that and listen to the user.
Instinct follows higher-level objectives. Instinct will learn to build trust with the user, learn to make the user genuinely feel safer with it, and learn to watch over the user and have their back when things might be dropped or other things like that.
When the user asks it to do something, if it is well-intentioned and well-meaning, one way to communicate safety and trust is just to do the task. So we end up doing a superset of what most other AI products are able to do.
I think focusing on higher-level objectives is very important here because it enables Instinct to be more robust to these edge cases. We don't want to influence users' behavior in a way that is not aligned with what the user truly wants.
Even on the business-model side of this, there's over $1 billion flowing through the platform every year, and we're just getting started. Honestly, we're growing at 10% day over day. So you imagine the transaction volume is also compounding at 10% day over day. That's not just $1 billion flowing through the platform now. That's $1.1 billion tomorrow, and then that's like $1.2-something billion the next day, and $1.3-something billion the next day.
It's still very early, but when I see transaction volume that is so high flowing through the platform, what I see is a very basic case. It's similar to Apple Pay, or it's similar to MX, or any other platform that provides distribution to underlying services and a great user experience.
Apple Pay is a great experience, right? You can go anywhere and scan your card, and the user doesn't have to pay for it. The user is getting a free, great experience, and the merchants on the other end who are benefiting from the business are paying to be a part of that platform.
So I see a blanket transaction take rate being enforced across the platform, which is just us exchanging distribution for being able to serve products on behalf of merchants.
Can you say a little bit more about how that runs into the existing world? I can imagine the layers being that you could be a card issuer, you could be something like Stripe, or you could be Visa or Mastercard. There are sort of rails that have been built in the payments world that create convenience, reduce friction, and take a vig as a percentage of the transaction. Those are some amazing businesses, to be sure.
How do you think about which of those are partners and which of those are potential things you would displace? How does that vision of a small take rate on the transaction volume on Instinct, because it's a free product, slot into the existing world?
Just to put that into context, those are great businesses, and there are so many people along the line. You make 1 digital transaction, and there are like 40 people along the line who make money on that.
And just to take a step back, those 40 people making money at each layer of the stack are really sharing 2.5%, 2%—it depends on where the transaction is coming from. So it's actually a very small piece of the pie. We're not primarily interested in doing what MX does best and having the network to do, or even what Stripe does on the internet, or some of the underlying payment infrastructure, because it's such a small piece of the pie, and I think they're providing real value.
Where I see the majority of the value, just thinking on the business side, is if you look at most of the other major platforms in the world and the take rates that they're able to achieve. Again, it's free for the user. It's a great free experience for the user, and the merchant is now recognizing the distribution source. You have Shopify, which I believe is between 2%, 2.5%, and 3%, providing its services and taking some take rate from those businesses. I'm not sure where Stripe is. I think it's maybe on the lower end of that. Then you have Amazon, which has a great platform and takes upwards of 10%. And, of course, the premier one is Apple, where any in-app purchase is 30%.
I'm not saying that we're going to be 30%. I think that's very unrealistic for a lot of businesses. But what I'm saying is that we still don't know where along this curve of distribution power versus take rate we're going to be. I just want to call this out again because I don't want it to be misunderstood: it's a free experience for the user. It's a free, great experience for the service that we're providing to all of the underlying providers. I'm not focused on finding 30 bps on the 2.5% with some partnership with some payment provider. I'm looking at whether we can provide so much value that we're on the upper end of this scale.
The reason why I think this is possible is very practical. Fifty percent of the transaction volume flowing through the platform is travel alone. We all know the travel industry and some of the rates that we've kind of—
Yeah, the OTA rates are high.
It's very high for flights. Maybe it's on the lower scale, but how many single-digit percentages can you take off of a flight? For hotels, some of these boutique hotels are offering to pay up to 30% for every transaction that you're able to deliver for them. I'm not saying, again, that we're going to be at 30%, but you can see the range. These are existing business models that we can bootstrap off of in the early days.
5. How existing businesses will adapt
But what would it be like to take that and extend it to effectively every major industry? We only have the ability to do this if it's true that most digital behavior then moves to these new types of interfaces, which is, again, no interface. It is only the case if there's significant distribution power. So I think there's actually quite an interesting intellectual question here about how this evolves over time.
It seems to me like there's going to be a serious corporate agent war. You've already seen this with Amazon and Muse[?]. You've got all these established players with tremendous vested interests and relationships with customers that Instinct and other agents could really disrupt in a major way. What are you thinking about that? How do you interface with the other great services out there?
I'll pick a random one. I use Uber Eats a lot. I order from Uber Eats all the time. It's kind of a pain in the ass to click through the thing. I can imagine a much better experience on a snappy WhatsApp connection with Instinct or something saying, “Hey, I want my usual from this restaurant,” and that's it. There's nothing else; it has its computer and goes on to Uber Eats and orders, or whatever.
At what point does Uber Eats not like that anymore? At what point do these great businesses that have been built up start to be adversarial against agents? What do you think are the most likely sources of conflict and reconciliation? It's going to be really interesting to watch.
I'm serious when I say—and you're alluding to this, too—that I think most digital services or industries are going to be not disrupted, but changed and transformed. I was running through an exercise the other day of pulling effectively every digital service, business, or application from various different verticals and industries, and plotting it along a line: what proportion of the user's attention and experience on the product is proportionate to revenue, and then what proportion of it is delivering the underlying service? Meaning, if the user didn't use the app, what proportion of that transaction would remain? Or what portion of their revenue is due to the underlying good or service being provided?
I think your question is about that upper end of the scale, where the majority, or some significant amount, of revenue is due to attention on the application, advertising, or things like that. So yes, that's Uber. I actually don't know what Uber's number is. That's a lot of the restaurant or food-delivery services, the travel agencies, and even Amazon itself, with the upselling that it does on its platform.
There's one simple way to look at this, which is that it's 70% ads and 30% good, and therefore you're going to slash the 70% and they're going to be a 30% business moving forward. There's another way to look at it, which is that I would take any business. Let's take Uber Eats or DoorDash as an example. I don't know the numbers, but as they reduce the number of clicks needed to check out—to order food to your house or to order an Uber—the transaction volume increases because it's reduced friction to get the same underlying good.
We're taking something like Instinct and making the friction to do anything almost zero. It is literally zero in the case that there is proactive behavior for Uber. Let's say for ridesharing, if Instinct has access to your calendar and owns your calendar, is booking all of these events, and knows where you need to be in person here and in person there, what if Instinct just always had a car lined up for you every time you need to be somewhere? It would be out of your mind to think, “In 5 minutes I need to order this Uber because I need to be in this place in 45 minutes and there might be traffic. Let me check the app to see how much traffic there is and when I need to order the ride.”
Instead, what if it were just proactive behavior? A car will always be lined up, and you will never be late because it's going to calculate traffic and figure these things out. What would that do to Uber's business, or to any rideshare business, if the default for existing riders who love Uber, Lyft, or any other rideshare business were for productivity to make the friction to experience or access that good or service nearly zero?
You talk about food-delivery services. Maybe you're coming home late off of an airplane, you need to be at your house, and you haven't eaten yet. It just texts you and says, “Are you hungry? Do you want the same thing that you ordered yesterday or 5 days ago? I can send it to your house if you'd really like it.” I know you're in the Uber because it ordered the Uber, too, so you'll get home at this time. Then the user just says, “Yeah, go with that. That's great,” or, “Thank you. Yeah, please order that.”
With very little friction, what will that do to the total transaction volume, or the number of times that a user interacts with the business? I actually think it'll go up. So it's this interesting game. I think it's this interesting transition period between users spending a lot of time on apps—painfully spending a lot of time on apps—and that being a monetizable surface because that's where the user's attention is, to the user actually interacting with the business even more. That's counterintuitive: actually interacting with the business more because the friction to do so is much less.
I think it'll be this interesting game. All these various industries are going to move in a different way, and we hope, on our end—I think the way to approach any big change like this is not to come in hot. We're still a new company. We're just getting started. We shouldn't come in hot and immediately start disrupting certain businesses, but go to them and just say, “Hey, this is what we think. This is what we think your business looks like.”
You guys certainly know what your business looks like. This is how users on Instinct are interacting with your business already. What can we do in collaboration to make that a better experience for both sides, where it makes sense for your business, makes sense for our business, and is simply a much better experience for the end user?
That is something that we're exploring and learning across so many different industries right now. If you think about the things that traditional companies could be doing now to prepare themselves for an agent-rich world, let's pretend half of Americans have an agent that's doing all the stuff you just described, which sounds incredible, magical, and very democratizing. I think that's something to highlight: This is going to bring to everyone capabilities that have been rare or expensive. I think that's a really cool feature of agents in general, but we can come back to that.
What kinds of businesses are going to thrive in that world? What should businesses think about doing to prepare for that world and be successful in it, do you think?
I think it comes back to that breakdown. That's why I was doing that exercise the other day: Where is your revenue coming from? What proportion of that is from the user spending time in your application? What proportion of that is from the user having access to the underlying service?
It's very clear: If your business benefits from more transaction volume, not at the cost of—or even with the cost of—less time on the application, then you're going to be in a really great spot, because Instinct is going to make it 100 times easier to do that.
If you're in a spot where nearly 100% of your revenue is due to the user's attention, I think that, in a lot of cases, is against the will of the user and what they want to do. There are so many games, or so much malicious product building, that tries to convince the user to use the application more against their will.
We're thinking about a lot of these social media companies where the user doesn't want to be on the app. They don't feel happy when they're on the app, but they're unwillingly giving their time to it. They can't get off, and they keep scrolling. That's because the underlying business benefits from the user's attention.
It's almost liberating for the user to be able to—and this is why it's so important for Instinct to act on behalf of what's best for the user—deliver experiences where you can actually liberate the user from being sucked into these infinite-scrolling moments.
I would say any blanket advice is probably not well thought out. I think it's a case-by-case basis and certainly differs across industries.
Very practically, we're finding early partners with very innovative CEOs or other executives who are really thinking ahead and are willing to be early partners. What we're discovering is almost like a playbook for how every business, no matter where it exists along that risk curve, can discover what the risks are, honestly, so that they have a little bit of data to work with. Then we can work together on finding ways where we can land in a happier spot for both sides.
One of those things is that you don't have to go all-in. You don't have to say, "Let's just turn it on," and suddenly see 70% of your revenue go to zero, leaving you stuck in this odd place. You can mitigate the risk. You can scale down the experiment. They can run A/B tests to figure out, "If we enable this certain thing across 1% of users, how do they interact with the business?"
Honestly, does the user like it more? Is it a more enjoyable experience? For the brand side, does the transaction volume go up? Does the willingness to buy or access the product itself increase?
There's a lot of work in product discovery, too. A lot of times, the user doesn't know they want to buy something, but they can't find it. Does this actually increase the user's ability to find exactly what they want?
I think running scaled experiments here is a really great playbook, because you can scale the risk accordingly. At the end of it, you get proportional data, so you don't have to run the experiment across your entire user base. That's what we're finding is really working so far. Again, it's still very early, so we're still discovering this in real time.
6. Designing a personal assistant people love
Before I ask more questions about the world-reordering nature of personal agents, I'd love to take a little side quest in the conversation and talk about what it takes to do all this—to provide all this. I want to hear about what's been hard about building the technology itself. I want to hear about compute. I think you said to me at some point that you spend a big chunk of your time just thinking about compute right now. Maybe that's different 5 years from now, but certainly in the moment, when you're growing fast, this is a really important thing.
I'd love to hear your thoughts on that. Talk us through what it's been like to build Instinct itself and the key, hard things to do. Before we do that, just because I'm remembering all of our conversations, it would be helpful for you first to frame up what you want it to feel like and why it's so important to you that it has this distinctive quality and performance before we talk about how you deliver those things.
Maybe first, just say a quick word on that: What do you want Instinct to feel like as a product?
I could talk all day about this, because I think that is so important. Honestly, as a product builder, it's a new muscle to flex and build. Really thinking beyond capabilities is something I want to push here.
Over the last 3 years, we've seen all these different product launches and new products saying, "AI can now do this," or, "AI can now do that," or, "Did you know that it can do this thing because of this small technical thing that happened under the hood?" I think the consumer is first fatigued by all this. They don't know how to access it. Second, I think we're missing the point.
Early on, one principle that we held was: Let's not focus on capability. Let's only focus on understandability. How much does a user understand about what's happening? What is their ability to predict what will happen when they ask this, do this, or interact with it in this way?
Honestly, I think that is one of the major factors that has led to engagement numbers that are completely off the charts and viral word-of-mouth growth that's happening at 10% a day. It's understandable. It should just feel good in some way, in the way that it communicates to you.
The purpose of communicating or sending a text message is not just the meaning of the text itself. It's down to underlying details, even including the shape of the text message. How will the user feel when they see that? If you see a big blob of text that requires the user to scroll to find what the next message is, versus front-loading some of the information so that the user really gets it in the first 30% and can optionally read the rest, that's an important consideration.
You have to think about how the user might read, or even about reading patterns. Most people scanning big chunks of text might read 80% of the first line, then maybe 50% of the next line, and then it tapers off, so it looks like a flag. That's a consideration that should be made.
If Instinct has the ability to think about things like this, it can understand that the user is going to spend most of their time reading the first 2 lines, and certainly the first part of each line, too. How does it craft its message to deliver it in the lowest-fatigue way, or with the least amount of cognitive load placed on the user?
There's a lot of consideration put into this in terms of product building. There's quite a bit of work we do on the infrastructure side to make it fast and affordable to serve, but this is another area that I think is so important and is going to be a differentiator.
How much of that is your personal taste and the team's taste versus being the result of a quantitative process—an iterative, quantitative process where you A/B-test emoji reactions versus short versus long, and this is the optimal thing? How much of it is a data-driven optimization exercise versus your own sensibility and your team's sensibility?
Honestly, it's all of the above. When you're thinking about building evals, or evaluation and other testing frameworks, to test for these very soft qualities—or these actually very long-term qualities, such as whether a user trusts Instinct 2 or 3 weeks in—how do you measure that? How do you run simulated evaluations to test whether the user is going to feel trust in 2 or 3 weeks?
A lot of it is staged rollouts over time.
So, I might come in and build a slightly different experience, and then I'll release it to myself. I'll play around with it for a little bit and see how I feel about it. I'm very opinionated about these types of things.
Then, if I feel comfortable with it, I'll send it out to the team. I'll say, “Hey, you guys should try this.” We'll see how they feel about it, and then they'll send it out to our smaller early-access group. They'll play around with it and see how it feels. If we're confident, and there are any tweaks that we need to make, we'll do that, and then we'll eventually roll it out to the general public.
I think this is very important because I think that Manus is quite capable and is one of the most capable products out there. I'm not saying that, over the next couple of months or years, others aren't going to come and deliver the same seemingly similar experience. But I think this deep focus and priority on how the user feels, and how to make it the most enjoyable experience beyond the words that it's saying—the way that it feels—is so important.
7. Growth, compute, and competing with Big Tech
One of the great things about the history of technology is this race between incumbents getting quality and innovation, versus upstarts like you getting distribution. Obviously, you've built an incredible product. The feel of it, like you said, is the worst it'll ever be. How do you think about that challenge—your speed of scaling and what your ambition is for how to get big really, really quickly?
Do you think this is a winner-take-most or winner-take-all market? What will the market shape of agents be? I'm really curious how you're thinking about this. You've got this foothold, you're growing 10% every day, and you do that math—it gets really big, really quickly. You need a lot of compute, and it's a free product. You're a very chill guy, but it just seems like a stressful situation to be in, facing down how big this could get as quickly as it could get. Talk us through that many-headed monster of a problem.
I think you just described everything all at once in about 20 seconds there. Maybe we can think about the growth story so far, just to describe where we're coming from. We are, famously or infamously, serving an invite-only product, which is honestly not meant to be an exclusive thing. Some users are treating it like that, but that's really not the intention.
The goal here is to scale this thing as fast as possible. I'm ambitious, but I also want to do it in a responsible way that enables us not to wake up one morning and have 10 times the number of users, only to find that 80% of them can't talk to it because there's not enough compute.
The interesting thing, both the benefit and the challenge, is that when we first started this program, we gave it to about 200 people. They were close friends and family members, and we just said, “Go try it out.” The next day, about 5 people came onto the platform because they had referred it to somebody.
Just to describe it a little bit more, it's an invite-only platform, and every user gets 5 invites. We rolled it out to 200 people. The next day, it was 205, and then the day after that, it was 210. A couple of people were sharing it with one person, so we thought, “Okay, cool.”
But that started to accelerate. It was not 1% or 2%; it became 3% or 4%. Once we hit a couple thousand users, some people started sharing it online—just natively sharing a cool use case they had with it. That accelerated the growth. It turned into 6%, 7%, 8%, 9%, and now I believe we're at 10% or 11% day over day.
To call that out, it's not that we're doing some creative marketing event every day. We've spent $0 on marketing so far. It's not that we're doing something every single day to support this growth. Every day, about 10% of the audience—or slightly less, because users can refer multiple people—are deciding to give up one of their 5 valuable invites to somebody else. That is happening every single day at a 10% rate.
I think there's something very significant there when I talk about the strength of word of mouth. It's honestly a little surprising. It's one of the strongest cases of word-of-mouth growth I've seen. I find all these stories of people saying, “Hey, actually”—they'll ask me, and they'll be ashamed. They'll email me and say, “Can I please get an invite? I think I have a friend who has it, but I don't know if I make it into his 5 friends.”
And then there are also people who are bragging, “I got 3 invites left,” and, “I'm holding on to them right now.” I saw the other day that some invites were selling on eBay, too. I don't know if you saw this. It was like $300—people were buying these invites on eBay.
It's just to control the growth. The big question that most people are asking right now is that you have much bigger players that are able to distribute to 1 billion or 2 billion people on the planet immediately. They may not be compounding naturally as fast, but they have such great top-of-funnel distribution. Then you have us, compounding at a very fast clip every day, but we don't own a major service with 2 billion or 3 billion people that we can distribute to immediately.
It's this interesting question: Where does the curve line up, and where's the inflection point? There's another problem that comes with this, too, which is an interesting scaling problem. This is the core problem that I spend about 40% of my time worrying about.
It's not like the traditional consumer products that grew very fast. To double the number of users on Instagram or Facebook, for example, would mean this many additional requests going through the platform. There's a scaling story there, and it's certainly hard infrastructure work. We have that, too, to be fair. But what do you do when the underlying compute also needs to grow by 10% every day?
We've been doing this for several weeks now. What does it mean when the amount of compute you need access to is effectively doubling every week? Do we buy 2 times the compute we have right now? We're going to consume that in a week. Then do you buy 5 times? We're going to consume that in less than 3 weeks. Do you buy 10 times? Now you're 10 times leveraged, if you can even stomach what it's like to buy 10 times ahead. But then you're going to consume that in a couple of weeks.
That's the hard problem. It's thinking about how far ahead to buy. It's going at a faster rate than Claude Code or Codex, or some of those other applications where they also had to reason about similar exponential problems.
The other subtlety here is that it's not like a SaaS business where, when you double the number of users, you can buy 2 times more resources to power it. The resource has a lead time of several months. You can't just go out tomorrow and start buying compute because you get charged 3 or 4 times what it costs. If you're wrong, you're wrong by 3 or 4 times.
But then, over 2 or 3 months at 10%—let's say it slows down, and we don't actually do 10% for a very long, sustainable period. Let's say it's 5% to 8%. Five to 8% compounding day over day for 3 or 4 months, which is the lead time to bring compute online—and that's aggressive in itself—is 100 million users. So then do you buy compute for 100 million users?
Those are the types of questions I'm wrestling with. If you're wrong, you're very wrong. You get charged 3 or 4 times.
What about if you zoom in on the individual user and the cost to serve them on a day-to-day basis? Do you have a sense of the cost per week? I don't know what the right metric is—how much does my using Manus cost in inference per day, or something like this? Do you have a sense of that scale? What is that scale?
One thing I think is to our advantage is that we've figured out how to serve the product. No matter how you evaluate it—whether it's A/B tests, internal evaluations, or tracking engagement across users who might be on one model or the other—we're able to deliver the same performance as Opus 5, which is now, I guess, dating ourselves. Opus 5 is frontier-level intelligence.
We're able to serve the same engagement rate, the same A/B-test performance, and the same internal-evaluation performance, but at a very low cost that is actually very affordable. We're running this program where every user has the product for free, and our goal is to deliver it at an affordable price. I'm not going to commit to free for a lifetime for now, but it is my personal goal to deliver this product for free to everyone for a lifetime.
It’s just hard infrastructure work. To give one example, if you use frontier-level APIs from some of the main providers, you’re taking a blanket cost on a certain request and on all of the requests that might be needed to power that product for that month. But a lot of the work that happens through proactivity throughout the day doesn’t actually need to finish in hundreds of milliseconds. It needs to finish in minutes or even hours.
There is batch work that consumes a lot of content and can be served with deployment shapes that are 3×, 5×, or 8× more efficient, with the same underlying compute. When you customize these inference deployments to perfectly shape the data and the workloads that you’re serving, you’re able to find 30% here, 5× there, 6× there, and 10% here. All of those compound to a rate where we’re able to serve at a very low cost.
How do you think about solving the bigger problem of how far ahead to buy? If I think about this at true scale—if you get to a scale of 1 billion users or something like this—how much new compute demand do you think this will represent? Code generation has obviously created an enormous amount of demand, but ground us in some sense of scale, whether it’s per user or some other way of understanding how much new compute this will require.
I’m just saying it’s going to be a lot. If you ignore compute and just think about how many tokens are flowing through the platform, the breakout products from a couple of months ago in the code-generation space required the user to prompt them, then they would go and run something, come back to the user, ask for something else, and then the user would send something again. A lot of that is background work, so quite a few tokens are consumed there.
But Manus has the ability to wake up and sleep at any moment during the day. That might sound a little odd, but its architecture enables it to do that. If you have a meeting that you’re running late to and need to order an Uber, or if the Uber has been ordered and the user isn’t showing up, Manus can help the user through those moments.
Proactivity is going to continue to expand over time. There’s this big compute buildout with the earlier breakout products in the AI space that are just scratching the surface of productivity or background work. Now we have something that is almost natively proactive. It’s a smaller subset that is actually interactive.
I think the amount of compute that’s going to be needed is honestly going to be orders of magnitude more than what we thought we needed. With more proactivity comes more care for the user and more time to think about certain things that could go wrong. There’s just so much happening under the hood.
Maybe Instinct wakes up at 6:00 a.m. because it knows that you wake up at 7:00 a.m. It scans everything and makes sure that everything is ready for the day. Then it realizes, “Actually, now is not a great time,” so it goes back to sleep or does one thing in the background without notifying the user.
Then it realizes that something is coming up at 4:00 p.m. and that there’s value it could provide. Maybe the user doesn’t even know how to interact with Instinct in that way, but Instinct thinks it’s well-meaning and valuable to the user. It might wake up at 4:00 p.m., do the task, and contact the user. The user could then lean into it and complete that task.
You can see how much work is happening in the background. With these coding products, you don’t really have that much background or proactive work happening. I don’t mean to quote a number here; I’m just saying that the shape of the product and the workload that’s going to be run means we’re only scratching the surface in terms of how many tokens will be needed.
What do you think of Muse? What do you think of the product?
I think it’s a great product. I was playing around with it for a little bit, and I think it’s interesting. It’s a different take. You can say that some of the underlying architecture might be similar, but I think it’s fundamentally different.
Instinct is meant to be simple and very easily accessible. It has the soft qualities we were talking about earlier: really thinking from the user’s standpoint about what’s important and what’s not important, and how to make the task easier and easier to read. Muse is more like a new application and a new interface. Maybe over time we’ll have an application that also delivers on a different set of tasks, but I think it’s a great product.
I spend very little of my time thinking about competition and other players in the space. You walk outside, go to a nearby café, and think about how many people in that café are actually using AI in the way they imagined they would or the way they want to be using it. I would say very few. You go to other countries or other cities, and it’s certainly even less the case.
I think it’s still an open space and an exciting time. It’s an interesting game that’s going to be played and rolled out over the coming months and years. I’m just focused on building the best product experience.
What have been the blunders so far? What’s gone wrong, and what have you done about it? I’m sure more things will go wrong. This is going to be an explosion of emerging properties and mistakes, and the end state is going to be really high. How do you think about the things to guard against proactively and the things to react to? Talk us through the darker side or the harder side of building this.
I think it’s important never to be reactive and always to be proactive—to look ahead at what new surface areas are being introduced and what new risks might arise. This is the reason why we ran the early-access invite program from the start. An early version of the product did not have firewalls in place, active monitors, and many other pieces of infrastructure designed to get ahead of these issues and proactively secure an agent with that much access to existing systems.
There was an early version of the product that had some of these qualities, or some of these mistakes, and we addressed them. We went above and beyond and didn’t just patch the problem. We built a different system to systematically solve these types of problems.
The key here is that the user should always be in control. The user should always be in control of their data. The user can share as much as they’d like or as little as they’d like, and if they ever change their mind at any point, they can always retract access to those services.
It makes me wonder what the future of security looks like. Even if you become the best in the world at this—which maybe you’ll have to become—you’re going to have so much information and context on so many people. Security is a big problem across the world. All these great hacking examples that we’ve studied show how wild these things can be, and the capabilities are going to get stronger.
Do you have a general philosophy? I’m curious for you to riff on the future of security and safety and guarding against this. In early technology revolutions, there have always been enormous hacks and data breaches. It’s hard to imagine that this technology revolution won’t have them, too. How do you think about this and the responsibility of providing safety, given how much you’ll know about people?
It’s the most important problem. I think it comes down to building security and safety into the core values of the company and its people, and into the way that you build the product.
This is topical because anytime you look at the past 20 or 30 years, whenever a new breakout or consumer experience has been revealed, there’s always been immediate backlash: “Whoa, this is different and confusing.” Different gets conflated with unsafe, and we’re seeing some of that now.
You also have to stay true to the principles that you hold. The user is always in control of their data. They should never feel out of control. Then you have to be proactive about the systems that you put in place to get ahead of these types of things.
If I call out one example, there was a hallucination case. Language models hallucinate all the time, but with a product like this, you don’t want a language model to be hallucinating.
So we put in place a more systematic solution that will detect, before an action is actually taken—before a thinking trace is executed as a tool call, before any action might be taken—that it is validated and scrutinized by something decoupled from the same incentive system as the underlying agent.
It’s like a watchdog.
Yeah, a filter that’s able to find, “Hey, this proper noun was just generated out of nowhere due to some sampling error in the underlying model.” It’s very easy, in hindsight, to capture these types of mistakes.
These systems are hardened. We have world-class security teams that are constantly working proactively to find harder and harder adversarial cases and edge cases here and there. It’s becoming rarer and rarer over time.
One subtle part about building a platform like this is that it’s getting better over time as we continue to do more adversarial testing and become more creative about certain edge cases. These models are just getting better and better over time at being robust to these types of attacks.
8. What’s next for Instinct and personal AI
On the other side of the ledger, everyone always shows that beautiful visual of each generation of the iPhone, and you can see it getting better and more refined over time. What’s that arc for you? It’s so interesting because it’s not an application. It’s not a device. It’s an interaction through existing communication channels: WhatsApp, iMessage, and so on. What are the things that you’re adding, and envision adding over time, that will make the platform and the product more powerful than it is today?
The product experience right now is very simple—very, very simple. Simplicity is one key that we focus on, but I actually think that it becomes even simpler over time.
We may roll out an application in the near future, but I actually think that we’re going to trend more toward a simpler interface. Do you even need to open the iMessage application, type in a piece of content, send the message, and then look at the response afterward?
There’s a certain subset of users that interact with Instinct only through voice. More than 90% of the messages they send to Instinct are primarily through voice. I even have this action button—the action button on your phone—and it’s paired so that I can click on the action button and say, “Hey, say hi to Patrick in 2 hours from now. I think you have his email address; you can go send him an email.” Then I can just send it, and it goes and sends the email. I don’t even need to unlock my phone, open the iMessage application, and type the message in.
You can even imagine real-time voice with voice recognition that understands what your voice sounds like and has the discretion to know when you’re addressing it and when you’re not. You could have an AirPod—not a new type of AirPod, just an AirPod, because it’s good enough—and it’s just on.
You might be going on a hike, a bike ride, a walk, or a run, and you’re just catching up: “Hey, this contract needs your review.” “Okay, I need to review that.” “Hey, this news article just came in, and here’s the headline and the takeaway. This new project was released, and you should take a look at this.” You’re just hearing it in your ear: “Okay, got it. Got it. Put that on my calendar. That’s 15 minutes.” “Okay, yeah, that’s not important, so go clear that part out.” “This person needs my help. I know the answer to that, so just tell them that it’s this.”
It’s going to be much, much easier over time. I think there are long-term and short-term considerations here. In the long term, I think the interfaces become much, much simpler.
In the short and medium term, maybe there are more expressive interfaces to be able to share information. We have a files feature that enables Instinct to send effectively entire sites—full web applications—to show much more, whether it’s a trip itinerary, a wedding plan, or something that requires much more creativity and surface area. It can generate those on the fly and show them to you.
There’s something interesting there when you look at the last 20 years of application building or product building. Anytime you needed a new interface to showcase some piece of information, you needed to build an application for it. It’s this long software development life cycle to produce the application, ship it out to people, get feedback—“Hey, I wish it looked like this”—make the improvement, and ship out a new version. This is on the order of months.
Over the last 20 years, all these applications have been built for so many different purposes and so many different things. Now the consumer is fatigued. Every time I need something, I’m asking, “Is there an app that does that?” I don’t need that. I think all of that collapses down in the future. I think all of software is going to collapse down into a single, very easy-to-use interface.
Don’t get that confused with capability being limited. I think capability is going to expand.
What’s the craziest thing you can imagine, capabilities-wise?
When we move away from daily tasks—“I want to do this. Can you do it?”—and that friction goes down to zero, I think it will start pursuing higher-level objectives that are aligned with the user.
The user is able to describe not just, “Hey, can you track this workout for me? I did this, and I performed here, and this is how I did,” but more, “Hey, over the next 3 or 4 months, can you work with me to make sure that I hit these certain goals?”
A lot of people are starting to experience or discover that today. You want to gain this many pounds, lose this many pounds, or hit a certain mile time. Being able to specify objectives and goals, and then work with it to achieve those goals and objectives, is powerful.
I see small businesses being run that are native to Instinct itself, meaning the entire business is run on Instinct. The entire back office is functioning on top of Instinct, and even parts of it are fully autonomous.
You can talk about high-level objectives: “It’s in this certain area. Make sure that this inventory level doesn’t drop below this and doesn’t go above this.” Now it’s not saying, “Hey, please order this.”
Static.
Yeah, it’s, “Pursue this higher-level objective, and use the tools and devices that you have available to accomplish that.”
I think that’s the future. There are so many objectives that you could state and have it pursue for an unbounded amount of time that we could brainstorm here. But I think the meta-level picture is that moving away from individual use cases toward higher-level objectives is where interaction will progress.
Do you want it to have a personality that’s distinctive for each person? I’m thinking now of the movie Her, where there’s this relationship that forms between this omniscient, omnipotent agent and the user. What do you think about that aspect? I think of Instinct as this very seamless, almost quiet, extremely capable, reliable thing, but not as having a cheeky personality or something. How do you think about that component of the product?
Well, you mentioned Her, I guess, comically. Relationship-building, or anything in that area, is certainly not something that we want Instinct to do, pursue, or develop with users.
If I describe more about what it should act like, feel like, and represent within a person’s life, it’s almost like a socially aware operator that knows, no matter what room they’re in, what’s best for different people and what the best interaction pattern is. It learns that over time.
It’s possibly one of the most customizable apps ever, if you even call it an app, because it’s able not just to serve different versions of itself, but to evolve over time. That’s why we put so much care into the software pieces. It picks up on when the user didn’t like something being said in a certain way, or if there’s a lower response rate because there’s too much text or because they don’t want to look through a huge file to understand what you’re talking about.
It’s going to learn those things over time and just become easier and more delightful for the user to use.
I think we're not trying to impose certain experiences onto users. We want to solve this from a higher level: the ability to adapt to exactly what the best communication style and task-execution style are.
Why is it called Instinct?
I liked Instinct as a name for so many different reasons. I think the main thing is that Instinct should not be this playful thing that you might bully once in a while and look down on, or this thing that only takes the dirty work off your plate. I think it's this new, creative, exciting, competent actor, where there's mutual respect and trust. You feel safe, and you trust that it has your back—that it's intelligent, competent, and socially aware enough to know how to act in certain situations.
It's a bet, but I like that Instinct isn't named after someone's name or something like that to try to personify it as a human. It's more of—I don't know. I don't even know what an instinct is. I developed my feel for what it is, and for what the brand is, through just using it and thinking about what it should feel like.
I think it's a benefit that every user is able to come into it with a fresh slate. There's no prior idea in their mind about what it looks like or what it's named or anything like that. It's purely through the product experience itself.
Who are your enemies and allies? I can imagine that on WhatsApp or something, you could get shut off. That's a Facebook product, and it's a key channel for you. iMessage is an Apple-controlled product. Who are your friends? Who are your enemies? How do you think about some inevitable competitive realities here?
I would focus less on whether Instinct is an iMessage app or Instinct is a WhatsApp app, and more on, from first principles, what interfaces users trust today, what they're familiar with interacting with today, and how we can deliver the experience of Instinct through the channels they already interact with every day.
We're not pinned to iMessage. Actually, more than 50% of our traffic doesn't run on iMessage itself, but users who are comfortable with iMessage are very familiar with it. I think it's more about thinking about what is most practical for the user. We're going to shapeshift as we see fit toward delivering that very simple, easily accessible, zero-friction experience that it is today.
I know your latest round is something like $1 billion at roughly a $10 billion valuation, with some of the best investors in the world—Zoya, Benchmark, and Kotu—as the leaders of the round. How do you think about the future capital needs of the business alongside this? And how did you pick the partners that you picked?
I'm very lucky to be able to work with some of the most supportive partners around the table, who have been through different but similarly shaped technology transformations in the past. I would focus less on the numbers and more on the demand in the space—the demand for a product like this, or the demand for an experience like this. It really is delivering on that AI experience, or that product experience, that I think we've all been waiting for.
It's also capital-intensive. We're a new company, and we're trying to create something that's hopefully going to be distributed to billions of people on the planet every day at a very affordable cost. When I say affordable, I mean that compute costs are high, and we're trying to deliver it in a very affordable way to our users.
There has to be some bootstrapping, especially if you think about the business model that we're chasing after. We could very easily just say, "Hey, it's going to be a certain subscription, and everybody on the platform needs to pay $100 a month," right? There are shorter-term rewards that we can chase after.
Or we could raise a little bit of capital and use that toward what venture capital is meant for: taking certain calculated risks. If there are small speed bumps that require a little bit of capital to get ahead of, we can use it to prove transaction volume or to prove the experience across certain industries. That's how you escape these local optima of the subscription plan or other things like that.
When I do these, I ask everyone the same traditional closing question: What's the kindest thing that anyone's ever done for you?
I've been able to surround myself with and build such great relationships with so many different people, whether it's in the industry or outside the industry, at work or in my personal life. I'm very grateful to be around them, and I certainly hope I can extend the same level of kindness and thoughtfulness toward them as well.
I know I'm not answering your question. It's more on a higher level. It's just the qualities of things that I notice and really appreciate.
No, you're building a fascinating company, maybe the most fascinating company in today's AI ecosystem. Thanks so much for your time.
Thank you so much.