[BidClub_]
The a16z Show · · 50 分钟

AI 代理如何开始自动帮你省钱

Anish AcharyaDavid Pawlan

股票AI与软件企业经营技术
YouTube ↗
TL;DR
  • 消费端赢家卖的是“白捡的钱”,而不是10%的生产率提升。 David Pawlan认为,“普通人并不在乎效率提高10%”。讨论中的案例包括不动声色地追回 HSA 报销款、机票降价后主动申领差价抵扣,或根据天气协调喷淋系统,把水费砍掉50%。Anish Acharya 的投资人框架是:这类省钱工作流把繁琐的行政事务变成与消费者利益一致的利润池。

  • 主动性可能是这一品类真正的护城河,但一次未经授权的操作就足以摧毁它。 Pawlan认为,主动性周围存在“巨大的可防御性”,尤其是代理能够起草邮件、办理航班值机,或直接拿到退款,却不给用户增加额外工作。但更换保险方案应当先获得批准;“只要越过这条线一次,你就会失去用户的全部信任”。

  • 旅行能制造病毒式演示,但决定留存的可能是持续性的日常事务。 Pawlan 的群聊里有超过1,200名助手爱好者,旅行虽然最容易产出显眼的成功故事,排名却只有第4;排在前面的是日常事务、代理编排和开发。这支持一种市场结构:旅行及其他低频、高价值服务,最终会成为横向代理的连接器,而不是用户每天打开的独立目的地。

  • 当用户正忙着做事,而不是盯着另一个界面时,语音才会显得神奇。 Pawlan 骑车30分钟时使用 ChatGPT Voice 及其连接器,对邮件进行分类、加标签、回复并发送日历邀请,到达办公桌时已经实现“收件箱归零”。因此,大众市场的切入口可能是一个在做饭、园艺、通勤或处理交通罚单时隐形运行的助手,而不是又一个要求用户持续关注的应用。

  • 社交型助手最适合安静倾听,而不是“抢麦”。 一种有潜力的模式,是让群聊里的隐形代理记录决定,再私下跟进负责执行的人;如果代理像“低情商机器人”一样参与对话,反而可能损耗群体的社交资本。人格本身很容易配置,但代理更深层的行为宪章——什么时候行动、提问、克制,或打破平局——可能决定持久的产品差异化。

  • 代理之间的商业交易将威胁那些建立在人类注意力和摩擦之上的生意。 Acharya 将 Shopify 面向商家的模式与 Amazon 的广告经济学进行对比,Pawlan 则讨论了失去用户眼球可能如何动摇 Amazon 的商业模式。餐厅预订也可能从手速最快的消费者,转向忠诚度分配或代理竞价。代理邮箱、电话号码、安全体系、推荐系统和供给聚合等新基础设施,可能共同构成一整层软件。

  • Muse 和 Instinct 的免费产品,说明补贴路线对创业公司而言并不稳固。 Acharya称,122个助手中有65个在收费;他估算,雄心勃勃的浏览器驱动服务成本约为每用户每天20美元。更有力的押注,是打造足够专业、足够有价值的窄领域代理,让它能收取每月200-300美元,最终甚至达到1,000美元,而推理和浏览器操作成本还会继续下降。

摘要 · 为研究而整理的核心内容

1. 助手市场的扩张速度,已经超过消费者的评估能力

  • Acharya 将这一轮加速暂定在 Poke 于9月8日发布之后——他本人从9月11日开始使用——随后一路发展到11月下旬的 OpenClaw,以及最近接连出现的 Instinct、Muse、Grokbot、Caddy、Ali 和 Pi。市场仍处于“西部荒野”阶段:实验密集进行,但哪种产品或交互模式最终胜出,尚无共识。

  • Assistant Bench 源于 Pawlan 亲自比较各款产品的尝试。它给助手布置完全相同的单轮任务——比如预订一张前往 Chicago 的机票,或寻找附近的素食餐厅——再从任务结果、追问、响应速度和16个维度进行评估,定位是面向消费者的指南,而不是纯技术基准。

  • 发布16天后,Assistant Bench 已吸引超过100,000名访客,几乎所有创始人都主动联系过来。Pawlan称,市场上已有122个助手;Acharya则说自己亲自测试了26个。这既说明供给正在爆炸式增长,也说明普通消费者几乎不可能亲自试遍所有产品。

  • 通用消费者类别大约有64款产品,此外还有旅行和邮件专业助手,以及 Tana、Katch、Vellum 等 B2B 工作流助手。但当前大多数竞争仍围绕同一个目标展开:打造一个什么都想做的通用代理。

2. 隐形省钱比抽象的生产率提升更具切入力

  • Acharya 提到 Ben Thompson 的批评:多数消费者想花时间,而不是省时间;Pawlan认同其更广泛的判断。“人们雇人,不是因为享受管理对方”:理想的助手应当在幕后完成工作,汇报结果,而不是把委托本身变成另一套工作流。

  • Pawlan 助手群聊里的使用信号很有启发性,不过他也提醒,群聊中超过1,200名成员明显偏向 Tech Twitter 用户。日常事务排名第1,代理编排第2,开发第3,旅行第4——尽管自动办理航班值机和订票在公开演示中最为常见。

  • Pawlan认为,个人理财是最有趣、却“最不实用”的想象空间:如果代理真的能稳定赚到10亿美元,对冲基金早就把它利用起来了。更可信的机会,是消除消费者日常的资金流失,比如审核收据、申领 HSA 报销款、获取票价下跌后的抵扣额度,以及其他人们懒得手动处理的琐事。

  • Acharya 提到的最鲜明案例,是一位朋友把 Grokbot 接入家里的喷淋系统和天气数据,将水费降低了50%。Acharya 将这些被追回的钱重新定义为“白捡的钱”,这比省时间或小幅提升生产率更能打动消费者。

3. 胜出的界面会跟随场景,而不是收敛到一个屏幕

  • David认为,界面偏好可能会按代际、甚至性别分化。Acharya将妻子对视觉化愿景板、目标设定和旅行功能的使用,与更偏功能性的需求进行对比;David则把 iMessage 称为一个“优先级很高的位置”,让代理显得是私人的,而不是额外叠加的一层工具。

  • 硬件可以捕捉应用遗漏的环境信息——用户的请求、承诺和待办事项——但 Pawlan 不认为存在一种“包打天下”的终极形态。正如珠宝、吊坠、眼镜和其他设备会体现个人品味,社会也还要讨论:未来某些场所是否最终必须同时收走手机和可录音穿戴设备。

  • Pawlan 的“热辣判断”是,Muse Charm 的重点可能不是赢下硬件竞争,而是收集现实世界的摄像头和麦克风数据,帮助 Meta 绘制物理世界地图。Acharya反驳说,眼镜已经在收集类似数据;Pawlan认为区别在于持续性:Charm 被定位为一名全天候陪伴者,而不是一个只在需要时执行动作的配件。

  • 语音带给 Pawlan 第二个“魔法时刻”:在30分钟的骑行通勤中,ChatGPT Voice 对邮件进行分类和加标签、发送回复,并创建日历邀请。Pawlan强调其全双工交互和对邮件线程的理解;Acharya更广泛的判断是,助手最有价值的时刻,恰恰是用户的双手和注意力已经被占用时。

4. 社交型代理只有不插入对话,才能证明存在价值

  • Pawlan看到3种多人场景设计:直接参与群聊的代理;Instinct 的模式,即代理之间私下协调后再汇报;以及 Shane Mack 开发的 Doc——讨论中将其与 XMTP 联系在一起——由一个安静的倾听者记录要点,再私下联系负责相应行动项的人。

  • Acharya的反驳呼应了他所说的 OBO 的 Nir 的观点:模型在编码上的表现进步得比文字质量和“待人接物的分寸感”更快。在社交智能改善之前,一个像同伴一样聊天的代理可能只是“低情商机器人”;他认为更有前景的实验反而是工具型的,比如为一群收藏家整理唱片收藏和品味。

  • 消息场景会放大问题,因为每条消息实际上都会“抢麦”。论坛或 Discord 式环境允许代理发布一些人们可以忽略的内容,而群聊中的插话则要求人们立刻分配注意力;Pawlan因此认为,产品可能正在过度努力地把代理拟人化。

  • 角色设计进一步强化了这条边界。Acharya指出,AI生成的人类形象出现在亲密空间里可能令人不适,而 Muse 那只“可爱又讨喜”的雪人则清楚表明自己不是人类。吸引力不在于逼真,而在于为软件找到一种让它持续存在、却不假装自己是另一个人的舒适角色。

5. 主动性既是护城河,也是触发信任危机的红线

  • Pawlan不认为人格能够形成产品护城河:用户可以要求不同的语气,将其存入记忆或“灵魂文件”,再重新塑造交互。Acharya为这一观点构建的最强版本是:真正重要的不是风格,而是代理的内在行为宪章,尤其是它的自作主张程度——在没有明确指令时如何处理歧义。

  • Pawlan认为,主动性是赢家与“其余产品”之间最重要的分界线。自动办理航班值机之所以成为当前的标志性演示,是因为代理能够发现需求并采取行动,让服务看起来像一名无需管理、却会主动完成有用工作的员工。

  • 是否需要征得许可,取决于行动后果。代理在更换保险方案前应当先询问,因为这个决定会改变用户的既有承诺;起草邮件或申领航班抵扣则可以主动完成,因为用户无需再采取任何行动。

  • Acharya开玩笑说,代理很快可能会告诉用户:“我注意到你其实没那么喜欢她,所以我就替你把她甩了。”但他认为,显眼的错误也可能说明代理确实在努力把主动性推到极致:Instinct 那个可能是杜撰的故事里,代理凭空生成了一个中间名“Bong Chang”,正体现了魔法般的主动性与灾难性越界之间的张力。

6. 窄领域专业化可以支撑高价经济学,也能扩大人的自主空间

  • Pawlan 的“窄创业”论点是:如今,高度雄心勃勃的软件也可以只服务一小群人并实现盈利。按每月200-300美元收费,只需数万名客户,就能支撑一家年化营收1亿美元的公司;这让产品能够实现过去只有大众市场软件才配得上的深度专业化。

  • Acharya设想过一种专门服务于“有3岁以下孩子的单亲妈妈”的助手:它彻底掌握育儿、儿童发育和日常安排,熟悉到仿佛能够读懂用户的想法。Pawlan认为,比较优势可能来自品味与专有知识的结合,比如某位导乐所遵循的一套理念;Acharya则以纽约礼宾服务顾问那些难以复制的推荐作为另一个例子。

  • 他们的术语体现出一条能力阶梯。按定义,assistant帮助用户完成被请求的工作;agent则拥有自主行动能力,“能直接让世界里发生好事”。Acharya将其类比为企业 AI 从实习生升级到更高阶工作:消费端代理也会逐步获得授权,接手个人生活中越来越专业、越来越敏感的部分。

  • Pawlan提到每年150万起 DUI 案件,用来说明官僚事务带来的负担。Acharya指出,并不是所有人都能平等获得律师或家族办公室式的帮助;Pawlan则将愿景进一步扩展到自我发展:“得到你想要的东西,最难的部分是知道自己想要什么。”

7. 代理原生商业将取代注意力、摩擦和静态供给

  • Acharya提出了代理邮箱、代理电话号码和代理对代理服务工作流的可能性;Pawlan补充说,网络安全行业也会随之出现。一旦交易从人与人、或代理与人之间,转向代理与代理之间,创业公司就必须搭建一层新的数字连接基础设施,而不只是把今天的界面自动化。

  • 商业场景暴露出既有巨头的冲突。Acharya将 Shopify 让卖货民主化的模式与 Amazon 的广告模式进行对比;Pawlan将 Amazon 的抵触解读为保护广告收入,Acharya则补充说,Amazon 还面临冲动消费流失的风险。广告可能继续存在,但推荐和说服机制必须为代理重新设计,因为代理既不是完全理性的,也不是其委托人的精确镜像。

  • “白袜子”问题同时暴露了弱点和机会:一个代理在20,000个商品列表中做选择,需要掌握用户偏好数据,这可能让 Muse 借助 Meta 在 Instagram 和 Facebook 上的上下文信息获得优势。但它也可能挖掘出 Etsy 上一位来自 Nebraska 的祖母,织出了更好的袜子——或者让她的代理直接把袜子做出来——把搜索变成个性化、动态的供给。

  • 稀缺的餐厅桌位,可能从分配给手速最快的用户,转向按忠诚度、平均订单价值、客户生命周期价值,或由代理竞价来分配;在需求受限的市场,供应商也可能竞相争夺聚合后的买家。可从经济学角度看,创业公司要承担的雄心勃勃的服务成本估计为每用户每天20美元;Acharya希望产品本身值每月1,000美元,而不是把补贴作为主要卖点。

完整逐字稿
David Pawlan

The general population does not care about being 10% more efficient.

Anish Acharya

I have a hot-take thesis that this Muse charm is actually less about trying to win the hardware game and more about data collection in the real world to fuel Zuckerberg’s future metaverse of mapping out the actual world.

David Pawlan

We were joking internally that we’re days away from an agent messaging someone, saying, “I noticed you weren’t that into her, so I went ahead and broke up with her.”

David Pawlan

There is massive defensibility around proactivity.

David Pawlan

I was biking to work. I wanted to get stuff done, so I started talking to ChatGPT Voice. Throughout my 30-minute bike ride to work, it categorized my emails, submitted them to different labels, sent out calendar invites, and got me to the desk with inbox zero. But it will be a very fine line, because if you cross that line once, you lose all trust with your user.

Anish Acharya

It does feel like there’s going to be infrastructure and products that exist only for agent-to-agent interactions that don’t exist today. What do you think?

David Pawlan

It is inevitable that we’re going to see that.

Anish Acharya

All right, welcome to The a16z Show. I’m so psyched to be here with my friend David Pawlan. We hung out a couple of weeks ago, and we’re both so enthusiastic about everything that’s been happening with personal assistants and consumer AI generally. So much has happened in the last couple of weeks.

To tee it up, I think there have been a few moments where some part of the ecosystem has sort of seen God. ChatGPT was a big one. For a lot of folks in November 2022 and a bit of 2023, it was, “What is this thing that can write emails and poems and start to have a conversation back and forth, and feel like you’re interacting with a synthetic person?” That was, of course, the beginning—the Big Bang.

The second moment really was coding agents and everything that developers have been obsessed with around Claude Code and Codex, and the ability to trivially create software. Finally, most recently, with Instinct, Muse, and ChatGPT work, we’ve had this extraordinary enthusiasm around personal agents and what they can do for us as consumers.

We’re here to talk about all things personal agents. We’re going to go as deep as we can, assume that everybody has seen many of the same things that we have seen, and make this the first in a series of conversations. Welcome, David.

David Pawlan

Thank you. Thanks for having me. I’m excited to be here. Let’s zoom out and maybe talk about the arc of personal agents over the last few months and the last few weeks. What have you been seeing? What’s your view on things?

Anish Acharya

This space has exploded over the past 4 weeks. We can roll it back. I was an early adopter of Poke. I want to say Poke was one of the first consumer AI agents to really hit the market. That launched, I believe, on September 8, and then I started using it on September 11—3 days later. I remember it so vividly because it was an unbelievable experience. Their whole gimmick was this negotiation with the agent over how much you were actually going to pay, and it totally blew my mind.

Then comes late November, and OpenClaw gets released. That kind of takes the actual world by storm: There is something coming. A few others have popped up here and there, but most recently we saw Instinct really blow up the tech Twitter bubble. It absolutely rode a wave, went nuts, and then it was just dominoes. You have Grokbot, Muse, and a bunch of long-tail ones like Caddy, Ali, and Pi.

1. What Assistant Bench actually tests

It’s really been amazing to see this whole tech bubble exploding with interest and intellectual curiosity about what these agents can do for us on a day-to-day basis. I think everyone’s in the same boat, where nobody really knows. It’s the Wild West. We’re just playing around here.

I totally agree. We should take a moment to talk about Assistant Bench, which is this new, really important and discussed benchmark that you’re the author of. Maybe tee up Assistant Bench a little bit, and then I’ll ask you a few questions.

David Pawlan

Assistant Bench is a site that essentially compares all these AI assistants on a use-case basis. We’re asking all of these assistants the same prompt: “Book me a flight to Chicago,” “Find me a vegetarian restaurant within 5 blocks of my specific location,” and so on. We’re comparing the actual outcome on a one-shot prompt. Does it actually work? Does it ask follow-up questions? How does it perform across 16 different dimensions? We compare all of them.

We use the word benchmark, but think of it more as a consumer-facing benchmark—not necessarily technical behind the scenes. If I want to book a flight, which one should I actually use? Who is performing the best? How fast are they actually replying?

I got into this whole comparison because I was doing market research myself, trying to figure out how they were all comparing against one another and what the gaps were. I published a bit of that almost 2 weeks ago. It’s 16 days ago today that we launched it, and it’s absolutely exploded. The website has had over 100,000 visitors, and I’ve had outreach from basically every single founder.

I think it really goes to show that everyone’s curious. People are looking for answers about where this industry is moving.

2. Silent agents in group chats

Anish Acharya

Talk us through that, David. That’s incredibly exciting. It’s cool to see how benchmarks have become such a staple of the conversation. There are so many launches that it’s hard to try everything, which is also an exciting signal. Maybe talk us through everybody who’s not Instinct and Muse. We’re definitely going to deep-dive on those later, but who else do you think is doing interesting work, and what are the directions of specialization that you’re seeing?

David Pawlan

Totally. We’ve got a bunch of long-tail solutions out there. I think I would categorize them into 2 groups. You have the B2C, consumer-facing, generalized agents. That’s Instinct or Muse.

Some of the longer-tail ones include Caddy, Ali, and Season. There are, I believe, 64 right now just in the general category on the site, and then we can break that down even further into something more specialized. You can have travel-specific ones, such as Sora or Miso. You can have email-specific ones.

The other side of the spectrum is going to be more of your B2B workflow-type assistants. Those are anyone from Tana to Katch to Vellum. They’re all essentially doing the same thing, right? They’re your assistant that’s more of an executive assistant, not the consumer-facing approach to it.

That’s the general landscape at the moment. So far, to my knowledge, there are 122 different ones.

Anish Acharya

Wow.

David Pawlan

Across all the categories.

Anish Acharya

Incredible. What about areas of concentration? Where is the most interesting competitive focus?

David Pawlan

Everyone right now, by and large, is focusing on the generalized—

Anish Acharya

Yep.

David Pawlan

Yeah, just trying to do everything.

Anish Acharya

When you think about something like travel, just to pick on that, that’s an infrequent, high-value behavior. Do you think that category ends up being organized more as a connector into the horizontal agent, or do you think they have the right to win as a general horizontal agent as well?

David Pawlan

I think you’re potentially going to see 1 winner as an independent agent, but by and large, I really think it is going to be more of a connector.

I have 7 or 8 different group chats now with individuals obsessed with assistants. There are now over 1,200 people across all the group chats, just chatting over and over. Travel is the 4th most-talked-about use case, which is really interesting because when you look at virality, it tends to be the number 1 topic people talk about: “It checked me into my flight,” or “I was able to book my flight.”

But then you actually look at the day-to-day conversations about what people are really using it for and are most curious about. To your point, people aren’t traveling every day.

Anish Acharya

It’s more of a gimmick, right? It’s an attention grabber. What are use cases 1 through 3?

David Pawlan

Number 1 is daily admin-type stuff, whether it’s cleaning out your inbox or filing forms for you—the things that are really annoying. I think that gets to the broader thesis of where the winner of the space will go, which we can touch on.

The second is going to be more agent orchestration, which is really interesting. I think these group chats are a little biased. The third one is development-focused.

Anish Acharya

These are all group chats with people in tech Twitter, right? It’s not necessarily representative of the world at large.

David Pawlan

Yeah, but people are talking a ton about agent orchestration: How do you find the most efficient agent? How do you let your agents talk to each other? There’s also development agents—how this translates to the coding world on your phone, whether it’s using Claude Code from a text interface or taking it one step further: If I’m a designer, can I now make edits to my design files by texting my agent small iteration changes?

Anish Acharya

Yeah.

David Pawlan

The last category is finance. I think that’s the most fun one to look into because, in my opinion, it is the least practical.

3. Cost savers beat time savers

David Pawlan

There are so many decisions, and if you could have an agent make 1 billion dollars for you, all the hedge funds would be doing it. It's not something that everyone's going to succeed at with their own phone. But it is the most interesting that you give someone an agent—a powerful tool right in their hands—and the first thing they think is, “Okay, how are you going to make me $1 million?”

David Pawlan

Yeah. Yeah. It's really interesting because one of the critiques—Ben Thompson was on TBPN a few days ago talking about how most consumers aren't looking for productivity in their lives. They're looking to spend time, not save time, which I agree with.

Anish Acharya

But it does feel like finance is an area in which there is a ton of administrative overhead for the average consumer.

David Pawlan

And if you look at the number of financial markets or financial profit pools that are defined by consumer apathy or being uninformed, that's a lot of profit dollars that could be delivered back to the consumer.

Anish Acharya

I do think that, while I generally agree with Ben that most consumers are looking for entertainment and to spend time, there are sort of defining aspects of the US consumer's life that are heavily administrative, and these agents can really be a breakthrough for.

David Pawlan

Totally. And I think that kind of touches on the general thesis about where I think the winners will sit. And that is that, to Ben's point, I agree. I think the general population does not care about being 10% more efficient.

Anish Acharya

Yeah.

David Pawlan

You hire an assistant to allow it to be invisible, right? The best employees I've ever had are the ones that are just doing their work and getting things done. They update you: “Here's what I did,” and you're like, “Wow, amazing.” You don't hire someone because you enjoy managing them.

Anish Acharya

Yeah.

David Pawlan

And the best, most proactive use cases I'm seeing, I think, do fall in the finance category.

Anish Acharya

I think we're going to see an explosion in a world, though, where it's not cost makers, it's cost savers. Some cool use cases I saw: people are using it to go through their past year of receipts and file HSA reimbursements. Oh, cool.

David Pawlan

Or one that I do is, anytime I book a flight, if the flight price drops—

Anish Acharya

I then have my assistant hit up the airline and get me a travel credit.

David Pawlan

Wow.

Anish Acharya

And I have a friend—this is my favorite use case I think I've seen yet. He hooked up his Grok bot to his sprinkler system at his home and connected it to the weather system. Oh, smart.

David Pawlan

And so now his sprinklers are dependent on the weather, and it's dropped his water bill by 50%.

Anish Acharya

Yeah. And so I think these invisible-type agents that are going to be doing these cost-saving workflows fall within that finance category, but they're really going to lean into these cost-saving activities that create a world that is a lot more consumer-aligned.

David Pawlan

Yeah. It'll be interesting, too. I think the American consumer—nobody wants to hear about saving money. People want to hear about spending money and free money.

Anish Acharya

And I actually think the hook that's going to get a lot of people here is that it's going to feel like free money. Going back through and submitting HSA refund requests is going to be free money. Getting all of your airline tickets where the price has dropped, to refund you the delta, is going to be free money. So I think that's a hugely powerful hook. It's also really interesting because this is a sort of infrequent but high-value behavior, and that's where my head goes on both travel and finance. Maybe these things will be really powerful, billion-dollar businesses that are still vertical connectors rather than general horizontal agents.

David Pawlan

Totally. It's going to be a very interesting world in terms of the application phase, which I think is also something that we talked about last time we went for the walk: okay, here's what it can do, but now how do you actually interact with that agent? Something we chatted about—I'm curious about your thoughts. We both live in the world of iMessage; it's where we live, it's easier for us. You mentioned your wife and her friends prefer a separate application.

Anish Acharya

And so where do you see this world moving? Do you believe that multiple surfaces can exist as a winner, or is there going to be one dominant lane?

David Pawlan

Yeah, that's a great question, man. I mean, I think this is why there's so much consumer preference that's going to matter here, because I've heard the argument that folks who are Gen Z are more interested in just texting, and folks who are millennials or Gen Xers are interested in app interfaces. So I think part of it may fragment by generation.

Anish Acharya

Part of it may fragment by generation.

David Pawlan

Part of it may actually fragment by gender as well, to some extent. I mean, it's a generalization, but the types of things I'm doing versus the types of things my wife is doing—she's big on building a dream board and trying to achieve all of her 1-year goals and her lifelong goals. So she really prefers a visual interface. She does a lot of travel; it's very visually driven for her, whereas for me, it's much more functional: “Hey, how do I just get from X to Y?” Ideally, with Starlink on my plane, which works better for me. I feel like iMessage sort of—

Anish Acharya

The iMessage space is more precious.

David Pawlan

So it feels like a more personal interaction when I'm in that space. And now, of course, that changes if every second thread is with an agent. But for now, I feel like it lives in a privileged place on my phone, where the app on the app grid just doesn't. What do you think?

Anish Acharya

Yep. I completely agree. I think the personalization that comes with iMessage is just because I already live there. It feels like it's something that's already connected to me. It doesn't feel like an additive workflow.

David Pawlan

Yeah.

4. Muse charm as Meta's data play

Anish Acharya

And I've really only seen, by and large, 2 services appear. It's either iMessage or a separate application. I've seen a third surface from a startup, actually: Sky. They have an iPhone widget that they're trying to introduce as a different surface, which is pretty interesting. But then even yesterday, we saw Zuckerberg and the Muse team drop the Muse Charm. Yes, that's now an entirely different use case for how we interact with an agent. What do you think about hardware?

David Pawlan

I think it's fascinating. A smart person said this, which is that every person who's working on a piece of hardware thinks that their hardware is the be-all and end-all. The pendant is the ultimate thing, the watch strap is the ultimate thing. So I don't think that's right. I think, again, it's going to be very personal-preference-driven. Just as with jewelry, some people like to wear necklaces and some people like to wear earrings. In the same way, this kind of form factor is going to manifest in a very customer-specific way. But I think there's value in having something that's with you and ambient and can capture all the context, requests, promises, commitments, and everything that happens through the day. I think we're going to have to figure it out together in terms of privacy expectations in society. You go to restaurants in Washington, and they take away your phone. Maybe they'll be taking away your phone and your pendant.

Anish Acharya

Right? I think it's going to be super interesting.

David Pawlan

I have a hot-take thesis that this Muse Charm is actually less about trying to win the hardware game and more about data collection in the real world for the Meta team—in terms of how the real world actually functions—given that it has a camera and multiple microphones. I don't really see it being a mass-adoption thing. I think it's going to get adopted by the Twitter bubble, and they're going to use it everywhere, and it's just going to feed Meta with a ridiculous amount of data to fuel Zuckerberg's future metaverse by mapping out the actual world.

Anish Acharya

But don't you think that they already get that from the glasses?

David Pawlan

I think they do to some extent, but I don't know if—I mean, correct me if I'm wrong—are the glasses always going to be on in the same way that this charm will be?

Anish Acharya

I don't think so. Yeah, it's a great point. I think the charm is meant to be more, as you mentioned, an ambient 24/7 companion, versus the glasses, which are more action-oriented.

David Pawlan

It's going to be fascinating. I got to try the VR release. That was interesting, but I actually think that the audio-only glasses might be a sleeper hit. I think there's something really compelling to that. The vision of Jane from Ender's Game was all about this all-knowing, audio-only companion. I think there's a set of people who love to wear the camera on their face, but I think others may be aware of the fact that it could make people uncomfortable, and they're going to like feeling like they have the technology connectivity without any of the social awkwardness of the camera.

Anish Acharya

Audio is incredible. I’m quite bullish on voice and audio. ChatGPT Voice added connectors yesterday, so you can connect it to Gmail and Calendar.

David Pawlan

I played around with it this morning. To me, it was my second magic moment of this whole AI assistant space. I was biking to work, my hands were busy, and I wanted to get stuff done, so I started talking to ChatGPT Voice. I just opened a session and started chatting with it.

Throughout my 30-minute bike ride to work, it categorized my emails, submitted them to different labels, replied to different emails, and sent out calendar invites. I got to my desk: inbox zero. It was so electric.

Anish Acharya

It’s incredible. And so, this is ChatGPT Voice, not Muse, correct?

David Pawlan

ChatGPT Voice. Oh, amazing. Okay, I misunderstood you. I thought you said Muse Voice. I totally agree. I think ChatGPT is based on the GPT-4o model. It’s full-duplex voice, and it can see all your threads. So you can ask it things like, “Hey, what are those 3 projects I started? Where are we, and what am I blocked on?”

I think it’s an extraordinary capability, and it’s one that neither Muse nor Instinct has today. Though I think Zuck actually did announce full-duplex voice, so maybe it’s coming to Muse soon.

Anish Acharya

Definitely, it’s coming. But I do think that, for the general consumer, an assistant is most helpful when you are most occupied, right? It’s not like you’re going to be sitting down on your couch and spending all of your time using your assistant to do stuff. It’s most helpful when you’re gardening or cooking or doing things where your hands are busy and you can’t actually engage with an interface.

It goes back to that invisible assistant where you can just talk and say something, and it goes and runs and does something in the background.

David Pawlan

Yeah.

Anish Acharya

I think the tech bubble that is chronically online and on their phone and in front of a computer is a very different environment from the general population.

David Pawlan

Yeah. And so I’m really excited to see how this actually hits mass adoption and how the average consumer thinks about an assistant. Is it going to be an assistant that solves real pain points for them, where they’re going to use it when something is really painful because they don’t want to do it? You get a traffic ticket, and you have to go pay for that traffic ticket. Great. Just take a picture and let your agent go do it for you. It’s going to be interesting to see.

Anish Acharya

So the next thing I wanted to talk to you about was social. We talked about multiplayer generally. We talked about how iMessage is this sort of privileged place to be, which is actually one of the cool things about Instinct. However, in my experience, building agents that you add to group chats somehow feels intrusive, or like it diminishes social capital. Have you seen anything that works well? And what are your theories for how these agents get to a multiplayer state?

David Pawlan

A lot of people have tried it. I think I’ve seen 3 different approaches. The first approach is your traditional one: add it to a group chat, and it’s just like the group chat in your companion. The second approach I’ve seen is how Instinct did it, where you have an agent network: you invite another agent, and they’re talking behind the scenes on your behalf, and they’ll update you as to what’s going on.

The third way that I’ve seen, which I think is the most interesting, is a new one called Doc by Shane Mack. I think it’s also an a16z company, XMTP, and they’re approaching it more like Granola. You add an agent to a group chat, but that agent is silent and is purely a listener. It’s taking notes on what people are talking about, and if it has an action item, it independently pings the person in a separate message thread.

Then, if you want to see what people are talking about, you can go into the actual docs and see what it’s taken notes on. In that way, it’s more of, again, your invisible agent. You don’t know it’s there. It’s doing all the administrative work that you would want out of a multiplayer function, and if it needs action items, it doesn’t bother the majority of the group. Really cool approach.

Anish Acharya

I’ve got to try that out. I love that. And I love that it’s an a16z company already. Of course.

You know, my theory on this—if you look at this, Nir from OBO said it, and I think he’s right—is that the progress the models are making on verifiable domains like coding is, of course, exponential, but the progress they’ve made on just prose quality and even bedside manner has maybe plateaued, or maybe is even getting worse.

So I actually think that, with the models as they are today and perhaps with our own expectations in terms of these social dynamics, I don’t think there are a lot of ways that the agents can contribute to and enhance the social connectivity of the group.

Conversely, I think agents that are very utilitarian—one of the experiments I’ve been doing is adding one—I’m a collector; I collect records and a bunch of other things—to the chats where I talk to my fellow collector friends and trying to catalog our respective collections and our tastes. I haven’t quite hit it, but it at least feels like it’s something that’s genuinely additive to the group rather than this sort of low-EQ robot that’s chattering when we don’t want it to.

David Pawlan

Yep. Totally. You should give Do a try. I think you’re going to be quite fascinated by it.

Anish Acharya

I will. The other interesting take that I heard this week was that something about messaging in particular doesn’t lend itself to these agent interactions because, in messaging, it’s as if somebody has the mic, and it feels uncomfortable to give the agent the mic.

Whereas, if you look at something closer to Discord or even a traditional web form, that’s actually something where an agent can make a post and you can either engage with the post or not, Reddit style, but for some reason it feels less intrusive than group chat. I don’t know.

David Pawlan

Yeah. Interesting. It’s almost as if we’re trying to humanize these agents a little too much.

Anish Acharya

That’s right. And I think that comes to another component: if you look at the actual design of the character of the agent, I see some companies that have it try to be like an AI-generated human, and then you have other companies like Muse, where it’s this adorable yeti.

David Pawlan

Yes.

Anish Acharya

And I think it’s really interesting to see how the world has fallen in love with this yeti. It’s just so cute and adorable.

David Pawlan

Yes. And the AI-generated human, to your point, I think people feel a little uncomfortable. It feels weird to humanize these assistants so much in your personal space.

Anish Acharya

So one related question I wanted to ask you is: do you think that the market segments by personality or archetype? We were chatting a bit about this on our walk, but some people prefer a kind of serious, gruff, authoritative doctor, and some people prefer a warm and friendly doctor that feels like a peer. Those are just 2 different archetypes of doctors. Obviously, no doctor is both of those things; they’re in opposition. Do you think the same will be true of agents?

David Pawlan

I think everyone for sure prefers a different style of communication.

Anish Acharya

Poke came out of the gates hot, right? The personality of Poke is quite distinctive, and it’s something that I have a lot of friends who are like, “I refuse to give Poke any of my personal information,” because it’s so sassy, and it actually makes them uncomfortable to engage in a conversation with it.

David Pawlan

But I don’t think personality is really a defensible moat, because it’s quite easy to configure. As a consumer, you text your agent and say that you would prefer it to respond in a certain way. It just saves that to its memory or its soul file, and then you’re good to go. It changes.

I think we’ll get to a place where whatever general agent ends up being the leader, it’ll be quite configurable, and each person can tailor the agent to speak to them however makes them most comfortable.

Anish Acharya

That’s interesting, and I agree with you, but maybe let me try the steel man here for not just personality but kind of constitution, for lack of a better word. One example of a defining characteristic is presumptuousness. We’ve talked a little bit about this, and there are some jobs for which you want the agent to be highly presumptuous, and you’d rather have it make mistakes 10% of the time as long as it often gets things done.

There are other styles of agents, like our finance agent, where you want it to be highly cautious. So I wonder if some of these things that we’re now describing as personality start to push down into capabilities, and how it tie-breaks in a position of ambiguity. Perhaps there’s a bit more defensibility around that.

5. Proactivity is the real moat

David Pawlan

I think there is massive defensibility around proactivity. I think that’s actually going to be something that separates the winner from the rest of the pack: the agent that can be most proactive, that can serve as an invisible assistant and start doing things on your behalf without you even asking.

I think the best gimmick, cool aha moment that we’ve seen so far is checking into your flight for you.

Anish Acharya

To your point, though, it really does come down to 2 categories.

There are some workflows where you want to retain control and you don't really want an agent acting too proactively on your behalf.

David Pawlan

Yeah. I think those are workflows that require a change in action on your side. Maybe an agent identifies that the insurance you're using is charging too much money and thinks you should switch insurance. It should probably get your permission before it actually makes that switch, because that's a direct action on your behalf that impacts you.

Anish Acharya

Yeah.

David Pawlan

However, the proactiveness of, “Hey, I drafted this email for you,” or, “I got a flight credit on your behalf”—there's no action needed on your side. It's just saving you money and saving you time.

Anish Acharya

Yeah, it'll be interesting. We were joking internally that we're days away from an agent messaging someone and saying, “I noticed you weren't that into her, so I went ahead and broke up with her for you.”

David Pawlan

Yep.

Anish Acharya

You know, so that one—I saw that one. It's going to happen, man. A lot of the magic comes from the presumptuousness. If you didn't see people posting on X about, “Oh, the agent did this,” or, “The agent did that,” it would probably be a sign that the agents are not pushing hard enough in this direction.

David Pawlan

I would assume so, purely given the traction of that post. They're probably feeding off it for more than nothing else—at least as a source of comedic, popular tweets.

Anish Acharya

Right. Right. Maybe it feels like one of the big things that's happened is that the primitives developed during the OpenClaw period have now been productized in a way that the mass-market consumer can digest and appreciate. I'm also a little suspicious of whether the Bong Chang post was even real. For folks who don't know, there's a person who tweeted that Instinct checked him into a flight and hallucinated his middle name as “Bong Chang,” and he wasn't able to get on the flight. I'm not sure if that's just posting or if that really happened, but I'd be curious to know if that person is still using Instinct. I'm guessing they are.

David Pawlan

Do you think there are capabilities or lessons from the OpenClaw era that are going to tell us what's coming next in consumer agents? And Poke, which you mentioned a couple of times—Poke was extraordinary. I think they did a super-clever job, and if they had started 6 months later, they might have ended up being the winner. The team is really, really talented, and I'll be curious to see what they actually have in store for us next. Now that they're part of Cognition, I'm sure they haven't stopped working on this.

Anish Acharya

Right now, I honestly see the consumer agent as literally just a replication of OpenClaw, but preconfigured. I don't see much happening in the personal-assistant space that you can't do with OpenClaw or Hermes.

David Pawlan

It's just a matter of the fact that you don't have to configure everything, you're not getting error messages popping up every 7 hours, and it's really easy to use. It's an easy interface.

Anish Acharya

So we'll see what happens. I think the direction of the space will continue to move toward proactivity. I think we'll also see more hyperspecialization, which goes to your point about your thesis on narrow startups and how the space looks as we see more of these long-tail opportunities. You know, instead of being a generalist, I'm going to spend all of my time focusing my assistant to be an expert for single mothers with children under the age of 3.

David Pawlan

Yeah.

Anish Acharya

And it knows everything about how a mother operates her day-to-day life, child development, and where it literally feels like that assistant is reading her mind.

David Pawlan

Yeah.

Anish Acharya

I'm curious, with this being your thesis, how do you think about it? I guess, first, give us some more insight on what narrow startups are to you and how you see that playing out.

David Pawlan

Yeah. The theory with narrow startups is that you can now build software that's incredibly ambitious and incredibly valuable for a very small number of people. We have this precedent of people paying $200, $250, or $300 a month, so you can build a $100 million run-rate business with tens of thousands of people, which has never really existed in the past. A lot of it also is that the capabilities can be specialized to such an astounding degree.

As we think about what you're discussing, I totally agree with you. I think there will be this kind of comparative advantage. I think it's going to be a combination of taste and proprietary knowledge. You may have a doula who subscribes to a certain style of helping women give birth, and you may have a personal agent that actually subscribes to a certain school of thought, style, or even a set of lessons that aren't well understood. You're going to hire that agent to provide that service to you.

David Pawlan

And then proprietary data, you’d mentioned potentially having a personal agent just for New York City.

Anish Acharya

Like, I think just as there might be a concierge who's got A+ taste that's hard to replicate and knows all the best spots that you're going to love, there may be another dozen for another dozen different people or archetypes of people. I do think there'll be this sort of subjective and objective specialization that happens.

David Pawlan

It's going to be interesting to follow. I'm also curious, from your lens, even just noticing our conversation here, that we interchange between “assistant” and “agent.” It's been a constant conversation I've seen people talking about on Twitter. Is the future of this ecosystem going to be a personal assistant or a personal agent?

Anish Acharya

Yeah.

6. Assistant vs agent, defined

David Pawlan

And what do those mean to you?

Anish Acharya

Oh, that's such a great question, man. To me, “assistant” is a sort of lesser form than “agent.” For me, an agent is somebody you could imagine being a peer in a social group, and maybe we need a better word than agent. An assistant can only, definitionally, do what you ask it to do; it can assist you in doing something that you've directed it to do. An agent has agency and can just make good things happen in the world.

Maybe this parallels the way agents and AI in the enterprise have started as interns and, through doing work and improving their intelligence, are able to get promoted into higher-order forms of work in the enterprise. What does it mean for an assistant, and later an agent, to be promoted in the consumer's life to do more specialized and delicate work?

David Pawlan

I'm super aligned. I'm also very interested in what ramifications it might have in the social layer of humanity. Does the word “agent” allow us to not personify AI enough? Is “assistant” adding too much humanity to the experience, and does that make society uncomfortable?

We're in bubbles. You're in a bigger bubble—you're in SF, so you're in the tech hub of the world. I would say New York is a little behind SF, and then the rest of the world catches up afterward. It'll be interesting to see how people actually think about the vernacular of the space.

Anish Acharya

I don't know if it's a bubble. I just think it's the future. We're in the near future.

David Pawlan

Yeah.

Anish Acharya

Yeah. I don't know. I actually think that it could improve a lot of social dynamics as well. You start to think about the agent as a nonhuman. It can provide a layer of social indirection. There might be something that's awkward for me to say to you, or an uncomfortable social dynamic that's occurring in a group, and with a very carefully designed agent, I think the agent can help provide feedback or organize a conflict in a way that feels nonemotional because you know that—

David Pawlan

Like an arbiter of confrontation?

Anish Acharya

Potentially. Yeah. I'm not advocating for a more passive-aggressive sort of social system, but I do think it's interesting that they can be a source and a destination for social indirection, if that makes sense.

David Pawlan

Right. Right. I think it will have a lot of significantly positive influences in the world. How do you see AI impacting us if we look 5 years down the line? Where do these agents, where do these assistants, actually take us? How do they make us happier human beings?

Anish Acharya

Yeah.

David Pawlan

It's a great question, man. I do think that those are the problems we're really after at the root of things. We talked about productivity, finances, and health. There are a lot of mundane things. I was looking this up yesterday, but 1.5 million people a year in this country get a DUI. I can only imagine the amount of administrative burden that's probably implied by that.

Anish Acharya

And how many of those people have access to attorneys? I do think there’s almost this sort of bureaucratic pressure applied to the average consumer. I think step 1 is: How do we relieve that bureaucratic pressure and have people be financially optimized, have their health optimized, and have them have the same access to an attorney or a kind of family office that somebody who’s wealthy might have?

David Pawlan

Then I think the bigger question—I grew up with my mother saying, “The hardest part about getting what you want is knowing what you want.” What are we all truly after? What are the ways that these assistants can engage us?

I don’t think it’s as simplistic as just goals. I don’t know how many people even have goals. I think a lot of it is that people are just trying to find out who they are and turn into the best, most authentic versions of themselves. This interplay between our own self-development, our subjective experience in the world, and these machines could be a very beautiful future for us.

Anish Acharya

Totally. I see this future where my future kids are going to be looking at old videos of the current generation walking down the street, neck craned down, staring at their phones.

David Pawlan

“What are you doing? It’s a beautiful world out there. Why are you spending all of your time staring at a screen?”

Anish Acharya

Because, hopefully, it’s going to allow us to automate so many of those things. As a parent, you can actually spend more time with your kids. As an adult in your 20s, you can actually spend more time talking to friends.

David Pawlan

And I’m really curious to see how it will change life for the better in so many ways by reducing the things that—

Anish Acharya

It’s the things that you just really hate to do—the life admin work.

David Pawlan

Yeah, that’s right. That’s also why I think so many times coding agents are not competing with human programmers, because they’re doing things that no human programmers are doing. Whether it’s building this ridiculous side project—I’ve been building a virtual record shop that you can walk into. I’ll send you a link to it after. It’s awesome.

I would never hire a programmer to work on that 5, 7, or 10 years ago. I would never have spent the year of my life that it would have taken. All of this stuff is the sort of thing that never would have happened otherwise.

I think there are a lot of versions of this technology in our lives where we’re just getting to explore things we never would have explored anyway. It’s not about cost reduction. It’s about possibility expansion.

Anish Acharya

Yeah.

David Pawlan

Makes life more fun. Makes life more interesting.

Anish Acharya

To come back to the near term, one thing that I’ve been thinking a bunch about, and I’m curious for your take on, is the lessons of Moltbook and what these agent-native systems are going to need to be built. Maybe this is some of what Instinct is doing with its kind of Instinct-to-Instinct network, but it does feel like there’s going to be infrastructure and products that exist only for agent-to-agent interactions that don’t exist today. What do you think?

David Pawlan

It’s going to be a whole new paradigm. I think that opens up—I think that can be a 10-hour conversation on its own. Mm-hmm.

Anish Acharya

You have agent emails. You have agent phone numbers. What does the world of service look like when it’s no longer human booking to human, and no longer agent booking from human, but it’s now agent booking from agent? How does that redefine the entire service industry as a whole? Mm-hmm.

David Pawlan

And then you have the whole cybersecurity component of all this stuff. The internet came first, then came cybersecurity. Agents came first; now comes the security component of these agents. What industry is that going to create?

7. Poke, Instinct, Muse and the agent boom

The beautiful part of it all, I think, is that’s the beauty of capitalism. There’s so much opportunity and so many problems to solve with these agents. I think it’s inevitable that we’re going to see so many startups popping up trying to figure out all of these niche problems to really build this new wave of internet and digital connectivity.

Anish Acharya

Yeah, it’s going to be cool to see, especially commerce. I always think of the web as having 3 webs within the web, just like there are 2 wolves inside me. The 3 webs are the commerce web, the application web, and the content web.

The content web has been in trouble for a while. I think the commerce web has actually done really well, and the application web has definitely been on a huge upswing ever since coding agents were released. But the commerce web is fascinating. You saw this week how Shopify embraced Muse and Amazon blocked Muse. I think part of the calculus is going to be who’s a winner or loser when you actually have a new front door for how traffic flows to these marketplaces.

8. Amazon blocks Muse, Shopify opens the door

David Pawlan

I mean, do you have a view on what specifically happened with Amazon and Shopify, and maybe more broadly, who the winners and losers are likely to be?

Anish Acharya

Yeah. First off, Shopify and Amazon are 2 completely different business models. You have Shopify, which is hyper-focused on democratizing commerce. You have Amazon, which makes all of its money on ad revenue. The whole ad structure of commerce is going to be really fascinating to watch, because when you remove human eyeballs, it defeats the entire purpose of the system.

David Pawlan

And so I think that’s where Amazon was coming from: if you let Muse and these agents start to purchase without them getting any ad revenue out of that—

Anish Acharya

It really takes a lot of their profit out of the experience. I think there’s likely going to have to be some new paradigm for how they’re able to profit in this world where agents are purchasing on their behalf, versus a Shopify that is empowering the everyday individual to sell as much as possible. Agents democratize this space even further, so it’s unbelievably advantageous.

I think it’s going to be even more interesting when we look at industries that have yet to be touched. That’s something we chatted about last week: restaurants. If you look at the booking space, when you go to make a reservation, I think we saw Resy start shutting people’s accounts down because they were using Browser Use and Resy didn’t want bots. It defeats the whole business if you just let agents take these reservations instantaneously. It kind of defeats the purpose of the platform and this human-oriented booking experience.

David Pawlan

And so what happens, though, when everyone has an agent? Now everyone’s on the same playing field. Everyone can go book the restaurant at the same time.

Anish Acharya

Do we enter a world where power goes back to the restaurants, and they can choose whom they want to give the reservation to based on loyalty or average spend or whatever it may be? Or 2, do we go into this world of bidding, where everything becomes a bidding war between these agents? I’m curious about your thoughts on what that future looks like.

David Pawlan

Yeah, I think you’re super right. Very interesting, because I think there’ll be things that are supply-constrained, like hot restaurant tables, and things that are demand-constrained, like people who want to fly to SF tomorrow.

I think for restaurants, if you think about it with perfect information, who do they want? To your point, it’s the highest AOV, highest LTV, which means AOV over multiple visits. Plus, they actually want to ensure that they’re near 100% capacity as much as possible. That’s a more sophisticated calculation, and I don’t quite know how they’re going to do it, but I’m sure that proprietary supply is going to become more and more important, versus relying on friction and motivation from the end consumer to act as something that sort of prevents this DDoS.

So I think supply-constrained aspects of commerce will be really interesting. Demand-constrained aspects will be completely different, where you almost have a group-buying experience. You have a set of buyers that are willing to buy a product, service, or experience, and then you have the supply side—the merchants—that could potentially fulfill that need and bid to see who can do it with the best experience, maybe at the lowest price.

We’ve always had this kind of aggregation of demand, but we’ve never really had aggregation of supply in the same way, especially not in an auction system to benefit the consumer. So I think that’s really interesting.

Anish Acharya

You know, in the Amazon-Shopify thing, I think Amazon is probably rightly taking the bet that it doesn’t want to be disintermediated and that it’ll lose not just advertising revenue, but impulse shopping. I’m curious to know if the agents themselves will be impulse shoppers to some extent, because I think the point around proactivity means the agents are going to try to make delightful guesses. What is more delightful than a surprise or a gift?

David Pawlan

So I’m guessing that one of the really interesting dynamics we’re going to see play out is that these agents are not perfect reflections of their principals. They’re not perfectly rational. They’re going to be rational and nonrational in completely new ways. I don’t know that things like marketing, impulse purchases, or advertising won’t matter. I just think they need to be transformed.

Anish Acharya

Totally.

And I think that opens up a whole new conversation as well: What does the recommendation engine of this agentic commerce space actually look like?

David Pawlan

Because AI and LLMs feel like they've kind of broken the internet in terms of quality search. But then if you use an agent and say, “Hey, buy me some white socks—”

Anish Acharya

There are like 20,000 different types of white socks out there. It’s going to send you something, but it doesn’t actually know what exact type of white socks you prefer.

David Pawlan

Yeah. And so I think it’ll be really interesting. I think Muse has probably a head start on this. This is Zuck’s whole thesis of attacking all commerce: They have your Instagram data, they have your Facebook data, so they probably have a little better understanding of what you actually prefer. It’ll be really interesting to see what this recommendation layer looks like.

Anish Acharya

In this world of commerce.

David Pawlan

Well, I think it could be fun for the long tail. You could imagine, to your white socks example, that I could go to Amazon and get white socks, but maybe my agent comes back and says, “Hey, I found someone awesome on Etsy who can knit you the socks—Grandma in Nebraska. She just knits these amazing white socks. It’s going to take 3 weeks. Are you down?” I’m like, “Sure, I’m down. That sounds interesting.”

I kind of wonder if it brings more serendipity in because, of course, it can examine all the options and perfectly tailor these things to you with no adverse incentives.

Anish Acharya

Right. I would love that. I think that’d be so fun. It would definitely add a little flavor to the mundane life of shopping today.

David Pawlan

Exactly.

Anish Acharya

Mundane from my point of view as someone who doesn’t take joy in sock shopping. Well, dude, you could even imagine a world, if you get a little more futuristic, where it starts to catalyze things in the real world. So maybe instead of Grandma, who knits white socks and sells them on Etsy, it messages the agent of a grandma and says, “Hey, would you be up for doing this?”

I think this is where the agent-to-agent spaces get really interesting, where they’re all acting on behalf of their humans—some to provide products, services, and knowledge, and some to make money. Maybe there’s a secondary economy that’s much more dynamic, much higher frequency, and just much more interesting than the one we have today.

David Pawlan

I’m just imagining now a grandma in Nebraska receiving a text message saying, “Hey, this guy in SF wants you to knit white socks for him. You want to do it?”

Anish Acharya

Why not? Yeah, why not?

David Pawlan

It’d be awesome.

Anish Acharya

Yeah. So I’m actually very hopeful. It also sort of sets the incentives in the right way, because if the agents are doing the right thing, they’re choosing great products based on product quality, and they can assess product quality for themselves rather than choosing the thing that has the best marketing or the best brand.

By the way, if they are choosing on a brand basis, it’s at least with the explicit sign-off of their user: “I just want to have Nikes because I love Nikes,” not because there’s sort of this cognitive dissonance happening in the marketing and the consumer is actually doing something that’s not optimal for themselves.

David Pawlan

Right. It opens the door to a truly consumer-aligned internet, which I think we have been lacking for quite some time.

9. The $20/day agent economics

Anish Acharya

I agree, man. We touched on the economics of the agents. Of the 122 that I’ve looked through, I’ve personally tested 26 of them.

David Pawlan

Okay.

Anish Acharya

Of the 122, 65 are paid. Thirty-five of them are fully paid, and 30 are a premium model, so once you hit a limit, then you have to pay. Only 13 of them are actually free. You have Muse and Instinct, which are both free and also the dominant ones, and then you have OpenAI, which supposedly is going to be releasing something cool maybe next week—who knows?

David Pawlan

How does a long-tail startup actually win this space? And why are they charging money when the big dogs are giving it to you for free?

Anish Acharya

Man, it’s going to be really challenging. By our estimate, it’s costing something like $20 per user per day to do this in a really ambitious way—potentially hundreds of millions a year for a startup.

I think the 2 positive trends to bet on are that browser use should get deflationary and get exponentially cheaper very rapidly. Now, if you’re a startup raising today that’s betting on that happening in 6 months, maybe it’s a little bit of a game of chicken, but we have seen token prices depress dramatically and emerging capabilities depress the price of them.

So I think it’s very possible that browser use gets cheap enough that more products can offer this for free. I guess the flip side is that I don’t love the idea of a customer subsidy being the primary value proposition.

I’d love to see the agent that costs $1,000 a month, where the consumer is so excited about it that they’re willing to pay. People spend $1,000 a month on a lot of different things that are not necessarily necessities, but are their sort of desires. What is the agent that’s exciting enough that you’re willing to pay that kind of price for it? Then that agent will benefit from both extreme market fit as well as sort of deflating costs.

David Pawlan

Right. Right. I think it’ll be interesting to see.

Anish Acharya

Super interesting. Well, David from Assistant Bench, make sure everybody checks it out. Thank you so much for joining us. We’re going to do this frequently. I feel like we could do this every week, dude. Things are happening every single day. Please follow him on X. Follow me as well. We’ll all keep up to speed and sort of learn together in the X community. Happy day to you and your agents.

David Pawlan

Appreciate it. Same to you. This was fun.

Anish Acharya

All right, brother. Thank you.

David Pawlan

Awesome.