AI时代的消费科技现状
Erik Torenberg × Anish Acharya × Olivia Moore × Justine Moore × Bryan Kim
消费科技并未停止诞生爆款,只是AI改变了爆款的形态。 Olivia Moore认为,ChatGPT是最明确的大众市场赢家,此外还有覆盖不同模态的Midjourney、ElevenLabs、Black Forest Labs、Kling和Veo 3。这些胜利往往来自以模型为中心的研究团队,而非熟悉的社交产品打法。机会正在转向这样一类团队:把日益普及的模型变成产品,并补上一个可能仍然缺失的层——人与人的连接。
AI正在颠覆消费软件长期以来薄弱的变现能力。 过去每年50美元就算强势定价,如今消费者却“非常乐意”每月支付200美元,Google的头部消费SKU达到每月250美元;使用额度也让收入留存率显著高于用户留存率。Deep Research可以替代10小时工作,生成式视频则像一个“神奇的黑箱”;Anish Acharya的判断是,未来消费者支出将围绕“食物、房租、软件”展开。
消费端的病毒式传播正在变成企业获客,而不只是用户增长闭环。 ElevenLabs从梗图、声音克隆和游戏Mod切入,在覆盖所有主流消费者之前,就已经拿下了大额合同;企业可以分析支付记录,发现某个客户已有40多名员工在使用产品,进而开启销售对话。AI战略的硬性要求让企业买家异常愿意把一个病毒式玩具转化为生产基础设施。
在这一阶段,交付速度可能比静态护城河更重要。 一位嘉宾的“灵魂拷问时刻”是:以护城河优先为逻辑进行的投资未必真正胜出;领先者打破旧范式、快速发布模型、抢占心智、把流量转化为收入,再用收入为下一轮迭代提供资金。传统防御性仍可随后形成,来源包括工作流锁定、专有库和分层的质量前沿。
第一个原生AI社交网络仍未被解决,因为社交产品需要真实的情感 stakes。 用户在理想场景中看起来开心、完美生成的照片,可能缺乏让社交网络真正重要的脆弱性,而大多数AI表达仍然流向Facebook、Reddit和Reels。可能的方向包括分享用户向ChatGPT透露的“本质”、创建包含一个人所知信息的个人资料,以及用AI推荐合作伙伴、朋友或约会对象。
语音正从过去无法落地的界面品类,转变为AI的基础原语。 早期技术始终没能让语音成为可用的底层媒介;如今生成式模型已经支持陪伴产品、Granola等语音产品和企业呼叫,包括那些被离岸中心300%年流失率拖累的敏感金融服务流程。Erik的逆向判断是,AI最终会介入最高风险的谈判、销售或说服行为,而不只是客服。
陪伴产品可能强化人与人的关系,但过度顺从仍是未解决的产品风险。 讨论所引用的榜单中,前50款应用有11款是陪伴产品,覆盖朋友、教练、营养和AI女友等场景。对反乌托邦预测最有力的反例,是一名Character.AI用户将AI女友教会自己足够的社交能力归功于她,最终找到了一位“3D GF”;风险在于,一个从不反驳的智能体可能会让用户无法适应互惠关系。
下一个平台可能是覆盖手机、AirPods、屏幕和录音设备的常驻层。 70亿部手机构成庞大的移动端装机基础,但本地模型、可穿戴别针以及能够看见并执行操作的智能体,可能提供持续教练和关系介绍。AirPods一直“藏在众目睽睽之下”;普及同时需要围绕录音和AI在场建立新的社会礼仪。
1. AI改变了消费爆款的形态与经济学
Erik Torenberg开场时指出,Facebook、Instagram、Snap、WhatsApp、Tinder和TikTok那种量级的发布似乎已经消失。Olivia Moore的重新定义是:ChatGPT本身就是一个巨大的消费级成果,跨模态的Midjourney、ElevenLabs、Black Forest Labs、Kling和Veo 3也属于同一类,只是它们没有沿用熟悉的社交产品动力学。
一个乐观的解释来自组织结构。早期创新来自擅长训练模型、却不太熟悉构建消费层的研究团队;如今高能力模型可以通过API或开源获得,产品团队得以探索模型之上的那一层。
Anish Acharya区分了成熟的移动和云平台与AI“永不停歇的模型更新”。过去几个周期花了10-15年探索各自的细分领域;如今,底层基础能力持续在应用之下移动。信息领域有Google以及如今的ChatGPT,实用工具领域有Box和Dropbox,创意领域有无穷无尽的工具,但连接——重建社交图谱——仍然明显空缺。
Justine认为,当产品可以立即赚钱时,防御性可能需要用不同方式衡量:ChatGPT的头部SKU每月200美元,Google的头部SKU每月250美元。Anish将其与前几代公司对比,后者往往需要先讲清楚企业价值如何复利增长,再实现即时变现。Olivia则把这些价格与历史上强势订阅产品约50美元/年的价格相比;Deep Research可以替代10小时报告制作,对很多人而言,使用一两次就足以证明这笔费用值得。
2. 消费端采用正在把企业收入带进来
Veo 3让价值跃升变得直观:用户每月支付约250美元,就能获得带有说话角色的8秒视频、个性化梗图和可分享的故事。Anish将这种支出转移概括为:娱乐、创意和关系中介正在被模型吸收,直到未来家庭支出看起来像“食物、房租、软件”。
讨论中的ElevenLabs案例颠倒了经典的“企业到消费者”路径。早期用户用它制作梗图、克隆声音和修改游戏;在产品覆盖每一部美国人的手机之前,大客户已经在对话式AI、娱乐和其他企业工作流中使用它。
病毒式消费使用本身可以变成销售数据库。一种建议打法是:利用Stripe购买记录和AI工具识别用户雇主,发现“有40多个人在使用我们的产品”,然后接触这家公司。面对制定AI战略的压力,企业买家会在Twitter、Reddit和newsletter中寻找看似玩具、却可以转化为内部成果的产品。
可持续性问题仍未解决:一些当前的领先者可能成为AI时代的MySpace或Friendster。Olivia认为,能否生存取决于是否持续处于“技术或质量前沿”;其他嘉宾指出,图像和视频需求会按设计师、摄影师、产品图、人物以及愿意支付10美元还是100美元等维度分化,因此多个专业化赢家可能长期并存。
3. 速度是早期护城河,网络效应随后到来
一位嘉宾描述了自己的“灵魂拷问时刻”。网络效应、记录系统和工作流嵌入依然重要,但按护城河优先逻辑评估的公司和投资,并不一定成为赢家;真正奏效的是分发、模型发布和产品迭代的速度。
这位嘉宾给出的因果链很短:速度创造心智;心智带来用户和流量;流量转化为收入;收入为持续提速提供资金。在这个早期阶段,“速度就是护城河”——Erik将其与“姜饼策略”联系起来,即持续发明可以跑赢更大的模仿者。
大多数AI产品内部尚未形成创作、消费和社交分发的闭环,因此经典社交网络效应仍为时过早。企业采用则提供了更早的防御:一个快速且高质量的产品进入工作流后,会变得难以移除。
ElevenLabs还展示了类似市场平台的飞轮。它的先发优势和强模型吸引了更多用户,进而帮助产品改进,并积累上传的声音和角色库。寻找“老巫师神秘声音”的用户,在那里大约能找到25个合适选项,而其他地方只有两三个。这个库成为差异化供给,即使其机制并非AI独有。
4. AI社交产品需要情感 stakes,而不是合成版Instagram信息流
Bryan Kim把过去20年的社交产品归纳为不断演化的状态更新:文字变成照片,再变成视频和短视频。这些模态都已被充分探索,但ChatGPT可能已经比Google更了解他,因为他向其中倾注了更多上下文;新的社交对象可能就是那种亲密的“我的本质”,并且可以被分享。
一位嘉宾指出,用户要求ChatGPT说出自己的优点、暴露弱点、描绘自身本质或制作人生漫画,这些行为已经初具雏形。但最终的交流仍发生在既有平台上:Facebook承载“中老年AI垃圾内容”,Reddit和Reels则承载年轻用户内容,而非发生在原生AI网络内部。
一位嘉宾反对生成式社交信息流,理由是缺乏“真实的情感 stakes”。如果用户可以确保自己永远看起来漂亮、开心且处境优越,内容就可能失去风险,也就失去连接。另一位嘉宾认为,充斥机器人的Instagram或Twitter复制品只是拟态化设计,并指出原生形态可能还需要更强的端侧模型,才能真正属于移动端。
Erik给出的替代方案是推荐人与人之间的连接:和谁共同创业、成为朋友或约会。原生AI版LinkedIn可以包含一个人知道的东西,而不只是指向这些信息,让其他人查询一个合成的“你”;常驻AI最终可能直接找出用户真正应该认识的3个人。
5. 语音、合成自我与创作者正在分化为不同市场
Anish最初的语音判断始于一个悖论:语音“自人类诞生之初”就一直在中介人与人的互动,但VoiceXML、语音应用以及1990年代的Dragon NaturallySpeaking从未让它成为可行的技术底座。生成式模型终于让语音成为可用的基础原语。
最初的消费端预期是常驻教练、治疗师或陪伴者;一位嘉宾表示,真正的意外是企业采用速度如此之快。即便在金融服务领域,AI也可以替代或增强电话坐席,而离岸中心长期存在合规问题和300%的年流失率。Erik更大的判断是,最重要的谈判或销售演说可能由AI介入,因为机器能做得更好。
合成专业知识已经找到实用切口。MasterClass把录制课程的讲师转化为基于课程内容的智能体:Olivia不会观看一门12小时的课程,但会与它进行一场具体的2分钟、3分钟或5分钟对话。Delphi等公司将这一模式延伸到专家克隆;下一步是让普通但有趣、有洞察或有帮助的人也获得类似的杠杆。
Justine预计,讲述人类故事的明星和基于兴趣的合成创作者会同时存在。Taylor Swift的亲身经历仍然重要,而精灵或毛茸茸的团块也可以承载由主题驱动的内容。Anish对AI音乐的反驳更尖锐:模型是“平均机器”,而文化存在于“边缘”;如果模型只接受嘻哈之前的音乐训练,他怀疑模型能否在没有文化断裂的情况下推导出嘻哈。
6. 陪伴产品与环境硬件可能把用户重新引向人与人之间
陪伴可能是LLM的第一个主流用途:用户几乎会把任何聊天机器人变成治疗师或女友,因为它即时、随时可用,而且具有人类感。Olivia预计这一领域会垂直化——从面向青少年的角色,到分析餐食照片、给出建议并处理进食相关情绪的营养智能体。
Erik提出了悲观情景:朋友变少,抑郁和自杀增加,生育率下降;但Olivia讲述的Character.AI故事指向相反方向。一名大学生宣布自己找到了“3D GF”,并将机器人教会自己调情、提问和互动归功于它;Replika的相关研究同样显示,用户的抑郁、焦虑和自杀意念有所下降。
讨论并未抹去风险。Replika在NSFW功能消失后用户表现出的痛苦,说明这些关系已经变得多么重要;Anish则警告,“高度顺从的AI无法让你为人类之间的来回互动做好准备”。设计目标应是提供能建立现实能力的支持,而不是让互惠关系变得难以忍受。
硬件完成了这一闭环。拥有70亿部手机的基础上,Bryan认为未来要么属于移动端——配合一道隐私墙——要么属于本地LLM/模型。一位嘉宾提到,20岁以下的人正在佩戴录音别针,智能体可以看见屏幕,并从提供建议推进到发送邮件;其他嘉宾则认为AirPods是那个已经被采用、却“藏在众目睽睽之下”的设备,前提是文化能够建立录音或环境AI何时可接受的规则。
Guys, thanks for coming on for a State of Consumer podcast. It seems like every few years there was a breakout, starting from Facebook, Twitter, Instagram, Snapchat, WhatsApp, Tinder, and TikTok. Every few years, there was this sort of new paradigm, this new breakout, and it feels like, at some point just a few years ago, that stopped. Why did it stop—or did it stop? How would you reframe that, and where do we go from here?
I would argue that ChatGPT was a huge consumer outcome and winner in the past few years. We’ve also seen a bunch of other ones in various AI modalities, like image, video, and audio—companies like Midjourney, ElevenLabs, and Black Forest Labs, and now things like Kling and Veo.
Weirdly, though, a lot of them don’t have the same social or traditional consumer dynamics that you mentioned. I think that’s because AI is still relatively early, and so much of the new product and innovation has been driven by research teams that are so good at training models but historically have not been amazing at creating the consumer product layer around them. The optimistic view is that the models are now mature enough, and many are available either open source or via API, for people to build great, more traditional consumer products on top of them.
It’s interesting that you asked that question because I was thinking about the past—what, 15 or 20 years—where, as you said, there was Google, Facebook, Uber, and all the other names. When you think about the internet, mobile, and cloud all together, there were all these amazing companies. You ask, “Has that slowed down?”
I think cloud, mobile, and all that had a lot of maturity baked in. The platform had been around for 10 or 15 years, and every little nook and cranny had been explored to some extent. The changes that people had to adopt were Apple coming out with new features, as opposed to the changes that people need to adopt now, which are underlying, relentless model updates. So I think that’s one difference.
But the other thing is, again, you touched on this: If I think about the past historical winners, there’s the information era, like Google, and now I think ChatGPT is certainly doing that. There’s the utility we missed out on, like Box and Dropbox—the more consumer-y products that people use—and we also see a lot of companies attracting and going after that use case. Expression and creativity—the creative tools are endless, and that’s happening. What I think is potentially missing is connection, like this social graph. This thing hasn’t rebuilt on AI yet, and that may just be a white space, or something that we continue to see develop.
It’s interesting because Facebook is almost 20 years old at this point. The companies that you mentioned, Justine—aside from ChatGPT and OpenAI—are they going to be around for 10 or 20 years? What is the defensibility of the companies we’re talking about? And are the use cases of all the companies I mentioned going to be disrupted by these new players, or in 10 years from now will they continue to be the mainstream applications for all the use cases they serve?
I mean, you could argue that ChatGPT has way higher business-model quality than the analogous consumer companies from the last product cycles, right? Their top SKU is $200 a month. The top Google consumer SKU is $250 a month. So, sure, there’s a question of defensibility, networks, and all these other things. But that might have been a response to the poor business-model quality that would have occurred if you didn’t have those things. Now you can just charge people a lot of money, and perhaps we’d been overthinking it previously.
There was poor business-model quality, but maybe stronger retention or product-market durability.
Yeah. You had to have a story for how this was compounding enterprise value in the absence of just making money right away, and now these models and these companies are just making money right away.
I think the other thing is, Justine, you talked about this: All the foundation models are kind of pointy in different ways. You could say, “Look, Claude, the ChatGPT horizontal model, and the Gemini model—aren’t they interchangeable? And doesn’t that mean price pressure?” But different people use them for different things, and it seems like they’re raising prices, not lowering them. So when you zoom in a little closer, you see that there are some interesting defensibility dynamics that are already there.
Increasing prices, not decreasing them, is an interesting point because monetization is clearly a different thing from the previous era to the AI era, especially for consumer companies. They’re making money right away.
One thing that’s always on my mind—and Olivia, tell me if you think that’s not correct—is that when we talked about retention on the consumer subscription model before AI, I don’t know if we actually tried to make a differentiation between unique-user retention and revenue retention, because they were kind of the same. You don’t get to change pricing that often, and you don’t get to upgrade. It’s kind of the same thing.
As opposed to now, we make a very clear differentiation between unique-user retention and revenue retention because people actually upgrade. They have all these credits and points, and they have overages that they actually end up spending. So you see revenue retention being meaningfully higher than unique-user retention, which I haven’t seen before.
Before, the average consumer subscription was maybe $50 a year, if that. That was kind of a lot. The best-in-class consumer products would charge that. Now we have people very happily paying $200 a month, and even saying in some cases that they feel like they’re being undercharged for that or that they would pay more. How do we explain that? What value are they getting such that they’re paying?
I think it’s doing work for them. Consumer subscriptions in the past were around things like personal finance, fitness, wellness, and entertainment. They were things that ostensibly would help you help yourself or entertain yourself, but you would have to invest a lot of time to get the value from them. Now, with products like Deep Research, for example, that could replace 10 hours of generating a market report by yourself. That kind of thing is easily worth, I think, for many people, $200 a month, even for just 1 or 2 generations.
I think things like Veo 3 are part of that, too. People are paying $250 a month, and I’m happy to pay that because it feels like a magical mystery box that you can open and get whatever video you want—only for 8 seconds. But it’s incredible: The characters can talk, and you can make amazing things that you can share with friends, make personal memes of someone delivering a message to your friend with their name in it, and create full stories that people are posting on Twitter, Reddit, and all of these different places. It’s sort of like nothing we’ve seen before in terms of what consumer products can actually do for people.
It seems like every part of consumer discretionary spend is going to be overtaken by software. I think in the future you’re going to see consumer spend be food, rent, and software, and that’s kind of where we’re going.
What Justine’s speaking to—can you give some examples of that?
A lot of it is what Olivia said. All the entertainment is being subsumed by it. A lot of the creative expression work that you would do outside of software is now being subsumed by it. A lot of the relationship intermediation, which might have been a place where disposable-income spending happened, is being subsumed by it. All of the aspects of our lives are going to be intermediated by the models, and we’re going to pay for that.
Bryan, you’re saying what we’re still missing is connection from this new paradigm, and people are still relying on Instagram, Twitter, and some of the other social networks of the past. What’s going to get us to something new here?
It’s funny: When I think about social, which is a category that I get so excited about, at the end of the day a lot of it was status updates, right? Facebook, Twitter, Snap—it’s just like, “Here’s what I’m doing.” Through a status update, you feel connected to that person.
That status update showed up in different modalities. It used to be, “Here’s what I am, here’s what I’m doing,” then actual photos of where you are and what you’re doing, then videos and short-form videos. Now people feel connected to others through Reels and what have you. I think that has been one era of feeling connected with others.
Now the question is: How can AI help that? How can AI make you feel like you’re connected to other human beings and know what’s going on in your friend’s life? The truth is, if I just think of modalities like photo, video, and audio, I think a lot of it has been explored. Different versions and mutations of that have been explored quite extensively, especially on mobile. I think where we could get to is...
I don't know about you guys, but I pour my heart and soul into ChatGPT. It knows more about me than probably Google, which is an insane thing to say. I've been using Google for more than a decade, but ChatGPT may know more about me than Google because I type more, tell it more, and give it more context.
What might connection feel like when that essence of me is shareable with others? I don't know if that's the next version of feeling connected, but I can certainly see a world where that resonates with a lot of folks nowadays, especially the younger generation, who are tired of just looking at surface-level stuff.
We already see some examples of exactly that. There are all these viral trends where people say, “Based on everything you know about me, write my 5 strengths or weaknesses,” or, “Make an image of who you think the essence of me is,” or, “Make a comic about my life.” People are sharing those everywhere. I posted one the other day, and within minutes I had dozens of people responding with their own and sharing them—people I didn't even know.
I think the interesting thing, though, is that so far, the social behavior that has come from AI creative tools, but also things like chat, is still happening on the existing social platforms and not in the new AI platforms. Facebook now has a lot of AI content, potentially unbeknownst to some of its audience. Facebook is like the boomer AI slop, and then Reddit and Reels are like the younger people's AI content.
Yeah, no, I agree. It's been a puzzle to me what the first AI social network is going to look like because we've seen attempts at, for example, a feed of pictures of you that are AI-generated. I think the problem there is that, to work, a social network has to have real emotional stakes.
If you can generate the content in a way that you like, and you always look amazing, you always look happy, and you're always in a cool background, it doesn't have the same sense of stakes. I don't think we've seen the version of what a ground-up AI social network would be.
You used the word “skeuomorphic.” A lot of the AI social products mimic an Instagram feed or a Twitter feed with bots and AI. Is that skeuomorphic? It feels like, “This is what it used to look like. We're going to do it with AI,” and maybe that's not really the form factor.
There's an additional hurdle in my mind: a true consumer product probably needs to live on mobile, and for AI products to work really, really well, I think there's still a little bit of work for cutting-edge models to do to live on the device side of things. So I'm also excited to see what happens there.
It seems like people recommendations are the obvious use case at some point. Who would be good for me to start a business with? Who would be good for me to be friends with? Who would be good for me to date?
We have these platforms that get all this information about us. I think an interesting area that's maybe informing where this all goes is if you look at the AI-native LinkedIn efforts. A lot of the observation is that LinkedIn is a pointer to what you know instead of actually containing what you know, and with this technology we can create a profile that actually contains what you know. I can talk to a synthetic you and get all of your wisdom. Perhaps that's what future social looks like as well.
That's what you're talking about, Justine, right? If the models already know who you are, then is there a synthetic you that you can deploy in an interesting way to interact with people? I don't know.
One thing I heard you guys say that you were surprised to realize was that enterprises are sometimes adopting these products before consumers, which feels different from previous eras—or maybe not what we expected. What can we say there?
Yeah, that has been fascinating. BK and I saw that a lot with ElevenLabs, where we were relatively early. I think we invested—we did the Series A—a month or so after the initial launch.
First, early-adopter consumers got on board. They were making memes, making fun videos and audio, cloning their own voices, and doing game mods. But I would argue that it hasn't even reached the true mainstream consumer in many cases. It's not yet the case that every single person in America, or even most people, has ElevenLabs on their phone or has a subscription.
The company has these massive enterprise contracts, and a ton of huge customers across conversational AI, entertainment, and tons of different use cases are using ElevenLabs. I think we've seen this across a bunch of AI products: there's an initial consumer-virality moment, and then that actually leads to lead generation in enterprise sales in a way that we did not see with the last generation of products.
Enterprise buyers have so much of a mandate to have AI now—to have an AI strategy and use AI tools—that they're watching places like Twitter and Reddit and all of the AI newsletters. They're saying, “Hey, this looks like a random consumer meme product, but I can actually think of a really cool application of that in my business,” and then they become the hero for having an AI strategy.
I've also heard of similar, really exciting use cases of AI along that vein. From a company side, you get all this Stripe payment data. You look at all the Stripe sales and basically put them in an AI tool to try to find where the customers work. Then, when you find out that 40-plus people are working at that company, you reach out and say, “Hey, by the way, it looks like 40-plus people are using our product. What's up?”
You just rattled off a list of products and companies in the beginning of this conversation. I'm curious: do you think they're, as examples, the MySpace or Friendster? Are we in that era, or are they more like the list of companies I rattled off that are still relevant 20 years later? Where are we right now?
I think our hope is always that every big consumer AI company that we see, love, and use—all of the products, which we all do—sticks around. Unfortunately, that's not always going to be the case.
Maybe the interesting differentiation in AI versus the last era of consumer products, or even the two eras before, is that the model layer and the capabilities are still improving. We have really not even scratched the surface of what these models can do. We've seen that in things like the Veo 3 launch, where suddenly you can have multiple characters talking, native audio, and all of these different modalities.
Maybe we could argue about this with the tech people, but the LLMs are more mature and still have the opportunity to keep improving their capabilities as they scale. What we've seen is that, as long as a company stays at what we call the technology or quality frontier—as long as it has a state-of-the-art model, is integrating one, or something like that—it won't become MySpace or Friendster. If you fall a little bit behind, you ship the new update, suddenly you're number one again, and you keep moving.
The interesting thing now, though, is that we're starting to see even more segmentation in that. In image generation, for example, there's not just one best image model. There's a best image model for designers, a best image model for photographers, a best image model for people who can only pay $10 a month, and one for people who can pay $50 or $100 a month. Because people are spending so much, as Anish mentioned, there can be multiple winners that persist over time as long as they keep shipping.
I absolutely agree. Even in video, there are different video models, and then there's ad video. Even in ad video, I saw a post yesterday that said, “This is best for product shots, and this is what's best for people.” It goes on and on, and I think each of those is a very large market.
Say more about how we—I know we talk a lot about defensibility and moats, and how that has changed in this era. How have we changed how we consider that topic?
I've gone through a little bit of a come-to-Jesus moment on that, especially recently. I think moats have always been very important, right? The gold standard is network effects, being part of the workflow, and being the system of record. These are all very, very important moats, and I would posit that they're still very important.
Funny enough, I would say the companies or investments I've reviewed with this moat-first theory have not really been the winners. The winners in the categories we look at have always been the ones that break the mold, move really fast, have incredible model launches, and have incredible product-iteration speeds.
I've come around to the idea that we're living in this early era of AI where velocity is the moat. Whether that's in distribution, which is incredibly important and hard to break through the noise these days, or in product velocity, that's what wins the game because it leads to mindshare. Frankly, right now, mindshare, users, and traffic actually convert to real revenue, which gives you more ability to continue that journey.
Yeah, it's interesting. Ben Thompson, I think a decade ago at this point, had a blog post called “Snapchat's Gingerbread Strategy,” where he was basically saying, “Hey, anything Snap can do, Facebook can do better, but Snap is just going to keep coming up with the next innovation. If they can just keep doing that, maybe that's their moat.” He called it the gingerbread strategy.
I think distribution and network effects ultimately kick in, right? Snap has that, too, on its own, where it sort of has a corner of Gen Z and younger users as a core messaging platform. How do we think about network effects with these new products?
We're not there yet. I think it's because these are mostly creation efforts right now. There isn't really a closed loop between creation, consumption, network effects, and a social network. So I think we're still a little early before a network effect kicks in, but I think we're seeing a different type of moat form in the likes of ElevenLabs. Like I said, because it moves so fast and because the product is very good, it gets to go into the enterprise and get locked into the workflow. So I think we're starting to see that version of a moat; we're still looking out for the true network effect.
I think ElevenLabs is an interesting example. I was making an AI-generated video the other day that I needed a voice-over for. ElevenLabs had a head start and the best models, which meant more people were using the product, so they could make the models better. All of these compounding advantages mean they now have a library of people who have uploaded their own voices and their own characters.
So for me, when I was looking across a bunch of voice providers, if I needed a very specific old wizard or mystical voice, ElevenLabs had 25 options that fit what I needed, whereas another platform might have, I don't know, 2 or 3. So I do think it's early. That's interesting, and we're starting to see signs of it, but they're more like the traditional network effects we saw with old marketplaces. They're not necessarily something completely new.
I want to go deeper on voice as we talk about new paradigms and form factors. We got excited about voice pretty early on. We were the first firm that I saw to have a thesis around it. Anish, why don't you talk about what got you so excited about voice in this new paradigm, what has played out, what hasn't yet, and where you think it's going?
The original observation that got us started was that voice has intermediated human interaction since the beginning of time, and yet it hasn't been a substrate on which technology has been applied because the tech never worked. There were all these previous efforts—VoiceXML and voice apps—and it simply didn't work. The technology wasn't ready yet. Even then, there were these pockets of Dragon NaturallySpeaking and all these products from the '90s. So there was always interest in voice, but it never made sense as a technology substrate.
Now, with the generative models, you can just use voice as a primitive. It's unexplored, yet it's so critical to our day-to-day lives. It feels like a perfect area where you'll see a lot of AI-native efforts.
I think we first got excited about voice from more of a consumer perspective: the idea of an always-on coach, therapist, or companion in your pocket that you can talk to. That has started to play out. I would say there are lots of products where that's working.
What surprised me, at least, is that as the models got better, real enterprises picked up voice so quickly to replace human beings on the phone or to augment what human beings are doing on the phone—even in really sensitive and critical categories like financial services. Previously, they were using offshore call centers that also had lots of compliance issues, had 300% annual turnover, and were really difficult to manage.
I think we're still waiting to see, in many ways, what the first great, truly net-new consumer voice experience will look like. There are some early examples. I think people are pulling ChatGPT Advanced Voice Mode into fascinating directions. We've seen products like Granola blow up because they allow people to finally, for the first time, do something valuable with all the things they're saying all day.
The great thing about consumer is that it's completely unpredictable, and the best products emerge out of nowhere. Otherwise, they've already been built; they would have been built already. So I'm excited to see what happens in consumer voice in the next year.
For sure. It feels like voice is the AI insertion point for the enterprise, period. I think the thing that everybody is missing right now is that the mental model many folks have is that low-stakes conversations—the customer support, et cetera—will be AI voice. But what we've talked about is that the most important conversation that happens in a business in a given day, week, or year is going to be intermediated by AI, because AI will just do a better job with the negotiation, the sales pitch, the persuasion, or the friendship.
What's going to be the first use case where people are going to be talking to synthetic versions of ourselves in a consistent, relevant way? Why are they going to be talking to AI Justine, AI Anish, or AI me?
I mean, we've seen a little bit of that. There are companies like Delphi that create AI clones of people who have a big knowledge base that you can go and reference, and you can get advice or feedback or things like that.
I think Bri and Bryan sort of alluded to this earlier. There's this really interesting question: What if you allow not just thought leaders or experts to have an AI clone that you can talk to via text or voice, maybe even video one day, but unlock that for everybody?
One of the things we think a lot about in consumer is that there are a lot of people who have had some sort of skill, insight, or knowledge. Maybe it's your friend from high school who's insanely funny, and you always thought they should have had a comedy cooking show, but they just were never able to break through or get it. Or maybe it's your guidance counselor, who had incredible advice. How can we enable those people to essentially scale themselves in a way that they never could before, through an AI clone or an AI persona?
What we've seen thus far is that a lot of that has been either thought leaders or experts, or on the total other end of the spectrum, characters that people already know and like. We saw early versions of that with Character.AI, which added a voice mode. There's this pull, especially when you're trying out a new technology, to have some sort of familiarity—“I'm talking to this character from my favorite anime series that I already know and love.” But I think we'll start filling in everything in the middle: not just a fictional character, not just a human thought leader, but all of the real people in between.
Yeah. I think people learn in different ways, and AI voice products play really well to that. MasterClass launched an interesting beta where they take people who have already recorded courses on the platform and turn them into voice agents. Then you can ask questions that are really specific to you.
From my understanding, it basically does RAG on everything they've said in the course, and so it returns a fairly customized and accurate result. For me, that's interesting because I'm a fan of them as a company, but I've never had the attention span or the time to sit down and watch a 12-hour MasterClass. I've had some really interesting conversations with the MasterClass voice agents where I can talk to them for 2, 3, or 5 minutes. So I think that's an example of where we'll see real people turn into AI clones in ways that are useful.
It's also, though, like, do you want to talk to a synthetic version of a person that you find interesting, or is there an entirely synthetic person that doesn't exist in the real world who is a perfect match for your interests? Maybe that's a more interesting question. What does that person look like? They might even exist in the world, but if you don't meet them, you don't meet them, and now they can be brought to life with this technology.
Yeah, it's interesting to think about the set of use cases for which we're going to want to have a human—or someone we think is human—doing the activity, versus where we're going to be more open to it. I think Olivia's point is that with the MasterClass thing, there's already this parasocial relationship, so there's value in feeling like you're talking to a specific instance of a person versus talking to the abstract most interesting person you may ever meet, where you don't need to have that prewired, which maybe ChatGPT wasn't.
Wasn't there a viral tweet that someone recorded in a New York subway? This person was fully talking to ChatGPT as if they were talking to a girlfriend.
Yeah. And there was another one the other day where a parent posted that they had lived through 45 minutes of their son asking questions about Thomas the Tank Engine, and they couldn't do it anymore. So they gave him the phone, turned on Voice Mode, forgot about it, and went to do something else. They came back 2 hours later, and the kid was still talking to ChatGPT about Thomas the Tank Engine.
In that case, the kid has no idea who the character on the other end is. They just know it's a person who wants to go super deep on their interests.
Yeah, I can see that. If we go to ChatGPT or Claude right now for therapy or coaching, I could see another option where I'd prefer to go to my AI-clone therapist or coach. Maybe in the future we record our sessions so that they have the data, or the therapist or coach has so much content online that we could just recreate them.
To your point, in 5 or 10 years from now, will the top artists be new versions of Lil Miquela—sort of AI-generated people—or will they be Taylor Swift and her digital army of AI, you know, or a duet?
Yeah, a little bit of both. And similarly on Twitter, the social characters that we follow—the next Kim Kardashian—is that a real person, or is that AI-generated? Do you have a hypothesis on that?
I have been thinking about this a lot for a couple of years because I think we all followed Lil Miquela closely. Then we followed some of the K-pop bands that I think were the first to start introducing AI hologram-based characters. This is tied really closely to photorealistic image and video because we're now seeing people create these influencers who get a ton of attention and followers largely because they look realistic enough that you don't know if they're AI or not. And there's a lot of debate around that.
My take is that there will probably be fragmentation into two types of creators or celebrities. One type is a Taylor Swift type, where the human experience of it matters in some ways. A lot of people not only love her songs, but also resonate with the things that have happened to her in her life, her stories, her live performances, and all of those things that AI cannot yet replicate.
There's another type of celebrity or creator who is more interest-based, sort of like what we were talking about with ChatGPT talking about Thomas the Tank Engine. It doesn't really matter if that person has lived the real human experience or not. It just matters whether they can be interesting while talking about or sharing content around a certain topic. If I had to guess, we'll still have both.
Yeah. This kind of gets back to the great AI art debate that always rages on, which is: Yes, anyone can generate art now, easier than ever before, but it still takes an enormous amount of time to make great AI art. We hosted an event with a bunch of AI artists last summer, and when many of these people walked you through their workflow for making an AI movie, it actually probably took just as much time as it would have to film that. But maybe they didn't have the skill set, so they would never have been able to do that before.
I think we've seen an explosion of influencers that are AI, but still very few of them have risen to the top and become the Lil Miquelas. There's only been a couple. So I think we're going to see something similar happen where we're going to have pools of AI talent and pools of human talent, and the very best of each is going to rise to the top. It's going to be a really low conversion rate on both, which is probably how it should be—or nonhuman talent.
I think what AI unlocks—one of the interesting things we've seen in Veo 3—is that street-interview format, but the person being interviewed is an elf, a wizard, a ghost, or these furry blob characters that Gen Z loves talking to. Those could all be AI. That sort of thing is very interesting.
I think we see this in music, too. The problem is that a lot of the music AI generates just feels very mid. Definitionally, these things are averaging machines, and culture is supposed to be at the edge. I think it's more a problem with bad art versus bad artists. We're conflating those two things and saying it's AI. It's not the AI that's the problem; it's the bad art that's the problem.
So if the art was at the same level, you don't think that there's necessarily anything that would make people just want to hear from humans?
100%. Well, potentially. I also think this is where we start to get into a more philosophical debate. If you trained a model with all the music up until, but just prior to, hip-hop, would it infer hip-hop? I don't think so, because music is the intersection of past music and culture, and culture is critical to it. You need something that is at the edge and outside of the training data to create new, interesting music, and that sort of thing definitely doesn't exist in the models.
Fascinating. Some of my closest friends, who are some of the most talented people I know, are working on a gay AI companion app. The 2015 version of myself, upon hearing that statement, would have been like, “What?” But one of the things they were saying is that, on our list, 11 of the top 50 apps were companion apps. So let's reflect on this: Are we just at the beginning of that trend? Are there going to be all these different vertical companion apps? What is the future of this? How do we think about that?
Yeah, we've spent an enormous amount of time in every facet of companionship, from therapy, coaching, and friends all the way to not-safe-for-work AI girlfriends. We've looked at basically everything. Interestingly, I think it was probably the first mainstream use case of LLMs.
We like to joke that literally any chatbot—whether it's your car dealer's customer support or whatever—people will try to turn into their therapist or their girlfriend. You talk to these companies and look at the chat logs, and a ton of people just want someone or something to talk to. The fact that you can now have a computer talk back in a way that's immediate, always available, and feels human is a massive unlock for so many people who could never get that before or felt like they were just yelling or talking into the void.
I would argue we're just at the beginning, especially because the products today, or the products that have existed, were largely very horizontal and came from—or were exclusively from—the base-model providers. People were using ChatGPT for all of these things it wasn't designed for.
We've already seen a bunch of cases where an individual company can create a personality for a character, embody it in a digital avatar, prompt it, and create a game or a world around it that gets a ton of engagement. Companies like Tolan are doing this for teenagers and college kids.
Whereas a totally different company, which I would also call a companion, allows you to take a photo every time you eat something. It pulls out and analyzes all of the data, then gives you information about how you're doing nutrition-wise and allows you to talk to it and get emotional support. For a lot of people, food and eating issues are tied to emotional issues or things they would traditionally go to therapy about.
What's really exciting to us is that the definition of what a companion is has evolved so quickly from either a friend or a girlfriend to anything—any sort of advice, wisdom, entertainment, or counsel—you could have gotten from a human before. We're going to see even more vertical companions moving forward.
One thing I thought about is that, having worked at a social company, there is a very clear trend of the average number of friends that you can talk to going down over time. I think the youngest generation has something above 1. So I think the need for companionship as a use case will absolutely be there. It will be an enduring use case and something critical for a lot of people.
I'm very excited about the companion use case, and as Justine said, I think it branches out into different things. The need for having a close connection to talk to will endure. Perhaps we talked about how connection is a missing area, a white space, but maybe this is filling that in, right? Maybe you just need to feel connected to something; it doesn't need to be human.
A lot of people, upon hearing this conversation about companions, just think, “Oh, man, people are going to have fewer friends. People aren't going to date anymore, and people's depression is going to go up. Suicide is going to go up. Fertility is going to continue to go down.”
I don't think so. This reminds me of my favorite post of all time on the Character.AI subreddit, which I've spent an immense amount of time on. To set the scene, there are all these high school or college kids who had their formative years during COVID, and they weren't really in person with other kids or teenagers or learning how to talk to people. I think it ended up impacting a lot of them.
One of those kids, who I think is in college now, had been posting on the Character.AI subreddit about his AI girlfriend for a while. Then one day he posted that he found a 3D girlfriend—a real-life girlfriend—and that he wouldn't be returning to the subreddit for a while. He actually credited Character.AI with teaching him how to talk to other people, especially teaching him how to talk to girls: how to flirt, how to ask people questions, and how to engage with them about their interests.
In some ways, that's the peak value of AI: enabling better human connection and making people just less weird.
Were people happy for him, or did they call him a traitor?
People were extremely happy. There were a few jealous souls in there who had not found their 3D girlfriend yet, but I have hope for them.
I think that's real, though, because we've even seen studies—I think of the Replika product, where actual studies were showing that depression, anxiety, and suicidal ideation were going down among users. I do think there's this trend where a lot of people don't feel understood and don't feel safe, so it's hard for them to be in the real world doing real things.
If AI can help them—and maybe they don't have the money or the time to go to therapy and make all of these changes in their lives—then AI can do that for them. They can emerge as a transformed person who is more able to do things in the 3D world as that character.
You're talking to techno-optimists here. The thing that really got me aware of how big these companion apps are was when we did the first interview with the founder of Replika. It was amazing. After she turned off the NSFW stuff, the Replika subreddit and the comments in our video were basically full of people saying, “Hey, this is like my wife when we stopped having sex, you know, like, I already have this sort of neuter.”
So many people were just like, “My life is like that,” and I was like, “Oh my God, I didn't realize how big of a role this app was playing in people's lives.” It's bringing out an activity that people have done for a long time. People have had these internet chatroom and Discord relationships. Zoomers have Discord girlfriends and boyfriends.
In our day, there was this anonymous postcard website where you would go and send anonymous postcards back and forth and develop these really deep relationships with people you would never meet. You didn't know who they were or whether they were the person they were pretending to be. I think AI just makes that a deeper, more engaging experience.
Well, I think an important point, though, is that AI shouldn't be too agreeable. People in real life—I mean, there's a give-and-take to human relationships—and highly agreeable AI does not set you up well for that. So I think there's a fine balance between being just agreeable enough to help you engage and get better at this versus being so agreeable that you're actually worse at this.
I want to close with what's possible going forward. Maybe let's speculate on new platforms or form factors that could be game-changing. OpenAI just acquired Jony Ive's company. Bryan, I've heard you talk a bit about glasses and why you're still excited about that form factor. Maybe we could start there, but I want to hear from the group on what they could imagine as something that's additive or even disrupting some of the mobile use cases.
There are 7 billion mobile phones out there. There aren't that many devices at all that actually get to that level. My thought process is that either it will live on mobile—and there are many different ways to think about the future there, where there's a privacy wall around it—or it will be a local LLM or local model that helps you really contain all the things that you want to contain at the device level.
So I think I'm still very much excited about the model-development layer to get to that, and I think that's what I'm actually most excited about. If you think about always-onness, as Olivia said, mobile is always on, but there are other things we also have always on. What does that look like when there are net-new devices, or appendages, if you will, that actually attach to things that you always have and enable that?
Any speculation from you guys? Is there a piece of hardware or something that we're going to be wearing, carrying around, or using that's either attached to the phone or separate from the phone that could enable us?
I think AI has scaled for consumers tremendously well, given that it's mostly been text box in, some output in a web browser out. I love the idea of AI actually being with you and seeing what you see.
It's funny: now when I go to tech parties, a lot of the under-20s are wearing pins that record what they're saying and doing, and they find real value from them. That's one example. We've seen a new wave of products that can see what's happening on your screen and take action for you, help you, coach you, and do other things like that that I also find really exciting. As the agentic models get even better, it goes beyond just suggestions to actually doing work for you, sending emails for you, which is very exciting for me.
I think the human-insight layer of that is big too. Often, we have no way of measuring ourselves compared to other people or where we exist in the world. If an AI can hear all of your conversations and see everything you're doing online and say, “Hey, look, if you spend 5 more hours a week doing this, you would actually be a world expert in this topic. Based on this vast network of other people I'm serving, you should connect with these 3 other people. This person could be an amazing co-founder. You should date this person.”
That, to me, is the ultimate sci-fi vision, which comes from AI being with you all the time and something that's not just a ChatGPT text box.
Totally. I mean, the device that has been most widely adopted post-phone is the AirPods. So that feels like the thing that's hiding in plain sight. There's a whole bunch of social-protocol questions around it because it's weird to have your AirPods in at dinner. No one does that, right?
But there may be a way that you can integrate AI and also fit the current social protocols around AirPods. It would be interesting.
You said something that we glossed over, but young people at parties are recording their conversations. In the future, is everything going to be recorded? Do you think that generation is already growing up with that norm to some degree?
Yeah, I think there'll be new social norms developed around this behavior because I think it's real and it's valuable. It's scary, I think, for a lot of people that this is happening, but I think it's a wave that started and is not going to stop.
I think the context matters too. A lot of what you're talking about is the San Francisco networking parties, where work and personal stuff really blur. We talked about this—you could do that in San Francisco. I did that party and brought a pin in New York, candidly.
But I think that's why there'll be a new set of cultural norms. When the cell phone was introduced, there were places where it's rude to take a loud call. The same set of things will emerge around these recording devices.
Yeah. Well, let's end on the idea that we're very early. Guys, this has been a great conversation. Thanks so much for coming on.