2025年末,消费者AI走到哪一步了?
- 2025年末,消费者AI呈现出“赢家拿走大头”的格局:ChatGPT拥有8亿–9亿周活跃用户,而在ChatGPT、Gemini、Claude和Cursor中,只有9%的消费者为不止一款产品付费。 今年大部分时间里,访问过其他头部供应商的ChatGPT用户不到10%。Olivia还提到,Gemini在web端新增规模估计相当于自身规模的35%,移动端约40%;Claude、Grok和Perplexity则都在8%–10%左右。Anish Acharya给出的品牌定位很简单:「ChatGPT 就像 AI 界的 Kleenex」(“ChatGPT is like the Kleenex of AI”)。
- Gemini之所以构成现实威胁,是因为病毒式传播的创意模型与加速增长同时出现:其桌面用户同比增长155%,ChatGPT为23%。 在Android上,Gemini的移动端规模约为ChatGPT的一半,但在iOS上只有17%——它“无处不在”,却仍然“无处落脚”于消费者习惯中。Justine Moore认为,只要持续推出图像和视频产品,Gemini就有机会追上;但ChatGPT的引导式模板让用户第一次创作远比Gemini的空白输入框容易。
- 今年消费者模型的突破,在于图像和视频模型把真实感、推理、检索与多媒介能力结合到了一起。 ChatGPT 4.0 image 的 Ghibli 时刻、Sora 2、Veo 3和Nano Banana证明,准确细节、搜索支持的Logo、一致的人物,以及与视频结合的音频,都能制造病毒式需求。下一代架构将是“输入任何东西,输出任何东西”,有望把文本智能、图像、视频和编辑能力合并进一个模型。
- 实验室的分发能力并不会自动带来成功的垂直产品,这构成了讨论中最明确的2026年创业机会。 Pulse、Atlas、群聊、Sora、Stitch、Gems和Opal都还没有成为爆款的独立消费者界面;NotebookLM是显著例外。Bryan Kim的保留意见是,只要产品核心仍是文本输入、文本输出,高频助手就很难被取代。
- Sora 2证明了市场需要AI视频创作,但还没有证明市场需要AI原生社交网络。 一小批创作者为TikTok、Instagram、X和Reddit生产内容,但应用内消费、二创和评论似乎没有最初想象的那么强;更准确的类比是“CapCut”,而不是TikTok。Bryan的看多逻辑是,幽默可能通过提示词技巧和文化敏感度,创造一种新的身份地位游戏。Anish则追问:如果视频最终都要导出,那么用Sora视频做TikTok,是否会“严格更好”?
- 近期最有防御性的市场可能是专业消费者和企业工作流,因为使用深度足以反转传统的消费者经济学。 据称,ChatGPT企业端使用量同比增长约8–9倍;Claude和Comet则展示了持久化工作流与跨工具上下文的价值。订阅费之外的用量收费,已经催生出收入留存率超过100%的消费者AI产品:“也许整个AI行业其实都是高级用户的故事。”
- 算力仍是战略约束:实验室必须在训练与推理之间、娱乐流量与编程智能之间分配资源,而专注于应用的公司则没有这种内部冲突。 Anish表示,据他所知,xAI“可能是唯一一家”没有受到算力瓶颈制约的模型公司;只做自研模型的实验室,也给服务高级用户的多模型产品留下了空间。如今模型质量已经足以“构建一个真正可规模化的应用”,讨论最后的希望是:2026年会成为消费者应用创业者极其重要的一年。
1. ChatGPT掌握使用习惯,Gemini掌握增长势头
Olivia Moore开场给出的数据,让市场集中度变得具体:只有9%的消费者为多款头部AI产品付费,2025年大部分时间里,访问过其他头部供应商的ChatGPT用户不到10%。ChatGPT拥有8亿–9亿周活跃用户。Olivia还提到,Gemini在web端新增规模估计相当于自身规模的35%,移动端约40%;Claude、Grok和Perplexity则都在8%–10%左右。
这张静态快照掩盖了方向上的急剧变化。Gemini桌面用户同比增长155%,ChatGPT为23%;在Android上,Gemini的移动端规模约为ChatGPT的50%,而在iOS上只有17%——这说明Google的分发能力有效,只是消费者习惯尚未跟上。
对于Gemini能否超过ChatGPT,Justine的回答是“能”,但前提是持续执行到位。顶尖的图像和视频模型会从专业用户与病毒式趋势中制造“几乎无限的需求”,把用户吸引到并不熟悉的Google产品中;Anish则用ChatGPT作为“AI 界的 Kleenex”来概括其品牌优势。
2. 创意模型从审美生成走向有依据的推理
Justine认为,今年消费者模型领域的几款爆款包括ChatGPT 4.0 image及其Ghibli时刻、Sora 2、Google的Veo 3和Veo 3.1,以及Nano Banana和Nano Banana Pro。OpenAI基本把功能留在ChatGPT内部;Google则把产品分散到Gemini、AI Studio、Labs和独立网站,并配以更专业化的界面。
Midjourney依然凭借审美取向独树一帜,尤其适合不擅长提示词的用户;但前沿方向已经转向真实感与推理能力:背景中的行人与车辆能够正确运动,多张图片和文本能够合成为一个完整设计,信息图也取代了过去“至少能把一个字母画对”的胜利。
Anish认为,一个被低估的突破是通过搜索提升准确性。历史准确的场景、真实的产品摄影、市场地图、正确的公司名单和Logo,都同时需要检索与视觉生成;同样,Veo 3的病毒式传播突破点,也在于把音频与视频结合起来这个并不显而易见的决定。
多步骤构图的局限仍然清晰可见。小组提出的测试是:把Monopoly上的每一处地产都替换成AI实验室和创业公司,同时不能遗漏、重复、重叠或放错名称;GPT Image 1.5仍然会在这个任务上受挫。但人物与风格的持续性,已经把重复生成变成了分镜创作:模型一旦做出有用的东西,“你就会想继续生成”。
3. 产品引导与持久化工作流,与模型同样重要
Bryan把Gemini里Nano Banana的空白提示框——“我不知道该做什么”——与ChatGPT类似TikTok的菜单进行了对比:热门风格、一步式变换和后续创意。这些“产品细节”能帮助用户完成第一次创作;他用Snap与Meta的关系作比,暗示Google可以复制成功的交互方式,再与自身分发能力结合。
Bryan认为Pulse被低估了,因为它以每周约25次ChatGPT使用为基础,尝试提供主动摘要和提醒,向西方版“万能应用”靠近。但反弹很快出现:Bryan说自己不是Pulse用户,Anish则基本把它关掉了;产品执行似乎不对,而且“99%的人不会靠日历安排人生”。
如果ChatGPT或Claude能可靠连接邮件、日历和文档,它们就可能“占据专业消费者工作空间”,但现实仍不稳定。Olivia给出的更强案例是Perplexity的Comet浏览器:它提供可重复的代理式工作流,即使分发能力弱于ChatGPT Atlas,持续流量仍高于Atlas,并且有望发展成专门面向专业消费者的界面。
对于复杂的通用工作,Olivia仍然更偏好Claude,因为它“以一种有趣的方式坚持己见”。Artifacts、Skills、文件创建和Claude Code都很强,但包装方式更偏技术用户;据报道,美国青少年使用Character.AI的人数是使用Claude人数的3倍,这说明MCP、Skills和命令行能力还没有转化成主流可及性。
4. AI视频先找到了分发,尚未找到原生社交闭环
Bryan的“起源理论”把产品背后的情感需求拆开来看。ChatGPT最终对应的是“帮我变得更好”;TikTok及类似社交产品则满足“娱乐我——我想看我的小丑”和“我很孤独,我想被看见”。加入群聊,并不会自动把一个生产力产品从第一类推入后两类。
Sora 2的cameo功能是一次强力下注,但实际行为显示,它更像创作者工具。Bryan的内容流大约已有三分之二由AI生成,其中一半以上来自Sora;但创作者集中导出作品,原生消费、二创和评论似乎没有最初想象的那么强。更准确的类比是CapCut。
这场分歧值得保留:Bryan认为AI生成媒体会削弱个人“身份地位游戏”,但他的看多逻辑是,幽默可以通过提示词技巧和文化敏感度创造另一种地位游戏。Anish回应称,这已经是另一种产品,并追问:如果视频最终需要导出,那么用Sora视频做TikTok,是否会“严格更好”。
Justine认为Meta最强的AI工作是SAM 3分割,而非消费者产品;Bryan则强调Instagram的AI翻译功能,它可以复制创作者的声音并翻译成5种语言,同时完成口型同步。Grok在图像和视频上的“斜率最陡”,迅速加入文生视频、音频、口型同步和15秒片段。Elon已经表示,希望在明年年底前实现互动游戏和电影。
5. 企业应用与多模态能力定义下一场平台竞争
据称,ChatGPT企业端使用量同比增长约8–9倍。强制性的职场使用可能进一步强化消费者习惯;其Apps SDK和Apps Directory则可能把ChatGPT变成跨工具的工作流层,这对SaaS供应商有直接影响,而不只是新增一个消费者分发渠道。
Justine对架构的预测是“输入任何东西,输出任何东西”:图像、视频、文本、模板和参考材料进入同一个系统,经过编辑或重新生成后输出媒体。根据她与各家实验室的交流,它们正试图把此前分开的文本推理与生成工作合并成“一个超级模型”,对设计领域的影响尤其大。
Olivia对行业的宏观判断是“更多相同的东西”。实验室会继续改进模型和核心助手,但数十次带有鲜明立场的消费者界面尝试都没有成功;在大约20个Google实验中,NotebookLM或许是唯一的成功案例之一。这给创业公司留下了机会:把能力越来越强的底层模型垂直化。
Bryan的保留意见是,纯文本输入、纯文本输出的创业公司很难对抗高频助手。他仍预计实验室会自行做出常见的应用类型,但Opal这类产品“悄无声息地上线”,本质上只是单模型尝试。Anish强调,创业者可能拥有优势,因为推广机制更偏好围绕核心指标的安全延伸,而不是高风险、强判断的产品。
6. 高级用户正在改变产品栈与消费者经济学
最不性感、却最现实的约束是算力:实验室要在训练与推理之间分配稀缺产能,再在Ghibli式娱乐内容与编程智能之间做取舍。Anish的谨慎判断是,据他所知,xAI“可能是唯一一家”没有受到算力瓶颈制约的模型公司;应用创业公司则不需要面对同样的内部资源分配。
多模型产品也有结构性机会,因为实验室和大科技公司仍然只使用自研模型。单一模型可能已经能提供“你所需要的80%”,但高级用户的变现深度足以让人相信,“也许整个AI行业其实都是高级用户的故事,其他人只是流量”。订阅费之外的用量收费,已经催生出收入留存率超过100%的产品。
Justine推荐Google Labs的Pameli,认为它展示了代理与生成结合后的形态:输入一个商业网站URL,它会抓取产品和品牌照片,总结品牌的审美与定位,然后生成3组覆盖文案、帖子、传单和产品图像的营销活动。Bryan披露a16z投资了Krea,但仍更喜欢通过Krea使用Nano Banana Pro,因为可复用的人物、物体和风格,省去了反复上传参考图的麻烦。
Anish每天使用ElevenReader,把保存的文档转换成步行时收听的音频;Olivia则选择Gamma制作演示文稿、Granola记录带上下文的会议笔记,以及Comet搭建易用的AI原生工作空间。Olivia还推荐Wabi的受限式应用生成能力,以及在Codex或Cursor中使用GPT-52——即便是知识工作场景也适用。最后,Bryan表示,今天的模型已经足以支撑“一个真正可规模化的应用”。
For most of the year, less than 10% of ChatGPT users even visited another one of the big LLM providers.
You open Gemini, and it has a pop-up that says, “We got Nano Banana. Would you like to do something with it?” There’s a little bit of pain where you have to type something.
Yeah. I don’t know what to do.
These are product nuances that I think make people actually take the first step.
The models have gotten to the level of quality where you can build a real, scalable app on top of them. The hope is that 2026 will be a huge year for consumer builders.
Today we’re talking about who won consumer AI in 2025. This was arguably the year that we saw the big model providers—OpenAI and Google, more than anyone else—make a major push of their own into consumer, both in terms of new models they released and in terms of new products, features, and interfaces that target the mainstream user.
You might wonder, why does it matter who is in the lead here? There are some early signs that the general LLM assistant space might be trending toward winner-take-all, or at least winner-take-most. Only 9% of consumers are paying for more than 1 of the group of ChatGPT, Gemini, Claude, and Cursor.
And for most of the year, less than 10% of ChatGPT users even visited another one of the big LLM providers, like Gemini. If we had to call it now, ChatGPT is currently in the lead by far, at 800 million to 900 million weekly active users. Gemini has added an estimated 35% of its scale on the web and about 40% on mobile, and everyone else trails significantly. Claude, Grok, and Perplexity are all at about 8% to 10% of the usage.
But especially in the last 3 to 6 months, things have been changing very quickly. With the launch of new viral models like Nano Banana, Gemini is now growing desktop users 155% year over year, which is actually accelerating even as it reaches more scale. That’s pretty crazy to see. ChatGPT is only growing 23% year over year.
We’re starting to see players like Anthropic almost specialize within consumer, owning different verticals, like the hypertechnical user. So today we’ve brought together the a16z consumer team to recap what we saw this year from the big model companies in consumer and also to predict what might be ahead of us in 2026.
Cool. Well, thank you, Olivia. It’s been a super fun year. If we wind the timeline back to last January, maybe we should start with what we saw launch—products, what worked, and what didn’t. Justine, tell us what you saw this year. OpenAI, Google—what are you paying attention to? What have you changed your mind on?
Those two in particular had a ton of consumer launches, like Olivia mentioned. From a model perspective, I would argue that their most viral models this year, at least among consumers, were in image and video.
For OpenAI, it was the ChatGPT 4.0 image, the Ghibli moment, which is crazy that that was this year.
It’s crazy that this was this year. It feels like it was years ago.
And then Sora, obviously—Sora 2. For Google, it’s Veo, Veo 3, and Veo 3.1, and then Nano Banana and Nano Banana Pro in image models, which went insanely viral, probably comparable to, if not beyond, the Ghibli moment for OpenAI.
I think in terms of the product layer, what we saw was that OpenAI tended to keep more things in the ChatGPT interface. polls, group chats, shopping, research, and tasks all launched inside ChatGPT as the core. The exception there is obviously Sora, as a standalone video app.
Whereas Google tended to launch more things as standalone products. They shipped a lot through Google AI Studio, Google Labs, Gemini, and the plethora of Google services available for launching a product. But they would also ship things as standalone websites that you could go to and visit, which basically allowed for a more custom interface for each type of product—not just the kind of chat entry, chat exit, or image video exit.
Well, Justine, I have a question for you on that. It felt like 18 months ago we were talking about Midjourney, and most of the multimodal models were defined by aesthetics and realism. Is that still true? What changed this year?
There were definitely different styles still. Midjourney, when you talk to people really deep in image and video, still kind of stands apart for this aesthetic sensibility that a lot of the models don’t have if you don’t know how to prompt for it.
But I would say this year in particular, we made a lot more strides on realism and also on reasoning within both image and video. That includes all of the little details that make an image or a video actually seem real. For example, if you have a person walking and talking, the people and the cars in the background, if they’re on a street, should be moving in the correct direction. They shouldn’t be morphing and looking strange.
In image, we were able to have multiple input images and text and reason across all of those uploads to create a cohesive design or something like that, which was not something we saw happening last year for sure.
Yeah, I remember when we were excited about having a letter show up correctly in images. Now we have insane infographics. We can just put up an amazing YouTube video and say, “Give me an image that explains this.”
Yes.
That’s incredibly different. Nano Banana Pro can even generate market maps. I would tell it, “Generate a market map.”
It’s incredible. It either has or will go do the web research within the image model, which is crazy, to get the correct list of companies and then pull their correct logos. That’s insane.
There’s one benchmark left that the reasoning image models have not cracked. I tested GPT Image 1.5 yesterday, and it sometimes struggles with both reasoning and multistep reasoning.
What I’ve been testing is that you upload a picture of a Monopoly board and say, “Remove the names of all the properties and replace them with names of AI labs and startups.” GPT Image 1.5 is actually the closest, but it’s very hard for the models to do all of those steps: remove the names, come up with the new names, put all of the new names in the correct places, make sure there aren’t overlaps, make sure you haven’t mentioned one thing 3 times, and make sure you haven’t left out another big player.
So there’s still some room to go on the image eval. But it’s interesting, especially with the image model from ChatGPT, where you can actually see persistence: it carries a character over into multiple image generations, with the same style.
Yeah, and I thought that was, “Oh, this is actually very interesting. We’re storyboarding.”
Totally. And once it does that, it makes you want to generate more.
For me, it felt like the most underhyped aspect of Nano Banana was the integration with search. It feels like there’s realism, which is physics and other things that make us feel like we’re at the uncanny valley. There’s reasoning, which is applying modifications that are adherent to what the user asked for. But then there’s also accuracy.
A good example of this is product photography. If you say, “Hey, generate a photo of this album cover,” or “Generate a historically accurate photo of this moment in time,” you actually need the search integration. That was nonintuitive, but it’s actually very useful.
Totally. It’s kind of like the Veo 3 moment. I don’t think it was intuitive to people that video would necessarily be cracked by bringing audio together with video in the same place, and that ended up being the thing that made AI video go viral.
Since Veo 3—and now Sora maybe dominates—my social feeds have been full of really realistic-looking AI videos all the time.
One-fifth of my feeds are AI-generated.
Amazing. There have been so many launches this year, and many of them went well, like Veo 3 and Nano Banana. What do you guys think is underhyped, or what products do you think didn’t get enough attention? Bryan?
It’s a good question. I think Pulse is probably still underhyped.
We’re talking about OpenAI and Google, which, to me, fall under the productivity category. If you go to the App Store today, 5 of the top 10 productivity apps are all Google. That’s insane. And ChatGPT is number 1.
We’re talking about a productivity category that helps you do things, and I feel like a lot of people are trying this from different angles: How do I actually ingest your data, your schedule, or your email to make it more helpful and give you more proactive notifications?
A lot of people are working on it. Given the frequency with which people use ChatGPT—which I think is, what, 25 times a week?
Pretty good. Pretty good—3 to 4 times a day.
It feels like it’s in a really good position to actually give you proactive nudges and summaries and help your life in general.
The everything app was always a myth in the Western world. I think OpenAI is trying to move in that direction, where it’s ingesting enough information, and enough people are going there often enough, to start giving really useful proactive nudges. That’s a space I’m excited about.
It’s interesting. But are you a DAU?
I am not a DAU.
Pulse?
Not a Pulse user.
Similarly, I tried Pulse for a while and have largely turned off of it. But I would agree with you that Pulse and a couple of other examples that OpenAI launched this year are new primitives or ideas that feel underhyped, even though the execution is a little off.
The execution—the UX is off. Another example that would be similarly about personal context is their connectors. Now you can connect your calendar, your email, and your documents. You can do this on Claude as well.
And so hypothetically, you could say to ChatGPT, “Read all of my memos over the past 6 months and summarize what's most interesting and least interesting.” I think when that works, it's really exciting. I have found it to be a little bit unreliable so far, but I think as the models get better, they have a real chance to own the prosumer workspace if they get that right.
Prosumer is the perfect category, because we talk about this sometimes, but 99% of people don't run their life on a calendar. We do, but that's what I'm thinking about: the actual average frequency of using ChatGPT. If it's 24 times a week, that's a pretty good place to start. Olivia, I feel like you are the ultimate power user. What are you still using? What's your stack?
That's a great question. From all of the larger model companies, actually, I would have to say the thing that I'm still using the most—and was maybe the most impressed by this year—was the Perplexity Comet browser. I was not using Perplexity as my core general LLM assistant; I use ChatGPT and Claude much more. But I think they really executed on it in a first-class way, in terms of both the agentic model within the browser and, perhaps more importantly, all of the workflows that you can set up that allow you to basically run the same task over and over, either at a preset time or when you trigger it on a certain webpage.
To me, that was a really exciting launch. If you look at the data, the spike at launch and the sustained traffic for Comet was actually much higher than for ChatGPT's own browser launch, Atlas, which is kind of crazy given how much more distribution ChatGPT has than Perplexity. I think they also launched an email assistant this year, and they made a couple of acquisitions of really strong agentic startups. What I would love to see from them next year is more of these dedicated prosumer interfaces. I feel like that would be an awesome direction for them to double down in. They do feel like the startup that has the biggest breadth of ambition, alongside the labs and big tech. It's very impressive, just the number of things they've shipped this year.
Yes, definitely. One thing I wanted to ask you, Justine, is Gemini feels like it's having a real moment because of all the image and video models. Do you think it can overtake ChatGPT? Is there truly that much demand for these types of models?
I think yes. What I've seen, basically, is that there is always nearly infinite demand for the best-in-class image or video model, because then you have a mix of tons of different people seeing it and wanting to use it. If you're using it professionally—in marketing, entertainment, storyboarding, or whatever—you always want to be using what's at the forefront of the field. So you're totally fine going somewhere other than ChatGPT and Sora to get access to Veo.
Even if you're an everyday consumer, so many new viral trends are created around new capabilities of the best-in-class image and video models. That ends up driving users into different products that they may have never tried before. You might be downloading the Gemini app or accidentally ending up on Google AI Studio, which I know they're trying to make more for developers, to use Nano Banana Pro, which a lot of users experienced in the past couple of months.
The interesting thing about Gemini to me is that, hypothetically, they benefit from the massive Google distribution advantage. If you look at Android, Gemini is at about 50% of ChatGPT's scale on mobile, whereas on iOS it's about 17%, so clearly something is working there. They launched a little Gemini widget within Chrome recently that encourages you to use it. They're launching it within Google Docs and Gmail and other things.
Yeah, but I think that most average people are still just using one AI product. ChatGPT is like the Kleenex of AI. It is the brand that has become synonymous with the category. I think Gemini still has a pretty big hurdle to overcome just in terms of that, but if they keep doing what they're doing with these amazing viral consumer creative-tool launches and model launches, they could get there next year.
What do you think about this? It's really interesting when you look at Gemini, which is everywhere.
Yeah, but yet nowhere, to some extent, right? When you look at the actual usage, people still think of Kleenex and go to ChatGPT. The interesting thing is also the product sensibility. This morning, I had 2 panes open: OpenAI's image model and Google's Gemini. I basically used the image functionality. When you open Gemini, it's a blank screen with a pop-up that says, “We got Nano Banana. Would you like to do something with it?” It's the little pane where you have to type something. I don't know what to do.
ChatGPT, you go in and it has a very TikTok-like style: “Here are the trending themes that you might want to generate.” You click on “I want a sketch pen,” or whatever, and they just use one of their pictures and create something amazing. Then it says, “Would you like a holiday card? Would you like a...” These are product nuances that I think make people actually take the first step to generate something. Once you have it, you have character consistency, so you keep going.
That's interesting, in that I think OpenAI and ChatGPT have proven that there's deeper product sensibility. This is a funny, maybe slightly non-kosher thing to say: I worked at Snap. When you look at Meta versus Snap, famously, Evan Spiegel was chief product officer at Snap. I wonder if there is a world where the ChatGPT team innovates on the product front again and again, and Google, with distribution, looks at them like, “That's cool. Let's just integrate it and keep going,” and actually plays that game.
The interesting thing there is that ChatGPT Images just launched yesterday, when we're filming this, in ChatGPT. Brand new. They had image models for years, and it took them that long to come up with a separate, relatively basic interface for generating images. I would almost argue that the application-layer companies, like Krea, Hedra, and Higgsfield, popularized that template format, did it first, and did it better. They are ChatGPT's product people.
Exactly. Flawless. Flawless. Well, maybe going in a slightly different direction, BK, I'm very curious for your take on OpenAI's social features, because it does feel like that's something where you really have to get product execution right, but also network design. There are some efforts around Sora 2—we should talk about that. There's also group chats within ChatGPT. You're our social guy, or have been historically. Bullish, bearish? Where's your head at?
Bearish for now.
To me, the reason is 2-fold. Historically, we look at products based on what I call inception theory: you go 3 to 4 layers down to figure out what the 1-liner is, which is, “I want my dad to love me.” When they think of a product, they're like, “That's for me,” as well as for a lot of you. I look at some of these products, like ChatGPT. Ultimately, when you peel the onion 5 times, I think essentially it's, “Help me be better. Help me get that information. Help me be more productive. Help me be more efficient.”
And then when I think about social features—Meta, Instagram, whatever, or even TikTok—the 2 layers of information or emotion that it's trying to address, to me, are, for TikTok, “Entertain me. I want my clown. Entertain me.” And then the other layer is, “I'm lonely. I want to be seen. I want to connect with people.” To me, these are 2 very different parallels in the product direction.
OpenAI's product is incredible. It's magic. It's amazing. But it's ultimately a “see me” or “help me” category, which is essentially why it's number 1 in the productivity category. Now we're trying to take this and shove it into people's lives and say, “Guys, connect. Connect better and actually feel like you're being seen.”
Even the group chat function, which I love, will be so good to plan a trip and actually have that common pain. But I think it still stops at probably a headcount of 2 to 3 people planning something in a “help me” way, versus, “Oh, I feel like I understand Anish so much better because I've sort of done that.” Largely, over time, I think that's the reason for that division. But that is not to say you can't build a separate product that completely addresses that.
I think Sora 2 was the other big social push this year from all the consumer AI companies. It was basically like a TikTok feed, but with all AI-generated video, and you could make cameos of your friends.
The cameos were a very good bet. It was a strong bet.
Yeah, and I think what we've seen in the retention data and how we're seeing it used is that it was massively successful as a creator tool. Now my feed is probably 2/3 AI slop, if not more, and over 50% of it is now Sora, whereas before it was all Veo and some claim. But it has not been as successful as a social app.
People—a small number of creators—are creating a ton of content and then bringing it out to TikTok, Instagram, X, and Reddit, where it's going massively viral. But it doesn't seem like there's as much consumption happening in the app, as much remixing, or as much commenting, especially as there was initially.
You know, in a funny way, the way I think about it is that Sora's competition or analogy isn't actually TikTok. It's actually CapCut.
Hmm. It's a funny way. It's almost like a creative tool.
Yeah, interesting. I think what I was going to say is that it goes back to your earlier point: the kind of motion that drives social apps is both these positive and negative feelings of, “Oh, I'm publishing this thing of myself that's kind of sensitive, or that I want people to think is this or that or this other thing.” That's what drives participation on the app.
Yeah. The status game.
Yeah, it's exactly the status game. And when it's AI-generated content, and people know it's not real—like, a real representation of you as a human being—the status game is lost a little bit.
Absolutely lost. I think the status game then becomes: can you prompt something very cool? But that's a different type of product, and that's why I think it goes viral on Twitter and all these other existing platforms.
My counterpoint, or bull case, for Sora 2 is that I actually think the status game was about humor more than anything else. And humor is the intersection of knowing how to prompt and being culturally aware. So I think that if they iterated on that, that's a direction that nobody has captured before.
Yes, but if you can export those videos, isn't it true that TikTok with Sora videos on it is strictly better than Sora? We talked about it so much: the ultimate social product is where consumption and creation both live together, or where the output of it is not native to other platforms like TikTok or YouTube Shorts.
So what do folks think of the challengers? We're talking about Sora 2. I mean, Meta—it's crazy to talk about Meta as a challenger. I guess in this context they are, but I think Claude, Perplexity, and Grok are the more obvious names for challengers. Olivia, what's your take?
I love Claude. I talk to Claude all the time. Claude has replaced ChatGPT for me as my general LLM. I think Claude is opinionated in an interesting way. I also love Claude because I'm willing to invest time into building out AI workflows. I think Claude actually launched a lot of really powerful things this year around Artifacts and Skills, where you can essentially set up tasks or workflows to run over time.
I do think the reason it hasn't hit the mainstream yet is that even the way they built those things is geared toward a technical user or an engineer. I think they tried to make Skills as easy as they could to create, and it still was not anywhere near easy enough for the mainstream consumer.
Another example would be that they were actually the first of the big players to launch file creation, slide-deck creation, and editing, and they branded it as file generation and analysis or something. It was a toggle feature within a setting bar of a setting bar or something, so very few people used it. And yet, to me, it's still the best product across all of them at doing that kind of complex work.
So I love Claude, but I think if they want to be a true mainstream consumer product, they need to dumb it down even more in terms of accessibility.
There was that survey you found recently of U.S. teens.
Yeah, I think it was 3 times more U.S. teens have ever used Character.AI than have used Claude.
Yeah. So I think that shows that Claude is a pretty broad thing.
Yeah. Claude is beloved amongst tech people, but outside of tech people, I think they are maybe struggling to pick up relevance.
It is interesting, though: if you look at the aesthetics, the product design, and the craft, 3 things that Anthropic did were MCP, Skills, and the command-line interface, Claude Code. Those are 3 surprising bets, especially Claude Code. I would have said, “Command-line interface, really? Is this the way that people want to interact?”
You were going to talk about taking over air mail and the thinking cap.
Yeah, that too. So, 3 things—where's the thinking cap? But it's sort of very high-minded design. I don't know if it's mass-market, or maybe that's apologetic on their behalf, but I think it is that it's opinionated and it's great.
Yeah, yeah. I do need to hear Justine's take on both Meta and Grok, as I feel like they both had fascinating years.
Yes. Meta hired all those researchers. I think their strongest models are actually not consumer-facing models. It's their SAM 3 series—Segment Anything for video, image, and audio.
Basically, for video, for example, you can upload a video and describe in natural language, “Find the kid in the red T-shirt,” and it will find and track that person across the entire video, even if they're coming in and out of the frame. It will let you apply effects like blurring them out or removing them or whatever. And you can imagine a similar thing with audio, with different stems, and then with images, with different objects in an image.
I think we're going to see next year, hopefully, some incredible consumer products built on top of those models, but today they're more of a playground for developers than they are a consumer-facing product.
Yeah, given the DNA of the company, the one good consumer feature I think they've launched this year with AI is Instagram AI translations. When you're uploading a Reel now, you can opt in to enable translations, and it will clone your voice, translate it into 5 different languages, apply the translation with your voice, and then redub it with lip sync.
Wow.
So it basically makes it seem like you're a native speaker in whatever language. I would love to see more of that stuff come to the Meta products.
Grok had a crazy year with the companions, with all of the LLM progress and the coding progress. I think their image and video progress is probably the steepest slope I've seen of any of the companies. It was probably 6 months ago that they didn't even have image and video models, and they're shipping so fast to launch new features.
Initially, it was just image-to-video; they added text-to-video, they added audio, then they added lip sync with speech, and then they added 15-second videos. They're just not slowing down the speed of progress, and Elon has made a bunch of statements about wanting more interactive, video-game-type content out of Grok and wanting movies out of Grok by the end of next year. So let's hope it continues to go at that pace.
Do you feel like it's a pincer movement where, on one hand, there's a very infrastructural model layer of, “Let's get to the top of the LLM arena charts,” and then the other one is, “Let's go, entertainment”?
I think that's a little bit of a bifurcated move. Right, the entertainment and the smart side.
Absolutely, but entertainment in a way that we're talking about Anthropic and ChatGPT's general population, but you just said Character.AI is way more popular.
Yes. So then how do we think about that? And I think it's a very interesting strategy in my mind. In the image and video app, since pretty early on, Grok has had templates of popular things, like you're standing somewhere and suddenly a rope drops from the ceiling, and you grab onto it and it swings you out of the scene. They have some really good ones that go viral regularly on TikTok and other places.
Yeah, really, really interesting. Well, maybe switching gears from 2025 to 2026, what are some of your predictions for next year? What do you think we'll see—hardware, models, commerce we haven't spoken about yet? What do we think will play out?
I know we're talking about consumer, but one of the things that's been really underrated for me about ChatGPT that we might see more of next year is that they've really made a push into the enterprise, both with the traditional enterprise licenses and then working with specific companies, even training models for them.
And I think when we think about the fact that most consumers only use 1 general LLM product, ChatGPT Enterprise usage—they published a big study—but it's up somewhat like 8 or 9× year over year. And so we're entering a world now where people have to use ChatGPT for their company or as part of their work. That could really translate into consumer usage, or maybe they become the workspace with the connectors and some of the other things that they're investing in, and someone else owns the consumer use cases.
I think, to that end, we have to talk about their push into apps, and I think whether or not that works is going to be the defining question for them next year.
Yeah, and I think we've all discussed the importance of the Apps SDK and the Apps Directory, as they're calling it, and it's going to be a huge new channel for consumers. I think what's less discussed is that it's hyper-relevant to enterprise.
Where ChatGPT shines is where it's able to operate across a number of tools for 1 workflow. And if you think about the number of things you do in your business day-to-day that operate across many tools, it's most of those things. So I think that will have very interesting implications for the SaaS ecosystem, and it's a part of the app store we're not talking about as much.
Yeah, maybe less of a prediction, but thinking through 2025 and talking about all the big moves from the big labs and from the startup point, I think one of the biggest trends we've seen is app generation. And I think there's a real world where we see the big labs, with the distribution and the frequency of usage of people coming in, start to say, “Look, maybe there is a common type of product and apps that we could actually help you generate within the confines of the big lab products.”
Yeah. I think that’s one of the interesting things which, again, going back to the supply chain of ideas and research, maybe that’s one thing. Again, nothing groundbreaking, but as we know, Studio Ghibli broke the internet. My cousin, who knows nothing about tech, sent me a Studio Ghibli photo. Well, let’s not send this to your cousin then.
And I think that goes to show that templates matter, that style matters.
Yeah. And I think about video, and it’s pretty freaking good.
Yeah. And it’s possible that we’re already at a point where it’s not necessarily just about the capabilities of the models of the big labs, but the stylistic things, the template. Think of TikTok. The core capability is largely still the same: music, trend, dance, go. Except the trend and format keep changing and keep it extremely fresh.
So, I feel like there’s a real world where the repurposer, or a team, or what have you, can start thinking about ways to really build video-first products, entities, live models. I think the cost will go down enough for people to try it out, and I’m excited to see that.
Yeah, I think what I’m most excited about is, along those lines, basically everything becoming multimodal. I call it anything in to anything out, which is basically—initially, especially with these image and video models, you put in a text prompt and get an image or a video out. You couldn’t really do much with it. And now we’ve started to see this with image-edit models, like Nano Banana, FLUX, and the new OpenAI model, where you can put an image in now and get another image out. You can put an image in with a text prompt and a direction, or put an image in with a template and another reference image, and get another image out.
What happens when you can put a video in and get images out that are related to the next iteration of the video? Or you can put a video in and a text prompt about what you want to edit and get the edited video out? From my conversations with the labs, a lot of them are trying to combine all of these largely separate efforts they’ve had across text reasoning and intelligence, the LLM space, and image and video into—what if we can merge those all into a mega-model that can take a lot of different forms of content and produce much more? I think it’s also going to have huge implications for design, because if you think about it, a lot of design is combining images with text, with video, and with different elements in interesting ways.
Yeah. I guess if I think about a macro-level prediction, I think it’s actually going to be more of the same. When we talk about what all of the labs have launched for consumers, they’ve done a great job with models, and they’ve done a great job with incremental things that improve the core experience of using something like ChatGPT or Gemini. In my opinion, we’ve gone through dozens of things that they’ve launched or tried as new consumer products or new consumer interfaces: group chat, Pulse, Atlas, Sora. Google has had a long tail, like Stitch, Gems, Opal—tons.
Yeah. None of those are really working, and I think it’s because it’s not the core competency of these companies anymore to build opinionated standalone consumer UIs. Out of all of those, I think the product that’s working the most is NotebookLM, and that’s one of maybe 20 things that Google has tried or experimented with.
So, I think it’s actually very positive for consumer startups in that the models will keep getting better, which the startups can use, and they’ll keep making ChatGPT better and better. But I don’t necessarily think that ChatGPT verticalizes into all of these other amazing use cases or products, and there’s still room for startups to be building there.
I have a yes-and to that: absolutely. But, however, when the input and the output are text—yeah—ChatGPT and Gemini shine the most. No matter how deep you go, no matter how specific you think your text output is going to be, essentially, given the frequency of use by users of the main big-lab products, I think it’s going to be really hard to stitch that and get that away from that usage if your product is mainly text in, text out.
So, I do think you have to be creative around what the angle is that you can use to steal people away from.
You know, I love that you used the word opinionated, because I think that for labs—certainly for big tech, and perhaps increasingly for labs—the priorities get set in their promo committee always. And if you’re a PM, and it’s always the sort of mid-career PMs—and I’ve been one of these—the incentives are always to get promoted, and the way to get promoted is to build something safe that extends a core metric and a core feature.
So, building opinionated products is a very risky way to manage your career, because they’re probably not going to work. They’re probably going to have a bunch of implications for legal and compliance, and the CEO might yell at you. I just think that they are so structured to do incremental things. The more founders do opinionated things, the more advantage they have.
I think, honestly, the big thing we haven’t discussed here, too, is compute, which is that the labs have this inherent tension: there’s a limited amount of compute, and they either spend it on training models or they spend it on inference. And even with inference, there’s this split between the entertainment Ghibli use cases and the coding-intelligence use cases.
I think xAI is probably the only model company that is not bottlenecked on compute, from my understanding, whereas the others have to make really serious and significant calls, like: if we release Nano Banana and it goes super viral, it may slow down the next big LLM we’re trying to push forward. Startups that focus on the app layer don’t have that problem, because there’s no tension there.
Absolutely. Yeah. We’ve talked about this before. I also think that there are categories in which being multimodal allows you to deliver a better proposition to the customer, and the labs and big tech are always going to be, definitionally, first-party-model-only.
So, I think as all the models get better, perhaps 80% of what you need can be achieved from a single model, but for the power users—and so much of AI is a power-user story—you always said that power users are just power users, and I think that’s true in a pre-AI world. But now the depth of value and the depth of monetization is so much higher that maybe all of AI is actually a power-user story. Everyone else is just traffic. Yes.
Yeah. Which is why we’re also seeing consumer products, for the first time ever, have more than 100% revenue retention.
Yes. And that’s what separates the good from the great from the exceptional in consumer AI. To be clear, how that happens is that they charge for usage, often in addition to a subscription. So, you can use beyond whatever your quota is for the month, given your subscription, and pay more.
An upgrade of the tier, or actually buying tokens or more usage. Yeah. That’s what differentiates it. If you told me, pre-AI, that we saw a consumer company with 100%-plus retention in revenue, I’d be like, “That doesn’t make any sense. That doesn’t compute.” Yeah. No pun intended. Exactly.
Well, guys, okay, maybe let’s start with specific recommendations. After this pod, what are the products people should download, or the features or the models? What should folks be using today?
On the multimodal point, I think one really underhyped product that people should check out—not because they’ll use it every day, but because it shows what is possible when you combine an agent with images and text—is Pameli. This is the Google Labs product where you put in the URL of your business, and it has an agent go to the website, pull all of the product and brand photos, summarize what it thinks your brand’s aesthetic is, what it stands for, and what kind of customers it’s targeting, and then generate 3 different ad campaigns for you.
It will generate not only the text, but also the Instagram posts. It will generate the flyer. It will generate the photo of your product wherever it thinks it should be, based on your customer. Very cool product. It would be hard to become a giant standalone product within Google, I think, but it shows the future of what happens if we combine agents with generation models that have a really deep understanding of context that an image model or video model normally wouldn’t have.
Startup products, though. Do you have a favorite startup product in creative tools?
I think we’re investors in Krea, so this is biased, but I think they’ve really done an exceptional job of being the best place to use every model, or every quality model, across every modality, and also building more of the interface on top of these models.
I now prefer to use Nano Banana Pro on Krea because Krea allows you to save elements, which are essentially characters, styles, or objects that you can tag to reprompt, versus having to drag the same image reference into Nano Banana over and over again.
It’s a good one. I suppose it falls under the startup category. Again, shilling companies, but the one that I use the most is actually ElevenReader. And the reason is, we’ve seen an explosion in podcasts, and there’s a reason for that, right? People are a lot more on the go. The reading capability of us reading, I think, is going down over time.
And so, let's not fight the reality. Let's embrace it. Let's find written material, translate it into listening, and do that. I used to be a power user of tools like Pocket. I didn't have time to read everything that I wanted to read, and it's a saving behavior, right? You're going around and saving all the things you eventually want to consume.
But I think what I do now is similar: I go get all the things I want to read, and I either PDF them or put them in ElevenReader. Once in a while, when I'm on a walk and I have 3 or 4 minutes, at 1.5x speed or 2x speed, I listen to one of these and get the gist of it. I think that's been a good way to use a little bit of time as a sort of semi-normal person.
Well, first of all, I love this question because I am strongly opinionated that by far the best way to get up to speed on AI is just to try a ton of products, and you get opinionated really quickly. I'm actually on Twitter for the whole month of December, publishing one new consumer product a day for people to check out.
So that's one way. I'll name 3 others that I think are especially relevant or interesting that people can plug into their workflows. One would be Gamma for slide deck generation. You can go from text prompt to slide deck, or from document to slide deck. I use it for everything.
Also, the slides have flexible sizes, so you're no longer editing every little pixel in your Google Slides to get it to fit into one, which is great. Another is Granola for note-taking. You might not have any meetings over the holidays, but in the new year, it just gets better and better the more meetings you have on it because it has the context of what you talked about before.
And then, lastly, I'm still going to plug the Comet browser. If you want to try an AI-native workspace, I think that's one of the most accessible ones to start with. I mean, for me, I've spent my whole year obsessed with coding and AI code. It's just been so tremendously fun. By the way, Bryan would take the other side of your argument that the big labs or big tech will win at app generation.
I think they just lack the focus. Products like Opal have been released with a whimper, and they're one model only. I didn't say they would win it. I think that we will see them doing it.
Yes, yes. I think that's true, but I think for the pure consumer side, of course, Wabi is really fun and really capable, and I think they're creating the right sort of constraints on app generation so that you can get a really satisfying, functional result. I think so far there's been a lot of overpromising in app generation, which has discouraged the early users.
I also think if you haven't tried GPT-52 in Codex or in Cursor, it's worth trying. Even for nontechnical people, it's just amazing. I think almost being technical is sort of a constraint because you have a pre-existing idea for what these models can do, and they can do a lot more. I'm hearing increasingly about people doing knowledge work and writing essays in Cursor instead of just writing code.
Wow. Just one thing I'm going to do at year-end, just to plug in a popular trend I see on TikTok: there are people who say, “What is the most unhinged thing I said this year?” And it actually does a review of all the things that you said.
But I think, similarly, it'll be a good thing. I'm going to do this at year-end: “Tell me how to live a better life next year. Give me actual, unvarnished opinions and some direction.” I think it'll be helpful.
I love that idea.
I'm going for a worse life next year. [Laughter.]
Fantastic. Let's go full degen, guys. Any closing thoughts?
That was a good one. I mean, the obvious one is that we are very actively investing in consumer companies, and I genuinely—I think a lot of people say this—I genuinely believe that the models have gotten to the level of quality that you can build a real, scalable app on top of them. Wabi is a great example of this.
And so the hope is 2026 will be a huge year for consumer builders, not just consumers being consumers of a product.
Yeah.
Yes. Well, thank you all for a super-fun year in consumer and AI. We'll be back with more next year, and Merry Christmas, guys. This is a wrap.
Yeah. Happy holidays.
Happy holidays.
[Laughter.] Happy holidays.