[BidClub_]
The a16z Show · · 36 分钟

全球 Top 100 GenAI 产品:排名与解读

Steph SmithAnish AcharyaOlivia Moore

YouTube
TL;DR
  • GenAI 榜单在边缘位置仍不稳定,但核心格局正变得越来越稳固。 2025年1月的 Similarweb 访问量和 Sensor Tower 月活用户数据显示,新增17个网页产品;其中16个网页产品已连续出现在全部4期榜单中。对投资者而言,产品仍可能“一夜之间重写榜单”,但反复出现的排名已经说明,持久的消费品牌和企业级业务正在形成。

  • Smith指出,消费端活动通常滞后于研究“6到9到12个月”,而 Acharya 仍认为许多消费级 AI 品类处于早期采用阶段。 当前视频系统能稳定生成优秀的3秒、5秒或6秒片段,但还做不出完整电影;消费级语音、浏览器智能体和理解屏幕内容的产品,在应用层仍相对不成熟。Acharya 更深层的判断是,直觉性假设可能被彻底反转:AI 可能“比人类更像人”,也可能负责组织和委派工作,让人来执行,而不只是接收人的指令。

  • DeepSeek证明,另一个横向助手仍能在极短时间内获得大规模分发。 大规模免费推理和可见的思考过程帮助它仅用10个1月中的日子就登上网页榜第2名,规模约为 ChatGPT 的10%;移动端仅统计5天便排到第14名,Moore称再多5天就能升至第2名。其30日留存率为7%,低于 ChatGPT 的9%;不过,在用户只能“选 DeepSeek,或什么都没有”的市场,比较结果可能存在偏差。

  • 随着新模型把偶发试用转化为多个反复发生的工作场景,ChatGPT 的增长重新加速。 从2023年2月到2024年2月,网页访问量基本持平;Moore当时看到的数据表明,学生贡献了50%以上的流量,此后访问量已翻倍。ChatGPT 自己公布的数据则显示,周活用户在6个月内从2亿增至4亿,而此前一次翻倍用了9个月;o1、GPT-4o、Advanced Voice Mode、Deep Research、Operator 和 Canvas 进一步扩大了产品用途。

  • 陪伴型产品已经具备相当规模,而多模态能力仍远未被充分挖掘。 Smith指出,前10名中有3个陪伴型产品,其中2个面向 NSFW 内容;部分用户把它们当作互动式同人小说。Acharya表示,一旦产品加入更丰富的多模态能力和语音,潜在需求可能大幅增长。

  • AI 视频正在从赢家通吃的模型市场,转向模块化的生产栈。 榜单新增 Hailuo、Kling 这两个中国模型,以及 Sora。Google 的 Veo 2 在测试中表现“像是下一个层级”,外界期待它能在未来3个月或6个月内发布。不同模型分别擅长人物、风景、动漫或超写实内容,这为 Krea 等聚合产品提供了空间:它们可以减少工作流之间的断点,并有机会整合10个或15个每月20美元的独立订阅。

  • Vibe coding 在第一批病毒式下游应用出现前,已经暴露出巨大的创作需求。 Cursor 和 Bolt 进入主榜,Lovable 进入 Brink 榜;Moore引用 Gary Tan 的说法称,“YC 公司中95%或差不多这个比例”如今都在使用这类工具开发。Lovable 的创作站点流量显著高于托管作品,验证了“个人软件”和“一次性软件”的需求,也预示着一个“将陷入混乱”的 App Store。

  • 使用量与变现能力,描绘的是截然不同的消费级 AI 市场。 Sensor Tower 的移动端收入榜与月活榜只有40%的产品重合;据称,严格将核心功能锁在订阅后的专业消费者产品,仅凭100万或200万用户就能做到5000万至1亿美元 ARR。付费获客可以用1至2倍的回收倍数制造出1000万用户,但行业最终仍遵循产品优先:没有注意力和留存,增长就只是一个“漏水的桶”。

摘要 · 为研究而整理的核心内容

1. 使用数据表明消费级 GenAI 仍处早期,市场假设依然脆弱

  • Moore强调,这份榜单依据的是用户行为,而非行业热度:网页前50名由2025年1月的 Similarweb 月度访问量决定,移动端则采用 Sensor Tower 月活用户数据。两套数据随后都筛选出以 GenAI 为核心的产品;单列的移动端收入榜,记录的是 Sensor Tower 能测量到的收入,通常包括订阅和应用内购买,可能不包括广告。

  • 流失与持续存在同时发生。共有17家公司新进入网页榜,而有16家公司每一期都在榜;两期各含5个产品的 Brink 榜,则记录了 Runway、Otter 和 Umax 这些前主榜产品被挤出,同时 Krea 和 Lovable 正在上升。

  • 用户采用的时间线早于 ChatGPT:2022年,Midjourney 和 Character.AI 已经吸引了各自的细分社区。随后 Snapchat 的 My AI 触达约1.5亿人;Balenciaga 教皇图像、“BBL Drizzy”,以及大量由 AI 生成的 Coca-Cola 圣诞广告,帮助图像、音乐和创意 AI 进入消费者与企业的认知范围。

  • Smith指出,消费端活动通常滞后于研究“6到9到12个月”。Acharya认为,市场正处于基础设施建设与应用建设之间,许多品类仍在早期采用阶段。他标志性的反转判断是有意保留假设性的:AI 可能更擅长处理某些关系型互动,因为它们有耐心,而且“永远不会有糟糕的一天”;与此同时,人类可能会享受执行由 AI 组织好的工作。

2. AI 视频正在成为模块化生产栈

  • 共有3个视频模型入榜:Hailuo 和 Kling 这两个中国模型,以及 Sora。Google 的 Veo 2 在测试中表现“像是下一个层级”,Moore希望它能在未来3个月或6个月内发布。Acharya原本预计,风格迁移会成为可规模化的视频路径,因为它更容易实现、推理成本也更低;但研究人员和产品开发者最终仍在推进原生文本生成视频。

  • 讨论嘉宾认为,中国模型的优势部分来自“对版权没那么敏感”的训练数据——这被顺带称为“一个很棒的委婉说法”——Moore认为,这与其更逼真的输出和更强的提示词遵循能力有关。她还提到,中国更容易获得视频标注劳动力,也拥有更大的图像和视频研究人才池。Sora让部分用户失望,而中国模型相对于其融资规模则超出预期。

  • 模型和界面正在变得更具明确偏好:用户可以指定摄像机角度、景别和运动方式,不同系统则分别擅长人物、风景、动漫或超写实内容。Acharya认为,Krea 的聚合能力“整体大于部分之和”:用户可以在 Midjourney 或 FLUX 中生成图像,先进行放大,再把这张图作为视频的首帧,全程不必跨越不同产品的工作边界。

  • Acharya特别提到 Ideogram 独特的文字生成能力和审美风格。Moore还介绍了它的图像转文字功能:用户可以把一个梗图或受版权保护的图像转成文字提示词,再据此生成新图像。

3. Vibe coding 在病毒式产出出现前,已经开始变现潜在创作需求

  • 面向技术用户的智能体 IDE Cursor,以及通过提示词生成网页应用的 Bolt,进入主榜;Lovable 进入 Brink 榜。所谓主流突破仍需打折看待:许多用户依然是技术人员,他们快速做出原型、导出代码,再自行完善。

  • Acharya将正在形成的品类称为“DIY 或个人软件”和“一次性软件”。一个 Bolt 或 Lovable 项目,可能只是一个有吸引力的原型、解决某个个人细分痛点的工具、一次20分钟的体验,也可能最终成为一家由从未学过传统编程的人打造的风险投资级公司。

  • Acharya最初以为,是少数创作者流量极高的网站解释了这些平台的崛起。数据却显示情况相反:人们进行创作的 Lovable.dev,访问量显著高于托管作品的 Lovable.app。迄今还没有大规模病毒式应用浪潮出现,因此 Moore预计,等这轮浪潮到来,市场认知还会进一步提升:“App Store 将陷入混乱。”

  • Smith询问,开发者是否更可能针对具体用途定制 Deep Research,而不是使用横向工具;Acharya回答:“不是更可能,但我认为这块还没有被充分挖掘。”市场报告是 Deep Research 的既定用途,但在追溯一个梗图的来源时,团队发现了一个比 Know Your Meme “好100倍”的版本;受约束的应用可以解决 Deep Research 面临的“空白页问题”。

4. 助手通过实用性、差异化与陪伴扩张

  • ChatGPT 的网页流量从2023年2月到2024年2月基本持平;Moore当时看到的数据表明,学生贡献了50%以上的访问量。此后流量翻倍,而 ChatGPT 自己公布的周活用户则在6个月内从2亿增至4亿,此前一次翻倍用了9个月。

  • Moore将加速归因于一系列能够解锁不同工作场景的能力:o1 推理、GPT-4o 多模态、Advanced Voice Mode、Deep Research、Operator 和 Canvas。她的使用频率从每周一次或更少,提升到开车聊天、研究备忘录和头脑风暴等场景中的每日使用;尤其是推理能力,让一些敏感工作不再像过去那样有风险——那时 ChatGPT 连“strawberry”里有几个字母 R 都无法稳定数对。

  • 陪伴型产品是另一个意外:Smith指出,前10名中有3个此类产品,其中2个面向 NSFW 内容;她还说,用户把它们当作互动式同人小说。Acharya则意外于榜单中没有更多多模态产品;他预计,一旦产品把角色设定与更丰富的语音及其他模态结合起来,陪伴需求会进一步上升。

  • Acharya认为,Claude 并不是那个质量只有十分之一、因而成为传统意义上第2名的产品;它是被更少用户喜爱的产品,但在创意写作、似乎更强的个性,以及“奇怪地好得多”的编程能力上更突出。这种差异化意味着,ChatGPT、Claude、Gemini 以及潜在的其他模型,有空间相互增强,而不只是彼此替代。

  • Moore在1月截取的数据只覆盖 DeepSeek 的10天表现:它在网页端以约 ChatGPT 10%的规模升至第2名,随后凭5天数据在移动端排到第14名;她表示,再多5天,DeepSeek 就能排到移动端第2名。它的产品钩子由两部分组成:以此前需要 ChatGPT 高级订阅才能使用的 o1 Pro 为参照,大规模免费提供推理模型;同时提供极具吸引力的可见思考过程。其30日留存率为7%,低于 ChatGPT 的9%;但受限制的市场可能会扭曲结果,因为用户的替代选项是“DeepSeek,或者什么都没有”。

5. 收入更奖励集中的实用价值,而非单纯触达规模

  • Sensor Tower 的收入榜与使用榜仅有40%的重合。大类仍然相近——图像和视频生成或编辑、美颜增强,以及类似 ChatGPT 的助手——但收入领先公司往往不同于月活用户领先公司。移动端使用也更偏好原生手机场景:头像应用依赖用户已存储的自拍,语音优先产品在手机上更自然,作业辅导应用在移动端的表现也强于网页端。

  • 更小的用户群带来了显著更高的单用户收入。Speak、Otter 和 Captions 这类严肃的专业消费者产品,可以把大部分功能锁在订阅之后;Moore表示,一些公司可能仅凭100万或200万用户就做到5000万至1亿美元 ARR,因此它们未必会出现在月活用户领先者之列。

  • 移动端独立开发者和国际应用工作室,可以利用 App Store 广告及其他相对低成本的获客渠道购买用户。如果能收回1倍或2倍的获客支出,这种模式就具有吸引力;它们可以借此触达1000万用户,但收入未必达到更小、收入更高的产品水平,最终利润也可能更低。

  • 最优模型取决于产品本身。一款植物识别应用可能永远无法覆盖1亿部手机,但坚定的植物或鸟类爱好者可能愿意每年支付100美元,并每天或隔天使用。Moore最后的规则是:围绕痛点或独特体验构建产品——即便使用的是更老的模型,或只有一个 AI 功能——因为技术复杂度无法挽救一个“漏水的桶”式产品。

Steph Smith

Consumer activity typically lags by 6 to 9 to 12 months. What's happening on the research side?

Anish Acharya

I think, compared to where we're going to be, we're still incredibly early. So many of these assumptions—and that's why they're assumptions—seem intuitively correct, but are going to turn out to be incorrect.

Olivia Moore

We are finally on the verge of AI video starting to work—to really work. It sort of follows the trend of AI decreasing the cost of creation in every way.

Anish Acharya

95% of YC companies or something are now building using those tools. The App Store is going to be chaos.

Steph Smith

Yeah. We're back for the 4th edition of the GenAI 100 list. You guys have been working hard and tracking the consumer landscape for years now, but specifically for the last 2.5 years, since we really had that kind of ChatGPT moment. Tell me more about how you're tracking that ecosystem and how that comes through in this list.

Olivia Moore

Yeah, it's super fun. This is one of my favorite reports that we put together a couple of times a year. We track the consumer AI landscape through what we do every day, which is meeting with consumer AI startups that come to pitch us and seeing what goes viral on Twitter.

But there's a whole separate set of companies and products that might be reaching the true mainstream consumer. They might not even be marketing themselves as AI products, but they're powered by and made possible by AI. The whole original purpose of this report was to see how much overlap there is between those 2 categories and what the actual everyday person who might not know that they care about AI is using in their day-to-day.

Steph Smith

That's great. So talk about the methodology. What makes it onto this list or not? Because, to your point, there are certain household names that you might see on Twitter or have that viral moment. But I think some people might be surprised to see what made it onto this list. So let's start with the methodology and what it requires.

Olivia Moore

Yeah. It's entirely based on data. We have 2 lists here: the top 50 on web and the top 50 on mobile. For the top 50 on web, we use a data provider called Similarweb, which tracks every single website globally. We essentially go down in descending order of how many visits they get each month. For this report, it was January 2025, and then we pick the first 50 of those that have the most monthly visits and are GenAI-first products.

We do something similar on mobile, but with a different dataset from Sensor Tower. For mobile, we look at monthly active users on the app. Again, we picked the top 50 that are GenAI products. For the first time ever, we actually looked at the top 50 on mobile by revenue, which we hadn't done before. It was a really interesting experiment because the lists were pretty non-overlapping.

Steph Smith

We've been in this AI ecosystem for a few years now. In your eyes, what were the pivotal moments that led up to this point in time where we have, like you said, 50 on mobile, 50 on desktop, and a whole lot more in the wider ecosystem?

You often say that it's usually the papers that are written, then the models are developed, and then applications are built on top of them. Consumer activity typically lags by 6 to 9 to 12 months. What's happening on the research side?

Olivia Moore

So maybe, just from a consumer awareness or behavioral perspective, there are a couple of moments for me. Midjourney and Character.AI both came out before ChatGPT, which I think a lot of people don't know. There were maybe these early niche communities of early adopters using both of those products in the summer and fall of 2022, leading up to ChatGPT.

Then, post-ChatGPT, there were things that just brought AI to consumer consciousness. I even remember Snapchat's My AI, with that little bot that appeared at the very top of your feed, and 150 million people used it. For a lot of younger consumers, that was probably their first real chance to have a conversation with an LLM.

On the image side, I think of the Balenciaga Pope, which was also, I think, in spring 2023. Such a cultural moment. It made a lot of people realize for the first time that they should even be interested in AI images because they could be that good and that convincing.

The first big AI music moment for me was the BBL Drizzy song, which I think was in spring 2024. That also went mega-viral. I think one of the moments where creative AI really shifted into almost enterprise consciousness was at the end of last year, when Coca-Cola did its Christmas ad, and a lot of that was generated by AI.

Anish Acharya

And then, of course, there was the DeepSeek launch earlier this year. DeepSeek was so interesting because I think it had become settled wisdom that it would be very hard for a horizontal model to get to mass consumer scale quickly again. ChatGPT had kind of done it—ChatGPT had become a verb—and that opportunity had already been explored. Now we see DeepSeek growing as quickly as it did, and there are actually a couple of interesting nuances to DeepSeek.

One important nuance is the fact that they released their reasoning model for free at scale. Previously, you had to use o1 Pro, and you had to pay for ChatGPT's premium subscription to get access to it. The other thing was the product execution around chain-of-thought, which we've talked about a lot and I think is pretty well understood. But the fact that it showed you its thought process in real time was super captivating, and now it's something that every model takes as a step.

I think it just really illustrates how early we are. We, as sophisticated users and investors, are looking for further and further refinements, and once in a while something like DeepSeek comes out of the clear blue sky and just blows away all assumptions.

Steph Smith

That word—specifically, “assumptions”—is so key when you talk about these pivotal moments. I feel like you could actually match each pivotal moment with an assumption. An assumption being, “Oh, well, AI could never trick me into thinking a picture is real when it's not,” right? Or, “I would never actually listen to a top-100 song that's generated by AI.” Or, “ChatGPT has cornered the market; no one else can penetrate it,” right?

All of these assumptions have people saying, “Okay, sure, I was wrong about that prior one, but this one I'm pretty sure about.” We're seeing just months being the delta between assumptions being broken. So, to your point on the arc of the market or the industry, I could see an argument where people are like, “Oh, you know, we're actually pretty far along because we've already slashed all of those assumptions.”

But then, on the other hand, I'm hearing that we still have a long way to go. So, kind of, put us along that arc if we were to compare to the mobile era, the cloud era, or previous technology eras. Are we in that early innovator stage still, or are we somewhere else?

Anish Acharya

Yeah, I think we're still very much in the early-adopter phase in many of these categories. We're arguably still in the infrastructure-building era and kind of moving into the application-building era. It depends on the modality. LLMs—maybe people thought that was a solved problem, but then again, DeepSeek came in and upended all of that.

There are a lot of things that are definitely not fully solved. AI video right now can generate great 3-, 5-, or 6-second clips, but hopefully, years from now, we have AI video that can generate minutes-long or even hours-long movies. Compared to where we're going to be, we're still incredibly early.

Here are 2 assumptions that I think are interesting because it may turn out that reality is the exact opposite. One is that AI will be very good at transactional interactions, but humans will still be the ones to build relationships and connections. An example of that would be: What kind of phone calls is AI going to be best at?

The assumption was that they would be great at scheduling and logistics and exchanging information and facts. But we've heard over and over that, in many cases, AIs are more human than humans. They just have more patience and more nuance. They're never having a bad day. They're never hung over. That's an interesting area of exploration.

The other one that I think is interesting is the idea that humans will delegate work to the AIs, and the AIs will do it. What if the AIs are the ones delegating the work to us? Perhaps AI is really good at organizing work, and we're really good at doing it and also get a lot of joy out of doing it. So I think so many of these assumptions—and that's why they're assumptions—seem intuitively correct, but are going to turn out to be incorrect.

Steph Smith

Totally. If we think about the report, maybe one important data point is the fact that we see so many newcomers still. If we were in that later part of the innovation curve, you might expect more stagnancy. You might expect to see the same players. But every time you build out this report, we're seeing all of these newcomers.

This particular time, in the 4th report, we saw 17 new companies on the web rankings in particular. You actually have this quote where you say, “A few unexpected players rewrote the leaderboard overnight.” Can you just speak to that and the movement that we're seeing?

Olivia Moore

One of the biggest trends among the newcomers is that we're finally on the verge of AI video starting to work—to really work. We had 3 new video models on the list this time: Hailuo and Kling, which are both Chinese models, and then Sora, which was OpenAI's model that was announced more than a year ago at this point and finally was released.

I think we'll see even more of a shakeup here because Veo 2 is the new Google model that's even next-level beyond that, from what we've seen in testing. That's probably finally going to hopefully come out in the next 3 or 6 months.

The other big category of newcomers was vibe-coding products. Cursor made the list; it's more of an agentic IDE for a technical audience. Bolt made the list as well, and it's for a nontechnical audience, where you basically go from a text prompt to a fully functioning web app.

Even though they made the list, I think we've still seen a really significant portion of their users are people who are in tech and are actually technical. They might be using something like Bolt or Lovable, which made our Brink list, to prototype something more easily and then export the code and play with it themselves. So I think we haven't quite seen the vibe-coding products hit the true mainstream user yet, in terms of someone who's never worked in tech or developed an app.

Steph Smith

I love this category. It's so fun and so satisfying to actually see your ideas come to life. In the case of Bolt and Lovable, sometimes they are just compelling interactive prototypes more than they are full-fledged products, but that's usually enough to get a feel for whether this is something you want to invest in more deeply.

It follows the trend of AI decreasing the cost of creation in every way and people just trying more ideas. Just think about what that says about the untapped market of people who wanted to build things with code—that this is on the top 50 list, right?

Honestly, neither of them has had many apps built on them yet that have gone super viral. When that happens—and I'm sure it will—those will become stories of their own, which will then increase awareness of the products among the true mainstream audience. I think we're going to see a really interesting diversity, or range, of products built on these. It might just be, “This is my app that I used for my very specific niche pain point,” or there might be people who never learned how to code who want to build a venture-scale product on something like Bolt or Lovable. Seeing how that plays out will be very cool.

Anish Acharya

I think there are 2 phrases I've heard that I like. One is DIY, or personal software. It never made economic sense to design software for 1 person, really. The other is disposable software. Just as Suno and Udio made it possible to make a song just to capture a joke that would be irrelevant the next day, these products make it possible to create a product or an experience that may have an extremely short shelf life—20 minutes, a week, or any other time period.

Steph Smith

Let's talk about the Brink list, because that's completely new to this year's fourth edition. What is the Brink list, and why add it?

Olivia Moore

The Brink list is essentially the 5 companies that almost made the list and were right below the cutoff, again purely based on the data. We pulled the 5 websites and the 5 mobile apps, and I think honestly we were just curious to see what it would capture.

The takeaway for me is that it does reflect how fast things are changing. There were a couple of companies on the list—like Runway, Otter, and Umax, across web and mobile—that have been on the core top 50 rankings in the past. Maybe they got edged out by DeepSeek launching this time, so they lost their spot for this ranking, but they might be on there in the next one, and they still have massive usage.

The other trend it caught was the rise of more recent products. Krea made the list, and Lovable made the list; both are very much on a consistent upswing. If that continues, we might see them on the main rankings, and they haven't made the main rankings before.

Steph Smith

What did you predict that you would see on the list that you didn't really see there? Were there any surprises on that end?

Anish Acharya

One thing I thought we'd see more of is style transfer as an approach to scalable video, because style transfer is a much more tractable problem and has a much lower cost of inference versus raw text-to-video. But researchers and product developers seem to be really going for text-to-video, and we've seen more of that than I would have expected.

I think the other things that we didn't see on this list that we have seen at the model level—which means maybe they'll be on the next list—are consumer voice products. There are a few of them, but not a ton. Some of the new models, like the Gemini Flash model, can see what's going on on your screen and interact with you. I built something to yell at me if I go on Netflix or something: “It's time to get back to work now. You've got this and can accomplish all of your goals.”

Or take the new OpenAI Operator, which can interact with things at the browser level on your computer and get tasks done for you, like pay a bill, make a graphic design, or hire someone to landscape your yard. I think there's always a lag, because the models have to be released to developers, and they have to be tuned by the developers, so it takes a while. But I would expect to see an explosion of fun, unique, and interesting products built on models like that on the next list or 2, hopefully, because it feels like we've really seen an explosion on the model side, and it's right there in terms of manifesting at the app level, too.

One example of this is Deep Research. If you've played with it, it's completely magical, but it's a primitive, right? It's not a product. It's something to build other things with. So it's really unclear whether Deep Research is going to be used to write a college thesis or to find the perfect meme to match a joke you want to make. That's all going to be up to the app developers.

Steph Smith

Just to double-click on that, you could see maybe a world where Deep Research is just this broader horizontal application, or you could see what you just described, where developers are tailoring that to specific end-use cases. Are you basically saying that you think the latter is more likely in terms of the progression of these models and apps?

Anish Acharya

Not more likely, but I think it's underexplored. If you come to Deep Research today, you kind of have the blank-page problem. I'd love to see developers create some constraints.

Steph Smith

Yeah, that leads to unexpected outcomes.

Olivia Moore

The known or prescribed use of Deep Research right now is basically market research reports, and it's amazing for that. I've used it for that a lot of times. But if you try other things—one day we were trying to trace the origin of a meme—Deep Research is like a 100x better version of the Know Your Meme website that goes through the history.

Steph Smith

Yes, and the etymology—or however you describe it—of how a meme comes to be.

Anish Acharya

That should be an app.

Olivia Moore

So there are lots of other use cases that aren't market research reports that could really benefit from an incredibly obsessive, compelling model that will go and read every website on the internet until it finds the answer.

Steph Smith

Those are the things that you thought might be on the list but you didn't actually see there. What about the opposite?

Olivia Moore

I think the fact that vibe-coding products like Bolt, Cursor, and Lovable made the mainstream consumer list is a testament to how widely they're used by the technical audience. They've reached saturation so quickly. I think Gary Tan had some tweet that 95% of Y Combinator companies or something are now building using those tools. It's something that nearly every developer is probably using, which was maybe a surprise to me: how quickly we reached saturation.

Steph Smith

We've talked about this, but a continuing surprise—I don't know if it counts as one, but I'm still surprised every time—is how many companion products are on the list and also how many of them rank so high. I think we had 3 companion products in the top 10. 2 of them were NSFW-oriented, which may not be surprising when you think about traffic on the internet in general outside of AI, but a lot of people are even using them as interactive fan fiction. Some of the biggest fan-fiction sites in the world are also top 100 or top 200 global sites, so it makes sense in that way.

I guess my last surprise would be that there's actually quite a bit of consistency in the list over the past 4 versions. There are always new entrants, which is really exciting, but across the 4 lists, there are now 16 companies on the web rankings that have made it every single time and have kept the streak going. That's pretty remarkable when you think of how early we are in AI, but I think it's a testament to how those companies have cemented their brands, their products, and their status in consumer consciousness. It's also a testament to the fact that real businesses have already been built in consumer AI.

Anish Acharya

Now, to add to that, one of the surprises for me in companions was not seeing more multimodality.

The first glimmers of that at scale were Grok, and Grok added a bunch of voices with some real aesthetics and points of view.

Steph Smith

That's a good way to put it.

Anish Acharya

But it's interesting that this feels—and of course Character.AI got voice mode and more—but it feels like character and companionship are such a horizontal category. There's so much latent demand, and it will really increase once you have multimodality.

The other interesting thing is that, with a lot of the text-to-code work, my assumption was that there was a small number of people creating sites that were heavily trafficked, and that explained the rise of them. But actually, the majority of the traffic, correct me if I'm wrong here, Olivia, is from people doing creation, not just consuming other people's creations.

Olivia Moore

It really shows how much demand there is to make things, even if people are not that interested in consuming them. Yeah, you can track the traffic of apps that people have launched on Lovable.app versus visits to Lovable.dev, which is where people go to make a Lovable product. Lovable.dev has significantly more usage—or more visits—than traffic to Lovable.app, which gets back to what I was saying before: we have not even seen the first wave of viral products built on top of Lovable and Bolt. When that happens, I think the awareness of these types of platforms is going to go significantly up, too.

Anish Acharya

I mean, the App Store is going to be chaos.

Steph Smith

Yeah, it is going to be chaos. We're going to need an AI just to solve the AI app-management problem completely. And to that end, you talked about the fact that there are some consistent players. One of those players is ChatGPT, which we've talked about as kind of the starting gun of some of this application development. ChatGPT has been at the very top of the list. Has it been that way for every single iteration of web and mobile? But maybe what would surprise people is that the traffic to ChatGPT hasn't always followed the same trajectory. So maybe can you talk about that, and what did we see this time around?

Olivia Moore

Yeah, it was basically flat for a while, which I think was surprising to a lot of people. Between February 2023—basically for a whole year through February 2024—it was essentially flat in monthly visits to the website. At that point, from the data that I've seen, 50%+ of the traffic was students who were using it for essays or homework problems, but the vast majority of other people, me included, to be honest, had not maybe found a daily active use case for ChatGPT yet. And it's completely reversed more recently. They've doubled the number of visits on the web since then. They actually made their own announcement, too, where they counted across web and mobile, and in the past 6 months they grew from 200 million to 400 million weekly active users.

Steph Smith

Wow.

Olivia Moore

Which is especially surprising because it took them 9 months to double before that, and it usually gets way harder to double at scale, not easier. I think from our perspective—and if you even plot it on the graph—you can track the increases to the release of new models that unlock new use cases. The new o1 reasoning models, the GPT-4o models, which were kind of multimodal for the first time, and then Advanced Voice Mode. They've also launched some new products, like Operator, which can perform tasks on your computer, and Canvas, where you can write more naturally.

Anish Acharya

Mhm.

Olivia Moore

So it's both bringing in new users who never tried it and then taking people like me. Honestly, I was maybe a weekly, if not less-than-daily, active user. And now I'm a daily active user across several use cases. Some days I'm driving and talking to it in Voice Mode. Some days I'm working on a memo and generating something with Deep Research. Some days I'm doing some random other project and brainstorming ideas with it. I would expect that to continue as they release new models.

Steph Smith

Have you heard from the ecosystem in terms of what more frequent use cases have emerged, kind of like yours? If before it was a lot of students writing research reports, is there now a sense of understanding of what those newer use cases are?

Olivia Moore

Yeah, I think it's gotten better at some things related to coding. It's gotten better at data analysis. And the reasoning models—it's hard to overestimate them, because in the past you couldn't even rely on ChatGPT to tell you how many R's were in “strawberry” accurately. So it was hard to feel good about really tasking any sort of delicate or serious work to it. I think there are probably a long tail of use cases that people have just migrated over now that they have more confidence in the models.

Anish Acharya

You know, what's interesting to add to that is that Claude is not a traditional number-two player. Typically, the number-two player has 10% of the market share and just 10% of the product quality. Instead, Claude sits in this very interesting place where it seems like it's more beloved by a smaller number of people. It's better at creative writing, and it seems to have more of a personality, which is interesting because, at least I think, it's designed to be more constrained.

Olivia Moore

Yeah.

Anish Acharya

And then it's also strangely, much, much better at coding.

Olivia Moore

Yes.

Anish Acharya

Why? I don't know. But it's very interesting to see there's a place for both ChatGPT and Claude and Gemini, and potentially other models, all to augment each other.

Olivia Moore

To me, the really interesting thing about this list, when it came to general LLM-assistant usage, was that we only had 10 days of data for DeepSeek for January because it launched at the end of the month, and it shot up from literally nothing to number 2 on the list—10% of ChatGPT's scale on the web within a week, a little more than a week. On mobile, it had even less than that: 5 days, and it was number 14. If it had had 5 more days, it would have been number 2, and the gap is even narrower there between DeepSeek and ChatGPT. So, again, to Anish's point, that was a surprise: we could see kind of a broad-based LLM product go so viral still and capture so many users.

Anish Acharya

Wow.

Steph Smith

And DeepSeek was obviously the story when it came out. What have we learned about retention since then? Is that learning specific to DeepSeek, or are we seeing that learning applied across the ecosystem on a retention basis? How many users are coming back to the app at exactly 30 days, exactly 7 days, exactly 60 days?

Olivia Moore

It's just slightly below ChatGPT. We're looking at 7% day-30 retention for DeepSeek and 9% day-30 retention for ChatGPT. It's too early to call on the web because it's kind of hard to track usage. Part of my theory here is, if you look at DeepSeek usage, a lot of it is in the U.S., but a lot of it is in China and other countries where ChatGPT is either basically unavailable, or they try to make you not use it and you can only get by with a VPN. In those markets, it's not ChatGPT versus DeepSeek versus Perplexity; it's DeepSeek versus nothing. In those markets, I think they have a structural advantage from the retention side that might skew the overall sample.

Anish Acharya

Totally. Next time, we should add a different cut for DeepSeek U.S.

Steph Smith

Yes, exactly. A geographic breakdown. Talking about trends that we're seeing on the list, you mentioned AI video before, but is there anything else you want to call out there in terms of its presence on the report?

Olivia Moore

Yeah, 2 of the video models were Chinese video models, which is super interesting. The models are less copyright-sensitive in their training data.

Anish Acharya

And a great euphemism.

Olivia Moore

They're maybe more realistic and more prompt-adherent in their outputs as a result. But also, just in China, it's easier to hire people to caption videos. They have maybe a greater volume of researchers doing image and video stuff versus other stuff, so that makes sense. Sora, in some ways, was a little disappointing for some people, whereas the Chinese video models were maybe better than a lot of people expected, given the relative lack of capital that they've raised.

Anish Acharya

That's right. I think an interesting trend is just seeing Krea on the Brink list. Krea is the single best place to access all the models and all the tools. And the nice thing that they do is stitch all of these things together to make them greater than the sum of their parts.

Steph Smith

Mhm.

Anish Acharya

Insofar as we live in this sort of multipolar world of models—image models, video models, language models—there'll be a role for aggregators like Krea to put them all together in a thoughtful way.

Olivia Moore

Totally, especially because people who are deep in AI video understand this, but each model is known for being good at specific things: shots of people, shots of landscapes, anime, hyperrealistic. And so it can rack up very quickly on $20-a-month subscriptions if you're paying for 10 or 15 different models independently, versus having one canvas to work with all of them.

Anish Acharya

You also typically use the products together. You usually generate an image in Midjourney or FLUX, then you take that image and upscale it, and then you put it as the beginning frame in a video. So you really don't want to have seams between all of those products.

Steph Smith

Are we seeing these video models, in particular, become more opinionated? What I mean by that is, we see that in image models, right? Midjourney might be good at this, and then you might see another model better at something else. Users will gravitate toward either models or applications that provide them with that specificity or opinion. Are we seeing that in video yet?

Olivia Moore

I'd say they're both becoming more opinionated on the model level, but also the applications—the choices that they're making, even the choices that the model companies are making—are becoming more opinionated. If you've used Runway or Kling or something, you can now prompt basically the camera angles or the wideness of the shot, or all of these things a human cinematographer would do. You can prompt how the video sweeps over the surface of a screen. And so that's also a big factor in what you use for maybe different parts of one video, which is interesting.

Anish Acharya

To comment specifically, I still think Ideogram is one of the most unique models for what it does—for what it's great at—which is text generation and the sort of aesthetic that it has. It just sits in a very unique place in the ecosystem.

Olivia Moore

We did an internal competition where we had to generate a bunch of 30-second videos, and Ideogram was amazing for that because you could not get that layer of specificity anywhere else. Then you could take what was generated in Ideogram and put it into another model to animate it, to swerve, or to do whatever you needed to do.

They also have a fun feature that essentially is image-to-text. So if you have a meme or a copyrighted image that you want to replicate, or at least be inspired by, you can use their image-to-text feature and then use that text as the prompt to create a new image. I also found that fascinating because, as you learn when you're prompting with AI in general, you learn that you don't know what you're looking for. When I would prompt Ideogram, it would modify your prompt while generating the image, and then you could actually go interrogate that and be like, “Oh, that's why I'm getting X, Y, or Z.”

Steph Smith

Video, in general, tends to be more of a mobile-first phenomenon, right? We saw tons of applications, even before AI, that focused on creators being able to edit and splice video. What are we seeing in terms of the difference between what's working on mobile and what's working on desktop?

Anish Acharya

A lot of the things that are working on mobile are either things you want to use on the go, or where the asset—the underlying asset you're working with—is easily captured by phones. The avatar apps will work on mobile because you have 10 selfies of yourself sitting on your phone. A lot of the voice-first consumer products that we are seeing work are actually on mobile versus web because it's easier and more natural to talk into your phone for language learning, companionship, or other use cases than it is to talk into your laptop. The same is true with homework-helper apps; those are really blowing up on mobile as compared to web.

Steph Smith

Another interesting breakdown that represents where we are in the innovation curve is not just what is getting views, but what's actually making money, and how those aren't always one-to-one mirrored. What is making money today? What are we learning there? And is that the same as what's getting traffic?

Olivia Moore

For the first time, we actually ranked the top 50 by what Sensor Tower can measure as mobile revenue, which is typically in-app purchases and subscriptions—probably not ads—and we ranked those separately from what has the most monthly active users. There was only 40% overlap between the 2 lists.

The surprise to me was that the main categories are the same in terms of what's making money versus what people are using: photo and video generators, photo and video editors, beauty filters and beauty enhancers—a massive standalone category—and then the realm of ChatGPT copycat apps. They're all making a ton of money and getting a lot of users, but the companies within those categories are very different in terms of who's making money and who has the most usage.

We actually plotted revenue per user versus number of users, and we found the apps that had smaller user bases out of this sample set were much more likely to be making significantly more money on a per-user basis. So apps like Speak, apps like Otter, Captions, and video-editing apps—there are a lot of reasons for this. One is that if you are making a lot of money per user, you're probably more of a serious prosumer app, and so you've probably actually gated the usage pretty significantly. You have to subscribe to use the product.

There are companies on here that might be making $50 million or $100 million in ARR off of only 1 million or 2 million users. So they wouldn't make the ranks, ironically enough, for monthly active users, but they rank really, really high on a revenue basis, which is exciting.

As anyone who looks at the mobile list knows, there's a lot of, for the tech audience, seemingly random products on there. “I've never heard of this. Is this from a startup?”

On mobile especially, there's a very precise game that you can play with App Store ads and other paid but fairly low-cost acquisition channels. If you're doing this as an indie developer or maybe an app studio running internationally, you're not looking for the 10x payback on acquisition costs that we might be looking for as venture investors. So if you make back 1 or 2x your money on a user, that's amazing. You can get to 10 million users mostly by paying for them, but you're probably not going to make as much revenue or ultimately as much profit, maybe, as some of the companies that are lower-usage but higher-revenue.

Steph Smith

And is there a learning there in terms of—you mentioned how, by nature, if you start gating certain features or an application entirely, you are potentially stifling the growth of the overall user base. Is there a learning in terms of how AI founders should be thinking about that trade-off today?

Olivia Moore

I think it depends. Some of these markets are naturally maybe not mainstream behavior. One example of a category that did appear on the mobile revenue list, but was not on mobile usage, was several plant-identification apps.

Steph Smith

I love those.

Olivia Moore

You take a picture, save down the plant, and it tells you exactly what it is. If you've seen that plant before, is that an app that 100 million people will have on their phones? Maybe, maybe not. But if you're one of the—I can think of a few relatives who love plants or love birds—totally, they'll pay $100 a year for that, and they'll use it every day or every other day.

So I think it's more about founders optimizing for the type of product you have and how mainstream it can be.

Steph Smith

All right. So there's a lot of information here that we've covered. We've covered desktop, we've covered mobile, we've covered revenue versus users, and then we've also talked about the stickiness of some of these players, right? You said there were—what was it?—16 that showed up on every single list. So what can we learn from the last few lists?

Olivia Moore

I feel like the biggest thing now, being a consumer investor for close to a decade, is almost like the more you know, the less you know in some cases, because it all just comes back to the product at the end of the day. Technologists or investors can have opinions on the best monetization strategy or the best growth hacks, but in the end, if the product isn't capturing users' attention and isn't retaining them, the business is just going to be a completely leaky bucket of users in, users out.

I would say the biggest lesson is that we often meet with these amazing Ph.D. researchers, best-in-class in the whole world in terms of their technical understanding of a model or capability, and they can struggle building in consumer sometimes because the more complicated thing is not actually the thing with the highest utility, or the most delightful or helpful thing for a consumer user.

We never like to be prescriptive on consumer products, but in general, we see that when teams focus on either the pain point they're trying to solve or the unique experience they're trying to create and build toward that. If that means the old model is actually better than the new model, use that. If that means it's just 1 AI feature instead of the whole product being built on AI because it's not stable enough, do that. I think in consumer, you really have to let the data be your guide there.

[Music]

Thank you so much for listening to the A16C podcast. If you've made it this far, don't forget to subscribe so that you are the first to get our exclusive video content. Or you can check out this video that we've hand selected for you.

[Music]

全球 Top 100 GenAI 产品:排名与解读 — 文字稿与摘要 | BidClub