本周 AI 你错过了什么(Google、Apple、ChatGPT)
- Google 的 Veo 3 做到了 Olivia Moore 所说的「AI 视频的 ChatGPT 时刻」:一条提示词即可生成带原生音频、多个说话角色的完整视频。 这种完整性推动视频获得百万级播放,也让“无脸频道”在数天内吸引数十万订阅者。限制依然明显:每次生成只有8秒,图生视频不支持音频,生成成本约为每秒0.75美元。
- 语音模型的竞争焦点已经从“听得懂”转向“像人一样不完美”。 ChatGPT 升级后的 Advanced Voice Mode 更自然、更有表现力,加入升调疑问句、填充音等人类化细节;在 Sesame、Gemini、Grok 和 NotebookLM 让其原有体验显得“不再那么先进”之后,ChatGPT 终于追了上来。ElevenLabs v3 则把这条路线推进到生产工具中,允许用文本标签提示耳语、情绪、音效、打断和多角色对话。
- Apple 的 AI 故事在主持人看来依然谨慎且令人失望,目前宣布的功能中,最有吸引力的是通话和 FaceTime 的实时翻译。 Justine 认为 Apple 正把大量“真正的 AI”外包给 ChatGPT,同时在通知摘要混乱引发反弹后,收缩了原生 AI Siri 的路线。最具象征性的失败是:Siri 无法判断明天是不是本月第2个周一,反而建议用户去搜索 ChatGPT。
- 消费级 AI 颠覆了历史上的初创公司收入曲线:a16z 数据集中,消费公司在成立12个月后达到的 ARR 中位数为420万美元,底部四分位数为290万美元,顶部四分位数为870万美元。 这些数字是 AI 时代 B2B 基准的2倍,彻底逆转了 AI 之前“消费公司要等3到5年才能变现”的假设。正如 Olivia Moore 所说:“消费者回来了。”
- 推理成本迫使消费级初创公司提前收费,而产品效用也支撑起平均每月22美元的用户付费额,是主持人提到的 AI 前订阅均值的2倍以上。 尽管存在大量免费用户“AI 旅游”,付费用户留存率大致与 AI 前消费软件相当。积分包还带来了类似企业软件的收入扩张:重度用户会在订阅续费前额外支付10美元、12美元或50美元。
- 自然语言创意工具正在把品牌开发从专家工作流压缩成由提示词驱动的工具栈。 Justine 用 ChatGPT 做创意,Ideogram 设计字体和包装,再用 Krea 上的 FLUX Kontext 保持产品与门店图像的一致性,在不到几小时内创建了虚构的冷冻酸奶品牌 Melt。她更大的判断是,未来创业者无需掌握 Photoshop 等工具,也能组装出涵盖产品设计、应用、广告、虚拟形象、网红和一件代发商品的“全栈 AI 品牌”。
1. Veo 3 把完整视频场景变成一条提示词的输出
Olivia 的判断是明确的:Veo 3 是“AI 视频的 ChatGPT 时刻”。Veo 2 已经建立了更高质量的场景、一致的角色和物理效果;Veo 3 增加了原生音频,让一条文本提示词同时生成画面、对白和多个说话角色。
这种完整性解释了传播上的跃升。Veo 3 视频获得了数百万次播放,完全由生成视频组成的频道也在数天内获得数十万订阅者。Stormtrooper vlog 之所以有效,是因为戴面具的角色、雪怪和水豚对细小的视觉不连续不那么敏感,而模型可能本来就知道这些角色或生物应该长什么样。
硬限制是8秒。音频只支持文生视频,不支持图生视频,这让较长的人类故事难以保持一致;创作者因此使用面具或非人脸角色。由此诞生了一类新的“无脸频道”,不再要求创作者本人出镜。
访问方式也迅速变化:发布时,用户必须通过 Flow 订阅 Google 每月250美元的 AI Ultra 方案;之后模型通过 API 开放,也可在 Hedra、Krea 等平台使用,相关方案约为每月10美元,或通过 fal 和 Replicate 按用量计费。按每秒约0.75美元计算,Olivia 预计 Google 会继续推进更大、更长的视频模型,但一致性和价格将成为挑战;她希望看到更多经过优化或蒸馏、成本更低的模型。
2. 合成语音靠犹豫、情绪和打断变得像人
升级后的 ChatGPT Advanced Voice Mode 更自然、更有表现力,包括疑问句升调以及“um”“uh”等填充音。在 NotebookLM、Sesame、Gemini、Grok 和开源供应商抬高预期后,主持人开玩笑说,它经历了“从 Advanced Voice Mode 到 basic voice mode,再回到 Advanced Voice Mode”的过程——现在终于又先进了。
他们尚未解决的问题是:其他供应商已经取得进展,为什么这件事可能花了6个多月?一种假设是,团队在“Her”引发争议、担心陪伴型 AI 取代人际关系后变得谨慎;另一种可能是,资源优先投向了文本 AGI、Sora、图像和推理。
ElevenLabs v3 把表演指导变成可编辑的文本标签。创作者不必先录下哭泣、耳语或带口音的表演再进行转换,而是可以给台词加上“悲伤地”“认命地”或“耳语”等标签,加入音效,还能安排一个角色打断另一个角色。
Justine 的农场演示结合了浓重的德州口音、哞叫的奶牛,以及一个打断 Austin、指责他假装口音的第二角色。她的核心判断是,提示词驱动的打断让广告和叙事对白听起来像“自然对话,这是 AI 语音以前从未做到过的”。
3. Apple 最强的 AI 功能是翻译,而不是智能 Siri
主持人对 Apple Intelligence 依然失望,因为他们期待的“手机上的真正个人助理”迟迟没有出现。Siri 连“明天是本月第几个周一”都答不上来,随后还建议用户搜索 ChatGPT,这暴露了基础用户预期与 Apple 当前系统之间的差距。
Justine 认为,Apple 似乎正在把许多实质性的 AI 任务外包给 ChatGPT;而在通知摘要把多条通知合并、整理得一团乱后,公司也开始收缩相关路线。Apple 强调了 Genmoji 和通话转录,但通话与 FaceTime 实时翻译才是最突出的功能:这是“自然而明显的使用场景”,也是即时价值最清晰的功能。
一位主持人仍然“抱着对 Genmoji 的希望”,因为 Olivia 看到过一条病毒式传播的 Gen Z TikTok,内容就是 Genmoji。对比很明显:Apple 展示的是渐进式界面功能,而本期其他模型已经打开了新的媒体生产方式。
4. 消费级 AI 比 AI 时代的 B2B 基准更快实现数百万收入
主持人分析了 a16z 在生成式 AI 出现后大约前22到24个月内接触到的公司,并从开始变现时点计算增长。消费公司在第12个月达到的年化收入中位数为420万美元;底部四分位数为290万美元,顶部四分位数为870万美元。
这些基准是可比 AI 时代 B2B 数据的2倍。在 AI 之前,面向企业销售的 B2B 初创公司如果第一年 ARR 达到100万美元,已经算是“惊人、行业最佳”;消费初创公司则通常要花3到5年积累用户,再通过广告、平台交易或其他后置收入实现变现。
机制始于推理成本:与传统软件不同,每次额外的 AI 查询都可能产生几美分或几美元的成本,一个活跃用户每月可能带来数十美元成本。公司“某种程度上被迫”收费,随后发现消费者愿意平均每月支付22美元,是主持人引用的 AI 前订阅均值的2倍以上。
这种付费意愿建立在创作、陪伴、语言学习、阅读辅导、营养和教练等场景的实际效用之上。每月22美元的订阅,可以替代或扩大过去需要人工服务的能力,而人工服务的收费可能达到每小时50美元,甚至更高;例如视觉模型可以把餐食照片转化为卡路里和蛋白质估算,并给出每日或每周饮食分析。
免费用户存在大量“AI 旅游”,但付费用户的留存中位数大致与 AI 前的消费公司一样强。积分耗尽会带来10美元、12美元或50美元的追加购买;ElevenLabs 等公司还可以从每月10美元的个人方案,快速升级到高 ACV 的企业合同,远快于 Canva 所经历的从消费业务走向企业业务的5到7年历程。
5. 提示词驱动的工具可以在数小时内组装出完整品牌
按 Justine 的说法,FLUX Kontext 的差异化在于保持一致性:它可以把人、产品或 logo 移到新的环境中,同时比 GPT-4o 图像模型保持高得多的一致性。这让自然语言编辑——“用自然语言提示词操作 Photoshop”——真正适用于产品摄影和营销物料。
她为 Melt 设计的流程从 ChatGPT 开始,用来确定冷冻酸奶概念、名称、配色和 logo 方向;随后用 Ideogram 生成字体和品牌杯身;在 Krea 中,她又用 FLUX Kontext 探索把杯子放进餐厅或公园、改变包装颜色,并制作紫薯口味版本。她还生成了一张门店图片,再把 logo 叠加到图片上。
下一步实验是用 Veo 3 或 Higgsfield 让这些素材动起来,并测试模型是否理解冷冻酸奶如何融化,以及把杯子抛到空中后落地会是什么效果——它是否会像现实中那样“啪地一下落下”。Justine 在不到几小时内完成了原型;Olivia 认为这个品牌看起来比很多专业品牌更有吸引力,并表示它可以用于广告代理商的营销活动概念稿。
Olivia 认为,这指向了“全栈 AI 品牌”。Justine 进一步设想,AI 可以设计 logo、产品照片甚至产品本身;用 vibe coding 制作网站或移动应用;向消费者一件代发;再用 AI 虚拟形象制作社交广告,由 Veo 3 生成、实际上并不存在的 AI 网红进行推广。在 Justine 看来,真正的关键变化是:“你只需要在文本提示词里说出自己想要什么”,然后不断迭代,直到满意为止。
AI video completely taking over our social feeds in the span of a week, which is absolutely insane.
V3 was sort of like the ChatGPT moment for AI video.
The next generation of entrepreneurs are going to be completely AI-assisted.
Like a world of possibilities has been opened up for AI storytelling, especially in video form.
Yes, it's an exhausting time for AI creatives. It's great, but exhausting.
[Music]
I'm Justine.
1. Meet the Hosts: Justine and Olivia
I'm Olivia, and this is our very first edition of This Week in Consumer AI. We're both partners on the investing team here at a16z, and we're also identical twins.
2. Veo 3: The Game-Changer in AI Video
Very confusing. Extremely confusing, but it should be fun for a podcast. We're excited to chat about some of the cool things we saw in the wild world of consumer AI this week, starting with Veo 3, Google's video model. Then we're going to talk through the ChatGPT Advanced Voice Mode updates and Apple's big AI announcements. We'll also cover ElevenLabs' new voice model, some data our team recently put out about how fast consumer AI startups are ramping revenue, and Flux's new editing model, Kontext, and how Justine used it to make her own froyo brand.
And stay tuned for the end because we have a cool tutorial and some demo footage on how to make your own brand.
Things are moving so quickly that it feels like we went from exciting but maybe not super-realistic AI video to AI video completely taking over our social feeds in the span of a week, which is absolutely insane.
Yeah, I've been following AI video for a few years now. You probably remember that I've been an early user of all these models, and I've wanted them to work and make cool things that everyday people would like for so long. I would say Veo 3 was sort of like the ChatGPT moment for AI video, where we were suddenly seeing all of these Veo 3 generations blowing up with millions of views and channels featuring only Veo 3 videos getting hundreds of thousands of subscribers within days.
What's actually different about Veo 3?
I should give the overview first. Veo 3 is Google DeepMind's latest video-model effort. They released Veo 2 late last year, which was the first breakthrough showing that you could get really high-quality video: a consistent scene, consistent characters, physics, and things that just looked good.
Veo 3 is the next iteration of that model series. What's very different about it is that it generates audio natively at the same time as it generates video. You can prompt it with a text prompt to say something like, “A street-style interview where a man and a woman are talking about dating apps.” Or you can be even more specific and say, “A street-style interview where a man walks up to a woman and asks her, ‘What dating apps are you on?’ She replies, ‘Why are you asking?’ and gives him a suspicious look.”
You no longer have to go to another platform for an audio voice-over or anything like that. You can get a full-featured talking-human video with multiple characters in one place.
It feels like a real unlock to me as someone who's been following AI video less closely. People are now able to generate, in one prompt, a full vlog, a full talking-head video, or something that looks like a podcast.
Yes, in one go. I think that's why we've seen things like the Stormtrooper vlogs completely blowing up on TikTok and Instagram.
“I told you, Greg. I told you not to touch the nav system.”
“I followed the route.”
“You plotted it upside down, Greg.”
The interesting thing about Veo 3 is that it's limited to 8-second generations only. It also doesn't generate audio if you start from an image-to-video prompt—only if you start from text—which means that it's really hard to have longer than an 8-second clip with character consistency, unless your text prompt references a character that the model already knows.
And so that's why we've seen all of these hacks in the viral vlogs featuring Stormtroopers or a yeti. You can't see their faces because they're covered by a mask. The model knows what the yeti looks like, or a capybara. If it's not a human face, I think we're less sensitive to little changes between the 8-second clips, so you have people generating minutes-long videos that look like a consistent vlog character.
Yeah, they've been super fun to watch. How do you actually use Veo 3? It feels like there's been some confusion.
When Veo 3 first came out, it was only available on the Google AI Ultra plan through Flow, Google's new Creative Studio, and you had to be on the $250-a-month plan. So there was a lot of hype and a lot of FOMO.
Now the model is available via API. What that means is that a bunch of consumer video platforms, like Hedra or Krea, are offering access to Veo 3 on their $10-a-month plans. Some of the more developer-oriented API platforms, like fal or Replicate, are offering generations where you pay per video. It's priced at around 75 cents per second today.
Still pretty expensive. You have to be careful about how you prompt it, but the results are amazing. What do we expect next, either from Google or from creators? What does this mean for AI video?
On the creator front, we've already started to see this explosion of what people have called faceless channels. Now you don't have to put your own face behind a camera or on a screen to be able to talk about a topic, film a vlog, or something like that. You can have a fully AI-generated character essentially telling your story or acting out your narrative for you, which is huge.
People are using it to tell extremely funny jokes and create narrative storylines, like Greg, the incompetent Stormtrooper who's crashing all of the missions. People are getting really invested in things like that.
In terms of the model providers and the companies, Veo 3 is clearly very expensive to run. I would imagine Google will want to train the next model that's even bigger and able to generate longer videos, but it will struggle with things like coherence and, honestly, the pricing of the model. Hopefully, we'll see more condensed, optimized, distilled models that are able to do similar things at a lower cost.
3. ChatGPT's Advanced Voice Mode Updates
Yeah, I'm excited for it. There was a lot of news last week, so this got kind of lost, but I heard there was a big update to ChatGPT's Advanced Voice Mode.
Yes, they announced it on Saturday, which was an interesting choice. Weird time to drop.
I think they actually dropped the improvements last Thursday or Friday. It was first only for all paid users, and now I think it started rolling out across the broader user base. Essentially, they made Advanced Voice Mode a lot more human.
The really interesting thing was that ChatGPT was the first one to do what I would call real-time consumer voice, where you could have a conversation. That was last September in the ChatGPT app. But then they didn't really improve the product or the model that much, at least from my perspective.
We saw Sesame and other open-source providers come out with arguably better and way more humanlike models. We saw Gemini and Grok launch voice products that were much more realistic. So I think it was kind of a question mark for a lot of people what ChatGPT was doing with consumer voice.
And so what makes it better now, or what were the main upgrades?
You should actually just talk to it and see. It's live demo time, folks.
Hey, ChatGPT, are you there?
Hey there. Yep, I'm here and ready for the demo. What can I help you with today?
Did they make some improvements to your voice over the last week?
Yeah, they did. They've been rolling out some updates to make my voice sound more natural and expressive. I'm glad you noticed.
Yeah, I can hear your voice maybe flexing upwards when you're approaching a question. Sometimes you'll say “um” or “uh,” or something that sounds like a mistake but actually makes you sound more human.
Exactly. Those little touches are all intentional to make the conversation feel more natural and relatable. It definitely makes things a bit more fun and engaging, I think.
Very cool. Amazing. It's crazy to hear. It's always wild when a voice model coughs, or even does funnier things like taking on an accent or speaking in another language. But the pure realism of the voice that it demonstrated is also extremely impressive.
It's so funny, too, because when Advanced Voice Mode first came out, my feeling was, “Wow, this is amazing. This is incredible. This is so humanlike.” But then, a month or two later, NotebookLM came out, and that was the first real voice experience that put in those ums, pauses, and other things that are so humanlike. It felt like such a huge upgrade.
Then when you used Advanced Voice Mode, you were like, “This is not that advanced anymore.” Now it's finally there, which is super exciting.
So it went from Advanced Voice Mode to basic voice mode to Advanced Voice Mode. It's Advanced Voice Mode again.
One of my questions has been: What took them so long? They're on the cutting edge of so many models, and yet it feels odd to me that it took maybe 6-plus months for them to roll out improvements that we saw from other model companies much faster.
Yeah, I honestly think a big part of it might have been when they first released Advanced Voice Mode. If you remember all the controversy around “Her.”
Yes.
Was this going to be a companion that replaced humans, and some of what people would think of as the scary implications of that? It seemed like that maybe spooked them a little bit, and so they didn't want to put anything out there that sounded too human.
Yeah. I mean, that, and then also, OpenAI has been super busy. This has always, I think, been the question about the frontier LLM labs: How do they balance priorities between the north star of text-based AGI, what they’re doing in video with Sora, all the image stuff they did—which we’ll talk about a bit later—with the 4o image model, reasoning, and all of those sorts of things?
4. Apple's AI Announcements and Siri's Shortcomings
Yeah, totally. It reminds me a little bit of another—I now say big tech, since OpenAI is big tech in some ways. It counts. But the other big tech consumer update this week was Apple’s developer conference and all of the things that they announced around AI, or didn’t announce. I think people have been somewhat disappointed by Apple Intelligence, which is its bundled set of AI features.
Yeah, I think we’ve all been waiting for the AI version of Siri or some kind of true personal assistant on mobile.
Yeah, I asked Siri—I had this the other day where I asked Siri, “Okay, tomorrow’s Monday. What Monday is it of the month?” Because of San Francisco street cleaning, I had to know if it was going to be the second Monday of the month. And it said, “I can’t—I don’t know that. Can I search ChatGPT for you?” And I was like, “Siri, how can you not answer this basic question?”
It does seem like, from a lot of Apple’s updates, they’re kind of outsourcing a lot of the true AI features to ChatGPT just running on your phone. It seemed like a similar story when they rolled out those AI-powered notification summaries, where they would group 3 or 4 sets of notifications into 1 and they got a little jumbled and people got upset. It seems like that spooked Apple a little bit, and they keep retrenching on the timeline for releasing AI Siri.
So we’ll see what happens. At least in the announcement yesterday, they were leaning into things like updates to Genmoji and call transcription. I think the coolest thing I saw was real-time translation of calls and FaceTime across languages.
Yes. I’ve been surprised we haven’t seen more on that, because that feels like a really natural and obvious use case.
Yeah. I think Google might have done real-time translation, but I haven’t seen a ton of adoption yet. I did see, for the first time, a viral Gen Z TikTok featuring Genmojis, which I’m surprised took that long to hit, because Gen Z loves Genmojis.
Holding out hope for those to make it really big.
5. ElevenLabs' New Voice Model: 11 V3
Okay, and before we get too far off voice, should we talk about Eleven v3?
Yes. So, ElevenLabs, the text-to-speech company—actually, a broader AI voice company—released its third-generation model, called Eleven v3. It can do things like:
“We’re off under the lights here for this semifinal clash. The stadium is buzzing with anticipation.”
What makes Eleven v3 really special is that it does a bunch of stuff with voice that you used to have to do via speech-to-text-to-speech. Before, if you wanted to have a character who was crying while talking, had some sort of emotion, or even had a weird inflection, you would have to record yourself saying it like that, upload it to ElevenLabs, and then they would translate it into the AI voice.
Now, they essentially take all of the weird inflections, emotion, and even accents, and turn them into text prompting through things called tags. Basically, the ElevenLabs interface is an editor where you can take a sentence that you want the character to say, pick your voice, write your sentence, and then tag it as “sadly,” “resigned,” “whispering,” or something like that.
“Liam, have you tried the new ElevenLabs v3?”
“Just got it. The emotion is amazing. I can actually do whispers now, like this.”
And you can do sound effects too, right?
That is huge. So, actually, should I bring up my example of this? I don’t know if it’s going to play or not. Let’s see. This is a 20-second clip I made of 2 characters talking back and forth.
And what’s the prompt on it?
Oh, it’s a text prompt. It’ll say, “Hey y’all, my name is Austin. I’m coming to you live from our family farm in Fort Worth.” Then he’s going to walk through milking a cow, and someone’s going to interrupt him. Great.
“Hey y’all. My name is Austin, and I’m coming to you live from our family farm in Fort Worth. Today I’m going to walk through what it’s like—”
“Austin, are you faking an accent again?”
“It’s not faking. I was born here.”
“And everyone knows you don’t talk like that.”
So my favorite thing about that is it showcases a couple of things about the model. It can do bad accents. It can do terrible accents. It can do great accents. That was two different characters. At first, I prompted the Austin character of having a thick Texas accent. Then I prompted the cows mooing. And then you can also prompt interruptions, which is really cool. So a tag is literally like starts talking and gets interrupted. And then the next character that comes in, you can say cuts the other character off. And so for narrative, storytelling, for ads, marketing, anything, it makes it sound like a natural conversation, which we've never had with AI voice before.
Yeah, it feels like between this and Veo 3, a world of possibilities has been opened up for AI storytelling, especially in video form.
Yes. It’s an exhausting time for AI creatives. It’s great, but exhausting, because there’s just way too much fun stuff to test.
Yeah. Any other favorite use cases of this?
One that I made was 1 of those car extended-warranty videos that you could actually make sound as menacing as they somewhat feel.
I think ElevenLabs is actually doing a competition now where they’re soliciting the best examples of people using v3 from all around the world.
So I’m going to be very curious to see how the professional narrative builders and storytellers are using it. We’ve made all sorts of fun stuff, but I think we’ve just scratched the surface of what’s possible here.
6. Report from a16z: AI Revenue Growth
Amazing. Okay, so you put out some data last week about AI revenue ramp and how fast companies are growing. Let’s chat through the main takeaways from that.
Yeah. So, basically, the methodology here—or, maybe, to even back up, the purpose here—was that I think we all have this idea in mind, or maybe we have that idea because we’ve heard it a billion times, that we’re in a new era of growth thanks to AI. Companies are scaling faster than ever before, right? But my question was: What does that really mean, and how fast is that? Is it 20% faster? Is it 50% faster than what we saw pre-AI?
We’re blessed to get to meet tons of companies here every day. We meet dozens of companies a week, so we went back and essentially pulled all the data from companies we’ve met in the generative AI era, which I would say is the last 22 to 24 months. We looked at, once they started monetizing, how fast they were growing.
Pre-AI, if you were a B2B startup selling to enterprises, if you got to $1 million in ARR in the first year, that was amazing and best-in-class. That was the rule of thumb. I remember that was the known metric.
Very exciting.
If you were a consumer startup, you would not make money for 3 to 5 years, maybe longer.
Yes. The whole idea was to build up a user base and then probably monetize them directly via ads or transactions for a marketplace, maybe down the line.
Yes, and there were counterexamples to that—some subscription companies—but that was definitely not the dominant model.
That has fully shifted in the AI era, and most companies are now making money directly from consumers via subscription. What we found was actually pretty surprising: The median ARR, or annualized revenue run rate, is now $4.2 million at month 12 for consumer startups. The bottom quartile is $2.9 million, and the top quartile is $8.7 million.
Wow.
A median B2C company in the age of AI is getting to $4 million ARR after a year, and the best-in-class companies are getting to $8 million in a year, upwards of $8 million. In the pre-AI era, we never would have seen anything like that. The even more surprising thing is that those numbers are twice as high as the B2B benchmarks in the AI era. Consumer companies are actually ramping revenue faster, which again is a total reversal from what we saw before.
I think there are a couple of reasons why this was happening. First, why have consumer AI companies adopted subscription? They were kind of forced to, because especially in the early era of models, they were so expensive that you, as a company, had pretty high cost of goods sold.
The inference cost, you mean, of running the model?
Historically, the benefit of software was that there was no marginal cost. You made an app, and there was no additional cost to serving the next user. In AI, that is actually not true at all, especially if you’re running inference on a model. It costs you cents, maybe even dollars, for each query. So each user could be costing you dozens of dollars a month.
Yeah, absolutely. So a lot of companies had to at least try to charge.
Right. And it turns out that these new, AI-native products are so powerful that consumers are willing to pay. We also ran some additional data analysis that shows that, on average, consumer AI startups are charging an average user $22 a month, which again is more than double what subscription companies were able to charge on average pre-AI.
And why do we think that is? Do we have theories?
I mean, on the creative tools side, what I’ve seen is that for people who weren’t creative, AI tools allow them to make photos, images, or art for the first time. They can make videos and animations. And for creative people—we have a cousin who’s a creative—they can genuinely use this to supercharge their workflows and do their job a lot faster.
So they're willing to pay for it. Have we seen examples of that outside of creative tools yet?
It's a good question. We've seen some of it around companion apps, where the products are just so powerful that having a friend with you 24/7 makes people excited to pay. We've also seen this around categories like language learning, teaching your kid how to read, and other things that previously you'd have to pay a human being $50 an hour, if that, to get access to. Now, $22 a month with AI feels pretty cheap.
Totally. I'm even thinking some of the things I've seen monetized well are in nutrition or coaching, where, for the first time, thanks to vision models, you can take a picture of what you're eating and have a vision model pull out how many calories are in it, how much protein, and then, at the end of a day or week, summarize insights about what you should be eating more or less of. That's something that, pre-AI, you could maybe do by taking a picture and uploading it to a forum, but you'd have to find a nutritionist, wait weeks to book an appointment, and maybe get a referral from your doctor. It would take forever. So, it's really exciting. I think it's monetizing people who never would have paid for that before.
Yes. And then people who would have paid for it are switching over to the AI version, or they're willing to pay even more, which is super exciting. The other thing that I think people have questions about, or maybe doubts about, would be, like, okay, great, they're growing fast, but they're not retaining a lot of users. We did some analysis on this, too.
There's definitely a lot of tourism—we call it AI tourism behavior—in terms of free users, which means you get a lot of hits on your website, essentially, and most of those users don't stick around. But if you look at it on a paid-user basis, once you actually subscribe, consumer AI companies are retaining, at the median, pretty much just as well as pre-AI consumer companies, which is really exciting.
I feel like what we've seen, especially on revenue retention, which is fascinating, is that you might have more tourists, so you might have more people who subscribe and then cancel, but for the first time you have real upsell activity in consumer subscriptions. You're not just paying $10 a month for the app. You're paying $10 a month for the image model, but then, if you love it and you run out of credits, you're paying another $10, $12, or $50 for additional credit packs before the next month of your subscription starts. And so that means you see revenue expansion opportunity now in consumer that we only used to see in enterprise.
Absolutely. Or, honestly, in games, what they call monetizing the whales—the people who are the really high spenders—which is, to me, one of the most exciting things about consumer AI products right now.
Yeah. And we're seeing, I would say, companies convert consumer revenue to enterprise revenue way faster than they ever did before. Companies like Canva previously took 5, 6, 7-plus years to really move from consumer/prosumer to enterprise, right? And now we're seeing companies like ElevenLabs. It's a great example: it starts at $10, and then they convert to a really high ACV enterprise contract, which is super exciting.
I mean, I feel like we even saw that in the very early days of consumer AI. I remember our friends at ad agencies or entertainment companies would tell us that they were using Midjourney to mock things up, or even using the images in their final work product. So it was a true enterprise use case, but growing bottoms-up, which is sort of a fascinating motion.
Yeah, it's exciting. Consumers are back.
7. Demo of the Week: AI in Brand Creation
All right, awesome. We're moving on to our demo of the week. One fun fact about us is that we genuinely love—at least for me, it's probably my number 1 hobby now—trying out all of the AI creative tools, especially, but also AI consumer products more broadly, figuring out how to make cool things and then sharing the workflows with other people whose number 1 hobby is not doing this. So this week, we're going to talk about brand creation and ideation using AI.
I made this new frozen yogurt brand called Melt that I iterated on with ChatGPT. Then I took it to Ideogram, and then I took it to Krea to do the final touches and make these really cool product photos and even store photos. I think the initial idea about this was seeing FLUX Kontext come out, which is the new image-editing model from Black Forest Labs, hosted on Krea.
FLUX Kontext is kind of like the GPT-4o image model, where you can upload an image and then say, “Make this Ghibli-style,” which was the viral example. You can also say, “Take the person from this photo and put them in a new environment,” or, “Take the logo and change it slightly.”
Yeah. Add or remove objects. I've seen it described as kind of like Photoshop, but with natural-language prompts.
Yes, you can edit with words for the first time. And I think that is what makes it different from the GPT-4o image model: the consistency with which it retains the item or the character, whatever, is much, much better. We'll show some examples here, but basically, if you're taking a photo of yourself and uploading it to GPT-4o and saying, “Put me in a podcast studio,” you will likely end up looking completely different in the new photo than you did in the initial photo, or maybe have some similar features, but quite different. Whereas this model does an amazing job at maintaining consistency.
And so that sparked this idea for me: that means this can actually be used for brands to do product photos or other sorts of marketing collateral, because the logos and the products can be consistent.
Awesome.
And, you know, I'm a huge froyo fan. I feel like froyo has gotten an unfair shake in recent years. It's kind of seen as the little-kid thing. And so I wanted to make a cool, hip, modern 20s New York froyo brand.
I went back and forth with ChatGPT on this idea to land on the name Melt, which I love, and to land on the branding: this is sort of what the font of the logo will look like, and then this is the color of the packaging. Then I took that prompt for the logo to Ideogram, which is an image-generation and sort of editing canvas, and it's super good, I think, at logos, typography, and anything sort of product- or word-related. I had it generate this photo of a froyo cup floating in the air with the Melt logo and branding.
Then I downloaded that photo and took it to Krea, where I used the new FLUX Kontext editing model to run all different kinds of scenarios. What's really cool about that is you can upload the photo and then say, “Take this froyo and put it sitting on a counter at a trendy restaurant. Put it in the hand of a woman at a park.” Or even, “Make the froyo cup white instead of blue, and give it a pink border. Make the froyo itself purple if they're having an ube froyo special.”
Yeah.
And then I think the next step, which I didn't do here—I kind of stopped at the product images—and I actually made an image of the store, too. I took the logo and superimposed it over the store I generated.
I know. You want to go there.
But the next step even further would be video. My idea is to take all of those product shots, take them to Veo 3 or Higgsfield, which does really cool special-effects stuff, and have the froyo cup in action. Have it actually melt over the side. Have it melt. It has to melt.
I'm very curious to see: do the models understand the physics of froyo? If it tosses the cup in the air, how does the froyo land? Does it kind of plop like we all know froyo would in real life?
Obviously, this is just a fun experiment for me. I'm not, unfortunately, actually going to be starting a froyo brand, but it sort of makes you start thinking about this: if you work at an ad agency, for example, and you were mocking up a deck for your client about your latest campaign, why would you not use something like this to show them what it might look like?
And you did it in less than a couple hours. Honestly, the branding does look more exciting than a lot of professional brands that we see out there. It makes me think about the next generation of entrepreneurs, who are going to be completely AI-assisted in a lot of these assets that they're putting together.
I think they're going to be able to make full-stack AI brands. There are also products where you can design with AI. You can make ads with AI. I think there'll be no reason for any person not to have their own product line, small business, or open a store if they want to. AI is assisting with these kinds of things, too.
Totally. Yeah. I think we'll see brands that are logo, product photo, maybe even product itself, designed by AI; vibe-coded or vibe-designed website or mobile app; and then kind of drop-shipped to the end consumer. Social media ads also generated with an AI avatar that holds it up for you and sells it on TikTok or Instagram, promoted by AI influencers generated by Veo 3. They don't actually exist.
I think that sort of thing is going to be really fascinating to see because you no longer have to know how to work all of these technical tools that you had to be able to use. Even Photoshop has so many buttons; it's very complicated. And now you can just ask for what you want in a text prompt, get something generated, and iterate on it until you end up with something that you really love, which I think is crazy powerful.
Awesome.