[BidClub_]
The a16z Show · · 30 min

What You Missed in AI This Week (Google, Apple, ChatGPT)

Justine MooreOlivia Moore

YouTube
TL;DR
  • Google’s Veo 3 delivered what Olivia Moore calls “the ChatGPT moment for AI video,” pairing generated footage with native audio and multiple speaking characters in one prompt. That completeness helped drive million-view clips and “faceless channels” gaining hundreds of thousands of subscribers within days. The constraint remains severe: eight-second generations, no audio from image-to-video, and roughly $0.75 per generated second.
  • Voice models are competing on human imperfection, not merely intelligibility. ChatGPT’s upgraded Advanced Voice Mode now sounds more natural and expressive, with rising question inflections, filler sounds, and other human-like touches, after Sesame, Gemini, Grok, and NotebookLM made its original experience feel “not that advanced anymore.” ElevenLabs v3 pushes the same frontier into production tooling, using text tags to prompt whispers, emotions, sound effects, interruptions, and multiple characters.
  • Apple’s AI story remains cautious and disappointing to the hosts, with the most compelling announced feature being real-time translation across calls and FaceTime. Justine suggests Apple is outsourcing much of its “true AI” to ChatGPT while retrenching on an AI-native Siri after jumbled notification summaries caused backlash. The emblematic failure: Siri could not determine whether tomorrow was the month’s second Monday and instead offered to search ChatGPT.
  • Consumer AI has inverted the historical startup revenue curve: the median consumer company in a16z’s dataset reached $4.2 million of ARR after 12 months, versus $2.9 million at the bottom quartile and $8.7 million at the top quartile. Those figures were twice the corresponding AI-era B2B benchmarks, a sharp reversal from the pre-AI assumption that consumer companies would wait three to five years before monetizing. As Olivia Moore put it: “Consumers are back.”
  • Inference costs forced consumer startups to charge early, but product utility is supporting an average user payment of $22 per month—more than double the pre-AI subscription average cited by the hosts. Paid-user retention is roughly comparable with pre-AI consumer software despite heavy free-user “AI tourism.” Credit packs also introduce enterprise-like revenue expansion as power users spend another $10, $12, or $50 before their subscriptions renew.
  • Natural-language creative tools are collapsing brand development from a specialist workflow into a prompt-driven stack. Justine created the fictional Melt frozen-yogurt brand in under a couple of hours using ChatGPT for ideation, Ideogram for typography and packaging, and FLUX Kontext on Krea for consistent product and store imagery. Her larger call is that future entrepreneurs can assemble “full-stack AI brands”—including product design, apps, ads, avatars, influencers, and drop-shipped goods—without mastering tools such as Photoshop.
Digest · the substance, structured for research

1. Veo 3 turns complete video scenes into one-prompt outputs

  • Olivia’s framing is categorical: Veo 3 was “the ChatGPT moment for AI video.” Veo 2 had established higher-quality scenes, consistent characters, and physics; Veo 3 added native audio, letting one text prompt generate dialogue and multiple speaking characters alongside the footage.

  • That completeness helped explain the distribution jump. Veo 3 clips attracted millions of views, while channels composed entirely of generated videos gained hundreds of thousands of subscribers within days. Stormtrooper vlogs work because masked characters, yetis, and capybaras are less sensitive to small visual discontinuities, while the model may already know what those characters or creatures look like.

  • The hard boundary is eight seconds. Audio works only from text-to-video, not image-to-video, making longer human stories difficult to maintain consistently; creators therefore use masks or nonhuman faces. The result is a new class of “faceless channels” that no longer requires a creator on camera.

  • Access also shifted quickly: launch required Google’s $250-per-month AI Ultra plan through Flow, while the model later became available via API and on platforms including Hedra and Krea through plans around $10 monthly, or through fal and Replicate on usage pricing. At roughly $0.75 per second, Olivia expects Google to pursue larger, longer-video models but says coherence and pricing will be challenges; she hopes to see more optimized or distilled models at lower cost.

2. Synthetic voice gets human through hesitation, emotion, and interruption

  • The upgraded ChatGPT Advanced Voice Mode sounds more natural and expressive, including rising question inflections and “um”/“uh” sounds. After NotebookLM, Sesame, Gemini, Grok, and open-source providers raised expectations, the hosts joked that it had gone “from Advanced Voice Mode to basic voice mode to Advanced Voice Mode”—and is finally advanced again.

  • Their unresolved question is why this took perhaps six-plus months while other providers advanced. One hypothesis is caution after controversy around “Her” and fears of a companion replacing human relationships; the other is prioritization across text-based AGI, Sora, images, and reasoning.

  • ElevenLabs v3 turns performance direction into editable text tags. Instead of recording someone crying, whispering, or using an accent and then converting that performance, creators can tag a line as “sadly,” “resigned,” or “whispering,” add sound effects, and script one character interrupting another.

  • Justine’s farm demo combines a thick Texas accent, mooing cows, and a second character cutting off Austin for allegedly faking the accent. Her key claim is that prompted interruption makes ads and narrative dialogue sound like “a natural conversation, which we’ve never had with AI voice before.”

3. Apple’s strongest AI feature is translation, not an intelligent Siri

  • The hosts remain disappointed by Apple Intelligence because the awaited “true personal assistant on mobile” has not arrived. Siri’s inability to answer what Monday of the month tomorrow would be—before offering ChatGPT search—captures the gap between basic user expectations and Apple’s current system.

  • Justine suggested that Apple seems to be outsourcing many substantive AI tasks to ChatGPT and retrenching after notification summaries grouped several notifications together and got jumbled. Apple highlighted Genmoji and call transcription; real-time translation across calls and FaceTime stood out as the “natural and obvious use case” with the clearest immediate value.

  • One host is still “holding out hope” for Genmoji after Olivia saw a viral Gen Z TikTok featuring Genmojis. The contrast is telling: Apple showed incremental interface features while the episode’s other models opened new forms of media production.

4. Consumer AI reaches millions in revenue faster than AI-era B2B benchmarks

  • The hosts’ analysis covered companies a16z met during roughly the first 22 to 24 months of generative AI, measuring growth from the start of monetization. The median consumer company reached $4.2 million in annualized revenue at month 12; the bottom quartile reached $2.9 million and the top quartile $8.7 million.

  • Those benchmarks are twice the comparable AI-era B2B numbers. Pre-AI, $1 million of first-year ARR was “amazing, best-in-class” for a B2B startup selling to enterprises, while consumer startups commonly spent three to five years building users before monetizing through advertising, marketplace transactions, or other later revenue.

  • The mechanism begins with inference costs: unlike conventional software, each additional AI query can cost cents or dollars, leaving an active user costing dozens of dollars monthly. Companies were “kind of forced” to charge, then discovered consumers would pay an average of $22 per month—more than double the cited pre-AI subscription average.

  • Utility supports that willingness across creative work, companionship, language learning, reading instruction, nutrition, and coaching. A $22 subscription can replace or broaden access to services that previously required a human charging $50 an hour, if that; vision models, for example, can turn meal photos into calorie and protein estimates plus daily or weekly dietary analysis.

  • Free users exhibit substantial “AI tourism,” but median paid-user retention is roughly as strong as in pre-AI consumer companies. Credit exhaustion creates upsells of $10, $12, or $50, while companies such as ElevenLabs can move from a $10 individual plan into high-ACV enterprise contracts—far faster than Canva’s cited five-to-seven-year consumer-to-enterprise journey.

5. Prompt-driven tools can assemble an entire brand in hours

  • FLUX Kontext’s distinction, in Justine’s telling, is preservation: it can move a person, product, or logo into a new setting while maintaining much greater consistency than the GPT-4o image model. That makes natural-language editing—“Photoshop but with natural language prompts”—usable for product photography and marketing collateral.

  • Her Melt workflow began with ChatGPT for the frozen-yogurt concept, name, colors, and logo direction; Ideogram produced typography and a branded cup; and, at Krea, FLUX Kontext let her explore placing the cup in restaurants or parks, changing packaging colors, and creating an ube variation. She also generated a store image and superimposed the logo onto it.

  • The next experiment would animate those assets with Veo 3 or Higgsfield and test whether the model understands how frozen yogurt melts and how it would land if the cup were tossed in the air—whether it would “kind of plop” as it would in real life. Justine made the prototype in less than a couple of hours; Olivia said the branding looked more exciting than many professional brands and suggested this could support agency campaign mock-ups.

  • Olivia says this points toward “full-stack AI brands.” Justine then imagines AI-designed logos, product photos and perhaps products; vibe-coded websites or mobile apps; drop-shipping to consumers; and social ads featuring AI avatars, promoted by AI influencers generated by Veo 3 that do not actually exist. The enabling shift, in Justine’s view, is that “you can just ask for what you want in a text prompt” and iterate until you love the result.

AI video completely taking over our social feeds in the span of a week, which is absolutely insane.

V3 was sort of like the ChatGPT moment for AI video.

The next generation of entrepreneurs are going to be completely AI-assisted.

Like a world of possibilities has been opened up for AI storytelling, especially in video form.

Yes, it's an exhausting time for AI creatives. It's great, but exhausting.

[Music]

Justine Moore

I'm Justine.

1. Meet the Hosts: Justine and Olivia

Olivia Moore

I'm Olivia, and this is our very first edition of This Week in Consumer AI. We're both partners on the investing team here at a16z, and we're also identical twins.

2. Veo 3: The Game-Changer in AI Video

Very confusing. Extremely confusing, but it should be fun for a podcast. We're excited to chat about some of the cool things we saw in the wild world of consumer AI this week, starting with Veo 3, Google's video model. Then we're going to talk through the ChatGPT Advanced Voice Mode updates and Apple's big AI announcements. We'll also cover ElevenLabs' new voice model, some data our team recently put out about how fast consumer AI startups are ramping revenue, and Flux's new editing model, Kontext, and how Justine used it to make her own froyo brand.

And stay tuned for the end because we have a cool tutorial and some demo footage on how to make your own brand.

Justine Moore

Things are moving so quickly that it feels like we went from exciting but maybe not super-realistic AI video to AI video completely taking over our social feeds in the span of a week, which is absolutely insane.

Olivia Moore

Yeah, I've been following AI video for a few years now. You probably remember that I've been an early user of all these models, and I've wanted them to work and make cool things that everyday people would like for so long. I would say Veo 3 was sort of like the ChatGPT moment for AI video, where we were suddenly seeing all of these Veo 3 generations blowing up with millions of views and channels featuring only Veo 3 videos getting hundreds of thousands of subscribers within days.

Justine Moore

What's actually different about Veo 3?

Olivia Moore

I should give the overview first. Veo 3 is Google DeepMind's latest video-model effort. They released Veo 2 late last year, which was the first breakthrough showing that you could get really high-quality video: a consistent scene, consistent characters, physics, and things that just looked good.

Veo 3 is the next iteration of that model series. What's very different about it is that it generates audio natively at the same time as it generates video. You can prompt it with a text prompt to say something like, “A street-style interview where a man and a woman are talking about dating apps.” Or you can be even more specific and say, “A street-style interview where a man walks up to a woman and asks her, ‘What dating apps are you on?’ She replies, ‘Why are you asking?’ and gives him a suspicious look.”

You no longer have to go to another platform for an audio voice-over or anything like that. You can get a full-featured talking-human video with multiple characters in one place.

Justine Moore

It feels like a real unlock to me as someone who's been following AI video less closely. People are now able to generate, in one prompt, a full vlog, a full talking-head video, or something that looks like a podcast.

Olivia Moore

Yes, in one go. I think that's why we've seen things like the Stormtrooper vlogs completely blowing up on TikTok and Instagram.

“I told you, Greg. I told you not to touch the nav system.”

“I followed the route.”

“You plotted it upside down, Greg.”

Justine Moore

The interesting thing about Veo 3 is that it's limited to 8-second generations only. It also doesn't generate audio if you start from an image-to-video prompt—only if you start from text—which means that it's really hard to have longer than an 8-second clip with character consistency, unless your text prompt references a character that the model already knows.

Olivia Moore

And so that's why we've seen all of these hacks in the viral vlogs featuring Stormtroopers or a yeti. You can't see their faces because they're covered by a mask. The model knows what the yeti looks like, or a capybara. If it's not a human face, I think we're less sensitive to little changes between the 8-second clips, so you have people generating minutes-long videos that look like a consistent vlog character.

Justine Moore

Yeah, they've been super fun to watch. How do you actually use Veo 3? It feels like there's been some confusion.

Olivia Moore

When Veo 3 first came out, it was only available on the Google AI Ultra plan through Flow, Google's new Creative Studio, and you had to be on the $250-a-month plan. So there was a lot of hype and a lot of FOMO.

Now the model is available via API. What that means is that a bunch of consumer video platforms, like Hedra or Krea, are offering access to Veo 3 on their $10-a-month plans. Some of the more developer-oriented API platforms, like fal or Replicate, are offering generations where you pay per video. It's priced at around 75 cents per second today.

Justine Moore

Still pretty expensive. You have to be careful about how you prompt it, but the results are amazing. What do we expect next, either from Google or from creators? What does this mean for AI video?

Olivia Moore

On the creator front, we've already started to see this explosion of what people have called faceless channels. Now you don't have to put your own face behind a camera or on a screen to be able to talk about a topic, film a vlog, or something like that. You can have a fully AI-generated character essentially telling your story or acting out your narrative for you, which is huge.

People are using it to tell extremely funny jokes and create narrative storylines, like Greg, the incompetent Stormtrooper who's crashing all of the missions. People are getting really invested in things like that.

In terms of the model providers and the companies, Veo 3 is clearly very expensive to run. I would imagine Google will want to train the next model that's even bigger and able to generate longer videos, but it will struggle with things like coherence and, honestly, the pricing of the model. Hopefully, we'll see more condensed, optimized, distilled models that are able to do similar things at a lower cost.

3. ChatGPT's Advanced Voice Mode Updates

Justine Moore

Yeah, I'm excited for it. There was a lot of news last week, so this got kind of lost, but I heard there was a big update to ChatGPT's Advanced Voice Mode.

Olivia Moore

Yes, they announced it on Saturday, which was an interesting choice. Weird time to drop.

I think they actually dropped the improvements last Thursday or Friday. It was first only for all paid users, and now I think it started rolling out across the broader user base. Essentially, they made Advanced Voice Mode a lot more human.

The really interesting thing was that ChatGPT was the first one to do what I would call real-time consumer voice, where you could have a conversation. That was last September in the ChatGPT app. But then they didn't really improve the product or the model that much, at least from my perspective.

We saw Sesame and other open-source providers come out with arguably better and way more humanlike models. We saw Gemini and Grok launch voice products that were much more realistic. So I think it was kind of a question mark for a lot of people what ChatGPT was doing with consumer voice.

Justine Moore

And so what makes it better now, or what were the main upgrades?

Olivia Moore

You should actually just talk to it and see. It's live demo time, folks.

Hey, ChatGPT, are you there?

Speaker 1

Hey there. Yep, I'm here and ready for the demo. What can I help you with today?

Olivia Moore

Did they make some improvements to your voice over the last week?

Speaker 1

Yeah, they did. They've been rolling out some updates to make my voice sound more natural and expressive. I'm glad you noticed.

Olivia Moore

Yeah, I can hear your voice maybe flexing upwards when you're approaching a question. Sometimes you'll say “um” or “uh,” or something that sounds like a mistake but actually makes you sound more human.

Speaker 1

Exactly. Those little touches are all intentional to make the conversation feel more natural and relatable. It definitely makes things a bit more fun and engaging, I think.

Justine Moore

Very cool. Amazing. It's crazy to hear. It's always wild when a voice model coughs, or even does funnier things like taking on an accent or speaking in another language. But the pure realism of the voice that it demonstrated is also extremely impressive.

Olivia Moore

It's so funny, too, because when Advanced Voice Mode first came out, my feeling was, “Wow, this is amazing. This is incredible. This is so humanlike.” But then, a month or two later, NotebookLM came out, and that was the first real voice experience that put in those ums, pauses, and other things that are so humanlike. It felt like such a huge upgrade.

Then when you used Advanced Voice Mode, you were like, “This is not that advanced anymore.” Now it's finally there, which is super exciting.

Justine Moore

So it went from Advanced Voice Mode to basic voice mode to Advanced Voice Mode. It's Advanced Voice Mode again.

One of my questions has been: What took them so long? They're on the cutting edge of so many models, and yet it feels odd to me that it took maybe 6-plus months for them to roll out improvements that we saw from other model companies much faster.

Olivia Moore

Yeah, I honestly think a big part of it might have been when they first released Advanced Voice Mode. If you remember all the controversy around “Her.”

Justine Moore

Yes.

Olivia Moore

Was this going to be a companion that replaced humans, and some of what people would think of as the scary implications of that? It seemed like that maybe spooked them a little bit, and so they didn't want to put anything out there that sounded too human.

Justine Moore

Yeah. I mean, that, and then also, OpenAI has been super busy. This has always, I think, been the question about the frontier LLM labs: How do they balance priorities between the north star of text-based AGI, what they’re doing in video with Sora, all the image stuff they did—which we’ll talk about a bit later—with the 4o image model, reasoning, and all of those sorts of things?

4. Apple's AI Announcements and Siri's Shortcomings

Olivia Moore

Yeah, totally. It reminds me a little bit of another—I now say big tech, since OpenAI is big tech in some ways. It counts. But the other big tech consumer update this week was Apple’s developer conference and all of the things that they announced around AI, or didn’t announce. I think people have been somewhat disappointed by Apple Intelligence, which is its bundled set of AI features.

Justine Moore

Yeah, I think we’ve all been waiting for the AI version of Siri or some kind of true personal assistant on mobile.

Olivia Moore

Yeah, I asked Siri—I had this the other day where I asked Siri, “Okay, tomorrow’s Monday. What Monday is it of the month?” Because of San Francisco street cleaning, I had to know if it was going to be the second Monday of the month. And it said, “I can’t—I don’t know that. Can I search ChatGPT for you?” And I was like, “Siri, how can you not answer this basic question?”

Justine Moore

It does seem like, from a lot of Apple’s updates, they’re kind of outsourcing a lot of the true AI features to ChatGPT just running on your phone. It seemed like a similar story when they rolled out those AI-powered notification summaries, where they would group 3 or 4 sets of notifications into 1 and they got a little jumbled and people got upset. It seems like that spooked Apple a little bit, and they keep retrenching on the timeline for releasing AI Siri.

Olivia Moore

So we’ll see what happens. At least in the announcement yesterday, they were leaning into things like updates to Genmoji and call transcription. I think the coolest thing I saw was real-time translation of calls and FaceTime across languages.

Justine Moore

Yes. I’ve been surprised we haven’t seen more on that, because that feels like a really natural and obvious use case.

Olivia Moore

Yeah. I think Google might have done real-time translation, but I haven’t seen a ton of adoption yet. I did see, for the first time, a viral Gen Z TikTok featuring Genmojis, which I’m surprised took that long to hit, because Gen Z loves Genmojis.

Justine Moore

Holding out hope for those to make it really big.

5. ElevenLabs' New Voice Model: 11 V3

Olivia Moore

Okay, and before we get too far off voice, should we talk about Eleven v3?

Justine Moore

Yes. So, ElevenLabs, the text-to-speech company—actually, a broader AI voice company—released its third-generation model, called Eleven v3. It can do things like:

“We’re off under the lights here for this semifinal clash. The stadium is buzzing with anticipation.”

What makes Eleven v3 really special is that it does a bunch of stuff with voice that you used to have to do via speech-to-text-to-speech. Before, if you wanted to have a character who was crying while talking, had some sort of emotion, or even had a weird inflection, you would have to record yourself saying it like that, upload it to ElevenLabs, and then they would translate it into the AI voice.

Now, they essentially take all of the weird inflections, emotion, and even accents, and turn them into text prompting through things called tags. Basically, the ElevenLabs interface is an editor where you can take a sentence that you want the character to say, pick your voice, write your sentence, and then tag it as “sadly,” “resigned,” “whispering,” or something like that.

“Liam, have you tried the new ElevenLabs v3?”

“Just got it. The emotion is amazing. I can actually do whispers now, like this.”

Olivia Moore

And you can do sound effects too, right?

Justine Moore

That is huge. So, actually, should I bring up my example of this? I don’t know if it’s going to play or not. Let’s see. This is a 20-second clip I made of 2 characters talking back and forth.

Olivia Moore

And what’s the prompt on it?

Justine Moore

Oh, it’s a text prompt. It’ll say, “Hey y’all, my name is Austin. I’m coming to you live from our family farm in Fort Worth.” Then he’s going to walk through milking a cow, and someone’s going to interrupt him. Great.

“Hey y’all. My name is Austin, and I’m coming to you live from our family farm in Fort Worth. Today I’m going to walk through what it’s like—”

“Austin, are you faking an accent again?”

“It’s not faking. I was born here.”

“And everyone knows you don’t talk like that.”

Justine Moore

So my favorite thing about that is it showcases a couple of things about the model. It can do bad accents. It can do terrible accents. It can do great accents. That was two different characters. At first, I prompted the Austin character of having a thick Texas accent. Then I prompted the cows mooing. And then you can also prompt interruptions, which is really cool. So a tag is literally like starts talking and gets interrupted. And then the next character that comes in, you can say cuts the other character off. And so for narrative, storytelling, for ads, marketing, anything, it makes it sound like a natural conversation, which we've never had with AI voice before.

Olivia Moore

Yeah, it feels like between this and Veo 3, a world of possibilities has been opened up for AI storytelling, especially in video form.

Justine Moore

Yes. It’s an exhausting time for AI creatives. It’s great, but exhausting, because there’s just way too much fun stuff to test.

Olivia Moore

Yeah. Any other favorite use cases of this?

Justine Moore

One that I made was 1 of those car extended-warranty videos that you could actually make sound as menacing as they somewhat feel.

Olivia Moore

I think ElevenLabs is actually doing a competition now where they’re soliciting the best examples of people using v3 from all around the world.

Justine Moore

So I’m going to be very curious to see how the professional narrative builders and storytellers are using it. We’ve made all sorts of fun stuff, but I think we’ve just scratched the surface of what’s possible here.

6. Report from a16z: AI Revenue Growth

Olivia Moore

Amazing. Okay, so you put out some data last week about AI revenue ramp and how fast companies are growing. Let’s chat through the main takeaways from that.

Justine Moore

Yeah. So, basically, the methodology here—or, maybe, to even back up, the purpose here—was that I think we all have this idea in mind, or maybe we have that idea because we’ve heard it a billion times, that we’re in a new era of growth thanks to AI. Companies are scaling faster than ever before, right? But my question was: What does that really mean, and how fast is that? Is it 20% faster? Is it 50% faster than what we saw pre-AI?

We’re blessed to get to meet tons of companies here every day. We meet dozens of companies a week, so we went back and essentially pulled all the data from companies we’ve met in the generative AI era, which I would say is the last 22 to 24 months. We looked at, once they started monetizing, how fast they were growing.

Pre-AI, if you were a B2B startup selling to enterprises, if you got to $1 million in ARR in the first year, that was amazing and best-in-class. That was the rule of thumb. I remember that was the known metric.

Olivia Moore

Very exciting.

Justine Moore

If you were a consumer startup, you would not make money for 3 to 5 years, maybe longer.

Olivia Moore

Yes. The whole idea was to build up a user base and then probably monetize them directly via ads or transactions for a marketplace, maybe down the line.

Justine Moore

Yes, and there were counterexamples to that—some subscription companies—but that was definitely not the dominant model.

Olivia Moore

That has fully shifted in the AI era, and most companies are now making money directly from consumers via subscription. What we found was actually pretty surprising: The median ARR, or annualized revenue run rate, is now $4.2 million at month 12 for consumer startups. The bottom quartile is $2.9 million, and the top quartile is $8.7 million.

Justine Moore

Wow.

Olivia Moore

A median B2C company in the age of AI is getting to $4 million ARR after a year, and the best-in-class companies are getting to $8 million in a year, upwards of $8 million. In the pre-AI era, we never would have seen anything like that. The even more surprising thing is that those numbers are twice as high as the B2B benchmarks in the AI era. Consumer companies are actually ramping revenue faster, which again is a total reversal from what we saw before.

I think there are a couple of reasons why this was happening. First, why have consumer AI companies adopted subscription? They were kind of forced to, because especially in the early era of models, they were so expensive that you, as a company, had pretty high cost of goods sold.

Justine Moore

The inference cost, you mean, of running the model?

Olivia Moore

Historically, the benefit of software was that there was no marginal cost. You made an app, and there was no additional cost to serving the next user. In AI, that is actually not true at all, especially if you’re running inference on a model. It costs you cents, maybe even dollars, for each query. So each user could be costing you dozens of dollars a month.

Justine Moore

Yeah, absolutely. So a lot of companies had to at least try to charge.

Olivia Moore

Right. And it turns out that these new, AI-native products are so powerful that consumers are willing to pay. We also ran some additional data analysis that shows that, on average, consumer AI startups are charging an average user $22 a month, which again is more than double what subscription companies were able to charge on average pre-AI.

Justine Moore

And why do we think that is? Do we have theories?

Olivia Moore

I mean, on the creative tools side, what I’ve seen is that for people who weren’t creative, AI tools allow them to make photos, images, or art for the first time. They can make videos and animations. And for creative people—we have a cousin who’s a creative—they can genuinely use this to supercharge their workflows and do their job a lot faster.

Justine Moore

So they're willing to pay for it. Have we seen examples of that outside of creative tools yet?

Olivia Moore

It's a good question. We've seen some of it around companion apps, where the products are just so powerful that having a friend with you 24/7 makes people excited to pay. We've also seen this around categories like language learning, teaching your kid how to read, and other things that previously you'd have to pay a human being $50 an hour, if that, to get access to. Now, $22 a month with AI feels pretty cheap.

Justine Moore

Totally. I'm even thinking some of the things I've seen monetized well are in nutrition or coaching, where, for the first time, thanks to vision models, you can take a picture of what you're eating and have a vision model pull out how many calories are in it, how much protein, and then, at the end of a day or week, summarize insights about what you should be eating more or less of. That's something that, pre-AI, you could maybe do by taking a picture and uploading it to a forum, but you'd have to find a nutritionist, wait weeks to book an appointment, and maybe get a referral from your doctor. It would take forever. So, it's really exciting. I think it's monetizing people who never would have paid for that before.

Olivia Moore

Yes. And then people who would have paid for it are switching over to the AI version, or they're willing to pay even more, which is super exciting. The other thing that I think people have questions about, or maybe doubts about, would be, like, okay, great, they're growing fast, but they're not retaining a lot of users. We did some analysis on this, too.

There's definitely a lot of tourism—we call it AI tourism behavior—in terms of free users, which means you get a lot of hits on your website, essentially, and most of those users don't stick around. But if you look at it on a paid-user basis, once you actually subscribe, consumer AI companies are retaining, at the median, pretty much just as well as pre-AI consumer companies, which is really exciting.

I feel like what we've seen, especially on revenue retention, which is fascinating, is that you might have more tourists, so you might have more people who subscribe and then cancel, but for the first time you have real upsell activity in consumer subscriptions. You're not just paying $10 a month for the app. You're paying $10 a month for the image model, but then, if you love it and you run out of credits, you're paying another $10, $12, or $50 for additional credit packs before the next month of your subscription starts. And so that means you see revenue expansion opportunity now in consumer that we only used to see in enterprise.

Justine Moore

Absolutely. Or, honestly, in games, what they call monetizing the whales—the people who are the really high spenders—which is, to me, one of the most exciting things about consumer AI products right now.

Olivia Moore

Yeah. And we're seeing, I would say, companies convert consumer revenue to enterprise revenue way faster than they ever did before. Companies like Canva previously took 5, 6, 7-plus years to really move from consumer/prosumer to enterprise, right? And now we're seeing companies like ElevenLabs. It's a great example: it starts at $10, and then they convert to a really high ACV enterprise contract, which is super exciting.

Justine Moore

I mean, I feel like we even saw that in the very early days of consumer AI. I remember our friends at ad agencies or entertainment companies would tell us that they were using Midjourney to mock things up, or even using the images in their final work product. So it was a true enterprise use case, but growing bottoms-up, which is sort of a fascinating motion.

Olivia Moore

Yeah, it's exciting. Consumers are back.

7. Demo of the Week: AI in Brand Creation

Justine Moore

All right, awesome. We're moving on to our demo of the week. One fun fact about us is that we genuinely love—at least for me, it's probably my number 1 hobby now—trying out all of the AI creative tools, especially, but also AI consumer products more broadly, figuring out how to make cool things and then sharing the workflows with other people whose number 1 hobby is not doing this. So this week, we're going to talk about brand creation and ideation using AI.

I made this new frozen yogurt brand called Melt that I iterated on with ChatGPT. Then I took it to Ideogram, and then I took it to Krea to do the final touches and make these really cool product photos and even store photos. I think the initial idea about this was seeing FLUX Kontext come out, which is the new image-editing model from Black Forest Labs, hosted on Krea.

FLUX Kontext is kind of like the GPT-4o image model, where you can upload an image and then say, “Make this Ghibli-style,” which was the viral example. You can also say, “Take the person from this photo and put them in a new environment,” or, “Take the logo and change it slightly.”

Olivia Moore

Yeah. Add or remove objects. I've seen it described as kind of like Photoshop, but with natural-language prompts.

Justine Moore

Yes, you can edit with words for the first time. And I think that is what makes it different from the GPT-4o image model: the consistency with which it retains the item or the character, whatever, is much, much better. We'll show some examples here, but basically, if you're taking a photo of yourself and uploading it to GPT-4o and saying, “Put me in a podcast studio,” you will likely end up looking completely different in the new photo than you did in the initial photo, or maybe have some similar features, but quite different. Whereas this model does an amazing job at maintaining consistency.

And so that sparked this idea for me: that means this can actually be used for brands to do product photos or other sorts of marketing collateral, because the logos and the products can be consistent.

Olivia Moore

Awesome.

Justine Moore

And, you know, I'm a huge froyo fan. I feel like froyo has gotten an unfair shake in recent years. It's kind of seen as the little-kid thing. And so I wanted to make a cool, hip, modern 20s New York froyo brand.

I went back and forth with ChatGPT on this idea to land on the name Melt, which I love, and to land on the branding: this is sort of what the font of the logo will look like, and then this is the color of the packaging. Then I took that prompt for the logo to Ideogram, which is an image-generation and sort of editing canvas, and it's super good, I think, at logos, typography, and anything sort of product- or word-related. I had it generate this photo of a froyo cup floating in the air with the Melt logo and branding.

Then I downloaded that photo and took it to Krea, where I used the new FLUX Kontext editing model to run all different kinds of scenarios. What's really cool about that is you can upload the photo and then say, “Take this froyo and put it sitting on a counter at a trendy restaurant. Put it in the hand of a woman at a park.” Or even, “Make the froyo cup white instead of blue, and give it a pink border. Make the froyo itself purple if they're having an ube froyo special.”

Olivia Moore

Yeah.

Justine Moore

And then I think the next step, which I didn't do here—I kind of stopped at the product images—and I actually made an image of the store, too. I took the logo and superimposed it over the store I generated.

Olivia Moore

I know. You want to go there.

Justine Moore

But the next step even further would be video. My idea is to take all of those product shots, take them to Veo 3 or Higgsfield, which does really cool special-effects stuff, and have the froyo cup in action. Have it actually melt over the side. Have it melt. It has to melt.

I'm very curious to see: do the models understand the physics of froyo? If it tosses the cup in the air, how does the froyo land? Does it kind of plop like we all know froyo would in real life?

Obviously, this is just a fun experiment for me. I'm not, unfortunately, actually going to be starting a froyo brand, but it sort of makes you start thinking about this: if you work at an ad agency, for example, and you were mocking up a deck for your client about your latest campaign, why would you not use something like this to show them what it might look like?

Olivia Moore

And you did it in less than a couple hours. Honestly, the branding does look more exciting than a lot of professional brands that we see out there. It makes me think about the next generation of entrepreneurs, who are going to be completely AI-assisted in a lot of these assets that they're putting together.

I think they're going to be able to make full-stack AI brands. There are also products where you can design with AI. You can make ads with AI. I think there'll be no reason for any person not to have their own product line, small business, or open a store if they want to. AI is assisting with these kinds of things, too.

Justine Moore

Totally. Yeah. I think we'll see brands that are logo, product photo, maybe even product itself, designed by AI; vibe-coded or vibe-designed website or mobile app; and then kind of drop-shipped to the end consumer. Social media ads also generated with an AI avatar that holds it up for you and sells it on TikTok or Instagram, promoted by AI influencers generated by Veo 3. They don't actually exist.

I think that sort of thing is going to be really fascinating to see because you no longer have to know how to work all of these technical tools that you had to be able to use. Even Photoshop has so many buttons; it's very complicated. And now you can just ask for what you want in a text prompt, get something generated, and iterate on it until you end up with something that you really love, which I think is crazy powerful.

Olivia Moore

Awesome.

What You Missed in AI This Week (Google, Apple, ChatGPT) | BidClub