本周 AI:GPT-5 发布,4o 被撤回,Grok Imagine 走向社交化
GPT-5 的发布表明,基准测试实力并不能保证消费者偏好。 它在编程、调试、数学和医疗问题上更强,但用户想念 GPT-4o 富有表现力的个性——“把那个老玩具还给我们”。遭遇反弹后,据报道 Sam Altman 表示 GPT-4o 将回归付费用户。主持人认为,市场对有娱乐性、类似陪伴者的模型存在巨大需求,它们不必拥有最高 IQ。
Grok Imagine 的优势在于分发和延迟,而非前沿级输出质量。 图片“基本即时”生成,视频也很快到达;X 用户长按自己或他人的照片,就能将其制作成动画或编辑。这个紧密的社交闭环、对手机相册的调用,以及生成真人内容的意愿,可能让 Grok 成为社交原生 AI 创作的重要试验场。
OpenAI 的 GPT-5 直播重点展示医疗辅助,Justine 将其解读为从默许非适应症使用走向公开背书。 直播展示了一名癌症患者用 ChatGPT 解读文件、讨论治疗方案,同时强调 GPT-5 在由250多名医生参与构建的 HealthBench 上排名第一。Illinois 则走向相反方向,广泛限制没有持证监督的 AI 治疗;一些公司已暂停在当地开展新业务或关闭新用户注册,而主持人质疑,私人聊天实际上能否被监管。
Genie 3 将生成式视频从短片推进到按需出现的交互式世界。 用户可以在由文字、图片或 Veo 3 视频生成的场景中移动,既可能录制可控电影、加速专业游戏开发,也可能生成个人小游戏。更深层的基础设施机会,是为数字智能体和机器人批量生成 RL 环境;不过主持人指出,这套系统成本高,而且大概率运行缓慢。
ElevenLabs 的全版权音乐模型,瞄准把素材来源视为采购前提的买家。 制作生日歌或表情包配乐的消费者可能不在意训练数据来源,但广告主、影视工作室、游戏公司和其他企业在意。来自全版权数据的强劲输出,挑战了“高质量必然依赖法律争议性抓取”的假设。
Vibe coding 已验证需求,但安全性和用户分层仍未解决。 Olivia 用几小时做出一款把用户与 Jensen at NVIDIA 合成自拍的应用;约3,000人一夜之间使用,耗尽了她100美元的 API 预算,但随后有外部人士发现 API key 暴露、用户上传照片的存储保护不足。她表示问题已经修复。主持人预计,今天“什么都服务所有人”的平台最终会分裂成有安全护栏的消费级工具,以及高度可配置的企业级产品。
1. Grok Imagine 把 AI 生成变成社交动作
Grok Imagine 已通过 Grok app 提供服务,网页版即将上线,同时嵌入 X:长按一张已发布的照片——无论是自己的还是别人的——即可编辑,或将其制作成视频动画。
主持人明确提醒,Grok Imagine “不是最强的”图片或视频模型,音频也只能算过得去。它的优势是速度——图片“基本即时”生成——这鼓励用户快速迭代,不到一周就让它成为移动端首选生成器。
调用手机相册后,制作表情包和老照片动画变成“一键操作”,不到1分钟即可完成。相比其他产品显著限制名人内容,Grok 允许生成真人,进一步扩大了表情包创作空间,也体现其相对不设限的立场。
2. GPT-5 能力提升,却失去了 GPT-4o 的消费级亲和力
讨论重点包括 GPT-5 在前端代码生成和调试上的优势,这与发布时将编程定位为经济价值来源的叙事一致。但多数消费者主要想要聊天;在这件事上,GPT-5 显得没那么有表现力:感叹号、emoji、全大写反应少了,也少了“它不只是好。它很棒”这类修辞。
两者必须区分:减少让模型变得无法信任的“过度迎合”是好事,但轻松、及时响应的个性是另一项功能。主持人认为,GPT-5 在后一个维度上“退了一步”。
GPT-4o 被直接移除,而不是作为选项与 GPT-5 并列,也让主持人感到意外,尤其是在 OpenAI 已围绕 GPT-4o 的图片生成搭建了一系列界面功能之后。用户“彻底炸锅”后,据报道 Sam Altman 在 Reddit 回复称,GPT-4o 将回归付费用户。
更大的结论是:客观基准上的“最聪明模型”,未必是人们最想聊天的模型;这为陪伴型和娱乐型产品留下空间,它们优化的是趣味性,而非最高 IQ。
3. 医疗 AI 获得产品背书,监管同步收紧
Illinois 刚刚广泛限制没有持证专业人士监督的 AI 治疗;持续支持或针对情绪问题提供个性化建议,也可能被纳入范围。一些心理健康公司暂停在 Illinois 开展新业务或关闭新用户注册,主持人则质疑:私人聊天中的规定能否执行,以及这些限制最终是否会伤害消费者。
与此相反,OpenAI 的 GPT-5 直播展示了一名癌症患者上传文件、讨论诊断和治疗方案,并强调 GPT-5 在由250多名医生参与创建的 HealthBench 上的表现。Justine 认为,OpenAI 不再只是默许医疗领域的“非适应症”使用,而是在“真正为它背书”。
4. Genie 3 与全版权音乐拓宽创意模型栈
Google 尚未发布的 Genie 3,可以根据文字、图片,甚至 Veo 3 视频生成交互式世界。用户向左转或穿过场景时,画面会即时重新生成,就像走进一幅画,或操控一款个人电子游戏。
主持人指出,这套系统成本高,运行时间可能也很长,但仍梳理出两条游戏路径:开发者可以生成并冻结世界,用于传统多人共享游戏;也可以让每名玩家创建个人小游戏,随着探索不断重新生成。在这些世界中移动并录屏,也比提示词生成固定视频片段更适合制作可控电影。
Genie 3 还可能大幅降低生成海量 RL 训练环境的门槛,为数字和实体智能体提供人工制作模拟之外的补充。市场需要的是能让系统学习移动、导航和与物体互动的世界。
ElevenLabs 则针对另一项瓶颈,将音乐模型训练在“全版权音乐”上。对于生日歌和表情包配乐,这种素材来源没那么重要;但对广告、电影、电视和游戏而言,企业需要在商业使用时避免引发版权方追责。
5. Vibe coding 的下一批赢家,将围绕用户风险完成专业化
Olivia 发布的第一款 vibe coding 应用使用 Lovable、fal.ai 和 FLUX.1 Kontext,把用户与 Jensen at NVIDIA 合成到自拍中。应用在几个晚上的时间里完成,随后一夜吸引约3,000名用户,并耗尽她自行设定的100美元 API 预算。
预算用完后,应用不再调用模型,而是生成一张“看起来像2005年 Microsoft Paint 做的”拼接图。更严重的是,有人提醒她 API key 已暴露,用户上传照片的存储桶也不是私有的;她表示已经修复问题。
Justine 和 Anish Acharya 的判断是,现有平台不可能永远“什么都服务所有人”。面向消费者的“带辅助轮版本”应阻止不安全部署,牺牲灵活性,支持移动端,并在5分钟内产出可用结果。
专业用户和企业用户需要相反的产品:能够控制语言、数据库和技术栈,并接入设计系统、CRM 和邮件平台。它们在企业内的获客可能来自自上而下的销售,也可能来自产品驱动增长;消费产品则通过 TikTok 和 Reels 扩散。早期赢家已经是 AI 应用中增长最快的一批,但这个市场仍“太早”。
Welcome back to This Week in Consumer. I’m Justine.
I’m Olivia.
And we have a bunch of fun topics we want to cover this week, starting in the creative tools ecosystem with Grok Imagine. Then we’re also going to talk about Genie 3 and the ElevenLabs music model. Then we’ll cover GPT-5 and the deprecation of GPT-4o, and we’ll cover our new vibe-coding thesis.
So this week, we are going to start with Grok, which has had a bunch of big updates over the last month or so. Obviously, Grok 4 came out. The Grok companions caused a huge stir, particularly Ani and Valentine. But I think more recently, what’s been really interesting is all of the image and video generation features on Grok Imagine.
Yeah. So Grok released an image and video generation model called Imagine, which is offered standalone through the Grok app. They’re also bringing it to the web, and it’s now embedded into the core X app as well, which is really exciting. I think that’s one of the things that’s really unique about it. I would say it’s not the most powerful image or video generation model that exists. Elon has tweeted a bunch about how they’re training a much bigger model, but I think what’s really cool about it is that it’s one of the first really truly social forays into AI image and video generation.
What you mean by it being integrated into the X app is that now, when you post a photo on X, you can long-click and press and immediately turn it into an animated video in the Grok app—
Or even if you see someone else’s photo posted on X, you can turn it into a video or edit the image with Grok, which is really exciting.
Totally.
I think one of the coolest things about Grok Imagine—to your point, it’s not the most powerful model. It’s not Veo 3. On video, I would say the audio generation is okay, but not great. But it’s fast.
So fast—
Which, I think, for a lot of people has been the real barrier to doing image and video generation more seriously. You put in a prompt, press go, and sometimes you’re waiting—
30, 60, 90 seconds for a generation.
Yeah. Often, it’s minutes for a generation, honestly.
And Grok images are basically instant, and the videos are pretty fast as well. I found myself iterating very frequently. In less than a week, it’s become my go-to tool for image generation on mobile.
Yeah.
And even the video, I would say, is getting there, especially if they’re training a better model.
Totally. I think it’s also that many people aren’t professional creators, so they don’t want to make an image in one place, download it, and then port it into another website, because very few other tools are on mobile, especially for video generation. I think it’s huge that you can click one button and get a video on your phone in less than a minute. That feels like a massive step forward for consumer applications of AI creative tools.
Elon and a bunch of folks on the xAI team have been tweeting about this. One of the big use cases is animating memes, animating old photos, or animating things that you already have on your phone, because you can access the camera roll so quickly through the Grok mobile app.
Yes.
Elon has been tweeting many Imagine-generated photos and videos of himself.
Yes.
But it will also generate real people, which I think is another big differentiator. It’s something we’ve really only seen from Veo 3, and even then, it’s mostly characters versus celebrities. But it comes from Grok’s uncensored nature, which is pretty—
I think it’s cool, and it unlocks a whole bunch of new use cases.
Yeah, for sure. I think that allows the meme generation, and even with Veo 3, half the time I try to do an image of myself, it’ll say, “Blocked due to our prominent-person thing,” and I’m like, “I’m not a—what do you mean? I’m not a prominent person?” But in that photo, I guess I look too much like some celebrity or prominent person, and it decided to block it. I’ve never had that problem on Grok, which makes it so fun and easy to play around with.
Yeah, I’m excited to see where they take it. It feels like we’ve seen Meta experiment a little bit with AI within its core products. They’ve done the AI characters you can talk to, as well as uploading a photo to get an avatar where you can generate photos of yourself. But none of it felt quite right, I would say—
In terms of baking it into the core existing experience on Instagram or Facebook. Grok feels a little bit different, so I’m excited to see where they go with it. I’d say none of the existing social platforms have leaned that heavily into AI creative content. A lot of the AI creative tools can, should, and will integrate more social features, but today most of them have just done relatively basic feeds and liking—not really comments, and not really a following model.
Agreed. The other big model news of this week, which was not just big for consumers but for pretty much all of AI, was the GPT-5 release and the corresponding deprecation of GPT-4o, which I think ended up being even bigger news in consumerland specifically.
Yeah. This was fascinating because, obviously, it’s been a while since OpenAI had a major LLM release. Since GPT-4, people had been very eagerly awaiting GPT-5. But as soon as I got access to GPT-5, I wanted to compare the outputs to GPT-4, and I immediately noticed GPT-4 was gone.
Yeah.
I’ve seen a lot of posts with people up in arms about GPT-4o disappearing. How would you describe the main differences between the models, at least in how they’re manifesting in user experiences?
I’ve talked to a bunch of folks about this. I think a widespread conclusion is that GPT-5 is really good at front-end code. A lot of the model companies are focusing on coding as a major use case, a major driver of economic value, and something they can be really good at. You can tell in the results from GPT-5—
And they emphasized it in the livestream pretty significantly.
You can see from the examples people use that it’s much better at generating things and much better at debugging. But a lot of consumers aren’t using it for code. A lot of consumers just want to chat with it, and there are a bunch of examples of how it’s a lot less expressive, emotional, and fun. It doesn’t really use exclamation points or emojis. It doesn’t send things in all caps like it used to.
It doesn’t do the classic, “It’s not just good. It’s great.”
Yes, exactly. I think there are 2 separate issues here. One is the glazing, the excessive validation. It would say, “You’re the best. You should totally do that. That’s the best decision for everything you said,” even if it was ridiculous. That’s a problem that I’m glad they’re working on and getting rid of, not to mention everyone’s concerns about GPT psychosis or whatever. You just can’t trust something that always tells you you’re right.
The second thing is whether it has a fun and engaging, more casual, human-feeling personality. I think that actually maybe took a step back from GPT-4o to GPT-5, and that is what people like. If you look at the ChatGPT subreddit—
People are freaking out, and I think that’s why Sam rolled it back. He may have announced this on Reddit, in a comment responding to all of the backlash, where he said, “We hear you guys. We’ll bring back GPT-4o for paid users.”
I was surprised they even got rid of GPT-4o. I know there had been a lot of jokes about what a pain it is to have to select the model and how the dashboard was always getting bigger. But they had even started building some UI around GPT-4o image generation. They had preset templates you could use. So the fact that they didn’t just add GPT-5 as an option, but took away your ability to use every other model, was a little surprising to me.
Yeah, I imagine there’s image generation on GPT-5, right? I imagine some of the templates and editing tools are just going to move over between the models. They may not have gotten there yet.
I think it’s funny because, if you imagine yourself in the shoes of one of these researchers, you’re thinking, “We trained what is, on the benchmarks, clearly a much better model. It’s smarter, it’s better at math, it’s better at coding, and it can answer medical questions now,” which they really focused on in the livestream. Of course everyone will love and embrace this with open arms. It’s a step forward in model intelligence.
Move toward AGI.
Exactly. Of course, classic consumer is like, “No, we don’t want that. Give us the old toy back. Give us our fun friend who mirrored the way we spoke to it and was over the top and sometimes crazy, but was really fun to chat with.”
I think, to me, honestly, this exemplifies something I’ve suspected for a long time, which is that I don’t necessarily think—
The smartest model that scores the best on all these objective benchmarks of intelligence will be the model that people want to chat with. I think there will be a huge market for more of these companionship, entertainment, and just-having-fun-type models. They don’t need to be the highest-IQ person, you know.
Yeah, I agree. I do want to spend 30 seconds on that mental health and health overall use case, though. It's interesting timing because, also last week, the state of Illinois just passed a law banning AI for mental health or therapy without the supervision of a licensed professional. It's pretty interesting because the law is wide-ranging, to the extent that some AI mental health companies have already shut down new operations in Illinois or prohibited new users from signing up. It's basically anything that's ongoing support or even personalized advice around specific emotional and mental issues is now counted as therapy and is technically illegal in Illinois.
Yeah, I am confident ChatGPT, honestly, is doing it well for a lot of people. I guess my question is: to what extent is this ever going to be enforced, because they can't see people's individual chats? I feel like Illinois always does weird stuff. We've been consumer investors for too long, and I remember in 2017 and 2018 we would literally talk to social apps—consumer social apps—that were like, "We've launched everywhere except for Illinois," because they had all these crazy regulations around people's data and sharing and all of these things. Obviously, it's good to have those, but they went way beyond other states, to the point where it made it difficult for apps to operate there, which is, in my opinion, bad for the consumer.
I think there are a lot of people now grappling with this question of what it means for AI to offer medical support or mental health support. I don't expect we'll see the other states go in the direction of Illinois, partially because it's just so hard to regulate. How can you control what someone is talking to their ChatGPT or Claude or whatever about?
Well, and especially because GPT-5 was trained, or at least fine-tuned, with data from real physicians. Is that right?
Yeah. They talked about this a lot in the livestream, and I was surprised they leaned in on this. I'm sure we've all seen the viral Reddit posts about "ChatGPT saved my life." My doctor said, "Wait for this imaging scan." It turns out I had this horrible thing that I was able to get treated immediately. Sam Altman and Greg Brockman had been retweeting these posts for a while, which I thought was interesting because, from a liability perspective, you would think they'd avoid that.
But they had a whole section of the GPT-5 livestream where they talked about someone who had cancer and was using ChatGPT to upload all of her documents, get suggestions about treatment, and talk through the diagnosis and what she could do next. They talked about how GPT-5 was the highest-scoring model on something called HealthBench, which is a benchmark they developed with 250-plus physicians to measure how good an LLM is at answering medical questions. I think it's a really big statement that OpenAI has leaned into this space so heavily, versus being like, "Hey, there's a lot of liability around medical stuff. Our AI chatbot is not a licensed doctor. We're going to let people do this off-label, but we're not going to endorse it."
Yeah.
It seems like now they're really endorsing it.
I'm excited.
Me, too. I enjoy uploading all sorts of stuff and getting all kinds of advice. It can be really smart and really helpful in a lot of cases.
I agree. There were 2 other big creative tool model releases this week: Genie 3 from Google and a new music model from ElevenLabs. So maybe let's start with Genie 3. What is it? I've seen the videos, but what is it?
Yes, Genie 3 took Twitter by storm. Google has a bunch of different initiatives around image, video, and 3D worlds. I think various teams, like Veo 3 and the Genie team, are working toward this idea of an interactive world model, which is basically that you're able to have a scene that you can walk through or interact with in real time, and that generates on the fly. You can imagine it like a personal video game.
Yeah. I saw some of the videos of taking famous paintings and, for the first time, being able to step into them, swivel around, and move around in the world, almost like you have a VR headset on and you're turning around and seeing the full environment.
Those were really cool. And it's not just famous paintings. They've shown a bunch of examples: from a text prompt, you can create a world; from an image, you can create a world. They've even shown taking Veo 3 videos and creating a world around them with Genie 3.
The cool thing about Genie 3 is that there are controls where you can move the character around. You can say, "Now go to the left," and then the scene regenerates to show you what you would see on the left. It's incredible. They haven't released it publicly yet, but they invited some people to try it out at their office, and those people were sharing results. They've shared a bunch of clips, and I'm personally really excited to get my hands on it.
The natural question we've all had with this use case and seeing the demos is: this looks amazing—what are we going to do with it?
Exactly. Yeah.
And it's expensive and probably takes a long time.
I think there'll be a couple of use cases. Video is an obvious one: if you're generating the scene in real time and then controlling how you, or any character or objects, are moving through it, that enables much more control over a video. You could then screen-capture what is happening, which gives you more control than you would get from a traditional video.
So you're almost recording the video as you move through the 3D world model, which then becomes a movie or a film, essentially.
Our portfolio company World Labs has a really cool product out that a number of folks and I are on, which does this. Martin on our team shares a bunch of really cool examples of stuff he makes with exactly that use case.
Very cool.
So, much more controllable video generation, which is huge. I think in gaming there are 2 paths this can go, and it could go both ways. One is that it allows real game developers to create games much more quickly and easily, where you don't have to code up and render an entire world. It can just generate from the initial image or text prompt you provide and the guidance you give it.
Could a game developer freeze that world?
Yes, and allow other people to play it like a traditional game. The game then would be the same for every person. In the first example, the game almost regenerates for everyone as they move through it.
Right. And then I think the second gaming example is more like what you're alluding to, which is more personal gaming. Every person puts in an image, video, or text prompt and then creates their own minigame, wandering through a scene. That's a totally new market that I think a lot of people will love.
Yeah. And then the third example, which is a little out of our wheelhouse, is that a lot of folks are talking about how creating these interactive, dynamic worlds are really good RL environments for agents to be trained on how to interact with the world: how things move, going around scenes, and interacting with objects. It's been a big space of conversation right now, and there's a desperate need for more. There are so many companies now selling these RL environments for agents that they're manually creating. Something like Genie 3 could make that much easier and allow you to generate unlimited environments for agents to wander through and learn.
I could see that for digital agents, but even physical agents operating within robots or something like that.
Totally. I think for all sorts of agents or self-learning systems, it's going to be fascinating. So I'm eagerly awaiting that one to come out. And then, yes, our portfolio company ElevenLabs also released its music model.
It's super exciting. I did not know they were working on music.
Yes, it's been in the works for a bit. The really interesting thing about it is that it's trained on fully licensed music.
Music is one of those spaces where the rights holders are extremely litigious.
And so, compared to things like image or video, it's been harder for music companies to avoid stepping on rights holders' toes.
You can't just scrape data from the internet; the record labels will come and sue you.
Yes, and the artists. It's often a very complicated ecosystem of who owns the rights to a specific song, or to an artist's voice, or something like that. A lot of folks have thought that you couldn't get a good-quality music model training on licensed data because it's hard and expensive, it takes a long time, and it's hard to get rights holders to agree to license you the data. But from what I've seen and from my own experiments, people have been really impressed by ElevenLabs' output.
Floating on a midnight plane. Jazz in my veins. Let it rain. Loose in the haze. We feel no pain. Loop to the sound. Break the chain.
And so what does the licensed data open up in terms of use cases for the music model, do you think?
Yeah.
So, I think a lot of consumers basically don't care if they're using a music model that's trained on licensed data or not because they're not really monetizing—or many of them are not monetizing—the stuff that they make with this music.
They're generating a birthday song for their friend, a meme clip, or something like that.
Or background music for their AI video.
Yep.
Whereas businesses, enterprises, big media companies, and gaming companies care. They need to be able to say, “This music model we used was trained on fully licensed data,” so they don't open themselves up to liability issues.
So, they could hypothetically use this music in advertisements, films, TV shows, or things like that.
Exactly, which I think is a big step forward for AI music as a whole. I think we should expect to see more from ElevenLabs on this front, which is very exciting.
Awesome. And then our last big topic of this week is vibe coding, which continues to explode. I think we have 2 things to talk about here. One would be our own experiments in the world of vibe coding, which relates to a piece that you and Anish Acharya put out this past week about how we're seeing the vibe coding market start to fragment. Your experiment is the more interesting part, so let's start with that.
Yeah. Maybe to give a real-world example, for the first time I vibe-coded an app that I fully published and made available to the internet. Essentially, what I did was think, “Hey, I'm seeing on my X feed all the time that everyone has a selfie with Jensen at NVIDIA.”
How did they get this?
In his classic leather jacket. He must be spending all of his time taking selfies now because everyone has one and I don't.
Yes.
And so I was thinking, there are all these new, amazing models out there, like FLUX.1 Kontext, that can take an image—say, of Jensen taking a selfie with someone else—and put myself in there instead.
You should have been in the photo.
I should have been in the photo. Exactly. So, I did that. I generated that myself on Krea, and then I thought, “I bet other people might feel like me and might want this.” I wanted to create an app where anyone could upload a photo and get a selfie with Jensen.
Yes.
And so I thought, “Okay, I can vibe-code this.” I vibe-coded on Lovable an app that connected to fal.ai to pull in the FLUX.1 Kontext API. You could upload your own photo, and it would generate the selfie with Jensen, which you could then download. It was great. It worked.
I published it on Twitter, and a lot of people used it. It was used by about 3,000 people overnight, to the point where, when I woke up, I had exhausted my self-imposed budget of $100 to spend on API calls.
Yes. And because you were funding it—you were funding it—you weren't making people pay for it or put in their own API key.
I was not making anyone pay for it or put in their own API key. So, instead of calling the model, it was just stitching together half of your photo with half of Jensen's photo to produce a really 2005 Microsoft Paint-looking output, which has its own charm.
Yeah.
Anyway, the surprise was, first, that someone who's completely nontechnical can build something that thousands of people can use. I did it in a couple of hours one evening, if that, and a couple thousand people used it overnight. That's amazing and so exciting.
Yes.
My second learning was that we're still early in vibe coding because these products are definitely built for people who are already technical.
Yeah, there were some issues we should talk about.
You should not expose your public API key.
The problem is, you didn't even know you were exposing your public API key until some nice man DMed you and told you that the vibe coding platforms, I think, assume that you have a certain level of knowledge already. So, if you go to publish a website, they won't stop you and say, “Hey, here's a security issue. Here's a compliance issue. Fix this before you publish.”
And so it was a really interesting learning experiment for me. I think—and this is what you got at in your blog post—that there'll hopefully be a V2, V3, or V5 of these vibe coding platforms that are built for people who don't know these things already.
So, 2 things people flagged to you that the vibe coding platforms did not were, first, that your API key was exposed. Second, you had not created protected private buckets for the photos that were uploaded. So, if you knew how, you could access the selfies that were uploaded.
Yeah. I fixed that, to be clear. I've had similar problems vibe-coding a lot of apps, where I feel like they assume you have a level of technical knowledge to be able to fix things or even know what a potential problem could be.
Anish Acharya was actually an engineer, and he and I published a post about how we think vibe coding will evolve in the future. I think today you have a bunch of awesome platforms that are trying to be everything to everyone. They're saying an engineer at a company can use this to develop internal tools, someone can use this to build a SaaS app that scales to hundreds of thousands of users, and a consumer can also use this to create a fun meme app.
But I think the truth, in terms of what we've seen at least, is that those are very different products, both in terms of the use cases and integrations and the level of complexity required. There probably should be, for example, a platform that's like the training-wheels version of vibe coding for consumer, non-developers like us, one that does not allow you to make mistakes like exposing the API key.
Yes, even if it then means less flexibility in the product.
Exactly. I wasn't super opinionated about what it looked like or about any of the specific features. I just wanted it to work.
Yeah, you probably weren't super opinionated about the coding language, exactly what database it was using, or all the backend details of what it was built on top of. You didn't really care. You just wanted it to work.
Whereas there are many enterprise or true developer use cases where people very much want to control every element of the stack, and that level of inflexibility would just not work for them.
I think what we're hoping to see is specialized players emerge that offer the best product for a particular demographic of user or for a particular use case. That will probably imply very different product experiences, product constraints, and go-to-market strategies.
If you're allowing any consumer to vibe-code a fun app to share with their friends or partner, you probably want to be going viral on TikTok and Reels. But if you were building a vibe coding platform for designers to prototype new features in an enterprise, or for engineers to make internal tools, you might want to have top-down sales, or at least product-led growth within businesses. You might also want to invest in deep integrations into core business systems and other things like that.
The consumer version might actually just let people vibe-code on mobile and get something that works in 5 minutes. And that's a great point, too: consumer users often just want something to look cool and work, without security issues. More business-oriented users often need it to integrate with what already exists for the business, whether that's a design system and aesthetic or their CRM, the emailing platform they use, or all of these different external products that it needs to connect to.
I think the conclusion of the piece was that we're seeing early winners in vibe coding already. These are some of the fastest-growing companies in the AI application space.
But we probably expect to see even more because it feels like we're so early.
Totally. And many of the users of these products are probably still pretty technical.
Yes.
There'll be a version of vibe coding that's truly consumer-grade.
Yes.
That's something I'm personally very excited to unlock. And I think we've seen this in a lot of AI markets, because these markets are large enough to have multiple specialized winners.
We've seen this with LLMs: OpenAI, Anthropic, Google, Mistral, and xAI. There are all of these companies that have models that are really good at particular things. We've also seen this a ton in image and video, which I think has a lot of corollaries to vibe coding.
Based on what type of user you are and what you care about, how much do you need to reference an existing character or an existing design format? Do you want it on your phone and super fast, or do you want it in the browser, slower, and at the highest quality? There are many companies doing well by focusing on different segments or verticals of this giant market.
Yeah, super exciting. Well, thanks for joining us this week. If you've tried out any of these creative models or had any vibe coding experiments yourself, we'd love to hear from you. Please comment below and let us know. And also, please feel free to ping us here or on Twitter if you have ideas of what we should cover in a future episode.