Gaurav Misra 与 Dwight Churchill——构建 Captions——[Invest Like the Best,第405期]
Gaurav Misra 将 AI 分为无边界的智能竞赛与有边界的渲染问题,这一区分会改变资本需求与护城河的持久性。 AGI 模型可能会持续让前代模型过时,“永远”没有尽头;而视频已有明确终点,因为只要预算足够,CGI 就能制作任何想象中的场景。AI 可能让渲染“容易100倍”;一旦质量接近完美,渲染就会变成耐久资产,业务也会越来越像软件。
Captions 的核心护城河不是模型质量本身,而是由产品生成视频数据的飞轮,并计划逐步转向完全授权的训练数据。 这款用2天做出的字幕应用登上 App Store 榜首,第二天早上每分钟约生成“600条视频”,从第一天起就内置了数据采集机制,用于改进未来模型。业务从脚本、录制、剪辑扩展到分发后,可用于训练的数据会持续复利;Gaurav 认为,未来可能由大众消费产品供给数据,再驱动付费 B2B 模型。
以人为中心的视频生成,可能比大多数用户预期更早接近实拍质量。 Gaurav 预计,再过1年到1年半就能生成“非常、非常好”的视频,并称与特定物体互动的能力将在6个月内实现,“基本可以保证”。视频扩散模型目前仍只有数百亿参数,而文本模型约为4000亿参数,随着规模扩大,性能仍有很大提升空间。
Captions 认为,AI 视频潜在用例中仅有1%-5%已经被解锁,因此其高速增长部分得益于在替代方案出现前进入市场。 付费产品 AI Creator 和 AI Edit 的受欢迎程度大致相当,绝大多数用户都在付费;两者结合后,用户只需输入几个词,就能得到一支完整视频。Gaurav 明确表示,这段没有竞争的窗口会随着更多用例变得可行而关闭。
AI 应用的定价均衡尚未形成,但用户已经展现出高于传统视频应用每月7.99-12.99美元区间的付费意愿。 Captions 可以收取每月25美元,一些消费级 AI 订阅甚至达到每月2000美元;不过 Gaurav 认为,可比模型稀缺可能支撑了这些价格。Dwight Churchill 警告,不要急于按被替代的人力成本给 AI 定价:CFO 本来就希望降低这项支出,而“传统订阅模式仍可能持续贡献alpha”。
当精致视频和虚构人物变得充足后,经济价值可能转向受信任的身份与高端制作元素。 Gaurav 预计,普通生成式形象的价值基本会归零;相反,那些已被大规模受众“认识、信任、理解”的身份会更值钱。高质量视频本身仍不可或缺——就像设计精良的网站最终成为标配,而不是失去价值——Dwight 则预计,低预算电影人将获得制作更复杂作品的能力。
投资者可能过度聚焦于巨型智能实验室的经济性,却低估了由异常精简团队构建有边界应用的潜力。 Gaurav 表示,做到世界级成果可能只需要“12人或更少”,一个有边界的基础模型投资大概需要数亿美元,之后的微调成本则相对低廉。Dwight 预计成熟期毛利率会非常高,但警告高毛利业务可能成为其他创业者的“完美攻击目标”,并以80%-90%毛利率的 CRM 行业为例。
1. 有边界的渲染可以成为资产;智能竞赛可能永无止境
更强的硬件、transformer、扩散架构和训练技术推动了模型规模扩大。随着规模持续提升能力,真正的约束变成了可持续数据:互联网并非无限,而高质量视频比文本或音频更庞大、更稀缺、处理成本也更高。
智能是一个没有终点的目标。模型可能超越最聪明的人类,却又被下一代模型取代;Gaurav 认为,资本被淘汰的周期“可能永远持续”,因为没人知道智能竞赛会在哪里结束。
视频生成不同,因为渲染在原理上已经解决:只要预算足够,制片厂可以制作人类、风景甚至巨龙。Patrick 的框架进一步明确了商业前提——想象与产出之间的摩擦不在于能不能实现,而在于成本;AI 则可能让制作“不是容易一点,而是容易100倍”。
2. Captions 从第一天起就把产品设计成数据飞轮
Captions 起步于高度商品化的剪辑市场,成本竞争让差异化变得困难。Gaurav 找到的切入口,是当时被低估的语音转文本准确率:第一款产品只是把文字放到视频上,用2天“打补丁式拼”出来,却意外在一夜之间登上 App Store 榜首。
第二天早上,Gaurav 告诉 Dwight,平台每分钟大约生成“600条视频”。即使是那个周末做出的原型,也已经布置了数据采集机制,让用户行为能够改进模型、带来更好的体验——飞轮不是获得增长后补上的护城河,而是产品最初的设计。
此后,Captions 从字幕扩展到脚本、录制、剪辑和分发,在每个环节收集有用信号。Gaurav 设想过一种类似 Google 或 Facebook 的结构:大众消费产品生成数据,再为付费 B2B 模型提供动力,而不完全依赖抓取公共视频库。他表示,Captions 计划大幅转向完全授权的数据,并认为在成熟且竞争激烈的市场中,这种保证可能更重要。
3. 扩散模型当前成本高昂,但规模化路径清晰
Gaurav 将扩散解释为反复去噪:从类似电视雪花的噪声开始,输入“穿蓝色衬衫的男人”等文本条件,每一轮都揭开一层清晰度。这与 LLM 根据此前上下文预测下一个词不同。
文本模型的参数规模已经达到约4000亿,而扩散模型仍处于100亿、200亿和300亿区间;Gaurav 认为 Meta 的 Movie Gen 大约有300亿参数。即使把 Captions 的全部训练视频语料下载下来,也要花费约100万美元,这体现了视频带来的存储与处理成本。
但质量曲线看起来非常陡峭:从臭名昭著的 Will Smith 吃意大利面案例开始,视频很快从“非常糟糕”变得令人信服。Gaurav 预计,1年到1年半内就能生成“非常、非常好”的内容,并“很容易想象”届时出现几乎无法与实拍区分的视频——甚至可能更早;他称这已经是最保守的情形。
Patrick 问,完美视频是否需要核动力 GPU 集群;Gaurav 的回答是“谁知道呢”,随后给出一个有边界的判断。渲染成本是已知的,扩散模型不必保留100个去噪步骤,而蒸馏可以把推理压缩到几个步骤,效率可能提升一个数量级,约为10倍。
4. 超高速增长来自解锁此前不存在的市场
Gaurav 说,现在做产品的即时回报在于因果关系清晰:工程师上线一个功能,第二天就能看到影响。更重要的是,每项新能力——无论是广告还是更高质量的生成——都会打开一个新的客户群,让 Captions 暂时成为“唯一能做某件事的公司”。
没有替代方案,有助于解释异常迅速的采用率和付费意愿,但 Gaurav 并不认为这种状态会永久持续。Captions 估计,潜在用例图谱中只有1%-5%已经被解锁,未来仍有多年扩张空间,但没有任何排他性的保证。
免费的传统产品覆盖录制和时间线剪辑;付费产品则以 AI Creator 和 AI Edit 对应这一套功能。Creator 可以生成一个说话的人——可以是授权形象、用户提供的演员,也可以是不存在的人;Edit 则用基础模型组装故事,取代关键帧、动画曲线和时间线。
绝大多数用户都在付费,Creator 和 Edit 的受欢迎程度大致相当。两者按顺序使用时,用户只需输入几个词,就能从零得到一支完整视频;未来的 Edit 提示词应当像对真人剪辑师下指令,例如“把它剪短到大概30秒”,或让画面“更有感觉”。
5. Captions 将模型收窄到沟通场景,而非所有视频类型
Gaurav 有意拒绝覆盖整个视频市场。Captions 瞄准的是围绕人物沟通展开的 A-roll,应用于营销、销售和教育,而不是通用素材或“兔子在火星上跳来跳去”。他说,Captions 目前是唯一一家专门针对这类 A-roll 生成训练基础模型的公司。
与品牌物体互动,需要采集人物已经拿着物体的视频、识别物体,并加入条件控制。文本可以生成一罐普通 Coke,但一瓶 FIJI Water 可能需要图像条件控制或多个角度;Gaurav 预计,早期版本将在数月内出现,6个月内实现这项能力,“基本可以保证”。
人体结构是专门的技术攻坚方向:手指、手臂、喝水、跳舞,以及其他需要协调的动作。Captions 针对人物进行训练,并计划加入骨骼条件控制——“我想让你跳的就是这个 TikTok 舞蹈”——让模型更容易学会正常的人体结构并执行指定动作。
Dwight 将这场竞赛定义为持续领先客户当下的需求,新版本在“第0天”就实现商业化。Gaurav 认为,交互设计仍未充分发展:用户不应只能“按按钮,等输出”,而是可以预览扩散过程中的中间步骤,并在生成尚未完成时改变方向。
6. 供给充足让精致视频成为标准,可信身份变得稀缺
Gaurav 将视频当前的变化比作2010年代的 Canva 和 Figma。简单建站工具让好看的设计变得普遍,但没有让它失去价值:一个好网站仍然是必需品,尽管 Dwight 开玩笑说,带有1990年代风格的网站“又变酷了”。视频可能沿着同样的路径,最终成为标配。
合成形象的逻辑不同。企业可以拥有虚构的代言人,但当任何人都能生成一张有吸引力的脸时,“形象本身的价值就会归零”。稀缺性会转移到那些被数千人或数百万人认识并信任的身份上;即使是虚构人物,也可能随着时间积累这种声誉。
Dwight 用电影作类比,区分制作成本与观众价值:一部成本约2.5亿美元的 Michael Bay 电影,与一部低成本电影,可能都能收取大致25美元的票价。生成成本下降后,小型电影人也许能尝试更复杂的作品;创作本身会改变,而真正依赖实体制作或高端制作的元素可能更具区分度。
7. 分发伙伴想要 Captions 的产出,而 Gaurav 认为 ByteDance 正在攻击产品
Captions 与社交网络处于半合作关系,因为它每天提供“数十万、数十万”条原创且无水印的内容。Gaurav 将其与 Instagram Reels 早期的问题作对比:当时平台上的重复视频明显带着 TikTok 水印。
他提出的竞争版图颇具挑衅性:Google 和 Facebook 已经不再是习惯性抄袭者,ByteDance 接过了这个角色。TikTok 的管理层很早就注意到 Captions;据 Gaurav 称,对方多次试图在整个品类上“抢走、杀死、摧毁”Captions,而不是展开合作。
当被问到“试图杀死”具体指什么时,Gaurav 指称对方复制了 Captions 的 App Store 描述、网站措辞、新闻稿语言,甚至完全相同的品牌色。他的直接结论是,ByteDance 的产品仍然平庸,但受益于 TikTok 的分发;Captions 的应对方式是把产品做得更好,而不是让路线图由模仿者决定。
8. 极小研究团队可以取胜,但产品市场匹配可能掩盖错误决策
生成模型人才仍然稀缺,因为专业能力需要多年积累,而技术每周都在变化。尽管如此,Gaurav 说这“并不需要一支大军”:只要招对每个人、配齐所有技术要素,可能12人或更少就能做出世界级成果,并击败规模大得多的组织。
Dwight 的招聘方案很实际:提供算力、专有数据,以及让研究人员真正发布成果的环境。大型 AI 实验室里有些人无法上线自己做出的东西;给他们资源和商业落地的即时反馈,可能让招聘变得没有头条所说的那么复杂。
Gaurav 从 Snap 得到的教训是,产品市场匹配可以在错误行为持续存在的情况下延续,让员工误以为增长证明了每一个决策都是正确的。他认为 Snap 的 CEO 拥有很强的产品直觉,也把他带入了一个由设计驱动的圈子。Dwight 说,即使公司已经拥有数千名员工,那个设计团队仍只有约10-12人,而参与其中塑造了他的设计职业生涯。
9. AI 今天支撑更高订阅价格,但按人力定价可能令人失望
Gaurav 的答案是,在用例覆盖率仍只有约3%-5%的情况下,无法知道定价均衡在哪里。传统消费级视频应用的月费集中在7.99-12.99美元;Captions 收取过25美元,有时甚至不给一次免费使用,但客户仍会回应:“给你钱,继续吧。”
在 AI 产品中,消费者每月最多可能支付2000美元,但 Gaurav 保留了限定条件:高质量模型仍然稀缺。如果市场上出现大量可比的“Tesla 车型”,竞争可能压低价格,即使底层能力继续提升。
对 B2B 买家而言,数据授权可能只有在多年后的饱和终局中才成为决定性因素。Captions 计划使用通过自有平台收集的完全授权数据进行训练;在竞争性交易中,企业可能会选择这种保证,并愿意为此支付溢价。
Dwight 反对把自动化劳动力当作特殊的定价锚点。CFO 本来就希望降低人力成本,移除人工环节反而可能带来更大的降价压力。按产出定价很有吸引力,但企业可能正在“急着转向这种模式”;相比流行的席位费与人力成本之争,传统订阅模式可能仍有更多alpha。
10. 有边界的 AI 可以获得软件毛利率,并打开全新产业
Dwight 认为,投资者过度关注巨型实验室解决智能问题,以及它们的研发和资本开支何时产生价值。有边界的自动化应用在经济性上不同;他认为,几乎每家成功公司都在探索替代工作或“用更少资源做更多事”的工具,而这些机会往往远离公众对模型实验室的讨论。
编程体现了这种区别:Gaurav 将编程从穿孔卡片、汇编、C++一路追溯到 Python,随后称英语是“新的编程语言”。将意图转化为代码可以是一个有边界的问题,无需意识、梦想,或一个会自主决定创办公司的实体。
Captions 估计,解决其生成问题可能需要数亿美元,之后增量微调会便宜得多,推理成本也会持续下降。护城河在于更好的数据和持续的模型表现;一旦竞争对手复现基础层,竞争就会回到工作流、API、产品包装、B2B 销售和消费级软件执行力。
Dwight 认为,随着新一代 GPU 降低成本,成熟期毛利率可能非常高,但警告高毛利业务会吸引颠覆者;他以80%-90%毛利率的 CRM 行业为例,认为这正是新公司发起攻击的机会。“任务完成”也不是终点:Gaurav 认为,今天的模型只是社交网络、电影和电视、教育、配音、后期制作,以及一系列新基础模型的起点。
My guests today are Dwight Churchill and Gaurav Misra, co-founders of Captions, which uses AI to generate and edit talking videos and has grown to significant scale at remarkable speed. We explore a key distinction in AI: tackling bounded problems like video generation versus unbounded problems like general intelligence, and what this means for building sustainable businesses. We also explore their unique data flywheel, why video generation could reach Hollywood quality within 18 months, and why building advanced AI products doesn't require huge teams.
So guys, the topic on everyone's mind, I think, is this shift from AI as this incredible technology—that's decided. Everyone understands how amazing this is—to, "Okay, great, what are we going to do with it? And how can we build enduring, generational businesses with this technology at the core?"
You were very early in building a business that charged customers very early on using this technology. Maybe you can begin by riffing on the lessons that you've learned so far about building an AI business that are distinctive from a normal software business or something. Also get into some of the open questions that you yourselves have as you're trying to evolve your business model. I just think this is becoming the important question in the marketplace right now, and you're one of the earliest adopters, so you're the perfect people to answer.
1. Data Determines AI Winners
Getting into it, I think the first question behind the question that comes to mind is: What exactly did we achieve with this AI revolution? What is actually the difference? AI existed before, and it exists today. Obviously, there's something magical about what exists today.
When you get into it, you realize that it's really about the ability to train larger and larger models. That's a combination of having better hardware to do it and better machine-learning architectures. There are transformers, diffusion models, and all these new types of architectural unlocks that we've created. Then there are other techniques that we've created that allow us to train larger and larger models.
It turns out that the larger and larger we make these models, the more problems they can solve and the better they can be at solving problems like text generation, moving toward AGI, video generation, or media generation in general. When you realize that, what you get to is that what really matters is the data at the end of the day.
A lot of companies are scraping the Internet, and the Internet is also limited in some ways. There's only so much information on the Internet, even though that's growing every day. But beyond that, we're going to have to find sustainable sources of data that can continue to grow bigger and bigger models. I think that's going to be the fundamental question behind who actually ends up winning in a lot of these different areas where AI is excelling today.
For us, being on the video-generation and video-editing side, it comes down to video data, which is much heavier and much rarer to find. It's not as common as text or even audio, and it's potentially much more expensive to train on as well—much more limited in terms of being created in the world.
That tends to be a big challenge. One of the big things that we're thinking about is how we create a flywheel where we can ingest data on a continuous and growing basis, and that data can create bigger and bigger models for us and keep us at the forefront.
I also want to call out that there's a pretty fundamental difference between different types of AI companies that are out there. If you look at a lot of the text-generation companies, they're not solving text generation. We don't call it text generation. They're actually solving a totally different problem, which is intelligence.
Intelligence is an unsolved problem. No one's figured that out yet. Yes, we're achieving some levels of intelligence in these models, and there's a long way to go. It may not end at human intelligence. There are people in the world who are really smart, and there are people in the world who are not so smart. They both exist.
Clearly, there's a range of intelligence that's possible. There's not one value for whether you're intelligent or not. Is there a chance that there's an ability to go beyond the smartest human? It's possible. But that's a frontier that we've never reached, so it's solving this unsolved problem.
If you think about audio generation, video generation, music generation, or these types of things, it's a little less about solving an unbounded intelligence problem and a little more about rendering a solved problem. Video, for example—CGI exists. We can make fake things. We can make fake humans, fake scenery, and dragons. This is a solved problem. We know that there are solutions to these problems, and with AI, we're just making it easier to solve them—not just a little bit easier, but 100 times easier.
In the end, that means it's more accessible, there's a larger market, and more people can use these types of technologies. I think that's one of the fundamental differences. If you look at business models for artificial-intelligence companies that are really working on AGI, you have to think about this unbounded problem: We put in a bunch of capital, create a model, and then that model is beaten by the next model, becoming essentially useless and obsolete. Then there's the next model after that.
How long does this go on for? We don't know. It may go on forever. There may be no end to this intelligence race. Whereas if you look at the media-generation companies, they're actually creating an asset. There might be very soon a point where you say, "Wow, it's just really good. It's just perfect or close to perfect, and we've solved it." Then it's an asset.
After that, it's just a software company. The asset is really expensive to create, but once it exists, it generates value, and it doesn't lose value that easily. What is going to make those models better and better? I think it's going to be fine-tuning with more data, fine-tuning for specific use cases, and different types of things you want to generate—different types of visuals, whatever it might be.
There are use cases where it's going to be used in ads, movies, social media, or something else. There may be a point where you say, "Wow, this is pretty good. It's realistic." I think that's a pretty important thing we're thinking about right now: How do we bootstrap that data flywheel to reach that level?
What is it like to work with video data, where I imagine the terabytes or petabytes—however you measure it—of data that you have is insane? How do you think about something that might just get as good as it can get?
I love the point that if you give a Hollywood studio or Wētā enough money, they can literally create any visual that you can imagine. The friction between imagination and output is already gone. It's just really, really expensive. So what you're doing is making something cheaper. When do you think that could be achieved?
2. Video Models Reach Hollywood Quality
I think it's pretty soon, honestly. At the rate at which video models are growing, you probably remember seeing the Will Smith spaghetti thing. Everyone's seen this meme.
Right.
It went from really horrible to, "Wow, this is actually good." I think really, really good is probably around a year to a year and a half away.
I only say this because if you compare text models to video models, text models are already in the 400-billion-parameter range. People understand better how to scale large language model technology today because more money and more time have been put into it. Diffusion models are still in the tens of billions. It's still early, not even close to the text models.
As that grows, there's just no doubt it's going to get better and better. The experts know that this is all possible.
It’s just that very few companies in the world have the funding and the expertise to actually go after this, so it just takes some time. It’s not like some unsolved problem. People know what needs to be done; it’s just that we’re all getting there, we’re all moving toward it, and we’ll see those models getting better and better, especially on the video side.
I could easily see, within a year and a half or so, something coming pretty close to being indistinguishable from a real recording—maybe even sooner. That’s the worst case.
Yeah, I don’t think people are entirely able to grasp that yet. I think the way that influences how they do their work every day, the workflows that end up getting reinvented, and new paradigms around all of that are arguably part design problem and part product problem in general.
The timelines Gaurav is talking about are pretty close. People are experimenting today. It’s extremely early, and I think with company adoption and everything around that, we’re not far off at all from really reinventing a lot of how people end up doing their everyday work.
Patrick O’Shaughnessy
Can you describe the stages that you’ve gone through as a company? Maybe we’ll use the Tesla analogy. One of the beautiful things about their model is that the cars, by virtue of being driven, are gathering data all the time. The product itself naturally generates data exhaust.
I think you’ve had a somewhat similar story on the video side, and so you don’t need YouTube or some massive proprietary library of video to do what you’re doing. Can you take us back to day 1 of the business—what it was like to start, why you started there, and how it’s progressed since?
3. Captions Builds Its Data Flywheel
It has been a pretty interesting journey, and we’ve been through some interesting twists and turns. But I think if you connect the dots end to end, it’s interesting.
When we started the company, the first app that we made was Captions. We launched it. Why did we make it? The goal was to get content creators to create content on a video creation platform of some sort. That’s not easy. I was at Snap before this, and Snap had tried this many times. They launched apps, and video is kind of a commodity. Video editors are commodities. A lot of these companies are actually foreign, and that’s because we’re just trying to minimize costs at this point, and it’s really difficult to compete in.
Our thought was that the way we were going to crack this was to use AI to help create video somehow. That was going to be our differentiator. That’s why people were going to come to us.
We saw that there was a need around speech-to-text. It was a technology that, by the way, at that point was pretty good. In tech circles, people were like, “Of course, speech-to-text—we understand that is pretty good at this point.” But I think the average person actually didn’t understand how good the technology had gotten or how accurate it was with names, obscure terminology, and all kinds of things.
We built the first product where it was literally just, “Hey, put text on the videos.” By the way, this was built in 2 days on a weekend, really just band-aided together. We put it on the App Store and went to sleep. The next morning, it was at the top of the App Store. There was no explanation. We didn’t do anything to make that happen. Somebody saw it, posted it somewhere, and it blew up.
I woke up and texted Dwight. I said, “Hey, I think there are 600 videos per minute being created on the app.” That was kind of an instant success.
But even in those 2 days of work, we had already instrumented the app in such a way that we would be able to continue training better and better models so that we could deliver better value to the user. The idea was that the app is an AI app: people come in, they use the app, we use the data to make the model better, and we deliver even better experiences the next time the person comes back. That was done from day 1, literally. That was the original plan.
After the launch of the app, we added so many more features over time and expanded the offering so much more. Now we cover the entire space, from scriptwriting to recording to video editing and distribution, as well as how AI can transform each of these different areas. There are applications in all of them, and there’s data that can be collected across all of those that can improve those models.
That’s what makes our offering really unique, because all the other companies aren’t really thinking about the data-collection side. They’re just generating outputs, and that’s why they have to scrape the internet to make their models better. For us, it’s more about growing a user base so that the data can actually power better and better models.
A lot of that comes through video—video being funneled directly into video-generation models. That gives us a significant advantage. That’s potentially a possible way in which a future business model could be set up.
Actually, it’s kind of familiar. It seems to me similar to the Facebook or Google business models, where you have a mass-consumer, free product, basically, and the data is used to power essentially a B2B paid product.
If you think about the literal process of training, maybe you can explain to people who are curious how this actually works. You have raw video. A lot of it has voice in it, so you can start by translating that voice into text.
But let’s say you’re trying to train a model. I like how you guys referred to many people like Asora focusing on what we’ll call B-roll—background video or just a landscape video—and your focus has been on A-roll, like a human being in an iPhone-like video talking.
How do you train a model where the output is A-roll like that? Just imagine a portrait video of someone reading an ad read or something like that that’s indistinguishable from an actual live video taken on an iPhone. What is the literal training process? What is the target of the model as it’s training? How similar or different is this from just next-token prediction?
What’s the mental model for next-X prediction or something in a video? How do you think about the literal, actual training process of what’s happening?
4. Diffusion Models Start With Noise
It’s interesting to think about, because the models that we train are diffusion models. They actually work by starting from noise. It starts from literal noise, like the static you see on TV. At every step, based on the text that’s provided, it looks at the noise and tries to predict a layer of clarity in that noise.
It says, “A man wearing a blue shirt.” So it starts to draw a little bit of a man wearing a blue shirt out of the noise. Then, every pass it takes through the model, it discovers a little bit more of what a man wearing a blue shirt looks like. That’s how the text conditioning helps it decide how to reach the destination of what a man wearing a blue shirt looks like.
That’s how diffusion models work, which is slightly different from how a next-token prediction model like GPT works. A next-token model is, as you might think about it, just predicting the next word based on all the previous words that have been spoken, which are considered the context.
These models are different. We are still early in the diffusion-model training path. We’re still in the 10-billion, 20-billion, 30-billion range. Meta’s Movie Gen was, I believe, 30 billion parameters. People haven’t really scaled this up. We actually don’t know how big OpenAI’s Sora is. They didn’t, I think, release that information.
A lot of the work is going to go into scaling these things up. Video is obviously really heavy, and that’s what makes it different from text. It consumes a ton of space and a ton of processing. For us, even if we were to download all of our training videos, it would cost us $1 million. That’s a whole different regime from text. It brings different types of challenges to training these models.
What does that mean in terms of the strain on resources that video models will represent relative to text models? One of the big discussions in public and private markets is how big the GPU farms need to get. Are these video models, to get to that point of perfection, necessarily more consumptive of GPUs than text models would be?
What’s your 2 cents on this big question? Do we need to build nukes next to data centers to train the perfect Lord of the Rings model or something?
5. Video Compute Gets More Efficient
You never know. But honestly, I think what will save us on the video-model side is the fact that it’s an easier problem than the text problem. The text problem is intelligence, as we were talking about, and the video problem is more rendering.
We already know how much rendering costs. We already know it’s GPU-intensive. If you were to literally CGI-render a scene, it would spend some time on the GPU. There’s no doubt. Can we be more efficient than that? It’s possible. It may not be the most efficient today. Maybe there are better ways of doing it. Maybe AI will be cheaper and faster than regular rendering. I think if that’s the case, then that’s a good thing.
But I think we know that it shouldn’t be worse than that. We should be able to solve it with fewer resources than that, potentially, or at least the same. We generally understand where it’s going to fall.
It’s still early. Just like on the training side, we’re still scaling up these models, and it’s still, “Oh, it’s 10 billion parameters, 20 billion parameters,” whatever. On the inference side, similar learnings are happening simultaneously. We don’t need to do 100 steps of diffusion for inference—100 denoising steps—to reach a clear picture.
We can distill models and have them work with a few steps of diffusion now. I think we're definitely the most inefficient we'll ever be, and it's only going to get more and more efficient. It could be a factor of at least an order of magnitude, like 10× or something.
Can you talk about the felt experience of having a business—we won't quote how big the business is—but it's very big, and it's grown ridiculously fast? One of the things you hear that's a common idea now is that a new technology like this unlocks distribution. Distribution used to be really expensive. Late in the mature SaaS cycle, people had their tools, but when tools are just 10× better or more, 100× better, distribution for a time becomes really easy.
I think you've been a beneficiary of that unlocking of distribution. Just talk about what that is like. What is it like to see revenue and users and all this stuff scale at this pace? It seems like the revenue growth rates of some of these AI application companies are faster than anything we've ever seen, and I would just love you to riff on that a little bit—first, describe what it was like, but also just reflect on anything that it teaches us.
6. AI Unlocks New Markets
It's definitely the most exciting thing for anybody who's working on the engineering or product side. There's nothing more exciting than seeing the direct results of doing a thing: I did a thing, and the next day it caused an impact. People cared. There's just nothing more exciting than that, and I think we see that, which is great. That's why we've been able to build a great team and hire all this great talent, which has really set us up for success.
But I think maybe the most interesting part of it for me is how, as you're expanding the use case, you can almost see that it's actually growing the potential market, and that potential market has no competitors. As you expand the use case, you see, "We're doing ads now," or "We're doing higher-quality video even on that axis," and you see entire new areas of the market unlock where there's actually no competition. That's what causes the fast growth. It's nothing other than the fact that we're the only company that can do something for a period of time, and that will change.
I think that's why it's going to be interesting to see: as more and more use cases unlock, at some point all of it is going to be unlocked, and all of it is going to have competition. That'll be a different time. It might be years from now. I don't know when it'll be. But at least for now, what we're seeing is this ability to expand the use case.
We really think that the use case unlocked so far is somewhere in the 1% to 5% range. We've barely scratched the surface of what's possible. As that grows, we see these entire new markets unlocked: "Wow, this is a whole new set of people who can now do something actually useful with this." They're completely willing to pay. They're running to us; we don't even need to sell it, and we're the only option. It makes growth really fast. I think that's been probably the most exciting thing for me.
Can you level-set what the platform can do today—the major use cases? Everyone can imagine feeding it a video and getting a captioned video back. That's very simple. Can you lay out the other ones and give us a sense of their relative popularity? What is the revealed preference for how people want to use a platform like Captions?
When you think about us, we actually divide the product into 2 areas. There's the traditional video editing and video recording, which is just as you would expect: it's a video editor and video-recording software, and it's built for consumers and is completely free. The play for us is to provide a service to a large number of people—a kind of freemium business model—in a way that's already familiar to them. But the goal is to upsell them into the AI use cases.
You actually don't need to spend all this time video editing and recording; you can just generate it. On the flip side of that, we offer the AI suite, which is 2 products: AI Creator and AI Edit. These exactly mirror recording and editing. AI Creator literally just makes videos of people talking, whether that's you, an actor we've provided, or anybody you might choose whom you have the license to use.
We can make them say whatever you want and deliver whatever message you want, and we can even create people who don't exist. In between that, you get a bunch of optionality in how you want your message to be delivered. A lot of the use cases, like marketing and sales and these things, are very close to revenue. Then there's AI Edit.
Just a recorded video isn't exactly enough to create value. You want it to be edited in some way to tell the story that you want to tell. That's why we have AI Edit. Its purpose is to take a video in and edit it for you. You don't have to worry about key frames and animation curves and timelines and all these concepts. Video editing is not easy, and a lot of people avoid it because they just don't want to deal with this complexity.
Our thesis is that we have a foundation model that just does the editing for you, so you don't have to worry about editing anything. So that's the suite of products, basically: the traditional versus the AI. In the AI, we have AI Creator and AI Edit.
In something like Edit, am I prompting it by saying, "I want it to do this specific thing," and is it basically prompting?
That's where it will go in the future. Currently, it's more about style preferences that you provide to it, so it's in the early days of that. A lot of what will happen in the future is that, as we get more video-editing data from our traditional products, we're going to use that to give our foundation model the ability to be prompted with text for whatever you want to say.
You might say things like, "I don't like these images. We want different images with a better vibe," or, "Let's cut it down to 30 seconds; 45 is too long," or, "It sounds a little slow. We want to tell the story at a faster pace." These are general prompts, what you might actually say to an actual video editor, and the type of thing that someone who doesn't have detailed, intricate knowledge of video editing might say.
What is the relative breakdown of the tools people use—the free versus the paid, and the editor versus Creator? How does it shake out?
Today, a vast majority of our users are paid users. Between AI Creator and AI Edit, they're both about equally popular. Some people just use AI Edit. Some people just use AI Creator, depending on the use case. Then there's a bunch of people who use both one after the other.
Using both lets you basically get from absolutely nothing to a fully edited video with just a couple of words typed, which is a great first-time experience. Some people might want to just record their own video, or they might be editing on somebody else's behalf. They might take a real video and pass it to AI Edit: "Okay, I already have a video. Edit this for me."
On the AI Creator side, some people don't want the editing, or they have a very specific use case for what they're trying to do with it. They just want to figure it out on their own, so they just do the AI Creator part. A lot of the time, it's using their own likeness, so they can mass-produce videos of different types, but it can also be using one of our actors.
A lot of that is marketing content and things like that, things that go on social media, but also ads and anything that might be marketing-related. So those are the relative popularities. I would say they're about equal.
Does it feel like you're in an arms race right now with other companies?
To an extent, yes. I think the most interesting thing that I've seen is a lot of new companies popping up. All of them are trying to do the same thing. I'll give you an example: I was at Snap before this. Literally, 5 other people have left Snap and tried to start the exact same company.
"Yeah, it's working. We should be doing that thing. It makes sense." I don't blame anybody. I think it's great that they're doing it. But what I like about it, in a way, is that people are copying us. I think it's a great sign. It means we're doing the right things, and we avoid looking at other companies too much.
Our product strategy and what we build and what we do are really decided by our mission and vision and where we see the future being. It shouldn't be decided by what somebody else is doing, because they may not have a strategy at all. We don't know. Their strategy might just be looking at us.
A lot of times we'll look at competitors only to the extent of understanding, "Okay, this is what they're doing." What we really focus on is thinking about our North Star and where we see the future being, and whether we're building toward that future—not just from a technology perspective, but from a product perspective and a user-experience perspective.
I think that's the fun part. I think that's so much fun. When do we get a chance in history to actually invent the entire stack from the bottom to the top, all the way from the hardware level? There are bugs in the NVIDIA drivers. There are bugs at the hardware level. It's crazy, and we get a chance to literally invent the UX: How are people going to interact with these things?
I think people are not even thinking enough about this yet. They're just literally taking models and throwing them on a UI and being like, "Press button. Output." What if it were more interactive? What if you could see the steps of diffusion, or preview things in the middle of the diffusion process, and change things according to what you want it to generate? There's just so much that's still to be unlocked.
Every function—whether it's design, learning about how the technology works, or technology people learning about how the marketing is going to work—is going to get so much more evolved and integrated, and that's what we focus on.
I think the arms race is ensuring that we're always delivering way in front of what our customer even needs today. Whenever we're releasing something, it gets commercialized on day zero, immediately. We're not testing it with a bunch of people, seeing what they need, and seeing if we're actually solving anything.
No, we're building this for their work. We're incredibly ingrained in how they do their work, whether you're a large enterprise or all the way down to the free consumer. Ultimately, to Gaurav's point, by inventing those design patterns and the way someone can interact with these new models, we're literally paving the way for how people even think about doing their work, and that's the really exciting stuff. That is the arms race in my mind, but that's not necessarily against another company.
What are the trade-offs that you've had to choose one way or another as you build? Video's a big category. That could mean I get to make a Lord of the Rings-quality movie, or it could mean something much more provincial than that.
7. The Talking Video Advantage
We've actually niched down quite a bit on purpose because, as you said, video is huge. It's a massive market, and it's almost too many problems to solve. I don't think we'd solve all of these things if we tried to focus on everything. So our focus is very much on videos oriented around communication. These are talking videos, people saying stuff. A lot of it tends to be marketing, sales, and education. These are the big categories, or maybe communication is just the next step.
It's about generating those types of videos. It's about editing those types of videos. I think generating stock video is fine. I think that's a great thing to solve, but our goal isn't to create stock video. It's actually to create A-roll video, telling the actual story of whatever it is you're trying to convey. So not just bunnies jumping around on Mars, but more like telling a story, pitching a product, or whatever that might be—something really communicative and informative.
That's where we've seen a lot of our product-market fit. We're actually the only company training a foundation model to do this type of thing today, to generate A-roll. There are a couple of technical reasons why that's the case. There are other companies in the space, but they're not training foundation models. So we'll see how the space evolves in the future. I think it will actually tend more toward what we're doing.
What are the surprising, hard limitations of what the models can do today or might be able to do in a year? I'm imagining we're sitting at this table. There's a bunch of stuff on the table, my specific brand of water bottle or something. I want to be able to tell the thing to hold it this certain way. I want to be able to direct an object that's not the person, but that interacts with the person. Is something like that relatively straightforward?
Yeah, I think that will happen within 6 months, guaranteed. We'll probably start seeing the first versions of this coming out within months.
How does that work? Are you creating a 3D representation of this thing somehow? What are the steps that go into the ability to create something like that?
You have to find training videos where people are already interacting with objects—you drinking a can of Coke or whatever it might be—and then you have to be able to identify those objects and provide them as conditioning. For example, it might be text conditioning. If you can adequately describe this particular can of Coke in text, that might be enough. But it may also not be, right? A FIJI Water bottle has a very particular design. Unless the model has seen one before, it may not be able to precisely recreate it, and text might not be enough to describe what it looks like.
So you might imagine image conditioning: here's a picture of a FIJI Water bottle, and then text that says, “Man in blue shirt holding FIJI Water bottle.” Then it'll be able to figure out the rest from there because it's seen bottles in general and it'll understand what bottles look like. If it sees it from one angle, it can predict what it looks like from the other. So if you're rotating it around and moving it around, it'll guess what it probably looks like on the other sides, but it'll be pretty accurate because it can see the bottle from one angle.
You could imagine a world in which we provide multiple angles of the bottle just to make it a little bit more accurate. Maybe there's something on the other side that isn't visible in one image that you want to make sure is clear to the model. Those are the types of things that are just obvious. This will be the first of what's going to happen.
How do you think the value of these things will change over time as the cost and frictions to create them fall? Humans are really good at scarcity and assigning value to scarce things, and so a beautiful video that shows a product was valuable because it was costly to create, in some sense, probably.
How does the availability of perfect, high-fidelity, unbelievable-quality video at a moment's notice change the value of the video itself? I'm curious about other knock-on effects of what you're doing that you've thought about.
One comparison you can draw with this is the 2010s generally. It was a phase of design really taking off. Companies like Canva and Figma were created in this decade, and there were a lot of companies saying, “Make a website with a few clicks. It looks awesome. Great-designed websites just one click away.”
This wasn't AI. There was a huge movement to, if you want to sell something on the internet or have a business of any sort, create a great-designed website. If your website looks like it's from the 1990s, no one's going to buy anything from there.
I think that's cool again, though.
Yeah, it is cool now. Which is crazy, how fashion moves, right?
It all moves in cycles, right?
That's right. There's almost nobody that has a bad website anymore, but that doesn't mean that having a good website is not valuable. It's still valuable. If you don't have a good one, then you might still suffer today, even though it's essentially a commodity—everybody should have it.
Video is more what's taking off this decade. I think we'll see more and more people adopt it. It feels like there are a lot of people adopting it today, but I think it'll be even much larger than that because the portion of creators within the video ecosystems will grow. More people will be creating it and potentially even more people consuming it.
I actually think that the value of the video will not shrink exactly. High-quality video will be high-quality video, and it'll be a requirement if you want to market, sell, or whatever you're doing. But I do think that there are going to be other things about video that are going to become more valuable.
For example, if you think about likeness, if models can just generate likenesses of people that don't exist at a whim, and they look like great people you would want to represent your brand, you could even own a likeness as IP of your company—of a person that doesn't exist—and have them be the spokesperson of the company. That sounds awesome. That sounds great. But that means that the value of the likeness is just going to zero. The average likeness is not worth anything because anyone can make one out of nothing.
And what does that mean for the cost of likenesses in general or on the high end? I think it's going to be determined by who's known. A likeness that is actually known by people, trusted, and understood by thousands, hundreds of thousands, or millions of people is valuable now and much, much more valuable all of a sudden. And by the way, that person may not have existed either. Someone might create a completely fabricated person, post videos and stuff, and become famous.
Yeah. Little Michaela was way ahead of his time.
Exactly.
Shout-out to Trevor McFedries.
Way ahead of the times, yeah. So that doesn't sound crazy in that world. I think you can go crazy with this stuff.
What are the surprising limitations of these things? What would people be surprised that they have an especially hard time doing?
We've all seen video models struggle with people at the end of the day.
Fingers.
Yes. Fingers, arms.
Drinking.
Yeah. Olympics.
Spaghetti.
Yeah.
Spaghetti.
Yeah. I think we're taking the unique angle on this generally, which is that we are training specifically on people. Our data is all people, and we are specifically generating people.
We're also going to have conditioning—the ability to provide, for example, a skeleton. This is the exact animation I want to play out. This is the exact TikTok dance I want you to do, for example. It'll just make it happen.
And that actually makes the model much more likely to be able to learn what human anatomy looks like and what's normal and what's abnormal. People do have 6 fingers. It does happen. The model doesn't know that. Obviously, it's not that the training data is causing it, but it may not fully realize that if it hasn't been given enough training data showing hands in all kinds of configurations and doing all kinds of things.
So our goal is to solve that human-generation problem—just actors, essentially, in general.
The scarcity aspect, too, is that some of these are not new problems. The corollary around movies is that a Michael Bay film with a $250 million budget or something like that blows up half of L.A.—Transformers or something, I don't know. Tons of people go out and see it, a blockbuster film, but all of those people are paying $25 per ticket or something. The same thing happens for a low-budget film if they can get into the box office, but the ticket price is exactly the same.
I'm actually very excited about a world in which lower-budget filmmakers, or video creators in general, can create more and do more complex things without necessarily having the same budget constraints. That's a massive hurdle for film creators and creators in general, and I think it just elevates everyone. I think the craft may shift a little bit, but those high-budget films, as Gaurav mentioned, are technically generated. They're not synthetic, or they're not real. Some of those things are real, I realize, but it may even create more of a premium on some of those aspects.
What does it feel like in the competitive landscape to have established something so successful and important? I love the idea that companies pass some level of maturity when someone else tries to kill them for the first time. Have you had that experience yet? I'm really curious about the sharper, rougher-elbows part of building something so fast. Any experiences like that that are interesting?
Definitely. I think with all these types of things, we're always saying, "Let's go with our mission and not worry about what others are doing." But yes, a lot of people care about what we're doing. In fact, in terms of bigger companies, I think we're seeing an interesting evolution. We fall in an interesting spot where we semi-collaborate with a lot of social networks because we're beneficial to their growth. We create content, and all social networks need content. We have unwatermarked content, content that is original.
This was a big problem for Instagram. If you remember, when they launched Reels, everything had a TikTok watermark on it, and it was essentially recycled TikTok content. But we have a lot of that type of good content being generated on our platform—hundreds and hundreds of thousands a day, by the way—that's going to social media. So we end up being a valuable partner for a lot of social networks, and we've seen the social network landscape evolve in that sense.
A lot of VCs ask the question, "What if Facebook copies you? What if Google copies you?" I think what we're starting to see is that Google and Facebook are not the copying companies anymore. They're not copying anything; they're just doing their own thing. The copying company is actually TikTok, or ByteDance more generally. I don't know how this shift exactly happened.
Facebook suddenly became the good guys. I think Mark Zuckerberg is a hero now for putting all these models out and making all this open-source stuff. Suddenly, his vibe has completely shifted. And then I think TikTok has become essentially what Facebook was: capture, kill, and destroy everything that exists in every market that exists. Don't collaborate with anybody. I think it'll be interesting to see how that plays out. Obviously, there are many talks happening about a ban and all this kind of thing. We'll see where all that goes. But their leadership is very well aware of our existence, and they have tried many, many times to kill us. To their credit, they were the first to be aware of our existence.
What does that look like, them trying to kill you? Literally just copying the product?
Blatant copying. They literally went to the extent of copying our App Store description and our website exactly, putting that in their press release word for word, copying our brand colors—exact, precise brand colors—pretending to be us. It was beyond anything you would imagine, and it's just crazy that a company of that size would even try these types of tactics. But at the end of the day, the software that they create is just very mediocre, and it works because they have great distribution through TikTok. I think we win because we just have a better product.
It seems like, in the early days of all these models getting built, research talent was one of the most important scarce resources, in extremely short supply. Can you talk through what you've learned about that and how it's changed? Is it still a handful of people—you really need a couple of them to be able to build the cutting-edge thing? What is the role of extreme research talent in building the models that fuel all this great product?
The talent side is still evolving. I don't think it's completely solved, for what it's worth. As the use cases are growing, as more and more people are realizing what's possible, and as more companies are getting started trying to solve similar problems, there's only going to be more and more of a shortage of talent. Talent isn't created overnight, right? It takes years and years of experience before someone can be considered experienced in an area, and I think we will still see continued pressure on the talent side, especially for building generative models, foundation models, and things like that.
Obviously, the more VC dollars that are poured into this area, the more of an effect that'll have. But I do think, interestingly, it doesn't take an army to build this type of stuff. It takes a few good people, and it might take an army to scale it and really make it big. But delivering world-class results can be done with maybe a dozen people or less, and you can beat everybody in the world with a team that small if you have the right ingredients in place.
A lot of the challenge just becomes finding those people. You can't get it wrong. You want a very specific set of people and a specific skill set. This is all new, so very few people have experience in it. It's all cutting-edge. There are discoveries and inventions happening every day, every week. So the more you care about it, the more you want to find people really close to the cutting edge who really know what's happening today and what the small wins and techniques are that will give us the edge over everybody else. It still is a challenge.
What our duty ends up being, then, is that we have to give them all the resources and everything they need to be able to do their work. There are a lot of these folks in AI labs who can't release anything they're working on, for better or worse, at least in that place's opinion. When you're able to bring some of the ingredients that we have, whether it be compute, data, or the environment, it ends up not being that complicated in terms of recruiting, negotiation, and stuff.
How do you think these products will price over time? This is always the weird question with software, where the marginal cost of delivering it is nothing. A lot of people have been talking about, let's look at Accenture's market cap or something like that. It's a $250 billion company basically selling labor—very expensive labor, important labor.
Do you think that AI applications will take labor budgets and be priced like heavily discounted labor because that's what they're doing and replacing, or is it just going to end up pricing like all software does? We've run these playbooks for 20 years, and we kind of know how to do it. What's your sense of how people should think about pricing AI software applications and what its equilibrium state will be in the future?
8. AI Business Models Take Shape
I don't know if we completely understand it yet. I think it's almost too early to tell because we aren't completely able to replace labor in all its different aspects, so we don't know what people are going to be willing to pay for it yet. In the use-case graph, we're at 4%, 3%, 4%, 5%—whatever, something in that range. It's just early, and we aren't able to fully replace certain workflows or very operationally heavy processes that might exist in companies. We will get there. Slowly and steadily, we're moving toward that, and I think we'll see what people will be willing to pay for it. So I think we'll figure it out.
One of the big questions there is how that splits between consumer and B2B. I think consumer pricing is evolving pretty clearly. We're starting to see what that looks like. It seems like it's coming down to consumer subscriptions, and it also seems like people are willing to pay a little bit more than they would have otherwise.
Traditionally, for video-related apps on the App Store, web apps, Android—whatever—the standard price is somewhere in the $7.99-to-$12.99 range. That's just considered normal, and there's a freemium business model to it. What we've seen that's different is that, for us, for example, we've been completely premium for a long time. There's no free product. You cannot even use it once for free, and that worked just fine. People were like, "Okay, whatever. Here's money. Let's move on."
That wouldn't have worked in an older world. Without the newer technologies and everything, people would've been like, "Yeah, I'm not paying for this. I'm moving on to the next one," right? I think the other thing we're starting to see is that we can charge $25 a month. Yes, we can. People are paying that. People are clearly willing to pay much higher prices. If you look at a lot of different AI companies out there, the video-generation companies and so on, they're going across this range, too, and people are paying all these prices.
People are paying up to, like, $2,000 a month for a consumer subscription. I think there's a lot more ability to go higher on subscription pricing than there was previously. Now, that might change. I think a big factor might be that there's just not enough competition still. There might still be only 1 or 2 models in the world that are of that quality, and people care about that quality, so you really don't have a lot of choice.
Maybe if there's a ton of these Tesla Model floating around, the price comes down in the future. So that's what we're seeing on the consumer side. On the B2B side, I think that's where we'll figure out a lot of it. I think some of the big things that need to be solved there are: Will businesses buy models that are trained on unlicensed data? That's an open question, and, yeah, they are to an extent. We'll see how all that plays out.
We're planning to go much more on the fully licensed side. That's going to be one of our main differentiators because we are uniquely positioned for that. We actually collect data at massive scale, so we can train fully licensed models. My feeling is that toward the endgame—not today, but as this area gets very saturated, maybe many years from now—having fully licensed models will factor in because you'll be able to win on that very easily in a competitive deal, and people will care about those types of things. People might even be willing to pay more for that type of guarantee, or just for the reputation that it's licensed.
Besides that, it really just comes down to how much of the use case we'll be able to cover, and that's the big question. Okay, we're at 5% today, but is the limit 100%? Is it 75%? Is it 50%? Where does this stop? My guess is we can go all the way to 100%, or at least very close to it, just because it's a solved problem. We know that this is solvable.
I think if we can get there, a lot is going to change about how video workflows work in the world.
And in the pricing around labor, a hot topic right now is seat licensing versus labor, or aligning to labor costs or something. I think people are maybe rushing into the labor argument, but it actually has a very similar path, or has had a very similar path. It turns out the CFO would like that number to go down. It's not some special number or something. If you remove the human element from it, my guess is that probably only puts more downward pressure on it. It's like, "Great, we can do more with less," and it's like, "Perfect."
Whatever the software is doing, whether it's writing code or it's an automated SDR, there is downward pressure against those things. I think people are getting a little excited and are running toward that. Don't get me wrong, pricing toward output and such is pretty cool. I'm sure there is something there, and there's some equilibrium we'll find. But I do think people are rushing to it maybe a little faster than they should be, in that there actually might be more continued alpha in the typical subscription. Sure, maybe that's not the Salesforce seat, the classic comparable, but there's some market exploration that needs to happen, and we probably haven't fully seen that yet.
I think the entire world of investors—VCs, growth equity investors, public investors—pretty much every single one of them is trying to figure out how to think about AI and its implications for their companies, prospective companies, equity valuations, and all the normal important questions. How would you advise them from the other side of the table? You've talked, I'm sure, to a lot of the great investors, and you have several of them that have invested in your company. What do you think investors understand about AI well? What do they feel like they're categorically not understanding in as much detail as you do from the builder side? Give us your lay of the land of how you think investors are doing. Give them a grade or something for their understanding.
Yeah, if I have to grade it, maybe from the public equity side, there are a lot of smart people out there, so I'm not going to give them too hard a time. I don't think it's fully being appreciated how much this is changing. Everyone's saying that and all, but I think there's continued talk of, "Oh, there's all this R&D spend or capital expenditure," and it's like, "Where is the value and stuff?"
I think there's just so much attention on the large AI labs that are effectively what Gaurav was talking about before: solving intelligence. That's a very, very different mission from the company that's maybe creating the automated software developer. Those are 2 very different worlds. Paying more attention to things outside of that is pretty important to really understand how this is changing things inside their companies.
I think it would be hard to find a successful company today that hasn't, or isn't, exploring an AI tool of some sort to either completely replace an activity inside the company or, quote unquote, "do more with less" in another capacity. I think that's true of every function of essentially every successful company today, and that's where you're seeing a lot of the adoption.
Even discussing the foundation model versus some of these companies that are just fine-tuning an open-source model, there's a ton of alpha in getting these tools inside of their company. If you're talking to someone who's asking, "How do we do this and roll it out as a larger enterprise?" I think there are already examples. There are massive enterprises where you go in—someone was telling me the other day that L'Oréal, the beauty company, has an internal GPT, basically, an internal LLM of some sort. Any employee can ask any question.
I don't know how much that's really being baked into their thought process. I think there's just so much attention toward these particular AI labs and the way that they're running their businesses, which is extremely different from some other companies, particularly if they have the backing of Microsoft or something. It's inherently just being driven differently. If you were to go into those companies, I think your viewpoint would potentially change, and that could definitely inform a better understanding of how this is actually going to change work.
I love this framing of the unbounded-problem nature of intelligence versus the bounded-problem nature of video or some of these other things. It's kind of a fascinating bifurcation.
I actually think that applies to the tech side, too. Even on text, we've already created essentially what is a tool for intelligence. It's like intelligence in a box—intelligence you can just apply to something to solve a bounded problem.
Whether that's coding, for example, think of it in the coding context. I think, as Dwight was saying, engineers are smart people. Does that mean we need AGI to solve coding? Not necessarily, because what it's really doing is just translating. Think of how computers have evolved over time. We used to literally do the punch-card thing. Then we were writing assembly language. Who knows that anymore? Then we were doing—
I think Aaron does.
C++, right? Exactly.
Just you.
Yeah.
Yeah.
Then we were writing C++, and then there were these higher-level languages like Python. We're coming to the modern era.
Scott from Cognition was the guest today, so—
Oh, perfect. Yeah.
He's built the next layer.
Perfect. Yeah. And then we're just saying, "Hey, the new programming language is English." That's not a crazy jump. It's actually a very bounded problem. It's a problem of inventing a new programming language, essentially—a programming language that's even more understandable to people because they already know it. It's a language that they already know.
Intelligence is a special case.
Exactly. The general intelligence idea of, "Oh, we're creating consciousness. It's a thing that's going to exist, go around, do things, have its own thoughts, and have its own dreams and hopes and stuff, and maybe it'll start a company at some point"—that's a whole different mission from solving intelligence in a box, which essentially already exists and is getting better and better.
I'd love to just extend the analogy one step further to the business model. Most of the commentary on AI businesses has been focused, again, on foundation model companies that have had 2 problems: huge CapEx outlays to train the models and then huge inference bills, so often, early on, really negative gross margins just to service their $20-a-month subscription product.
Inference has fallen 100× in cost in the last 18 months or something crazy. These costs are going down, but those were the 2 criticisms of the business model: "Oh my God, this unbounded race. I've got to spend 10× every time to build the next thing. When am I ever going to make some money?"
It seems like this other category of more bounded problems has pretty normal, great business models. Is that right? Is your sense that you guys are just going to have really high gross margins like a normal software company? You have to spend money training your foundation models, but it's not $10 billion. Walk me through the business-model expectations, margins, CapEx, things like that, and what the J-curve looks like in these businesses. Educate us a little bit on this second category.
The way we think about it for our business specifically is that there is a bounded cost that actually solves this problem.
That bounded cost is probably in the hundreds of millions of dollars, but it actually gets us to a solution. It gets us to something that is reasonably good at generating anything that a CGI studio might be able to do, and that is the level that we need to be at.
Now, will that evolve? Yes, it will need to be fine-tuned, but fine-tuning is generally cheap. It's not even close to as expensive as training a foundation model from scratch. New data will come in, which we already have a flywheel we're building for, and it's going to be massive amounts of data. We're going to be continuously training the model and making it aware of what's happening today and what things people might want to generate today.
But that's just incremental fine-tuning, and it's going to be a low cost underlying the business. On top of that, inference costs are going down, so I think it's going to start looking more and more like a traditional software business. Initially, with these types of models existing, whoever truly solves this problem will have a moat for a while, as long as they are ahead. For us, we're also trying to build that data moat simultaneously so that we are permanently ahead.
Once enough data is out there, enough people have raised enough money and have tried the exact same playbook and built these models—and this could be many, many years in the future—it's going to become a software race: building the workflows and all the traditional stuff that we know about. Pricing and packaging, all this stuff is going to become really important. We've seen it all. People are going to do APIs, B2B, consumer, all this stuff. There are going to be all these use cases.
I think that's where the real competition will happen, and there are going to be winners. Our theory and strategy on this is that the winners are going to be really determined by who has the best model that's consistently outperforming everybody else. All that comes down to data acquisition, the flywheel, essentially, and the ability to constantly improve the model.
I do think this won't be the end, though. I think new problems will get unlocked, and we already have line of sight into what those other problems look like. Those problems will have their own foundation models and their own data to be collected. Essentially, you could imagine a series of foundation models solving a family of problems across a whole set of workflows that are broad across video and maybe even other types of media—different types of use cases like film, television, whatever you want, basically. Maybe it's dubbing, maybe it's—
Hands.
Yeah, hands, post-production, lots of different possible use cases. As always, that will happen. There's no doubt about that. You can see that these models will reach a point of maturity.
On what this ends up looking like at the mature-business end, call it some threshold, I genuinely believe that these can look like very high-margin businesses. There's the deflationary behavior of GPUs and compute in general. It's incredibly early—we're talking about the video's latest chip, et cetera. You're already seeing the cost come down on H100s as the H200 architecture and related technology are being rolled out.
Throughout history, these prices have never gone the other way. It's highly deflationary as each next generation is rolled out because, ultimately, that's their business model. They make them more efficient, they make them more powerful, whatever it is. I think, generally speaking, that is 100% guaranteed, at least from my perspective.
The interesting thing, though, is that when you're earlier-stage, and companies are earlier-stage—just talking about startups in general—higher-margin businesses actually sound to me like perfect attack vectors for another entrepreneur. I think you should be very wary of companies that are operating at really high margins at that particular earlier stage.
For these types of businesses, there are a ton of margin-expansion opportunities. That also goes for later-stage companies, though. I think you're seeing it right now with CRM companies operating at 80% or 90% margins. It's great, typical SaaS kind of stuff. They seem like great opportunities for companies to go right after and reinvent some of the things those companies do. Those companies don't have the same pricing power that they had 15 or 20 years ago.
At the same time, the great ones are thinking about that right now and reinventing themselves. As we discuss some of these business-model changes, it does feel like a bit of shifting ground.
If you think about the future now, what is on the other side of the mission-accomplished banner? You just did all video, and you can create anything you can imagine in CGI with a $100 million budget. Now you can do it in Captions. Then what? What do you think you would do then?
If we actually achieve that within a reasonable timeframe, I think that would be just the beginning, because you could go so much beyond that. These industries are massive. You could imagine a social network based on something like this. You could imagine film and television being dominated by these types of technologies. You could imagine education being completely transformed.
The list is essentially endless. This would be the starting point of a potential complete transformation across multiple industries. Today, we're really excited about accomplishing this particular mission, but I think the possibilities beyond that are practically endless.
Any major lessons from your time at Snap, which strikes me as a very unique culture and an extremely product-centric, good place to train for product, maybe? What lessons do you take from your time there, and what lessons do you leave behind?
9. Snap Shaped Their Product Thinking
Snap had, as any company does, a lot of good things and some bad things. The great things that I was able to get from Snap were the ability to work with a lot of great people.
I think Snap was in a tough spot in many ways. They were in one of the most competitive businesses you can exist in, monopolistic by nature, where it's really difficult to get something started and very easy to get killed, and only the biggest one actually wins and survives. In that arena, they were able to make a place for themselves mainly because of innovation.
This comes down to the CEO. He was able to, out of all the random noise, see something and understand, "Yes, this will work, and nobody else will see it, but I know why it will work." At the core of it, he had an understanding of the product and the customer in a way that nobody else did, and nobody even came close to it.
There were many moments in the company's history where he was like, "We're going to do this," and everybody was like, "No, we shouldn't do that. This is a bad idea." He'd be like, "I don't care. We're doing it." We did it, and it was the best thing we ever did. That's the level to which his intuition was there.
Snap was famous for constantly innovating. Stories came out of Snap, and the old maps, location-sharing product and idea came out of there. There were so many things that we were innovating on.
I think they missed a little bit on the public TikTok thing, but that actually was something that didn't fit into their core vision and strategy, basically because they're a private-sharing platform. Their whole purpose was low abuse. People don't feel like they can even reshare posts, because that's a way to embarrass somebody by reposting that thing to other people who aren't supposed to see it.
Everything was designed around feeling good, having fun, and sharing with friends, which was really everything that people cared about at that time. I think they missed the TikTok thing because it was the exact opposite of that. It was actually sharing with everybody.
Interestingly, it created similar dynamics, where sharing with everybody actually made you feel more private because there were so many people that the people you know would never see it. Somebody else would see it, but that's the detail.
On the downsides of Snap, one of the interesting things—and one of the learnings there—is that product-market fit often doesn't have a lot to do with what people are doing day to day within the company. Once it exists, it can stay there despite the actions of the people.
What ends up happening sometimes in bad cases is that people think the wrong actions they're taking are actually the contrarian right view because the company is growing, so of course whatever they did was the right thing. But the company is actually growing despite the wrong thing that was going on at the time.
It's difficult to tell what is actually causing the company to grow, what's the good thing, and what's the bad thing. A lot of people walk away from these types of high-product-market-fit companies thinking that all the things they did were good things and there were no bad things because the company grew. But the reality is that the company was growing despite those actions.
Identifying those things was a skill that I had to really work on building: understanding how we can truly measure what we're launching and what we're building, and understand what's a good thing and what's a bad thing.
What I'm really grateful for from that time is the ability to work with the CEO there. He really brought me into the circle. He had built a great design team, and a lot of the decision-making was driven through the design team.
It was a small set of people, around 10 to 12 people on that team, even when the company had many thousands of people overall, post-IPO. So being a part of that team and learning from the great people on that team, I evolved my design career through this process. I think his ability to identify that this was a person who would fit in well and would be able to learn and figure these things out—props to him. He was definitely doing something right.
The closing question I ask everyone on this show—it’s fun to get to do this twice today—is: What is the kindest thing that anyone’s ever done for you?
It’s hard for me not to say the kindest thing is probably my wife and us starting this company. We were already married. We had our first kid. It’s pretty hard not to call it that. It could have obviously not gone that way. I decided not to start the company, not to do a bunch of this stuff, and she enabled me to take more risk.
Now I can’t use that answer.
Yeah, exactly.
Yeah, I mean, I think—
It’s a little unfair.
Besides that, likewise, for me, if I were to give you a different answer, I think it would be my parents making sure that I was born in the US. Literally, because I was only here for the first couple of years of my life. I was born while my dad was doing his PhD at Northeastern. He was studying economics, so he was there for four or five years, and then I was born in the middle of that. They moved back to India after that, but I had US citizenship. Without that, I’d still be in India.
Simple and powerful.
Yeah.
Guys, thank you so much for your time.
Thank you.
Thanks.