No Priors 第143期|与 ElevenLabs 联合创始人 Mati Staniszewski 对谈
ElevenLabs 以一款拥有500万 MAU的创作者产品叠加一项营收接近一半的企业级 agents 业务,在350名员工的规模下实现了3亿美元 ARR。 自助订阅和创作者业务仍约占收入的50%;数千家企业客户覆盖“从财富500强到增长最快的一批 AI 初创公司”。经营层面的信号是业务足够宽,但没有放弃最初的研究优势。
创始切口来自波兰糟糕的配音体验:所有角色都由同一个平板旁白配音,但公司的判断已扩展为,声音才是计算最自然的交互界面。 ElevenLabs 的目标,是让“原声、原本的情绪、原本的语调”跨越语言传递,这一点已在 Lex 对 Narendra Modi 的采访中得到展示。Mati 更大的判断是,键盘和屏幕交互“感觉已经坏掉了”,因此机会远大于传统配音市场的支出规模。
Mati 的运营模式是研究先行、产品并行,再围绕具体问题组建小型跨职能实验室。 一个约5人的声音实验室先解决了接近真人的叙述问题,随后 ElevenLabs 扩展到有声书和配音;之后的 agent 实验室则把语音转文字、LLMs、文字转语音、集成、测试和监控组合起来。产品团队能看到研究路线图,只有在相关突破看起来还要超过3个月时,才通常会自行开发替代方案。
客户支持是最先大规模落地的 agent 类别,但声音交互正从被动处理工单,走向发现、结账、教育、媒体和政府服务。 Meesho 的 agent 已从退款和物流追踪推进到产品引导,并可能进一步覆盖结账;MasterClass 让学习者与 Chris Voss 练习谈判;乌克兰正在推进 Mati 所称的“第一个 agentic government”。这使商业价值从节省人力成本,扩展到创造收入和全新的体验。
声音质量不是一个单一标量基准,因为即使底层模型质量不同,用户偏好也会随声音身份、语言、受众和表达方式改变。 一名客户要求声音“尽可能机器人化”;而一项覆盖日本和韩国的部署,则希望面向年轻来电者使用兴奋的声音,面向老年用户使用平静、缓慢的声音。因此,ElevenLabs 会提供声音指导,并预计最终针对每个人、每种场景动态选择声音。
ElevenLabs 假设基础模型终将商品化,并将研究领先视为只有6到12个月的优势。 Mati 说,“研究只是起跑优势”;长期价值必须通过分发、声音、集成、工作流,以及连接模型与业务逻辑的产品层不断累积。开源模型的叙述质量已经很强,但可控性和实时编排仍有差异化;一个可能只有50-100名顶尖音频研究员的集中人才池,仍能带来架构级跃迁。Mati 估计其中约10人可能在 ElevenLabs。
近期路线图瞄准可控的多模态创作,以及更低延迟、更具情绪感知能力的 agents,同时保留级联系统服务可靠的企业工作负载。 据称,Scribe v2 在 FLEURS 的前30种语言上实现了低于150毫秒的延迟和93.5%的准确率;融合式 speech-to-speech 可能更富表现力,但也可能引入幻觉,且可见性更低。Sarah 估计,接近真人的对话交互至少还需要1年,但可能在1年内到来;实时配音或翻译则可能在2年内实现。个性化教育——“随时拥有一名专属老师”——是最大的尚未兑现的应用。
1. 糟糕的波兰配音,揭示了一个全球性界面市场
Sarah Guo 先从规模切入:ElevenLabs 拥有350名员工、3亿美元 ARR、超过500万月活创作者用户,以及数千家企业客户。收入约为“五五开”,一半来自自助式创作者订阅,另一半来自日益以 agents 为核心的企业业务。
她当初对投资的质疑值得保留:ElevenLabs、Midjourney、Suno 和 Hunyuan 这类创作工具,看起来可能只服务于一小部分人;传统配音市场本身也并不算大。Mati 的信念来自波兰:外国电影经常让所有男性和女性角色都使用同一个平板旁白,“体验非常糟糕”。
他们提出的替代方案不是简单翻译台词,而是保留表演者本身:“让原声、原本的情绪、原本的语调跨越语言传递。” ElevenLabs 已在 Lex 对 Narendra Modi 的采访中展示这一理念,但 Mati 进一步把它延伸到整个计算领域:人类仍把时间花在键盘和屏幕前,尽管语音才是“最自然的界面”。
2. 以问题为中心的实验室,把基础研究接入产品
ElevenLabs 最初尝试优化现有模型,用于叙述和配音,但输出过于机器人化,无法让人真正享受。Mati 将关键突破归功于联合创始人——两人相识已有15年——以及早期研究团队,他们做出了产品所需要的基础改进。
最初的“声音实验室”约有5人,由研究员、工程师和运营人员围绕一个共同使命协作。研究先行,随后叠加一个简单可用的产品层,再逐步扩展出有声书、电影旁白和配音的完整工作流。
接近真人的输出实现后,“agent 实验室”开始解决按需获取知识的问题,负责编排语音转文字、LLMs 和文字转语音。产品要求很快从延迟和准确率,扩展到旧系统集成、函数调用、生产部署、测试、监控和评估。
一些实验室是由明确需求推动的:客户希望在生成语音旁边加入音乐,于是公司开发了完全授权的音乐模型,最终又结合合作伙伴的图像和视频模型,形成更广泛的创作套件。配音则更多来自判断和信念;Mati 预计,跨越语言障碍进行实时沟通的“完整 Babel fish 理念”将成为一个巨大的市场。
3. 声音质量无法被压缩成一张排行榜
Sarah 指出了买方难题:多数企业决策者不是机器学习科学家,即使研究人员,也没有覆盖所有音频场景的基准。客户可能认得出一个令人信服的声音克隆,但未必知道如何选择或评估最佳声音。
ElevenLabs 的应对方式,是配备专门的声音专家,把品牌和受众要求转化为声音选择,并提供包括 Sir Michael Caine 在内的标志性人才资源。需求差异极大:一家欧洲大型企业要求声音“尽可能机器人化”;一项覆盖日本和韩国的部署,则为年轻来电者匹配兴奋的声音,为年长客户匹配平静、较慢的声音。
Mati 认为,更棘手的问题在于,替换声音可能实质性改变用户对模型的比较,即使技术质量本身存在差异。标准化标注更擅长记录“说了什么”,却难以记录“怎么说”——情绪、口音和表达方式——因此 ElevenLabs 不得不自行构建这套定性音频标注能力。
4. Agents 正从客服走向交易与体验
客户支持的规模化速度最快,但 Mati 看到的变化,是从“我遇到了一个问题”转向覆盖客户全流程的主动协助。在印度电商公司 Meesho,agent 已从退款和包裹追踪扩展到产品发现、礼物推荐、导航,并可能进一步覆盖结账;他还提到 Square,说明语音交互正从下单走向发现。
互动媒体让静态知识产权变得可以对话。ElevenLabs 曾与 Epic Games 合作,将 Darth Vader 放入 Fortnite,让数百万玩家实时与这一角色互动;Mati 预计,这一模式还会扩展到书籍、游戏和其他故事世界。
教育是他最坚定的应用判断。Chess.com 可以让 Hikaru Nakamura、Magnus Carlsen 或 Botez sisters 成为老师;MasterClass 则让学习者直接致电 Chris Voss,练习一场谈判,而不是单纯观看他的课程。模型由此把内容从顺序播放,变成了互动练习。
在乌克兰,ElevenLabs 正与数字化转型部合作,推进 Mati 所称的“第一个 agentic government”。拟议系统将整合围绕福利、就业和办事流程的公民支持、主动公共信息服务,以及个人辅导;中央数字化转型职能则与各部委的工程负责人配合。
5. 平台广度,而非单点方案,是 ElevenLabs 的竞争答案
Sarah 梳理了企业客户的选择:通过 Palantir 或大型咨询公司部署,使用 ElevenLabs 或 OpenAI 这样的平台,或者购买 Sierra 这样的场景专家方案。Mati 坦率承认,对于一个孤立问题,ElevenLabs “很可能”不是最佳选择。
这一平台更适合那些要在客户支持、内部培训、销售和新客户体验中部署对话式界面的组织。客户可以只使用部分组件,与现有内部投入组合,也可以引入 ElevenLabs 的工程师;国际化声音、语言和集成能力则是另一项明确优势。
面对 Google 和 OpenAI,Mati 认为音频既需要专注的研究,也需要满足客户实际需求的产品层——包括那些不那么光鲜但不可或缺的工作。他声称公司在文字转语音、语音转文字和编排方面处于基准领先位置,并补充说,音频对单纯规模的依赖小于对架构突破的依赖:“人数并不重要”,真正重要的是顶尖人才。
他的让步很明确:“从长期看,模型会商品化”,时间可能是2年、3年或4年。开箱即用的叙述质量已经在开源和商业模型之间趋同;可控性仍然更难。长期路径将是研究、产品,然后是由分发、声音、集成和工作流构成的生态系统。
6. 研究争取时间,界面转向导师与机器人
Sarah 提出了一个不太舒服的判断:技术优势可能持续1年,也可能持续10年,但不可能无限期守住。Mati 表示同意:研究能带来“6到12个月的优势”,让客户更早获得某项能力,同时给 ElevenLabs 时间构建周边产品。团队与公司内部共享的研究项目协作,并以3个月作为“自行开发还是等待”的大致分界线。
当前工作覆盖可控的文字转语音、近100种语言的语音转文字、完全授权的音乐模型,以及音频与视觉模型的组合。在 agents 方面,据称 Scribe v2 在 FLEURS 前30种语言上实现了低于150毫秒的延迟和93.5%的准确率;即将推出的编排机制则应能降低端到端延迟,并纳入情绪上下文。
未来1年,面对可靠性要求较高的企业部署,Mati 仍选择语音转文字、LLM、文字转语音的级联系统,因为每一步都可检查,也能调用工具。融合式 speech-to-speech 可能更富表现力,但伴随幻觉风险。Sarah——而不是 Mati——估计,类似他们这场对话的交互仍至少需要1年,但可能在1年内实现;她预计实时配音或翻译将在2年内到来。
Mati 预计陪伴型产品会变得重要,但他个人更偏好“Jarvis”式超级助手,而非社交替代品:它理解他的上下文,能拉开百叶窗、播报天气和播放音乐。他最确信的未来仍是“随时拥有一名专属老师”,同时也明确保留不受技术打扰的人际互动;声音将成为“agents 十年”以及随后“机器人十年”的关键界面。
Hi listeners, welcome back to No Priors. Today I'm here with Mati Staniszewski, the co-founder and CEO of ElevenLabs, which was founded to change the way we interact with each other and with computers with voice. Over three short years, they've skyrocketed to more than 300 million in run rate. Mati and I talk about the future of voice, education, customer experience, and the other applications of this voice, as well as how to build a multi-segment company from self-serve to enterprise and combine research and product. Welcome, Mati.
Thanks for having me.
Thank you for doing this at 7 in the morning. It's great we finally got to do this together. I think a lot of our listeners will have used or played with ElevenLabs at some point, but for everybody else, can you reintroduce the company?
Definitely. At ElevenLabs, we're solving how humans and technology interact and how you can create seamlessly with that technology. In practice, we build foundational audio models: models that help you create speech that sounds human, understand speech in a much better way, or orchestrate all those components to make it interactive. Then we build products on top of those foundational models.
We have our creative product, which is a platform to help you with narrations for audiobooks, voiceovers for ads or movies, and dubbing those movies into other languages. Our agent platform is an offering to help you elevate customer experience, build agents for personal AI and education, and create new forms of immersive media. It's all under the umbrella of that mission: solving how we can interact with technology on our terms, in a better way.
You started the company in 2022.
That's right.
You've had amazing, rocket-ship growth since then. I'm sure it has felt different at different times, and I want to ask you about that. Can you give us a sense of the scale of the company today?
We've grown to 350 people globally. We started in Europe as a remote company, and we're still remote-first, but we have hubs around the world, with London being the biggest, New York being the second-biggest, and offices in Warsaw, San Francisco, Tokyo, and Brazil.
We're at $300 million in ARR, roughly 50/50 between self-serve—which is a lot of subscriptions and creators using our creative platform—and the enterprise side, which is approaching 50% and uses our agents platform. That's on the classic, sales-led side. We serve more than 5 million monthly active users on the creative side, and on the enterprise side, we have a few thousand customers, from Fortune 500 companies to some of the fastest-growing AI startups.
I think this is such an interesting company because it is very unintuitive to many people, and investors in particular. We were both there in 2022, and there's a class of companies that enable creation in some way. I would put ElevenLabs, Midjourney, Suno, and Hunyuan in this category. There's an overall sense of, “Who really wants to do this?”
What was your initial read of how many people would want to make voices? What made you believe it was going to be much broader than, for example, dubbing? Dubbing isn't a huge market.
It's very tricky to do both the product and the research. I'm in a lucky position because my co-founder and I have known each other for 15 years. I think he's the smartest person I know, and he's been able to create a lot of that research work to build the foundation and then elevate the experience.
Both of us are from Poland originally, and the original insight came from Poland. It's a peculiar thing, but if you watch a foreign movie in Polish, all the voices—whether it's a male voice or a female voice—are narrated by a single character. You have a flat delivery for everything in a movie.
A terrible experience.
It is a terrible experience. As soon as you learn English, you switch over and you don't want to watch content in this way. It's crazy that it still happens today for the majority of content.
Combining that with the fact that I worked at Palantir and my co-founder worked at Google, we knew that would change in the future and that all information would be available globally. As we started digging further, we realized it could be done in every language, in a high-quality way. That was the starting point, and the big thing was: instead of having it just translated, could you have the original voice, original emotions, and original intonation carried across?
Mhm.
Imagine having this podcast, but people could switch it over to Spanish and still hear Sarah and Mati, with the same voice and the same delivery. That's exactly what we did with Lex when he interviewed Narendra Modi, and you could immerse yourself in that story a lot better.
That was the original insight. As we started digging further, we realized that so much of the technology we interact with will change, whether this is how you create. It's still relatively tricky to bring voice alive. You need to go through the expensive process of hiring voice talent, having a studio space, and using expensive tooling to then actually adjust it. The tooling isn't intuitive, so that whole creation process will and should change to make it easier for new people with a keen interest to bring that to life.
A lot of the technology wasn't possible for you to recreate a specific voice or create it in a high-quality way. Then, of course, as we dug further and shifted away from the static piece, the whole interactive piece was still crazy in the way it functioned. Most of us have seen this technological evolution over the last decades, but you still spend most of your time on the keyboard. You look at the screen, and that interface feels broken.
It should be possible to communicate with devices through speech, through the most natural interface there is—one that started when humanity began. We realized we wanted to solve that. Fast-forward from 2022, and I feel like many people now carry that belief too: voice is the interface of the future. As you think about the devices around us—whether it's smartphones, computers, or robots—speech will be one of the key interfaces. In 2022, that wasn't the case, but if you think about the market, whether for the creative side or the interactive side, it was very clear that it would be a huge one.
Even when you think about the research part of your business, you have products for at least 2 different markets and then this larger mission. A lot has changed in the last 5 or 10 years, but it used to be a strongly held traditional belief that one must do one thing well in a startup and that there was no other path.
You're treating this like an interaction company, a platform company. How did you think about sequencing the research and product effort? Does that make sense, or were you thinking about new markets?
Wrapped up in that question is: where are we in terms of quality on voice? If the models aren't good enough for certain use cases, it doesn't make sense to build products around them.
I think that's right. It's almost exactly how we thought about it when we started. We initially tried to use existing models that were in the market and optimize them for our first use case, which was a combination of narration and dubbing on the creative side. We realized pretty quickly that the models that existed produced such robotic and poor speech that people didn't want to listen to it. That's where my co-founder’s genius came in: he was able to assemble the team and do a lot of the research himself to create new versions of that technology.
To your question, the way we're organized internally, and how we think about sequencing, came from looking at the first problem and then creating a lab around that problem—a combination of researchers, engineers, and operators—to go after it. The first problem was the problem of voice: how can we recreate the voice? As you say, it needs research expertise to do that well.
We started with what was effectively a voice lab, with the mission of asking, “Can we narrate the world in a better way?” It was a combination of roughly 5 people doing that work. We sequenced the research first and then built a simple layer on top of that work to allow people to use it. From there, we expanded it into a holistic suite for creating a full audiobook, and then creating a full movie narration and movie dubbing.
Then we moved to the next problem, which was the realization that we had solved voice for making content sound human. For that to be useful in interacting with technology, the next problem was solving how you bring knowledge on demand into it.
So we effectively started the second team, which was a second lab—an agent lab, effectively. It was a team that would combine researchers, engineers, and operators once more, trying to solve the problem: Okay, we have text-to-speech; how do we now combine this with LLMs and speech-to-text and orchestrate all those components together while integrating them with other systems to make it easier?
Similarly, you expand from looking just at the voice layer into how those systems work together. Here, too, you need research expertise to do that in a low-latency, efficient, and accurate way. At the same time, there’s a product layer that starts forming. It’s not only the orchestration that matters; it’s also the integrations—how you link up to legacy systems, how you build functions around it, and how you deploy it in production and test, monitor, and evaluate it over time.
Do you feel like you were creating new use cases when you built the tools? Did people know that they wanted to do this already? Because one argument that I remember hearing was, “Enterprises don’t know what to do with voice. How many people really want to do it?” And then you’re serving essentially perhaps the creator-publisher side of your business. Yeah.
It’s definitely a combination of initiatives that we believe will happen in the world and a response to a lot of that. As I think back, of course, the internal voice lab—or agent lab—kick-started so many of the other labs in response to the problems we started with.
We started a music lab because people wanted to create music with ElevenLabs using a fully licensed model. People wanted to use and create speech, but they also wanted to add music in a simple way. We wanted to deliver that, and then, of course, that came together through the question: How do we combine music, audio, and sounds?
We’re now integrating partner models from image and video into that suite. How could you combine all of that in one? A lot of that was in response to the market saying, “Hey, we would love this,” and then you have completely different use cases even in that space.
Let’s say dubbing. Dubbing is a use case where we didn’t feel there was a big push for it, but we knew that, in the ideal world of the future, you would be able to have content delivered naturally across languages while still carrying the original meaning. I still think this market will be immense, because it’s not going to be only the static delivery of movies.
If you travel around the world and want to communicate in real time, the full Babel fish idea from The Hitchhiker’s Guide to the Galaxy will happen. Breaking down language barriers—the barriers to communication and creation—all of that will break. That will be the foundational real-time dubbing concept. I’m super excited about that part.
Similarly, on the agent side, there are obvious things that customers or partners we work with will want to integrate. They’ll say, “We want integrations with XYZ systems.” But there are other parts that might not be as easy to predict. As you interact with technology, you of course want to understand what’s happening, but you also want to understand how things are being said and bring that into the fold.
That’s something we try to prioritize on our side, so that when people actually interact with the technology, they realize, “Oh, expressing a thing is actually so much more enjoyable, beneficial, and helpful.”
So I want to ask a question about this that relates to quality. I work with a series of companies where we’re selling a product to buyers who are generally not machine learning scientists.
Right. Right.
And even the scientific community does not have the full suite of evaluation benchmarks to understand every domain. It’s a well-known problem, but I imagine that for a lot of your customers, they don’t know how to choose a good voice. How do you deal with that problem? Is it, “Hey, I make a clone, and that sounds like me, so I believe it. I’m going to try all of these different options”? Or are you actually teaching people to do evaluations?
It’s a great question, because I think there are 2 big problems. One is: How do you benchmark the general space in audio, where it’s so dependent on the specific voice? Let alone if you are training it for interactive use, then it’s even trickier.
The second piece is, as you’re working on a specific use case, how do you select a voice? I’ll take the second front first. We have a voice specialist, effectively. As we work with enterprises, we deploy that person to work with them and help them navigate the process. That person is like a voice coach and has an incredible voice themselves. Now we have a team under that person that will partner with you to help you find the right branding voice.
And now you have a celebrity marketplace.
And now we have a celebrity marketplace to help you get iconic talent in there, like Sir Michael Caine. That piece was important because, of course, the voice will depend on the use case that you’re trying to build. The language and all of that will have an impact on what the right voice is for your customer base.
We have a voice person helping those companies, and some companies will be very opinionated about what they want. Sometimes they’ll select it themselves; sometimes they’ll give us a brief: “Hey, we want a voice that sounds professional, neutral, and calming.”
We recently had one of the biggest European companies give us a very unusual brief. They wanted as robotic a voice as possible.
Okay.
It was counterintuitive.
You’re like, “We can’t do that anymore.”
Almost. We were trying to go backward: How do we do that? But I think we got a good result.
Recently, we had a company in Japan and Korea where they wanted to serve different voices depending on the customer who was calling in. They have an older population and a very young population. For the younger one, they wanted one of the famous voices in the market that’s very excitable and happy. For the older one, they wanted a calm, slow-speaking voice. We help a lot with that.
It’s like a personalized choice, and then it can even be dynamic for a customer.
Yes. Okay. Exactly. Exactly.
And then maybe in the future, it will fully depend on your interaction. You’ll have a voice created as we understand the preferences of what people want. Let’s say you’re tired in the evening and want a slightly different voice—or maybe not. Maybe that’s the best focus time you have, and you want a voice that’s giving you that energy. It’s probably different when you wake up and it gives you the morning news or tells you what’s happening with the weather. All of those could be different.
Yesterday, we had dinner with some of our partners, and one of them said, “Hey, I have a new request for you. I want a New York voice with a Long Island accent,” which I never knew was a thing. Apparently, it is a thing. So we have that.
On the first piece, I think it’s still an unsolved problem. You have good benchmarks, of course, in LLMs. In image space, I think they’re pretty good. In voice space, you have speech quality, but so much of whether or not you like the speech depends on the voice.
If you compare Model A to Model B and serve them different voices, even if the quality is very different, the voice itself can make that sort of difference. We’ve seen this. I don’t know if you know the Artificial Analysis benchmarks, but I think they’re pretty good. Just switching the voice makes such a big impact.
That’s so interesting. And I wonder if, as you said, this is the most dominant interaction mode we’ve had for millennia—over all of human history, right? I think so.
I think so.
We’re just very sensitive to it. And I think people are going to be very sensitive to their own personalization as well.
100%. I think there’s also a third piece, which maybe is not directly related to your point. We’ve also realized that, even beyond the benchmarks and finding the right voice for your audience, the understanding of how you describe audio data is still lagging in the industry.
When we initially started, we of course went to the traditional players to help us label not only what was said, like transcription, but also how it was said—what the emotions were and what the accent was. Most people just weren’t able to do that work effectively, because you need to hear it and have a little bit of a skill set for how you would describe a specific delivery.
So we needed to create that ourselves. I think there’s that piece as well: How do you effectively interpret audio data on a more qualitative basis?
That’s trickier. Can you talk about what’s happening on the agent platform side? What is challenging for businesses or even creators that are trying to build agents, and what are the surprising or high-traction use cases?
I think everybody's aware of the idea of agent-based customer support, but I imagine you're doing many things beyond that.
Yeah. Exactly, customer support is probably the one that's kicking off the quickest, and it's the one that we see overtaking so many use cases, whether it's with Cisco, Twilio, or TELUS Digital. All of them are elevating that to a high extent.
I think the second exciting piece within that domain is the shift from effectively reactive customer support—you have a problem and reach out to customer support—into more of a proactive part of the customer-support experience. To make it explicit, we work with the biggest e-commerce shop in India, Meesho, where they started on the customer-support side: “I want a refund. I want to see the tracking of the package.”
That's now shifting to having an agent be a front part of the experience. If you go to the website, you have the widget, and you can engage with it through voice. You can ask it, “Hey, can you help me navigate to item X or item Y?” Or, “Can you explain what's the right thing for me to get as a gift at this time?”
It will help you based on your questions and what's on offer, show you those items, navigate to the right parts of the site, and maybe go all the way through checkout. I think this will be a phenomenal way of elevating the full experience, where that's more of an assistant across the whole thing.
We kicked off our work with Square to enable businesses to do that. The exact same pattern started with voice ordering: How can this now be part of the full discovery experience, too, where you get items shown to you and can have a lot more explanation? I think that will be a phenomenal piece, where effectively, from beginning to end, you have that experience.
So that's one category. The second one is the wider shift from static to immersive media, where there's so much incredible storytelling and IP that today exists in effectively one way of delivery, and now you'll be able to interact with that content in a completely new way.
I think one of the incredible use cases was working with Epic Games. We worked with them on bringing the voice of Darth Vader—and Darth Vader himself—into Fortnite, where millions of players could interact with Darth Vader live in the game. You had a full experience of Darth Vader in a new way.
I think this will be a theme across the board, whether it's talking to a book or talking to the character that you like. The whole space is shifting.
The one that I'm most excited about, for the world and for the shift, is going to be education, where you'll be able to have a personal tutor on your headphones and actually study something in an amazing way. I'll give you 2 quick examples.
One is that we recently worked with Chess.com. I'm a huge fan of chess. I'm a true chess fan.
Okay, great.
So you can learn chess, but you can have Hikaru Nakamura or Magnus Carlsen be your teacher for how you deliver that, which is amazing. You can even have the Botez sisters, or the whole plethora of different players who engaged with that, which I think is great.
Maybe a last one is MasterClass, which we worked with to shift from—you can of course have the content go through step by step—but you can also have an interactive experience. The best example of that was working with Chris Voss, the FBI negotiator and one of the top negotiators, who has a MasterClass lesson. You can actually call him and have a practice negotiation, which is crazy.
Yeah. Got to get that hostage out. We'll definitely try it.
Can I add one more? I think the last one combines all of them together, which I realized just recently and which was crazy.
Recently, I went to Ukraine, where we are working with the Ministry of Digital Transformation. They are effectively creating the first agentic government. The crazy thing is, they have all of those government—agentic government—systems. They want to change how they run all the ministries.
Okay.
It sounds like a big, ambitious, lofty goal.
No, I think the baseline is here. I'm on board with that immediately.
Yeah. The crazy thing is, I think they are so far ahead in actually doing that. There are 2 concrete things there.
One, they combine all those use cases. We're looking into how they can have effectively customer support for the government, whether it's asking about benefits or employment, or about the process of how you leave the country. All of that can be run through a digital app.
Then, 2, they can have a proactive way of informing citizens about things that might be happening, while also having an education system that runs through this personal tutoring experience. All of that is happening.
That was incredible to see. The second amazing thing was the way they've done it. They have the digital-transformation piece, but they also have engineering leaders in each of the ministries who lead those efforts and then bring them back to one central piece.
That is incredible to see, and I'm also proud to be able to be working with them on that shift. Despite everything that's happening, they're so far ahead.
That's amazing. That's really encouraging. Can I ask you a business-model question here? Looking at the strategic landscape—and I actually have many questions here—one observation I'd have is that if I look at one of these rich voice-and-action agent experiences, there are a lot of Fortune 500 and Global 2000 leaders who listen to the pod. I think a lot of them are going to buy the idea of, “I want this amazing, automated, real-time, available-24/7, every-language experience for my customer that's consistent and high quality.”
The ways I might get there include working with Palantir or a large consulting firm, working with ElevenLabs or a platform technology company like OpenAI, or working with a more use-case-oriented company like Sierra. Let's talk about that. How do you think about how people are making that decision, or how they should make that decision?
My past is also in Palantir, so I started from exactly that side. We do blend a lot of the forward-deployed engineering inside the company, too.
As I think about our offering and the choice customers are making, if you're looking for a single-purpose solution and only that one, then likely we aren't the best choice. If you're looking to deploy that across a plethora of different experiences—whether it's customer support, internal training, or elevating your sales efforts and actually increasing the top line with new experiences for how you engage customers beyond that reactive piece—then it's a great platform to build on.
We effectively combine that platform work with our engineering resources to help companies deploy on it. We also increasingly see this in Fortune 500s and Global 2000s, where companies want to build parts of things themselves because they already have a lot of investment in the platform, while engaging us on some of the new pieces and combining those.
I think our model, and the way it's different from a lot of the use-case-specific companies, is that our platform is relatively open. You can use pieces of that platform and not all of them for the different use cases.
Palantir, of course, or some of the consulting companies will have a lot more resources to go into the wider digital-transformation journey. In our case, it's very specific conversational agents. If you're looking for a new interface with customers, that's the best way.
Companies like Sierra are phenomenal, of course, in how they're thinking about the specific, pointed use case.
The other piece is that, as we think about our work, it depends on what you're optimizing for. We have a lot of international partners. If you have a wider geographic user base, that's great—that's what we optimize for.
Our voices, our languages, and our support for integrations internationally are so much broader. Depending on your exact scope, this will be a big factor. I would summarize it this way: If you're looking for a solution across a set of different use cases, and you want our engineering help to deploy it, then we are the right solution and probably the best solution.
I want to talk a little bit about OpenAI and the foundation-model companies. One of the reasons I called this podcast No Priors is because people are making a lot of assumptions all the time about how the market is going to work, and, lo and behold, many of those assumptions end up being nonsense. You can't—you have to very much decide your own narrative at this point in time.
Correct me if I'm wrong, but in 2022 and 2023, you probably heard a lot of people say, “Google can do this and OpenAI can do this. Why do you persist in working on voice anyway as a general capability?” What's the answer?
That also adds another element to a couple of the other previous questions. Whether it's agents' work or the creative work, to deploy the value in that work, you need a very strong product layer. You need the integrations, and you need to help people deploy the work, which is the most common piece.
Our superpower and our focus for a long time was building the foundational models to actually make that experience seamless. As I think about a lot of the companies in the market, they will optimize for a lot of other things, and that will be the differentiator.
In our case, we make the whole experience—especially with voice—seamless and human-controllable in a much better way.
So fundamentally, you would argue that the labs just aren't going to focus on this—and haven't?
Exactly. I think most of those companies—and that's the thing about the long term—will produce incredible research and incredible products that meet customers where they are and work backward from there.
I don't think the labs will focus on building that product layer that's so important. But I think the part of the question that you're asking is how—and why—they haven't done even the research part to the quality that we've been able to. Here, I'm also biased, but we're happily beating them on benchmarks with text-to-speech, speech-to-text, and the orchestration mechanisms. Credit to my co-founder and the team: they've been able to do it.
It's just mighty researchers continuing their work. But I think the main difference in the audio space is that you don't need scale as much as you need architectural breakthroughs and model breakthroughs to really make a dent. We've been able to do that a couple of times, and I think the number of people doesn't matter, but the people that you do have matter. We think there are maybe 50 to 100 researchers in the audio space who could do it. We think we have probably 10 of them in the company, and they're some of the best ones.
I think this obsession with those people working across the company, and actually giving them the full focus of the company to make it work and bring their work to production, then seeing how users interact with it and feeding that back, was so important. That's how we've been able to create models better than some of the top companies out there. But the truth is, to a large extent, why they weren't able to do it is also interesting. We don't know. They have incredible talent there, too.
How do you think about open-source models?
Anyone you ask in the company, I think, will say the same thing. That's a narrative we think about: in the long term, models will commoditize, or the differences between them will be negligible for some use cases. They will still matter for most use cases, but they will be broadly available. We don't know whether that's 2 years, 3 years, or 4 years, but it's going to happen at some stage.
Then, of course, you'll have a fine-tuning layer that will matter a lot on top of those models. But the base models, I think, will get pretty good. That's why, from the company perspective and also from the value perspective, the product piece is so important for us. If you have a model that's great, actually connecting your business logic and knowledge to it, and having the right interface for creating an ad for your work or completely new material, is a very different exercise.
And they'll be broadly available.
They will be broadly available, exactly. If I split it into two, for more of that asynchronous content narration, I think narration is pretty much solved. Open-source is great, commercial models are great, and the differences are getting smaller in out-of-the-box quality. What most of the models haven't figured out—and I think where we are—is how to make them controllable.
So that's the narration piece. I think the whole interaction piece—how you orchestrate the components together, whether that's a cascaded speech-to-text, LLM, and text-to-speech approach, or whether in the future it's a fused approach where you train them together—I think this is good for customer support or customer experience, but it's still a long way from a conversation like we have and from passing that Turing test.
I think this is still at least a year away—within a year. Then you'll have real-time dubbing, a variation of real-time translation and conversation, and I think that's maybe 2 years away, within 2 years. You know, a very uncomfortable belief that I feel comfortable having, but that I think is uncommon in the market right now, is that most advantages in technology could last you a year or they could last you 10, but they're not infinitely defensible.
If you think about that from a model-quality perspective or a product perspective, those advantages allow you to serve the customer better, build momentum, and build scale for some period of time. That's really powerful over time, but it's not a clean, forever answer. I think that makes businesspeople and investors uncomfortable.
And I mean, it's very true as well. [Laughter]
The way we think about it, research is a head start. This gives us the ability to give customers an advantage earlier, and it's 6 to 12 months of advantage. That's also a way for us to build the right product layer for you to get the best of that research. Frequently, we do that in parallel. The moment the research is out there, you have the product because we know our initiatives and we know what the product is.
You have research and product in parallel, and that extends the advantage. But the thing that will really give you long-term value is the ecosystem that you create around it—whether that's reach and distribution, the collection of voices you can have, the collection of integrations you can build, or the workflows that you can build. That's how we sequence it in our mind: research, product, and ecosystem. Research is a head start, allowing you to accelerate the future a little bit closer.
I think that's a really powerful insight, especially if the research team and the company team believe that internally as well.
I think the interesting piece for us—and I think this is the big question for all companies that do research and product—is whether you wait for research or make a product change. Even for research-product companies, do you wait for someone else to do the research? The timeline for that isn't clear. Is it 3 months, 6 months, or 12 months? You don't know exactly what it will do.
That's the hard choice: do I invest in the product layer, or do I just wait longer for the research? In our case, we internally let all the product teams know the research initiatives so we can parallelize that work, but we don't hold them back. If a product team thinks we should deliver value to the customer by doing something different, they can.
The rough rule of thumb is 3 months. If we think it's going to be longer than 3 months, we'll probably build it. If it's less than that, we probably won't.
Can you talk about some of the research that you're doing now, and how you think about the cadence of delivery and what's worth working on?
We now have a number of different initiatives across the audio space, and there are 2 big buckets. Roughly, they relate to the creative and agent sides.
On the creative side, this means text-to-speech models that are controllable. We then added a speech-to-text model that transcribes with high accuracy, including across low-resource languages, covering almost 100 languages. Then we created a music model, a fully licensed music model.
As you think about the future, it's also about how those models will interact with the visual space. There's a lot of effort going into how you can get the best of audio and potentially combine that with existing video that you have to really have the best delivery.
On the agent side, of course, it's how you optimize real-time speech-to-text and real-time text-to-speech. We just released our speech-to-text model, Scribe v2, which is under 150 milliseconds, with 93.5% accuracy across the top 30 languages on FLEURS. It's only the top 30 here because we serve so many others, but most of them don't. It's beating all the models on benchmarks.
As you think about the future, it's also the orchestration piece of how you bring speech-to-text, an LLM, and text-to-speech together. We'll be releasing, over the next couple of months, a new orchestration mechanism that will lower the end-to-end latency, we think, in a great way.
The second thing, which is so hard, is that it's not only going to allow you to combine those pieces, but also add the emotional context of the conversation, so you can actually respond with the model in a more expressive and better way. In the future, something we're investing in is parallelizing a speech-to-speech, more fused approach as well.
Of course, depending on the use case, if you have an enterprise, reliable use case, the cascaded approach is the approach for the next year or two.
It has more structure.
Yeah, more structure. You have more visibility into each of the steps, it's reliable, and you can call tools. If you want something more expressive and can tolerate hallucinations, speech-to-speech might be the choice. Maybe over time you'll see them go one over another depending on the industry.
That's a huge investment on our side. The foundation of the whole platform, and the main part that we're continually investing in, is a plethora of different models that combine the best of audio with some of the best of the other modalities.
I want to take our last few minutes and ask you a few questions about the future that I think you'll have a really good point of view on, given that you think about voice and audio all the time. What do you think of AI companions?
I think they will be a big thing and exist in a big way. It's not something I'm personally excited about or something that we spend much time on, but I think the whole line between an assistant, a companion, and a character that you enjoy as part of an experience will become blurry and blend to a large extent.
They can be very common, but you're not personally enthusiastic about them?
I'm more excited about the Jarvis version of that, or more of a super-assistant superpower.
It’s like the Jarvis version versus the social version.
The social version—I think it would be such an incredible unlock. It also involves blending into a person’s context. I would love to start the day with someone who understands me, tells me what’s relevant to me, opens the blinds, tells me what the weather and sunshine are like, and plays music straight away.
It’s going to happen.
It’s going to happen. That’s what I’m excited for. I think the companion use cases will mean solving loneliness, and in that part, I think that’s one way—maybe there are different ways of engaging people back. I do think there will be an interesting future, even if you think about education, where you will have a superpower for learning from AI tutors. But on the flip side of that—and this is my personal take—you will have a good percentage of time spent with AI tutors, but then an explicit percentage of time spent without any technology, human to human.
So you can kind of learn that part too.
Yeah, I think this is the correct model, both in terms of emotional guidance and coaching and guardrails, as well as peer-to-peer.
Exactly. What do you think about dictation, or what happens in terms of how we control technology that isn’t necessarily personified as well? Or does it just all become personified?
I think not all of it will be personified. Some things—communicating with an oven and a home—will probably stay pretty static.
Or code.
Yeah, exactly. You probably don’t need that much additional emotional input. But I think it’s going to be a huge part where, in a way, what I hope will happen is that you will have the ability to stay more immersed in real life, with the devices going back into the pocket, back into some version of an attached element, assuming that’s in the right setting, and that kind of acts on your behalf.
In many ways, let’s say dictation—as Karpathy says, “a decade of agents.” Let’s call it a decade. Then you’ll have a decade of robots. If you are interacting with robots, of course, voice will be the input and the output as one of the key interfaces. So you will need that dictation as a huge part.
I think the robot’s going to be personified.
Yeah, 100%. No, I think most of the use cases will be personified.
Okay, last one. What’s one thing that you’ve seen already exist today—or, if you project out a few years, will change about how we interact with content? Maybe it’s personalized voice content, or just something people are going to do with AI voice that they don’t do today or that not everybody knows about.
I think this is still the biggest one that hasn’t yet kicked into the system: how education will be done. I think learning with AI, with voice, where it’s on your headphones or in a speaker, is just going to be such a big thing. You’ll have your own teacher on demand who understands you, is very personified, and delivers the right content through your life. I think this will be one of the biggest use cases, and I don’t think it has happened yet.
I think we’ve seen, of course, some of the commercial partners, but schools and universities—how that’s deployed in a safeguarded way, in a way that supports the other part of education, the social part of education—I think all of that will evolve. Maybe there’s a cool version of that where you have Richard Feynman or Albert Einstein deliver those lecture notes, or other teachers that you love. It will be sick.
It’s a great note to end on. Thanks for doing this, Mati.
Thanks so much.