Sarah Guo
Hi listeners, welcome back to No Priors. Today I'm here with Mati Staniszewski, the co-founder and CEO of ElevenLabs, which was founded to change the way we interact with each other and with computers with voice. Over three short years, they've skyrocketed to more than 300 million in run rate. Mati and I talk about the future of voice, education, customer experience, and the other applications of this voice, as well as how to build a multi-segment company from self-serve to enterprise and combine research and product. Welcome, Mati.
Mati Staniszewski
Thanks for having me.
Sarah Guo
Thank you for doing this at 7 in the morning. It's great we finally got to do this together. I think a lot of our listeners will have used or played with ElevenLabs at some point, but for everybody else, can you reintroduce the company?
Mati Staniszewski
Definitely. At ElevenLabs, we're solving how humans and technology interact and how you can create seamlessly with that technology. In practice, we build foundational audio models: models that help you create speech that sounds human, understand speech in a much better way, or orchestrate all those components to make it interactive. Then we build products on top of those foundational models.
We have our creative product, which is a platform to help you with narrations for audiobooks, voiceovers for ads or movies, and dubbing those movies into other languages. Our agent platform is an offering to help you elevate customer experience, build agents for personal AI and education, and create new forms of immersive media. It's all under the umbrella of that mission: solving how we can interact with technology on our terms, in a better way.
Sarah Guo
You started the company in 2022.
Mati Staniszewski
That's right.
Sarah Guo
You've had amazing, rocket-ship growth since then. I'm sure it has felt different at different times, and I want to ask you about that. Can you give us a sense of the scale of the company today?
Mati Staniszewski
We've grown to 350 people globally. We started in Europe as a remote company, and we're still remote-first, but we have hubs around the world, with London being the biggest, New York being the second-biggest, and offices in Warsaw, San Francisco, Tokyo, and Brazil.
We're at $300 million in ARR, roughly 50/50 between self-serve—which is a lot of subscriptions and creators using our creative platform—and the enterprise side, which is approaching 50% and uses our agents platform. That's on the classic, sales-led side. We serve more than 5 million monthly active users on the creative side, and on the enterprise side, we have a few thousand customers, from Fortune 500 companies to some of the fastest-growing AI startups.
Sarah Guo
I think this is such an interesting company because it is very unintuitive to many people, and investors in particular. We were both there in 2022, and there's a class of companies that enable creation in some way. I would put ElevenLabs, Midjourney, Suno, and Hunyuan in this category. There's an overall sense of, “Who really wants to do this?”
What was your initial read of how many people would want to make voices? What made you believe it was going to be much broader than, for example, dubbing? Dubbing isn't a huge market.
Mati Staniszewski
It's very tricky to do both the product and the research. I'm in a lucky position because my co-founder and I have known each other for 15 years. I think he's the smartest person I know, and he's been able to create a lot of that research work to build the foundation and then elevate the experience.
Both of us are from Poland originally, and the original insight came from Poland. It's a peculiar thing, but if you watch a foreign movie in Polish, all the voices—whether it's a male voice or a female voice—are narrated by a single character. You have a flat delivery for everything in a movie.
Sarah Guo
A terrible experience.
Mati Staniszewski
It is a terrible experience. As soon as you learn English, you switch over and you don't want to watch content in this way. It's crazy that it still happens today for the majority of content.
Combining that with the fact that I worked at Palantir and my co-founder worked at Google, we knew that would change in the future and that all information would be available globally. As we started digging further, we realized it could be done in every language, in a high-quality way. That was the starting point, and the big thing was: instead of having it just translated, could you have the original voice, original emotions, and original intonation carried across?
Sarah Guo
Mhm.
Mati Staniszewski
Imagine having this podcast, but people could switch it over to Spanish and still hear Sarah and Mati, with the same voice and the same delivery. That's exactly what we did with Lex when he interviewed Narendra Modi, and you could immerse yourself in that story a lot better.
That was the original insight. As we started digging further, we realized that so much of the technology we interact with will change, whether this is how you create. It's still relatively tricky to bring voice alive. You need to go through the expensive process of hiring voice talent, having a studio space, and using expensive tooling to then actually adjust it. The tooling isn't intuitive, so that whole creation process will and should change to make it easier for new people with a keen interest to bring that to life.
A lot of the technology wasn't possible for you to recreate a specific voice or create it in a high-quality way. Then, of course, as we dug further and shifted away from the static piece, the whole interactive piece was still crazy in the way it functioned. Most of us have seen this technological evolution over the last decades, but you still spend most of your time on the keyboard. You look at the screen, and that interface feels broken.
It should be possible to communicate with devices through speech, through the most natural interface there is—one that started when humanity began. We realized we wanted to solve that. Fast-forward from 2022, and I feel like many people now carry that belief too: voice is the interface of the future. As you think about the devices around us—whether it's smartphones, computers, or robots—speech will be one of the key interfaces. In 2022, that wasn't the case, but if you think about the market, whether for the creative side or the interactive side, it was very clear that it would be a huge one.
Sarah Guo
Even when you think about the research part of your business, you have products for at least 2 different markets and then this larger mission. A lot has changed in the last 5 or 10 years, but it used to be a strongly held traditional belief that one must do one thing well in a startup and that there was no other path.
You're treating this like an interaction company, a platform company. How did you think about sequencing the research and product effort? Does that make sense, or were you thinking about new markets?
Wrapped up in that question is: where are we in terms of quality on voice? If the models aren't good enough for certain use cases, it doesn't make sense to build products around them.
Mati Staniszewski
I think that's right. It's almost exactly how we thought about it when we started. We initially tried to use existing models that were in the market and optimize them for our first use case, which was a combination of narration and dubbing on the creative side. We realized pretty quickly that the models that existed produced such robotic and poor speech that people didn't want to listen to it. That's where my co-founder’s genius came in: he was able to assemble the team and do a lot of the research himself to create new versions of that technology.
To your question, the way we're organized internally, and how we think about sequencing, came from looking at the first problem and then creating a lab around that problem—a combination of researchers, engineers, and operators—to go after it. The first problem was the problem of voice: how can we recreate the voice? As you say, it needs research expertise to do that well.
We started with what was effectively a voice lab, with the mission of asking, “Can we narrate the world in a better way?” It was a combination of roughly 5 people doing that work. We sequenced the research first and then built a simple layer on top of that work to allow people to use it. From there, we expanded it into a holistic suite for creating a full audiobook, and then creating a full movie narration and movie dubbing.
Then we moved to the next problem, which was the realization that we had solved voice for making content sound human. For that to be useful in interacting with technology, the next problem was solving how you bring knowledge on demand into it.
So we effectively started the second team, which was a second lab—an agent lab, effectively. It was a team that would combine researchers, engineers, and operators once more, trying to solve the problem: Okay, we have text-to-speech; how do we now combine this with LLMs and speech-to-text and orchestrate all those components together while integrating them with other systems to make it easier?
Similarly, you expand from looking just at the voice layer into how those systems work together. Here, too, you need research expertise to do that in a low-latency, efficient, and accurate way. At the same time, there’s a product layer that starts forming. It’s not only the orchestration that matters; it’s also the integrations—how you link up to legacy systems, how you build functions around it, and how you deploy it in production and test, monitor, and evaluate it over time.
Sarah Guo
Do you feel like you were creating new use cases when you built the tools? Did people know that they wanted to do this already? Because one argument that I remember hearing was, “Enterprises don’t know what to do with voice. How many people really want to do it?” And then you’re serving essentially perhaps the creator-publisher side of your business. Yeah.
Mati Staniszewski
It’s definitely a combination of initiatives that we believe will happen in the world and a response to a lot of that. As I think back, of course, the internal voice lab—or agent lab—kick-started so many of the other labs in response to the problems we started with.
We started a music lab because people wanted to create music with ElevenLabs using a fully licensed model. People wanted to use and create speech, but they also wanted to add music in a simple way. We wanted to deliver that, and then, of course, that came together through the question: How do we combine music, audio, and sounds?
We’re now integrating partner models from image and video into that suite. How could you combine all of that in one? A lot of that was in response to the market saying, “Hey, we would love this,” and then you have completely different use cases even in that space.
Let’s say dubbing. Dubbing is a use case where we didn’t feel there was a big push for it, but we knew that, in the ideal world of the future, you would be able to have content delivered naturally across languages while still carrying the original meaning. I still think this market will be immense, because it’s not going to be only the static delivery of movies.
If you travel around the world and want to communicate in real time, the full Babel fish idea from The Hitchhiker’s Guide to the Galaxy will happen. Breaking down language barriers—the barriers to communication and creation—all of that will break. That will be the foundational real-time dubbing concept. I’m super excited about that part.
Similarly, on the agent side, there are obvious things that customers or partners we work with will want to integrate. They’ll say, “We want integrations with XYZ systems.” But there are other parts that might not be as easy to predict. As you interact with technology, you of course want to understand what’s happening, but you also want to understand how things are being said and bring that into the fold.
That’s something we try to prioritize on our side, so that when people actually interact with the technology, they realize, “Oh, expressing a thing is actually so much more enjoyable, beneficial, and helpful.”
Sarah Guo
So I want to ask a question about this that relates to quality. I work with a series of companies where we’re selling a product to buyers who are generally not machine learning scientists.
Mati Staniszewski
Right. Right.
Sarah Guo
And even the scientific community does not have the full suite of evaluation benchmarks to understand every domain. It’s a well-known problem, but I imagine that for a lot of your customers, they don’t know how to choose a good voice. How do you deal with that problem? Is it, “Hey, I make a clone, and that sounds like me, so I believe it. I’m going to try all of these different options”? Or are you actually teaching people to do evaluations?
Mati Staniszewski
It’s a great question, because I think there are 2 big problems. One is: How do you benchmark the general space in audio, where it’s so dependent on the specific voice? Let alone if you are training it for interactive use, then it’s even trickier.
The second piece is, as you’re working on a specific use case, how do you select a voice? I’ll take the second front first. We have a voice specialist, effectively. As we work with enterprises, we deploy that person to work with them and help them navigate the process. That person is like a voice coach and has an incredible voice themselves. Now we have a team under that person that will partner with you to help you find the right branding voice.
Sarah Guo
And now you have a celebrity marketplace.
Mati Staniszewski
And now we have a celebrity marketplace to help you get iconic talent in there, like Sir Michael Caine. That piece was important because, of course, the voice will depend on the use case that you’re trying to build. The language and all of that will have an impact on what the right voice is for your customer base.
We have a voice person helping those companies, and some companies will be very opinionated about what they want. Sometimes they’ll select it themselves; sometimes they’ll give us a brief: “Hey, we want a voice that sounds professional, neutral, and calming.”
We recently had one of the biggest European companies give us a very unusual brief. They wanted as robotic a voice as possible.
Sarah Guo
Okay.
Mati Staniszewski
It was counterintuitive.
Sarah Guo
You’re like, “We can’t do that anymore.”
Mati Staniszewski
Almost. We were trying to go backward: How do we do that? But I think we got a good result.
Recently, we had a company in Japan and Korea where they wanted to serve different voices depending on the customer who was calling in. They have an older population and a very young population. For the younger one, they wanted one of the famous voices in the market that’s very excitable and happy. For the older one, they wanted a calm, slow-speaking voice. We help a lot with that.
Sarah Guo
It’s like a personalized choice, and then it can even be dynamic for a customer.
Mati Staniszewski
Yes. Okay. Exactly. Exactly.
And then maybe in the future, it will fully depend on your interaction. You’ll have a voice created as we understand the preferences of what people want. Let’s say you’re tired in the evening and want a slightly different voice—or maybe not. Maybe that’s the best focus time you have, and you want a voice that’s giving you that energy. It’s probably different when you wake up and it gives you the morning news or tells you what’s happening with the weather. All of those could be different.
Yesterday, we had dinner with some of our partners, and one of them said, “Hey, I have a new request for you. I want a New York voice with a Long Island accent,” which I never knew was a thing. Apparently, it is a thing. So we have that.
On the first piece, I think it’s still an unsolved problem. You have good benchmarks, of course, in LLMs. In image space, I think they’re pretty good. In voice space, you have speech quality, but so much of whether or not you like the speech depends on the voice.
If you compare Model A to Model B and serve them different voices, even if the quality is very different, the voice itself can make that sort of difference. We’ve seen this. I don’t know if you know the Artificial Analysis benchmarks, but I think they’re pretty good. Just switching the voice makes such a big impact.
Sarah Guo
That’s so interesting. And I wonder if, as you said, this is the most dominant interaction mode we’ve had for millennia—over all of human history, right? I think so.
Mati Staniszewski
I think so.
Sarah Guo
We’re just very sensitive to it. And I think people are going to be very sensitive to their own personalization as well.
Mati Staniszewski
100%. I think there’s also a third piece, which maybe is not directly related to your point. We’ve also realized that, even beyond the benchmarks and finding the right voice for your audience, the understanding of how you describe audio data is still lagging in the industry.
When we initially started, we of course went to the traditional players to help us label not only what was said, like transcription, but also how it was said—what the emotions were and what the accent was. Most people just weren’t able to do that work effectively, because you need to hear it and have a little bit of a skill set for how you would describe a specific delivery.
So we needed to create that ourselves. I think there’s that piece as well: How do you effectively interpret audio data on a more qualitative basis?
Sarah Guo
That’s trickier. Can you talk about what’s happening on the agent platform side? What is challenging for businesses or even creators that are trying to build agents, and what are the surprising or high-traction use cases?
I think everybody's aware of the idea of agent-based customer support, but I imagine you're doing many things beyond that.
Mati Staniszewski
Yeah. Exactly, customer support is probably the one that's kicking off the quickest, and it's the one that we see overtaking so many use cases, whether it's with Cisco, Twilio, or TELUS Digital. All of them are elevating that to a high extent.
I think the second exciting piece within that domain is the shift from effectively reactive customer support—you have a problem and reach out to customer support—into more of a proactive part of the customer-support experience. To make it explicit, we work with the biggest e-commerce shop in India, Meesho, where they started on the customer-support side: “I want a refund. I want to see the tracking of the package.”
That's now shifting to having an agent be a front part of the experience. If you go to the website, you have the widget, and you can engage with it through voice. You can ask it, “Hey, can you help me navigate to item X or item Y?” Or, “Can you explain what's the right thing for me to get as a gift at this time?”
It will help you based on your questions and what's on offer, show you those items, navigate to the right parts of the site, and maybe go all the way through checkout. I think this will be a phenomenal way of elevating the full experience, where that's more of an assistant across the whole thing.
We kicked off our work with Square to enable businesses to do that. The exact same pattern started with voice ordering: How can this now be part of the full discovery experience, too, where you get items shown to you and can have a lot more explanation? I think that will be a phenomenal piece, where effectively, from beginning to end, you have that experience.
So that's one category. The second one is the wider shift from static to immersive media, where there's so much incredible storytelling and IP that today exists in effectively one way of delivery, and now you'll be able to interact with that content in a completely new way.
I think one of the incredible use cases was working with Epic Games. We worked with them on bringing the voice of Darth Vader—and Darth Vader himself—into Fortnite, where millions of players could interact with Darth Vader live in the game. You had a full experience of Darth Vader in a new way.
I think this will be a theme across the board, whether it's talking to a book or talking to the character that you like. The whole space is shifting.
The one that I'm most excited about, for the world and for the shift, is going to be education, where you'll be able to have a personal tutor on your headphones and actually study something in an amazing way. I'll give you 2 quick examples.
One is that we recently worked with Chess.com. I'm a huge fan of chess. I'm a true chess fan.
Sarah Guo
Okay, great.
Mati Staniszewski
So you can learn chess, but you can have Hikaru Nakamura or Magnus Carlsen be your teacher for how you deliver that, which is amazing. You can even have the Botez sisters, or the whole plethora of different players who engaged with that, which I think is great.
Maybe a last one is MasterClass, which we worked with to shift from—you can of course have the content go through step by step—but you can also have an interactive experience. The best example of that was working with Chris Voss, the FBI negotiator and one of the top negotiators, who has a MasterClass lesson. You can actually call him and have a practice negotiation, which is crazy.
Sarah Guo
Yeah. Got to get that hostage out. We'll definitely try it.
Mati Staniszewski
Can I add one more? I think the last one combines all of them together, which I realized just recently and which was crazy.
Recently, I went to Ukraine, where we are working with the Ministry of Digital Transformation. They are effectively creating the first agentic government. The crazy thing is, they have all of those government—agentic government—systems. They want to change how they run all the ministries.
Sarah Guo
Okay.
Sarah Guo
It sounds like a big, ambitious, lofty goal.
Mati Staniszewski
No, I think the baseline is here. I'm on board with that immediately.
Mati Staniszewski
Yeah. The crazy thing is, I think they are so far ahead in actually doing that. There are 2 concrete things there.
One, they combine all those use cases. We're looking into how they can have effectively customer support for the government, whether it's asking about benefits or employment, or about the process of how you leave the country. All of that can be run through a digital app.
Then, 2, they can have a proactive way of informing citizens about things that might be happening, while also having an education system that runs through this personal tutoring experience. All of that is happening.
That was incredible to see. The second amazing thing was the way they've done it. They have the digital-transformation piece, but they also have engineering leaders in each of the ministries who lead those efforts and then bring them back to one central piece.
That is incredible to see, and I'm also proud to be able to be working with them on that shift. Despite everything that's happening, they're so far ahead.
Sarah Guo
That's amazing. That's really encouraging. Can I ask you a business-model question here? Looking at the strategic landscape—and I actually have many questions here—one observation I'd have is that if I look at one of these rich voice-and-action agent experiences, there are a lot of Fortune 500 and Global 2000 leaders who listen to the pod. I think a lot of them are going to buy the idea of, “I want this amazing, automated, real-time, available-24/7, every-language experience for my customer that's consistent and high quality.”
The ways I might get there include working with Palantir or a large consulting firm, working with ElevenLabs or a platform technology company like OpenAI, or working with a more use-case-oriented company like Sierra. Let's talk about that. How do you think about how people are making that decision, or how they should make that decision?
Mati Staniszewski
My past is also in Palantir, so I started from exactly that side. We do blend a lot of the forward-deployed engineering inside the company, too.
As I think about our offering and the choice customers are making, if you're looking for a single-purpose solution and only that one, then likely we aren't the best choice. If you're looking to deploy that across a plethora of different experiences—whether it's customer support, internal training, or elevating your sales efforts and actually increasing the top line with new experiences for how you engage customers beyond that reactive piece—then it's a great platform to build on.
We effectively combine that platform work with our engineering resources to help companies deploy on it. We also increasingly see this in Fortune 500s and Global 2000s, where companies want to build parts of things themselves because they already have a lot of investment in the platform, while engaging us on some of the new pieces and combining those.
I think our model, and the way it's different from a lot of the use-case-specific companies, is that our platform is relatively open. You can use pieces of that platform and not all of them for the different use cases.
Palantir, of course, or some of the consulting companies will have a lot more resources to go into the wider digital-transformation journey. In our case, it's very specific conversational agents. If you're looking for a new interface with customers, that's the best way.
Companies like Sierra are phenomenal, of course, in how they're thinking about the specific, pointed use case.
The other piece is that, as we think about our work, it depends on what you're optimizing for. We have a lot of international partners. If you have a wider geographic user base, that's great—that's what we optimize for.
Our voices, our languages, and our support for integrations internationally are so much broader. Depending on your exact scope, this will be a big factor. I would summarize it this way: If you're looking for a solution across a set of different use cases, and you want our engineering help to deploy it, then we are the right solution and probably the best solution.
Sarah Guo
I want to talk a little bit about OpenAI and the foundation-model companies. One of the reasons I called this podcast No Priors is because people are making a lot of assumptions all the time about how the market is going to work, and, lo and behold, many of those assumptions end up being nonsense. You can't—you have to very much decide your own narrative at this point in time.
Correct me if I'm wrong, but in 2022 and 2023, you probably heard a lot of people say, “Google can do this and OpenAI can do this. Why do you persist in working on voice anyway as a general capability?” What's the answer?
Mati Staniszewski
That also adds another element to a couple of the other previous questions. Whether it's agents' work or the creative work, to deploy the value in that work, you need a very strong product layer. You need the integrations, and you need to help people deploy the work, which is the most common piece.
Our superpower and our focus for a long time was building the foundational models to actually make that experience seamless. As I think about a lot of the companies in the market, they will optimize for a lot of other things, and that will be the differentiator.
In our case, we make the whole experience—especially with voice—seamless and human-controllable in a much better way.
Sarah Guo
So fundamentally, you would argue that the labs just aren't going to focus on this—and haven't?
Mati Staniszewski
Exactly. I think most of those companies—and that's the thing about the long term—will produce incredible research and incredible products that meet customers where they are and work backward from there.
I don't think the labs will focus on building that product layer that's so important. But I think the part of the question that you're asking is how—and why—they haven't done even the research part to the quality that we've been able to. Here, I'm also biased, but we're happily beating them on benchmarks with text-to-speech, speech-to-text, and the orchestration mechanisms. Credit to my co-founder and the team: they've been able to do it.
It's just mighty researchers continuing their work. But I think the main difference in the audio space is that you don't need scale as much as you need architectural breakthroughs and model breakthroughs to really make a dent. We've been able to do that a couple of times, and I think the number of people doesn't matter, but the people that you do have matter. We think there are maybe 50 to 100 researchers in the audio space who could do it. We think we have probably 10 of them in the company, and they're some of the best ones.
I think this obsession with those people working across the company, and actually giving them the full focus of the company to make it work and bring their work to production, then seeing how users interact with it and feeding that back, was so important. That's how we've been able to create models better than some of the top companies out there. But the truth is, to a large extent, why they weren't able to do it is also interesting. We don't know. They have incredible talent there, too.
Sarah Guo
How do you think about open-source models?
Mati Staniszewski
Anyone you ask in the company, I think, will say the same thing. That's a narrative we think about: in the long term, models will commoditize, or the differences between them will be negligible for some use cases. They will still matter for most use cases, but they will be broadly available. We don't know whether that's 2 years, 3 years, or 4 years, but it's going to happen at some stage.
Then, of course, you'll have a fine-tuning layer that will matter a lot on top of those models. But the base models, I think, will get pretty good. That's why, from the company perspective and also from the value perspective, the product piece is so important for us. If you have a model that's great, actually connecting your business logic and knowledge to it, and having the right interface for creating an ad for your work or completely new material, is a very different exercise.
Sarah Guo
And they'll be broadly available.
Mati Staniszewski
They will be broadly available, exactly. If I split it into two, for more of that asynchronous content narration, I think narration is pretty much solved. Open-source is great, commercial models are great, and the differences are getting smaller in out-of-the-box quality. What most of the models haven't figured out—and I think where we are—is how to make them controllable.
Sarah Guo
So that's the narration piece. I think the whole interaction piece—how you orchestrate the components together, whether that's a cascaded speech-to-text, LLM, and text-to-speech approach, or whether in the future it's a fused approach where you train them together—I think this is good for customer support or customer experience, but it's still a long way from a conversation like we have and from passing that Turing test.
I think this is still at least a year away—within a year. Then you'll have real-time dubbing, a variation of real-time translation and conversation, and I think that's maybe 2 years away, within 2 years. You know, a very uncomfortable belief that I feel comfortable having, but that I think is uncommon in the market right now, is that most advantages in technology could last you a year or they could last you 10, but they're not infinitely defensible.
If you think about that from a model-quality perspective or a product perspective, those advantages allow you to serve the customer better, build momentum, and build scale for some period of time. That's really powerful over time, but it's not a clean, forever answer. I think that makes businesspeople and investors uncomfortable.
Mati Staniszewski
And I mean, it's very true as well. [Laughter]
The way we think about it, research is a head start. This gives us the ability to give customers an advantage earlier, and it's 6 to 12 months of advantage. That's also a way for us to build the right product layer for you to get the best of that research. Frequently, we do that in parallel. The moment the research is out there, you have the product because we know our initiatives and we know what the product is.
You have research and product in parallel, and that extends the advantage. But the thing that will really give you long-term value is the ecosystem that you create around it—whether that's reach and distribution, the collection of voices you can have, the collection of integrations you can build, or the workflows that you can build. That's how we sequence it in our mind: research, product, and ecosystem. Research is a head start, allowing you to accelerate the future a little bit closer.
Sarah Guo
I think that's a really powerful insight, especially if the research team and the company team believe that internally as well.
Mati Staniszewski
I think the interesting piece for us—and I think this is the big question for all companies that do research and product—is whether you wait for research or make a product change. Even for research-product companies, do you wait for someone else to do the research? The timeline for that isn't clear. Is it 3 months, 6 months, or 12 months? You don't know exactly what it will do.
That's the hard choice: do I invest in the product layer, or do I just wait longer for the research? In our case, we internally let all the product teams know the research initiatives so we can parallelize that work, but we don't hold them back. If a product team thinks we should deliver value to the customer by doing something different, they can.
The rough rule of thumb is 3 months. If we think it's going to be longer than 3 months, we'll probably build it. If it's less than that, we probably won't.
Sarah Guo
Can you talk about some of the research that you're doing now, and how you think about the cadence of delivery and what's worth working on?
Mati Staniszewski
We now have a number of different initiatives across the audio space, and there are 2 big buckets. Roughly, they relate to the creative and agent sides.
On the creative side, this means text-to-speech models that are controllable. We then added a speech-to-text model that transcribes with high accuracy, including across low-resource languages, covering almost 100 languages. Then we created a music model, a fully licensed music model.
As you think about the future, it's also about how those models will interact with the visual space. There's a lot of effort going into how you can get the best of audio and potentially combine that with existing video that you have to really have the best delivery.
On the agent side, of course, it's how you optimize real-time speech-to-text and real-time text-to-speech. We just released our speech-to-text model, Scribe v2, which is under 150 milliseconds, with 93.5% accuracy across the top 30 languages on FLEURS. It's only the top 30 here because we serve so many others, but most of them don't. It's beating all the models on benchmarks.
As you think about the future, it's also the orchestration piece of how you bring speech-to-text, an LLM, and text-to-speech together. We'll be releasing, over the next couple of months, a new orchestration mechanism that will lower the end-to-end latency, we think, in a great way.
The second thing, which is so hard, is that it's not only going to allow you to combine those pieces, but also add the emotional context of the conversation, so you can actually respond with the model in a more expressive and better way. In the future, something we're investing in is parallelizing a speech-to-speech, more fused approach as well.
Of course, depending on the use case, if you have an enterprise, reliable use case, the cascaded approach is the approach for the next year or two.
Sarah Guo
It has more structure.
Mati Staniszewski
Yeah, more structure. You have more visibility into each of the steps, it's reliable, and you can call tools. If you want something more expressive and can tolerate hallucinations, speech-to-speech might be the choice. Maybe over time you'll see them go one over another depending on the industry.
That's a huge investment on our side. The foundation of the whole platform, and the main part that we're continually investing in, is a plethora of different models that combine the best of audio with some of the best of the other modalities.
Sarah Guo
I want to take our last few minutes and ask you a few questions about the future that I think you'll have a really good point of view on, given that you think about voice and audio all the time. What do you think of AI companions?
Mati Staniszewski
I think they will be a big thing and exist in a big way. It's not something I'm personally excited about or something that we spend much time on, but I think the whole line between an assistant, a companion, and a character that you enjoy as part of an experience will become blurry and blend to a large extent.
Sarah Guo
They can be very common, but you're not personally enthusiastic about them?
Mati Staniszewski
I'm more excited about the Jarvis version of that, or more of a super-assistant superpower.
Sarah Guo
It’s like the Jarvis version versus the social version.
Mati Staniszewski
The social version—I think it would be such an incredible unlock. It also involves blending into a person’s context. I would love to start the day with someone who understands me, tells me what’s relevant to me, opens the blinds, tells me what the weather and sunshine are like, and plays music straight away.
Sarah Guo
It’s going to happen.
Mati Staniszewski
It’s going to happen. That’s what I’m excited for. I think the companion use cases will mean solving loneliness, and in that part, I think that’s one way—maybe there are different ways of engaging people back. I do think there will be an interesting future, even if you think about education, where you will have a superpower for learning from AI tutors. But on the flip side of that—and this is my personal take—you will have a good percentage of time spent with AI tutors, but then an explicit percentage of time spent without any technology, human to human.
Sarah Guo
So you can kind of learn that part too.
Mati Staniszewski
Yeah, I think this is the correct model, both in terms of emotional guidance and coaching and guardrails, as well as peer-to-peer.
Sarah Guo
Exactly. What do you think about dictation, or what happens in terms of how we control technology that isn’t necessarily personified as well? Or does it just all become personified?
Mati Staniszewski
I think not all of it will be personified. Some things—communicating with an oven and a home—will probably stay pretty static.
Sarah Guo
Or code.
Mati Staniszewski
Yeah, exactly. You probably don’t need that much additional emotional input. But I think it’s going to be a huge part where, in a way, what I hope will happen is that you will have the ability to stay more immersed in real life, with the devices going back into the pocket, back into some version of an attached element, assuming that’s in the right setting, and that kind of acts on your behalf.
In many ways, let’s say dictation—as Karpathy says, “a decade of agents.” Let’s call it a decade. Then you’ll have a decade of robots. If you are interacting with robots, of course, voice will be the input and the output as one of the key interfaces. So you will need that dictation as a huge part.
Sarah Guo
I think the robot’s going to be personified.
Mati Staniszewski
Yeah, 100%. No, I think most of the use cases will be personified.
Sarah Guo
Okay, last one. What’s one thing that you’ve seen already exist today—or, if you project out a few years, will change about how we interact with content? Maybe it’s personalized voice content, or just something people are going to do with AI voice that they don’t do today or that not everybody knows about.
Mati Staniszewski
I think this is still the biggest one that hasn’t yet kicked into the system: how education will be done. I think learning with AI, with voice, where it’s on your headphones or in a speaker, is just going to be such a big thing. You’ll have your own teacher on demand who understands you, is very personified, and delivers the right content through your life. I think this will be one of the biggest use cases, and I don’t think it has happened yet.
I think we’ve seen, of course, some of the commercial partners, but schools and universities—how that’s deployed in a safeguarded way, in a way that supports the other part of education, the social part of education—I think all of that will evolve. Maybe there’s a cool version of that where you have Richard Feynman or Albert Einstein deliver those lecture notes, or other teachers that you love. It will be sick.
Sarah Guo
It’s a great note to end on. Thanks for doing this, Mati.
Mati Staniszewski
Thanks so much.