[BidClub_]
The Cognitive Revolution · · 65 min

a16z on AI Voices: Call Centers, Coaches, and Companions with Olivia Moore & Anish Acharya

Nathan LabenzOlivia MooreAnish Acharya

YouTube
TL;DR
  • Voice AI’s first durable commercial wedge is vertical B2B, not a standalone consumer app. Businesses already pay people to answer phones, making after-hours coverage and otherwise unanswered calls easy entry points; the strongest startups then expand into workflow ownership. Labenz says some are among “the fastest-growing B2B startups we’ve seen in 10 years.”

  • Conversational viability is largely here, but human likeness still depends on more than transcription accuracy. Latency is now typically below half a second, while Sesame showed how pauses, “ums,” and vocal inflection can turn a polished synthetic voice into one “that could be mistaken for a human.” Emotional adaptation, multi-party turn-taking, and interruptibility remain material gaps.

  • Happy Robot demonstrates why specialized conversation quality can unlock higher-value work. Its agents disclose that they are AI and befriend, disagree with, and negotiate with truckers for freight brokers; one tactic inserts a five-second “let me talk to my supervisor” delay before returning with a concession, which Moore says, with some uncertainty, produces a much higher acceptance rate. The investable insight is that better delivery earns permission to handle persuasion and pricing, where commodity voice does not.

  • The application moat is the vertical system around the voice, not voice generation alone. Enterprises need integrations, customer-specific context, evaluation, guardrails, and recovery when systems of record fail—“the capability gets you in the conversation but isn’t sufficient to get you to the other side.” That favors vertical platforms over horizontal agents, even as base models improve.

  • Apple’s postponed Siri overhaul illustrates an incumbent disadvantage, not a lack of technical progress. Acharya calls Siri “a stick in the eye five times a day” and argues that large companies struggle to embrace AI’s messy humanity; Moore adds that Apple must ship safely to hundreds of millions of users, unlike a startup serving self-selected beta testers. Labenz presents Google’s failure to commercialize Deep Research before ChatGPT became associated with it as a parallel missed opportunity.

  • Voice automation is producing coaching and task substitution, but not yet the 90% call-center headcount collapse Labenz tested as a scenario. Real-time coaching can justify hundreds of dollars per month when it influences a $10,000 HVAC upsell, while recruiting agents can return roughly 20 hours a week for work with five priority candidates. Despite call-center turnover reaching 300% annually, the guests had not observed order-of-magnitude job losses, and Acharya would not confidently translate an 18-month technology horizon into labor-market timing.

  • The consumer frontier extends from seniors and children to increasingly personalized companions. Multimodal voice could give seniors patient technical help, provide children with tutors or socially positive Minecraft partners, and produce companions ranging from sympathetic listeners to challenging “East Coast mode” personalities. Acharya’s longer-term framing is an “emotional bicycle” that extends people emotionally as computers extended them intellectually.

  • Safety policy must reconcile demonstrated impersonation risk with a market for licensed identity. Labenz reported that two calling platforms still let him clone Donald Trump and make scalable calls a year after he disclosed the problem, while Moore said people were currently more frustrated by model restrictions than by being cloned. Her twist on a do-not-clone registry is economically constructive: let people prohibit impersonation while explicitly licensing their voices or avatars for approved uses.

Digest · the substance, structured for research

1. Demand-side behavior is often the clearest product signal

  • Moore’s scouting funnel spans founder-heavy Twitter, newsletters, and meetups, but also Instagram, TikTok, and especially YouTube—the “number one mobile app and the number two website in the world.” For many consumer and prosumer AI products, YouTube tutorials are the largest source of social referrals.

  • The team watches ordinary users—“usually teenage girls, to be frank”—force ChatGPT into roles such as therapist, friend, or coach. Moore’s framing: “Consumer is so random and magical that we try to let the data tell us,” because pedigree cannot rescue a consumer product that misses timing, onboarding, or one decisive feature.

  • Moore’s less glamorous edge is simply using the products: Operator, Deep Research, DeepSeek, o1 Pro, and Krea were examples of tools many supposed insiders had not tried. Her memorable demand test updates an old social-app joke: “Every large language model is being tortured into being a therapist.”

2. Voice opens a technology surface that screens never addressed

  • Acharya starts from first principles: voice intermediates most human relationships, yet technology historically lacked the infrastructure to address it. Unlike computing surfaces with decades of product experimentation, voice remains “a complete blank piece of paper,” creating both product and distribution opportunities.

  • Moore sees more traction in B2B agents because even small businesses pay one to three people to answer phones. Once models approach human performance, using an agent for after-hours calls or calls otherwise sent to voicemail becomes an obvious substitution; many consumers may already have encountered one unknowingly.

  • Consumer exposure has instead flowed through ChatGPT, Grok, and the viral Sesame demo. Moore reads 1-800-ChatGPT as a clue that many people’s first meaningful AI experience—whether or not that product itself succeeded—may arrive by telephone rather than through a new app.

3. Seniors show why multimodality matters as much as conversation

  • Moore’s early-90s mother already asks Alexa to play music, making Alexa+ a plausible bridge into conversational computing. Voice could also unlock older technologies she never learned to operate, from email interfaces to the television remote.

  • Labenz’s reliable support instruction—“read everything on the screen from the top to the bottom”—already worked as a text-only GPT-4 prompt. A patient model could reproduce that help for people who lack an “infinitely patient” relative.

  • Acharya points to Google’s release—“I think it was late last year, in December”—of Gemini models that could see what was on a screen and interact in real time; OpenAI has a similar capability. His sharper point is that the relevant context is often physical: pointing a phone at the remote is more natural than verbally translating every button to Nathan or Google.

4. Sub-half-second latency solved entry-level voice, not conversation

  • Acharya says “solved” is too strong, but basic latency and understandability moved close to solved within the past year. Most models now respond in less than half a second—the difference between exchanging audio and being able to sustain a conversation at all.

  • Sesame’s advance was to preserve apparent imperfections: extra pauses, fillers such as “um,” and expressive inflections that a conventional system might treat as errors. Those details move output beyond a better Alexa or Siri toward a voice “that could be mistaken for a human.”

  • Emotionality remains unfinished. Founders want agents to recognize content and adapt tone—brighter for exciting news, lower and slower for sadness—while interruptibility remains awkward, especially in groups. Moore notes that humans themselves have not fully solved the moment when two people begin speaking together.

  • Labenz distinguishes voice models from conversation models: people coordinate turns through subtle audio and visual signals and often begin speaking before knowing the exact sentence. An AI that already knows its entire response can feel uncanny, suggesting that native conversation behavior needs more than a speech-to-text-to-LLM-to-speech pipeline.

5. Specialized delivery earns the right to negotiate

  • A16z portfolio company Happy Robot supplies voice AI to freight brokers. Acharya argues that its deeper technical work makes conversations feel materially better than commodity alternatives, which gives customers confidence to delegate “persuasion, negotiation, disagreement”—not merely information retrieval.

  • Moore’s standout example is deliberately added latency. The agent says, “Hold on, let me go talk to my supervisor,” waits about five seconds, then returns with a slightly better price, despite already knowing its permitted range; the human feels they earned a concession, and Moore thinks the final-offer acceptance rate is much higher.

  • Labenz’s reaction—“I don’t know how to feel about that”—preserves the tension between effective design and manipulation. The agent discloses that it is AI, yet truckers fall into familiar conversational rhythms anyway; intellectual awareness does not override the “reptilian brain” trained by negotiation rituals.

  • The broader claim is that AI can be “more human than the humans”: the same friendly agent answers each time, listens patiently, and can spend all the time in the world with the caller. Superhuman patience and low-to-no wait times are operationally valuable even before the model becomes superhuman at reasoning.

6. Incumbents face a cultural and distribution handicap

  • Apple’s announcement that a major Siri update would wait until 2027 feels absurd against rapidly improving startups. Acharya says the mismatch between poor everyday Siri performance and Apple Intelligence advertising degrades trust: “It’s like a stick in the eye five times a day.”

  • His diagnosis is institutional: AI works best across a messy surface of human interaction, while incumbents are organized to remove humanity and unpredictability from technology. Committees, lawyers, and efforts to “neuter the AI” create what he calls a “very difficult spiritual problem.”

  • Moore offers the counterweight: Apple must deliver something natural and correct to hundreds of millions of users across ages and use cases. Startups can ship imperfect beta products to voluntary early adopters; she guesses that public reaction to generated notification summaries spooked Apple.

  • Google Labs offers a partial model through gated experiments such as NotebookLM, but commercialization remains slow. Labenz’s sharpest example is Deep Research: it originated as a Gemini product that Google should have dominated, yet ChatGPT became known for the capability—“one missed opportunity after another with incumbents.”

7. Vertical context and integrations form the real enterprise moat

  • Native voice-to-voice models are earlier and somewhat more expensive; Acharya calls Gemini Flash perhaps the best current option, while noting that interruptibility still lags. His standing hedge is that “the models are the worst that they’re ever going to be right now,” so today’s multi-stage stack may look antiquated within a year.

  • Acharya sees reasoning models as a separate primitive despite their familiar interface. Probabilistic language behavior helps with friendship and rapport, while pricing and other factual decisions demand accuracy; orchestrating reasoning and language models can allocate each part of the conversation to the appropriate system.

  • Tool use makes voice generation only the opening layer. As Acharya puts it, “The capability gets you in the conversation but isn’t sufficient to get you to the other side”; durable products require workflows, integrations, outcome measurement, guardrails, and long-tail handling of backend failures.

  • Moore says this burden is especially severe for traditional enterprises that may struggle to build once, let alone keep models and systems of record synchronized. She emphasizes that vertically focused platforms can handle the long tail of industry conversation types and integrations; Labenz frames customer-specific context as the scarce “last mile” that a base model cannot infer from generic knowledge.

8. Coaching works now, while mass displacement remains unproven

  • Acharya says real-time coaching is already working for call-center staff, salespeople, and HVAC technicians. If one nuanced question determines a $10,000 upsell, even hundreds of dollars per month for an individual coach can deliver obvious returns; jobs with physical or deeply personal components remain particularly suited to augmentation.

  • Automation is also removing undesirable tasks. Call centers can experience roughly 300% annual turnover, while recruiting agents can conduct initial screens and return about 20 hours weekly for a recruiter to spend with five preferred candidates, including persuasion and candidate care.

  • Labenz rejects an easy reskilling story: displaced call-center workers will not necessarily move into higher-value jobs at the same employer. He asks whether the technology could support a 90% headcount cut and suggests that such an effect might be a 2026-plus phenomenon, while acknowledging that the question is about order of magnitude rather than a literal zero-human call center.

  • Acharya pushes back empirically: they had not seen such reductions, because jobs bundle screening with interviews, salary negotiation, onboarding, and even taking an employee to a baseball game. “Even if the technology is 18 months away,” he says, its labor-market effect remains hard to predict; abundance may shift the question from jobs toward purpose.

9. Consumer voice expands from SMB receptionists to emotional infrastructure

  • For restaurants, spas, and home-services companies, Acharya recommends vertical rather than generic agents. Owners generally are not firing the core employee who answered phones; they redirect that person toward customer experience and growth while the agent handles missed and routine calls.

  • Creator tools already span ElevenLabs voice cloning and described-voice generation, likely Delphi-style interactive digital clones, and HeyGen avatars. Moore thinks fully synthetic podcast hosting may remain “a couple years away,” although she can imagine a world in which a scripted episode requires neither camera nor microphone.

  • Labenz cites Synthesis, Ello, and Super Teacher as examples of personalized tutoring, while Moore’s strongest example is a positive AI Minecraft companion replacing “toxic teenagers,” alongside a multimodal classroom observer giving parents and teachers feedback unavailable in under-resourced schools.

  • Companion demand has also defied assumptions: the guests see more AI-boyfriend and interactive-fiction behavior, often serving women, than a market dominated by “frisky young dudes.” Companions might supplement relationships, teach conversation or flirting, absorb emotional strain, or challenge rather than flatter users—the jokingly named “East Coast mode.”

10. Identity licensing could pair protection with economic opportunity

  • Labenz cites research from Stanford discussed in an earlier Replika episode: users reported reduced suicidal ideation and, more often than not, greater engagement with the outside world, though he questioned some of the data and noted that part of it was self-reported. His conclusion is conditional: beneficial companions exist, but predatory and addictive versions can also be built.

  • His red-team evidence makes impersonation concrete: two reasonably well-known calling platforms still allowed rapid Donald Trump voice cloning and scalable outbound calls a year after he reported the vulnerabilities. He therefore favors disclosure rules and a do-not-clone registry sooner rather than later.

  • Moore sees the registry’s positive-market counterpart: people could license approved uses of their identity. ElevenLabs’ voice collections already create opportunities for voice-over artists, while an influencer with 5,000 followers might monetize an extensible avatar even when a large brand would ignore the human creator.

  • Acharya warns against reflexive paternalism, arguing that consumers have learned not to trust everything in books, online, or on social media. Moore’s long-term vision is voice across AirPods, glasses, computers, and every product—sometimes two-way, sometimes transcription-only—while Acharya calls it an “emotional bicycle” that extends human emotional capacity.

Nathan Labenz

Hello and welcome back to The Cognitive Revolution. Today, I'm speaking with Olivia Moore and Anish Acharya, partners at Andreessen Horowitz and fellow AI scouts who are constantly tracking emerging technologies and consumer behaviors, and who in recent months have really distinguished themselves as keen observers and eager early adopters of AI voice platforms and products.

In this conversation, we review recent developments in voice AI technology and explore how voice AI is already starting to transform business operations, user experiences, and human-computer interaction more broadly. On the technology side, we discuss how recent multimodal models have simplified the model stack and, with it, the application-development process; how reduced latency and improved interruptibility are enabling far more natural conversations than ever before; and how products including past guest Hume AI's Octave model, Google's interactive version of NotebookLM, and the very viral and incredibly natural-sounding Sesame are all beginning to demonstrate a remarkable level of emotional intelligence, both in their ability to understand the user and in the expressiveness with which they can communicate back.

On the impact side, we cover a bunch of really interesting real-world applications and trends, including the company Happy Robot, which uses voice AI to handle complex negotiations and build rapport with truckers in the context of freight brokerage; vertical solutions for restaurants and other types of SMBs; and how business owners are often employing these systems to handle after-hours and other calls that they can't answer themselves.

We also discuss how enterprises are using the technology to facilitate meetings and provide real-time coaching, why we haven't yet seen a major impact on call-center jobs and how soon that might happen, why Apple is now saying that Siri won't get a major update until 2027 even as so many other things are clearly starting to work, and, of course, the ongoing rise of AI companions for kids, seniors, and the lonely.

Along the way, we touch on a couple of philosophical questions relating to human-labor displacement and the delicate balance between protecting consumers and encouraging innovation. As someone who's been a career entrepreneur and genuinely loves using OpenAI's Advanced Voice Mode as, among other things, a biology tutor and a real-time video-game guide for me and my kids, I am super excited about this technology, but also legitimately concerned about just how scammy the world could quickly become if the models get even a little bit more lifelike.

As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, write a review on Apple Podcast or Spotify, or leave a comment on YouTube. We welcome your feedback, too, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network.

It had been a while since my last voice-AI red-teaming exercise, but with this conversation in mind, I did go back to 2 reasonably well-known AI calling-agent platforms to which I had previously reported flagrant vulnerabilities. I found that, sadly, a year later, both still allow me to quickly clone Donald Trump's voice and scalably call anyone and say anything in that voice, all with no meaningful controls.

This, to me, strongly suggests that we need new rules for these sorts of products sooner rather than later—both to protect the public and to protect the AI industry from its most careless developers. As it turned out, Olivia offered a brilliant twist on my “do not clone” registry idea, which could both expand economic opportunity and protect the public from AI impersonation, and which I feel it is truly time to build.

For now, I hope you enjoy this exploration of the rapidly evolving world of AI voice interactions with Olivia Moore and Anish Acharya, consumer-technology investors, AI scouts, and partners at Andreessen Horowitz. Olivia Moore and Anish Acharya from a16z are here to talk about the future of AI voice interactions. Welcome to The Cognitive Revolution.

Olivia Moore

Awesome. Thanks for having us. I'm excited.

Nathan Labenz

You guys are part of a group that I affectionately refer to as AI scouts: people who are out there on the edges of what exists and exploring it in a lot of different ways. Before we get into the object-level stuff of what you've found and what you think is coming in the realm of voice, I'd love to get a little bit of the meta lessons or specific alpha tips for how you do such a good job of this. Where do you go for information? What's the top of funnel for you? How do you know you're onto something? I think people could learn a lot from your example in that respect.

Olivia Moore

Yeah. In many ways, it's our job to be chronically online and tracking every new thing that happens, especially as consumer investors. It's something that we've tried to hone over the last few years in particular.

It's so interesting because there's a pretty big delta, I think, between where AI scouts, AI experts, and early adopters are spending their time and talking, and where normal consumers are. So we try to be in both places. I would say Twitter, of course, is where most founders and AI founders are announcing new companies, new models, and new breakthroughs. AI newsletters have been massive, as have meetups.

But in terms of where real people are sharing what they do with AI, it's mostly on places like Instagram, TikTok, and YouTube. YouTube is a shocking one. YouTube is actually the number-one mobile app and the number-two website in the world. We find that for many companies in the consumer/prosumer space, if you look at traffic, by far their number-one referral source from social is YouTube.

There's this whole separate economy of YouTube influencers or YouTube creators who are making how-to content about using different AI tools. So I would say we try to track all those places.

In terms of what's an early signal to us, often we'll see normal people—usually teenage girls, to be frank—trying to manipulate ChatGPT into doing something, like being a therapist, a friend, or a coach. Once we see something like that, it's like, “Okay, the consumer pull is strong enough that there probably can and will be a couple of standalone, more focused products here.”

I think the other thing we try to do a lot of is just use the products. That sounds obvious, but it's surprising how few people seem to have actually tried Operator, Deep Research, DeepSeek, o1 Pro, or Krea. These are not super-obscure, long-tail products, and consumers find their way to them. But among the folks who are insiders and are paid to be doing this work, surprisingly few actually look at those products. It's just a great way to build your intuition.

Nathan Labenz

Yeah, that's my number-one piece of advice too: just get hands-on. You can't really go wrong with that. The one interesting thing there was that it sounds like you're looking as much, or maybe even more, for demand-side pull as you are at the people coming forward with their technologies to offer. You're looking for people who are specifically trying to meet a need that maybe nobody's met yet and figuring out what that implies.

Olivia Moore

Yeah. Consumer is so random and magical that we try to let the data tell us. You can have the most tenured, pedigreed team in the world building a consumer app, and if it doesn't hit, it doesn't hit. That can be for a variety of different reasons. Maybe it's bad market timing. Maybe they got the product insight wrong but some specific feature right, or the product insight right but some specific feature wrong. Then no one completes the onboarding.

I would say we try to let the data, in terms of what people are actually using, tell the story to us. Sometimes looking at data such as people pulling ChatGPT into these off-label use cases will give us a little warning signal or a heads-up that, “Okay, this is something—this is a behavior that's working,” and so we should keep an eye out for products that are targeting this behavior.

A great example of this is the old joke in pre-AI that every social app would collapse into being a dating app. In the same way, every large language model is being tortured into being a therapist. It's a funny thing to say at a dinner party, but it's also a leading indicator of what consumers want these models to do and some of the things that we want to see in the future.

Nathan Labenz

Cool. Well, let's talk about voice. To start, I would love to get the top highlights: What products and experiences have you seen that are the absolute best user experiences out there today? Hopefully, I've tried them, but we're about to find out.

Anish Acharya

Maybe I can frame it, and then Olivia can talk about some of the specific products. I think the thing to zoom all the way out on is grounding ourselves in the fact that voice mediates every human interaction and relationship, largely, right? Here we are, obviously, having voice mediate our relationship and our conversation and the way that we get to know each other.

It really is the original and most important form of human communication, but it's just been completely unaddressable by technology because we've never had the infrastructure. It's very interesting because so many of the other substrates that we're applying AI to are areas where we've also had a lot of historical technology exploration, whereas voice is just a complete blank piece of paper. That's why I think we're as excited about the product implications as we are about the distribution implications of this sort of technology surface.

Olivia Moore

Totally. Yeah. I would say there have been a couple of surprising things for us in terms of where voice is working now, at least on the startup side. A lot of the startups that are getting real traction in terms of net-new companies and products are actually more B2B-oriented, just because there are so many businesses that are now running off call centers or paying for 1, 2, or 3 people, even for small businesses, to answer the phone all day. Once you're at a point where voice models can be anywhere in the realm of human performance there, it kind of makes all the sense in the world to at least have the voice agent be doing your after-hours calls or your calls that would go to voicemail.

My guess would be that a lot of people have maybe interacted with an AI voice agent and not quite known it, because it's been a business calling them, or it's been the receptionist when they've called to schedule an appointment or something like that. On the consumer side, it's been so far maybe a little different than we expected. I think most consumers have interacted with AI voice through something like ChatGPT or Grok, which are incredible voice experiences. More recently, something like Sesame was a massive breakthrough, and that is still just a web demo, an early version of what's to come there.

My guess is, when we see the Sesame team open-source the model, and when we see models like that spread and become more accessible for app builders to build on top of, we'll see a corresponding explosion in consumer-focused, voice-first tools. I mean, a crazy thing that happened—and things are moving so quickly, I think it's easy to forget these things—is 1-800-ChatGPT. What was that? And I think, whether it failed or succeeded, it pointed to an important insight, which is that maybe the first way most people in the world will actually experience AI is via voice, both as consumers and as consumers consuming business offerings.

Yeah. I think of my dear mama all the time. She's now in her early 90s and lives alone, and is sharp but not an early adopter of new technologies. For her, I think it's going to be Alexa+ that's going to be the big transition. She already sits there and asks it to play music for her, but will she engage it in conversation? How natural will she find it? I don't know, but it's clearly going to be that form factor for her that could unlock a whole new set of things.

What's very funny is, ironically, Mama perhaps is now exploring new technology, but also isn't that familiar with old technology, which is maybe why she's calling you to get tech support and help using her existing products and devices. I think the potential applications for voice as applied to seniors are super interesting. We've discussed it a lot, and it's not just access to the new things; it's access to the old things as well that they just never developed the skills to interact with.

If she can figure out the TV remote, then we'll really be in business.

Nathan Labenz

Yeah, funny enough, one of my first GPT-4 tests, going way back to the red-team days, was tech support for seniors. The prompt I found to work really well is exactly what I say to her when what you're describing happens—when she calls me and says, “My friend emailed me, and I can't find it.” I always tell her, “Read everything on the screen from the top to the bottom.”

She'll literally go, “Okay, Verizon, the time,” and then eventually we get down to where the issue is. That same thing basically worked out of the box with GPT-4. Of course, that was text-only at that time, but I do see that as a huge unlock for all sorts of different stuff that she struggles to access right now.

Anish Acharya

Oh, exactly. Well, I think most people don't have an infinitely patient person like you to walk them through how to do that. We even saw recently—I think it was late last year, in December—Google release the Gemini models that could see what was on your screen and interact with you in real time. It feels like we're right on the brink of models like that. OpenAI has one as well, becoming API-available and ready for builders to actually capitalize on. Once we see something like that become actually usable, it's going to be massive.

It's also so interesting because it points at something maybe Google and other search players should have done pre-AI: How do you take everything on the internet and apply it? The most important context is the context that's around me in my physical space. The idea of being able to point your phone at the remote, in the case of Nana, and be able to debug the problem that way, instead of trying to translate what you're seeing in the physical world line by line to either Nathan or to Google—it just, that interaction pattern doesn't make sense.

Nathan Labenz

Hey, we'll continue our interview in a moment after a word from our sponsors.

Even if you think it's a bit overhyped, AI is suddenly everywhere. From self-driving cars to molecular medicine to business efficiency. If it's not in your industry yet, it's coming and fast. But AI needs a lot of speed and computing power. So, how do you compete without costs spiraling out of control? Time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure or OCI. OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all of your AI and machine learning workloads. OCI costs 50% less for compute and 80% less for networking. So you're saving a pile of money. Thousands of businesses have already upgraded to OCI, including Vodafone, Thompson Reuters, and Sunno AI. Right now, Oracle is offering to cut your current cloud bill in half if you move to OCI for new US customers with minimum financial commitment. Offer ends March 31st. See if your company qualifies for this special offer at oracle.com/cognitive. That's oracle.com/cognitive.

The Cognitive Revolution is brought to you by Shopify. I've known Shopify as the world's leading e-commerce platform for years, but it was only recently when I started a project with my friends at Quickly that I realized just how dominant Shopify really is. Quickly is an urgency marketing platform that's been running innovative, timelimited marketing activations for major brands for years now. We're working together to build an AI layer which will use generative AI to scale their service to longtail e-commerce businesses. And since Shopify has the largest market share, the most robust APIs, and the most thriving application ecosystem, we are building exclusively for the Shopify platform. So if you're building an e-commerce business, upgrade to Shopify, and you'll enjoy not only their marketleading checkout system, but also an increasingly robust library of cuttingedge AI apps like Quickly, many of which will be exclusive to Shopify on launch. Cognitive Revolution listeners can sign up for a $1 per month trial period at shopify.com/cognitive where Cognitive is all lowercase. Nobody does selling better than Shopify. So visit shopify.com/cognitive to upgrade your selling today. That's shopify.com/cognitive.

So one of the things I noticed in your presentation about this is that you said it's basically solved, I think was the phrase. I'm wondering what you see as remaining, if anything. The Sesame model might have even come out since that presentation, and it certainly takes another step forward in terms of the overall natural sound of the voice.

I find the interruption mechanics still a little bit gnarly, especially if there's a multi-way conversation. Sometimes I try to demo Advanced Voice Mode for people, just to bring them up to speed. My goal is to make it a quick way to bring people up to speed—I'm like, “This is where AI is at now if you haven't been paying attention.” But those demos often go a little bit sideways on me because the interruption mechanic is still a little weird.

If I'm doing it one-on-one, it's okay; I kind of know how to use it, and it seems to be optimized for that one-on-one. But in the group setting, it doesn't do super well in many cases. Anyway, that's just me identifying my remaining pain points, but what do you think are the hardest things to get right, or the most important things that still need to get solved?

Anish Acharya

Yeah. I think saying that things like basic latency and understandability have gotten solved is too strong, but they're very close to solved. Over the past year, those are the things that have gotten very close to solved, and that's the difference between being able to have a conversation and not have a conversation. In most cases, I think the models are getting those right now. Latency is less than half a second on most of the models, which feels very human-like.

I think the realization that a lot of people have had who've tried the Sesame demo is that, beyond latency, there are so many speech-pattern nuances that an actual human will say, which might actually sound like an error to the model: extra pauses, saying “um” or “you know,” and vocal inflections. When you add those things in, which the Sesame team has done, it goes from being a voice that sounds much better than Alexa and Siri, but is still kind of robotic in some ways, or so clearly AI, into something that could be mistaken for a human.

I think the remaining things for me—there's still a lot to do around emotionality. I talk to a lot of founders building voice agents who want the models to be able to understand what they're saying and vary the tone and the inflection based on that.

Olivia Moore

So if the AI voice agent is going to say something happy or exciting, the voice should reflect that. If it's going to say something sad, the voice should be lower in tone and pitch and a little bit slower. That's something that we still need to solve. And then interruptibility is huge.

I think of it as humans having also not solved interruptibility in conversations. We still have the issue where 2 people start talking at once and you have to be like, “No, no, you go ahead.” So we need a clever way for voice AI to be able to solve that in a way that humans maybe have not yet.

Nathan Labenz

I think this is why we've gotten where we need to get to for voice models to work. But I think what Sesame showed us is that conversation models may be very different from voice models, or maybe an extension of voice models. Even the way the 3 of us are coordinating on who's going to talk next involves so much nonverbal communication, whether it's video or even audio.

The models have been trained really well to do speech-to-text and text-to-speech. Interestingly, fewer of the companies we're seeing are using native voice-to-voice models. So that's 1 opportunity, and then there's just generally understanding, as Olivia noted, some of the nuances of conversation and having that be natively programmed in.

For example, when I start talking, I'm not exactly sure what I'm going to say. It sort of comes—it like—and that's true for all of us, whereas the AI knows exactly what it's going to say, which can sometimes be a little bit creepy. It does not quite get out of the uncanny valley.

There's a recent paper from Meta that you're bringing to mind, where they're doing brain reading and have sort of established, in actual signal-understanding terms, the kind of phrase-level formation that happens about 2 seconds before you actually talk, and then how it kind of gets down to the literal next syllable that's just a fraction of a second before you spit it out.

There is something that feels like maybe some of these things might be better served in some ways by a diffusion structure, if you could kind of go from coarse to fine, as opposed to always being 1 token at a time at the end. But anyway, that's more speculation.

Anish Acharya

You mentioned the stack. Actually, can I add 1 more thing? You actually do see varying levels of performance and conversation quality, which is a very big predictor of driving business outcomes. For example, we're investors in a company called HappyRobot, which is voice AI for freight brokers.

If you just look at the quality of their text-to-speech and how the conversation feels, it feels way better than many of the off-the-shelf things. This is because the team is more technical and has done more work under the hood. So yes, there are a bunch of other competitors who can provide a fine commodity voice experience, but if you're a little bit more specialized in the technology, you can do something that feels more human, and that gives you permission to move into higher-value conversations.

Persuasion, negotiation, disagreement—those are pretty nuanced conversations. And for the business to trust you with those conversations, you've got to have a voice model that not just says the right things, but says them in a way that feels compelling.

Nathan Labenz

Yeah, negotiation in particular is a lot to ask somebody to delegate to an AI. Is this company actually doing that—doing actual negotiations?

Anish Acharya

Negotiates, befriends, disagrees. We'll send you the demo. You should embed it. It's amazing. It's a real aha moment, I think, in terms of what's possible.

Olivia Moore

It's really interesting because, to the point about the LLM always knowing exactly what it could or should say, there is a version of an AI voice agent that does negotiation that would just respond to the human and say, “No, this is my best price. This is what I'm going to offer you,” and just say it over and over.

But if you launch that kind of experience to a user, they're going to try to circumvent it and talk to a human agent. It's not going to feel like an actual negotiation to them, where they've given it their best and they've gotten a concession from the other side.

What HappyRobot has done, which is really smart, is introduce extra latency by saying, “Okay, the voice agent will say, ‘Hold on, let me go talk to my supervisor,’” and put them on hold for about 5 seconds. Then they come back with a slightly better price.

Of course, the voice agent knows, “Here is my actual max. This is how much I can go up or down or move the price for the end customer.” But they found that the acceptance rate of that kind of final offer is, I think, much higher in cases where the human feels like they've gone through an actual negotiation, because the voice agent has simulated a situation that feels satisfying to them.

Nathan Labenz

Yeah, I don't know how to feel about that, to be honest. It's genius, and it's a little bit far out. Staying on this for a second, do people know that they're talking to an AI when they're talking to this? Does it say that upfront, or how do they—

Anish Acharya

Yes. It discloses it. What's actually most surprising is that people don't mind.

These are truckers driving all over the country in their big rigs, so these aren't exactly Stanford technology enthusiasts. They don't mind at all because our reptilian brain is so trained to react to these interactions in a specific way that, once you get into the conversation, even if you intellectually know that it's an AI, you fall into the rhythm, expectations, patterns, and cadence of human conversation very quickly.

Olivia Moore

It reminds me of something Anish says a lot, which is that in the best cases, AI can be more human than the humans. The HappyRobot example is a good one: every time you call in and reach the voice agent, you get the same person. They're friendly. They'll listen to you talk about your day or what happened. They'll be very patient and sympathetic.

They'll spend all the time in the world with you on the phone if you want to. So it's actually, in many cases—assuming the voice agent can answer your question, which almost all of them can now—a better experience for the end consumer than if you get what can sometimes be a grumpy actual human being on the other end of the phone line.

Nathan Labenz

Yeah. Superhuman patience turns out not to be that hard to achieve and is quite valuable. Low-to-no wait times are also a huge driving force for value. I definitely see it. I've been really excited about voice for a long time, even though it's probably going to put me, as a podcaster, out of work before too long as well.

You mentioned Siri a minute ago, and obviously they've recently made headlines in a seemingly negative way by saying they're not going to have an update until 2027, which just feels like possibly the other side of the singularity from where I'm sitting. I don't know if you can make any sense of that. I guess another way to come at that is: How reliable do these things need to be?

I think there's often, in AI in general, this faux comparison or imagined perfection that people compare an AI solution to. Ethan Mollick, I think, has a great phrase for this as well: the best available, or the best hirable, human for the job is really the comparison that you should be making. So is that what Apple is getting wrong here, or how do you make sense of that development?

Anish Acharya

I have so many thoughts on this. One, I think for any consumer that interacts with AI products and also uses things like Siri, it's like a stick in the eye 5 times a day because Siri is still so bad at just the most basic things. And then, juxtaposed with all the advertising that Apple's doing about Apple Intelligence, it's not awesome. I think it really degrades consumer trust.

And 2, I think that AI does best when it is exploring the surface of human interaction, which is a little messy, and these large incumbent corporations are designed to take the humanity out of every technology product. So there's a sort of almost irreconcilable tension between the 2.

The more they try to neuter the AI, the more dissatisfied consumers are with it. Some of the Genmoji stuff—it's a valiant effort, but it just looks terrible. I don't know; maybe some people think it looks good. I don't.

I think it's going to be a very difficult spiritual problem for these companies to resolve because, between the committees and the lawyers and the whole posture that large incumbents have, it's going to be hard for them to embrace the messiness of AI.

Olivia Moore

Yeah, I mean, we saw the reaction to the AI-generated text and notification summaries on iPhones, and I would guess that kind of spooked Apple a little bit. To Anish's point, for Apple to launch a new AI product, it has to be production-ready to land on hundreds of millions of iPhones for people of all ages and all sorts of use cases.

It needs to feel both natural and also be correct, whereas a startup has the luxury of not having to meet that bar. The people who seek out and try a new startup product are the natural early adopters, and they know that it's beta, they know that it's AI, and they know that it's a test product.

I think—not to give too much credit here—but I think what we've seen Google do well has been the new Google Labs experiments. That's where a lot of the best, in my opinion at least, AI Google products have come out of, like NotebookLM and some of the video models, like Veo 2.

Essentially, they've taken the approach that if you're an early technology tester, you sign up and get on the waitlist, beta-test, and use these products, and then ideally they get them production-ready to launch to a slightly larger audience incrementally.

Nathan Labenz

But even as we've seen with NotebookLM, because it's Google, once it does make its way to the public, the pace of innovation there is a lot slower than it would be at a net-new company that doesn't have tens of thousands of people to continue to employ. It's like 10 VPs for every engineer working on NotebookLM right now.

And look, another example of this is Deep Research. Deep Research actually was originally a Gemini product, and it's obviously a product that Google should be the best in the world at, yet they never commercialized it in the right way for whatever reason. Now ChatGPT is known for its deep research capabilities. So it's just one missed opportunity after another with incumbents. We'll circle back to that in a minute.

I wanted to get a little bit deeper on the stack and the sort of balance between allowing the AI to handle things and putting your trust and confidence in its decision-making versus trying to maximize accuracy, which typically is going to mean more control measures and more stilted, or let's say less natural, interactions. The stack, for those that aren't familiar with this, basically has been the sort of audio in, transcribe that to text, feed that into a language model, feed that response back into another text-to-speech model, and then send the speech back to the user, roughly speaking.

You can complicate that if you want to, but that pipeline basically now does work fast enough in many cases to be viable. But then there's the fully multimodal model, voice-to-voice, all through one set of weights. When you said that builders are mostly still building on these more multi-model, pipeline-style technology stacks, is that because they're getting better results from it, or is it because that's the only thing that's economical for their use case right now? What's driving that, and do you see that changing?

Anish Acharya

Yeah, I think the voice-to-voice models are definitely, and unsurprisingly, probably a little bit more expensive right now. But also, I think they're just earlier. I talk to a lot of voice agent founders who have probably tried all of the models available, and I think Gemini Flash is maybe the best of the bunch if you're going to try to do full voice-to-voice. But in general, the interruptibility is not quite there yet on those models.

But as we see and say pretty much every week, the models are the worst that they're ever going to be right now. I'm sure, similar to how last May, when we first released our voice report, sub-1-second latency was hard to imagine, by this time next year everyone is going to be using a model stack that's much sleeker and is hard for us to even imagine right now.

I also think that reasoning models are a new primitive, and I think that they're maybe being underappreciated by many because they have the same interface as language models. But if you think about it, there are aspects of interactions that really benefit from this probabilistic nature of the outputs of language models—all the friendship, negotiation, et cetera, that you have in a business conversation, even.

And then there are things that really do not benefit from that, where you want high degrees of accuracy, like a negotiation: 1 price is definitively better than another for 1 of the parties and the counterparty. So in that case, you can orchestrate reasoning models and language models to handle the right parts of the conversation to get the desired output. So I think you have fewer issues around aspects of the conversation that demand accuracy and the models not being well suited to that.

I think the other broad philosophical point Marc has mentioned a lot is: Is the bar to be as good as humans, or is the bar to be perfect? If the bar is to be perfect, all these technologies are not there and maybe never will be. If the bar is to outperform humans, I think we're actually there in many ways.

Nathan Labenz

Hey, we'll continue our interview in a moment after a word from our sponsors.

What does the future hold for business? Ask nine experts and you'll get 10 answers. Bull market, bare market, rates will rise or fall, inflation's up or down. Can someone please invent a crystal ball? Until then, over 41,000 businesses have future-proofed their business with Netsuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, and HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data. When you're closing books in days, not weeks, you're spending less time looking backward and more time on what's next. As someone who spent years trying to run a growing business with a mix of spreadsheets and startup point solutions, I can definitely say don't do that. Your all-nighters should be saved for building, not for prepping financial packets for board meetings. So whether your company is earning millions or even hundreds of millions, Netswuite helps you respond to immediate challenges and seize your biggest opportunities. And speaking of opportunity, download the CFO's guide to AI and machine learning at netswuite.com/cognitive. The guide is free to you at netswuite.com/cognitive. That's netswuite.com/cognitive.

Yeah, it's all happening very fast. How is the tool use? Because that's one other thing that I could see being a challenge, although I could see it cutting either way, depending especially on how idiosyncratic or esoteric your tool-use context is. For example, if you're doing calls into some obscure freight management system, you might need to have potentially even the ability to fine-tune that model to get those calls to work quite right.

So, broadly speaking, how big of a deal is the back-end interaction of tool use, and what's working or not working there today from what you've seen?

Anish Acharya

It's a good question. I think what we're broadly seeing is that you've got to build a lot more product than just the voice capability. I think the voice capability alone is insufficient, and there may be more room to explore voice-only in areas like agents, where there are different types of conversations, different outputs that you want, and different considerations around price point and what APIs you want to consume for what fidelity of output.

But even then, there's an enormous amount of integrations and workflows you probably have to deliver to build a traditional moat. So, without commenting on a specific use case in something like freight, our general observation is that the capability gets you in the conversation but isn't sufficient to get you to the other side.

Olivia Moore

Yeah, I would agree. Especially when you think about a lot of these companies that voice agents are selling into, they're very much traditional enterprises. They're not the Apples or the Googles of the world. For them to build or launch in a more horizontal way, or even to try to build a voice agent themselves, in many ways it's a miracle if they can do it once, let alone keep it updated as the models get better and as there are new options that are going to be a better experience for the customer.

When the integration breaks with the back-end system of record, what do they do? I think that is why we have seen so much excitement on the customer side for more vertically focused platforms that, to your point, are maybe both fine-tuned for the types of conversations that these customers are having, but also have done the work to build out the long tail of integrations and the long tail of, honestly, just conversation types—how do you manage piecing together different tools for different types of tasks that you need to complete?

Nathan Labenz

Yeah, I always think about Tyler Cowen saying, “Context is that which is scarce,” and that has never been more true than in the AI era. I think he said that before the AI era, but it's never been more true. So much of what I see standing in the way between businesses that want to use AI and actual successful use is literally just assembling the context, and sometimes getting the context out of their heads and into some documented form that the AI can process.

It's not too surprising that you would see base models, even as they are quite powerful and extremely versatile, certainly relative to anything that came before them, not know the intricacies of not just the freight business in general, which they might even know, but the way you handle your freight business. That stuff is the last mile; it trips a lot of people up.

Do you have any observations or synthesis of what's going on there? I honestly still struggle, and I do some amount of—not a lot, but enough—hands-on consulting with businesses where I've seen this repeatedly: people have a really hard time assembling the content. Maybe this is what the verticalized startups are going to solve for us, but do you have a sense of what's going on there? It seems like it should be easier.

We should be making faster progress in AI adoption than I feel like we're actually seeing.

Olivia Moore

Yeah, it's funny. I think we've definitely seen an explosion of AI budgets within enterprises and within end customers.

Anish Acharya

And in some cases, this was especially true 6 months ago. It’s less true now, which I think is good, but they’re kind of looking for things to buy because, to your point, they’re not spending all day thinking about what’s latest in AI, so they’re not exactly sure how to use it in their day-to-day. I think we even saw this with ChatGPT, which launched with a massive splash—the fastest product ever to 100 million users.

But then people weren’t really sure how to use it every day, and the usage was flat for basically a year. Only now, in the past year, with more models and more obvious things to do with it, has the usage kind of picked back up. So I think it’s exactly to that point: on both the consumer side and the business side, if people don’t know what to do with the product, and if it’s more than a few steps to get up and running, there’s going to be a decreasing funnel, unfortunately, of people who make it through and are actually becoming paying customers. So that’s part of the reason that I think we’re seeing companies that are so vertically focused see the most success here.

Nathan Labenz

Totally. And actually, I wouldn’t underestimate the amount of growth these voice AI companies are seeing, especially on the agent side, but also on the scribe side. The businesses either don’t know how to use it or do know how to use it, and it’s growing explosively within these businesses because, for the businesses that can embrace it, it’s such a straightforward substitution for the humans that make the phone calls. Of course, they need to think, “Okay, we’re going to hold it to some sort of CSAT score, and we’re going to measure outcomes and things like negotiations.” So, of course, there are guardrails and integrations that need to be done, but in the cases where it’s working, it’s really, really working. We’re seeing some of the fastest-growing B2B startups we’ve seen in 10 years.

Perfect setup for a little lightning round on what’s really working across different corners of the economy. I was going to start with enterprise, actually. You mentioned scribes. It seems like that’s a pretty well-established use case at this point. We’re starting to see that bleed over into real-time coaching on calls and then, obviously, full-on agent substitution. Is the real-time coaching working? What do we know about that? And is the substitution actually happening to the point where you think we’ll see labor-market statistic effects this year, or do you think this will continue to be isolated examples that are the exception rather than the norm for a while still?

Anish Acharya

Yeah, these are great questions. I would say that coaching is definitely working, and it is an interesting transition point in that there are some jobs—for example, a call-center job—where, if you’re an AI product selling a coach to call-center workers, there’s a massive amount of demand for that right now. But what we might imagine is that in 5 years, or probably less, the AI agent is going to replace a lot of those workers.

I think where the coaching will really continue to exist is in these jobs that have a heavy in-real-life component or a heavy personal component. We’ve seen quite a few AI real-time coaches for salespeople, for example, and for HVAC technicians. Many of these jobs, whether or not you get the $10,000 upsell, comes down to the nuance of what you say or the question that you ask. So even if you’re paying hundreds of dollars a month as an individual user for an AI coach there, it’s absolutely worth it.

In terms of the economic impact of voice agents, I think, to be honest, there are some cases—a basic call center, for example—where an AI taking the calls will free up the human workers to do much better and more rewarding jobs. These are massively high-turnover, 300%-per-year, thankless jobs in many cases. And I think there are better things for people to be doing.

Then in other cases, recruiting is one example. We’ve seen quite a few voice agents that conduct initial screening calls for human recruiters. That means they can spend those 20 extra hours a week with the 5 candidates they’re really excited about, catering to them and really convincing them to engage in the process and take the job. So in many cases, I think we see it as amplifying the humans in their current roles versus replacing them.

Olivia Moore

Yeah, I think all of that is exactly right. Potentially more humans sort of move up the stack to do higher-value work. I think the other thing is that you could take a look at what’s happening and say, at the limit, we have 20% fewer jobs, or you could say, at the limit, we all work 4 days a week and we’re paid to be optimists. But I believe that there’s an opportunity for people to be more specialized and do more of the work that matters, and less of the administrative overhead stuff that seems to consume most of our day-to-day.

Nathan Labenz

Yeah, I’m with that. I guess what I’m trying to really zero in on, though, is that I think it’s true in some ways, what you’re saying, but then there are other ways in which I think retraining and reskilling—and everybody was going to become a programmer, and that really hasn’t happened. When we look at call centers specifically and the people that work there, we may free up resources and the company may grow and invest in other ways, but I think in many cases the people that have the call-center jobs are not going to be moved into other jobs at that same company. Instead, the AI is just going to do that job, and they’ll have a much lower headcount in the call-center operation.

They may reinvest. There may be more R&D. There could be all sorts of great things. I also think people should work less, and a new social contract that embraces that is high on my list of things people should be developing now. But, leaving aside the second-order effects of what happens, are we just at a point technologically where we could see, if enterprises wanted to do it, a 90% headcount reduction in call centers?

And it seems like the call-center thing, if I’m understanding your answer correctly, is probably at least another year before we’d see that sort of effect. Of course, these things are not binary either, but I put that 90% number out there just to say “order of magnitude,” even if it doesn’t go to zero humans in the call center. It sounds like you think that’s at least a 2026-plus phenomenon.

Anish Acharya

I mean, I don’t think so. We’ll see. So far, we’re not seeing it because, as Olivia said, nobody’s job at the call center is to just do initial phone screens. People have recruiting jobs, which involve initial phone screens that are annoying and can be overwhelming, as well as deeper interviews, salary negotiations, and ensuring that employees are successful once they’re onboarded.

So, yes, the AI is going after the initial phone screen, but we haven’t seen a reduction in headcount because all of the other work is so important. Frankly, in many cases, they don’t yet trust the AI to do the work, or the AI can’t do the work. The AI can’t take your employee to a baseball game a month after they’ve started and make sure they’re having a fantastic experience. So I totally understand the conceptual argument; it’s just not something we’ve seen yet.

I think we’ll also see the success of AI open up probably a new type of job that we haven’t imagined before for humans to do. One great example is that one of the fastest-growing jobs right now is basically contributing training data, doing online tasks, and doing other things that help the AI. It might be similarly paid, but it could offer a much better lifestyle than a call-center job in many ways. So it’ll be interesting to see what new types of opportunities open up for humans in the AI era.

Nathan Labenz

Certainly, the AIs are less abusive than the human callers. I think we can safely say that.

Anish Acharya

Yeah. Yes.

Nathan Labenz

Yeah, for sure. I’m not—to be clear on my perspective on this—I’m not anti-displacement or trying to stoke fear about that. I think we probably are going to see it and should get ready for it, and ideally it would be a good thing.

I often ask people outside of the Silicon Valley bubble—and I live in Detroit, Michigan—if you didn’t have to work the job you have to make the money that you need for the rest of your life, would you still work that job? The overwhelming answer is no. I consider myself very fortunate that I would probably continue to do what I’m doing even if I weren’t getting paid for it or didn’t need to get paid for it.

But I think it’s really important for the Silicon Valley set to keep in mind that most jobs are not jobs people are doing for the joy of the job. If they could have their needs met in other ways, they would happily take that trade. So I’m not a job preserver; I’m just trying to figure out at what point this wave of disruption is actually going to hit. How much time do we have to get ready for it?

Anish Acharya

Yeah. Yeah. Look, I think it's cold to tell everybody, "Hey, just go learn programming." That's not what I'm saying at all. I just think it's very hard to understand what the labor impact of these technologies will be. And I think it's easy to hypothesize about a world in which all the jobs just go away. But it's not what we're seeing yet. So even if the technology is 18 months away, I don't know that the labor market will change in the way that we're perhaps imagining as a result of the technology. I think we'll have to see.

I think a broad question, though, that you're speaking to is: What does it mean for our society when we have all this abundance, and is there kind of a lack of purpose? I have a big theory that people need purpose, and if they don't have enough of it, they create it. Sometimes it gets pointed in bad directions. That's a lot of why I think Google ends up not working culturally.

I think there's a lot of brilliant people there, but in a sense, it's a low-stakes environment from a purpose perspective because the business is almost too good. I do get nervous about mirroring that in society. So I think an even more interesting question to me is: In a world where we actually just do have all of this abundance, how do we think about ensuring that people have purpose versus ensuring that people have jobs and income and all those other necessities?

Nathan Labenz

Yeah, I'm a little more, maybe even optimistic on that dimension, but I certainly filed that under a good problem to have. Okay, so this is supposed to be the lightning round, so we'll go through these next ones faster.

SMBs: You've got the call answering, which seems pretty straightforward. Any highlights or, for SMB owners out there, where should they go to potentially get the best AI call answering today?

Anish Acharya

Yeah, I would say something we've been very excited about for SMBs is that even for SMBs, there's a vertical solution depending on what you're doing. So if you're a restaurant operator, a spa, or in home services, we put all of these in our market map, but there is a solution catered to you, which is fantastic.

This actually gets back to what we were talking about before, which is that SMBs typically have 1 or 2 people doing nothing but answering phone calls, which is incredibly expensive for a small business. When we talk to the SMB customers, when they switch over to a voice agent, they are not laying off that human, who is usually a core part of their business.

That person is now able to spend their time doing things that are much better for the customer experience, growing the business further, or pursuing other extensions of the business, which are really powerful and exciting.

Nathan Labenz

Okay, how about creators? This one is maybe a little less interactive, although maybe you're seeing interactive experiences that are sort of in the creator economy. But one question I had, because I might actually do an episode powered by AI voice in the not-too-distant future, is: Who has the best voice design today?

If I want to create somebody who's going to give me that sort of film-noir kind of read, where should I go for that sort of thing?

Olivia Moore

I think there are different answers to this question. One, there are platforms like ElevenLabs, where you can clone your voice. ElevenLabs also has a great tool where you can describe a voice, describe a sound, and have it created.

The other end of the creator part of AI voice, to me, is these digital clones, which we're seeing more and more of on platforms like Delphi, where you can essentially launch a version of yourself that your audience can interact with in your voice, or maybe via text or other modalities, which is fascinating.

I haven't seen AI replacing full podcast episodes for any podcasters yet. I think hypothetically we could get there, where maybe you just prompt the questions that you would ask, but we're probably still a couple of years away from that. It's worth, though, playing with HeyGen, Captions, and a bunch of these other products.

You can fine-tune a model of yourself, video and audio, and then give it a script, because I think there is a world in which you do this entire podcast without ever owning a video camera or microphone.

Nathan Labenz

Yeah. We did an episode with HeyGen, actually, and for some reason Josh's audio wasn't great, so we then redid his entire side of the conversation with his avatar from HeyGen. It's been 6 months.

It was pretty good. I mean, it was not quite as good as the original would have been, but it was fitting that it happened on that episode.

How about for kids? I've been playing classic Nintendo games with my kids recently, and I often turn on Advanced Voice Mode while playing Super Mario 64, the old open-world game, because I don't know where to go. Where's the star? What do I have to do?

I'll ask Advanced Voice Mode, “All right, this is the level I'm playing. Where do I go?” My kids are now to the point where they're like, “Daddy, ask AI any time that either I'm slow to do something or don't know what to do. So, Daddy, ask AI.”

That's cool. I would love to have something interactive and educational for my kids, but I'm also like, yikes. I don't necessarily want to trust anyone to implement AI effectively for my kids. So any winners, any early winners in that space?

Olivia Moore

I think we've talked a lot about this as a team. An area of exploration we're fascinated by is all the stuff around behavioral, social, and emotional development for kids. I think that is an area where AI is very naturally suited to deliver value, and there's just not a lot of technology there.

A great example is that my son loves to play Minecraft, but all of the other people he meets online are toxic teenagers. So why isn't there a companion that can play Minecraft with him and model positive social behavior?

Another example is observing the classroom. If your child goes to one of those great schools where they've got 2 teachers in every classroom—1 doing all the academic components and 1 doing all the social-emotional components—you already benefit from this. But a lot of kids don't go to schools like that.

So having a vision model, a multimodal model, that can observe the children interacting and give feedback to parents and teachers could be really valuable. I think there's a ton there. Of course, there's assignment generation, quiz generation, and helping kids learn in whatever ways are best suited to them.

That stuff will happen, and it'll be super important, but I think pairing it with all the emotional opportunities is where we get most excited.

Nathan Labenz

Yeah, I was going to say I feel like we've seen companies like Synthesis, Ello, and Super Teacher that are kind of like, what if every kid had a reading tutor or a math tutor sitting next to them all day who could understand how they learn best and cater to them?

Then, on maybe the other end of the spectrum, we've seen things like the Curio toys, which is: What if, maybe even more importantly than the tutor use cases for many kids, they had just a friend, a mentor, or a coach who was with them every day and could both track their progress and help get them on the right track—or even just be a completely sympathetic listening ear?

As a side note—sorry, I know we're in the lightning round—the companion stuff is so cool. There are, of course, so many amazing companies like Character.AI and many others, and a bunch around the top 100 and top 50, but it still feels like we're in the first inning, maybe the warm-up or something, of exploring this space.

There are so many contextual opportunities to do it.

Olivia Moore

And look, I think one aspect of it may be completely sympathetic. Another aspect of companions might be not that sympathetic—one that really challenges you, pushes you, and disagrees with you.

Even something as simple as that: We always joke and call it East Coast mode, like your companion that's a little more terse doesn't exist. Why doesn't it? I don't know. I think we're going to get to see those products in the next 2 years.

Nathan Labenz

Yeah, that's really interesting. All right, in the interest of time, we'll skip over legal and medical. Maybe I can just ask for a word from Munjal from Hippocratic AI, because I would love to talk to him about what he's doing in the medical space.

But how about companionship? Maybe the last one in the lightning round is just: It seems like the farthest-out edge right now is going beyond companionship and into relationship, and even into not-safe-for-work types of things.

I don't know if that's stuff that you guys would touch in investment terms, but I trust that you're at least scouting that territory somewhat. What do you see going on on the far fringes of romance with AIs today?

Anish Acharya

I mean, it's a good question. I think the thing that's surprising everybody—or at least, I had an assumption that most of the companion use cases would be frisky young dudes, and it actually hasn't been that. A lot of it has been an audience that caters a lot more to women and probably feels more like interactive fiction than it does what you might consider pornography.

So, 1, I think there are a lot of mistaken assumptions that even I had about how the products would be used and who they were going to be used by. I also think there are a lot of definitions of romance, and I think people are perhaps critiquing these products, saying that they're a substitute for traditional romance, when in fact they may make us so much more capable.

You've either got somebody—an AI—that helps you train to be better at things like conversation and even flirting, or an AI that can just be a vent for a lot of the frustration and emotional weight that people can sometimes bring to their in-person relationships.

Anish Acharya

So those are some of the more surprising things. What do you think, Olivia?

Olivia Moore

No, no, I agree. It's funny: whenever we pull the top 50 or top 100 list of AI apps and send it around to our team, every time without fail, people are like, “Oh my gosh, there's a ton of companion platforms on here,” and a lot of them are maybe more NSFW-oriented. But I think it's been exactly what Anish has said, in that it's actually much more of the AI boyfriend than the AI girlfriend use case, interestingly, and a lot more like interactive fan fiction than anything else.

But that's a part of the human experience, right? Sexuality is a part of the human experience, and we can't pretend it doesn't exist. When we do, you end up as Apple, which can't release a product for 5 years. So I think that we have to embrace that this is going to be a part of these products and just find ways to get behind that. Of course, there's always going to be products at the fringes that we'll never invest in and perhaps most people will never use, but those are almost the least interesting products to talk about because it's always been that way.

Nathan Labenz

Yeah. I've done 2 episodes, actually, with Eugenia from Replika. The more recent one was reviewing some research that folks at Stanford did that showed that not only did Replika reduce suicidal ideation in a substantial way for people who had that issue coming in, but also that, more often than not, it helped people get out into the world. People indicated that using Replika was not holding them back, but in fact encouraging them to get out into the real world. I questioned some of the data—some of it was self-reported—but I thought it was quite interesting.

I think it's amazing. I'm a big fan. I think, as with all this, as with all technology, but maybe even more so with AI, it feels to me like the shape of this, and the specifics of the shape of exactly what we build, is going to be really important. I have no doubt that you could make a predatory romantic AI that is addictive and exploitative in all sorts of ways, but I think we do see at least some existence proof that you can make really positive, or at least predominantly majority-positive, versions of these things.

And that brings me to a question on rules of the road. I know that it is early in this space. One rule that's been proposed is that AI must disclose that it's AI. That's a Yuval Noah Harari one that I like for its simplicity. I've also been thinking about the idea recently of a “do not clone” registry, which would be the modern version of “do not call.” You could go and say, “Here's my likeness and my voice; don't clone me on other AI platforms.” I'm wondering if you guys have any ideas for either emerging best practices or possible regulations that might keep all of this on the good side as much as possible for us.

Olivia Moore

Yeah, it's definitely early, but at least on my side, I've been surprised. It seems like more people these days are frustrated by the large model companies taking the approach of, “We're not going to let you do something,” versus people being frustrated that, “I'm being deepfaked or I'm being cloned.” That's not really happening to the average consumer right now. I think we've seen both the biggest startups and the biggest model companies be extremely careful about allowing you to do anything related to a public figure, let alone personal pictures or other things like that.

And so I personally am very intrigued by the idea of the directory or the registry, especially because it opens up this opportunity for people to license or allow their identity to be used for use cases that they are excited about. We're seeing platforms like ElevenLabs; they have these iconic voice collections of celebrities or people who will allow their voices to be used. But it's also been a massive boon to this industry of voice-over artists who maybe historically couldn't get a job in Hollywood. Now there are all of these voice-over jobs on ElevenLabs.

We could see something similar in the influencer-creator economy, where if you're an influencer with 5,000 followers, you're going to have a hard time getting a big brand to respond to you. But if the AI avatar version of yourself is even better, more powerful, and more extensible, then maybe you actually can get some of those big deals. So I'm really interested to see how people can extend themselves using the AI tools, versus—I at least have seen less to be concerned about—the everyday person who isn't a public figure getting deepfaked or anything like that.

Anish Acharya

I totally agree. I think every time there's a new technology, the talking heads try to get overly paternalistic. I just don't think that's a generous enough view of the average consumer, how smart they are, and how media-literate they are. Of course, every technology has the potential for misuse, so I'm not being glib about that. But I do think the paternalism is unwelcome and often unnecessary, because people have learned for 30 or 40 years that just because it's written in a book, it's not true; just because it's on the internet, it's not true; and just because it's on social media, it's not true. There's no reason that this technology will be any different. Whatever we do here, I hope that we're generous in our assumption that consumers are smart and savvy and will know how to use the products and technologies with the appropriate level of caution.

Nathan Labenz

Maybe just give me your medium- or long-term vision for where this voice-enabled computing is going to go. Are we all going to be walking around with AI in our earpiece and untethered from our devices? Maybe we've got glasses that pair with that. What's the tech-optimist view of life in this voice-enabled-computing future?

Olivia Moore

I think that we will see voice unlocked as a kind of modality feature on every product and in every interaction, in every device: AirPods, glasses, your computer. As we've dove into voice, especially from a consumer use case, you find that there are a lot of situations where maybe you don't actually want to be having a 2-way conversation, or you can't be having a 2-way conversation. You want it to be transcribing what you say, or vice versa: you can't talk and you want it to be talking back to you.

And so I think right now we're in inning 1 of AI voice, where we have a set of really compelling and exciting products, but 5 years from now they're going to look incredibly limited based on what we have then, where you can interact with voice in any way, at any time, for whatever is most useful and helpful to you.

Anish Acharya

Steve Jobs famously said that a computer is a bicycle for your mind. That meant that a computer extended us intellectually in ways that were unimaginable, and that's what technology has done for us for 40 years. I think we're now going to have the emotional version of that, the sort of emotional bicycle, where it extends us emotionally through products like companionship, but many, many more. And I think voice is going to be the primary catalyst and interface to that. So maybe that's a subject for our next conversation, but I think that's really the way it's going to impact us, and it's been a bit underestimated.

Nathan Labenz

Cool. I love it. Olivia Moore and Anish Acharya, thank you both for being part of The Cognitive Revolution.

Olivia Moore

Thank you.

Nathan Labenz

It is both energizing and enlightening to hear why people listen and learn what they value about the show. So please don't hesitate to reach out via email at tcr@turpentine.co or you can DM me on the social media platform of your choice.

a16z on AI Voices: Call Centers, Coaches, and Companions with Olivia Moore & Anish Acharya | BidClub