a16z谈AI语音:呼叫中心、教练与陪伴者——与Olivia Moore、Anish Acharya对谈
AI语音最先形成持久商业切口的,将是垂直B2B,而不是独立消费应用。 企业本来就会付钱请人接电话,因此夜间覆盖和原本无人接听的电话都是很自然的切入口;最强的初创公司随后会扩展到工作流主导权。Labenz称,其中一些属于“过去10年里我们见过增长最快的B2B初创公司”。
对话式AI的可用性基本已经到位,但拟人感仍不只取决于转写准确率。 延迟如今通常低于半秒,而Sesame展示了停顿、“嗯”等填充词和声音抑扬如何把一副打磨过的合成声音变成“足以被误认为真人”的声音。情绪适配、多方轮流发言和可打断性仍是明显短板。
Happy Robot说明,专业化的对话质量可以打开更高价值的工作。 它的智能体会向货运经纪商披露自己是AI,并与卡车司机交朋友、唱反调和谈判;其中一种策略是先插入5秒钟的“让我去和主管谈谈”,再回来让步一点,Moore认为这一做法可能带来高得多的接受率,但她对此并不确定。值得投资的启示是,更好的表达方式能赢得处理说服和定价的许可,而商品化语音做不到这一点。
应用层的护城河在于围绕语音搭建的垂直系统,而不只是语音生成本身。 企业需要集成、客户专属上下文、评测、护栏,以及系统记录失效时的故障恢复——“这种能力能让你进入对话,但不足以把你带到另一边”。这使垂直平台比横向智能体更占优势,即使底层模型仍在进步。
Apple推迟Siri大改版体现的是 incumbents 的劣势,而不是技术进展不足。 Acharya称Siri“每天五次都像往眼睛里戳一根棍子”,并认为大公司很难拥抱AI混乱而具有人性的部分;Moore补充,Apple必须安全地向数亿用户交付产品,不像初创公司只服务于主动报名的测试者。Labenz还把Google未能在ChatGPT与Deep Research绑定之前将其商业化,视为类似的错失机会。
语音自动化正在带来教练和任务替代,但还没有达到Labenz设想的呼叫中心90%裁员。 当实时辅导能够影响一笔1万美元的暖通空调增购时,每月数百美元的费用就可能合理;招聘智能体则能为招聘人员每周释放约20小时,让其服务5名优先候选人。尽管呼叫中心年人员流失率达到300%,嘉宾们尚未观察到数量级的岗位损失,Acharya也不愿自信地把18个月的技术时间线直接换算成劳动力市场的时间线。
消费级前沿正从老年人和儿童延伸到越来越个性化的陪伴者。 多模态语音可以为老年人提供耐心的技术帮助,为儿童提供辅导老师或具有积极社交属性的Minecraft伙伴,也可以打造从共情倾听者到会挑战用户的“东海岸模式”人格不等的陪伴者。Acharya更长期的框架是“情感自行车”:正如计算机扩展了人的智力,它也将扩展人的情感能力。
安全政策必须在已被证明的冒充风险与授权身份的市场需求之间取得平衡。 Labenz称,在他披露问题一年后,仍有两个呼叫平台允许他克隆Donald Trump并进行规模化呼叫;Moore则表示,当下人们对模型限制的挫败感,可能高于对自己被克隆的担忧。她对“禁止克隆登记”的改造具有经济建设性:人们可以禁止冒充,同时明确授权自己的声音或头像用于获批场景。
1. 需求侧行为往往是最清晰的产品信号
Moore的项目观察漏斗覆盖创始人密集的Twitter、newsletter和线下聚会,也包括Instagram、TikTok,尤其是YouTube——后者是“全球排名第1的移动应用、第2的网站”。对许多消费级和专业消费者AI产品而言,YouTube教程是社交推荐的最大来源。
团队会观察普通用户——“说实话,通常是十几岁的女孩”——如何把ChatGPT强行变成治疗师、朋友或教练。Moore的框架是:“消费端太随机、太神奇,所以我们尽量让数据说话”,因为如果产品错过了时机、上手流程或某个决定性功能,再好的出身也救不了它。
Moore不那么光鲜的优势,就是亲自使用这些产品:Operator、Deep Research、DeepSeek、o1 Pro和Krea都是许多自称业内人士的人尚未尝试过的工具。她用一个更新版的社交应用老笑话来测试需求:“每个大语言模型都在被折磨成治疗师。”
2. 语音打开了屏幕从未覆盖的技术界面
Acharya从第一性原理出发:语音介入了大多数人与人之间的关系,但技术过去没有处理这一界面的基础设施。与经历了数十年产品试验的计算界面不同,语音仍是“一张完全空白的纸”,同时创造了产品和分发机会。
Moore认为B2B智能体的进展更快,因为即使是小企业也会付钱请1到3个人接电话。当模型接近人类水平后,用智能体接听夜间电话或原本会转入语音信箱的电话,就成了显而易见的替代方案;许多消费者可能已经在不知情的情况下遇到过这类智能体。
消费端的接触则主要通过ChatGPT、Grok和病毒式传播的Sesame演示发生。Moore认为,1-800-ChatGPT暗示着许多人第一次真正有意义的AI体验——无论这个产品本身是否成功——可能会通过电话到来,而不是通过一个新应用。
3. 老年人说明多模态与对话同样重要
Moore那位90多岁的母亲已经会让Alexa播放音乐,因此Alexa+可能成为进入对话式计算的桥梁。语音也可以解锁她从未学会操作的老技术,包括电子邮件界面和电视遥控器。
Labenz有一条屡试不爽的技术支持指令:“从上到下,把屏幕上的所有内容读一遍。”这条指令已经可以作为纯文本GPT-4提示词发挥作用。一个有耐心的模型,能够为那些没有“无限耐心”的亲属的人复现这种帮助。
Acharya提到Google在“我想应该是去年年底,12月”发布了能看见屏幕内容并实时交互的Gemini模型;OpenAI也有类似能力。他更尖锐的观点是,关键上下文往往存在于物理世界:把手机对准遥控器,比把每个按键逐一用语言解释给Nathan或Google更自然。
4. 低于半秒的延迟解决了入门级语音,却没有解决对话
Acharya认为“解决了”说得太满,但基础延迟和可理解性在过去一年里已经接近解决。如今大多数模型的响应时间都低于半秒,这决定了系统只是交换音频,还是已经能够维持一场对话。
Sesame的突破在于保留了表面上的不完美:额外停顿、“嗯”等填充词,以及传统系统可能视为错误的表达性语调。这些细节让输出超越了更好的Alexa或Siri,趋近于“足以被误认为真人”的声音。
情感表达仍未完成。创始人希望智能体能识别内容并调整语气:面对令人兴奋的消息时更明快,面对悲伤时更低沉、更缓慢;可打断性仍然笨拙,尤其是在多人场景中。Moore指出,人类自己也还没有完全解决两个人同时开口时该怎么办。
Labenz区分了语音模型和对话模型:人类通过细微的听觉和视觉信号协调轮流发言,而且往往在还不知道完整句子之前就已经开口。一个早已知道全部回答的AI会让人觉得不自然,这说明原生的对话行为需要的不只是“语音转文字—LLM—文字转语音”的流水线。
5. 专业化表达赢得谈判资格
a16z被投公司Happy Robot为货运经纪商提供语音AI。Acharya认为,它更深层的技术工作让对话明显优于商品化替代方案,从而让客户有信心把“说服、谈判、反驳”交给它,而不只是让它检索信息。
Moore最突出的例子,是智能体刻意增加延迟。它会说:“等一下,让我去和主管谈谈。”停约5秒后,再回来给出一个略高一点的价格,尽管它早已知道自己的授权区间;人在主观上会觉得自己争取到了让步,Moore认为最终报价的接受率会高得多。
Labenz的反应是:“我不知道该怎么看这件事。”这保留了有效设计与操纵之间的张力。智能体会披露自己是AI,但卡车司机仍会进入熟悉的对话节奏;理性上的知情,并不能抵消谈判仪式训练出来的“爬行动物大脑”。
更广泛的判断是,AI可以“比人更像人”:同一个友好的智能体每次都能接待,耐心倾听,并且可以把全部时间都花在来电者身上。即便模型的推理能力还没有超越人类,超人的耐心以及几乎为零的等待时间也已经具备运营价值。
6. incumbents 同时面临文化和分发障碍
Apple宣布重大Siri更新要等到2027年,与快速进步的初创公司放在一起看显得荒谬。Acharya认为,日常Siri糟糕的表现与Apple Intelligence的广告之间的落差正在侵蚀信任:“每天五次都像往眼睛里戳一根棍子。”
他的诊断是制度性的:AI最适合在人类互动混乱的表面上工作,而 incumbents 的组织方式却是从技术中消除人性和不可预测性。委员会、律师,以及试图“阉割AI”的努力,构成了他所说的“非常困难的精神问题”。
Moore提供了另一面:Apple必须向数亿不同年龄、不同使用场景的用户交付自然且正确的产品。初创公司可以把不完美的测试版交给自愿尝鲜的早期用户;她猜测,生成式通知摘要引发的公众反应可能吓到了Apple。
Google Labs通过NotebookLM等受限实验提供了部分范本,但商业化仍然缓慢。Labenz最尖锐的例子是Deep Research:它最初是Gemini产品,本应由Google占据主导,却最终让ChatGPT因这一能力而闻名——“incumbents一次又一次错失机会”。
7. 垂直上下文和集成才是真正的企业护城河
原生端到端语音模型仍处于早期,而且成本略高;Acharya认为Gemini Flash或许是当前最好的选择,但也指出其可打断性仍然落后。他一贯的保留意见是:“模型现在正处于它们未来任何时点都不会更差的状态”,因此今天的多阶段技术栈一年后可能就会显得过时。
Acharya把推理模型视为一种独立原语,尽管它们拥有熟悉的界面。概率式的语言行为有助于建立友谊和融洽感,而定价等事实性决策要求准确;通过编排推理模型和语言模型,可以把对话的不同部分分配给合适的系统。
工具调用意味着语音生成只是第一层。正如Acharya所说:“这种能力能让你进入对话,但不足以把你带到另一边”;持久的产品需要工作流、集成、结果衡量、护栏,以及对后端故障长尾情况的处理能力。
Moore称,这一负担对传统企业尤其严重:它们可能连一次性搭建都很困难,更不用说持续让模型与系统记录保持同步。她强调,专注垂直领域的平台可以处理行业对话类型和集成的长尾;Labenz则把客户专属上下文称为稀缺的“最后一公里”,这是基础模型无法从通用知识中推断出来的。
8. 辅导已经奏效,大规模替代仍未得到证明
Acharya称,实时辅导已经在呼叫中心员工、销售人员和暖通空调技师中发挥作用。如果一个细微问题决定了一笔1万美元的增购,那么每月数百美元的个人教练费用就能带来显而易见的回报;涉及体力劳动或深度个人关系的工作尤其适合增强。
自动化也在移除不受欢迎的任务。呼叫中心的年人员流失率可能达到约300%,而招聘智能体可以完成初筛,每周为招聘人员释放约20小时,让其投入到5名优先候选人的说服、沟通和照护中。
Labenz不接受简单的再培训叙事:被替代的呼叫中心员工未必会在同一家雇主那里转入更高价值的工作。他追问这项技术是否可能支持90%的人员削减,并认为这种影响或许要到2026年及以后才会出现,同时承认问题关注的是数量级,而不是一个完全没有人的呼叫中心。
Acharya从经验出发反驳:他们尚未看到这样的削减,因为一份工作往往同时包含筛选、面试、薪资谈判、入职,甚至陪员工去看棒球比赛。“即使技术还有18个月才到位,”他说,它对劳动力市场的影响仍很难预测;当供给变得充裕,问题可能会从工作转向意义。
9. 消费级语音从中小企业前台扩展到情感基础设施
对餐厅、水疗店和家政服务公司,Acharya建议采用垂直而非通用智能体。老板通常不会解雇原本接电话的核心员工,而是把他们转向客户体验和增长,让智能体处理漏接电话和常规来电。
创作者工具已经覆盖ElevenLabs的声音克隆和描述式声音生成、可能类似Delphi的互动数字分身,以及HeyGen头像。Moore认为,完全合成的播客主持可能还要“几年时间”,但她可以想象这样一个世界:一档脚本化节目既不需要摄像机,也不需要麦克风。
Labenz列举了Synthesis、Ello和Super Teacher等个性化辅导案例;Moore最看重的例子,则是一个积极的AI Minecraft伙伴取代“有毒的青少年”,以及一个多模态课堂观察者,为资源不足的学校中的家长和教师提供原本无法获得的反馈。
陪伴者需求也打破了原有假设:嘉宾看到的AI男友和互动小说行为,往往更多来自女性,而不是由“精力旺盛的年轻男性”主导的市场。陪伴者可以补充现实关系,教授对话或调情,承接情绪压力,也可以挑战用户而不是一味奉承——也就是开玩笑称作的“东海岸模式”。
10. 身份授权可以把保护与经济机会结合起来
Labenz引用了斯坦福大学在此前一集Replika节目中讨论的研究:用户报告称自杀意念减少,而且更多时候与外部世界的互动增加,但他质疑其中部分数据,并指出一些结果来自自我报告。他的结论是有条件的:有益的陪伴者确实存在,但掠夺性和成瘾性的版本同样可以被制造出来。
他的红队测试让冒充风险变得具体:在他报告漏洞一年后,两家相对知名的呼叫平台仍允许快速克隆Donald Trump的声音并进行规模化外呼。因此,他主张尽早实施披露规则和禁止克隆登记。
Moore看到了登记制度对应的正向市场:人们可以授权自己身份的获批用途。ElevenLabs的声音收藏已经为配音艺术家创造机会;一个拥有5,000名粉丝的网红,也可能通过可扩展的头像实现变现,即使大型品牌会忽略这个真人创作者。
Acharya警告不要本能地采取家长式保护,认为消费者已经学会不去相信书籍、互联网或社交媒体上的一切。Moore的长期愿景是让语音遍布AirPods、眼镜、电脑和每一种产品——有时是双向交互,有时只是转写;Acharya则称其为一辆“情感自行车”,能够扩展人的情感能力。
Hello and welcome back to The Cognitive Revolution. Today, I'm speaking with Olivia Moore and Anish Acharya, partners at Andreessen Horowitz and fellow AI scouts who are constantly tracking emerging technologies and consumer behaviors, and who in recent months have really distinguished themselves as keen observers and eager early adopters of AI voice platforms and products.
In this conversation, we review recent developments in voice AI technology and explore how voice AI is already starting to transform business operations, user experiences, and human-computer interaction more broadly. On the technology side, we discuss how recent multimodal models have simplified the model stack and, with it, the application-development process; how reduced latency and improved interruptibility are enabling far more natural conversations than ever before; and how products including past guest Hume AI's Octave model, Google's interactive version of NotebookLM, and the very viral and incredibly natural-sounding Sesame are all beginning to demonstrate a remarkable level of emotional intelligence, both in their ability to understand the user and in the expressiveness with which they can communicate back.
On the impact side, we cover a bunch of really interesting real-world applications and trends, including the company Happy Robot, which uses voice AI to handle complex negotiations and build rapport with truckers in the context of freight brokerage; vertical solutions for restaurants and other types of SMBs; and how business owners are often employing these systems to handle after-hours and other calls that they can't answer themselves.
We also discuss how enterprises are using the technology to facilitate meetings and provide real-time coaching, why we haven't yet seen a major impact on call-center jobs and how soon that might happen, why Apple is now saying that Siri won't get a major update until 2027 even as so many other things are clearly starting to work, and, of course, the ongoing rise of AI companions for kids, seniors, and the lonely.
Along the way, we touch on a couple of philosophical questions relating to human-labor displacement and the delicate balance between protecting consumers and encouraging innovation. As someone who's been a career entrepreneur and genuinely loves using OpenAI's Advanced Voice Mode as, among other things, a biology tutor and a real-time video-game guide for me and my kids, I am super excited about this technology, but also legitimately concerned about just how scammy the world could quickly become if the models get even a little bit more lifelike.
As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, write a review on Apple Podcast or Spotify, or leave a comment on YouTube. We welcome your feedback, too, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network.
It had been a while since my last voice-AI red-teaming exercise, but with this conversation in mind, I did go back to 2 reasonably well-known AI calling-agent platforms to which I had previously reported flagrant vulnerabilities. I found that, sadly, a year later, both still allow me to quickly clone Donald Trump's voice and scalably call anyone and say anything in that voice, all with no meaningful controls.
This, to me, strongly suggests that we need new rules for these sorts of products sooner rather than later—both to protect the public and to protect the AI industry from its most careless developers. As it turned out, Olivia offered a brilliant twist on my “do not clone” registry idea, which could both expand economic opportunity and protect the public from AI impersonation, and which I feel it is truly time to build.
For now, I hope you enjoy this exploration of the rapidly evolving world of AI voice interactions with Olivia Moore and Anish Acharya, consumer-technology investors, AI scouts, and partners at Andreessen Horowitz. Olivia Moore and Anish Acharya from a16z are here to talk about the future of AI voice interactions. Welcome to The Cognitive Revolution.
Awesome. Thanks for having us. I'm excited.
You guys are part of a group that I affectionately refer to as AI scouts: people who are out there on the edges of what exists and exploring it in a lot of different ways. Before we get into the object-level stuff of what you've found and what you think is coming in the realm of voice, I'd love to get a little bit of the meta lessons or specific alpha tips for how you do such a good job of this. Where do you go for information? What's the top of funnel for you? How do you know you're onto something? I think people could learn a lot from your example in that respect.
Yeah. In many ways, it's our job to be chronically online and tracking every new thing that happens, especially as consumer investors. It's something that we've tried to hone over the last few years in particular.
It's so interesting because there's a pretty big delta, I think, between where AI scouts, AI experts, and early adopters are spending their time and talking, and where normal consumers are. So we try to be in both places. I would say Twitter, of course, is where most founders and AI founders are announcing new companies, new models, and new breakthroughs. AI newsletters have been massive, as have meetups.
But in terms of where real people are sharing what they do with AI, it's mostly on places like Instagram, TikTok, and YouTube. YouTube is a shocking one. YouTube is actually the number-one mobile app and the number-two website in the world. We find that for many companies in the consumer/prosumer space, if you look at traffic, by far their number-one referral source from social is YouTube.
There's this whole separate economy of YouTube influencers or YouTube creators who are making how-to content about using different AI tools. So I would say we try to track all those places.
In terms of what's an early signal to us, often we'll see normal people—usually teenage girls, to be frank—trying to manipulate ChatGPT into doing something, like being a therapist, a friend, or a coach. Once we see something like that, it's like, “Okay, the consumer pull is strong enough that there probably can and will be a couple of standalone, more focused products here.”
I think the other thing we try to do a lot of is just use the products. That sounds obvious, but it's surprising how few people seem to have actually tried Operator, Deep Research, DeepSeek, o1 Pro, or Krea. These are not super-obscure, long-tail products, and consumers find their way to them. But among the folks who are insiders and are paid to be doing this work, surprisingly few actually look at those products. It's just a great way to build your intuition.
Yeah, that's my number-one piece of advice too: just get hands-on. You can't really go wrong with that. The one interesting thing there was that it sounds like you're looking as much, or maybe even more, for demand-side pull as you are at the people coming forward with their technologies to offer. You're looking for people who are specifically trying to meet a need that maybe nobody's met yet and figuring out what that implies.
Yeah. Consumer is so random and magical that we try to let the data tell us. You can have the most tenured, pedigreed team in the world building a consumer app, and if it doesn't hit, it doesn't hit. That can be for a variety of different reasons. Maybe it's bad market timing. Maybe they got the product insight wrong but some specific feature right, or the product insight right but some specific feature wrong. Then no one completes the onboarding.
I would say we try to let the data, in terms of what people are actually using, tell the story to us. Sometimes looking at data such as people pulling ChatGPT into these off-label use cases will give us a little warning signal or a heads-up that, “Okay, this is something—this is a behavior that's working,” and so we should keep an eye out for products that are targeting this behavior.
A great example of this is the old joke in pre-AI that every social app would collapse into being a dating app. In the same way, every large language model is being tortured into being a therapist. It's a funny thing to say at a dinner party, but it's also a leading indicator of what consumers want these models to do and some of the things that we want to see in the future.
Cool. Well, let's talk about voice. To start, I would love to get the top highlights: What products and experiences have you seen that are the absolute best user experiences out there today? Hopefully, I've tried them, but we're about to find out.
Maybe I can frame it, and then Olivia can talk about some of the specific products. I think the thing to zoom all the way out on is grounding ourselves in the fact that voice mediates every human interaction and relationship, largely, right? Here we are, obviously, having voice mediate our relationship and our conversation and the way that we get to know each other.
It really is the original and most important form of human communication, but it's just been completely unaddressable by technology because we've never had the infrastructure. It's very interesting because so many of the other substrates that we're applying AI to are areas where we've also had a lot of historical technology exploration, whereas voice is just a complete blank piece of paper. That's why I think we're as excited about the product implications as we are about the distribution implications of this sort of technology surface.
Totally. Yeah. I would say there have been a couple of surprising things for us in terms of where voice is working now, at least on the startup side. A lot of the startups that are getting real traction in terms of net-new companies and products are actually more B2B-oriented, just because there are so many businesses that are now running off call centers or paying for 1, 2, or 3 people, even for small businesses, to answer the phone all day. Once you're at a point where voice models can be anywhere in the realm of human performance there, it kind of makes all the sense in the world to at least have the voice agent be doing your after-hours calls or your calls that would go to voicemail.
My guess would be that a lot of people have maybe interacted with an AI voice agent and not quite known it, because it's been a business calling them, or it's been the receptionist when they've called to schedule an appointment or something like that. On the consumer side, it's been so far maybe a little different than we expected. I think most consumers have interacted with AI voice through something like ChatGPT or Grok, which are incredible voice experiences. More recently, something like Sesame was a massive breakthrough, and that is still just a web demo, an early version of what's to come there.
My guess is, when we see the Sesame team open-source the model, and when we see models like that spread and become more accessible for app builders to build on top of, we'll see a corresponding explosion in consumer-focused, voice-first tools. I mean, a crazy thing that happened—and things are moving so quickly, I think it's easy to forget these things—is 1-800-ChatGPT. What was that? And I think, whether it failed or succeeded, it pointed to an important insight, which is that maybe the first way most people in the world will actually experience AI is via voice, both as consumers and as consumers consuming business offerings.
Yeah. I think of my dear mama all the time. She's now in her early 90s and lives alone, and is sharp but not an early adopter of new technologies. For her, I think it's going to be Alexa+ that's going to be the big transition. She already sits there and asks it to play music for her, but will she engage it in conversation? How natural will she find it? I don't know, but it's clearly going to be that form factor for her that could unlock a whole new set of things.
What's very funny is, ironically, Mama perhaps is now exploring new technology, but also isn't that familiar with old technology, which is maybe why she's calling you to get tech support and help using her existing products and devices. I think the potential applications for voice as applied to seniors are super interesting. We've discussed it a lot, and it's not just access to the new things; it's access to the old things as well that they just never developed the skills to interact with.
If she can figure out the TV remote, then we'll really be in business.
Yeah, funny enough, one of my first GPT-4 tests, going way back to the red-team days, was tech support for seniors. The prompt I found to work really well is exactly what I say to her when what you're describing happens—when she calls me and says, “My friend emailed me, and I can't find it.” I always tell her, “Read everything on the screen from the top to the bottom.”
She'll literally go, “Okay, Verizon, the time,” and then eventually we get down to where the issue is. That same thing basically worked out of the box with GPT-4. Of course, that was text-only at that time, but I do see that as a huge unlock for all sorts of different stuff that she struggles to access right now.
Oh, exactly. Well, I think most people don't have an infinitely patient person like you to walk them through how to do that. We even saw recently—I think it was late last year, in December—Google release the Gemini models that could see what was on your screen and interact with you in real time. It feels like we're right on the brink of models like that. OpenAI has one as well, becoming API-available and ready for builders to actually capitalize on. Once we see something like that become actually usable, it's going to be massive.
It's also so interesting because it points at something maybe Google and other search players should have done pre-AI: How do you take everything on the internet and apply it? The most important context is the context that's around me in my physical space. The idea of being able to point your phone at the remote, in the case of Nana, and be able to debug the problem that way, instead of trying to translate what you're seeing in the physical world line by line to either Nathan or to Google—it just, that interaction pattern doesn't make sense.
Hey, we'll continue our interview in a moment after a word from our sponsors.
Even if you think it's a bit overhyped, AI is suddenly everywhere. From self-driving cars to molecular medicine to business efficiency. If it's not in your industry yet, it's coming and fast. But AI needs a lot of speed and computing power. So, how do you compete without costs spiraling out of control? Time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure or OCI. OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all of your AI and machine learning workloads. OCI costs 50% less for compute and 80% less for networking. So you're saving a pile of money. Thousands of businesses have already upgraded to OCI, including Vodafone, Thompson Reuters, and Sunno AI. Right now, Oracle is offering to cut your current cloud bill in half if you move to OCI for new US customers with minimum financial commitment. Offer ends March 31st. See if your company qualifies for this special offer at oracle.com/cognitive. That's oracle.com/cognitive.
The Cognitive Revolution is brought to you by Shopify. I've known Shopify as the world's leading e-commerce platform for years, but it was only recently when I started a project with my friends at Quickly that I realized just how dominant Shopify really is. Quickly is an urgency marketing platform that's been running innovative, timelimited marketing activations for major brands for years now. We're working together to build an AI layer which will use generative AI to scale their service to longtail e-commerce businesses. And since Shopify has the largest market share, the most robust APIs, and the most thriving application ecosystem, we are building exclusively for the Shopify platform. So if you're building an e-commerce business, upgrade to Shopify, and you'll enjoy not only their marketleading checkout system, but also an increasingly robust library of cuttingedge AI apps like Quickly, many of which will be exclusive to Shopify on launch. Cognitive Revolution listeners can sign up for a $1 per month trial period at shopify.com/cognitive where Cognitive is all lowercase. Nobody does selling better than Shopify. So visit shopify.com/cognitive to upgrade your selling today. That's shopify.com/cognitive.
So one of the things I noticed in your presentation about this is that you said it's basically solved, I think was the phrase. I'm wondering what you see as remaining, if anything. The Sesame model might have even come out since that presentation, and it certainly takes another step forward in terms of the overall natural sound of the voice.
I find the interruption mechanics still a little bit gnarly, especially if there's a multi-way conversation. Sometimes I try to demo Advanced Voice Mode for people, just to bring them up to speed. My goal is to make it a quick way to bring people up to speed—I'm like, “This is where AI is at now if you haven't been paying attention.” But those demos often go a little bit sideways on me because the interruption mechanic is still a little weird.
If I'm doing it one-on-one, it's okay; I kind of know how to use it, and it seems to be optimized for that one-on-one. But in the group setting, it doesn't do super well in many cases. Anyway, that's just me identifying my remaining pain points, but what do you think are the hardest things to get right, or the most important things that still need to get solved?
Yeah. I think saying that things like basic latency and understandability have gotten solved is too strong, but they're very close to solved. Over the past year, those are the things that have gotten very close to solved, and that's the difference between being able to have a conversation and not have a conversation. In most cases, I think the models are getting those right now. Latency is less than half a second on most of the models, which feels very human-like.
I think the realization that a lot of people have had who've tried the Sesame demo is that, beyond latency, there are so many speech-pattern nuances that an actual human will say, which might actually sound like an error to the model: extra pauses, saying “um” or “you know,” and vocal inflections. When you add those things in, which the Sesame team has done, it goes from being a voice that sounds much better than Alexa and Siri, but is still kind of robotic in some ways, or so clearly AI, into something that could be mistaken for a human.
I think the remaining things for me—there's still a lot to do around emotionality. I talk to a lot of founders building voice agents who want the models to be able to understand what they're saying and vary the tone and the inflection based on that.
So if the AI voice agent is going to say something happy or exciting, the voice should reflect that. If it's going to say something sad, the voice should be lower in tone and pitch and a little bit slower. That's something that we still need to solve. And then interruptibility is huge.
I think of it as humans having also not solved interruptibility in conversations. We still have the issue where 2 people start talking at once and you have to be like, “No, no, you go ahead.” So we need a clever way for voice AI to be able to solve that in a way that humans maybe have not yet.
I think this is why we've gotten where we need to get to for voice models to work. But I think what Sesame showed us is that conversation models may be very different from voice models, or maybe an extension of voice models. Even the way the 3 of us are coordinating on who's going to talk next involves so much nonverbal communication, whether it's video or even audio.
The models have been trained really well to do speech-to-text and text-to-speech. Interestingly, fewer of the companies we're seeing are using native voice-to-voice models. So that's 1 opportunity, and then there's just generally understanding, as Olivia noted, some of the nuances of conversation and having that be natively programmed in.
For example, when I start talking, I'm not exactly sure what I'm going to say. It sort of comes—it like—and that's true for all of us, whereas the AI knows exactly what it's going to say, which can sometimes be a little bit creepy. It does not quite get out of the uncanny valley.
There's a recent paper from Meta that you're bringing to mind, where they're doing brain reading and have sort of established, in actual signal-understanding terms, the kind of phrase-level formation that happens about 2 seconds before you actually talk, and then how it kind of gets down to the literal next syllable that's just a fraction of a second before you spit it out.
There is something that feels like maybe some of these things might be better served in some ways by a diffusion structure, if you could kind of go from coarse to fine, as opposed to always being 1 token at a time at the end. But anyway, that's more speculation.
You mentioned the stack. Actually, can I add 1 more thing? You actually do see varying levels of performance and conversation quality, which is a very big predictor of driving business outcomes. For example, we're investors in a company called HappyRobot, which is voice AI for freight brokers.
If you just look at the quality of their text-to-speech and how the conversation feels, it feels way better than many of the off-the-shelf things. This is because the team is more technical and has done more work under the hood. So yes, there are a bunch of other competitors who can provide a fine commodity voice experience, but if you're a little bit more specialized in the technology, you can do something that feels more human, and that gives you permission to move into higher-value conversations.
Persuasion, negotiation, disagreement—those are pretty nuanced conversations. And for the business to trust you with those conversations, you've got to have a voice model that not just says the right things, but says them in a way that feels compelling.
Yeah, negotiation in particular is a lot to ask somebody to delegate to an AI. Is this company actually doing that—doing actual negotiations?
Negotiates, befriends, disagrees. We'll send you the demo. You should embed it. It's amazing. It's a real aha moment, I think, in terms of what's possible.
It's really interesting because, to the point about the LLM always knowing exactly what it could or should say, there is a version of an AI voice agent that does negotiation that would just respond to the human and say, “No, this is my best price. This is what I'm going to offer you,” and just say it over and over.
But if you launch that kind of experience to a user, they're going to try to circumvent it and talk to a human agent. It's not going to feel like an actual negotiation to them, where they've given it their best and they've gotten a concession from the other side.
What HappyRobot has done, which is really smart, is introduce extra latency by saying, “Okay, the voice agent will say, ‘Hold on, let me go talk to my supervisor,’” and put them on hold for about 5 seconds. Then they come back with a slightly better price.
Of course, the voice agent knows, “Here is my actual max. This is how much I can go up or down or move the price for the end customer.” But they found that the acceptance rate of that kind of final offer is, I think, much higher in cases where the human feels like they've gone through an actual negotiation, because the voice agent has simulated a situation that feels satisfying to them.
Yeah, I don't know how to feel about that, to be honest. It's genius, and it's a little bit far out. Staying on this for a second, do people know that they're talking to an AI when they're talking to this? Does it say that upfront, or how do they—
Yes. It discloses it. What's actually most surprising is that people don't mind.
These are truckers driving all over the country in their big rigs, so these aren't exactly Stanford technology enthusiasts. They don't mind at all because our reptilian brain is so trained to react to these interactions in a specific way that, once you get into the conversation, even if you intellectually know that it's an AI, you fall into the rhythm, expectations, patterns, and cadence of human conversation very quickly.
It reminds me of something Anish says a lot, which is that in the best cases, AI can be more human than the humans. The HappyRobot example is a good one: every time you call in and reach the voice agent, you get the same person. They're friendly. They'll listen to you talk about your day or what happened. They'll be very patient and sympathetic.
They'll spend all the time in the world with you on the phone if you want to. So it's actually, in many cases—assuming the voice agent can answer your question, which almost all of them can now—a better experience for the end consumer than if you get what can sometimes be a grumpy actual human being on the other end of the phone line.
Yeah. Superhuman patience turns out not to be that hard to achieve and is quite valuable. Low-to-no wait times are also a huge driving force for value. I definitely see it. I've been really excited about voice for a long time, even though it's probably going to put me, as a podcaster, out of work before too long as well.
You mentioned Siri a minute ago, and obviously they've recently made headlines in a seemingly negative way by saying they're not going to have an update until 2027, which just feels like possibly the other side of the singularity from where I'm sitting. I don't know if you can make any sense of that. I guess another way to come at that is: How reliable do these things need to be?
I think there's often, in AI in general, this faux comparison or imagined perfection that people compare an AI solution to. Ethan Mollick, I think, has a great phrase for this as well: the best available, or the best hirable, human for the job is really the comparison that you should be making. So is that what Apple is getting wrong here, or how do you make sense of that development?
I have so many thoughts on this. One, I think for any consumer that interacts with AI products and also uses things like Siri, it's like a stick in the eye 5 times a day because Siri is still so bad at just the most basic things. And then, juxtaposed with all the advertising that Apple's doing about Apple Intelligence, it's not awesome. I think it really degrades consumer trust.
And 2, I think that AI does best when it is exploring the surface of human interaction, which is a little messy, and these large incumbent corporations are designed to take the humanity out of every technology product. So there's a sort of almost irreconcilable tension between the 2.
The more they try to neuter the AI, the more dissatisfied consumers are with it. Some of the Genmoji stuff—it's a valiant effort, but it just looks terrible. I don't know; maybe some people think it looks good. I don't.
I think it's going to be a very difficult spiritual problem for these companies to resolve because, between the committees and the lawyers and the whole posture that large incumbents have, it's going to be hard for them to embrace the messiness of AI.
Yeah, I mean, we saw the reaction to the AI-generated text and notification summaries on iPhones, and I would guess that kind of spooked Apple a little bit. To Anish's point, for Apple to launch a new AI product, it has to be production-ready to land on hundreds of millions of iPhones for people of all ages and all sorts of use cases.
It needs to feel both natural and also be correct, whereas a startup has the luxury of not having to meet that bar. The people who seek out and try a new startup product are the natural early adopters, and they know that it's beta, they know that it's AI, and they know that it's a test product.
I think—not to give too much credit here—but I think what we've seen Google do well has been the new Google Labs experiments. That's where a lot of the best, in my opinion at least, AI Google products have come out of, like NotebookLM and some of the video models, like Veo 2.
Essentially, they've taken the approach that if you're an early technology tester, you sign up and get on the waitlist, beta-test, and use these products, and then ideally they get them production-ready to launch to a slightly larger audience incrementally.
But even as we've seen with NotebookLM, because it's Google, once it does make its way to the public, the pace of innovation there is a lot slower than it would be at a net-new company that doesn't have tens of thousands of people to continue to employ. It's like 10 VPs for every engineer working on NotebookLM right now.
And look, another example of this is Deep Research. Deep Research actually was originally a Gemini product, and it's obviously a product that Google should be the best in the world at, yet they never commercialized it in the right way for whatever reason. Now ChatGPT is known for its deep research capabilities. So it's just one missed opportunity after another with incumbents. We'll circle back to that in a minute.
I wanted to get a little bit deeper on the stack and the sort of balance between allowing the AI to handle things and putting your trust and confidence in its decision-making versus trying to maximize accuracy, which typically is going to mean more control measures and more stilted, or let's say less natural, interactions. The stack, for those that aren't familiar with this, basically has been the sort of audio in, transcribe that to text, feed that into a language model, feed that response back into another text-to-speech model, and then send the speech back to the user, roughly speaking.
You can complicate that if you want to, but that pipeline basically now does work fast enough in many cases to be viable. But then there's the fully multimodal model, voice-to-voice, all through one set of weights. When you said that builders are mostly still building on these more multi-model, pipeline-style technology stacks, is that because they're getting better results from it, or is it because that's the only thing that's economical for their use case right now? What's driving that, and do you see that changing?
Yeah, I think the voice-to-voice models are definitely, and unsurprisingly, probably a little bit more expensive right now. But also, I think they're just earlier. I talk to a lot of voice agent founders who have probably tried all of the models available, and I think Gemini Flash is maybe the best of the bunch if you're going to try to do full voice-to-voice. But in general, the interruptibility is not quite there yet on those models.
But as we see and say pretty much every week, the models are the worst that they're ever going to be right now. I'm sure, similar to how last May, when we first released our voice report, sub-1-second latency was hard to imagine, by this time next year everyone is going to be using a model stack that's much sleeker and is hard for us to even imagine right now.
I also think that reasoning models are a new primitive, and I think that they're maybe being underappreciated by many because they have the same interface as language models. But if you think about it, there are aspects of interactions that really benefit from this probabilistic nature of the outputs of language models—all the friendship, negotiation, et cetera, that you have in a business conversation, even.
And then there are things that really do not benefit from that, where you want high degrees of accuracy, like a negotiation: 1 price is definitively better than another for 1 of the parties and the counterparty. So in that case, you can orchestrate reasoning models and language models to handle the right parts of the conversation to get the desired output. So I think you have fewer issues around aspects of the conversation that demand accuracy and the models not being well suited to that.
I think the other broad philosophical point Marc has mentioned a lot is: Is the bar to be as good as humans, or is the bar to be perfect? If the bar is to be perfect, all these technologies are not there and maybe never will be. If the bar is to outperform humans, I think we're actually there in many ways.
Hey, we'll continue our interview in a moment after a word from our sponsors.
What does the future hold for business? Ask nine experts and you'll get 10 answers. Bull market, bare market, rates will rise or fall, inflation's up or down. Can someone please invent a crystal ball? Until then, over 41,000 businesses have future-proofed their business with Netsuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, and HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data. When you're closing books in days, not weeks, you're spending less time looking backward and more time on what's next. As someone who spent years trying to run a growing business with a mix of spreadsheets and startup point solutions, I can definitely say don't do that. Your all-nighters should be saved for building, not for prepping financial packets for board meetings. So whether your company is earning millions or even hundreds of millions, Netswuite helps you respond to immediate challenges and seize your biggest opportunities. And speaking of opportunity, download the CFO's guide to AI and machine learning at netswuite.com/cognitive. The guide is free to you at netswuite.com/cognitive. That's netswuite.com/cognitive.
Yeah, it's all happening very fast. How is the tool use? Because that's one other thing that I could see being a challenge, although I could see it cutting either way, depending especially on how idiosyncratic or esoteric your tool-use context is. For example, if you're doing calls into some obscure freight management system, you might need to have potentially even the ability to fine-tune that model to get those calls to work quite right.
So, broadly speaking, how big of a deal is the back-end interaction of tool use, and what's working or not working there today from what you've seen?
It's a good question. I think what we're broadly seeing is that you've got to build a lot more product than just the voice capability. I think the voice capability alone is insufficient, and there may be more room to explore voice-only in areas like agents, where there are different types of conversations, different outputs that you want, and different considerations around price point and what APIs you want to consume for what fidelity of output.
But even then, there's an enormous amount of integrations and workflows you probably have to deliver to build a traditional moat. So, without commenting on a specific use case in something like freight, our general observation is that the capability gets you in the conversation but isn't sufficient to get you to the other side.
Yeah, I would agree. Especially when you think about a lot of these companies that voice agents are selling into, they're very much traditional enterprises. They're not the Apples or the Googles of the world. For them to build or launch in a more horizontal way, or even to try to build a voice agent themselves, in many ways it's a miracle if they can do it once, let alone keep it updated as the models get better and as there are new options that are going to be a better experience for the customer.
When the integration breaks with the back-end system of record, what do they do? I think that is why we have seen so much excitement on the customer side for more vertically focused platforms that, to your point, are maybe both fine-tuned for the types of conversations that these customers are having, but also have done the work to build out the long tail of integrations and the long tail of, honestly, just conversation types—how do you manage piecing together different tools for different types of tasks that you need to complete?
Yeah, I always think about Tyler Cowen saying, “Context is that which is scarce,” and that has never been more true than in the AI era. I think he said that before the AI era, but it's never been more true. So much of what I see standing in the way between businesses that want to use AI and actual successful use is literally just assembling the context, and sometimes getting the context out of their heads and into some documented form that the AI can process.
It's not too surprising that you would see base models, even as they are quite powerful and extremely versatile, certainly relative to anything that came before them, not know the intricacies of not just the freight business in general, which they might even know, but the way you handle your freight business. That stuff is the last mile; it trips a lot of people up.
Do you have any observations or synthesis of what's going on there? I honestly still struggle, and I do some amount of—not a lot, but enough—hands-on consulting with businesses where I've seen this repeatedly: people have a really hard time assembling the content. Maybe this is what the verticalized startups are going to solve for us, but do you have a sense of what's going on there? It seems like it should be easier.
We should be making faster progress in AI adoption than I feel like we're actually seeing.
Yeah, it's funny. I think we've definitely seen an explosion of AI budgets within enterprises and within end customers.
And in some cases, this was especially true 6 months ago. It’s less true now, which I think is good, but they’re kind of looking for things to buy because, to your point, they’re not spending all day thinking about what’s latest in AI, so they’re not exactly sure how to use it in their day-to-day. I think we even saw this with ChatGPT, which launched with a massive splash—the fastest product ever to 100 million users.
But then people weren’t really sure how to use it every day, and the usage was flat for basically a year. Only now, in the past year, with more models and more obvious things to do with it, has the usage kind of picked back up. So I think it’s exactly to that point: on both the consumer side and the business side, if people don’t know what to do with the product, and if it’s more than a few steps to get up and running, there’s going to be a decreasing funnel, unfortunately, of people who make it through and are actually becoming paying customers. So that’s part of the reason that I think we’re seeing companies that are so vertically focused see the most success here.
Totally. And actually, I wouldn’t underestimate the amount of growth these voice AI companies are seeing, especially on the agent side, but also on the scribe side. The businesses either don’t know how to use it or do know how to use it, and it’s growing explosively within these businesses because, for the businesses that can embrace it, it’s such a straightforward substitution for the humans that make the phone calls. Of course, they need to think, “Okay, we’re going to hold it to some sort of CSAT score, and we’re going to measure outcomes and things like negotiations.” So, of course, there are guardrails and integrations that need to be done, but in the cases where it’s working, it’s really, really working. We’re seeing some of the fastest-growing B2B startups we’ve seen in 10 years.
Perfect setup for a little lightning round on what’s really working across different corners of the economy. I was going to start with enterprise, actually. You mentioned scribes. It seems like that’s a pretty well-established use case at this point. We’re starting to see that bleed over into real-time coaching on calls and then, obviously, full-on agent substitution. Is the real-time coaching working? What do we know about that? And is the substitution actually happening to the point where you think we’ll see labor-market statistic effects this year, or do you think this will continue to be isolated examples that are the exception rather than the norm for a while still?
Yeah, these are great questions. I would say that coaching is definitely working, and it is an interesting transition point in that there are some jobs—for example, a call-center job—where, if you’re an AI product selling a coach to call-center workers, there’s a massive amount of demand for that right now. But what we might imagine is that in 5 years, or probably less, the AI agent is going to replace a lot of those workers.
I think where the coaching will really continue to exist is in these jobs that have a heavy in-real-life component or a heavy personal component. We’ve seen quite a few AI real-time coaches for salespeople, for example, and for HVAC technicians. Many of these jobs, whether or not you get the $10,000 upsell, comes down to the nuance of what you say or the question that you ask. So even if you’re paying hundreds of dollars a month as an individual user for an AI coach there, it’s absolutely worth it.
In terms of the economic impact of voice agents, I think, to be honest, there are some cases—a basic call center, for example—where an AI taking the calls will free up the human workers to do much better and more rewarding jobs. These are massively high-turnover, 300%-per-year, thankless jobs in many cases. And I think there are better things for people to be doing.
Then in other cases, recruiting is one example. We’ve seen quite a few voice agents that conduct initial screening calls for human recruiters. That means they can spend those 20 extra hours a week with the 5 candidates they’re really excited about, catering to them and really convincing them to engage in the process and take the job. So in many cases, I think we see it as amplifying the humans in their current roles versus replacing them.
Yeah, I think all of that is exactly right. Potentially more humans sort of move up the stack to do higher-value work. I think the other thing is that you could take a look at what’s happening and say, at the limit, we have 20% fewer jobs, or you could say, at the limit, we all work 4 days a week and we’re paid to be optimists. But I believe that there’s an opportunity for people to be more specialized and do more of the work that matters, and less of the administrative overhead stuff that seems to consume most of our day-to-day.
Yeah, I’m with that. I guess what I’m trying to really zero in on, though, is that I think it’s true in some ways, what you’re saying, but then there are other ways in which I think retraining and reskilling—and everybody was going to become a programmer, and that really hasn’t happened. When we look at call centers specifically and the people that work there, we may free up resources and the company may grow and invest in other ways, but I think in many cases the people that have the call-center jobs are not going to be moved into other jobs at that same company. Instead, the AI is just going to do that job, and they’ll have a much lower headcount in the call-center operation.
They may reinvest. There may be more R&D. There could be all sorts of great things. I also think people should work less, and a new social contract that embraces that is high on my list of things people should be developing now. But, leaving aside the second-order effects of what happens, are we just at a point technologically where we could see, if enterprises wanted to do it, a 90% headcount reduction in call centers?
And it seems like the call-center thing, if I’m understanding your answer correctly, is probably at least another year before we’d see that sort of effect. Of course, these things are not binary either, but I put that 90% number out there just to say “order of magnitude,” even if it doesn’t go to zero humans in the call center. It sounds like you think that’s at least a 2026-plus phenomenon.
I mean, I don’t think so. We’ll see. So far, we’re not seeing it because, as Olivia said, nobody’s job at the call center is to just do initial phone screens. People have recruiting jobs, which involve initial phone screens that are annoying and can be overwhelming, as well as deeper interviews, salary negotiations, and ensuring that employees are successful once they’re onboarded.
So, yes, the AI is going after the initial phone screen, but we haven’t seen a reduction in headcount because all of the other work is so important. Frankly, in many cases, they don’t yet trust the AI to do the work, or the AI can’t do the work. The AI can’t take your employee to a baseball game a month after they’ve started and make sure they’re having a fantastic experience. So I totally understand the conceptual argument; it’s just not something we’ve seen yet.
I think we’ll also see the success of AI open up probably a new type of job that we haven’t imagined before for humans to do. One great example is that one of the fastest-growing jobs right now is basically contributing training data, doing online tasks, and doing other things that help the AI. It might be similarly paid, but it could offer a much better lifestyle than a call-center job in many ways. So it’ll be interesting to see what new types of opportunities open up for humans in the AI era.
Certainly, the AIs are less abusive than the human callers. I think we can safely say that.
Yeah. Yes.
Yeah, for sure. I’m not—to be clear on my perspective on this—I’m not anti-displacement or trying to stoke fear about that. I think we probably are going to see it and should get ready for it, and ideally it would be a good thing.
I often ask people outside of the Silicon Valley bubble—and I live in Detroit, Michigan—if you didn’t have to work the job you have to make the money that you need for the rest of your life, would you still work that job? The overwhelming answer is no. I consider myself very fortunate that I would probably continue to do what I’m doing even if I weren’t getting paid for it or didn’t need to get paid for it.
But I think it’s really important for the Silicon Valley set to keep in mind that most jobs are not jobs people are doing for the joy of the job. If they could have their needs met in other ways, they would happily take that trade. So I’m not a job preserver; I’m just trying to figure out at what point this wave of disruption is actually going to hit. How much time do we have to get ready for it?
Yeah. Yeah. Look, I think it's cold to tell everybody, "Hey, just go learn programming." That's not what I'm saying at all. I just think it's very hard to understand what the labor impact of these technologies will be. And I think it's easy to hypothesize about a world in which all the jobs just go away. But it's not what we're seeing yet. So even if the technology is 18 months away, I don't know that the labor market will change in the way that we're perhaps imagining as a result of the technology. I think we'll have to see.
I think a broad question, though, that you're speaking to is: What does it mean for our society when we have all this abundance, and is there kind of a lack of purpose? I have a big theory that people need purpose, and if they don't have enough of it, they create it. Sometimes it gets pointed in bad directions. That's a lot of why I think Google ends up not working culturally.
I think there's a lot of brilliant people there, but in a sense, it's a low-stakes environment from a purpose perspective because the business is almost too good. I do get nervous about mirroring that in society. So I think an even more interesting question to me is: In a world where we actually just do have all of this abundance, how do we think about ensuring that people have purpose versus ensuring that people have jobs and income and all those other necessities?
Yeah, I'm a little more, maybe even optimistic on that dimension, but I certainly filed that under a good problem to have. Okay, so this is supposed to be the lightning round, so we'll go through these next ones faster.
SMBs: You've got the call answering, which seems pretty straightforward. Any highlights or, for SMB owners out there, where should they go to potentially get the best AI call answering today?
Yeah, I would say something we've been very excited about for SMBs is that even for SMBs, there's a vertical solution depending on what you're doing. So if you're a restaurant operator, a spa, or in home services, we put all of these in our market map, but there is a solution catered to you, which is fantastic.
This actually gets back to what we were talking about before, which is that SMBs typically have 1 or 2 people doing nothing but answering phone calls, which is incredibly expensive for a small business. When we talk to the SMB customers, when they switch over to a voice agent, they are not laying off that human, who is usually a core part of their business.
That person is now able to spend their time doing things that are much better for the customer experience, growing the business further, or pursuing other extensions of the business, which are really powerful and exciting.
Okay, how about creators? This one is maybe a little less interactive, although maybe you're seeing interactive experiences that are sort of in the creator economy. But one question I had, because I might actually do an episode powered by AI voice in the not-too-distant future, is: Who has the best voice design today?
If I want to create somebody who's going to give me that sort of film-noir kind of read, where should I go for that sort of thing?
I think there are different answers to this question. One, there are platforms like ElevenLabs, where you can clone your voice. ElevenLabs also has a great tool where you can describe a voice, describe a sound, and have it created.
The other end of the creator part of AI voice, to me, is these digital clones, which we're seeing more and more of on platforms like Delphi, where you can essentially launch a version of yourself that your audience can interact with in your voice, or maybe via text or other modalities, which is fascinating.
I haven't seen AI replacing full podcast episodes for any podcasters yet. I think hypothetically we could get there, where maybe you just prompt the questions that you would ask, but we're probably still a couple of years away from that. It's worth, though, playing with HeyGen, Captions, and a bunch of these other products.
You can fine-tune a model of yourself, video and audio, and then give it a script, because I think there is a world in which you do this entire podcast without ever owning a video camera or microphone.
Yeah. We did an episode with HeyGen, actually, and for some reason Josh's audio wasn't great, so we then redid his entire side of the conversation with his avatar from HeyGen. It's been 6 months.
It was pretty good. I mean, it was not quite as good as the original would have been, but it was fitting that it happened on that episode.
How about for kids? I've been playing classic Nintendo games with my kids recently, and I often turn on Advanced Voice Mode while playing Super Mario 64, the old open-world game, because I don't know where to go. Where's the star? What do I have to do?
I'll ask Advanced Voice Mode, “All right, this is the level I'm playing. Where do I go?” My kids are now to the point where they're like, “Daddy, ask AI any time that either I'm slow to do something or don't know what to do. So, Daddy, ask AI.”
That's cool. I would love to have something interactive and educational for my kids, but I'm also like, yikes. I don't necessarily want to trust anyone to implement AI effectively for my kids. So any winners, any early winners in that space?
I think we've talked a lot about this as a team. An area of exploration we're fascinated by is all the stuff around behavioral, social, and emotional development for kids. I think that is an area where AI is very naturally suited to deliver value, and there's just not a lot of technology there.
A great example is that my son loves to play Minecraft, but all of the other people he meets online are toxic teenagers. So why isn't there a companion that can play Minecraft with him and model positive social behavior?
Another example is observing the classroom. If your child goes to one of those great schools where they've got 2 teachers in every classroom—1 doing all the academic components and 1 doing all the social-emotional components—you already benefit from this. But a lot of kids don't go to schools like that.
So having a vision model, a multimodal model, that can observe the children interacting and give feedback to parents and teachers could be really valuable. I think there's a ton there. Of course, there's assignment generation, quiz generation, and helping kids learn in whatever ways are best suited to them.
That stuff will happen, and it'll be super important, but I think pairing it with all the emotional opportunities is where we get most excited.
Yeah, I was going to say I feel like we've seen companies like Synthesis, Ello, and Super Teacher that are kind of like, what if every kid had a reading tutor or a math tutor sitting next to them all day who could understand how they learn best and cater to them?
Then, on maybe the other end of the spectrum, we've seen things like the Curio toys, which is: What if, maybe even more importantly than the tutor use cases for many kids, they had just a friend, a mentor, or a coach who was with them every day and could both track their progress and help get them on the right track—or even just be a completely sympathetic listening ear?
As a side note—sorry, I know we're in the lightning round—the companion stuff is so cool. There are, of course, so many amazing companies like Character.AI and many others, and a bunch around the top 100 and top 50, but it still feels like we're in the first inning, maybe the warm-up or something, of exploring this space.
There are so many contextual opportunities to do it.
And look, I think one aspect of it may be completely sympathetic. Another aspect of companions might be not that sympathetic—one that really challenges you, pushes you, and disagrees with you.
Even something as simple as that: We always joke and call it East Coast mode, like your companion that's a little more terse doesn't exist. Why doesn't it? I don't know. I think we're going to get to see those products in the next 2 years.
Yeah, that's really interesting. All right, in the interest of time, we'll skip over legal and medical. Maybe I can just ask for a word from Munjal from Hippocratic AI, because I would love to talk to him about what he's doing in the medical space.
But how about companionship? Maybe the last one in the lightning round is just: It seems like the farthest-out edge right now is going beyond companionship and into relationship, and even into not-safe-for-work types of things.
I don't know if that's stuff that you guys would touch in investment terms, but I trust that you're at least scouting that territory somewhat. What do you see going on on the far fringes of romance with AIs today?
I mean, it's a good question. I think the thing that's surprising everybody—or at least, I had an assumption that most of the companion use cases would be frisky young dudes, and it actually hasn't been that. A lot of it has been an audience that caters a lot more to women and probably feels more like interactive fiction than it does what you might consider pornography.
So, 1, I think there are a lot of mistaken assumptions that even I had about how the products would be used and who they were going to be used by. I also think there are a lot of definitions of romance, and I think people are perhaps critiquing these products, saying that they're a substitute for traditional romance, when in fact they may make us so much more capable.
You've either got somebody—an AI—that helps you train to be better at things like conversation and even flirting, or an AI that can just be a vent for a lot of the frustration and emotional weight that people can sometimes bring to their in-person relationships.
So those are some of the more surprising things. What do you think, Olivia?
No, no, I agree. It's funny: whenever we pull the top 50 or top 100 list of AI apps and send it around to our team, every time without fail, people are like, “Oh my gosh, there's a ton of companion platforms on here,” and a lot of them are maybe more NSFW-oriented. But I think it's been exactly what Anish has said, in that it's actually much more of the AI boyfriend than the AI girlfriend use case, interestingly, and a lot more like interactive fan fiction than anything else.
But that's a part of the human experience, right? Sexuality is a part of the human experience, and we can't pretend it doesn't exist. When we do, you end up as Apple, which can't release a product for 5 years. So I think that we have to embrace that this is going to be a part of these products and just find ways to get behind that. Of course, there's always going to be products at the fringes that we'll never invest in and perhaps most people will never use, but those are almost the least interesting products to talk about because it's always been that way.
Yeah. I've done 2 episodes, actually, with Eugenia from Replika. The more recent one was reviewing some research that folks at Stanford did that showed that not only did Replika reduce suicidal ideation in a substantial way for people who had that issue coming in, but also that, more often than not, it helped people get out into the world. People indicated that using Replika was not holding them back, but in fact encouraging them to get out into the real world. I questioned some of the data—some of it was self-reported—but I thought it was quite interesting.
I think it's amazing. I'm a big fan. I think, as with all this, as with all technology, but maybe even more so with AI, it feels to me like the shape of this, and the specifics of the shape of exactly what we build, is going to be really important. I have no doubt that you could make a predatory romantic AI that is addictive and exploitative in all sorts of ways, but I think we do see at least some existence proof that you can make really positive, or at least predominantly majority-positive, versions of these things.
And that brings me to a question on rules of the road. I know that it is early in this space. One rule that's been proposed is that AI must disclose that it's AI. That's a Yuval Noah Harari one that I like for its simplicity. I've also been thinking about the idea recently of a “do not clone” registry, which would be the modern version of “do not call.” You could go and say, “Here's my likeness and my voice; don't clone me on other AI platforms.” I'm wondering if you guys have any ideas for either emerging best practices or possible regulations that might keep all of this on the good side as much as possible for us.
Yeah, it's definitely early, but at least on my side, I've been surprised. It seems like more people these days are frustrated by the large model companies taking the approach of, “We're not going to let you do something,” versus people being frustrated that, “I'm being deepfaked or I'm being cloned.” That's not really happening to the average consumer right now. I think we've seen both the biggest startups and the biggest model companies be extremely careful about allowing you to do anything related to a public figure, let alone personal pictures or other things like that.
And so I personally am very intrigued by the idea of the directory or the registry, especially because it opens up this opportunity for people to license or allow their identity to be used for use cases that they are excited about. We're seeing platforms like ElevenLabs; they have these iconic voice collections of celebrities or people who will allow their voices to be used. But it's also been a massive boon to this industry of voice-over artists who maybe historically couldn't get a job in Hollywood. Now there are all of these voice-over jobs on ElevenLabs.
We could see something similar in the influencer-creator economy, where if you're an influencer with 5,000 followers, you're going to have a hard time getting a big brand to respond to you. But if the AI avatar version of yourself is even better, more powerful, and more extensible, then maybe you actually can get some of those big deals. So I'm really interested to see how people can extend themselves using the AI tools, versus—I at least have seen less to be concerned about—the everyday person who isn't a public figure getting deepfaked or anything like that.
I totally agree. I think every time there's a new technology, the talking heads try to get overly paternalistic. I just don't think that's a generous enough view of the average consumer, how smart they are, and how media-literate they are. Of course, every technology has the potential for misuse, so I'm not being glib about that. But I do think the paternalism is unwelcome and often unnecessary, because people have learned for 30 or 40 years that just because it's written in a book, it's not true; just because it's on the internet, it's not true; and just because it's on social media, it's not true. There's no reason that this technology will be any different. Whatever we do here, I hope that we're generous in our assumption that consumers are smart and savvy and will know how to use the products and technologies with the appropriate level of caution.
Maybe just give me your medium- or long-term vision for where this voice-enabled computing is going to go. Are we all going to be walking around with AI in our earpiece and untethered from our devices? Maybe we've got glasses that pair with that. What's the tech-optimist view of life in this voice-enabled-computing future?
I think that we will see voice unlocked as a kind of modality feature on every product and in every interaction, in every device: AirPods, glasses, your computer. As we've dove into voice, especially from a consumer use case, you find that there are a lot of situations where maybe you don't actually want to be having a 2-way conversation, or you can't be having a 2-way conversation. You want it to be transcribing what you say, or vice versa: you can't talk and you want it to be talking back to you.
And so I think right now we're in inning 1 of AI voice, where we have a set of really compelling and exciting products, but 5 years from now they're going to look incredibly limited based on what we have then, where you can interact with voice in any way, at any time, for whatever is most useful and helpful to you.
Steve Jobs famously said that a computer is a bicycle for your mind. That meant that a computer extended us intellectually in ways that were unimaginable, and that's what technology has done for us for 40 years. I think we're now going to have the emotional version of that, the sort of emotional bicycle, where it extends us emotionally through products like companionship, but many, many more. And I think voice is going to be the primary catalyst and interface to that. So maybe that's a subject for our next conversation, but I think that's really the way it's going to impact us, and it's been a bit underestimated.
Cool. I love it. Olivia Moore and Anish Acharya, thank you both for being part of The Cognitive Revolution.
Thank you.
It is both energizing and enlightening to hear why people listen and learn what they value about the show. So please don't hesitate to reach out via email at tcr@turpentine.co or you can DM me on the social media platform of your choice.