语音智能体如何学会对话节奏——Shawn Wen
Tim ScarfeTsung-Hsien (Shawn) Wen
- 企业语音仍是尚未解决的系统工程问题,即便语音组件越来越便宜、越来越易得。 早期 ChatGPT 封装产品只是把 ASR、LLM 和 TTS 串起来,但 Wen 认为,语音还引入了时间、噪音、打断和适应能力:这是“向真实世界迈出一步”(one step to the real world),智能不只是对说了什么进行推理,还要回应对方是如何说的。
- 最有胜算的架构,可能是在听取和轮次控制上采用端到端方案,同时保留文本作为企业可控接口。 Wen 的音频原生模型直接流式接收音频,判断用户是否已经说完,再输出答案或工具调用、引用,最后生成转写文本。文本输出保留了安全护栏和可审计性,也不必让系统等转写完成后再响应。
- 生产环境对话数据构成难以防守的优势,因为干净实验室音频几乎不会覆盖真正棘手的场景。 他们的平台在客户同意的前提下,从银行、公用事业、物流、餐饮、酒店、零售和外呼销售等联络中心部署中获取数据,删除个人信息并植入合成身份。训练会刻意加入串音,以及“背景里有婴儿哭声”时说出一个名字等组合场景;Wen 表示,旧式降噪流程可能反而让新模型变差,因为它制造了一个“过度清洁”的环境。
- 延迟的核心首先是轮次控制,而不是模型原始推理速度。 固定的端点阈值要么会打断说话较慢的人,要么让没耐心的用户干等;音频原生模型则可以结合语速和未完成的语义,识别句中停顿。公司还结合预计算、将模型自托管以降低 P95 网络延迟,以及“延迟预算推理”,判断多想1秒是否真的能显著改善答案。
- 语音身份会直接影响性能和信任,但企业仍警惕过度拟人化。 通用声音会继承失灵 IVR 系统的负面印象,而适度的地域特征——例如英国纽卡斯尔口音——可以传递熟悉的人类服务场景。公司曾发现,最自然的英语声音其实是让一个法语声音强行说英语:“正是那种不完美,才让它变得真实。”
- 企业价值层正在从模型获取转向评估与控制。 Wen 批评忽视延迟、工具调用、引用、姓名识别和任务完成度的语音基准;他的公司改用真实对话数据评估,并计划在“未来1-2个月”内发布该基准。企业可以在不同模型 API 之间切换,但很多企业仍希望拥有自己的测试框架,将品牌、工作流、安全护栏和审计轨迹编码其中。
- AI 供给过剩后,组织瓶颈正从生成转向监督。 Wen 的工程师已经面对多到难以舒适审阅的代码:“瓶颈已经从生产内容转向真正验证和审计内容。”长期市场可能分化为更具娱乐性的消费级语音和受监管的企业级语音,但采用速度仍取决于隐私规范、文化语境,以及智能体能否在不产生未经审计的“AI 垃圾”的前提下隐身于工作流之中。
1. 语音在日益商品化的语音组件之外,又增加了时间与适应能力维度
Wen 最初也认同 ChatGPT 之后市场普遍形成的预期:语音很快就会被解决。封装公司把 ASR、ChatGPT 和 TTS 连接起来,但真实部署暴露出一个缺失维度:语音是“向真实世界迈出一步”(one step to the real world),时间点会实质性地改变响应的含义。
文本智能受益于海量文档和在线论文;相比之下,机器对于具身互动和口语互动的观察数据少得多。Wen 的区分构成了本期节目的基础:这里的智能“并不真正关乎推理,真正关乎的是适应”。
Tim Scarfe 用机器人做类比,把难点说得很具体:人类轻松就能煮咖啡或倒啤酒,但要把这些动作明确描述给机器,却并不容易。对话也是同样的隐性复杂任务——人类每天产生海量对话,但“机器从来没能观察”其中相关的节奏、犹豫和调整。
2. 企业要的是人类级能力与机器人式可预测性的结合
企业需求中存在一个明确矛盾:客户希望品牌声音能准确读出公司名称,同时希望智能体“像人一样超级强大”,又“像机器人一样超级可控”。衡量标准往往是公司最优秀员工的表现,而不是平均运营水平。
Wen 认为,以效率为导向的部署是最现实的采用切入口。智能体一旦在边界明确的任务上取得成功,客户就会开始认为它“比我原以为的更好”,从而为更广泛的使用打开空间,而不必要求它从第一天起就在所有场景达到人类水平。
前沿的 speech-to-speech 系统可以做出让人感觉“奇点已经到来”的演示,但联络中心暴露的是另一类失败模式。一位年长来电者可能因为困惑而停顿,调得过于激进的智能体却继续压着她说——而这恰恰是经过打磨的消费级演示不会测试的体验。
级联式技术栈分别处理识别、语言和端点检测,工程师不得不调节脆弱参数,或额外加入“识别年长说话者”等规则。把这些功能合并后,一个模型就能感知慢语速、背景噪声、串音和犹豫,再从真实部署中的对话里学习恰当的响应方式。
3. 音频原生输入与文本输出构成可控契约
Wen 的模型是音频原生的,但并不是从波形直接生成波形:LLM 接收流式音频表示并输出文本。Wen 保留文本,是因为企业可以更可靠地对文本施加安全护栏;不受约束的语音输出“难处理得多”。
每个音频片段到达后,模型都会反复生成一个成本很低的首个 token,用来表示轮次状态:用户刚开始、仍在继续,还是已经说完。模型预测到结束时,音频已经编码进 embedding 空间,LLM 可以立刻流式输出答案或工具调用,无需重新处理整段话。
输出顺序有意颠倒了经典流程:先给出响应,再给出用于标明所调用知识库事实的引用,最后输出供调试和审计使用的转写文本——也就是说,“无需等转写完成”就能回复。
检索也从传统的“先转写再做 RAG”变成工具调用:模型需要外部信息时,直接生成搜索查询。Wen 认为,真正的创新在于“输入输出契约”(IO contract)——训练配对、数据准备和规定的输出格式;随着更强候选模型“每月、甚至每周”出现,开源基础模型可以被替换并重新训练。
4. 真实噪音是训练信号,而不是污染
他们的平台已经从银行、公用事业、物流、餐饮、酒店、零售和外呼销售等部署中获得相关音频。在 GDPR 和个人数据约束下,取得客户许可并不容易,因此公司会删除真实身份,生成合成姓名和地址;目标是学习对话机制,而不是客户事实。
音频通过 mel 频谱图表示,并按片段流式传入,使感知过程类似图像识别。Scarfe 回忆,受过训练的人可以从视觉上读出语音、泛音乐器、打击乐或椒盐噪声;神经网络同样可以从这种表示中学习模式。
噪音增强是有意为之。Wen 举例称,系统会把他普通话姓名的合成发音,与“背景里有婴儿哭声”等真实生产环境干扰结合起来,在不暴露私人对话的情况下,倍增模型可能遇到的环境。
反直觉的是,旧式降噪和流程清理可能损害集成模型,因为清理后的信号与部署现场的音频并不相似。面对串音,他们目前没有使用专门的说话人分离机制;融合模型可以根据训练和指令,被要求忽略背景说话者、跟随主要参与者,或加入多说话者对话。
5. 自然的延迟来自知道何时不该回答
在级联系统中,端点判断由一个粗略的等待时间参数控制。设得太短,智能体会打断语速较慢或年长的来电者;设得太长,语速较快的人会以为系统没有响应。单一全局阈值无法覆盖每个人不同的对话节奏。
集成模型同时看到语速和语义。如果说话者在句子语义尚未完成前停顿,模型可以继续等待;与此同时,增量编码会预计算部分工作,将模型部署在公司其余平台旁边则可以消除网络往返,而网络往返对 P95 延迟的损害尤其明显。
Wen 希望把节省下来的时间投入“延迟预算推理”(latency-budgeted reasoning)。模型会估计现在停止是否仍能给出好答案,并在训练中加入这样的约束:“你真的需要想这么久吗?能不能把它限制在,比如1秒?”
后端调用时间较长时,仍需要给用户可见的反馈。填充语和打字声一旦重复就会失效,而等待音乐可以为更深度的推理或响应缓慢的企业 API 争取时间。沉默尤其危险:如果 AI 黑屏“3或5秒”,来电者可能会恐慌,怀疑系统是否还在工作,尤其是在多年不可靠自动化的经历之后。
6. 熟悉的不完美往往比通用的精致更能赢得信任
大多数企业似乎并不希望助手被推向与真人完全无法区分;它们希望助手解决问题,同时带有克制的品牌身份。一家南美赌场曾考虑使用高度本地化的开场声音,但最终因“感觉有点太过了”而放弃上线,说明这种热情接近生产环境时会明显退潮。
通用声音往往表现更差,因为来电者会把它们与僵硬的 IVR 提示联系起来,而这类系统反复无法识别“付款”或“转账”。地域特征可以扭转这种预期:英国用户可能信任纽卡斯尔口音,因为许多联络中心员工来自那里,这传递出的信号是,用户可以直接说出问题,并让智能体解决问题。
公司早期的英语声音不如一个被强行用来讲英语的法语声音自然。后者略带口音,提供了有用的不完美。因此,Wen 把声音选择视为品牌和地区特定的问题,而不是一场追求最大流畅度、或寻找某种所谓中性声音的普遍竞赛。
7. 可信的语音基准必须衡量完整互动
Wen 认为,单独的 ASR 和 LLM 基准无法覆盖智能体的多维表现:响应延迟、转写、事实质量、推理难度、人名识别、指令遵循、引用和工具调用都很重要。“不考虑延迟的语音智能体基准不是真正的基准”;来电者不可能等待60秒。
公开数据集往往使用合成数据,因为真实语音涉及隐私,而现有测试也没有充分代表生产环境。公司基于消费者对话搭建了内部基准,Wen 表示其模型目前在该基准上表现最好;他的计划是在“未来1-2个月”内向研究人员和市场发布该基准。
主观体验比事实评分更难衡量,因为口音和语气会随文化与任务变化。在生产级评估中,Wen 更看重行为证据:用户是否参与、是否完成任务、是否获得满意的解决方案?这些结果评估的是完整的“模型加测试框架”系统,而不只是模型本身。
8. 企业会先拥有边界与工作流,再拥有模型权重
企业越来越把 AI 视为必须掌握的技术,但拥有模型需要调优能力、GPU 和预算,而很多企业并不具备这些条件。因此,它们可以接受在不同模型 API 之间切换,同时把测试框架——提示词、工具、品牌行为和流程逻辑——视为可实现的专有智能。
Wen 提出的边界是:让快速响应的语音智能体充当面向客户的“前门”,把困难工作交给企业自有的“幕后主脑”。他预计企业会发现,自己不必拥有每一层技术;但当任务足够关键时,一些企业仍会保留定制化测试框架。
B2B 对话尤其适合优化,因为它的规格比社交聊天更清晰:问题是否解决、客户是否满意、速度是否足够快?Wen 将其与人类闲聊作对比,后者的成功可能改变一段关系,也不存在可比的黄金标准。
自主性仍会与治理要求正面冲突。一家大型公用事业客户表示,未来每个 PR 都需要经过持续2周、可审计的审核流程,不能接受自动合并。因此,公司保留了流程可视化、受限工具集和逐响应引用;无代码式界面如今服务于审计查看,而不再是主要的构建界面。
9. 下一道瓶颈是人的注意力、上下文与采用意愿
Scarfe 将其称为“认知债务”:智能体可以在他走路时工作,但等他回到电脑前,可能才发现智能体重复执行、委派错误,或产生了一些无人监督的推理。Wen 在组织层面也看到了同样问题:AI 正向 Slack 和代码审查灌入大量内容——“不要把你还没读过的 AI 垃圾发给别人。”人类必须理解并认可自己传递出去的内容。
语音可以缓解这个问题,但也会引入另一个问题。他们在生产环境中的回复平均约50个输出 token,避免出现大段文字,但语音是低带宽、 “非常乒乓式”的沟通渠道。Scarfe 的广度优先工作流可以协调邮件、编辑和编码智能体,但 Wen 怀疑,仅靠语音无法传达完整情况,除非测试框架能够呈现简洁的高层要点。
让智能体隐身,需要更丰富的组织上下文,而不只是更多文本。书面价值观通常不会记录这些价值观为何形成,因此模型可能过度强调6项原则中的某一项。即便“垃圾信息”也取决于接收者:为 Wen 的语境优化的消息,对同事而言可能完全无关,从而形成一种循环前景——智能体为其他智能体生成丰富内容,再由后者进行个性化处理。
Wen 预计,端到端轮次控制将成为基础设施,随后分化为两条路径:更具娱乐性的消费级助手,以及专业、受监管、可审计的企业级语音。会议协调、内部生产力和“第二波物联网”可能随之而来,但他明确表示仍不确定:技术可能在10年内成熟,可围绕办公室语音、隐私、智能眼镜和交出控制权的社会不适感,仍可能拖慢采用。
完整逐字稿
They want the agent to be super powerful, like a human being, but they want the agent to be super controllable, like a robot. Voice AI is very important because that's the native way humans talk to each other, and it's the most natural interface. Computers haven't been at the same level as humans in terms of voice conversations yet, but we do foresee that this should be the future in the next couple of years.
Do you think that speech technology has become commoditized, or is that just completely wrong?
It has become a lot more commoditized in some ways, but in many other ways, it's still not. Modeling a real-time conversation is still hard, and it's not quite there yet. The bottleneck has already shifted from producing content to actually validating and auditing the content.
People think that voice assistants are a solved problem, but it's actually much harder than people think. Why is that?
1. Voice Is Harder Than Text
When ChatGPT came along, we thought that voice was probably going to be solved very soon. I think a lot of people had the same idea, and that's when you saw a lot of wrapper companies come along, wrapping around ChatGPT, piping ASR and TTS together, and trying to solve the problem.
The reality is that it's actually harder than we think. If you think about voice, it's actually one step closer to the real world. All the current innovation on the AI front is predominantly in text and the digital world, where solving a PhD-grade physics problem seems not to be a problem anymore because you have really good documentation and a lot of online papers you can research.
In the voice scenario, you have an additional axis, which is time. Time is a very critical matter, and I think we have not trained the model to understand that concept better yet. Because of that one step into the physical world, we started to realize that there's a lot more complexity when you add that time element.
With humans, intelligence here is not really about reasoning; it's really about adaptation.
When we spoke last time, you said that doing voice technology is a little bit like robotics.
Yeah.
Right? So it's incredibly easy for humans. There's the famous Wozniak coffee test: “I want a robot to make me a cup of coffee.” Just tell me a little bit more about that. Why is it so difficult?
Again, it gets back to the physical-world situation. In the physical world, humans collect data very naturally, and we have a lot of data because we have visual data and body movements. Your hands can reach out to things in the world.
However, computers don't have that yet. All the tools computers have are about how to use computers. I think we have trained the models to be really good at understanding text and understanding how to use tools in the virtual world. But in the physical world, the data is still very limited, and that applies to robotics.
That's why a lot of robotic companies right now, as you said, are trying to grab a coffee from one place and move it to another. I remember one of my really dear friends, and also a really good researcher from Cambridge, who was originally at Poly and has now left to start his own B2B business. He wanted to start a robotics company.
I asked him, “What is the first application you're trying to build?” He said, “I want the bot to be able to pour beer for me. That's test number one.” I was like, “Oh, wow, okay.” It's just really, really hard.
Then you look at self-driving cars. Elon has been promising the world that this is going to be fully commoditized. I think we're starting to see signs of that now. Voice is a very similar problem, except that we have so many conversations with one another, and machines have never been able to observe them.
What do enterprises actually want from a voice agent?
2. Enterprises Want Power And Control
Enterprises want a branded voice. They want it to be able to pick up their identity. They want the agent to say their brand name exactly the way they want it to. That's number one.
Number two, they want the agent to be super powerful, like a human being, but they want the agent to be super controllable, like a robot. It's actually a contradiction. I think that's the difficulty of deploying AI agents into enterprises, because people's expectation of AI is that it needs to be as good as a human—actually, as good as the best human on their team.
That's why the adoption curve is always harder to justify, unless you can find a particular area where you're saying, “Hey, this is more about an efficiency play.” Then people start adopting it and start to feel, “Okay, actually, the agent is better than I thought.” That's what we're seeing from enterprise adoption right now.
Why did you decide to build your own custom model rather than using some of the existing ones out there?
3. Why They Built Their Own Model
We've been doing research and building our own models for a very long time. We previously built our own speech-recognition model, and we also built our own large language models.
Around 1 year ago, we decided that cascaded systems were going to be outdated in the future, and that end-to-end models would be the future. We started doing a lot of experiments, and we decided not to invest in our speech-recognition models anymore because of that radical shift into the future.
When we started looking at the market, the reality was that all these latest speech-to-speech models were really good for demos. I can show you a really natural conversation, and you'll be like, “Wow, that's the singularity already.” But once we deploy them into contact centers, they break in so many different places.
The interruptions are sometimes too aggressive. If you're on the phone with an old lady who is talking to a bot for the first time, she might be constantly confused and stop, and then the agent just keeps speaking over her. That's a level of frustration.
We see a lot of these scenarios, and we say, “Hey, we probably have to train our own model for these use cases.” I think the frontier labs are not optimizing for what we care about: real conversational turn-taking for B2B business use cases.
For consumer use cases, I think you can build voice agents already. For enterprises, there are additional governance and enterprise requirements that we have to satisfy.
How adaptable could that be? For example, if I'm talking to an old lady on the phone, I would adapt my speech. I would slow down a little bit, and I would give her more thinking time. Could the models be adaptive in that way?
Yes. That's why you have to collapse all these modules together into an end-to-end model, or in some sort of end-to-end fashion. In the traditional cascaded system, you have a speech recognizer, a large language model, and turn-taking on top, which is responsible for detecting the end of a sentence.
These 3 models don't really work with each other very well, so you have to tweak the parameters. Traditionally, you can have another detector that says, “Oh, this is an old lady, so I have to tweak a parameter here and there,” but it's all very clunky. The reality is that it doesn't really work like that, and it doesn't work very well.
Therefore, it's almost necessary to collapse all these pipelines together so you have a large language model directly perceiving the audio. The audio can involve slower speech, background noise, crosstalk, and many other scenarios that you encounter when you're having a real conversation with a human over the phone.
Then you can collect that data so the model can adapt to it. I think collapsing everything is the first step, and then the initial model needs to be good enough so you can continue to iterate on it, because you use it to collect more data as well.
I suppose it all comes back to Rich Sutton’s The Bitter Lesson, which is that, in principle, maybe we could write the code to do this. But there’s an exponential number of failure modes—combinations of bad things that might happen at the same time. In this situation, we need to do this; here, there’s some crosstalk; or here, there’s some noise in the background. Just training a neural model end to end, you can cover a lot of those failure modes.
But I suppose it still needs to be a cascade in some sense, because it’s not possible to embed frontier-level intelligence in these voice models. But you could potentially draw a boundary around the voice intelligence model, and then it could call out to an agent when it needs more intelligence.
Yeah.
But I wanted to ask: can you sketch out the high-level architecture for this voice model?
Yeah. The high-level architecture is that this is basically an audio-native model. The base model itself is a large language model, so it actually outputs text. We’re not training models to do end-to-end speech yet. The reason is that, for enterprises, control is important: text is a lot easier to apply guardrails to, while speech is a lot more fiddly to deal with.
We train the model to output text, but at the same time, we allow the model to take ASR into its input channel. That means the model itself natively understands audio. The way you pipe audio into the model is by streaming audio chunks into it.
The way we structure the model is that it’s doing multiple things at once. First of all, based on the streaming audio, the model predicts a first token. The first token is basically the turn-taking signal. Turn-taking is basically saying, “Hey, is the speech just getting started? Is the user still speaking? Or has the speech already ended?”
That initial token is very cheap to generate. It’s just like the first token in the output. As audio comes in, we keep generating that token until the model generates, “Okay, this is the end of the speech.” Then we start having the LLM do the processing.
At that point, all the audio has already been preprocessed, because as we encode it and then predict the first token, it’s already in the embedding space. So you don’t need to encode that sentence from scratch in that case.
The next job of the model is to directly predict an output. The output is text, as we said, so it could be a response or a tool call, for example. That’s all streamed, like the typical ChatGPT API. We stream the response back, and at the end of that response, we then continue to predict a few things.
Mm.
A few things like citations. For example, if you have a knowledge base with a lot of different content, the model will generate a citation, saying, “Actually, for this response, I’m using this fact and that fact to make the prediction.” Then enterprises know that it actually has a reference there.
Finally, it also streams out the transcription. So we’re kind of reversing the order, because the model’s job is not to predict the transcription, but we still want it to predict the transcription for auditing purposes, debugging purposes, and things like that. But it comes last, because you don’t have to wait for the transcript to actually start responding to the user.
Even that’s quite an interesting trade-off to talk about, because many folks would have used the GPT live model, and that is a bidirectional waveform-to-waveform model. It’s really interesting. It just makes all of these weird pauses and sounds, and it sort of breaks the uncanny valley.
Right.
But actually, that might be great for me at home, but in the enterprise, you need to have constraints. You need to say, “No, you’re not allowed to talk about this, and you’re not allowed to talk about that.”
Conceivably, you could run separate ASR on the output, and then do retrieval-augmented generation and link to the knowledge base and stuff like that. But there are all of these trade-offs, and in enterprise, that’s not necessarily what you want to do.
Retrieval itself—you can also do retrieval with these audio-native models, because the audio-native model can handle it. But the retrieval mechanism is slightly different. You cannot use traditional RAG anymore, because traditional RAG means that you embed the user’s transcription and then do a vector search in your vector database, right?
But because now you don’t have that transcription, what we need to do is prompt the LM to generate a tool. The tool is basically a search query. It says, “Hey, now I think that in order to answer this question, I need to generate a search tool, and that search tool needs to come with that query.”
It’s pretty much like function-call-style search, which people call agentic search. That’s just another buzzword, but it’s basically just a tool call, really.
The important thing is that our innovation is not in the model itself. We’re innovating at the I/O contract. That means: how do we prepare the data, how do we prepare the input, and what is the mechanism for the data feeding into the model? But also, how does the model output that data? And then how do we train the model based on these input-output pairs to make it behave exactly the way we want?
To be honest, the underlying model almost doesn’t matter all that much, because we know that in this market, open-source models come out every month, even every week. So we want to maintain our data and our model, and then have these training scripts that are always ready. Whenever a new model comes in, we train on top of it to see which one has better performance.
And voice data is much harder to collect than text data. How are you doing that?
4. Real Conversations Train Better Models
For us, we don’t have to collect it, because our platform generates a lot of voice data. We’ve been deployed in enterprise contact centers.
Yeah.
We have banking clients, utility clients, logistics clients, restaurants, hotels, and a lot of retail outbound sales use cases. We already have a lot of that data, and that data is exactly the kind of data we want to use to solve the problem, because we serve predominantly enterprises at the moment. Therefore, we leverage a lot of this data.
Obviously, getting agreements with enterprise clients to train those models is nontrivial because of data privacy, GDPR, PII, and all of those issues. But kudos to our legal team: we actually get quite a lot of our customers to agree to that because they know that, if they agree, the model can continue to help improve things for them.
The way we train the model is that we redact all the PII data, so none of the customer data is actually exportable post-training. We also produce a lot of fake PII data, so the model can still learn how to identify names, addresses, and things like that. But really, the data is just to let the model learn the mechanics of the conversation rather than the actual user data.
Very cool. And I wonder: does the model currently take in pure waveform data, or does it take in featurized data? I just remember I was doing audio stuff for my PhD, and back in those days, it would have been unimaginable to put in waveform data. And now—
Yeah.
Now it’s simply a case of if the model’s big enough and you train for long enough, you can.
I think currently the standard audio encoder is that you feed in the mel spectrogram, right? It’s still like this image. You remember, there are different spectrums and different frequencies and things like that. It’s still that representation.
But you feed it into the model segment by segment, right? To the model, it feels more like an image-recognition task. I remember that you’re from a speech background, so I remember there was an MIT professor. I think his name was Victor Zue or something.
Back in the 1990s or something, he said, “You will be able to identify what human speech is by just looking at the spectrogram.” He was like, “This is exactly that for the particular syllable,” or whatever. And people were like, “What do you mean? There’s nothing you can identify there.” So, in fact, a machine can learn those patterns really well from it.
I love looking at spectrograms—
Yeah.
—because there are so many things that you can immediately see. For speech, or if it's a violin, for example, you just see these horizontal lines—
Yeah, it's really cool.
—and you can see that there's a harmonic series. If it's percussion, it looks like a sort of vertical bar in the spectrogram. But yeah, after a while, as a human, you know that you've been looking at these for a while, and you can just read it like a book.
Oh, totally.
It's fascinating. And noise, of course, just looks like salt and pepper.
That's right.
But on the subject of noise, this is quite interesting. How do you guys deal with that? Is this the kind of thing where you deliberately sample noisy data, or do you augment your data and add noise to it?
Yeah. It's a very deliberate decision to add a lot of different noise to the training data set. To be honest, our training data set by itself already contains a lot of different noise because it's real conversation data. But we typically apply another generative model on top of it.
Say that I wanted the model to pronounce my Mandarin name, Zhong Shen Wen. In doing that, I also ask these TTS models to add the additional audio data that we see in real production. The outcome of that generated data set could be a robot pronouncing my name while there's a background baby crying, for example. These things can be simulated so that you can have a lot of variety.
Thanks to modern generative AI models, you can create a lot of very realistic data. So the answer is yes, we do a lot of this, and it's a very deliberate approach because we want the model to be robust against all these different environments. But even the training data itself already contains a lot of that.
The funny thing is that, because with the previous technology we were doing a lot of denoising and pipeline sanitization beforehand, all those additional components in this new dialogue-reasoning model are actually making it worse, just because the environment is too sanitized. It hasn't seen that before.
I remember for many years there was this phenomenon called the cocktail-party problem—
Yeah.
—which is that when we're at a cocktail party and there are many people talking at the same time, we can focus our attention on one person, and our brain can filter out the other people. ASR models couldn't do this.
Right.
But increasingly, they can do it now because maybe internally they're learning, "This is speaker 1, and this is speaker 2," and they focus on speaker 1. Crosstalk has always been a problem. How are you doing with those failure modes?
Crosstalk in our model currently doesn't have a specialized mechanism for it, but the architecture is suited to tackle that problem. Speaker diarization technology was previously predominantly applied to speech recognition, and it always assumed a cascaded pipeline solution. You tag someone as speaker 1, speaker 2, or speaker 3.
But it's not a very natural way of doing it because, as a human, if you're at a party listening to so many people, you don't actively label them: "This is speaker 1, this is speaker 2, and this is speaker 3." You think, "That's another person. That's another person I shouldn't pay attention to," because there are so many of them. Some of them will stop speaking, and then you'll lose track of them for a while.
I think traditional speaker diarization is trying to solve a very different problem. With a dialogue system, what you really need to do is have the agent focus on the right subject—the person it is talking to. Because you now combine the language model with speech recognition, the language model can natively perceive that there's background crosstalk somehow.
To be honest, this language model is also promptable. You can prompt it by saying, "Don't respond to the background crosstalk apart from the main speaker you're speaking to." You can also say, "Whenever there's crosstalk in the background, take active participation in that conversation as well." Technically, you can do all of this.
The major difference is how you train the model. How do you collect the data? What kind of data do you collect, and what instructions do you give the model? I think you watched the GPT live demos where you have a lot of these old ladies coming along, and the model is able to identify a new speaker. I'm pretty sure that in their prompts they're saying, "Now you're in this party-like conversation, and you have to gradually identify speakers as they come along." Once you fuse these models together, they have the ability to do that. The major difference is how you train it and what you're training it to do.
The turn-taking thing in particular—we spoke about that earlier—is really, really difficult. I suppose we should imagine that, in an ideal world, in 10 years' time, when this technology has reached the utmost maturity, maybe we can have microphones in every single meeting room. When anything private is discussed, it will know not to pay attention to that.
There might be multiple people involved in the conversation. You can address the agent, the agent can get involved, and the agent knows who's talking. At the moment, we're kind of solving it in the one-on-one case—
Yeah.
—and we're solving this turn-taking problem. Is that a reasonable estimation of what the future might look like?
I think the technology will be complete. Your prediction is right. Ten years from now is a lot of time. Given the speed at which all this technology is improving, I think it's completely reasonable for the future.
The only thing is that voice modality—I mean, people say it's the most natural modality for interacting with a computer. The reality is that we've had several iterations of these kinds of personal assistants in the past. It started with Siri, then you had Google Assistant and Alexa. You have in-car voice assistants as well.
Maybe the technology wasn't really mature yet, but there's also the fact that voice is a lot more personal and private to you. I don't know whether consumer behavior will be there yet, because that means you have to surrender a lot of your control to the agent. A lot of people will say that's a great thing because the agent can do a lot for you. But I think there will be some consumer adoption curve that we still have to climb, even though the technology is here.
Can you tell me about how you did latency engineering? We're going to talk about this uncanny valley and all the things we need to do to make it possible for people to suspend their disbelief because they're talking to a voice agent and it's suddenly really good. But latency is probably one of the most important things. When you reduce latency, there's always a bit of a trade-off: you might be trading off some of the intelligence. You guys have best-in-class latency, so how have you done that?
5. Latency Shapes The Conversation
Latency, first of all, the major thing is that you need to use the right size of model. To be honest, if you're using the right-size model, the model processing itself isn't really a latency problem.
Historically, the cascaded system struggled with waiting to decide when the speech had ended. There's a very coarse parameter that you can tune: how long you want it to wait until the end of the speech. If you tune it to be too short, it could be too aggressive, and you're going to interrupt a lot of older people, for example. If you tune it to be too long, you'll be waiting for too long, and a lot of people will become impatient.
Everyone has a different threshold in terms of how fast they expect the conversation to be. I think that's the difficult part, because you're trying to set a programmatic metric for every single user speaking to your assistant.
With this kind of model, you have a lot more flexibility because the model itself is now predicting based on the audio frame. The audio frame contains how fast you speak, as well as the semantic information.
So if I'm pausing in the middle of a sentence, the semantics have not been completed yet. The agent knows, “Okay, I have to continue to wait until the person finishes the full sentence.” That is the major benefit of collapsing the ASR and the language model together: you have more natural turn-taking built into the model.
The model is now making that prediction by itself. A lot of the pre-cached processing, as we talked about earlier, is another way to optimize it. To be honest, the fact that we host our own models alongside the rest of our platform is another benefit, because you don't have network-travel latency impacting the long-tail use cases, where the P95 is the major problem for network latency.
We optimize a lot of this so that the agent can respond very fluently. However, my team is now adding more reasoning to the model as well. The good thing is that once you save a lot of time, you can use that time to do more things. Our team is looking into adding auto-reasoning to the model, so the model can decide, based on a particular user query, “If I stop my reasoning right now, how likely am I to give a good answer?”
We train our models on auto-reasoning in a way that we always challenge them: “Do you actually need to think for that much time? Can you cap it at one second and then respond to the user?” The reasoning we are training the model to do is very latency-budgeted reasoning. It's almost like going to a competition: someone gives you a question, and you suddenly start counting down.
It's so interesting because sometimes when I interview people and ask them a really difficult question, even though they're thinking, they still feel that they need to say something. Quite often, on a difficult question, the first sentence of their response is just filler because they're obviously thinking, and then they actually answer the question.
That's right.
Right. It's a similar thing here, isn't it? This auto-adaptive reasoning is really interesting because, from a user-experience point of view, how do we solve that problem? Do we make the model say, “Hmm, that's interesting,” and indicate that it's thinking?
That's right.
It might actually be a significant amount of thinking. Maybe it needs to delegate to an agent and do some things, and then come back to you.
There are a lot of tricks you can use in that scenario. In our real productions, you can add a lot of filler sentences, obviously. But if you repeat them too many times, it becomes very repetitive, and the user will know that you're just bluffing them.
Another way is to play a typing sound, so there's computer typing in the background to mimic the actual scenario. The real thing we're trying to solve in that scenario is that you always have to deal with API latency anyway, and some enterprise backends aren't the fastest by default. You have to bake in a lot of those things to make sure the user doesn't feel abandoned.
The reality is that users have been talking to a lot of these automated systems, and if the AI goes dark on them over the phone for even 3 or 5 seconds, they panic. They'll be thinking, “Is the AI actually working?” In order to have that trust, you have to provide constant feedback.
The good thing is that contact centers have a lot of ways to deal with this. An agent will put you on hold with music while they're thinking. In the old world, a human agent might actually go to different places and talk to their colleagues about how to solve the problem. The AI agent can put you on hold music while it's doing the reasoning.
I think there are a lot of tricks you can use there. But articulating things well is very important, because I think the biggest issue voice agents have in contact centers is that users don't trust them, simply because they've been burned many times in the past.
That's a really cool idea. Play elevator music while it's thinking.
Yeah.
I suppose this leads neatly to the next thing: what is the uncanny valley for a voice agent? We talk a lot about anthropomorphization and the risks of that. Is that something we actually want in the enterprise? Do we want it to sound indistinguishable from a human?
Different businesses would give you different answers, but I think the majority of businesses would probably hold the same position as Apple and Google, for example. Alexa is probably the consumer personal assistant that goes furthest in terms of personalization.
Mm-hmm.
They give Alexa an identity, and there's a very unique voice there. But that's as far as they want it to go, because if this personal assistant is representing a particular company brand, it's very hard to know how successful it would be. A lot of the time, companies still try not to attach too much personality to it. The goal is simply to make the assistant solve the problem first.
The majority of enterprises are still on the conservative side. We do have some customers—for example, there's a casino from South America that likes very local, very South American voices. They like the agent to use those voices to open a conversation, for example. In the end, they didn't launch that agent because they were like, “Oh, that feels a bit too much still.”
People would think about that once they felt more comfortable with the technology. But I think we're still in a phase where people are just building trust in the technology.
It's interesting how much we should or shouldn't chase the personality thing, because the reason it sometimes sounds egregious is that it might be normal in Texas or in a particular location, but it isn't normal to me. Do you think, in principle, there should just be a generic voice? Or should the agent get to know you, and when it knows how you talk and what you like, should it adapt to you?
I can speak for what's happening right now, but I can't really predict what people will want in the future. Currently, the majority of our customers don't like generic voices. Generic voices also perform much worse for consumers in general, because consumers identify that it's a robot, and once they identify that it's a robot, they immediately don't trust it.
We have these traditional IVR systems that have been deployed throughout America: “Say a few words. If you have a payment issue, say ‘payment.’ If you want to transfer money, say ‘transfer.’” You say it many times, and it still doesn't recognize you. Those voices are always very generic and a little robotic, because previous-generation technology couldn't really make a natural-sounding voice.
So consumers immediately connect a generic voice to the previous system they've talked to. They hang up, or they try to bypass the agent as much as possible. We've found that the greatest success comes when you don't use generic voices. You want them to have a little bit of a local, regional take.
In the UK, for example, a lot of contact centers and consumers like the Newcastle voice for obvious reasons, because many agents are from there. When they hear the Newcastle voice, they know it's a human agent. They can say their problem to them, and the agent can solve it for them.
I think those voice choices are quite nuanced but quite important. It depends on which brand it represents and what people expect the voice to be. The best-performing voices usually come with a little bit of a regional take.
The funny thing is that, in the early stage of shipping these voice agents, we found that all the English voices were not natural enough. The most natural voice was a French voice, and then you forced the French voice to speak English, so it took a little bit of a French accent. It's the imperfectness of that voice that makes it real.
So that's fascinating.
Yeah.
We should talk about benchmarks. For years, there's already been benchmark blindness, and many of the ASR companies embellish their benchmarks. They all use different internal tests, different public tests, and so on.
But now, in the regime of voice agents, it's even more difficult to test this.
Yeah.
So how are you guys doing that?
6. Voice Agents Need Better Benchmarks
Benchmarking the voice agent—I mean, benchmarking ASR and LLM—we have a lot of test standards. These are all more of a closed, particular module, and you evaluate them on that.
The voice is really difficult to evaluate because a voice agent has latency considerations. You have to reply fast enough. There's also the quality of the response. There's also a reasoning task: how difficult a task the agent can handle, and how accurately it can identify personal names and information like that. So it has too many axes.
I think the majority of the benchmarks that are publicly available out there don't really serve that purpose. Another problem is that a lot of these voice-agent benchmarks on the market are synthetic, to a degree, because voice is a lot more private. Releasing that voice data set for evaluation and then making it public is nontrivial as well.
Therefore, we have a lot of these benchmarks on the market, but they don't sufficiently represent what a good voice agent should look like. I think the industry as a whole is still figuring out how best to evaluate it. The reality is that voice agents are moving so much into real-world applications that sometimes it's also hard to say there's a really good benchmark that solves all the problems.
Our benchmark is built on the real data set we have from conversations with our consumers. We built it in a way that measures a lot of different important axes. We wanted to make sure the agent's response time is a factor in the benchmark, because a voice-agent benchmark without latency consideration isn't a real one. A user cannot wait on the phone for 60 seconds for the agent to come back with an answer.
That's a very important element. We evaluate understanding: would you be able to understand what the user says and then transcribe it accurately? ASR transcription is still one of them. Then there are the tool and traditional LM evaluations, instruction following, and the quality of the response. Are you using citations? Are you using tool calls and things like that? There are a lot of different evaluations we run.
Currently, our model is evaluated on that benchmark. It is the best model, of course, but it's an internal benchmark right now. Our plan is to release that benchmark in the next month or two to the market and to the research community, so they can also evaluate on it. We do believe that having a more realistic benchmark to evaluate on would be a good thing for the overall industry to continue moving forward.
One thing that I'm thinking of is that this might be an example of subjectivity. There are some objective things, right? Just quality and trust and tonality and stuff like that.
Yeah.
But it might be culturally specific and task specific. For example, in this particular application, someone in Newcastle might like it, while someone in France might not like it.
Right.
So what would you do there? Would you statistically aggregate over lots of different situations? How would you approach that?
Currently, the benchmark objectively looks at the quality of the response, more at a factual level. The actual end-user experience is an even harder problem to measure. For measuring that, we can argue that we need a benchmark data set to evaluate it.
But the reality is that if you're building an agent or a system toward that level, it's almost at production grade already. In production, you already have the end user to tell you. The way that we also measure this in real production is: does the user want to engage with the agent, and does the user actually finish the task in the end?
There are a lot of other production-related metrics you can measure there. But that is a lot harder to measure because it's not just about the model's performance. At that level, you have to couple the model with the harness together. So it's that entire agent evaluation, and that is an even harder task.
Can you talk me through the complexity of harness engineering? The reason this is interesting is that so many enterprises are adopting AI, and the main form of adaptation is harness engineering.
Yeah.
It's an incredibly good thing, and it's an incredibly bad thing, because everyone is building their own harnesses, adapting their skill surfaces and prompts. Essentially, everyone is reinventing the wheel and doing different things.
Right.
So there is an argument for having some kind of platform and standardized ways of doing this.
Yeah.
How do you deal with that problem?
7. Enterprises Are Building The Harness
This is what the market demands, right? I think enterprises acknowledge that artificial intelligence is going to be a very important technology for them to own, not just buy. They want to own it. They obviously want to own the entire brand and, ideally, the models as well, but they know they cannot own the models because they don't know how to tune them, and it costs a lot. They need to have a lot of GPUs, and they don't have the budget for it.
So they gave up on the models. They don't want to own the models anymore. They were like, "Hey, we can switch to different model APIs. That will be okay." But they still want to own the intelligence themselves. The only way for them to own it is to own the harnessing.
That's why I think a lot of enterprises are hoping to build their own harnessing, because everyone considers this such an important technology that they need to have in-house. Harnessing is the only thing they can currently bring in-house. So I think that's the market sentiment right now.
I do think not every harness should be owned by you yourself. It depends on how critical the task is. To be honest, I think voice agents—the voice-facing customer experience—is a harder problem than just general reasoning, because the adaptation itself is very important and very critical.
I don't think enterprises value that part as much. They value the reasoning part a lot more. What we like to tell our customers is, "You don't actually need to own the voice agent or the voice-agent harnessing itself. You can have this voice agent as the front door to your customers, and you can own your brand in the background."
Then you have the voice agent, which can respond very quickly, but it can delegate a task to your mastermind behind the scenes. I think this is probably what we're going to see in the next couple of years: enterprises will realize that they don't have to own everything, but they have to figure out where to draw the boundary—"This is how comfortable I feel giving things away."
Some people will probably always say, "I really need to own the harnessing because I want to do a lot of customization myself, and then I'm comfortable using the model." That's okay as well, but I don't think that will be the majority.
I love thinking of software engineering as a specification acquisition problem. Apparently, there's a whole theory around this, which is software engineering as theory building. Many folks at home will have experienced this: when you do vibe coding with agents, there's a kind of Rubicon moment where your software is sufficiently well specified—
Yeah.
—so that an autonomous agent, when it sees problems, can just fix things. It's almost like, if you can get it to the point where it's well enough specified, then it just converges from there. That's the magical moment.
Exactly. The beauty of working with these enterprise or B2B conversations is that there's kind of a good loop, right? It's actually relatively well specified because, in an enterprise conversation, the major thing is: does the customer's problem get solved in the end? Does the customer feel good in the end? Is it a good conversation? How fast can you solve the customer's problem?
The good thing is that, despite the fact that in human-to-human conversations there's no gold standard—a chitchat, if it's a good chitchat, would actually change the relationship between the two—in a B2B conversation, there's no good chitchat. It's just, "Did you solve my problem or not?"
So it's a much more well-defined problem. That's why I think we see a lot of success in applying coding agents on top of it, to be able to automatically optimize it.
Very cool. I guess one of the issues with agents—I mean, this has been a dream of mine for ages—is: what if we could almost build an artificial organization with some agents and some humans, and then solve the coordination problem between them?
Yeah.
So we have some agents that are specialized in front office and some that are specialized in back office and coding, and they're all just sending messages to each other, fixing things.
Yeah.
That's very exciting, but it's quite problematic as well because—
It is.
—there are all sorts of failure modes. How would you approach that?
It is, and I think enterprises don't feel comfortable about it yet. Despite the fact that we have a lot of vibe coders and a lot of AI-native companies rushing to build all these features, to be honest, a lot of these products are really, really cool. But I think enterprises are feeling very threatened by that.
Yesterday, we actually had one of our largest utility customers talking to my team. My team is really freaking out because they say, “In the future, all the PRs will get merged into the repository. We need to have a review, and the review cycle is 2 weeks, and then we have to make sure that everything that gets merged into production is auditable and reliable.”
In the past, it wasn't really the case because we were just extending the agent and making the agent really, really good for their use cases. But now they are applying that control on top. I think they have been holding off for a very long time, and now they feel it's time to do it because the agent is really good right now.
But just because enterprises want to have their visibility, their guardrails, and everything to be auditable, you kind of have to build around that as well. So a coding agent that directly spits out PRs and auto-merges would not fly with them.
That's why in our platform, we also build a lot of visualization components with the agents. It's like, “Hey, this is what the flow looks like. So if the agent enters here, this is what the process looks like.” There's a bunch of tools, and these tools are exactly what the agents can do. So anything outside of that tool set, the agent doesn't have the ability to do.
There is a citation module of the model, so every response comes with a citation and they can have an audit trail for that. I don't know how long we'll have to have this because we kind of live in a world where we have a lot of vibe coding, and coding agents can do things very fast, but we are still dragging these traditional SaaS enterprise no-code flows with us.
I think the good thing now is that these no-code flows are no longer the primary interface for people to build. They're just for you to audit the actual processes of the AI agent. I think we are kind of shifting towards that way, but that UI component and visualization is still very important for enterprises.
Adaptation is intelligence, so these models are being adaptive. But at the moment, we've been talking about skill-surface adaptation and code adaptation. The great thing about weight adaptation is that you can, in principle, share this model in the organization, because if you're adapting code and skill surface—
Yeah.
—and you've got agents everywhere, it's a little bit brittle in the sense that I can't really take my code here or my skill here and share it with an agent somewhere else in the organization.
Yeah.
So how do you think about this adaptation problem in general?
There's always this debate: where do you want it to change the behavior, and where should the customization go? To a certain degree, we agree with that principle because we do actually train our own models as well. We agree that there's just inherent behavior.
Of course, you can try to solve it from the agent-harness side, but even if you try to solve it there, you're probably going to put a lot of wrappers around it, and it's probably not going to be good enough for your use cases. If you have the ability to alter the weights, then you have a lot more control over what you are building.
So I think it's more of a philosophical question: to what degree do you think you should have control? Because there's definitely some danger in terms of going that deep. You say, “Yeah, you can control the weights,” but maybe you don't know how to do it the best way. You end up with a suboptimal sort of solution.
For experts like them, I'm pretty sure that they know what they are doing. But I would say that for a lot of enterprises, that's probably not their primary area of expertise. So I think that would require some help, even if they wanted to fine-tune the model.
Usually, what I would suggest to a lot of our customers is that they start with the harnessing, and if they really cannot get around that, then they can think about using different models or tuning their own models. But that should be a second resort. That should not be your number-one priority immediately.
8. Cognitive Debt Changes How We Work
Another thing I wanted to talk about is this notion of cognitive debt. I've been using OpenAI. You can now link your phone to it, and I can control all of my coding agents, and it's amazing.
I can go down the shops. I can take a walk along the road, and in some ways it's amazing; in other ways it's not. The reason it's not is that when I go back to my computer, I realize that sometimes the agent was doing the thinking when I thought the voice agent was doing the thinking. One of the underlying coding models should have been doing the thinking.
Sometimes it wasn't using my established agents; it was just creating new ones, and it was a little bit messy.
Yeah.
And cognitive debt is also about how it's really important sometimes just to see the visual information in front of you. So how do you address that problem?
It's a huge problem. I mean, it's a huge problem not just for product development and our customers; it's also a huge problem for us organizationally. All our teams are adopting Coco right now, and we have a lot of people integrating with Gmail, G Suite, and Slack channels.
People are now starting to send messages in Slack channels to each other. If people are not careful, they're sending a wall of text to the other person, and the other person will be like, “Oh, I have to spend so much time reading it.”
Over time, we see more people just say, “Okay, now have Coco read all the text and then summarize it back to me.” It's hard. That's the problem.
Humans don't have that much capacity. The bottleneck has already shifted from producing content to actually validating and auditing the content, and that has been the problem my engineers have been telling you about since the beginning of the year. We have so much code to review, and I don't trust agents to do code reviews yet because I don't feel comfortable with them just merging the code.
So we're definitely entering that stage, and we probably need a lot of education for people. Especially in the co-working scenario, you have to educate people: you have to help people summarize your point. And also, don't send people the AI slop that you haven't read yet. You have to read it yourself and then really understand and agree with what you send, not just throw that out to people. So I think it's a real problem.
I think in coding it's particularly acute because you're just dealing with something of very high complexity. There's lots of background knowledge. It feels like, in a customer service agent or something like that, both parties understand enough about what's being spoken about.
Yeah.
And is that part of it? Is it something like, if we both understand it, then it's fine because the voice agent is a little bit lower resolution, but it's pointing to things that we already know? But in the domain where we don't know something, then there's that cognitive debt problem.
Voice agents have less of the problem because, by nature, those voice conversations are a lot shorter. It doesn't actually generate a wall of text. Our average output token size in the production agent is probably about 50 tokens. It's not a lot, right? So consuming those is relatively simple.
It's the reasoning behind it. If you are prompting a larger agent to do the reasoning task, where it generates a lot of reasoning tokens, that's actually more of a problem. Usually, a voice agent by itself is not a problem.
Voice agents have a different problem: they don't have sufficient bandwidth in that communication channel, because voice requires very ping-pong-y, very short, that kind of conversation.
So it has a slightly different problem. I think the task you are doing is actually very interesting because I have been thinking about using a voice agent to do vibe coding as well. That would be really cool. I think that would probably require a little bit of harnessing into it to actually make sure that the voice agent always surfaces the necessary high-level bullet points, rather than just passing on a simple sentence per se. But I don't know whether it will be easy enough for you to really understand the full picture still, because that's a lot.
What I really love about voice agents, or just agents in general, is that I can do insane amounts of context switching.
Yeah.
I can be doing my email triage in one, video editing in another, and coding in another. It's really made me change my workflow. So now I have a breadth-first search workflow where I'm just popping off one, then popping off another one.
And that's good and bad because, in a sense, I'm not going really, really deep on things. I'm just doing many things at once. But the beauty of the agents is, as we were saying earlier, these domains have become well enough specified that the agent can have some autonomy, and I just sort of feed information to every single one. So it's kind of changing how I work. Have you noticed that?
I build a lot of these agents as well. Currently, my workflow every day is that I probably open only 2 or 3 pieces of software: Claw, Slack, and Chrome.
Sometimes I have to open Chrome so I can log into some software, so my agent can use the computer through it. It's not really me using it myself; it's the agent, so the agent can access the tool as well. So it's already changed my workflow completely.
Yeah. And do you think this is a mindset thing? Obviously, we're engineers, and what I did as quickly as possible was transform all the work I do into an agentic representation, basically. The million-dollar question is, we want to get to a point where voice agents are used for basically all work that can be done. How do you see that transition happening?
Voice is a very good input modality. Whether the voice agent itself is the best working intermediate layer, I'm not actually sure, to be honest. Let me give you an example. I've been using Whisper Flow myself. I'm pretty sure you use it as well. It's a very convenient ASR tool.
Oh, Whisper Flow.
Yeah.
Yeah, I love it.
It's amazing.
Yeah.
Right?
I'm in the top 1%.
Yeah, exactly. It's so convenient. You can actually just talk to the agent, and the agent, despite the fact that I'm not a native speaker, understands the errors in the transcription. The transcription is not always accurate, but those agents understand those errors still.
It's not like I send a voice memo to someone, and then they see that there's a lot of mistranscription and get confused. Agents don't have that problem. An agent will do the reasoning and say, “Oh, this word should mean that. This word should mean that. Okay, well, I'm going to do it that way.”
However, I am a user who uses it heavily. I use it even at work.
Yeah.
At work, you put your headphones on, and people are sitting next to you. There are already people taking calls at their desks. What's the problem with putting on headphones and talking to an agent? I got judged by so many people who said, “What are you doing? That's so weird.” I was like, “No, I'm just working.”
So I think voice, as I said, has this privacy layer where some people just don't feel natural doing that. Even me talking in the office, just talking to my agent, people find it weird. That's the same argument as the Meta glasses, right? You can wear them and actually talk to your agent. I got one as well. I started using it, and people were like, “What's going on?” I was like, “Hey, Meta, play the music. Take a picture.” They were just like...
People find it weird. So I think there's this user-adoption issue. I would say we're rare because we're really agent-built. We really like to work with that new modality. But I think a lot of people probably still aren't quite there yet.
We are engineers, and what I do is build tools to help tasks become agent-enabled. For example, you can't use an agent to edit something in DaVinci Resolve, or sometimes I make 3D animations in Blender. You can do that with an agent, but it's always going to go wrong.
So what I do is build a tool and a mental model—an abstract model of the domain—and then I build an agentic CLI. I put a skill prompt in there, and then my agent can do it. I'm always in this modality of tool-building to make an application amenable to agents.
It is.
And that feels like the gap at the moment because my mum wouldn't be able to do that. We want to get to the point where they can almost demonstrate. They can say, “This is something that I do,” put a demonstration in there, and describe abstractly what their thought process is. Then they can build an agent interface, and from there, the agent can do it, and you can correct it when it makes mistakes.
I have been building quite a lot of tools for my internal teams as well because people don't know how to do these things. I think one of the most useful tools I built for the team was an internal wiki to track all the product features, functionalities, sales, how we pitch to customers, and that kind of thing, and then continue to iterate on that.
I also built another agent that does online research to search for suppliers, vendors, competitors, and market trends, and then just builds that into the wiki. That's the context engineering I'm trying to do for them.
There are also a lot of tools you can build around it, like access to your email and your Slack channels. I'm going deep because Claude has a lot of these default connectors. I almost throw out all those connectors because, first of all, they're not fast enough. Second of all, they're not powerful enough. I want to go straight to the source, so now it's really powerful.
I think this just requires a lot of that kind of work. A lot of people just want to use it. They don't necessarily have the mindset of building something and using it themselves. They're kind of waiting for people to build it for them.
And what do you think the next decade of voice research is going to look like?
The voice agent needs to become increasingly more end-to-end. Number 1, turn-taking must be end-to-end and collapse into the entire pipeline, because that is the single most annoying part when it comes to engineering and the most difficult part to actually get right. This is already happening with a lot of these models.
In the voice modality, we'll probably have to start branching out. There should be consumer voice for consumer entertainment and day-to-day use cases—the Alexa and Google Assistant world. There are enterprise voice use cases, starting with the enterprise contact center, more customer support, and professional use cases.
It could also be meeting note-takers, and an agent to help you coordinate these meetings could be part of that scenario as well. Then it could also become the internal productivity layer of voice, between your coworkers. I don't know what that would look like yet, but I do think that the voice channels required for consumers and enterprises are very different.
One requires a lot more professionalism, regulation, and auditability. The other one is really about fun, because conversation is also supposed to be fun sometimes. So I think there should be these 2 different branches coming out.
Then we'll see more hardware devices that can carry these voice agents, because voice is a natural modality for IoT devices. We'll probably see a second wave of IoT coming along. How successful that will be, I don't know, but I'm really excited about that.
I actually bought an open-source robot recently, more like a voice-speaker robot. You can program a lot of it, but it's still very clunky. I think all these things will become a lot more prominent in the next 10 years.
I can imagine a future where the technology becomes so invisible that we can train the system. It’s almost like teaching a child. And then I suppose a really important thing, certainly in terms of organizations, is how we share those learnings.
Currently, it does actually feel like the technology is still very visible because there’s still a lot of learning curve. In our organization, people have been talking about AI slop because it’s very visible. More and more people are sending AI slop messages to each other, so it is still very visible.
I do actually agree that eventually it should just become invisible, in a way that it’s blended into your organization. I think that will probably still require a bit more iteration in terms of sharing the context, sharing who you are, and what the organization is about. Despite the fact that we’ve done a lot of internal knowledge architecting, there are still a lot of principles that we haven’t written into those MD files yet.
So I think there’s still a lot to do, and we’re probably running out of the low-hanging fruit now, because some of the things that are actually so personal to ourselves or to the organization are more spiritual. It’s about these kinds of constitutions or principles that you probably don’t print out everywhere.
Because the funny thing is that if you have the agent ingest these cultural values of your company right now, sometimes the agent just comes up with random ideas and overly focuses on one thing. But it’s not actually about that. Despite the fact that we write our cultural values down like this, there’s a reason why they came about, and those reasons a lot of the time aren’t injected into the context.
So agents oftentimes focus on the wrong things because, in their weights, somehow they think that out of these 6, this one is the most important one. But that’s actually not true for the organization. So we kind of have to share a lot more in order to get to that invisible state that you’re talking about.
The slop thing is very interesting, and my definition of slop is generation without competence. I’m using memory systems, adaptive specialization, skill adaptation, and so on. So it always has significantly more context. Increasingly, it knows what my preferences are. I’ve grounded it with the correct, technically referenceable materials, and so on.
So the fact of the matter is that there are 2 modes of AI. There’s AI that has sufficient context and knows what to do, and that AI does not produce slop.
Being slop, unfortunately, is probably a little bit more subjective, in a way. I engineer and context-engineer my agent so that the agent doesn’t feel like slop to me, but it may feel like slop to the other person who is reading my message because, in their context, the things I care about probably don’t really matter all that much to them. Or maybe they’re not aware that this is important yet.
So when they read it, they think it’s slop because they don’t actually see the context behind the reasoning. That is a bit of a subjective problem. So I think we have this internal debate or argument: should I care about whether I’m generating slop and then just sharing it with my colleagues?
The fundamental debate is whether a human being should be reading those things anyway, or whether a human being should just ask an agent to read it for them and then summarize it in a way that’s better suited for their focus. I almost landed on the idea that I should probably build a skill for everyone’s agent, based on their personal context, to summarize whatever message people send to them, so they can extract the useful information for themselves.
In that case, you want to encourage the output side to be as rich as possible. You send a wall of text, and that’s okay, because on the other side there’s an agent summarizing it for that particular individual. But I kind of feel like this is running in circles. It’s like we humans now need an agent in between to help us cherry-pick the useful information for us.
John, it’s been amazing having you on the show. Thank you so much for joining us.
Thank you. Yeah.