How a Voice Agent Learns the Rhythm of Conversation — Shawn Wen
Tim ScarfeTsung-Hsien (Shawn) Wen
- Enterprise voice remains an unsolved systems problem even as speech components become cheaper and more available. Early ChatGPT wrappers simply piped ASR, an LLM, and TTS together, but Wen argues that voice adds time, noise, interruption, and adaptation: it is “one step to the real world,” where intelligence means responding to how someone speaks, not merely reasoning over what they said.
- The winning architecture may be end-to-end for listening and turn-taking while retaining text as the controllable enterprise interface. Wen’s audio-native model streams audio directly, predicts whether the user has finished, then emits an answer or tool call, citations, and finally a transcript. Text output preserves guardrails and auditability without forcing the system to wait for transcription before responding.
- Production conversation data is a defensible advantage because the hard cases barely appear in clean laboratory audio. Their platform draws—with customer agreement—from contact-center deployments across banking, utilities, logistics, restaurants, hotels, retail, and outbound sales, redacting personal data and inserting synthetic identities. Training deliberately includes crosstalk and combinations such as a name spoken over “a baby crying in the background”; Wen says old denoising pipelines can make the new model worse by creating an environment that is “too sanitized.”
- Latency is principally a turn-taking problem, not raw model-inference speed. A fixed endpoint threshold either interrupts slower speakers or leaves impatient users waiting, whereas an audio-native model can use pace and unfinished semantics to recognize a mid-sentence pause. The company combines that with precomputation, self-hosting to reduce P95 network delays, and “latency-budgeted reasoning” that asks whether another second of thought will materially improve the answer.
- Voice identity directly affects performance and trust, but enterprises remain wary of excessive anthropomorphism. Generic voices inherit the baggage of broken IVR systems, while a modest regional character—such as a Newcastle accent in the UK—can signal a familiar human service context. The company once found its most natural English voice was a French voice forced to speak English: “It’s the imperfectness of that voice [that] make it real.”
- Evaluation and control—not model access—are emerging as the enterprise value layer. Wen criticizes voice benchmarks that ignore latency, tool use, citations, name recognition, and task completion; his company instead evaluates on real conversational data and planned to release that benchmark within “the next month or two.” Enterprises may switch among model APIs, but many still want to own the harness that encodes their brand, workflows, guardrails, and audit trail.
- AI abundance shifts the organizational bottleneck from generation to supervision. Wen’s engineers already face more code than they can comfortably review: “The bottleneck has already shifted from producing content to actually validating and auditing the content.” The longer-term market may split between playful consumer voice and regulated enterprise voice, but adoption still depends on privacy norms, cultural context, and whether agents can become invisible without producing unaudited “AI slop.”
1. Voice adds time and adaptation beyond increasingly commoditized speech components
Wen initially shared the widespread post-ChatGPT expectation that voice would soon be solved. Wrapper companies connected ASR, ChatGPT, and TTS, yet real deployments revealed a missing dimension: voice is “one step to the real world,” where timing materially changes what a response means.
Text intelligence benefits from abundant documentation and online papers; machines have far less observational data about embodied or spoken interaction. Wen’s distinction is the episode’s foundation: “The intelligence here is not really about reasoning, it’s really about adaptation.”
Tim Scarfe’s robotics analogy makes the difficulty concrete: making coffee or pouring beer is effortless for a person but difficult to specify for a machine. Conversation has the same hidden complexity—humans generate enormous amounts of it, but “machines just never [were] able to observe” the relevant rhythm, hesitation, and adjustment.
2. Enterprises demand human capability with robotic predictability
The enterprise brief contains an explicit contradiction: customers want branded voices that pronounce company names exactly as intended, and they want an agent “super powerful like a human being” yet “super controllable like a robot.” Expectations often benchmark the system against the company’s best employee, not its average operation.
Wen sees efficiency-focused deployments as the practical adoption wedge. Once an agent succeeds in a bounded task, customers begin to conclude that it is “better than I thought,” creating room for broader use without requiring human-level performance everywhere on day one.
Frontier speech-to-speech systems can produce demos that feel like “singularity already,” but contact centers expose different failure modes. An older caller may pause because she is confused, only for an aggressively tuned agent to keep speaking over her—precisely the experience a polished consumer demonstration does not test.
A cascaded stack separately handles recognition, language, and endpoint detection, forcing engineers to tune brittle parameters or bolt on rules such as detecting an older speaker. Collapsing those functions lets one model perceive slow pace, background noise, crosstalk, and hesitation, then learn the appropriate response from deployed conversations.
3. Audio-native input and text output create a controllable contract
Wen’s model is audio-native but not waveform-to-waveform: an LLM consumes streaming audio representations and emits text. Wen keeps text because enterprises can apply guardrails to it more reliably; unconstrained speech output is “a lot fiddlier to deal with.”
As each audio chunk arrives, the model repeatedly generates a cheap first token representing turn state: the user is starting, continuing, or finished. When it predicts the end, the audio is already encoded in the embedding space, so the LLM can immediately stream an answer or tool call without reprocessing the utterance.
The output order is intentionally reversed from a classic pipeline. The response comes first, followed by citations identifying the knowledge-base facts used, and finally the transcript for debugging and audit—meaning “you don’t have to wait for the transcript” before replying.
Retrieval also becomes a tool call rather than conventional transcription-first RAG: the model generates a search query when it needs outside information. Wen says the real innovation is the “IO contract”—the training pairs, data preparation, and prescribed outputs—while open-source base models can be swapped and retrained as stronger candidates arrive “every month, even every week.”
4. Real noise is training signal, not contamination
Their platform already receives relevant audio from deployments spanning banks, utilities, logistics, restaurants, hotels, retail, and outbound sales. Customer permission remains non-trivial under GDPR and personal-data constraints, so the company redacts real identities and generates synthetic names and addresses; the goal is to learn conversational mechanics, not customer facts.
Audio is represented through mel spectrograms streamed segment by segment, making perception resemble image recognition. Scarfe recalls how trained humans can read speech, harmonic instruments, percussion, or salt-and-pepper noise visually; neural models can likewise learn patterns from that representation.
Noise augmentation is deliberate. Wen’s example combines a synthesized pronunciation of his Mandarin name with realistic production disturbances such as “a baby crying in the background,” multiplying the environments the model can experience without exposing private conversations.
Counterintuitively, legacy denoising and pipeline sanitation can degrade the integrated model because the cleaned signal is unlike deployment audio. For crosstalk, their model currently uses no specialized diarization mechanism; the fused model can instead be prompted to ignore background speakers, follow the main participant, or join a multi-speaker exchange, depending on its training and instructions.
5. Natural latency comes from knowing when not to answer
In cascaded systems, endpointing is governed by a coarse wait-time setting. Set it short and the agent interrupts slower or older callers; set it long and faster speakers assume the system is unresponsive. One global threshold cannot represent every user’s conversational tempo.
The integrated model sees both pace and meaning. If a speaker pauses before completing the semantics of a sentence, it can keep waiting; meanwhile, incremental encoding precomputes work, and hosting the model beside the rest of the company’s platform removes network travel that disproportionately damages P95 latency.
Wen wants to reinvest that saved time in “latency-budgeted reasoning.” The model estimates whether stopping now would still yield a good answer and is trained with constraints such as, “Do you actually need to think that much time? Can you cap it to, like, one second?”
Longer backend calls still require visible feedback. Fillers and typing sounds work until they become repetitive, while hold music buys time for deeper reasoning or slow enterprise APIs. Silence is dangerous: if the AI goes dark for even “3 or 5 seconds,” callers may panic and wonder whether it is working, particularly after years of unreliable automation.
6. Familiar imperfection often earns more trust than generic polish
Most enterprises do not appear to want an assistant pushed far toward indistinguishability from a person; they want it to solve the problem while carrying a restrained brand identity. A South American casino explored a strongly localized opening voice but ultimately declined to launch it because it “feels a bit too much,” illustrating how enthusiasm recedes near production.
Generic voices often perform worse because callers associate them with rigid IVR prompts that repeatedly failed to recognize “payment” or “transfer.” Regional cues can reverse that expectation: UK users may trust a Newcastle voice because many contact-center employees are from there, signaling that they can state their problem and have the agent solve it.
The company’s early English voices sounded less natural than a French voice made to speak English. Its slight accent supplied the useful imperfection. Wen therefore treats voice selection as brand- and region-specific, not as a universal race toward maximum smoothness or a single supposedly neutral voice.
7. A credible voice benchmark must measure the whole interaction
Wen argues that standalone ASR and LLM benchmarks miss the multidimensional nature of an agent: response latency, transcription, factual quality, reasoning difficulty, personal-name recognition, instruction following, citations, and tool calls all matter. “A voice agent benchmark without latency consideration is not a real one”; a caller cannot wait 60 seconds.
Public datasets are often synthetic because real voice is private, and existing tests do not sufficiently represent production. The company built an internal benchmark from consumer conversations and Wen says its model is currently the best on that benchmark; his stated plan was to release it to researchers and the market within “the next month or two.”
Subjective experience remains harder than factual scoring because accent and tone vary by culture and task. At production grade, Wen prefers behavioral evidence: did users engage, complete the task, and obtain a satisfactory resolution? Those outcomes evaluate the full model-plus-harness system rather than the model alone.
8. Enterprises will own boundaries and workflows before they own weights
Enterprises increasingly view AI as technology they must own, but model ownership requires tuning expertise, GPUs, and budgets many lack. They therefore accept switching among model APIs while treating the harness—the prompts, tools, brand behavior, and process logic—as the attainable form of proprietary intelligence.
Wen’s proposed boundary is a fast voice agent as the customer-facing “front door,” delegating difficult work to the enterprise’s own “mastermind behind the scenes.” He expects companies to discover that they need not own every layer, though some will retain custom harnesses when the task is sufficiently critical.
B2B conversation is unusually amenable to optimization because its specification is clearer than social chat: was the problem solved, did the customer feel good, and how quickly? Wen contrasts this with human chitchat, where success may alter a relationship and no comparable gold standard exists.
Autonomy still collides with governance. One large utility customer said future PRs would need an auditable review cycle lasting two weeks; automatic merging would not be accepted. The company therefore retains flow visualizations, bounded tool sets, and per-response citations—no-code-style interfaces now serving as audit views rather than the primary construction surface.
9. The next bottleneck is human attention, context, and adoption
Scarfe describes “cognitive debt”: agents can work while he walks, yet returning to the computer may reveal duplicated agents, wrong delegation, or reasoning nobody supervised. Wen sees the same organizationally as AI floods Slack and code review: “Don’t send people the AI slop that you have not read yet.” Humans must understand and endorse what they transmit.
Voice limits that problem but introduces another. Their production replies average roughly 50 output tokens, avoiding walls of prose, yet speech is a low-bandwidth, “very ping-pong-y” channel. Scarfe’s breadth-first workflow can coordinate email, editing, and coding agents, but Wen doubts voice alone can communicate the full picture without a harness that surfaces concise high-level bullets.
Making agents invisible requires richer organizational context, not just more text. Written values omit why those values arose, so a model may overemphasize one of six principles. Even “slop” is recipient-dependent: a message optimized for Wen’s context may feel irrelevant to a colleague, creating the circular prospect of agents generating rich output for other agents to personalize.
Wen expects end-to-end turn-taking to become foundational, then split into two branches: playful consumer assistants and professional, regulated, auditable enterprise voice. Meeting coordination, internal productivity, and a “second wave of IoT” may follow, but he is explicit about uncertainty: the technology could mature within ten years while social discomfort around office speech, privacy, smart glasses, and surrendering control still slows adoption.
Full transcript
They want the agent to be super powerful, like a human being, but they want the agent to be super controllable, like a robot. Voice AI is very important because that's the native way humans talk to each other, and it's the most natural interface. Computers haven't been at the same level as humans in terms of voice conversations yet, but we do foresee that this should be the future in the next couple of years.
Do you think that speech technology has become commoditized, or is that just completely wrong?
It has become a lot more commoditized in some ways, but in many other ways, it's still not. Modeling a real-time conversation is still hard, and it's not quite there yet. The bottleneck has already shifted from producing content to actually validating and auditing the content.
People think that voice assistants are a solved problem, but it's actually much harder than people think. Why is that?
1. Voice Is Harder Than Text
When ChatGPT came along, we thought that voice was probably going to be solved very soon. I think a lot of people had the same idea, and that's when you saw a lot of wrapper companies come along, wrapping around ChatGPT, piping ASR and TTS together, and trying to solve the problem.
The reality is that it's actually harder than we think. If you think about voice, it's actually one step closer to the real world. All the current innovation on the AI front is predominantly in text and the digital world, where solving a PhD-grade physics problem seems not to be a problem anymore because you have really good documentation and a lot of online papers you can research.
In the voice scenario, you have an additional axis, which is time. Time is a very critical matter, and I think we have not trained the model to understand that concept better yet. Because of that one step into the physical world, we started to realize that there's a lot more complexity when you add that time element.
With humans, intelligence here is not really about reasoning; it's really about adaptation.
When we spoke last time, you said that doing voice technology is a little bit like robotics.
Yeah.
Right? So it's incredibly easy for humans. There's the famous Wozniak coffee test: “I want a robot to make me a cup of coffee.” Just tell me a little bit more about that. Why is it so difficult?
Again, it gets back to the physical-world situation. In the physical world, humans collect data very naturally, and we have a lot of data because we have visual data and body movements. Your hands can reach out to things in the world.
However, computers don't have that yet. All the tools computers have are about how to use computers. I think we have trained the models to be really good at understanding text and understanding how to use tools in the virtual world. But in the physical world, the data is still very limited, and that applies to robotics.
That's why a lot of robotic companies right now, as you said, are trying to grab a coffee from one place and move it to another. I remember one of my really dear friends, and also a really good researcher from Cambridge, who was originally at Poly and has now left to start his own B2B business. He wanted to start a robotics company.
I asked him, “What is the first application you're trying to build?” He said, “I want the bot to be able to pour beer for me. That's test number one.” I was like, “Oh, wow, okay.” It's just really, really hard.
Then you look at self-driving cars. Elon has been promising the world that this is going to be fully commoditized. I think we're starting to see signs of that now. Voice is a very similar problem, except that we have so many conversations with one another, and machines have never been able to observe them.
What do enterprises actually want from a voice agent?
2. Enterprises Want Power And Control
Enterprises want a branded voice. They want it to be able to pick up their identity. They want the agent to say their brand name exactly the way they want it to. That's number one.
Number two, they want the agent to be super powerful, like a human being, but they want the agent to be super controllable, like a robot. It's actually a contradiction. I think that's the difficulty of deploying AI agents into enterprises, because people's expectation of AI is that it needs to be as good as a human—actually, as good as the best human on their team.
That's why the adoption curve is always harder to justify, unless you can find a particular area where you're saying, “Hey, this is more about an efficiency play.” Then people start adopting it and start to feel, “Okay, actually, the agent is better than I thought.” That's what we're seeing from enterprise adoption right now.
Why did you decide to build your own custom model rather than using some of the existing ones out there?
3. Why They Built Their Own Model
We've been doing research and building our own models for a very long time. We previously built our own speech-recognition model, and we also built our own large language models.
Around 1 year ago, we decided that cascaded systems were going to be outdated in the future, and that end-to-end models would be the future. We started doing a lot of experiments, and we decided not to invest in our speech-recognition models anymore because of that radical shift into the future.
When we started looking at the market, the reality was that all these latest speech-to-speech models were really good for demos. I can show you a really natural conversation, and you'll be like, “Wow, that's the singularity already.” But once we deploy them into contact centers, they break in so many different places.
The interruptions are sometimes too aggressive. If you're on the phone with an old lady who is talking to a bot for the first time, she might be constantly confused and stop, and then the agent just keeps speaking over her. That's a level of frustration.
We see a lot of these scenarios, and we say, “Hey, we probably have to train our own model for these use cases.” I think the frontier labs are not optimizing for what we care about: real conversational turn-taking for B2B business use cases.
For consumer use cases, I think you can build voice agents already. For enterprises, there are additional governance and enterprise requirements that we have to satisfy.
How adaptable could that be? For example, if I'm talking to an old lady on the phone, I would adapt my speech. I would slow down a little bit, and I would give her more thinking time. Could the models be adaptive in that way?
Yes. That's why you have to collapse all these modules together into an end-to-end model, or in some sort of end-to-end fashion. In the traditional cascaded system, you have a speech recognizer, a large language model, and turn-taking on top, which is responsible for detecting the end of a sentence.
These 3 models don't really work with each other very well, so you have to tweak the parameters. Traditionally, you can have another detector that says, “Oh, this is an old lady, so I have to tweak a parameter here and there,” but it's all very clunky. The reality is that it doesn't really work like that, and it doesn't work very well.
Therefore, it's almost necessary to collapse all these pipelines together so you have a large language model directly perceiving the audio. The audio can involve slower speech, background noise, crosstalk, and many other scenarios that you encounter when you're having a real conversation with a human over the phone.
Then you can collect that data so the model can adapt to it. I think collapsing everything is the first step, and then the initial model needs to be good enough so you can continue to iterate on it, because you use it to collect more data as well.
I suppose it all comes back to Rich Sutton’s The Bitter Lesson, which is that, in principle, maybe we could write the code to do this. But there’s an exponential number of failure modes—combinations of bad things that might happen at the same time. In this situation, we need to do this; here, there’s some crosstalk; or here, there’s some noise in the background. Just training a neural model end to end, you can cover a lot of those failure modes.
But I suppose it still needs to be a cascade in some sense, because it’s not possible to embed frontier-level intelligence in these voice models. But you could potentially draw a boundary around the voice intelligence model, and then it could call out to an agent when it needs more intelligence.
Yeah.
But I wanted to ask: can you sketch out the high-level architecture for this voice model?
Yeah. The high-level architecture is that this is basically an audio-native model. The base model itself is a large language model, so it actually outputs text. We’re not training models to do end-to-end speech yet. The reason is that, for enterprises, control is important: text is a lot easier to apply guardrails to, while speech is a lot more fiddly to deal with.
We train the model to output text, but at the same time, we allow the model to take ASR into its input channel. That means the model itself natively understands audio. The way you pipe audio into the model is by streaming audio chunks into it.
The way we structure the model is that it’s doing multiple things at once. First of all, based on the streaming audio, the model predicts a first token. The first token is basically the turn-taking signal. Turn-taking is basically saying, “Hey, is the speech just getting started? Is the user still speaking? Or has the speech already ended?”
That initial token is very cheap to generate. It’s just like the first token in the output. As audio comes in, we keep generating that token until the model generates, “Okay, this is the end of the speech.” Then we start having the LLM do the processing.
At that point, all the audio has already been preprocessed, because as we encode it and then predict the first token, it’s already in the embedding space. So you don’t need to encode that sentence from scratch in that case.
The next job of the model is to directly predict an output. The output is text, as we said, so it could be a response or a tool call, for example. That’s all streamed, like the typical ChatGPT API. We stream the response back, and at the end of that response, we then continue to predict a few things.
Mm.
A few things like citations. For example, if you have a knowledge base with a lot of different content, the model will generate a citation, saying, “Actually, for this response, I’m using this fact and that fact to make the prediction.” Then enterprises know that it actually has a reference there.
Finally, it also streams out the transcription. So we’re kind of reversing the order, because the model’s job is not to predict the transcription, but we still want it to predict the transcription for auditing purposes, debugging purposes, and things like that. But it comes last, because you don’t have to wait for the transcript to actually start responding to the user.
Even that’s quite an interesting trade-off to talk about, because many folks would have used the GPT live model, and that is a bidirectional waveform-to-waveform model. It’s really interesting. It just makes all of these weird pauses and sounds, and it sort of breaks the uncanny valley.
Right.
But actually, that might be great for me at home, but in the enterprise, you need to have constraints. You need to say, “No, you’re not allowed to talk about this, and you’re not allowed to talk about that.”
Conceivably, you could run separate ASR on the output, and then do retrieval-augmented generation and link to the knowledge base and stuff like that. But there are all of these trade-offs, and in enterprise, that’s not necessarily what you want to do.
Retrieval itself—you can also do retrieval with these audio-native models, because the audio-native model can handle it. But the retrieval mechanism is slightly different. You cannot use traditional RAG anymore, because traditional RAG means that you embed the user’s transcription and then do a vector search in your vector database, right?
But because now you don’t have that transcription, what we need to do is prompt the LM to generate a tool. The tool is basically a search query. It says, “Hey, now I think that in order to answer this question, I need to generate a search tool, and that search tool needs to come with that query.”
It’s pretty much like function-call-style search, which people call agentic search. That’s just another buzzword, but it’s basically just a tool call, really.
The important thing is that our innovation is not in the model itself. We’re innovating at the I/O contract. That means: how do we prepare the data, how do we prepare the input, and what is the mechanism for the data feeding into the model? But also, how does the model output that data? And then how do we train the model based on these input-output pairs to make it behave exactly the way we want?
To be honest, the underlying model almost doesn’t matter all that much, because we know that in this market, open-source models come out every month, even every week. So we want to maintain our data and our model, and then have these training scripts that are always ready. Whenever a new model comes in, we train on top of it to see which one has better performance.
And voice data is much harder to collect than text data. How are you doing that?
4. Real Conversations Train Better Models
For us, we don’t have to collect it, because our platform generates a lot of voice data. We’ve been deployed in enterprise contact centers.
Yeah.
We have banking clients, utility clients, logistics clients, restaurants, hotels, and a lot of retail outbound sales use cases. We already have a lot of that data, and that data is exactly the kind of data we want to use to solve the problem, because we serve predominantly enterprises at the moment. Therefore, we leverage a lot of this data.
Obviously, getting agreements with enterprise clients to train those models is nontrivial because of data privacy, GDPR, PII, and all of those issues. But kudos to our legal team: we actually get quite a lot of our customers to agree to that because they know that, if they agree, the model can continue to help improve things for them.
The way we train the model is that we redact all the PII data, so none of the customer data is actually exportable post-training. We also produce a lot of fake PII data, so the model can still learn how to identify names, addresses, and things like that. But really, the data is just to let the model learn the mechanics of the conversation rather than the actual user data.
Very cool. And I wonder: does the model currently take in pure waveform data, or does it take in featurized data? I just remember I was doing audio stuff for my PhD, and back in those days, it would have been unimaginable to put in waveform data. And now—
Yeah.
Now it’s simply a case of if the model’s big enough and you train for long enough, you can.
I think currently the standard audio encoder is that you feed in the mel spectrogram, right? It’s still like this image. You remember, there are different spectrums and different frequencies and things like that. It’s still that representation.
But you feed it into the model segment by segment, right? To the model, it feels more like an image-recognition task. I remember that you’re from a speech background, so I remember there was an MIT professor. I think his name was Victor Zue or something.
Back in the 1990s or something, he said, “You will be able to identify what human speech is by just looking at the spectrogram.” He was like, “This is exactly that for the particular syllable,” or whatever. And people were like, “What do you mean? There’s nothing you can identify there.” So, in fact, a machine can learn those patterns really well from it.
I love looking at spectrograms—
Yeah.
—because there are so many things that you can immediately see. For speech, or if it's a violin, for example, you just see these horizontal lines—
Yeah, it's really cool.
—and you can see that there's a harmonic series. If it's percussion, it looks like a sort of vertical bar in the spectrogram. But yeah, after a while, as a human, you know that you've been looking at these for a while, and you can just read it like a book.
Oh, totally.
It's fascinating. And noise, of course, just looks like salt and pepper.
That's right.
But on the subject of noise, this is quite interesting. How do you guys deal with that? Is this the kind of thing where you deliberately sample noisy data, or do you augment your data and add noise to it?
Yeah. It's a very deliberate decision to add a lot of different noise to the training data set. To be honest, our training data set by itself already contains a lot of different noise because it's real conversation data. But we typically apply another generative model on top of it.
Say that I wanted the model to pronounce my Mandarin name, Zhong Shen Wen. In doing that, I also ask these TTS models to add the additional audio data that we see in real production. The outcome of that generated data set could be a robot pronouncing my name while there's a background baby crying, for example. These things can be simulated so that you can have a lot of variety.
Thanks to modern generative AI models, you can create a lot of very realistic data. So the answer is yes, we do a lot of this, and it's a very deliberate approach because we want the model to be robust against all these different environments. But even the training data itself already contains a lot of that.
The funny thing is that, because with the previous technology we were doing a lot of denoising and pipeline sanitization beforehand, all those additional components in this new dialogue-reasoning model are actually making it worse, just because the environment is too sanitized. It hasn't seen that before.
I remember for many years there was this phenomenon called the cocktail-party problem—
Yeah.
—which is that when we're at a cocktail party and there are many people talking at the same time, we can focus our attention on one person, and our brain can filter out the other people. ASR models couldn't do this.
Right.
But increasingly, they can do it now because maybe internally they're learning, "This is speaker 1, and this is speaker 2," and they focus on speaker 1. Crosstalk has always been a problem. How are you doing with those failure modes?
Crosstalk in our model currently doesn't have a specialized mechanism for it, but the architecture is suited to tackle that problem. Speaker diarization technology was previously predominantly applied to speech recognition, and it always assumed a cascaded pipeline solution. You tag someone as speaker 1, speaker 2, or speaker 3.
But it's not a very natural way of doing it because, as a human, if you're at a party listening to so many people, you don't actively label them: "This is speaker 1, this is speaker 2, and this is speaker 3." You think, "That's another person. That's another person I shouldn't pay attention to," because there are so many of them. Some of them will stop speaking, and then you'll lose track of them for a while.
I think traditional speaker diarization is trying to solve a very different problem. With a dialogue system, what you really need to do is have the agent focus on the right subject—the person it is talking to. Because you now combine the language model with speech recognition, the language model can natively perceive that there's background crosstalk somehow.
To be honest, this language model is also promptable. You can prompt it by saying, "Don't respond to the background crosstalk apart from the main speaker you're speaking to." You can also say, "Whenever there's crosstalk in the background, take active participation in that conversation as well." Technically, you can do all of this.
The major difference is how you train the model. How do you collect the data? What kind of data do you collect, and what instructions do you give the model? I think you watched the GPT live demos where you have a lot of these old ladies coming along, and the model is able to identify a new speaker. I'm pretty sure that in their prompts they're saying, "Now you're in this party-like conversation, and you have to gradually identify speakers as they come along." Once you fuse these models together, they have the ability to do that. The major difference is how you train it and what you're training it to do.
The turn-taking thing in particular—we spoke about that earlier—is really, really difficult. I suppose we should imagine that, in an ideal world, in 10 years' time, when this technology has reached the utmost maturity, maybe we can have microphones in every single meeting room. When anything private is discussed, it will know not to pay attention to that.
There might be multiple people involved in the conversation. You can address the agent, the agent can get involved, and the agent knows who's talking. At the moment, we're kind of solving it in the one-on-one case—
Yeah.
—and we're solving this turn-taking problem. Is that a reasonable estimation of what the future might look like?
I think the technology will be complete. Your prediction is right. Ten years from now is a lot of time. Given the speed at which all this technology is improving, I think it's completely reasonable for the future.
The only thing is that voice modality—I mean, people say it's the most natural modality for interacting with a computer. The reality is that we've had several iterations of these kinds of personal assistants in the past. It started with Siri, then you had Google Assistant and Alexa. You have in-car voice assistants as well.
Maybe the technology wasn't really mature yet, but there's also the fact that voice is a lot more personal and private to you. I don't know whether consumer behavior will be there yet, because that means you have to surrender a lot of your control to the agent. A lot of people will say that's a great thing because the agent can do a lot for you. But I think there will be some consumer adoption curve that we still have to climb, even though the technology is here.
Can you tell me about how you did latency engineering? We're going to talk about this uncanny valley and all the things we need to do to make it possible for people to suspend their disbelief because they're talking to a voice agent and it's suddenly really good. But latency is probably one of the most important things. When you reduce latency, there's always a bit of a trade-off: you might be trading off some of the intelligence. You guys have best-in-class latency, so how have you done that?
5. Latency Shapes The Conversation
Latency, first of all, the major thing is that you need to use the right size of model. To be honest, if you're using the right-size model, the model processing itself isn't really a latency problem.
Historically, the cascaded system struggled with waiting to decide when the speech had ended. There's a very coarse parameter that you can tune: how long you want it to wait until the end of the speech. If you tune it to be too short, it could be too aggressive, and you're going to interrupt a lot of older people, for example. If you tune it to be too long, you'll be waiting for too long, and a lot of people will become impatient.
Everyone has a different threshold in terms of how fast they expect the conversation to be. I think that's the difficult part, because you're trying to set a programmatic metric for every single user speaking to your assistant.
With this kind of model, you have a lot more flexibility because the model itself is now predicting based on the audio frame. The audio frame contains how fast you speak, as well as the semantic information.
So if I'm pausing in the middle of a sentence, the semantics have not been completed yet. The agent knows, “Okay, I have to continue to wait until the person finishes the full sentence.” That is the major benefit of collapsing the ASR and the language model together: you have more natural turn-taking built into the model.
The model is now making that prediction by itself. A lot of the pre-cached processing, as we talked about earlier, is another way to optimize it. To be honest, the fact that we host our own models alongside the rest of our platform is another benefit, because you don't have network-travel latency impacting the long-tail use cases, where the P95 is the major problem for network latency.
We optimize a lot of this so that the agent can respond very fluently. However, my team is now adding more reasoning to the model as well. The good thing is that once you save a lot of time, you can use that time to do more things. Our team is looking into adding auto-reasoning to the model, so the model can decide, based on a particular user query, “If I stop my reasoning right now, how likely am I to give a good answer?”
We train our models on auto-reasoning in a way that we always challenge them: “Do you actually need to think for that much time? Can you cap it at one second and then respond to the user?” The reasoning we are training the model to do is very latency-budgeted reasoning. It's almost like going to a competition: someone gives you a question, and you suddenly start counting down.
It's so interesting because sometimes when I interview people and ask them a really difficult question, even though they're thinking, they still feel that they need to say something. Quite often, on a difficult question, the first sentence of their response is just filler because they're obviously thinking, and then they actually answer the question.
That's right.
Right. It's a similar thing here, isn't it? This auto-adaptive reasoning is really interesting because, from a user-experience point of view, how do we solve that problem? Do we make the model say, “Hmm, that's interesting,” and indicate that it's thinking?
That's right.
It might actually be a significant amount of thinking. Maybe it needs to delegate to an agent and do some things, and then come back to you.
There are a lot of tricks you can use in that scenario. In our real productions, you can add a lot of filler sentences, obviously. But if you repeat them too many times, it becomes very repetitive, and the user will know that you're just bluffing them.
Another way is to play a typing sound, so there's computer typing in the background to mimic the actual scenario. The real thing we're trying to solve in that scenario is that you always have to deal with API latency anyway, and some enterprise backends aren't the fastest by default. You have to bake in a lot of those things to make sure the user doesn't feel abandoned.
The reality is that users have been talking to a lot of these automated systems, and if the AI goes dark on them over the phone for even 3 or 5 seconds, they panic. They'll be thinking, “Is the AI actually working?” In order to have that trust, you have to provide constant feedback.
The good thing is that contact centers have a lot of ways to deal with this. An agent will put you on hold with music while they're thinking. In the old world, a human agent might actually go to different places and talk to their colleagues about how to solve the problem. The AI agent can put you on hold music while it's doing the reasoning.
I think there are a lot of tricks you can use there. But articulating things well is very important, because I think the biggest issue voice agents have in contact centers is that users don't trust them, simply because they've been burned many times in the past.
That's a really cool idea. Play elevator music while it's thinking.
Yeah.
I suppose this leads neatly to the next thing: what is the uncanny valley for a voice agent? We talk a lot about anthropomorphization and the risks of that. Is that something we actually want in the enterprise? Do we want it to sound indistinguishable from a human?
Different businesses would give you different answers, but I think the majority of businesses would probably hold the same position as Apple and Google, for example. Alexa is probably the consumer personal assistant that goes furthest in terms of personalization.
Mm-hmm.
They give Alexa an identity, and there's a very unique voice there. But that's as far as they want it to go, because if this personal assistant is representing a particular company brand, it's very hard to know how successful it would be. A lot of the time, companies still try not to attach too much personality to it. The goal is simply to make the assistant solve the problem first.
The majority of enterprises are still on the conservative side. We do have some customers—for example, there's a casino from South America that likes very local, very South American voices. They like the agent to use those voices to open a conversation, for example. In the end, they didn't launch that agent because they were like, “Oh, that feels a bit too much still.”
People would think about that once they felt more comfortable with the technology. But I think we're still in a phase where people are just building trust in the technology.
It's interesting how much we should or shouldn't chase the personality thing, because the reason it sometimes sounds egregious is that it might be normal in Texas or in a particular location, but it isn't normal to me. Do you think, in principle, there should just be a generic voice? Or should the agent get to know you, and when it knows how you talk and what you like, should it adapt to you?
I can speak for what's happening right now, but I can't really predict what people will want in the future. Currently, the majority of our customers don't like generic voices. Generic voices also perform much worse for consumers in general, because consumers identify that it's a robot, and once they identify that it's a robot, they immediately don't trust it.
We have these traditional IVR systems that have been deployed throughout America: “Say a few words. If you have a payment issue, say ‘payment.’ If you want to transfer money, say ‘transfer.’” You say it many times, and it still doesn't recognize you. Those voices are always very generic and a little robotic, because previous-generation technology couldn't really make a natural-sounding voice.
So consumers immediately connect a generic voice to the previous system they've talked to. They hang up, or they try to bypass the agent as much as possible. We've found that the greatest success comes when you don't use generic voices. You want them to have a little bit of a local, regional take.
In the UK, for example, a lot of contact centers and consumers like the Newcastle voice for obvious reasons, because many agents are from there. When they hear the Newcastle voice, they know it's a human agent. They can say their problem to them, and the agent can solve it for them.
I think those voice choices are quite nuanced but quite important. It depends on which brand it represents and what people expect the voice to be. The best-performing voices usually come with a little bit of a regional take.
The funny thing is that, in the early stage of shipping these voice agents, we found that all the English voices were not natural enough. The most natural voice was a French voice, and then you forced the French voice to speak English, so it took a little bit of a French accent. It's the imperfectness of that voice that makes it real.
So that's fascinating.
Yeah.
We should talk about benchmarks. For years, there's already been benchmark blindness, and many of the ASR companies embellish their benchmarks. They all use different internal tests, different public tests, and so on.
But now, in the regime of voice agents, it's even more difficult to test this.
Yeah.
So how are you guys doing that?
6. Voice Agents Need Better Benchmarks
Benchmarking the voice agent—I mean, benchmarking ASR and LLM—we have a lot of test standards. These are all more of a closed, particular module, and you evaluate them on that.
The voice is really difficult to evaluate because a voice agent has latency considerations. You have to reply fast enough. There's also the quality of the response. There's also a reasoning task: how difficult a task the agent can handle, and how accurately it can identify personal names and information like that. So it has too many axes.
I think the majority of the benchmarks that are publicly available out there don't really serve that purpose. Another problem is that a lot of these voice-agent benchmarks on the market are synthetic, to a degree, because voice is a lot more private. Releasing that voice data set for evaluation and then making it public is nontrivial as well.
Therefore, we have a lot of these benchmarks on the market, but they don't sufficiently represent what a good voice agent should look like. I think the industry as a whole is still figuring out how best to evaluate it. The reality is that voice agents are moving so much into real-world applications that sometimes it's also hard to say there's a really good benchmark that solves all the problems.
Our benchmark is built on the real data set we have from conversations with our consumers. We built it in a way that measures a lot of different important axes. We wanted to make sure the agent's response time is a factor in the benchmark, because a voice-agent benchmark without latency consideration isn't a real one. A user cannot wait on the phone for 60 seconds for the agent to come back with an answer.
That's a very important element. We evaluate understanding: would you be able to understand what the user says and then transcribe it accurately? ASR transcription is still one of them. Then there are the tool and traditional LM evaluations, instruction following, and the quality of the response. Are you using citations? Are you using tool calls and things like that? There are a lot of different evaluations we run.
Currently, our model is evaluated on that benchmark. It is the best model, of course, but it's an internal benchmark right now. Our plan is to release that benchmark in the next month or two to the market and to the research community, so they can also evaluate on it. We do believe that having a more realistic benchmark to evaluate on would be a good thing for the overall industry to continue moving forward.
One thing that I'm thinking of is that this might be an example of subjectivity. There are some objective things, right? Just quality and trust and tonality and stuff like that.
Yeah.
But it might be culturally specific and task specific. For example, in this particular application, someone in Newcastle might like it, while someone in France might not like it.
Right.
So what would you do there? Would you statistically aggregate over lots of different situations? How would you approach that?
Currently, the benchmark objectively looks at the quality of the response, more at a factual level. The actual end-user experience is an even harder problem to measure. For measuring that, we can argue that we need a benchmark data set to evaluate it.
But the reality is that if you're building an agent or a system toward that level, it's almost at production grade already. In production, you already have the end user to tell you. The way that we also measure this in real production is: does the user want to engage with the agent, and does the user actually finish the task in the end?
There are a lot of other production-related metrics you can measure there. But that is a lot harder to measure because it's not just about the model's performance. At that level, you have to couple the model with the harness together. So it's that entire agent evaluation, and that is an even harder task.
Can you talk me through the complexity of harness engineering? The reason this is interesting is that so many enterprises are adopting AI, and the main form of adaptation is harness engineering.
Yeah.
It's an incredibly good thing, and it's an incredibly bad thing, because everyone is building their own harnesses, adapting their skill surfaces and prompts. Essentially, everyone is reinventing the wheel and doing different things.
Right.
So there is an argument for having some kind of platform and standardized ways of doing this.
Yeah.
How do you deal with that problem?
7. Enterprises Are Building The Harness
This is what the market demands, right? I think enterprises acknowledge that artificial intelligence is going to be a very important technology for them to own, not just buy. They want to own it. They obviously want to own the entire brand and, ideally, the models as well, but they know they cannot own the models because they don't know how to tune them, and it costs a lot. They need to have a lot of GPUs, and they don't have the budget for it.
So they gave up on the models. They don't want to own the models anymore. They were like, "Hey, we can switch to different model APIs. That will be okay." But they still want to own the intelligence themselves. The only way for them to own it is to own the harnessing.
That's why I think a lot of enterprises are hoping to build their own harnessing, because everyone considers this such an important technology that they need to have in-house. Harnessing is the only thing they can currently bring in-house. So I think that's the market sentiment right now.
I do think not every harness should be owned by you yourself. It depends on how critical the task is. To be honest, I think voice agents—the voice-facing customer experience—is a harder problem than just general reasoning, because the adaptation itself is very important and very critical.
I don't think enterprises value that part as much. They value the reasoning part a lot more. What we like to tell our customers is, "You don't actually need to own the voice agent or the voice-agent harnessing itself. You can have this voice agent as the front door to your customers, and you can own your brand in the background."
Then you have the voice agent, which can respond very quickly, but it can delegate a task to your mastermind behind the scenes. I think this is probably what we're going to see in the next couple of years: enterprises will realize that they don't have to own everything, but they have to figure out where to draw the boundary—"This is how comfortable I feel giving things away."
Some people will probably always say, "I really need to own the harnessing because I want to do a lot of customization myself, and then I'm comfortable using the model." That's okay as well, but I don't think that will be the majority.
I love thinking of software engineering as a specification acquisition problem. Apparently, there's a whole theory around this, which is software engineering as theory building. Many folks at home will have experienced this: when you do vibe coding with agents, there's a kind of Rubicon moment where your software is sufficiently well specified—
Yeah.
—so that an autonomous agent, when it sees problems, can just fix things. It's almost like, if you can get it to the point where it's well enough specified, then it just converges from there. That's the magical moment.
Exactly. The beauty of working with these enterprise or B2B conversations is that there's kind of a good loop, right? It's actually relatively well specified because, in an enterprise conversation, the major thing is: does the customer's problem get solved in the end? Does the customer feel good in the end? Is it a good conversation? How fast can you solve the customer's problem?
The good thing is that, despite the fact that in human-to-human conversations there's no gold standard—a chitchat, if it's a good chitchat, would actually change the relationship between the two—in a B2B conversation, there's no good chitchat. It's just, "Did you solve my problem or not?"
So it's a much more well-defined problem. That's why I think we see a lot of success in applying coding agents on top of it, to be able to automatically optimize it.
Very cool. I guess one of the issues with agents—I mean, this has been a dream of mine for ages—is: what if we could almost build an artificial organization with some agents and some humans, and then solve the coordination problem between them?
Yeah.
So we have some agents that are specialized in front office and some that are specialized in back office and coding, and they're all just sending messages to each other, fixing things.
Yeah.
That's very exciting, but it's quite problematic as well because—
It is.
—there are all sorts of failure modes. How would you approach that?
It is, and I think enterprises don't feel comfortable about it yet. Despite the fact that we have a lot of vibe coders and a lot of AI-native companies rushing to build all these features, to be honest, a lot of these products are really, really cool. But I think enterprises are feeling very threatened by that.
Yesterday, we actually had one of our largest utility customers talking to my team. My team is really freaking out because they say, “In the future, all the PRs will get merged into the repository. We need to have a review, and the review cycle is 2 weeks, and then we have to make sure that everything that gets merged into production is auditable and reliable.”
In the past, it wasn't really the case because we were just extending the agent and making the agent really, really good for their use cases. But now they are applying that control on top. I think they have been holding off for a very long time, and now they feel it's time to do it because the agent is really good right now.
But just because enterprises want to have their visibility, their guardrails, and everything to be auditable, you kind of have to build around that as well. So a coding agent that directly spits out PRs and auto-merges would not fly with them.
That's why in our platform, we also build a lot of visualization components with the agents. It's like, “Hey, this is what the flow looks like. So if the agent enters here, this is what the process looks like.” There's a bunch of tools, and these tools are exactly what the agents can do. So anything outside of that tool set, the agent doesn't have the ability to do.
There is a citation module of the model, so every response comes with a citation and they can have an audit trail for that. I don't know how long we'll have to have this because we kind of live in a world where we have a lot of vibe coding, and coding agents can do things very fast, but we are still dragging these traditional SaaS enterprise no-code flows with us.
I think the good thing now is that these no-code flows are no longer the primary interface for people to build. They're just for you to audit the actual processes of the AI agent. I think we are kind of shifting towards that way, but that UI component and visualization is still very important for enterprises.
Adaptation is intelligence, so these models are being adaptive. But at the moment, we've been talking about skill-surface adaptation and code adaptation. The great thing about weight adaptation is that you can, in principle, share this model in the organization, because if you're adapting code and skill surface—
Yeah.
—and you've got agents everywhere, it's a little bit brittle in the sense that I can't really take my code here or my skill here and share it with an agent somewhere else in the organization.
Yeah.
So how do you think about this adaptation problem in general?
There's always this debate: where do you want it to change the behavior, and where should the customization go? To a certain degree, we agree with that principle because we do actually train our own models as well. We agree that there's just inherent behavior.
Of course, you can try to solve it from the agent-harness side, but even if you try to solve it there, you're probably going to put a lot of wrappers around it, and it's probably not going to be good enough for your use cases. If you have the ability to alter the weights, then you have a lot more control over what you are building.
So I think it's more of a philosophical question: to what degree do you think you should have control? Because there's definitely some danger in terms of going that deep. You say, “Yeah, you can control the weights,” but maybe you don't know how to do it the best way. You end up with a suboptimal sort of solution.
For experts like them, I'm pretty sure that they know what they are doing. But I would say that for a lot of enterprises, that's probably not their primary area of expertise. So I think that would require some help, even if they wanted to fine-tune the model.
Usually, what I would suggest to a lot of our customers is that they start with the harnessing, and if they really cannot get around that, then they can think about using different models or tuning their own models. But that should be a second resort. That should not be your number-one priority immediately.
8. Cognitive Debt Changes How We Work
Another thing I wanted to talk about is this notion of cognitive debt. I've been using OpenAI. You can now link your phone to it, and I can control all of my coding agents, and it's amazing.
I can go down the shops. I can take a walk along the road, and in some ways it's amazing; in other ways it's not. The reason it's not is that when I go back to my computer, I realize that sometimes the agent was doing the thinking when I thought the voice agent was doing the thinking. One of the underlying coding models should have been doing the thinking.
Sometimes it wasn't using my established agents; it was just creating new ones, and it was a little bit messy.
Yeah.
And cognitive debt is also about how it's really important sometimes just to see the visual information in front of you. So how do you address that problem?
It's a huge problem. I mean, it's a huge problem not just for product development and our customers; it's also a huge problem for us organizationally. All our teams are adopting Coco right now, and we have a lot of people integrating with Gmail, G Suite, and Slack channels.
People are now starting to send messages in Slack channels to each other. If people are not careful, they're sending a wall of text to the other person, and the other person will be like, “Oh, I have to spend so much time reading it.”
Over time, we see more people just say, “Okay, now have Coco read all the text and then summarize it back to me.” It's hard. That's the problem.
Humans don't have that much capacity. The bottleneck has already shifted from producing content to actually validating and auditing the content, and that has been the problem my engineers have been telling you about since the beginning of the year. We have so much code to review, and I don't trust agents to do code reviews yet because I don't feel comfortable with them just merging the code.
So we're definitely entering that stage, and we probably need a lot of education for people. Especially in the co-working scenario, you have to educate people: you have to help people summarize your point. And also, don't send people the AI slop that you haven't read yet. You have to read it yourself and then really understand and agree with what you send, not just throw that out to people. So I think it's a real problem.
I think in coding it's particularly acute because you're just dealing with something of very high complexity. There's lots of background knowledge. It feels like, in a customer service agent or something like that, both parties understand enough about what's being spoken about.
Yeah.
And is that part of it? Is it something like, if we both understand it, then it's fine because the voice agent is a little bit lower resolution, but it's pointing to things that we already know? But in the domain where we don't know something, then there's that cognitive debt problem.
Voice agents have less of the problem because, by nature, those voice conversations are a lot shorter. It doesn't actually generate a wall of text. Our average output token size in the production agent is probably about 50 tokens. It's not a lot, right? So consuming those is relatively simple.
It's the reasoning behind it. If you are prompting a larger agent to do the reasoning task, where it generates a lot of reasoning tokens, that's actually more of a problem. Usually, a voice agent by itself is not a problem.
Voice agents have a different problem: they don't have sufficient bandwidth in that communication channel, because voice requires very ping-pong-y, very short, that kind of conversation.
So it has a slightly different problem. I think the task you are doing is actually very interesting because I have been thinking about using a voice agent to do vibe coding as well. That would be really cool. I think that would probably require a little bit of harnessing into it to actually make sure that the voice agent always surfaces the necessary high-level bullet points, rather than just passing on a simple sentence per se. But I don't know whether it will be easy enough for you to really understand the full picture still, because that's a lot.
What I really love about voice agents, or just agents in general, is that I can do insane amounts of context switching.
Yeah.
I can be doing my email triage in one, video editing in another, and coding in another. It's really made me change my workflow. So now I have a breadth-first search workflow where I'm just popping off one, then popping off another one.
And that's good and bad because, in a sense, I'm not going really, really deep on things. I'm just doing many things at once. But the beauty of the agents is, as we were saying earlier, these domains have become well enough specified that the agent can have some autonomy, and I just sort of feed information to every single one. So it's kind of changing how I work. Have you noticed that?
I build a lot of these agents as well. Currently, my workflow every day is that I probably open only 2 or 3 pieces of software: Claw, Slack, and Chrome.
Sometimes I have to open Chrome so I can log into some software, so my agent can use the computer through it. It's not really me using it myself; it's the agent, so the agent can access the tool as well. So it's already changed my workflow completely.
Yeah. And do you think this is a mindset thing? Obviously, we're engineers, and what I did as quickly as possible was transform all the work I do into an agentic representation, basically. The million-dollar question is, we want to get to a point where voice agents are used for basically all work that can be done. How do you see that transition happening?
Voice is a very good input modality. Whether the voice agent itself is the best working intermediate layer, I'm not actually sure, to be honest. Let me give you an example. I've been using Whisper Flow myself. I'm pretty sure you use it as well. It's a very convenient ASR tool.
Oh, Whisper Flow.
Yeah.
Yeah, I love it.
It's amazing.
Yeah.
Right?
I'm in the top 1%.
Yeah, exactly. It's so convenient. You can actually just talk to the agent, and the agent, despite the fact that I'm not a native speaker, understands the errors in the transcription. The transcription is not always accurate, but those agents understand those errors still.
It's not like I send a voice memo to someone, and then they see that there's a lot of mistranscription and get confused. Agents don't have that problem. An agent will do the reasoning and say, “Oh, this word should mean that. This word should mean that. Okay, well, I'm going to do it that way.”
However, I am a user who uses it heavily. I use it even at work.
Yeah.
At work, you put your headphones on, and people are sitting next to you. There are already people taking calls at their desks. What's the problem with putting on headphones and talking to an agent? I got judged by so many people who said, “What are you doing? That's so weird.” I was like, “No, I'm just working.”
So I think voice, as I said, has this privacy layer where some people just don't feel natural doing that. Even me talking in the office, just talking to my agent, people find it weird. That's the same argument as the Meta glasses, right? You can wear them and actually talk to your agent. I got one as well. I started using it, and people were like, “What's going on?” I was like, “Hey, Meta, play the music. Take a picture.” They were just like...
People find it weird. So I think there's this user-adoption issue. I would say we're rare because we're really agent-built. We really like to work with that new modality. But I think a lot of people probably still aren't quite there yet.
We are engineers, and what I do is build tools to help tasks become agent-enabled. For example, you can't use an agent to edit something in DaVinci Resolve, or sometimes I make 3D animations in Blender. You can do that with an agent, but it's always going to go wrong.
So what I do is build a tool and a mental model—an abstract model of the domain—and then I build an agentic CLI. I put a skill prompt in there, and then my agent can do it. I'm always in this modality of tool-building to make an application amenable to agents.
It is.
And that feels like the gap at the moment because my mum wouldn't be able to do that. We want to get to the point where they can almost demonstrate. They can say, “This is something that I do,” put a demonstration in there, and describe abstractly what their thought process is. Then they can build an agent interface, and from there, the agent can do it, and you can correct it when it makes mistakes.
I have been building quite a lot of tools for my internal teams as well because people don't know how to do these things. I think one of the most useful tools I built for the team was an internal wiki to track all the product features, functionalities, sales, how we pitch to customers, and that kind of thing, and then continue to iterate on that.
I also built another agent that does online research to search for suppliers, vendors, competitors, and market trends, and then just builds that into the wiki. That's the context engineering I'm trying to do for them.
There are also a lot of tools you can build around it, like access to your email and your Slack channels. I'm going deep because Claude has a lot of these default connectors. I almost throw out all those connectors because, first of all, they're not fast enough. Second of all, they're not powerful enough. I want to go straight to the source, so now it's really powerful.
I think this just requires a lot of that kind of work. A lot of people just want to use it. They don't necessarily have the mindset of building something and using it themselves. They're kind of waiting for people to build it for them.
And what do you think the next decade of voice research is going to look like?
The voice agent needs to become increasingly more end-to-end. Number 1, turn-taking must be end-to-end and collapse into the entire pipeline, because that is the single most annoying part when it comes to engineering and the most difficult part to actually get right. This is already happening with a lot of these models.
In the voice modality, we'll probably have to start branching out. There should be consumer voice for consumer entertainment and day-to-day use cases—the Alexa and Google Assistant world. There are enterprise voice use cases, starting with the enterprise contact center, more customer support, and professional use cases.
It could also be meeting note-takers, and an agent to help you coordinate these meetings could be part of that scenario as well. Then it could also become the internal productivity layer of voice, between your coworkers. I don't know what that would look like yet, but I do think that the voice channels required for consumers and enterprises are very different.
One requires a lot more professionalism, regulation, and auditability. The other one is really about fun, because conversation is also supposed to be fun sometimes. So I think there should be these 2 different branches coming out.
Then we'll see more hardware devices that can carry these voice agents, because voice is a natural modality for IoT devices. We'll probably see a second wave of IoT coming along. How successful that will be, I don't know, but I'm really excited about that.
I actually bought an open-source robot recently, more like a voice-speaker robot. You can program a lot of it, but it's still very clunky. I think all these things will become a lot more prominent in the next 10 years.
I can imagine a future where the technology becomes so invisible that we can train the system. It’s almost like teaching a child. And then I suppose a really important thing, certainly in terms of organizations, is how we share those learnings.
Currently, it does actually feel like the technology is still very visible because there’s still a lot of learning curve. In our organization, people have been talking about AI slop because it’s very visible. More and more people are sending AI slop messages to each other, so it is still very visible.
I do actually agree that eventually it should just become invisible, in a way that it’s blended into your organization. I think that will probably still require a bit more iteration in terms of sharing the context, sharing who you are, and what the organization is about. Despite the fact that we’ve done a lot of internal knowledge architecting, there are still a lot of principles that we haven’t written into those MD files yet.
So I think there’s still a lot to do, and we’re probably running out of the low-hanging fruit now, because some of the things that are actually so personal to ourselves or to the organization are more spiritual. It’s about these kinds of constitutions or principles that you probably don’t print out everywhere.
Because the funny thing is that if you have the agent ingest these cultural values of your company right now, sometimes the agent just comes up with random ideas and overly focuses on one thing. But it’s not actually about that. Despite the fact that we write our cultural values down like this, there’s a reason why they came about, and those reasons a lot of the time aren’t injected into the context.
So agents oftentimes focus on the wrong things because, in their weights, somehow they think that out of these 6, this one is the most important one. But that’s actually not true for the organization. So we kind of have to share a lot more in order to get to that invisible state that you’re talking about.
The slop thing is very interesting, and my definition of slop is generation without competence. I’m using memory systems, adaptive specialization, skill adaptation, and so on. So it always has significantly more context. Increasingly, it knows what my preferences are. I’ve grounded it with the correct, technically referenceable materials, and so on.
So the fact of the matter is that there are 2 modes of AI. There’s AI that has sufficient context and knows what to do, and that AI does not produce slop.
Being slop, unfortunately, is probably a little bit more subjective, in a way. I engineer and context-engineer my agent so that the agent doesn’t feel like slop to me, but it may feel like slop to the other person who is reading my message because, in their context, the things I care about probably don’t really matter all that much to them. Or maybe they’re not aware that this is important yet.
So when they read it, they think it’s slop because they don’t actually see the context behind the reasoning. That is a bit of a subjective problem. So I think we have this internal debate or argument: should I care about whether I’m generating slop and then just sharing it with my colleagues?
The fundamental debate is whether a human being should be reading those things anyway, or whether a human being should just ask an agent to read it for them and then summarize it in a way that’s better suited for their focus. I almost landed on the idea that I should probably build a skill for everyone’s agent, based on their personal context, to summarize whatever message people send to them, so they can extract the useful information for themselves.
In that case, you want to encourage the output side to be as rich as possible. You send a wall of text, and that’s okay, because on the other side there’s an agent summarizing it for that particular individual. But I kind of feel like this is running in circles. It’s like we humans now need an agent in between to help us cherry-pick the useful information for us.
John, it’s been amazing having you on the show. Thank you so much for joining us.
Thank you. Yeah.