[BidClub_]
No Priors · · 30 分钟

AI Agent 如何重塑客户支持:与 Decagon 的 Jesse Zhang 对谈

Jesse ZhangElad Gil

播客
TL;DR
  • 客户支持正成为 AI Agent 近期的“黄金用例”,因为部署可以从小规模启动,同时节省成本和服务质量仍能直接量化。 Decagon 同时追踪自动化处理的对话占比、客户满意度、NPS 和准确率,实际上为每位客户提供跨语言、全天候在线的“口袋私人礼宾”。

  • Bilt Rewards 提供了 Decagon 最清晰的经营杠杆证明:大约1个月内,随着 AI 接管大量自动化工作,Bilt 停止扩充客服团队。 近1年后,Bilt 已重组客服职能,并量化出约65名客服的人员成本节省;客户则表示,其客服“完全不像我们以前用过的任何 AI 或聊天机器人系统”。

  • Decagon 声称的差异化来自通用模型之上的软件,而不是对 GPT-4o、GPT-4 或 Claude Sonnet 的独家访问。 其编排层围绕客户特定的业务逻辑评估并组合模型;产品层则展示数据、决策步骤、知识缺口和对话类别。Zhang 的框架是:“大部分 alpha,或者说你真正构建的大部分特殊能力,都在模型之上。”

  • 对于客服 Agent,指令遵循比 o1 和 Sonnet 时代被反复强调的编程与数学推理能力提升更重要。 客服 Agent 必须“严格按 SOP 执行”,尤其是在受监管行业;更强的推理能力有帮助,但对 Decagon 而言,可靠地遵循工作流才是更关键的模型发展方向。

  • 语音扩大了可覆盖的工作流,但也带来持续存在的延迟与计算权衡。 端到端语音模型响应很快,但生产环境中的通话可能需要检索数据并多次调用模型;语音转文字、文本处理、再转语音会增加延迟,却能完成这些工作。Decagon 的务实过渡方案包括使用“好的,请稍等一下,我正在查询您的数据”这样的对话缓冲。

  • 组织的终局并不只是减少人手,而是让更多人“监督和编辑 Agent”。 Zhang 预计,管理者将监控、纠正并硬编码无限扩展系统的行为;屏幕上下文和计算机使用能力则可能让 Agent 真正操作产品,而不只是回答问题。不过在他看来,Anthropic 展示的 computer use“还没达到生产就绪”。

  • 相较于行业演示所暗示的前景,Zhang 对大多数近期 Agent 类别“更悲观”。 成功的市场既需要在接近完美之前逐步上线,也需要容易量化的 ROI:安全 Agent 面临非确定性问题,因为任何微小异常都可能必须被捕捉;text-to-SQL 则往往仍是需要人工监控的 copilot,经济价值难以定价。更好的模型未来可能解锁这些类别。

摘要 · 为研究而整理的核心内容

1. 客户支持为 AI Agent 提供了异常务实的切入点

  • 将自己的第一家公司卖给 Niantic 后,Zhang 与 Ashwin 探索 AI Agent 时没有过度设计这套命题:他们此前最大的经验是,创业者“不能把事情想得太复杂”。与客户交流后,他们将客户服务确定为“黄金用例”。

  • Decagon 成立于2023年8月,目前已为 Rippling、Notion、Duolingo、Eventbrite、Vanta、Substack 和 Bilt Rewards 等公司的客服运营提供服务。反复使用的评估框架很直接:AI 能完成多少工作,客户满意度提升了多少,以及——尤其是在受监管行业——准确率有多高。

2. Bilt 将指数级增长的客服需求转化为可见的人员杠杆

  • Elad Gil 引用了 Klarna 的案例。Klarna 是一家欧洲先享后付服务商,数据来自其 CEO 在 X 上发布的帖子:4周内处理230万次聊天,客户满意度与人工相当,相比人工重复咨询减少25%,解决问题平均用时2分钟,而人工客服需要11分钟。帖子还称,Klarna 已覆盖23个市场、35种语言,并提供7×24小时服务。Gil 表示,他认为 Klarna 已将700名全职客服转去从事其他工作。

  • 在 Bilt,Zhang 的因果链条很简单:咨询量大致随用户数线性增长,因此用户规模“基本呈指数增长”就会带来相应的客服压力。最初的诉求非常直接:“天啊,这么大的量把我们压垮了。AI 能不能帮忙?”

  • 大约1个月内,随着 Decagon 自动化处理更多工作,Bilt 停止扩充客服团队。近1年后,Bilt 已重组客服职能,并量化出约65名客服的人员成本节省——这是很容易表达的 ROI,同时还带来了更快的服务速度和社交媒体上的正面评价。

3. 可防御的产品建立在共享基础模型之上

  • Zhang 将 Decagon 视为一家软件公司,因为所有人都能访问相同的 GPT-4o、GPT-4 和 Claude Sonnet 模型。差异化来自“编排层,或者说围绕模型的软件”,而不是底层模型本身。

  • 编排意味着评估每项任务最适合由哪个模型处理,将这些模型组合起来,再围绕客户的业务逻辑塑造最终系统。周边软件让决策过程可检查:用户可以看到系统参考了哪些数据、采取了哪些步骤、生成了什么答案,并提供反馈。

  • 在大规模运行时,LLM 可以读取100万段人工团队无法逐一审阅的对话,发现知识缺口、划分需求类别,并汇报整体运行情况。买方可以用人工表现和业务指标对质量进行基准测试,再让 Agent 先处理1%的流量,逐步扩大规模——也可以同时测试另一种方案。Zhang 认为,可观测性、可解释性和控制能力是 Decagon 取得效果的关键。

4. 指令遵循与语音延迟成为下一阶段约束

  • 近期围绕 o1 和 Sonnet 的进展,重点集中在定量推理、编程和数学能力,但 Zhang 表示,这些并不是 Decagon 面临的最大瓶颈。客户支持依赖指令遵循:给定一套 SOP 或工作流,Agent 能否“严格按要求执行”?

  • 语音只是通过另一条渠道处理同一个客户互动问题。随着 ElevenLabs、OpenAI 和 Cartesia 提升语音的真实感与响应速度,Decagon 的客户正在测试语音 Agent;Zhang 的更大范围职责覆盖聊天、电子邮件、SMS 和电话,因为“渠道本身并不重要”,底层请求才重要。

  • 延迟仍是核心问题。端到端语音响应很快,但生产环境中的客服可能需要检索数据并多次调用模型;先转录、再用文本完成计算、最后重新生成语音则更慢。团队可以流式输出,或使用“好的,请稍等一下,我正在查询您的数据”这样的对话缓冲。Decagon 会围绕每个客户的优先级做出这些权衡。

5. 人的参与仍是经营杠杆的重要来源

  • Zhang 预计 Agent 数量会出现“合理的爆发式增长”,尽管部分用例需要更长时间才能落地。因此,客服工作会转向由人来监督和编辑 Agent,管理者负责监控行为、纠正偏差,并保留可见性和控制权。

  • 监督 AI 与管理员工不同:Agent 可以无限扩展,部分行为还可以直接硬编码。Zhang 认为,实时纠正的交互界面仍有很大改进空间,例如可以直接告诉 Agent:“你刚才这件事做错了。下次请这样做。”

  • 他的创始人网络也提供了另一种人的优势。过去5、6年里,参加数学和编程竞赛的同行从学术界或量化交易转向创业,随后在非正式层面交叉投资,并共享招聘、销售和薪酬数据。Zhang 认为,竞赛经历是“相当不错的信号”,但 Decagon 的招聘流程总体并无不同,也不要求候选人具备这类背景。

6. 渐进式上线与清晰 ROI,区分出可持续市场与演示效果

  • Zhang 表示,按照当前模型的发展阶段,绝大多数用例都不会实现真正的商业化应用。Decagon 刚开始时,他并不知道未来12或24个月内是否会出现真实用例;事后看来,他认为成功的用例需要渐进式上线,并且 ROI 可量化。

  • 安全领域体现了可靠性问题:AI 看起来很适合处理包含海量日志的环境,但实际工作可能要求捕捉每一个微小异常。由于生成式模型天然具有非确定性,Zhang 预计,企业采用 Agent 化安全系统的速度会“非常、非常、非常慢”。

  • text-to-SQL 体现了经济性问题。买方可能喜欢生成结果,但仍需要有人监控和编辑,最终 Agent 变成了 copilot;由于大多数团队并没有那么多数据科学家,替代价值难以基准衡量,也很难据此证明大额合同的合理性。

  • 近期真正胜出的模式,将编程 Agent 式的任务拆分与客服式的效果衡量结合起来:在边界明确的工作上部署,在达到完美之前先交付价值,并量化回报。Zhang 仍对许多当前类别持悲观看法,但也保留了这一判断:“随着模型改进,它们将解锁大量新的用例。”

Elad Gil

Hello, and welcome to No Priors. Today, I'm talking to Jesse Zhang, co-founder of Decagon. Decagon is an early-stage company building enterprise-grade generative AI for customer support. Founded in August of twenty twenty-three, their platform is already being used by large enterprises and fast-growing startups like Rippling, Notion, Duolingo, ClassPass, Eventbrite, Vanta, and more.

Jesse, welcome to No Priors.

Jesse Zhang

Of course. Thanks for having me, Eli.

Elad Gil

Absolutely. Maybe we can start a little bit with your background and what Decagon does. You're a serial founder. You started another company before this, Niantic Bot. Now you and Ashwin have started Decagon, and you've been working on it for a while and have seen some really interesting adoption from companies like Rippling, Notion, Eventbrite, Vanta, Substack, and many others, right? So you've really started to carve out a real space for the company. Can you tell us a little bit more about what Decagon does, how it works, and what the focus of the company is?

Jesse Zhang

Of course. So, quick background on me. I grew up in Boulder and did a lot of math contests and stuff like that growing up. I studied computer science at Harvard and, as you mentioned, started a company right out of school. That company was eventually bought by Niantic, and then I left to start this company.

Ashwin and I met through mutual friends. We officially met at this VC offsite. When we got together, we were like, “Okay, the biggest learning from my first company is that you can't really overthink things too much.” We started by just being interested in AI agents. It's very exciting technology, arguably the coolest thing from this generation.

1. Decagon Finds Its Use Case

We talked to a bunch of customers, like the ones you listed. I think over the years we've gotten a lot better at figuring out how to talk to folks and what questions to ask. Through that process, we arrived at the current use case, which is maybe what we think is the golden use case for these AI agents: customer interactions and customer service.

The use case is very tailor-made for what LLMs are good at, so we started building from there. We still weren't thinking too much about vision or anything yet. It was just, “All right, we had a lot of customers in front of us. How can we make it so that they're happy and they really like what we're building?” That led to where we're at now.

I would say that right now, as a company, Decagon ships these AI agents for folks to use on the customer service and customer experience side. The thing that's made us special so far is that we have a huge focus on transparency, I guess. When people use us, especially these larger companies, it's very important for them that the AI agent is not a black box. They want to feel like, even though LLMs are cool and there's a lot of things you can do with them, they can see how decisions are being made, what data is being used, how answers are generated, and that they can give feedback if they want to.

Currently, we're in production with a bunch of these large companies that have large support teams. Pretty much any company that has a sizable support operation is a good fit for us.

Elad Gil

That makes a lot of sense. It's interesting because I feel like one of the things that's been really striking over the last year in the AI world is that the CEO of Klarna posted on X about the impact AI has had on its customer support team. Klarna is a buy-now, pay-later service out of Europe.

His tweet basically said that in the first 4 weeks, they handled 2.3 million customer service chats. Customer satisfaction was on par with humans. There was a 25% reduction in repeat inquiries relative to people. It resolved customer errands or issues in 2 minutes versus 11 minutes for a human agent. They were instantly live 24/7 in 23 markets and 35 languages because AI supports so many things.

It had a huge impact on that company, and I think they shifted 700 full-time agents to do other work, right? In terms of the impact on Klarna itself as an organization, what sort of impact have you been seeing with your customers as they adopt this technology? How do you think through the lens of what you're really bringing to these customers and the satisfaction that their own end users have?

2. The Customer Support Payoff

Jesse Zhang

It's an interesting way to think about it. All these people are shipping this use case, right? There's a lot of evangelists out there, which is nice. The Klarna article is awesome; it gave a lot of tailwinds to the industry.

One interesting thing we've seen is that the benefits people get are all roughly in the same vein, but different people prioritize different things. For our customers, it's always the same: What fraction of total work—in this case, conversations—can the AI agent do? How much work is this saving us? And how much happier are our customers? What's the customer satisfaction score or NPS score?

Those two are often the leading metrics by far. As I said before, different people may value each one slightly differently. Then there are other things, like making sure there's accuracy. If we're in a regulated industry, this has to be very accurate for us.

Those are where the benefits lie. We're saving a bunch of money, as well as time and resources, but on the other side, we're making customers happier. That can lead to higher retention, more conversions, and a lot more upside.

It's like you're giving every customer a personal concierge in their pocket that they can chat with at any time, in any language, 24/7. That can be pretty transformational for a lot of businesses.

Elad Gil

Is there any example customer that you can talk about as a case study in terms of the impact this has had, how it's lifted their metrics, and the success they've seen using Decagon?

Jesse Zhang

Of course. We just did a big case study with a company called Bilt Rewards, which is a great use case for us. They have a very large user base that's growing very quickly. You use it to either earn points or make payments. A lot of my friends use the product.

As a result, when you have a large customer base, people have questions and things they need help with. The number of support inquiries basically grows linearly with the number of users. Because they're growing so fast—basically exponentially—that means the number of support queries is also growing exponentially.

When they first started using us, that was the main goal: “Holy crap, we're getting overwhelmed by all this volume. Can AI help here?” Within a month of starting to use us, they were able to stop scaling their team, and the AI would take over a lot of the automation, which just makes everything very smooth.

Now we're almost a year in. They've been able to really restructure their customer support team. We published a case study on this where they were able to quantify the savings. So far, it's around 65 agents' worth of headcount saved—a very tangible difference.

For us, it's also great because we're able to provide them with that value. It's a very easy ROI, but the customer experience is also a lot snappier. They get a lot of social media posts like, “Holy crap, I just tried Bilt Rewards support, and it doesn't feel like any sort of AI or chatbot system we've ever used before.” That makes us happy.

Elad Gil

Could you tell me a little bit more about what you've built from a technology and infrastructure perspective? I guess there's the core models that anybody can access, right? The GPT-4o models of the world or GPT-4, the Claude Sonnet models, et cetera. Then there's all the stuff you've built on top of them to actually make this work well for your specific use case and for customer support agents. Could you tell us a bit more about what you all have had to build over time?

3. The Agent Software Stack

Jesse Zhang

Of course. Like you said, everyone has access to the same models. We see ourselves very much as a software company. We're obviously doing a lot of work around AI and using the AI models a lot. But I would argue that most applications nowadays are real software companies, and AI models are tools that everyone can use.

Most of the alpha, or most of the special stuff that you build, is on top of models. It's either the orchestration layer or the software around it. For us, there's been a big focus on both.

The orchestration layer is how you can use all these different models together. You probably have evals set up that measure how good each model is at certain things. You put them together, and the whole goal is to mold them around the business logic of the customer. That’s part 1.

The other thing you build is just very classic software. You have this AI agent there. It’s all the things I was saying before: transparency is a big piece. You really don’t want this to feel like a black box that’s just there answering questions.

How can you build all the tooling to see what data the agent is using and what steps it’s taking? Can I analyze all these conversations that are coming in? If you have 1 million conversations, no one’s reading all those. How can you make it so that the AI, the LLM, can read every single conversation, tell you how things are going, find gaps in its knowledge, and give you a breakdown of the big categories you should care about? There’s been a trend here that’s been interesting.

That’s all the software around it that we’re building, and that’s typically how it’s structured. The orchestration layer, I think, is going to be different for every agent. Our agent versus a coding agent—the orchestration is going to look pretty different. But at the end of the day, you’re just building a structure on top of the LLMs.

Elad Gil

Yeah, it seems like we’re very early in the days of true agentic stuff. That includes the ability to sequence chains of events that include certain forms of reasoning. Obviously, there are things like o1 and other things that have been coming out to start to address this, but we seem quite early in the scaling curves.

What do you think are the main pieces of technology that are missing to really take your vision to the next level in terms of how these agentic systems should work?

Jesse Zhang

Yeah. One thing we were talking about the other day is that there are actually different types of intelligence with the AI models. A lot of the recent developments with o1 or Sonnet and things like that have been around, I guess, quantitative reasoning intelligence. They’ve gotten better at coding, and they’ve gotten better at math.

For us, those things help, but they’re actually not the biggest difference-makers. In our use case, the type of intelligence that matters the most, we would probably describe as instruction-following. You just have a bunch of instructions, and can you follow them to a T? I’m sure there are other types as well, but we’re excited to see developments in the other areas, too.

Everyone’s saying, “Oh, there’s a plateau happening with the core models and the intelligence.” I think when most people say intelligence like that, they’re probably talking about reasoning capabilities. For us and the agentic flows that we use, instruction-following is a huge piece. Just think about a customer service SOP or a workflow or something like that: you have to be very accurate about it. I know there’s research going on about this in the major labs, and I think that’s one thing we’re looking forward to next year.

Elad Gil

One other area that seems like it really touches on customer success, customer support, and user experience is voice-based support. One of the things that’s a little under-discussed in the AI world is that we keep talking about large language models and understanding text. Obviously, that stuff is crucial to everything else, but I feel like we almost under-discuss text-to-speech engines and the ability to understand spoken words and then respond with audio.

There are companies like Cartesia, ElevenLabs, OpenAI, Google, and others that are starting to provide some of these services and APIs. How much of an impact does that have on what you’re doing? Is that a separate type of product? How do you think about the voice component of these things?

4. Voice Agents Enter Production

Jesse Zhang

Great question. A huge impact. We have customers now trying our voice agents. If you just think about our space, the overall problem is the same: you have a bunch of customers, and they have questions or issues or things they need to talk about. The channel really doesn’t matter to them. Some people prefer voice, some people prefer chat, some people prefer email, and some people prefer SMS or something like that. Our job is to handle all of those.

Obviously, you start with text because that’s the easiest one. It’s easy to evaluate for the customer as well. I think just now you’re getting to the point where you have big companies that are very interested in voice. They’ve seen the results of a text-based agent, and they’re saying, “Okay, well, yeah, we should be able to generate voices and do the same thing for phone calls.”

None of this would be possible without the models that you just listed and those companies. ElevenLabs, OpenAI, and Cartesia are doing some cool stuff. I think there have also been huge strides this year with those models around how realistic the voices sound.

Latency matters a lot in our use case, because if you’re making a phone call, you expect things to feel very snappy. It’s a big topic for us, and as these companies get better, we’re working with them pretty closely right now on how you can actually build these things well at scale. As they get better, that’s also going to be huge for us to keep delivering these voice agents.

Elad Gil

Makes sense. My sense is that one of the issues is latency. It takes enough time to take an audio stream where somebody’s talking, translate that into text, feed that into a language model, and then output it as voice again that it feels like there are a lot of pauses where people have to wait. There are different things that people have been trying to do in the background, like streaming potential solutions back out and being able to shorten that latency timeline.

Do you feel latency is still an issue, or is it solved by integrating voice directly into the models in a deeper way for some of these services? When do you think latency becomes a solved problem for these sorts of application areas?

Jesse Zhang

Latency is a big deal here, of course, with voice models. Nowadays, you have the voice-to-voice models that we’re playing around with. OpenAI is doing a lot of work here. There are obviously a lot of trade-offs. Voice-to-voice latency is great.

Sometimes, though, with these production use cases, you do need the extra computation cycles to fetch data, do multiple model calls, or handle other reasons that you can’t do voice-to-voice. That’s one option that you would consider.

The other one is the one you described, where you’re transcribing—doing speech-to-text—and then doing all the computation within text and generating the voice at the end. That always causes a little extra latency, of course. As you mentioned, a lot of folks have figured out fairly clever ways to get around that. You can start generating stuff first.

In our use case, you can always do something like, “Hey, give me a second. I’m looking up your data.” These are all things we’re playing around with. For each customer that we work with, there are different trade-offs, so we’re really trying to base what we build on the things we’re hearing from them and the priorities that they have.

Elad Gil

One thing that I think is interesting is the number of companies in the AI world today that have been founded by people with Math Olympiad, IOI, or other sorts of backgrounds. I think you were involved with Math Olympiad stuff in high school. I think Decagon has actually hosted some Math Olympiad events for the team, which isn’t your typical happy hour.

There are other teams and companies. Before that, there was Ramp and things like that, but I think the Brain Trust team, the Pika team, and Cognition, which launched Evan, all have that common thread. Where do you think that comes from? Why do you think this community is now so active in AI?

5. The Math Olympiad Network

Jesse Zhang

That’s a good question. We’re actually all around the same age as well, so we’ve known each other since middle school and high school. It’s a great community. For us, we have a lot of people on the team with math-contest and coding-contest backgrounds.

I think it’s more that this community was always there. Math contests have been around for a while, and a lot of very smart kids go through them. It’s also a great way for folks to get to know each other, get connected, and build friendships.

I think the main thing is that now, in the last few years—maybe the last 5 or 6 years—startups have become a lot more mainstream. A lot of folks in this demographic have gravitated toward startups, as opposed to traditionally going into academia or quant trading and things like that.

So there’s just a big influx of these super-smart, super-talented people coming into the startup world. Because there’s this community aspect, folks can see what other people are doing, what sort of works, and the types of companies that people are building. That doesn’t say they’re all the same, but I think a lot of folks with these backgrounds are now working on startups, and that’s why there’s a lot of progress in the companies that folks have been building.

Elad Gil

Mm-hmm. Are there ways that you all have been supporting each other through the startup journey? I feel like every generation has a clique of people who built some of the more interesting companies, and they all interact. They provide advice, maybe they angel invest in each other. There’s a thriving community, and every 5 to 7 years, it shifts who it is.

I feel like the IOI, Math Olympiad, or coding competition communities are very engaged right now. Is there any formal version of that, or are you all just informally helping each other?

Jesse Zhang

Yeah. I angel invest in a lot of the companies you just listed, and a lot of their founders are angel investors in our company. It’s very informal, obviously. It’s just casual friends helping each other.

I think the main thing is that, with company building, there’s a lot of surface area. As you know, there’s hiring people, doing sales, building the thing, structuring compensation—I don’t know, there are infinite things. Having the other data points is obviously super helpful.

I hang out with them quite often and play games and card games. There’s a Chinese version of Bridge that I play with a lot of these folks. It’s fun; we just hang out. Everyone’s in roughly the same stage of life, so, like you said, there’s definitely a lot of camaraderie and help that goes around.

Elad Gil

Has coming from this background, from the Math Olympiad community, impacted at all how you think about hiring or your hiring practices at Decagon?

Jesse Zhang

A little. If someone else has the same background and has gone through the same contests or programs, obviously that’s a pretty good signal, since I have a good idea of what those people have done.

My co-founder, Ashwin, also has a similar background. He didn’t grow up in the U.S., but in India he did a lot of these contests as well. I think there’s some correlation with people who, as kids, did a lot of this stuff. Now we’re all adults, and there’s some sort of signal there when you’re talking about hiring.

For the most part, there are so many talented people here, whether or not you did math contests. In San Francisco, at Decagon, and at other companies, I think our hiring process has been more or less the same. It is a nice trigger for events, I guess. When you host these events, people come out, and you can get a nice community of folks who are interested in the same things.

We’re probably going to be hosting more. Not all of them are going to be contest-based, obviously—puzzles and things like that, where you get a lot of fun engineers and people bringing their friends. That’s pretty important to us.

Elad Gil

Mm-hmm. For AI writ large, what are you most excited about in the coming years? If you were to extrapolate out 12 to 24 months, what are you anticipating most keenly, or what are you waiting for?

6. Agents Reshape Human Work

Jesse Zhang

Obviously, the models getting better is awesome. The models getting better across different modalities is also awesome. We talked about voice. There are also other modalities that are tangentially interesting to us.

A lot of our customers have software products, so it would be awesome if, when you’re asking questions to AI agents, the agents had context of your entire screen and all the interactions you’ve done. You can even go a step further and have them actually help you navigate things. There’s just so much you can do there when you talk about other modalities or more advanced model capabilities.

We’ve seen the computer-use demo from Anthropic. Probably, in my opinion, it’s not production-ready yet, but as that gets better, there are a lot of cool things you can do there. On the model side, that’s one thing we’re excited about.

On the non-core-model side, one thesis we have is that, as the years go by, AI agents are going to become undeniable. There’s going to be a reasonable explosion of them, and they’ll be used for a bunch of different use cases. I think some use cases will take longer than others, but the value that they’re providing is pretty undeniable.

There are definitely going to be a lot of AI agents out in the world—in our use case, customer service, and in other use cases. One thesis we have is that the nature of the work of human agents and people like us is also going to change pretty drastically. One of the things that’s going to change is that there are going to be a lot more people supervising and editing agents.

That’s something we think about. We’re excited for a lot of the innovation there because, right now, a big part of what we care about is letting the human agents at our customers and their leadership teams go in and make changes, monitor the agents, and have a lot of visibility and control.

What does that look like? If you compare it to monitoring humans, you can give them feedback in real time. You can be like, “Oh, no, don’t do this. You did this thing wrong. Please do this next time.” When you’re doing that with an AI agent, there are a lot of different possibilities because they have some things that are different from humans. They’re infinitely scalable, and you can really hard-code things sometimes.

That’s the other area, probably going into next year, that we’re looking forward to.

Elad Gil

Mm-hmm. That’s really cool. Do you view that as a main area of differentiation for you relative to some of the other folks in the market providing customer-success and support tooling?

Jesse Zhang

Yeah, right now that’s probably the biggest thing.

The interesting thing about our space—and I think this will probably be true for a lot of AI agent spaces—is that the results are very quantifiable. You’re basically taking the agent and benchmarking it against how good a human would be, how much money it’s saving you, and how much better the quality of the customer experience is.

Because of that, when people evaluate us in our space, it’s a pretty quantitative evaluation. They’re like, “Okay, cool. This kind of works. Let me just put you out into production for 1% of the volume and build up from there,” and maybe do that with another option.

A lot of the old-school companies, like Salesforce, are going to be in this very exciting space for them, too, so they’re going to have alternatives. Then you just benchmark everyone: How good are the stats? How good are the metrics? How good of a job is everyone doing?

So far, we’ve been performing very well. The main reason for that is this transparency piece: giving people observability, explainability, and control over the AI. There’s still a long way to go in that field. There’s still so much more you could do, and that’s been our specialty so far.

Elad Gil

That’s great. I’ve had some conversations with your customers over time, as people have been trying some of these agents. They’ve called me to ask questions about different companies in the space and everything else. The 3 things they tend to point out are that you all ship really fast, you’re very responsive as a team and company, and, most importantly, the product tends to outperform. I think that’s really been great to watch over time.

How do you think about the areas where AI agents are going to be successful versus not successful in the short run?

7. AI Agent Adoption Stays Uneven

Jesse Zhang

One of the things that we’ve been thinking through—and this was pretty big for us when we were first starting out—is that there’s going to be a huge variance between the different types of AI agents, how successful they’ll be, and how quickly they’ll roll out.

When we were first starting the company, we were pretty open to what to build, and we knew that AI agents were exciting. At that point, we didn’t even know whether there would be any real use cases that would emerge in the next 12 or 24 months, but we were exploring.

I think our view is that, for the vast majority of use cases right now, there still isn’t going to be real commercial adoption with the current state of the models, because of a bunch of things.

One big thing is that, in a lot of spaces, there’s really no structure to incrementally build up. It has to be good—almost perfect—off the bat. If you think about a space like security, where you have all these SIMs out there, it makes sense. There are tons of logs, and that’s perfect for AI models, but the goal of that job is to catch any small thing that happens.

Because the models are inherently nondeterministic, it’s very hard for buyers to really trust a generative AI solution there, especially an agentic solution. I think adoption there is going to be really, really, really slow—a lot slower than people think—even though people have cool demos and see things seem to work. Getting real enterprise adoption is going to be very slow. That’s one interesting thing we’ve been thinking about.

The other side of that is that there are also a lot of spaces where, on the surface, it seems like AI agents will be perfect. But the follow-up is that it’s actually not that easy to quantify the ROI that’s happening. One example I would give is that there were a lot of text-to-SQL companies and things like that where you could see it working.

But basically, immediately, everyone’s reaction is, “Oh, this is cool, but we’re still going to have to have someone monitoring it and editing it,” and so it becomes kind of a copilot. Okay, cool. So then how do we measure how much we should pay for one of these agents? It’s very difficult because most teams don’t have that many data scientists anyway.

If you’re claiming that you have an AI agent data scientist, it’s like, okay, let’s benchmark you against a real one. You’re probably not going to be able to replace a real one. I think that’s the sort of thing where it’s very hard to quantify the ROI. You’re saving some people time, but because of that, if I’m a large company, it’s hard for me to justify, “Okay, I’m going to give you a large contract for this AI agent data scientist.”

Those are the things that we were thinking through. We weren’t thinking through them in the moment—we were obviously just asking customers what their willingness to invest in certain things was. But in hindsight, looking back on the last year, I think that’s been a big thing that’s been true: the use cases that emerge have to have two qualities.

It has to be something that can be rolled out slowly and doesn’t have to be perfect off the bat, but is already providing value. I think coding agents are a good example of this. You can just section off some tasks for them, and they’ll do them.

The other piece is the ROI. You have to really be able to easily quantify the ROI. In our case, luckily, you have these support agent teams and people who track metrics very closely. That’s something we’ve been thinking about.

I think the takeaway from that is that we’re probably more bearish on a lot of these AI agent use cases in the near term. But as these models get better, they’re going to unlock a lot of new use cases.

Elad Gil

Super interesting. Jesse, thank you so much for joining us today.

Jesse Zhang

Thanks, Elad. Thanks for hosting. It’s great seeing you.