No Priors 第97期|与 Decagon CEO 兼联合创始人 Jesse Zhang 对谈
客户支持是 Decagon 的 AI 智能体“黄金用例”,因为自动化程度和客户结果都能直接衡量。 Jesse Zhang 表示,买方会追踪智能体处理的对话占比、CSAT 或 NPS,以及——尤其在受监管行业——准确率。最终,企业可能获得一个支持任何语言、全天候在线的“私人礼宾”,在节省人力的同时提升留存和转化。
Bilt Rewards 采用 Decagon 后约1个月内就停止扩张支持团队,此后已记录节省约65名员工的编制。 随着用户基数快速扩大,其支持量也持续增长;自动化让 Bilt 得以重构运营,同时加快响应速度。Zhang 称,这个案例的 ROI “非常容易看清”。
真正的差异化在基础模型之上,而不在于独家获取基础模型。 Decagon 通过评测驱动的编排组合多个模型,围绕每个客户的业务逻辑进行塑形,并展示回答背后的数据、步骤和知识缺口。“你真的不希望它给人的感觉像一个黑箱。”
对客服智能体而言,指令遵循比模型讨论中占主导的编程和数学能力提升更重要。 Zhang 欢迎 o1 和 Sonnet 带来的进步,但表示决定性能力在于,模型能否拿到一份客服 SOP 或工作流并“严格照做”。即使核心模型继续进化,应用层仍有大量工作要完成。
近期的智能体赢家必须支持渐进式部署,并在达到完美之前就能带来容易量化的回报。 客户支持和编程满足这一标准;安全领域可能不满足,因为漏掉一个细微事件就可能不可接受,而 text-to-SQL 往往仍是需要人工监督的副驾驶,定价能力也不清晰。因此,Zhang 表示自己“对很多 AI 智能体近期用例更加悲观”。
语音、屏幕上下文和智能体监督,是 Decagon 未来的关键方向。 语音到语音模型可以降低延迟,但生产工作流仍可能需要检索数据并进行多次调用;Zhang 认为 computer use 可能还没准备好投入生产。更长期看,他预计会有更多人监督和编辑可无限扩展的智能体,使可观测性和控制成为产品重点。
1. 客户需求将客服支持选为 Decagon 的智能体切入口
Decagon 成立于2023年8月。Zhang 的第一家公司被 Niantic 收购;他和联合创始人 Ashwin 创办 Decagon 时,最大的心得是创始人“不能把事情想得太复杂”。他们最初对智能体有广泛兴趣,随后与客户交流,最终由这些对话筛选出客户服务这一最强的初始用例。
这一匹配建立在一个事实上:该用例非常适合发挥 LLM 的长处。如今,Decagon 面向拥有大规模支持业务的公司,客户覆盖高速增长的初创企业和大型企业。
透明度成为最核心的要求。客户需要查看哪些数据生成了答案、智能体采取了哪些步骤,以及自己能否提供反馈:“对他们来说,AI 智能体不是黑箱非常重要。”
2. 客服自动化带来异常清晰的经济回报
Elad Gil 提到的基准案例是 Klarna:4周内处理230万次聊天,满意度与人工持平,重复咨询减少25%,解决问题平均用时2分钟,而人工需要11分钟。AI 还让 Klarna 能够在23个市场、35种语言中提供全天候服务,同时将700名全职员工转去从事其他工作。
Zhang 的评估表有两个核心指标:智能体完成全部对话的比例,以及客户是否通过 CSAT 或 NPS 变得更加满意。受监管行业的买方还会增加准确率约束,但更广泛的收益包括降低成本、加快服务、提升留存和增加转化。
Bilt Rewards 提供了 Decagon 最清晰的案例。由于咨询量随着用户基数快速增长,Bilt 起初担心支持团队会不堪重负;但在约1个月内,它就停止扩张支持团队。近1年后,公开案例显示其节省的编制约为65名员工,同时客户体验也变得更加“利落”。
3. 应用层差异化来自编排、评测和运营控制
GPT-4o、GPT-4 和 Claude Sonnet 等模型人人都能使用,因此 Zhang 将 Decagon 描述为一家把模型当作工具的软件公司。他认为,真正的特殊之处在于编排以及模型周边的软件,而不是模型本身的获取权限。
编排层可以通过 evals 衡量模型在特定任务上的表现,将多个模型组合起来,并围绕每个客户的业务逻辑塑造最终系统。这种编排会因品类而异:客服智能体和编程智能体需要不同的结构,即使底层使用的是同一批模型。
周边软件还必须能够在规模化运营中理解业务。如果客户有100万次对话,不可能由人全部读完;系统应该识别主要类别、趋势、缺失知识和其他缺口,同时解释单项决策背后的数据和步骤。
买方可以先让智能体处理1%的业务量,将结果与人工表现及其他方案进行比较,再逐步扩大部署。Zhang 表示,Decagon 在这些基准测试中的优势来自“可观测性、可解释性、控制力”,但也承认距离成熟“还有很长的路要走”。
4. 指令遵循与延迟定义技术前沿
Zhang 将定量推理与指令遵循区分开来。近期模型在编程和数学方面进步显著,但客服更依赖忠实执行详细 SOP——“严格照做”。相比抽象推理能力,他更希望顶尖实验室推进的是这种能力。
语音是同一个底层客户问题的另一个渠道,与聊天、邮件和 SMS 并列。Decagon 起步时选择文本,因为客户更容易评估;但见过文本智能体效果的客户,如今已经开始测试由 ElevenLabs、OpenAI 和 Cartesia 等公司构建的语音智能体。
架构上的权衡是延迟与计算量。语音到语音模型响应迅速,但生产场景可能需要检索数据并进行多次模型调用;语音转文字、文本处理和重新生成语音都会增加延迟。一种实用的过渡方式是用对话填充时间——“稍等一下,我正在查你的数据”——同时让后台继续工作。
除了语音,Zhang 还希望智能体能够利用用户完整的屏幕内容和交互历史,然后直接帮助用户操作软件。Anthropic 的 computer-use 演示展示了这一方向,但他判断目前可能还没准备好投入生产。
5. 可行的智能体需要渐进式上线和可量化 ROI
Zhang 回顾时采用两个标准:智能体必须在接近完美之前就能创造价值,买方还必须能够量化这种价值。他认为,在当前模型条件下,“现在绝大多数用例”仍未达到商业化准备状态。
安全领域体现了完美主义问题。尽管大规模日志流看起来非常适合 AI,但这项工作可能要求捕捉每一个细微事件;模型的非确定性让企业买方不愿信任智能体方案。该领域的采用可能会“非常、非常慢”——远慢于那些令人印象深刻的演示所暗示的速度。
text-to-SQL 体现了衡量价值的问题。买方可能喜欢生成结果,却仍要求有人监督和编辑,最终系统变成副驾驶。如果所谓的 AI 智能体数据科学家大概率无法替代真正的数据科学家——而企业本来就只雇佣很少的数据科学家——供应商就很难证明一份大额合同的合理性。
编程是相反的正面案例:团队可以划出部分任务,让智能体尝试完成,并在不交出全部控制权的情况下获得有用结果。更好的模型将解锁更多品类,但 Zhang 对近期前景的结论仍然刻意保持谨慎。
6. 智能体监督成为一种新型工作
随着智能体不断增加,Zhang 预计人们会把更多时间用于监督和编辑智能体。不同于人类员工,智能体“可以无限扩展”,而且部分行为可以硬编码,从而为实时反馈、监控和控制创造新的可能。
Decagon 自身的产品方向也顺应这一变化:人工客服和管理团队都应该能够检查表现、介入处理并作出修改。Zhang 将这层控制能力视为公司当前的差异化所在,也是下一步创新的方向。
Zhang 还将数学和编程竞赛社区描述为一个非正式的创始人网络。成员之间会相互进行天使投资、交换运营建议并一起社交,包括通过中文版本的桥牌进行社交;这一背景是有用的招聘信号,但他强调,Decagon 的招聘流程总体上没有区别,人才来源也远不止竞赛参与者。
Hello and welcome to No Priors. Today I'm talking to Jesse Zhang, co-founder of Decagon. Decagon is an early-stage company building enterprise-grade generative AI for customer support, founded in August 2023. Their platform is already being used by large enterprises and fast-growing startups like Rippling, Notion, Duolingo, ClassPass, Eventbrite, Vanta, and more.
Jesse, welcome to No Priors.
Of course. Thanks for having me.
Absolutely. Maybe we can start with your background and what Decagon does. You're a serial founder—you started another company before this, and that company was eventually acquired by Niantic. Now you and Ashwin have started Decagon, and you've been working on it for a while. You've seen some really interesting adoption from companies like Rippling, Notion, Eventbrite, Vanta, Substack, and many others. You've really started to carve out a real space for the company. Could you tell us more about what Decagon does, how it works, and what the focus of the company is?
Quick background on me: I grew up in Boulder and did a lot of math-contest stuff growing up. I studied computer science at Harvard and, as you mentioned, started a company right out of school. That company was eventually bought by Niantic, and then I left to start this company.
Ashwin and I met through mutual friends. We officially met at a VC off-site, and when we got together, we thought, “The biggest learning from our first company is that you can't really overthink things too much.” We started by being interested in AI agents. It's very exciting technology—arguably the coolest thing from this generation—and we talked to a bunch of customers, like the ones you listed.
I think over the years we've gotten a lot better at figuring out how to talk to folks, what questions they ask, and what matters to them. Through that process, we arrived at our current use case, which is what we think may be the golden use case for these AI agents: customer interactions and customer service. The use case is tailor-made for what LLMs are good at.
We started building from there, and we still weren't thinking too much about the vision or anything yet. It was just, “We have a lot of customers in front of us. How can we make them happy, and how can we make sure they really like what we're building?” That led to where we're at now.
Right now, as a company, Decagon ships these AI agents for people to use on the customer service and customer experience side. The thing that's made us special so far is our huge focus on transparency, I guess. When people use us, especially larger companies, it's very important to them that the AI agent isn't a black box.
Even though LLMs are cool and there are a lot of things you can do with them, they want to see how decisions are being made, what data is being used, how the agent comes up with answers, and whether they can provide feedback. We're currently in production with a bunch of large companies that have large support teams. Pretty much any company with a sizable support operation is a good fit for us.
Yeah, it's interesting because I feel like one of the things that's been really striking over the last year in the AI world is that the CEO of Klarna posted on X about the impact AI had on its customer support team. Klarna is a buy-now-pay-later service out of Europe.
The post basically said that, in the first 4 weeks, they handled 2.3 million customer service chats. Customer satisfaction was on par with humans, there was a 25% reduction in repeat inquiries relative to people, and it resolved customer queries or issues in 2 minutes versus 11 minutes for a human agent. They were instantly live 24/7 in 23 markets and 35 languages because AI supports so many languages.
It had a huge impact on that company, and I think they shifted 700 full-time agents to do other work. In terms of the impact of Klarna's AI support on the organization, what sort of impact have you been seeing with your customers as they adopt this technology? How do you think through the lens of what you're bringing to these customers and the satisfaction their own end users have?
It's an interesting way to think about it. Everyone is shipping this use case, and there are a lot of evangelists out there, which is nice. The Klarna article was awesome and provided a lot of tailwinds for the industry.
One interesting thing we've seen is that the benefits people get are roughly in the same vein, but different people prioritize different things. At this point, it's not really much of a hot take to say that, in a couple of years, these agents are going to be super pervasive. People can use them for all these customer interactions, and they're going to be everywhere.
To your point, what is the benefit? For our customers, it's always the same: What fraction of the total work—in this case, conversations—can the AI agent do? How much work is this saving us? And, second, how much happier are customers? What's the customer satisfaction score or NPS score?
Those 2 are often the leaders by far. As I said before, different people may value each one slightly differently. Then there are other things, like accuracy. If we're in a regulated industry, this has to be very accurate for us.
Those are where the benefits lie. We're saving a bunch of money, time, and resources, but on the other side, we're making customers happier. That can lead to higher retention, more conversions, and a lot more upside. You're giving every customer a personal concierge, essentially, in their pocket that they can chat with at any time, in any language, 24/7. That can be pretty transformational for a lot of businesses.
Is there an example customer you can talk about as a case study, in terms of the impact this has had, how it's lifted their metrics, and the success they've seen using Decagon?
We just did a big case study with a company called Bilt Rewards. It's a great use case for us. They have a very large user base that's growing very quickly, and people are using the product to either earn points or make payments. A lot of my friends use the product.
As a result of having a large customer base, people have questions and things they need help with. The number of support inquiries grows linearly with the number of users. Because they're growing so quickly—basically exponentially—the number of support queries is also growing exponentially.
When they first started using us, the main goal was, “Holy crap, we're getting overwhelmed by all this volume. Can AI help here?” Within basically a month of starting to use us, they were able to stop scaling their team. The AI would take over a lot of the automation, which just makes everything very smooth.
Now we're almost a year in, and they've been able to really restructure their customer support team. We published a case study on this where they quantified the savings. So far, it's around 65 agents in saved headcount—a very tangible difference.
For us, it's also great because we're able to provide them that value. It's a very easy ROI. The customer experience is also a lot snappier, and they get a lot of social media posts saying, “Holy crap, I just tried the Bilt Rewards support system, and it doesn't feel like any AI or chatbot system we've ever used before.”
Could you tell me a little bit more about what you've built from a technology and infrastructure perspective? I guess there are the core models that anybody can access—the GPT-4o and GPT-4 models, Claude Sonnet, and so on—and then there's all the stuff you've built on top of them to make this work well for your specific use case and for customer support agents. Could you tell us a bit more about what you've had to build over time?
Like you said, everyone has access to the same models. We see ourselves very much as a software company. We're obviously doing a lot of work around AI and using AI models a lot, but I would argue that most applications nowadays are real software companies, and AI models are tools that everyone can use.
Most of the alpha, or most of the special sauce that you build, is on top of the models. It's either the orchestration layer or the software around it. For us, there's been a big focus on both.
The orchestration layer is how you can use all these different models together. You probably have evals set up that measure how good each model is at certain things. You put them together, and the whole goal is to mold them around the business logic of the customer.
The other thing you build is just classic software. You have this AI agent, and transparency is a big piece. You really don't want this to feel like a black box that's just there answering questions. How can you build all the tooling to see what data the agent is using and what steps it's taking? Can I analyze all these conversations that are coming in?
If you have 1 million conversations, no one is reading all of them. How can you make it so that the AI, the LLM, can read every single conversation, tell you how things are going, find gaps in the knowledge, and give you a breakdown of the big categories you should care about? Maybe there's been a trend here. That's all the software around it that we're building.
Typically, that's how it's structured. The orchestration layer is going to be different for every agent. Our agent versus a coding agent will have very different orchestration. At the end of the day, you're building a structure on top of the LLMs.
It seems like we're very early in the days of true agentic systems, including the ability to sequence chains of events that include certain forms of reasoning. Obviously, there are things like o1 and other models that have been coming out to try to address this, but we see them quite early in their scaling curves.
What do you think are the main pieces of technology that are missing to really take you, or your vision, to the next level in terms of how these agentic systems should work?
One thing we were talking about the other day is that there are actually different types of intelligence in AI models. A lot of the recent developments with o1, Sonnet, and things like that have been around quantitative-reasoning intelligence. They've gotten better at coding and math.
For us, those things help, but they're not the biggest difference-maker. In our use case, the type of intelligence that matters most is what we would probably describe as instruction following. You have a bunch of instructions, and you need to follow them to a T.
I'm sure there are other types as well, but we're excited to see developments in those other areas, too. People are saying that there's a plateau happening with the core models and intelligence. When most people say “intelligence” in that context, they're probably talking about reasoning capabilities.
For us, and for the agentic flows that we use, instruction following is a huge piece. Think about a customer service SOP, playbook, or workflow. You just have to be very accurate about it. I know there's research going on about this in the major labs, and that's one thing we're looking forward to next year.
One other area that seems to touch on customer success, customer support, and user experience is voice-based support. I think one thing that's a little under-discussed in the AI world is that we keep talking about large language models and understanding text, even though that stuff is crucial to everything else. But we almost under-discuss text-to-speech engines, the ability to understand spoken words, and then respond with audio.
There are companies like Cartesia, ElevenLabs, OpenAI, and Google that are starting to provide some of these services and APIs. How much of an impact does that have on what you're doing? Is that a separate type of product, or how do you think about the voice component of these systems?
Great question. It has a huge impact. We have customers now trying our voice agents. If you think about our space, the overall problem is the same: You have a bunch of customers with questions or issues that you need to talk about, and the channel really doesn't matter to them. Some people prefer voice, some prefer chat, some prefer email, and some prefer SMS. Our job is to handle all of those.
Obviously, you start with text because it's easier to evaluate for the customer. I think we're just now getting to the point where large companies are very interested in voice. They've seen the results of a text-based agent and are saying, “Well, you should be able to generate voices and do the same thing for phone calls.”
None of this would be possible without the models you just listed and the companies building them: ElevenLabs, OpenAI, Cartesia, and others. There have also been huge strides this year in how realistic the voices sound. Latency matters a lot in our use case because, if you're making a phone call, you expect things to feel very snappy.
It's a big topic for us. As these companies get better, it's going to be huge for us to keep delivering these voice agents. We're working with them pretty closely right now on how to build these things well at scale.
My sense is that one of the issues is latency. It takes enough time to take an audio stream of somebody talking, translate that into text, feed it into a language model, and then output it as voice again that it can feel like there are a lot of pauses. People have to wait.
There are different things people have been trying to do in the background, like streaming potential solutions back out and shortening that latency timeline. Do you feel latency is still an issue, or is it solved by integrating voice directly into the models in a deeper way for some of these services? When do you think latency becomes a solved problem for these types of applications?
Latency is a big deal here, of course, with voice models. Nowadays, you have voice-to-voice models that we're playing around with. OpenAI is doing a lot of work here. Voice-to-voice latency is great.
Sometimes, though, with production use cases, you need the extra computation cycles. You need to fetch data, make multiple model calls, or do other things that mean you can't use voice-to-voice. That's one option you have to consider.
The other option is the one you described, where you're transcribing—or doing speech-to-text—and then doing all the computation in text and generating the voice at the end. That always causes a little extra latency. As you mentioned, a lot of people have figured out fairly clever ways to get around that. You can start generating things first.
In our use case, you can always do something like, “Give me a second. I'm looking up your data.” These are all things we're playing around with. For each customer we work with, there are different trade-offs, so we're trying to base what we build on what we're hearing from them and the priorities they have.
One thing I think is interesting is the number of companies in the AI world today founded by people with math Olympiad, IOI, or other similar backgrounds. You were involved with math Olympiad stuff in high school, and I think Decagon has hosted some math Olympiad events for the team, which isn't a typical happy hour.
There are other teams and companies. Before that, there was Ramp and companies like that, but I think the Braintrust team, the Pika team, and Cognition, which makes Devin, all have that common thread. Where do you think that comes from? Why do you think this community is now so active in AI?
That's a good question. We're actually all around the same age as well, so we've known each other since middle school or high school. It's a great community. We have a lot of people on the team with math-contest and coding-contest backgrounds.
I think it's more that this community was always there. Math contests have been around for a while, and a lot of super-smart kids go through them. It's also a great way for people to get to know each other, get connected, and build friendships.
The main thing is that, in the last 5 or 6 years, startups have become a lot more mainstream. A lot of people in this demographic have gravitated toward startups, whereas traditionally they would have gone into academia, quantitative trading, or things like that.
There's been a big influx of super-smart, super-talented people into the startup world. Because there's this community aspect, people can see what others are doing, what works, and the types of companies people are building. That doesn't mean the companies are all the same, but a lot of people with these backgrounds are now working on startups, which is why there's been a lot of progress in the companies people have been building.
Are there ways you've all been supporting each other through the startup journey? I feel like, in every generation, there's a clique of people who build some of the more interesting companies and all interact with one another. They provide advice, maybe angel-invest in each other, and create a thriving community.
Every 5 to 7 years, it shifts who those people are. I feel like the math Olympiad and coding-competition communities are very engaged right now. Is there any formal version of that, or are you all just informally helping each other?
I've angel-invested in a lot of the companies you just listed, and a lot of their founders are angel investors in our company. It's very informal—just casual friends helping each other.
The main thing is that company building has a lot of surface area. As you know, there are questions like how to hire people, how to do sales, how to build the product, and how to structure compensation. There are infinite things. Having other data points is obviously super helpful.
We hang out quite often, play games, and play card games. It's a Chinese version of bridge that I play with a lot of these people. It's fun to hang out when everyone's at a relatively similar stage of life. Like you said, there is definitely a lot of camaraderie and help that goes around.
Has coming from this background, from the math Olympiad community, affected how you think about hiring or your hiring practices at Decagon?
A little. If someone else has the same background and has gone through the same contests or programs, that's obviously a pretty good signal, since I have a good idea of what those people have done. My co-founder, Ashwin, has a similar background. He didn't grow up in the United States, but in India, and he did a lot of these contests as well.
There's some correlation with people who, as kids, did a lot of this stuff. Now we're all adults, and there's some sort of signal there when you're talking about hiring. But for the most part, there are so many talented people in San Francisco, whether or not they did math contests, at Decagon and at other companies, that our hiring process has been more or less the same.
It is a nice trigger for events, I guess. When you host these events, people come out, and you can build a nice community of folks who are interested in the same things. We're probably going to host more. Not all of them will be contest-based—some will involve puzzles and things like that—but you get a lot of fun engineers and people bringing their friends. That's pretty important to us.
For AI at large, what are you most excited about in the coming years? If you were to extrapolate out 12 to 24 months, what are you anticipating most keenly, or what are you waiting for?
Obviously, the models getting better is awesome. The models getting better across different modalities is also awesome. We talked about voice, but there are other modalities that are tangentially interesting to us.
A lot of our customers have software products, so it would be awesome if you could ask questions to AI agents and they had context from your entire screen, all the interactions you've done, and things like that. You could go a step further and have the agent actually help you navigate things. There's so much you can do with other modalities and more advanced model capabilities.
We've seen the computer-use demo from Anthropic. In my opinion, it's probably not production-ready yet, but as that gets better, there are a lot of cool things you can do there. On the model side, that's one thing we're excited about.
On the non-core-model side, one thesis we have is that, as the years go by, AI agents are going to become pervasive. At this point, it's undeniable that there's going to be a reasonable explosion of them across a bunch of different use cases. Some use cases will take longer than others, but the value they're providing is pretty undeniable.
There are definitely going to be a lot of AI agents out in the world—in our use case, customer service, and in other use cases. One thesis we have is that the nature of the work done by human agents and people like us is also going to change pretty drastically. One of the things that's going to change is that there will be a lot more people supervising and editing agents.
That's something we think about, and we're excited about a lot of the innovation there. A big part of what we care about is letting the human agents at our customers and their leadership teams go in, make changes, monitor the agents, and have a lot of visibility and control.
What does that look like? If you compare it to a human, when you're monitoring a human, you can give feedback in real time. You can say, “No, don't do this. You did this thing wrong. Please do this next time.” With AI, there are a lot of different possibilities because agents have properties that are different from humans. They're infinitely scalable, and you can sometimes hard-code things. That's another area we're looking forward to next year.
Do you view that as a main area of differentiation for you relative to some of the other companies on the market providing customer success and support?
Right now, that's probably the biggest thing. The interesting thing about our space—and I think this will probably be true for a lot of AI agent spaces—is that the results are very quantifiable. You take the agent and benchmark it against how good a human would be, how much money it's saving, and how much better the customer experience is.
Because of that, when people evaluate us in our space, it's a pretty quantitative evaluation. They say, “Okay, this kind of works. Let me put you into production for 1% of the volume and build up from there.” Maybe they do that with another option as well.
A lot of the old-school companies, like Salesforce, are going to see this as a very exciting space, too, so they're going to have alternatives. Then you benchmark everyone: How good are the stats? How good are the metrics? How good a job is everyone doing?
So far, we've been performing very well. The main reason is this transparency piece—giving people observability, explainability, and control over the agent. There's still a long way to go in that field. There's so much more you could do, and that's been our specialty so far.
I've had conversations with your customers over time. People have been trying some of these agents and have called me to ask questions about different companies in the space. The 3 things I tend to point out are that you ship really fast, you're very responsive as a team and company, and, most importantly, the product tends to outperform. I think that's really been great to watch over time.
How do you think about the areas where AI agents are going to be successful versus unsuccessful in the short run?
One thing we've been thinking through—and this was pretty important to us when we were first starting out—is that there's going to be a huge variance between the different types of AI agents, how successful they'll be, and how quickly they'll roll out.
When we were first starting the company, we were pretty open to what to build. We knew AI agents were exciting, but at that point, we didn't even know whether there would be any real use cases that emerged in the next 12 or 24 months. We were exploring.
Our view is that, for the vast majority of use cases right now, there still isn't going to be real commercial adoption with the current state of the models, for a few reasons. One big thing is that, in a lot of spaces, there's really no structure for incrementally building up. It has to be good—almost perfect—right off the bat.
Think about a space like security, where you have all these SIEMs and tons of logs. That seems perfect for AI models. But the goal of that job is to catch every small thing that happens. Because the models are inherently nondeterministic, it's very hard for buyers to trust a generative solution there, especially an agentic solution.
I think enterprise adoption in that area is going to be very, very slow—much slower than people think. People have cool demos, and things seem to work, but getting real enterprise adoption is going to be very slow.
The other side of that is that there are a lot of spaces where, on the surface, it seems like AI would be perfect, but the follow-up is that it's actually not easy to quantify the ROI. One example would be text-to-SQL companies. You can see it working, but almost immediately everyone's reaction is, “This is cool, but we're still going to have to have someone monitoring it and editing it.”
It becomes a kind of copilot. Then the question is, how do we measure how much we should pay for one of these agents? It's very difficult because most teams don't have that many data scientists anyway.
If you're claiming that you have an AI agent data scientist, it's like, “Okay, let's benchmark you against a real one.” You're probably not going to be able to replace a real one. That's the sort of thing where it's really hard to quantify the ROI. You're saving some people's time, but if I'm a large company, it's hard for me to justify giving you a large contract for an AI agent data scientist.
Those are the things we were thinking through. In the moment, we were asking customers what their willingness to invest in certain things was, but in hindsight, looking back on the last year, that's been a big pattern.
The use cases that emerge have to have 2 qualities. First, they have to be something that can be rolled out slowly and doesn't have to be perfect right away, while still providing value. Coding agents are a good example: You can section off some tasks for them, and they'll do them.
The other piece is ROI. You have to be able to quantify it easily. In our case, luckily, you have these support agent teams, and people track metrics very closely. The takeaway is that I'm probably more bearish on a lot of these AI agent use cases in the near term. But as the models get better, they'll unlock a lot of new use cases.
Super interesting. Jesse, thank you so much for joining us today.
Thanks, Elad. Thanks for hosting. It's great seeing you.