Kavak 围绕 AI 重建公司的打法
- Kavak 已完成自我重构:每天实例化100至200,000个智能体——每位客户对应一个,每个智能体都拥有独立虚拟机、数年的交互记忆,并以最大化客户终身价值为长期目标。 AI 负责人 Alejandro Maza 表示,目前96%的互动和95%的交易已完全由智能体处理;其设计问题是:“如果到了2035年,拥有 GPT-10-level intelligence,我们会如何打造 Kavak?”从交易导向转向关系导向具有直接商业意义:数据库中已有10M名客户,“仅激活这批客户中的1%”就可能创造数亿美元价值。
- 这些智能体不是客服机器人,而是销售员;其转化率如今达到 Kavak 人工团队的2.1倍,NPS和客户满意度评分均升至3倍。 Maza 给出的机制是:一个“超级专家”取代分布在融资、保险和置换环节的15名人工专家,它“有无限耐心”;某个智能体一旦出错,其他智能体到次日就能吸取教训。在信贷业务中,Kavak 通常不到3分钟即可批出车贷,而墨西哥和部分新兴市场需要2个月甚至更久;利率、风险和贷款额度均可个性化设定。
- 制约 Kavak 迭代速度的不是谨慎,而是评测:投入评测的工程师时间、token 和资金,与投入智能体本身的大致相当。 “能跑多快?取决于评测质量”——这就是汽车刹车的类比;首轮检查直接看业务结果,包括转化、客户获得的价值、满意度,以及再次互动意愿。
- Opus 4.5 发布后,Maza 毁掉了已运行2年的多智能体架构,因为“这已经不是正确的范式”。 12月时,已有数千乃至可能数万个智能体在运营业务;他判断图结构和多智能体网格可能束缚更新一代智能,因此改为围绕虚拟机执行框架重建,加入记忆、评测、CLI和长期目标。他的建议很明确:“不要构建智能体工作流”(don’t build agentic workflows)。
- 一个 AI CEO 已接管 Cuernavaca 业务6周——虽未实现利润翻倍目标,却做到了1.5倍,即利润增加50%——它逐项盯住每个数字,并向线下作业员工发送每日计划。 Maza 表示,仍需大量人力的工作主要集中在物理世界:Kavak 在墨西哥约800名机械师都配有 El Mike,这个《Ratatouille》式助手可以协助验车;质保索赔下降约26%。
- 面向企业买家,Maza 提出一套供投资者衡量 AI 支出的 token 质量框架:3级 token 投向可量化单 token ROI 的智能体,2级 token 的回报可间接衡量,1级 token 则是直接使用 Claude Code、ChatGPT、Cowork 或类似工具——“这些会带来什么? 我不知道。”转型必须自上而下:“如果每个人都对战略各有主意,军队根本无法运转”;依靠黑客松和自下而上推动用例“行不通”。
- 宏观逻辑遵循 Schumpeter:浅层采用 AI 的既有企业可获得6–10%的提升,想追求10倍改善则必须深度重构组织。 他的电力类比是:Ford 生产线所需技术在1879年和1881年已经出现,但仅将燃煤发动机换成电动机,只带来约6%的改善;围绕电力重建工厂,则让生产率达到原来的约3倍。“这是工业规模上的创新者窘境”;他给创始人的建议是,“只需勾勒出 AI 能力的一条线性趋势”(just map a trend that’s linear),再围绕它构建。
1. 每位客户一个智能体,每天最多实例化200,000个——定义这家公司的核心押注
- Maza 的经历解释了这份笃定:2013年,他在 transformer 出现之前创办 Opi Analytics——“我们领先时代10年”——为 Fortune 500 企业提供风险算法、物流、预测和营销服务,之后加入纵向一体化的二手车交易平台 Kavak。拉美当时不存在服务客户所需的基础设施,因此 Kavak 还自建了金融科技业务、物流体系和 Carfax。
- 架构如下:客户一到,“就会为这位客户专门生成一个智能体,并配备独立虚拟机”;它记得数年交互——网页访问、2年前的一通电话——并设定最大化终身价值的长期目标。每天有100至200,000个智能体实例化,工作“有时3分钟,有时8小时,有时3天”,随后给自己设定闹钟,再次休眠。
- 3项决策推动了这场转型:其一,不是给员工配上 ChatGPT,而是重写多数 API,以智能体为中心重塑公司;其二,押注超人类智能体,让其直接处理难题,并由此生成数据与反馈闭环;其三,从交易指标——买入多少辆车、卖出多少辆车、订购多少刹车片——转向关系指标。Kavak 数据库中有10M名客户,其中大多数已分配智能体;公司估算,仅激活其中1%,价值就可能达到数亿美元。
2. 评测是刹车,智能体是销售员——以及受监管的信贷业务
- Maza 的运营规则是:投入评测和智能体的工程师时间、token 与资金大致相等。“我喜欢极快推进,但要快,就必须有刹车。”首轮检查直接看业务结果:客户是否转化、是否获得价值、是否满意,以及之后是否愿意再次互动。
- Kavak “从未构建客服或客户服务智能体,我们构建的是销售智能体。”在拉美卖1辆车,意味着要从20,000个 SKU 中选车,安排融资、保险和保障方案,并给置换车辆报价;过去这需要15个团队的15名专家协作。如今一个智能体就是横跨所有环节的超级专家:NPS和客户满意度升至3倍,转化率最初比人工团队高50%,如今达到其2.1倍;原因是智能体“永远不会累”,而且一个智能体犯的错会在次日传导给其他智能体。
- 主持人追问,受监管的金融服务能否由 AI 端到端完成。Maza 的回答是,体验“好太多了”:Kavak 通常不到3分钟即可批出车贷,而在墨西哥和部分新兴市场,这一过程需要2个月甚至更久。Kavak 会个性化设定利率、风险等级和最高贷款额度,同时从整体上优化贷款组合。若借款人无力继续偿付,纵向一体化使 Kavak 可以收回车辆,再提供一辆更便宜的车,从而降低月供。
3. AI CEO 实验与物理世界的边界
- 最难的能力问题——AI 能否胜任 CEO——直接进入实测。6周前,Kavak 将 Cuernavaca 业务单独划出,让一个智能体担任 CEO,首月目标是利润翻倍。Maza 表示,它未达标,但实现了1.5倍,即利润增长50%;方法是逐一审视每个数字和每位客户、做预测,并对执行进行微观管理。它向线下作业员工发送每日计划,并要求用语音消息汇报进度;库存、融资渗透率和客户满意度均有提升。
- Maza 认为,仍需大量人力的工作主要集中在物理世界,需要灵活操作与感官能力。Kavak 在墨西哥约800名机械师都配有 El Mike;Maza 将其比作《Ratatouille》里的老鼠,它会讲解如何验车、提供提示,并向机械师演示操作。验车和维修变得更快、更便宜,交付车辆质量提高,质保索赔下降约26%。
4. Jedi Academy、人类为智能体工作,以及给既有企业的建议
- Kavak 所有人都通过内部 Jedi Academy 重新受训。该项目由 Maza 设计并持续升级,因为“你不可能把这些人送去 Stanford 学这些东西。”从 CEO、AI 工程师到财务人员和机械师,所有人都要参加;6周后,他们便把最先进的 AI 智能体上线生产。Maza 的话很直接:要么为新现实接受训练,否则“如果这不适合你,也许就离开 Kavak”。他说,该项目强化了企业文化,也让员工掌握了与智能体协作的实用技能。
- 如今,组织主要由扁平、资深且获充分授权的团队组成,只做3类事:“构建智能体、为智能体工作,或身处物理世界直面客户。”有时智能体是人的上司,有时人负责设计智能体。生产环境中,卡住的智能体会调用“我需要帮助”API,由人类响应,而不是把案例丢给 Tier 2 后不了了之。这个机制闭合了反馈回路,并产生改进智能体所需的数据。
- Maza 给管理层的建议有2点:第一,转型必须自上而下,并清晰描绘公司3年或5年后应有的样子——“如果每个人都对战略各有主意,军队根本无法运转”;第二,token 质量比采用率更重要。3级 token 投向可量化单 token ROI 的智能体;2级 token 的价值可间接衡量,例如开发者的工作是否进入代码库并上线生产;1级 token 则用于 Claude Code、ChatGPT、Cowork 或类似工具,对其价值,“我不知道”。
5. 毁掉可运行系统、自我改进组织与 Schumpeter 的警告
- 本期最冒险的决策是:12月,已有数千——甚至可能数万个——多智能体系统在规模化运行并运营业务。Opus 4.5 随后发布,Maza 判断,图结构、多智能体网格和执行框架会限制更新一代智能。于是他毁掉了让 Kavak 实现盈利的2年成果,转而围绕虚拟机重建:智能体、记忆、评测和 CLI 构成核心,目标是利用递归式自我改进以及更新、更智能的模型。他的建议毫不含糊:“不要构建智能体工作流。”
- RSI 的重新框定是:“过去4,000年,人类社会的经济价值由组织而非个人创造。”企业应追求的闭环,是打造一个自我改进的组织,利用更好的模型与智能,让自身智能和创造的经济价值同时复利增长。
- Maza 认为,这轮价值将由创业公司而非既有企业攫取;其逻辑是 Schumpeter 的创造性破坏,电力史提供了例证。Ford 生产线相关技术在1879年和1881年已经出现,但在那座靠传动轴和皮带运转的4层工厂里,仅将燃煤发动机换成电动机,只带来约6%的效率提升;围绕小型发电机在平面上重建工厂,则让生产率达到原来的约3倍。今天,浅层采用者可能获得6–10%的提升;追求10倍改善则需要深度重构。“这是工业规模上的创新者窘境。”
- 给创始人的结语:这是“人类历史上最激动人心的时代”,几乎不花钱、或每月约20美元,就能使用强大工具和智能。无需做指数级外推:“只需勾勒出 AI 能力的一条线性趋势”,再围绕它构建。
I’m investing more today in tokens than in knowledge workers. We could build superhuman agents. This means that, by every dimension that matters, our agents would outperform the best human we had ever hired.
The most ambitious companies listening to this will decide to follow suit. You decided to build an agent per customer.
Yes. Every day, between 100 and 200,000 agents get instantiated specifically for this customer, each with its own virtual machine.
There are a lot of people worried about what the organizations of the future are going to look like and the role that humans are going to play.
If you haven’t faced fear before, if you haven’t felt it, then you haven’t tried AI. We launched a program inside Kavak called the Jedi Academy. From the CEO to AI engineers to mechanics, we train everyone, and after 6 weeks, they launch state-of-the-art agents to production.
Today, we have Alejandro Maza Ayala, the head of AI at Kavak. We’re going to discuss the transformation that Alejandro led within Kavak to turn it into an AI-native company. Thank you, Alejandro, for being with us today.
Thanks for having me.
Before starting at Kavak, you were running a company called Opi Analytics.
That’s right.
And you were very much into AI before ChatGPT. Do you want to tell us a little bit about that journey?
Yes, of course. We called it machine learning back then. It was a different family of algorithms, and we founded a company with this very ambitious vision that new machine learning models would be so powerful that they could solve any complex problem. This was pre-transformers, right? This was 2013.
So, we started building the company that way, and I think we were 10 years ahead of our time. But we built a great company. We served Fortune 500 companies around risk algorithms, logistics, forecasting, and marketing.
1. What Kavak Does & the Agent-Per-Customer Architecture
The power of transformers, and then the ChatGPT moment when it arrived, made it very clear that we could now build a whole new company and a whole new way of building companies. We joined Kavak and Carlos to build that.
Amazing. Let’s spend the bulk of this podcast talking about exactly how you transformed Kavak. But maybe just to start, what does Kavak do, and what is your role there?
Kavak started out as a used-car marketplace. We buy cars, refurbish them, and then sell and finance them. But to do that, we also had to build a fintech and a logistics company, as well as Carfax. Basically, all the infrastructure for this to work didn’t exist in LATAM, so we had to build everything vertically so we could serve our customers the right way.
I’m going to start with the framing of what the architecture looks like. A consumer comes in and says, “I want to sell my car.” How many agents do they touch? What does the harness look like? Ground us in how you designed this.
We bet the company on transforming into a company run by agents. The question we asked ourselves was: How would we build Kavak in 2035 with fable 10 or GPT-10-level intelligence? That company looks very different from what we had built or what we had back then.
When a customer comes in right now, an agent gets spawned specifically for that customer, with its own virtual machine. It’ll remember years of this customer’s interactions with Kavak: what they visited on the webpage or a call they had 2 years ago. It’ll remember everything in its memory, come up with a strategy, and set a long-term goal to maximize the lifetime value of this customer.
It will do whatever it takes to make the customer happy and convert them into all our different products over time. This is a completely new and groundbreaking architecture at scale. I think people are still building multi-agent systems with experts, and we realized that we should bet on long-running agents with hard goals, not just workflows, that could maximize our customers’ satisfaction and, obviously, their lifetime value.
2. Three Bets: Redesign the Company, Build Superhuman Agents, Change the Metrics
Awesome. Okay, so we’re going to jump to the nuances. Many companies say, “Hey, we want to be agentic,” and they try some workflows. You guys took the bet that we had to make this work. You had to downsize dramatically, and it didn’t work for a year.
Right.
So, do you want to talk through that? Obviously, you had to tune a lot of things to make that work. Describe the harness at that time and, specifically, what models you were using.
There were 3 main decisions that we had to make. The first—and this is where I think many companies are stuck right now—is that the first instinct is, “Okay, let’s adopt AI.” You basically leave your structure as it is and just give ChatGPT to your team, and then there are no efficiencies. Your customers have the same problems, and nothing happens.
You need to redesign your whole company around the agents and around future capabilities. This means really rebuilding most of your APIs and rebuilding your systems so the agents can use them to perform.
Then you need to start generating the data and the feedback loops to fine-tune these agents. The only way to really make them work is if you teach them. How do you teach them? You put them out in the open. You put them in front of customers, get that data, get those evals, and then train your agents.
This was the second bet that we made: that we could build superhuman agents. This means that, by every dimension that matters—conversion, lifetime value, customer experience—our agents would outperform the best human we had ever hired. We put them in front of the hardest problems.
Finally, you start to change how you measure the success of the company. Kavak was a transactional company. We used to measure how many cars we bought, how many cars we sold, and how many brake pads we needed to buy.
We moved to a relational company, where now I have 10 million customers in my database and agents assigned to most of them, with the task of maximizing their lifetime value. Now, we’re selling cars, personal loans, and very high-ticket items. Activating 1% of this customer base is hundreds of millions of dollars if we do it the right way.
It’s a bet that made sense for us because of our industry, because of the ticket, and because, at the end of the day, customers need to build trust with a company because they’re buying a used car. The way to build trust is to know them and to plan and nurture a long-term relationship.
I just wanted to double-click on something: evals over agent demos. You probably get pitched a lot of new agents, and it’s never been easier to build things. One of the questions is: How do you guys go about evaluating this? Not everybody tests them across 90% of customer interactions to see if they’re really working. And you guys—I believe it’s about 98% of the interactions, or something like that—are now handled by agents.
Yes, totally. To give you a sense of the scale, 96% of all interactions are handled by agents. There are no humans there. About 95% of all transactions are completely handled by agents.
Obviously, you meet a human when you pick up your car. There’s someone physically there to give you the keys. But the rest of the experience, the journey, is handled by an agent.
Every day, between 100 and 200,000 agents get instantiated. They wake up, they work—sometimes for 3 minutes, sometimes for 8 hours, sometimes for 3 days—and they set an alarm clock for their next task and go back to sleep. The scale of this is just amazing, and it’s working.
Now, how do you get this to work at scale? The answer—you mentioned it—is evals. I like to move extremely fast, but in order to move fast, you need to have brakes, right? Imagine a car: you’ll hit the gas only if you have the right brakes.
AI is super powerful, and I’ve seen many companies get this wrong because they try to go slow because they don’t have the right brakes. I thought about it the other way around: How fast can we go? It depends on the quality of our evals.
A good rule of thumb here is that we spend about the same amount of engineer time, tokens, and money on building the evals as on building the agents. This is how you get better and better and better—not treating evals as an afterthought.
What do we measure first and foremost? The results for the business. If my customer is happy, they’ll buy a car, they’ll get their loan approved, they’ll sell a car to us, and that’s the first check: Did it convert? Is it bringing value to the customer, and is the customer happy to re-engage with us after a while?
3. Agents That Sell: 2.1x Better Conversion Than Humans
Once you get those evals connected, then it’s just optimizing the right agentic architecture and giving the agent skills to scale this and cater to millions of customers. It’s really amazing. Related to this, once you create the right evals and it’s working, some companies still feel a little risk-averse about putting them in front of customers and being able to perform
In your case, that highest-leverage task would be selling.
Do your agents really sell to customers?
Yes. We never built customer support or customer service agents. We built sales agents. It's extremely hard to sell a car in Latin America. Imagine someone wanting to buy a car: They can choose among 20,000 SKUs. Then they need to pick financing and go through the financing process, insurance, and coverage. They're probably trading in their car, too, so we need to quote that car.
It's a process that, if someone does it—or the way Kavak did it back in 2020 and 2021—you need to be extremely good at 15 different things and have 15 different experts in 15 different teams. Usually, the person would speak with the expert in financing, the expert in car advisory, the expert in buying, and the expert in insurance. They'd build a package and buy a car. That's extremely hard to do.
The first thing we did was ask, can we get an agent to be better than the expert in each of these things, and then put it together and have a mega-expert that's an expert in insurance, financing, and everything else? That's who we put in front of the customer. The experience for the customer is amazing. We tripled NPS and customer satisfaction scores by putting the agent in front of the customer.
At first, it converted 50% more than our human team, and now it's converting 2.1× more. It's a completely different concept.
Your agents are better sellers.
Totally better. You get this right because they're experts, they're infinitely patient, they know all your history, they can plan for the long term, and they never get tired. If they make a mistake, they learn from it, and the next day, not just that agent but the other 200,000 agents will have learned from that mistake. That's the feedback loop we engaged, and that's showing in the growth, results, and satisfaction of our customers.
Yeah. One of the two very cool things I think about Kavak is that the world has gotten comfortable with the idea that AI can do customer service. It's still very hard to do well, but as Gabe said, there's still a view that customers aren't going to want to buy expensive things from AI, and you are proving them wrong.
The next layer on that is, well, you're not actually going to be able to do regulated financial services end-to-end with AI.
But if you walk through what you're doing, you are underwriting a thin- or no-file customer, pricing them correctly, and doing servicing. So maybe talk through how you wrote the evals to get comfortable with that, and then, versus going to a bank branch or even a fintech, how is that experience?
Yes. So much better.
4. Car Loans Approved in Three Minutes
The first financial product that we launched was a car loan. Usually, in Mexico and some emerging markets, it'll take 2 months or more to get a car loan approved. We usually approve it in under 3 minutes, which is pretty cool because we have all this data around the customer and the car.
If the customer can't pay for the car anymore, they'll just return it to us, and we can give them a cheaper car. Then they're paying a smaller amount each month, and they get out of the water, which is amazing about the vertical integration of the business.
When we started launching other financial products, we realized that this is a very important decision for the customer. They usually take 3 to 4 months to make up their mind about buying a car, getting a loan, or getting a large personal loan, which we also do. If you get to know your customer throughout this process and make the process easy for them, then your conversion and retention metrics start going through the roof.
It's not just the transaction; it's understanding each customer personally and getting them to convert when they're ready, with deep personalization of the interest rate, the risk, and the maximum amount of the loan. It's done in a way that makes sense for the portfolio as a whole, obviously, but is optimized to the customer's risk level and probably the other offers that the customer is getting.
And then maybe give us an example. Evals are always a very hot topic—you kind of led with that. What is an example of a hard-to-design area for evals, or one where you had to spend extra time, given that there's real money and PII at risk?
5. The AI CEO Experiment: 1.5x Profits in Six Weeks
When we decided to redesign the company around AI, we asked the question: Is AI going to be able to do this job—even the CEO's job, or jobs where the leadership is? The answer, honestly, is probably yes. At this rate of improvement, by 2035 it will be able to do it.
So we said, let's try it now. Let's try to build an AI CEO. We carved out a city in Mexico—it's Cuernavaca—and we put an agent in one of our harnesses as a CEO. It started learning, making decisions, and evaluating those decisions. It's only been running for 6 weeks now.
The goal of the first month was to double the profits of Cuernavaca.
It didn't reach that, but it was 1.5×—50% more profits—just managing the city, which is crazy. It's amazing. It's the CEO. People thought that was the last job AI was supposed to take, but it isn't. How did this happen?
It's like a very smart person—Fields Medal-level smart—going into every single number and every single customer, making the perfect forecast, and micromanaging every single thing that needs to be executed every day to reach a plan. It will literally send messages to all the physical workers in Cuernavaca with their plans for the day and ask them to send voice notes back to know their progress.
Customer satisfaction grew. We got better inventory, rotated it better, and had better financing penetration. Every KPI started to improve.
It's super cool and super exciting. What are the jobs where we think we're still training and hiring humans? Those are related to the physical world.
When we talk about mechanics, Kavak has around 800 mechanics in Mexico. There's a lot of dexterity and senses involved that's super hard to substitute. So there, we also build these agents with the exact same harness, and it's scaling. The mechanics have a sidekick.
I was telling you guys earlier that it's like the movie Ratatouille—the mouse that's actually a chef collaborating with a human. It's kind of like that. It's a sidekick. We call it El Mike, and it tells them how to inspect a car, gives them tips, and shows them how to do it.
The quality of inspections went through the roof. We're inspecting faster, repairing faster, and it's cheaper. Most importantly, we're delivering higher-quality cars. Warranties came down around 26% since we launched, and customer satisfaction went up again.
It's about this question: How would you design your organization from scratch with abundant superintelligence that's cheap? Just go build it.
Now, this is a good segue to a key topic right now in Silicon Valley, where there's a lot of people worried about what the organizations of the future are going to look like and the role that humans are going to play in this. I think you touched a little bit on that, so we'd love to hear how you guys are thinking about that.
Yes, and the organizations. Yeah.
6. The Jedi Academy: Training Mechanics to Ship Agents
Totally. We took that question very seriously 3 years ago, and the truth is that everyone's job will change. What we were doing a couple of years ago will probably be performed better by an AI agent.
What does this mean? We need to train everyone. So we launched a program inside Kavak called the Jedi Academy, where anyone from Kavak—from the CEO—
To—yeah.
And it's awesome. From the CEO to AI engineers to mechanics, everyone's going to the academy. It's super hard. I led them myself.
You designed the program.
I designed the program.
Constantly, because you need to be upgrading the program because everything's changing so fast. You can't send these people outside to Stanford to learn this because it's new stuff. So we train everyone, and after 6 weeks they launch state-of-the-art AI agents into production. It's mechanics and finance guys and engineers—everyone can do it.
What this generated is that maybe this person won't become an AI engineer—some of them have—but they know how to collaborate with this new technology. The way we looked at it was: There's no way back. This is the way Kavak is going, and this is the way the company will look. These are the changes for the engineering team, the finance team, the product team—this is what's going to change.
You have the choice to train and get the skills to perform in this new reality, in this new world, or maybe leave Kavak if this is not for you. But this is the way we're going.
It was great. We strengthened the culture, everyone was super excited, and people really know how to build these agentic systems. If you look at Kavak now, any process is really a collaboration of agents and humans. Sometimes agents are the bosses of humans, and sometimes humans are designing the agents.
I think we managed to really build this and change the company through the idea that we need to be learning every day and things will continue to change. The only way to continue being relevant is to upgrade your skills every month or every couple of months.
But you do have—or did have—thousands of people. Now agents do most things. What is the org structure of Kavak? Does the middle-management concept even exist anymore? What does your org look like?
Right. So the way it looks now is very flat teams, very senior teams, super empowered. If you look at a team, you'll have engineering, AI, operations—everything—and they're either building the agents, working for the agents, or being in the physical world in front of the customer.
Yep.
Most of our organization looks like that. So it's really built around the idea of how organizations will look in the future and around AI, really harnessing this new technology.
Obviously, this required lots of retraining because in 2023—or 2022—no one was building agents, no one was helping agents or taking orders from agents. And the way you cater to the physical world or the customers was different from how it is if an agent is telling you what to do or helping you make your job better. So it's a completely different structure from what we had just 2 years ago.
Yeah. Explain—we talked about this before—what working for the agents looks like. I think the way you described it was an agentic system, and then sometimes when it fails, it's kicked out to a kind of human queue, right? But then that's lost. So how have you brought that together?
We see humans in the loop in most of these agentic systems in production right now. Large-scale agentic systems usually, if an agent hits a wall or can't perform anymore, it'll send this case or this customer to Tier 2 support and forget about it. That doesn't really work because you don't close the loop, so you don't generate the data to train the agent to do this better.
What works right now is that we have an agent that's obsessed with each of the customers—millions of these. They have access to every single API, every single skill, and we have humans building those skills for them. If an agent hits a wall or cancels something, it'll call this API, saying, "I need help." On the other side, it's not an agent or software; it's a human helping them out.
But if you map this out in an org chart, it's really human teams that have an agent, and I'm getting better results. It's super clear; it makes sense.
That's actually a perfect segue. I know you get lots of leaders at larger institutions inbounding to you. Maybe this will save you many phone calls, but I think rationally many leaders of companies intuitively understand this. It's still very hard to deploy AI through their organization. The models are good enough that it's an org problem; it's a psychology problem. What advice do you have, or what have you seen?
I think it's 2 things. The first is it has to be top-down because if you just get adoption, it won't go anywhere. It's hard to generate this taste or strategy for people to decide bottom-up what to build and what not to build, and come up with something that works for the company. The transformation has to be top-down, and leaders need to adopt. Leaders have to have a very clear plan on what to build.
I've seen so many companies where it's just, "We're doing a hackathon. People are coming up with use cases. We're sponsoring some of these use cases." That doesn't work. Be very clear on what the company will look like in 3 or 5 years, and then start building that. Be very vertical in guiding your troops toward that. An army doesn't really work if everyone comes up with ideas on the strategy and tactics, goes to the battlefield, and does whatever they want. You need a very clear strategy, and that's what we need now. It's a transformation stage.
The second one is you need to measure what really matters, and it evolves—but it also evolves in the right way. I see a lot of companies spending huge amounts now, and they say, "Okay, I got adoption. I'm just spending hundreds of millions of dollars in tokens now."
What about that? There's quality in the tokens.
So I have a framework here that's also useful. Tier 3 tokens—the most valuable—are these agents where you can get the ROI of each specific token, and I can do that now. That's great news for me because I'm growing and because I know the ROI of each token, because it goes to agents that are performing the job of the organization. These are the best tokens.
Tier 2 tokens are things that you can measure indirectly. Do I see devs in the codebase? I can evaluate the value of these tokens at least indirectly and then push those to production. Tier 1, where most companies are, is people are just using Claude Code or ChatGPT or Cowork or whatever. What happens with those? I have no idea.
So it's not just about adoption. It's really about having a clear vision, measuring that each token you spend is bringing you those benefits, and iterating, iterating, iterating from there.
And I want to—and we touched on this a little bit, but I think it's worth a dive—as maybe the most ambitious companies listening to this will decide to follow suit—which is: You decided to build an agent per customer versus per task, and then discovered along the way that each one of those agents needs its own micro virtual machine. So maybe walk us through those decisions, that architecture.
Yes.
Those decisions, that architecture.
7. Destroying Two Years of Work: From Multi-Agent Graphs to One Agent Per Customer
And I think we're seeing these results now, but it was a really risky bet, because people usually go from workflows to graphs, functions, or objectives. If I could advise everyone, don't build agentic workflows. We built those. These are multi-agent systems that can perform a whole function for a complex goal, like the one I told you about: To sell a car, you need to do financing, purchasing, recommendations, et cetera.
We had thousands—tens of thousands—of these agents working at scale, running the business back in December. But then Opus 4.5 came out, and I realized this isn't the right paradigm anymore. The intelligence now doesn't need the graph and the multi-agent latticework and harness, because that will constrain this level of intelligence.
So we decided to destroy everything we had been building for 2 years—everything that was working, that brought us to profitability and amazing growth—and start over with a harness that we thought would be robust and scalable and leverage recursive self-improvement, or new, more intelligent models coming out every month.
The way this looks is a virtual machine with an agent, with access to memory and evals and the CLI, where they can access every tool and every API in my company. The long-term goal is that I instantiate hundreds of thousands of these each day with long-term goals like maximizing lifetime value.
Yeah. The self-improving organization.
Exactly, the self-improving organization. I think people are super obsessed with RSI now, and this will improve the models. But if you look at it this way, economic value in humanity for the past 4,000 years has been delivered by organizations, not individuals.
So what you want to self-improve and to engage in that loop is the organization that can deliver more economic value. That's the loop that I think companies will start to focus on, because if you get that loop working and it's an organization that is really self-improving and harnessing the newer models and the better intelligence that we're getting every couple of days now, then you hit the exponential—not just in intelligence, but in the value that you can generate as a company.
So that's really exciting. That's what we're working on.
You mentioned that because of all the challenges in adopting AI, you saw the biggest opportunity in net-new companies being formed, working on this new way, and then disrupting markets. Do you want to talk a little bit about that?
Yes. There's this concept in economics about creative destruction from Joseph Schumpeter. What it says is that the way innovation hits the economy isn't by companies adopting the new technology, but by companies remaining the way they were and incumbents with the new technology destroying the old companies.
This destroys value in the short term in the economy, but in the long term it's better for everyone, because these new, more efficient, more effective companies will provide better products and services for the economy as a whole. This has happened in the past, like in industrial revolutions. This has always happened, and this is a great opportunity for entrepreneurs and people today because it's hard to adopt AI deeply.
It's really hard for a CEO today, especially of a large company or public company, to go and say, "Hey, I'm betting everything on AI. The company has to look this way. I'll destroy and rebuild everything I've been building for the past 40 years to become an AI-native company." How many CEOs will do that in a company at scale?
8. Creative Destruction & Ford's Factory: Why Adoption Isn't Enough
So while they adopt, new companies can be formed that are built around the strengths of AI and take over and bring new products and services to the masses. This has happened before. This happened with electricity. This is a story I always tell my team.
The technologies for Ford's production line were developed in 1879 and 1881. Edison started commercializing electricity in New York and then London, and he invented a dynamo that was extremely efficient. So you could have built Ford's factory 40 years before Ford.
The technology was there; everything was there. But the way people adopted electricity and the dynamo was, “Okay, I’m going to leave my factory like 4 floors, shafts, and belts, and just change my coal engine for an electric engine.” This will bring you benefits, yes, but only like a 6% efficiency improvement.
What needed to be done was to destroy that factory, build it on a flat surface—not in the center of New York, but in Connecticut or New Jersey—and redesign your whole factory around small dynamos and electricity. Then you get like a 3× improvement in productivity that powered the US during the 20th century. The same happened again with the computer, and the same is happening again today. People want to adopt it, but they’re not willing to redesign the whole company; they just adopt it superficially.
In the end, that’ll give you a 6% or a 10% improvement, not a 10× improvement. It’s like the Innovator’s Dilemma at an industrial scale.
9. Advice for Founders: The Most Exciting Time in Human History
Again, I think you’ve just made an amazing case for any future founders out there that it’s time to build.
It’s time to build. Maybe a great place to end is—you’ve built and scaled your own company, and you’ve now turned Kavak fully agentic. What advice do you have for future founders or first-time founders who might be listening?
This is the most exciting time in human history. I believe that we’re living in the most exciting time in human history, and it’s the most exciting time to be a founder because it’s the first time that anyone has access to the most powerful tools and intelligence in the world, almost for free or for $20 a month.
Literally, the democratization of the tools for people to build has never been this way in human history. There are so many problems to be solved and a new reality to be built around this new paradigm. So just go for it, but go for it deep. Imagine what the future around AI will look like.
It’s not even exponential. Just map a trend that’s linear. If AI keeps getting better at a linear scale, build for that, and you’ll come up with wonderful ideas that will bring a lot of value to the world.
Amazing. Alejandro, thank you for joining us.
Thanks for having me.