AI agents 的崛起:CrewAI 的 João Moura
CrewAI押注,企业价值最终会沉淀在控制平面,而不只是构建 agent 的框架上。Moura预计,3年内企业将运营“数千个、甚至数十万个”agent,由此催生对规划、构建、部署、监控、集成、身份认证和访问控制的需求。否则,这些 agent 最终会变成数千套彼此割裂的遗留代码库,而不是可管理的企业资产。
应用会从低精度、人工复核的工作起步,再逐步向更多决策环节扩展。目前的应用横跨销售、营销、后台流程、编码、IRS表格、媒体剪辑和运行时定价。Moura定义未来价值的公式是“任务复杂度×自主性”,但“我们还没进入最上面的四分位”,关键流程目前仍无法脱离人工盯控运行。
高管支持、技术负责人和明确痛点,是区分靠谱企业买家与 agent 观光客的关键。如果一家公司来问 CrewAI“应该自动化什么”,通常是弱信号;成功客户往往清楚自己想改造哪项流程,也有工程师能够打通内部系统。一些财富500强客户已经从简单工作流推进到让 agent 在监控竞争对手的同时调整市场价格。
规模正在加速,但 CrewAI通过专业化来约束复杂工作,而不是放任无限自主。Moura称,仅1月就超过5000万个 agent,并介绍了一个由21个 agent 组成的团队:金融、品牌、定位和营销分别由专门 agent 负责,最终产出数十页的竞争情报报告。但面对一份约60页、背后配有720页手册的 IRS 表格,代码会先把问题逐个隔离,再由 agent 逐题研究和作答。
agent 经济正在同时冲击软件定价和模型经济学。当一个高生产率 agent 足以取代一个人类席位时,按席位收费的 SaaS 模式就会变得尴尬;不过Moura预计,面向 agent 优化的端点会比很多人想象得更晚出现,因为让 agent 学会使用现有的人类界面,就能“解锁整个互联网”。与此同时,o1和o3能力很强,但在 agent 工作负载中尚未大规模起量:构建者偏好的是只要能完成任务、就尽可能便宜且快速的模型,包括 GPT-4o mini。
开源模型具备战略势能,但目前 CrewAI 的大部分工作负载仍由闭源模型承担。受监管的金融和保险客户越来越倾向于在“气密容器”中自行部署模型;R1在美国的采用情况释放出混合信号,但进一步强化了Moura对开源模型继续推进的信心。他预计,微调将重新升温,更小的模型可能在 agent 工作负载上胜过今天的700亿参数模型。
短期劳动力影响更像组织杠杆,而不是全面替代。Moura设想,团队可能从4个人做重复性工作,变成1个人监督 agent、另外3个人转向其他岗位;未来几年内,agent 参与决策的程度可能进一步提高。他给出的谨慎建议是学习使用这些工具:这可能不够,也可能足够,但从统计上看,至少会改善个人处境。
1. agent 数量增长后,控制平面成为必需品
Moura的3年判断是,企业将使用“数千个、甚至数十万个”agent。战略分岔非常明确:这些 agent 要么变成数千套遗留代码库,要么由企业通过控制平面统一管理,成为一项资产。
生命周期远不止构建:企业还必须规划、构建、部署、监控和集成 agent,并处理身份认证、范围设定、访问权限、市场等相关需求。CrewAI希望覆盖这套技术栈;Moura提到,微调是一个足够复杂、可以由另一家公司解决的相邻问题,目前 CrewAI 并不提供该服务。
CrewAI的平台支持企业配置角色与权限、LLM连接和私有代理;既可以用开源框架或 CrewAI Studio 构建,也可以部署带负载均衡、自动扩缩容和 SSL 的生产级 API;同时监控输出质量、幻觉和自定义指标。用户还可以用两次点击,让 agent 针对新模型进行测试。
他的定义非常清晰:“agent 必须具备 agency”。固定的 RAG 流程如果每次都以同样方式检索和总结,只能算工作流;只有当 AI 能决定下一步做什么,并在路径 A、B、C 之间做选择时,才成为 agent。Moura把不加区分地给一切贴上 agent 标签称为“agent washing”。
2. 生产级应用青睐明确痛点、高层支持与专业化
已在运行的应用横跨销售、营销、定制化后台自动化、代码生成和 IRS 表格。在前沿应用中,Moura介绍了媒体 agent:它们跟踪直播体育画面中的球,剪辑片段、添加字幕和声音,再发布到社交媒体。他强调,许多公司仍在从更简单的用例起步,然后逐步扩大规模。
最强的客户筛选信号包括:高管积极支持、能够打通自研系统的技术人员,以及一个明确选定的用例。Moura认为,负面信号是公司反过来询问“应该做什么”或“别人都在做什么”,因为项目背后“未必存在明确的痛点”。
CrewAI官网称,财富500强中有40%在使用 CrewAI,其中相当一部分采用始于开源项目。Moura回忆,公司成立的“顿悟时刻”来自 Oracle 一位员工:对方说项目已经在生产环境运行,并要求提供支持。他还重点提到 GCP、金融和保险;尽管金融行业推进很快,但受监管影响,部署前仍需要更多沟通。
Moura称,官网显示的“超过1亿个多 agent 团队”这一数字需要更新。他表示,仅1月就超过5000万个 agent,不确定是否已经达到10亿,但称“可能正在接近”。他最看重的一种配置包含21个专业研究 agent,分别覆盖财务、品牌、定位和营销,之后汇总成数十页的竞争情报报告。
3. 内部数据创造价值,agent 使用则打破按席位定价
研究和抓取工具可以覆盖许多简单应用;当 agent 能够接入 Salesforce 等 CRM、SAP、自研系统、API 和数据湖时,更深层的价值才会出现。真正重要的突破,是访问企业内部数据和系统。
Biewald提出的质疑值得保留:当数千个 agent 访问 Salesforce 时,按席位收费到底意味着什么?Moura认为,席位不是正确方向,因为 agent 可能高效到根本不需要对应的人类席位,由此带来收入自我蚕食的问题。
面向 agent 优化的端点终将出现,但Moura预计其到来时间会晚于很多人的判断。让 agent 遵守原本为人类设计的界面更容易,因为一次适配就能“解锁整个互联网”;如果围绕 agent 重做每项服务,负担则会反过来落到服务提供方身上。他预计,定价模式也会调整,计费对象将从人类消费数据转向 agent 消费数据。
CrewAI有意移除了对 LangChain 的依赖,因为后者会覆盖某个被导入类约90%的内容,同时让代码 rebase 变得非常困难。Moura称这是“我们做过的最佳决定”,原因既包括技术限制,也包括客户在使用该依赖时遇到的问题。CrewAI仍支持 LangChain 和 LlamaIndex 工具、自有工具,以及 Composio 提供的300多个集成。
4. 可靠的自主性来自狭窄范围、硬性停止与记忆
对于开放式任务,CrewAI可以限制请求数量和执行时间,并在接受完成结果前执行程序化护栏。Moura举了一个刻意简单的例子:用 Python 检查单词“but”出现的频率,如果结果显示文本质量不够,就把输出重新送回系统。
高风险任务会被进一步拆解。面对一份约60页、说明手册长达720页的 IRS 表格,普通代码先提取每一页的问题;agent 再结合 RAG、内部数据库、数据湖和手册逐题处理。“agent 始终只在单个问题的范围内工作”,人工也可以在任何节点被强制拉入流程。
记忆正在成为“基本配置”。短期记忆是协作 agent 共用的沙盒,每次运行结束后重置;长期记忆则跨执行周期保留;CrewAI还提供实体记忆和用户记忆。用户记忆可以从一开始就预加载文档、PDF和其他材料。
由于 CrewAI 要求每项任务都设定预期输出,长期记忆可以比较实际结果与预期结果,并生成未来的规则或校验。运营方还可以设置60秒上限:时间到后,系统会要求给出当前最佳答案,结果可能成功完成,也可能失败,但不会无限运行。
5. 模型选择看任务经济学,而不是基准测试光环
Moura将 agent 价值定义为任务复杂度乘以自主性。当前最现实的模式是人机协作:过去由4个人完成的重复性工作,未来可能由1个人监督 agent,另外3个人转向更有意思的任务。他预计,未来几年内 agent 将更多参与决策,但距离全面取代工作还很远。
Biewald提出的问题是:为什么具备 Putnam 数学竞赛水平、还能通过 LSAT 的系统,在商业场景中仍然会出错?Moura的回答是,数学题有明确的正确答案,强化学习可以围绕它进行优化;现实商业决策则充满细微差别,需要判断当下什么才是最佳选择。在前一种任务上的能力“并不一定意味着”后一种任务中的判断力更强。
o1和o3都是“惊人的模型”,但Moura尚未看到它们在 agent 部署中大规模起量。主导工程原则仍然是:“我最低能用到什么程度?”如果 GPT-4o mini 已经能满足要求,它的速度和成本就可能胜过更大模型的额外能力。
尽管Moura相信开源最终会凭借可及性和规模“胜出”,但在 CrewAI 的开源和企业使用中,闭源模型目前仍最常见。Ollama 常用于本地执行,受监管的金融和保险用户则偏好自行托管。R1在美国的采用释放出混合信号,但它的蒸馏版本以及与其他模型的组合已经激发了新一轮工作,也可能推动开源推理模型大量涌现。
6. 无代码扩大分发范围,但代码仍具优势
Moura预计,无代码将拿下可观的市场份额,因为非程序员的数量远超程序员。但关键限制在于:代码的可定制性仍然更强,因此能解锁许多影响最大的应用;这就像即使模板已经足够强大,定制化网站仍然需要工程师。
CrewAI Studio不会把新用户直接丢进“30个节点”的迷宫。它的“渐进式演化界面”先从聊天开始,再把对话转换成由 agent 和任务组成的结构化表格,最后才展示节点视图。到这一步,用户已经理解自动化是如何组装的,可以继续对话、进一步定制,或者下载代码。
CrewAI自身也在用 agent 做营销、会议准备、客户支持和 pull request 审查,覆盖开源与闭源项目。它还会处理会议录音和文字稿,跟进后续行动与演示文稿。最典型的内部应用从新客户的姓名和邮箱开始:agent 调研客户账户,推断可能的应用场景,把这些假设写入 CRM,并据此定制产品本身——让用户忍不住问:“这个应用怎么知道我正准备往那个方向走?”
Moura对2025年的预测是,微调将重新回归,驱动力来自 DeepSeek-R1 在小型模型上的工作。他预计,更小的系统会变得“非常、非常强”,可能击败当时的700亿参数模型,并且更容易针对 agent 进行微调。至于职业风险,他的建议是学习这些工具——“可能不够,也可能够”,但这是个人能够控制的下注。
All right, João, thanks for taking the time. I really appreciate it. Can you start by describing the problem you’re trying to solve with CrewAI?
Yes. By the way, thank you so much for having me. I’m very excited to be here today.
The problem we’re trying to solve with CrewAI is that, 3 years from now, most companies—especially enterprises—are going to have thousands, if not hundreds of thousands, of agents working for them. Either that’s going to mean thousands or hundreds of thousands of legacy codebases, or there’s going to be something that allows them to manage those agents as an actual asset. We’re building the control plane that allows them to do that.
Interesting. What do you expect agents to be doing?
I think it’s going to be a little all over the place. We’re already seeing that there are no clear winners within these companies. It’s very much cross-horizontal: people are doing back-office automation, coding automation, support, and all kinds of other things.
It’s a little bit of everything, and if anything, that just makes the challenge more exciting to me. If you prove that something is doable in one specific horizontal, it’s very easy to start talking about all the different horizontals you can expand to.
1. Signals of success: What makes AI agent adoption work?
I think it’s going to start with what we call low-precision use cases, where you still have humans in the loop and people are reviewing things. Gradually, though, you evolve into more decision-making, and things are going to get very interesting by then.
What does the control plane do?
When you think about the life cycle for these agents, they start with planning. A lot of people are focusing nowadays on building, but building is just one component. There are frameworks out there—CrewAI is one of them, probably the best one, if you ask me and many other people—and many others as well.
But it’s not only the building. You have to plan, and then once you build, you have to deploy. Once you deploy, you want to monitor, and once you monitor, you want to integrate. That’s not even mentioning things like authentication, scoping, access, marketplaces, and everything in between.
What we’re trying to do is cover the entire stack and make sure that these companies are equipped to build, deploy, monitor, and integrate those agents.
2. What AI agents are actually doing in companies today
What’s working today? What are companies actually doing with agents right now?
I’ve got to tell you, it’s kind of impressive. If you look at the simpler use cases, there are a lot of people doing sales and marketing. Then you have more advanced companies doing back-office automation, where they’re automating more custom processes within their businesses.
Then you start getting into some of the cutting-edge work. You find companies trying to tackle very complex problems, like automating the entire life cycle of code creation or trying to automatically fill out IRS forms, which is something you don’t want to get wrong.
There are also people working on even more cutting-edge applications. We’ve seen a media company using CrewAI agents to automatically edit footage. You have live footage of games streaming on TV, and then agents that can track the ball, automatically edit, cut, add captions and sound, and push that into social media.
There’s a little bit of everything, but I want to say that it’s still very early days. A lot of these companies are early in their journey, starting with simpler use cases that they can then scale.
What kinds of companies are successful at deploying agents? What are they doing? When I talk to most companies, they’re interested in agents. There couldn’t be a hotter topic right now, but most of the companies I talk to haven’t successfully gotten agents to do anything except maybe a little bit of coding, chat support, or internal support.
Those are places where I see things working, but it seems like you’re probably talking to companies that are really nailing this. Why is that?
I think it’s a combination of things. It’s funny that you ask, because that automatically becomes part of our qualification criteria when deciding which companies we engage with more deeply.
3. The impact of R1 and open-source models on AI agents
As you said, it’s a very hot topic right now, so everybody and their mother wants to talk about AI agents. Because we’re the leading platform, everyone comes to talk to us. What we have to do as a business is figure out who is actually real and check a bunch of those boxes so that we can go deep with them and really bear-hug them.
Usually, the signals we look for in companies that translate into successful use cases are executive support. There’s an executive saying, “We really need to embrace this. We know this is where our business is going.”
There’s usually technical support as well. Even if the buyer isn’t a technical persona, there’s generally someone technical within the project and scope who can unlock internal integrations with homegrown systems and other things that really unlock the power of more custom use cases.
4. AI agents for research: A 21-agent team working on market intelligence
There’s also an understanding of which use cases they actually want to pursue. If a company approaches us asking what they should be doing and what they’re seeing other people do, that’s usually not a good signal. There isn’t necessarily a clear pain point or something they’re trying to do, and they haven’t spent much time thinking about it.
Those are some of the signals we look at. On a day-to-day basis, a lot of companies start small and then expand into other things. We have more advanced customers—Fortune 500 companies—that started with simple use cases and now have a lot of their pricing flows automated. They adjust prices using agents at runtime in marketplaces and other places, have agents monitoring competitors, and do a bunch of other things throughout their companies.
There are some very interesting use cases out there.
Do you have a specific way that you define agents? I feel like, with everything in AI, once it gets hot, the definition expands. Would you consider a RAG application an agent, or does it have to be more advanced than that?
I think it’s not about how advanced it is, but I would not consider that an agent. It’s funny that you say that, because we’ve been talking with some people from Gartner, and I think they refer to this as agent washing: companies stamping something as “agentic now” and trying to pass it off as an agent.
I think that’s a short-term detriment to the industry, but over the long term, things are going to consolidate and figure themselves out.
My definition is that agents require agency. You can have workflows, and you can have AI help with those workflows. I think that, at the point where the AI is actually controlling what happens next, you have an agent, because the AI is guiding the process and choosing between A, B, or C.
A RAG application wouldn’t count because it always uses the same tool for the lookup and then the same tool for summarization. It doesn’t choose its own path.
Interesting.
Exactly. It’s kind of like “if this, then that.” That’s not necessarily an agent. If you get that RAG output and do something with it, and depending on what you’re doing, different things might happen—you don’t necessarily know what that might be—then you have an agent.
5. The role of tool use in AI agent success
What about tool use? Tools are obviously a critical component for the adoption of agents. What kinds of tools do you see being used most frequently, and where is this going?
Tools are what make these agents very useful at the end of the day. You have to have tools that allow the agents to connect with internal data, external data, or whatever else you want.
There are simpler tools that you need to have, such as tools that help with research or scraping. That covers a lot of use cases, especially the simpler ones.
Where the value really gets unlocked is when you can tap into internal data. A lot of the time, that means some of the larger enterprise companies that we all know, like Salesforce and other CRMs, SAP, and a few other systems. It also means internal homegrown systems—systems that have been running in these companies for many years and that they need to access through an API or a data lake.
Those are some of the things we usually look for in terms of tools that really unlock value.
6. How Salesforce, LinkedIn, and others are rethinking their pricing models for agents
Do you see companies like Salesforce modifying their APIs or changing their pricing models because agents are using them as tools? For example, what does a per-seat price mean in the context of an agent? If I have thousands of agents and they all access Salesforce, do I have to buy thousands of agent seats for them?
I don’t think seats are the way to go. We’re seeing some companies experimenting with that. I think LinkedIn is doing something along those lines, and I’ve heard that it’s a tough problem for them to navigate, as it is for every other company out there.
Do you charge less for an agent seat? What if the agent is so good that you don’t need an actual seat? You’re basically cannibalizing your revenue a little bit. It creates all these interesting problems that you have to navigate.
I do think there are going to be endpoints that are more focused on agents and optimized for agents. That said, I think that will happen further down the road than most people believe.
It’s much easier to get agents to comply with the inputs and outputs that we humans already use. If you do that once, you unlock the entire internet, instead of trying to do the opposite and translate the entire internet into something agents can use.
I think it’s going to take a little longer, but there are definitely going to be pricing-model updates based on agents consuming data instead of humans.
7. How Crew AI reached 40% of the Fortune 500
Your website says that 40% of the Fortune 500 uses CrewAI. Can you talk about how you got that adoption? Also, what verticals or industries have the most adoption of agents in CrewAI?
The adoption is insane. A lot of it happens through open source first and then eventually migrates into the enterprise.
When I first created this project, I wasn’t expecting it to blow up the way it did. I’m very thankful for it. I think we have an amazing community now.
I still remember being here where I am now, at SHACK15 in San Francisco, and being back in a bar on the other side. Someone from Oracle approached me. Back then, CrewAI wasn’t a company; it was just an open-source project. This person said, “We’re using CrewAI in production. Can you help us? We need some help.”
That was a big aha moment for me, because I realized that this big company was actually using the open-source project and needed help. If I was going to provide them with the help they needed, I couldn’t do that through an open-source side gig. I needed resources, so it might be worth turning this into a company.
It was very organic. Me being enthusiastic about agents and doing a lot of work in public also helped drive adoption. A lot of the educational content that followed helped us penetrate many of those large enterprises.
Nowadays, we have major banks reaching out to me, and I learn that their CTO has been using the CrewAI open-source project. They say, “All right, I guess this is happening.”
In terms of verticals, there’s a little bit of everything. There are definitely a few verticals that are moving faster. I think GCP is impressive in how quickly it’s moving. Finance is also moving a little faster, although it’s a highly regulated industry, so it takes more conversations before they start deploying some of these things.
Those are the industries that have been impressing me. Insurance companies are also moving fast.
Interesting. So finance and insurance are moving the fastest? GCP?
GCP.
Interesting. What about the volume of usage? Your website also says that more than 100 million multi-agent crews have run using CrewAI. What have been the longest-running crews, or however you define that?
We need to update that number.
What’s the number now?
I can tell you that January alone was over 50 million agents. We had a presentation yesterday, and when I pulled the number up, January alone was over 50 million agents. That was insane.
I don’t know if we’ve gotten to a billion just yet, but we might be closing in. It’s insane to see the scale and how fast things are accelerating.
Sorry, I missed the question. The question was, what are the longest-running agents?
A lot of that is one we’re co-building with a customer now, and it will be a long-running one. More than the long-running ones, the most impressive ones I’ve seen are crews with a lot of agents.
I remember seeing a crew that had 21 agents on it, and that was impressive to me. I didn’t see that one coming.
8. AI agent memory: Short-term, long-term, and entity memory
How does that work? Why so many agents? What was it doing?
This one was great. It was a company that sells reports to larger, consumer-facing companies about their competition and market positioning.
The company would charge them a lot of money and conduct extensive research: here are all your competitors, here’s what they’re doing, here are photos of their actual stores, here’s how they’re positioning products in their stores, and a lot of other market and competitive analysis.
For some of that, you need people on the ground, and you can’t replace that just yet. But a lot of the research involved digging through a huge amount of material. You know the sources you want to check, but you also want to do some exploratory research.
That was basically what this use case was doing. The reason they had so many agents was that they went very specific and specialized with each agent and what they wanted each agent to research.
If one agent was researching the financial aspects of the market and competition, that was one dedicated agent. If they had another agent researching branding, positioning, and marketing, that was a separate agent.
At the end of the day, they would produce final reports that were tens of pages long. That was the deliverable they gave back to the customer.
9. How AI agents can interact with humans to avoid errors
What about the different levels of autonomy with agents? For autonomous vehicles, they talk about 5 levels of autonomy. You’ve talked about humans in the loop. How do you think about that? How do you set up agents to go back to humans and engage with them? How do you put checks and balances in place to make sure the agents are producing the intended results?
That’s a great question. There are a lot of different ways to do it. It depends on how much precision you want and how complex your use case is.
10. Open-source vs. closed-source AI models for agents
For most people who want to have a lot of agency and want their agents to go out and figure things out, there are a few different ways to add checks and balances. That might mean limiting the number of requests agents can make or how much time they have to get the work done. Those things can help keep them from going haywire.
You can also add programmatic guardrails. When an agent finishes something, instead of simply saying, “I’m done,” you can run its output through a programmatic guardrail. That might be actual Python code, if you want it to be, that checks something simple, such as how many times the word “but” appears.
If there are a lot of “buts” throughout the output, you know it’s not good enough, so you send it back.
If you’re talking about more complex use cases—for example, filling out IRS tax forms—that’s something you want to be much more careful about. In this particular use case, the form was around 60 pages long. But fear not: it comes with an instruction manual, and the instruction manual is 720 pages long.
For that use case, you want a lot of agents, but you don’t want them to have too much autonomy. We would use a flow. CrewAI flows allow you to intertwine agents and regular code.
You would use regular code to get each page of the form and extract all the questions from that page. Those questions would then go one by one to a group of agents. The agents would perform RAG over an internal database, check a few data lakes, and create an answer for that specific question.
As part of that process, they could consult the instruction manual. The agents are always working within the scope of one question, so they aren’t going crazy and hallucinating about all the other questions and all that context.
Once that’s done, they go to the next question. That’s one way to have much more control over complex use cases.
At any point in time with CrewAI, you can force a human into the loop. You can say, “There’s going to be a human in the loop right here,” and the agents will stop and wait for you to get back to them.
I guess that’s a good segue into talking about the CrewAI platform. You’ve given me some really good examples of agents and why they would be useful, but could you describe which parts of the solution you leave to the ecosystem or other vendors, and which parts of the solution you actually provide for your customers?
There are a lot of parts that belong on the vendor side. Anything involving AI—not even agents, just AI—is very complex. It’s not only about making an API call; that only gets you so far.
11. Defining AI agents: When is it real and when is it hype?
There are problems out there that are big enough for entire companies to exist to solve them. We’re already seeing that. One big cluster is everything around fine-tuning. That feels complex enough that you could have companies focused entirely on helping you get it right.
You have a lot of background in that. If you go back to your first principles, hyperparameter tuning and everything involved in it require a lot of work to do correctly. If you’re just using Jupyter notebooks on your local computer, you get lost very quickly.
I still remember not being able to replicate something and being so mad at myself. I couldn’t remember which hyperparameters I had used, and I couldn’t get the same result.
Fine-tuning is a good example of something we don’t offer right now. Maybe we will in the future, but we already have so much on our plates. It’s a big enough problem for another company to solve.
What do you do, actually? That’s probably how I should have asked the question.
On the platform, we try to cover all the different stages of agent building, deployment, monitoring, and integration.
You can go into the platform and configure everything for your company. That includes inviting the right people, setting up the right permissions and roles, configuring your LLM connections, and bringing any LLM you want. You can use a private proxy if you want to, so you can connect all those things there.
Once you’re set up, you go into building: “I want to build this agent.” In the platform, you can use the open-source framework, but you can also build with no code. We offer CrewAI Studio, where you can essentially chat your way into an automation.
You get the automation going, the agents are created, and you can deploy it right away. Or you can go back into code.
We talked about setting up, planning, and building. Once you build the agents, you need to deploy them, and you can deploy them right there as well. That automatically becomes an API, and it’s production-grade. It includes a load balancer, autoscaling, SSL, and all those different things.
Now you have an API that you can integrate with other systems. You also get a bunch of metrics. Those metrics are more agentic: you can see not only prompts and other basic information, but also the quality of the outputs, hallucinations, and how much of that is happening.
You can set up custom metrics if you want to track something specific and set alerts based on them.
Then we get toward the end of the life cycle, where you want to iterate on the agents. A new model might come out, like o3-mini, and you might want to test all your agents on it to see which ones perform better.
We offer ways for you to run your agents with a new model in 2 clicks. You can see how they perform, change the model, and pick and choose which agents you want to use it with. That makes it easier to constantly iterate on these agents.
The short version is that there’s a lot that goes into all of that. There are a bunch of metrics, a bunch of things around iteration, and other features as well.
What about memory? I feel like that’s a big topic right now. Do you help with memory at all?
For sure. When people ask me about memory, RAG, and things like that, I say, “Yes, of course they must exist, and your agents work better with them.” I’m finding that all of this is becoming table stakes very quickly. You have to have it; there are no questions about that.
There are many different ways to build it. It’s easier to build when you’re using CrewAI because it integrates with any vector database, if you want it to. You can integrate with whatever data source you want, which makes it extra easy.
In CrewAI, the open-source and enterprise versions have short-term memory, long-term memory, and entity memory. We also recently added a fourth type called user memory.
User memory is memory that you can preload into your agent. You can give the agent a set of documents, PDFs, and other materials that go into its memory from the get-go and get stored there.
Short-term, long-term, and entity memory are populated autonomously as the agents do their work.
How does short-term memory work differently from long-term memory?
Short-term memory is used during execution. You might have a few agents working together, and this memory acts as a sandbox where those agents can share information with each other.
They can delegate specific work to one another if you enable that, but they have this common place where they put some of their learnings and the things they’re doing. The agents working together can tap into it.
That gets reset on every run. You do a run, everything gets cleaned up, and then the new run happens.
Long-term memory is where agents store learnings from multiple executions. In CrewAI, we require you to specify the expected output for each task, so we have something to compare against.
We can take the actual output, compare it with what you expected, and see what the agents did right or wrong. Then they can autonomously create rules to follow in the future, such as, “You didn’t get this thing right, so let me create a validation for that.”
Over many executions, that makes sure your agents get better and better.
Do you have ways of automatically monitoring agents when they run for a long time to make sure they haven’t gone haywire?
You can not only monitor them, but also set specific hard stops. You can say, “You can only run for 60 seconds.” When it reaches 60 seconds, we make a final call saying, “You’re done. Give us your best answer right now.”
That either goes through, or it might blow up and fail. You’ve configured it not to go over the limit.
What do you see in terms of open-source versus closed-source adoption? Which is more popular? Do you have specific recommendations? If I came to you and said, “I don’t care; I just want the best model,” where would you guide me?
You’re talking about open-source models?
I’m a huge believer in open source. I’ve been living and breathing open source for many years. I’ve had quite a few projects throughout the years, and I’m a strong believer that, at the end of the day, open source will probably win—whatever “winning” means here.
I don’t think that necessarily means closed models won’t exist anymore. I think it means that models will become more accessible and easier for people to run at scale.
The models that are used most often today are closed-source models. Those are the most common in our open-source and enterprise products, but we’re seeing open source take off as well.
A lot of people are running agents locally using things like Ollama. It’s very common for us to see people doing that, and it’s more common in enterprises when you’re talking about highly regulated industries.
For example, finance and insurance companies usually want to self-host their models. They want the whole thing in an airtight container so that no information goes out, because they need to be very mindful of their data.
Have you started to see DeepSeek R1 taking off in your metrics?
I’m getting a lot of mixed signals. Yes, there are people using it, even in highly regulated industries. What I think would be a hard “no” is that I’m not sure it’s really going to take off in the United States specifically.
I’m grateful to R1 for the idea that open source can go even further than most people thought. That has inspired a bunch of people to try new things.
I think we’re going to start seeing a new influx of models, especially reasoning models, coming out of the open-source community. I’m very curious to see how that plays out.
The distillations we’re seeing now, and the cross-breeding of R1 with other models, are also very interesting.
How would you describe the boundary today of what agents can do and what they can’t do?
There’s more that agents can do than most people would think, but it’s not enough for people to worry about losing their jobs. I don’t think we’re close to that just yet.
If you take one step back, the value that companies get from these agents is basically a math formula in my mind: how complex is what the agents are trying to do, multiplied by how much autonomy they have in doing it.
If you can automate your most important process in the company and the agents can do it completely autonomously, that creates a lot of value. That’s why you see a lot of people aiming for code development, because that’s a very crucial process in a company. If agents can do it autonomously, that adds a lot of value.
I don’t think we’re in the top quartile yet, with the most amazing processes operating without any hand-holding. I also don’t think people are ready to have no hand-holding in the better half of their processes.
What we’re seeing is agents working with humans, and that seems to be taking off and getting more adoption. Where you previously had 4 people doing something, you now have 1 person overseeing agents, and the other 3 people have been reallocated to do more interesting work because the agents are handling the busywork.
Right now, we’re going to see a lot of efficiency gains from repetitive tasks that we’re going to automate. I think we’re going to see an influx of agents into decision-making over the next couple of years. That’s where I don’t think we are just yet, and I think it’s going to get very interesting once we get there.
12. Where AI agents still struggle and what’s missing today
What problems are hard for agents? If you dropped down on planet Earth right now, you’d look at this and think, “Wow, these agents can do Putnam-level math problems. They can pass the LSAT.” If I had an employee who could do that, I’d think, “Wow, you could probably do a lot of things around here.” Where do you see them break down?
We’re seeing these models, especially reasoning models, really excel at math problems because they have a hard right answer. They can do reinforcement learning based on that, which is why you’re seeing a lot of these models succeed in that area.
The real world, especially in business, is much more nuanced. You can say, “That was the right choice,” but at the point in time when you’re making a decision, there’s a lot more nuance. There’s an element of deciding what you believe is going to be best.
These agents are getting very good, due to reinforcement-learning techniques, at answering problems with hard yes-or-no answers or specific numerical answers. I don’t think they’re as well-versed in nuance just yet.
You can use some of that capability, but I think it’s harder for people to believe that a decision was exactly the right choice. That’s where we’re seeing the gap.
Even if you had a human who was amazing at solving math problems, you might assume that would translate into logical capabilities that would lead to better decision-making in nuanced situations. But that’s not how these models work.
They perform better in that situation because they have reinforcement learning that optimizes them for it. That doesn’t necessarily mean they’re better in nuanced situations just yet.
Do you see that changing with o1 and o3? Do those feel like big improvements in that regard, or not yet?
Not yet. They’re amazing models and definitely very capable, but especially for agents, we’re not seeing them take off that much.
At the end of the day, the same engineering first principles apply here: what’s the minimum I can get away with? That helps you optimize for speed, cost, and all those things.
A lot of people see themselves running agents on a model like OpenAI’s GPT-4o mini, for example. If that gets you what you need, you might be better off using a smaller model than a larger one that’s more expensive and a little slower.
I want to talk a little bit about your architecture. It seems like you visibly took LangChain out of your stack. Can you talk about why you did that?
Sure. In the early days, when we started CrewAI, using LangChain was a good way for us to get access to a bunch of tools from the ecosystem. Those tools helped people take some of those actions, so it made sense to use LangChain.
As we started to grow and build more and more logic into the agents, it became very hard to keep things in sync. Some of the decisions we were making in the framework started to diverge from how they were making decisions and changing code on their side.
It got to the point where we were overwriting 90% of the code we were importing. It was one specific class that we were importing, but we were overriding so much of it that every time we needed to update a version, we had to do crazy rebases.
I thought, “This isn’t going to work. We need to cut this off.” It wasn’t working.
I think it was the best decision we ever made, because it allowed us to grow into other things that we had previously been constrained from doing. We had to override certain methods in certain ways, but now we can do much more.
That’s one of the more technical reasons we decided to go that route. There were also commercial reasons, including some customers who were having issues with that specific dependency.
Are there integrations that are important to you that you’re definitely going to keep, or integrations that you want to add over time?
The overall tools ecosystem is something where the more you connect to all these other things, the better.
You can still use any LangChain tool with CrewAI, but we don’t depend on LangChain anymore. You can also use any LlamaIndex tool with CrewAI, and we don’t depend on LlamaIndex either.
You also have CrewAI tools that you can use directly. There are amazing companies like Composio, which has more than 300 tools that you can use with CrewAI.
We’re finding a lot of value in these dependencies because they help us avoid rebuilding things over and over. You build something once, and it can be used across the ecosystem.
That’s something we’re not going to change anytime soon. If anything, I want to go deeper on some of those integrations.
What AI tools do you personally use day to day?
A lot of people are having this Windsurf-versus-Cursor conversation right now. I’m a Cursor guy.
Me too.
It works so well for me. I tried Windsurf, and I think they’re doing amazing work, but I was using Cursor so much that I got the hang of it. I know how that thing ticks.
Moving to something else would just make my life a little harder. I’m not ready for that.
I use a lot of ChatGPT. Anthropic’s Sonnet models have also been in and out, at least in the user interface. I’ve been using them a lot on the coding side of things, but outside of coding, I don’t use the model interfaces that much anymore.
So you use Cursor with Sonnet?
I use Cursor with Sonnet, and I’m using o3-mini now as well. I do find a few use cases where I like it, but Sonnet is just—I don’t know what black magic they’re doing over there, but that thing works.
Are there any lesser-known Cursor features that have made your life easier lately?
I use a lot of the features that people use, but one thing I started using that I think most people don’t know about is custom Cursor rules. You can add custom rules per project.
You can create a file called “.cursorrules” and put strings in there to define custom rules that Cursor preloads for a particular project. For example, if you’re doing a Python project, you can add a bunch of Python rules there, and they’ll be embedded into your project.
That’s something I didn’t know about. Now I use it, and everyone on the team does too. I really like it.
Any other tools that you love right now?
Let me think about tools that I love in AI. I tried Perplexity for a while, but I wasn’t a huge fan. I’m now starting to use more of the deep-research tools, and I’m finding them pretty good. I think OpenAI did good work there.
I also use some tools for video editing. There’s a tool called Descript. It’s an amazing video editor—honestly, the best video editor ever.
I’m just trying to produce content. I want to make sure things look good, and I want it to be easy to cut and add captions. Descript is chef’s kiss.
How do you respond to or think about concerns that AI will replace human jobs?
When people ask me about that, what I tell them is that we like to assume we have more control over our futures than we actually do. Honestly, you don’t know what’s going to happen tomorrow. No one knows. Everyone is going through this for the first time.
You have a couple of options. You can choose to focus on the things that are in your control, or you can worry about the things that aren’t. If something isn’t in your control, then what is?
In my mind, the choice that people who are worried about this are going to make is, “Maybe I should learn how to use these tools.” That might not be enough, or it might be, but at the end of the day, I would rather be someone who knows how to use them than someone who doesn’t.
Statistically, that would put you in a better position. I try to reframe the conversation that way: maybe this is how we should be thinking about it.
13. Will AI agent building become completely no-code?
Do you think technical skills will help with building agents, or do you think that, over time, agent building will completely move to a no-code mode?
Long term, no code is going to become more and more popular and probably get a lot of the market share. That would be my take.
It’s not that people are going to get lazy. There are just many more people who don’t know how to code than people who do. Eventually, if you want this to run in every company out there, you’ll need to support that very well.
I do think coding will still allow more customizability, and that’s what will unlock most of the high-impact use cases. It’s similar to coding in general. Can you build a website with no code? Yes, absolutely. Are there amazing templates that can make it look super good? Yes, there are.
But if you want to do something highly customized, do you need an engineer? Yes, you do.
When you think about your product, do you think it will ultimately be aimed more at a no-code audience?
We already have no-code tools that allow people to do that, and I think we’re going to keep investing in them. There’s a broader audience we’re trying to serve that goes beyond regular engineers.
I also find myself using CrewAI Studio a lot because it makes me faster, similar to Cursor and other tools. We’re probably going to keep investing in no code as well.
What applications have worked well for you inside your own company?
In my company, the thing we use the most is CrewAI Studio. We took a very different approach from most no-code tools out there.
I have a point of view that I don’t love node-based UIs. I know that’s what a lot of people use for no-code tools. I find that they make for beautiful screenshots, but if you’re dropped into a page with 30 nodes, it’s impossible for someone without technical knowledge to understand what’s going on with ease.
What we decided to do to offset that was: you can go into app.crewai.com and create a free account and test it out. We have what I call a gradually evolving UI. You start with a chat interface similar to custom GPTs back in the day, so you start chatting your way into what you're trying to build. That is kind of like building that mental model in your mind.
Once you're ready, there's a button that pops up that you can click. It's kind of like, “All right, generate me a crew plan.” Once you click on that, the UI evolves into a table view where you can see our agents and tasks. You can still change them, but now it's more structured; it's not just plain text, and you can still keep chatting if you want to.
Once you get there, you have a new button to generate the crew, and that takes you into a node view. By the time you get to the node view, you have a full understanding of how you got there and what you're trying to build, so it's easier for you to customize it that way. That's kind of our approach to it. And then, from there, you can download the code if you want to.
But for you personally, inside your company, what's working? What are you actually using Crews to accomplish?
We're doing a lot of things. We're generating marketing content automatically. We're doing a lot of meeting prep, with agents that are automatically researching people, companies, and everything along the way. Anything related to pull request reviews—we have Crews reviewing every single pull request across the company, open source and closed source—and that has been pretty good.
We're also using agents to do a bunch of support, writing responses and all that. We have agents that are picking up recordings and transcripts of meetings and following up with next actions, presentations, and such. That is also pretty good.
One that I love is the one that happens with onboarding. As these customers onboard to the platform and provide their names and emails, we kick off agents that research them. But they don't stop there. Based on the research, these agents infer what they might use agents for, and then they take those ideas and push them into 2 different places.
One is our CRM, so now every marketing engagement that they get is super custom, based on why we believe they might want to build something. We also inject that into the product itself, so the product self-customizes based on our agentic hypothesis of what agents they will be using. Then it just works, and they're left wondering, “How does this app know that I'm trying to go that way?” That's kind of how the sausage is made.
14. How Crew AI uses agents internally for marketing, development, and automation
Interesting. Do you have any predictions for 2025? Any new applications that you think will become possible with better LLMs or better tools?
15. Joe Mora’s predictions for AI in 2025
I think we're going to see some more fine-tuning coming back. People have not been talking much about it, but I think you're going to start to see some fine-tuning coming back, especially because smaller models are going to start trending again. Some of the things coming out of DeepSeek-R1 are like, “Hey, can we reuse this to make small models that are amazing, and what would that look like?”
I think that's going to be a trend this year, where we're going to have smaller models that are very, very capable—better than 70-billion-parameter models nowadays—and we're going to be able to fine-tune those with ease to use them to run agents.
All right. Well, João, thank you so much. This has been a real pleasure to talk to you.
Thank you so much for having me. It was a lot of fun. I really appreciate it, Lukas. This was great.
Awesome.