No Priors 第142期|Harvey 联合创始人兼总裁 Gabe Pereyra
Harvey 的下注,已经从让单个律师更快,扩展到让整个律所更具盈利能力。成立短短3年半多,Harvey 已拥有近1,000家客户和500名员工。产品也从“律师的 IDE”演进为服务“编排、治理”和数千个客户事项的基础设施。
企业级机会正从律所延伸至法律服务的最大买方。Harvey 最近宣布 Walmart 成为客户,同时正与 AT&T、财富500强公司、私募股权机构和 Global 2000 企业合作,服务内部合同管理与法律运营,并提供与外部法律顾问安全共享数据和协同工作的“协作组织”。
法律智能体拥有天然的运行环境——客户事项,但奖励函数仍是开放问题。律师助理本来就以智能体的方式工作:研究判例、总结内容、起草备忘录、添加引证,再吸收合伙人的反馈。但查找控制权变更条款可以验证,最终判断一份并购协议是否足够好,却仍取决于资深合伙人说一句:“嗯,这看起来相当不错。”
AI 未来可能压缩律师助理的杠杆比率,但短期内不会大幅改变资深合伙人的角色。Elad Gil 提出的挑战是:如果律所只需要20名或50名律师助理,而不是100名,就可能无法培养出足够多能够晋升为下一代合伙人的人。Gabe Pereyra 预计基层职能会发生变化,但战略、委派、专业判断和客户互动目前仍大体稳固。
前置部署工程正在成为 Harvey 企业级服务和产品发现闭环的一部分。大型银行甚至可能没有法律文件管理系统,而 Blue Owl 等客户看到了大量可以接入生成式 AI 的流程,却需要帮助定义这些流程。反复出现的定制化工作可以沉淀进平台,律所本身也可能通过为客户部署 Harvey 获得实施收入。
Harvey 明确拒绝自建律所,而是选择搭建让每家律所都成为“AI-first”的平台。Pereyra 表示,Harvey 曾与 Atrium 的约30人以及 Sam 和 Jason 交流;核心难题在于,律所和科技公司是两种不同的业务,必须分别建设。拥有一家律所会制造利益冲突,也无法覆盖一个规模约1万亿美元的法律市场,以及规模“约3万亿至5万亿美元”的专业服务市场。
Pereyra 的非共识判断是,AI 最大的增益将来自组织重构,而不只是让个人提速。“让一个人编程快20%,并不意味着产品开发也会快20%”;下一层能力在于协调整个律所内的专业化人类与 AI,尤其是在模型持续变得远强于组织即时吸收能力的情况下。
1. Harvey 正在成为法律组织的运营层
Harvey 成立3年半多,如今已服务近1,000家客户,拥有500名员工。它最初的定位很直接——为大型律所和企业法务提供 AI,但价值创造的单元不断从单个律师扩展到整个律所及其客户。
主持人提出的基础问题——为什么不直接使用 Copilot、ChatGPT 或 Claude?——正好概括了 Harvey 的演进路径。GPT-4 早期开放使用时,直接与模型交互对高度依赖文本的法律行业极具价值,但律师很快就遇到幻觉和上下文缺失问题;因此,Harvey 花了约2年围绕这些模型打造“律师的 IDE”。
Pereyra 当前的表述更为宏观:“我们要解决的重大问题,不是如何提升单个律师的生产力。”真正的问题是,团队如何处理一个客户事项,以及律所如何在同时处理数千个事项时变得“更高效、更赚钱”——这越来越取决于编排、治理和企业级产品能力,而不只是模型本身的智能。
分发渠道也在通过律所本身拓宽。大约1年半前,律所开始向客户展示 Harvey,促成了最近宣布的 Walmart 签约,以及与 AT&T、大型私募股权机构和其他财富500强、Global 2000 企业的合作。这些客户既需要内部法律运营,也需要与外部法律顾问安全协作。
2. 法律工作流为智能体提供了天然场景
Pereyra 以设立基金为例,说明法律工作远不只是给律师发一封邮件。一只10亿美元的基金可能需要一份100页的有限合伙协议、约100名投资者以及反映不同投资者要求的附函,其中还包括税务影响;每次修改都可能产生下游影响,需要律师协调处理。
投资业务同样拥有高度密集的上下文。律师要检查公司的数据室,理解其合同,以验证收入是否按所宣称的方式构成;还要识别诉讼,并跨文件追踪义务。Pereyra 将其比作“理解一个代码库”,只不过这里的代码库由合同和法律工作组成,而在语言模型出现之前,这些工作流很难结构化。
Harvey 在接入 GPT-4 的第一天就出现了智能体直觉:Winston 花了14个小时复刻律师助理的工作——查找判例、总结内容、将总结输入起草流程,再为结果添加引证。“可以把律师助理看作智能体”:他们接收合伙人的战略,搜集证据,再交回一份备忘录。
按照 Pereyra 的强化学习类比,客户事项就是运行环境。处理基金设立、收购或诉讼的智能体,可以在文件管理系统中导航、检查数据室、研究判例并获取合伙人反馈,就像编程智能体与代码仓库交互并尝试通过测试一样。
3. 专业判断是核心,也是训练瓶颈
当输出从检索进入长篇起草,法律领域的强化学习就变得困难。找到全部控制权变更条款,可以支持传统基准测试;但生成一份并购协议,很难简单归类为“好”或“坏”。Pereyra 称,构建这一奖励函数是“真正重大的问题之一”。
他的答案是:“奖励函数就是合伙人。”律所拥有草稿、修改记录和合伙人反馈,这些都可能成为公开申报文件中不存在的训练材料。公开的 SEC 文件展示的是结果,而不是生成这一结果的决策过程。
Pereyra 认为,成熟的软件工程最终也会遇到同样的验证问题。单元测试在短期内有帮助,但生产环境中的成功可能意味着系统连续6个月服务了100万名用户且没有崩溃;一项并购的真正检验,可能要到3年后才出现——合并后的公司是否避免了意外诉讼。
Wachtell 合伙人、早期加入 Harvey 并担任顾问的 Gordon Moody,代表了这种缺失的专业能力。他曾参与 Michael Dell 将 Dell 私有化、重组业务并再次上市的过程,需要在多年期重组、史上最大规模的债务发行以及一种全新金融工具之间进行推演。这要求的是“理解如何架构这些交易的技术能力”,而不只是人脉关系。
4. AI 可能先改变律师助理金字塔,再改变合伙人
Gil 的反驳值得保留:律所传统上可能招募100名律师助理,明知最终只有约10人能够成为合伙人。如果 AI 最终把所需人选规模降至50人或20人,律所可能失去那条用于识别谁值得被托付复杂收购交易的经验漏斗,即便目前 AI 的作用仍主要是增强能力和扩大业务。
Pereyra 乐观地认为,模型可以加速培养过程。正如编程模型能让学习者翻译 Python,或询问某段代码为何这样编写,律师也可以先要求生成一份并购协议,再追问:“我们为什么要这样设计?”律所还可以把积累的合伙人反馈转化为训练数据。
组织重构无法用一套方案覆盖整个行业。诉讼、大型交易和中型交易拥有不同的工作流、人员配置和定价方式,因此 Harvey 正按业务领域逐一推进。例如,它会与基金设立团队及其私募股权客户共同工作,重新设计运营模式。
资深合伙人可能比初级职能变化更小。Pereyra 将他们比作杰出工程师:两者都负责定义战略和抽象框架、委派执行、识别细微的失效模式,并与客户沟通。“我不认为模型很快会取代他们现在做的事”,但他明确猜测基层工作会发生变化。
5. 部署工作正把组织摩擦转化为产品
Harvey 最初强调横向平台,以及 Workflow Builder 等工具,除 PwC 这类超大型客户外,很少进行客户专属开发。如今,训练律所专属模型和智能体需要在客户环境内接入文件仓库、计费系统、治理系统以及其他私有数据。
企业环境的异质性进一步放大了这项负担。一家大型银行可能会告诉 Harvey:“我们的法务部门没有任何文件管理系统。你们能不能直接帮我们建一个?”Blue Owl 同样看到了大量可能接入生成式 AI 的流程,但希望技术团队与其并肩探索这些系统究竟应该如何设计。
主持人将前置部署描述为 Oracle、Dell 或 IBM 式的经典企业服务打法:先提供平台,再围绕定制数据进行开发,并把反复出现的通用实施方案不断吸收到核心产品中。Pereyra 认为,Harvey 的做法更接近 Sierra 的智能体工程项目。
实施生态已经开始形成。律所可以向企业法务客户推荐 Harvey,帮助搭建工作流并实施部署,承担规模较小的法务部门无法自行配置的工作。这既可能创造新的收入来源,也会进一步强化 Harvey 的分发能力。
6. Harvey 要赋能每家律所,而不是拥有一家律所
Pereyra 表示,Harvey 曾与 Atrium 的约30人以及 Sam 和 Jason 交流。Atrium 参与者对这一想法持积极态度,但难题在于,Harvey 将同时建设一家律所和一家科技公司:“我认为一件事只能做好一件。”
更大的战略目标,是帮助“每家律所都成为 AI-first 律所”,在提升盈利能力的同时,为客户提供更快、更便宜的服务。拥有一家律所会制造利益冲突,也会阻碍模式扩张。
一项大型跨国收购展示了平台机会:可能有约100家外部法律顾问律所参与,因为当地问题需要专业机构处理。Pereyra 提到,Harvey “在新西兰大概有40家客户”,而这类交易还可能涉及投行、PwC、税务顾问和人力资源咨询公司。Harvey 瞄准的是连接这些参与方的安全数据共享和 AI 基础设施,覆盖规模约1万亿美元的法律市场,以及“约3万亿至5万亿美元”的专业服务市场。
7. 下一轮生产力突破属于组织,而非个人
Harvey 在 GPT-4 发布前就已启动,当时并排对比显示,其产品在 GPT-4 上运行正常,但在 GPT-3.5 上无法工作。Pereyra 表示,此前10年里他一直在尝试创办类似 Harvey 的公司,自己只是较早入场,并非一夜成功。他在 AI 领域的经历让规模化路径变得清晰:一旦研究人员找到通用方法,“通常只要继续扩展规模,这些东西就会持续有效”。
Winston 提供了法律洞察,Pereyra 则提供了对能力边界的判断。他们最关键的产品选择,是保持足够开放,以支持“任何类型的法律工作”,而不是去优化 GPT-3.5 时代的某个狭窄任务。这相当于打造一个支持任何编程语言的助手,而不是只检查 Python 漏洞的工具。
Pereyra 认为,法律行业很早就找到了 AI 的产品形态:上传文件,对其进行处理,再返回高度准确的引证。他认为,编程领域当时需要更强的基础模型和更好的 IDE 集成,这或许解释了为什么编程产品后来才出现,尽管 GitHub Copilot 已经证明了市场需求。
Pereyra 的前瞻判断超越了 Copilot:“让一个人编程快20%,并不意味着产品开发也会快20%。”他说,律所规模相较于计算机和互联网出现之前已经扩大到原来的10倍,并认为通过组织专业化的人类与模型,这一过程可能再次发生,而不是寄希望于某个足够聪明的 AI 独自完成一切。
Gabe, thanks for doing this.
Of course.
Yeah, thanks for coming. Maybe we can just start with, for anyone who hasn’t heard of Harvey, what is the company? Can you talk about the scale and who you serve today?
At Harvey, we’re building AI for law firms and large in-house teams. We’re almost at 1,000 customers and 500 employees. We started just over 3½ years ago, and we’ve been scaling quickly since then. You guys were some of our OG seed investors, so it’s good to be here.
From the most basic perspective on the product, why is it not just Copilot, ChatGPT, or Claude?
That’s how the product started. When we first raised from OpenAI, we got access to GPT-4. I think the jump from GPT-3 to GPT-4 was so significant that the intuition at the time was just to give the model to lawyers and have them play with it. The industry was so text-heavy that you got a lot of value from simply interacting with the models.
As soon as you gave it to lawyers, you also ran into all the sharp edges of the models: they hallucinate, and they’re not connected to a lot of the context you need. I would say the first 2 years of the company were about how to build, essentially, the IDE for lawyers around these models—connecting them to all the context you need to be productive as an individual lawyer.
But in the past year and going forward, the big problem we’re solving is not how to make individual lawyers more productive. It’s how to make a team of lawyers working on a client matter more productive and, more importantly, how to make an entire law firm working on thousands of these client matters more productive and more profitable. When you get to that scale, a lot of the problems we’re solving aren’t just model-intelligence problems. They’re orchestration, governance, and all the enterprise-product problems you run into at scale.
You’ve also been broadening from just law firms into enterprises—big companies using you in concert with both their in-house legal teams and external counsel. Can you talk more about that and how it’s been evolving?
We started selling to the largest law firms. Something that started happening about a year and a half ago was that these law firms started showing Harvey to their clients, and their clients both wanted to collaborate more effectively with their law firms and wanted to use it directly in their in-house departments.
We recently announced that we signed Walmart. We’re working with AT&T, a bunch of Fortune 500 companies, large private equity firms, and Global 2000 companies—the largest consumers of legal services. What we’re starting to build is a platform for in-house teams to do the work they do internally, things like contracting and the long tail of legal operations that you typically don’t send out to law firms.
There’s also the collaborative tissue: “I’m working on a large transaction or a litigation. I need outside expertise. I want to securely share this data with my law firm.” There are a lot of technical problems around security and data privacy that we want to solve so these law firms and their clients can collaborate effectively.
I think we have a largely technical audience, but most people don’t know exactly what legal workflow looks like. Before we really started working together, I imagined it as: I email my lawyer, he thinks about it, reads a document, and sends something back, right? There’s redlining involved somewhere, and maybe there’s negotiation. Can you paint a picture of what workflow means to you guys?
A lot of people, when they think about legal, think about consumer legal. They have a lease and need to get input on it. That’s completely different from what these massive law firms are doing.
A really good example of what these firms are doing—something you guys will be familiar with, and I think a lot of people in the startup space will be familiar with—is what law firms do for venture capital firms or private equity firms. VCs and PE firms do 2 main things: they raise money and they invest it. That process is actually important, but there’s less legal work there. The important things you need to do are fund formation.
Mhm.
How do I structure the entity that’s going to hold all that money? It sounds easy, but if you’re a large private equity firm, a sovereign wealth fund comes in and says, “We need to structure it in this way because of tax implications.” Then you have a pension fund that has other requirements.
It ends up being this incredibly complex process. How do you draft the limited partnership agreement, which can be 100 pages? You can have 100 investors, and they all have side letters that modify the agreement. You need to understand that if you modify it this way, it’s going to have these implications. A lot of it is also the project management that goes into coordinating all of these projects if you’re raising a $1 billion fund.
Once you’ve created that fund, there are all the investments you make out of it. For example, when we did any of our Series rounds, you need to get a data room. We share a bunch of data, and you look at that. You need to understand the contracts we have to make sure that the revenue we say we have is actually structured in the way we’ve claimed. Is there litigation? All these things.
It’s a massively complex process of understanding. One analogy is understanding a codebase, but the codebase is all of these contracts and all this legal work. I think the reason legal is so difficult is that the workflows aren’t structured. In programming, it was really hard before these models to build tools for programmers. You basically just had an IDE, and then programmers did things in all the different languages. You didn’t have, “Here’s a tool for Python; here’s a tool for C++.”
Legal is kind of the same way. A lot of why you’re seeing traction in programming and legal is that there are a lot of analogies in these workflows. They’re so text-heavy, and until you had these models, you couldn’t structure them in the way that I think you can now.
One of the directions people are going in on the coding side is to build things that are being called agentic. It’s very early in terms of what “agentic” means, but basically, it’s being able to deconstruct a logic tree in terms of the set of actions you need to take in a certain situation, and then having the AI agent go back and check each item, do it, go on to the next item, and double-check it against the prior one. Do you do that from a legal perspective, or is that a little further in the future relative to where code is today?
We’re starting to do this now. When I was at DeepMind, a lot of the reinforcement-learning research I did was on that. When we first got access to GPT-4, we had the very strong intuition that you were going to be able to string together a bunch of these model calls or eventually do things like reasoning models, where the full agent is differentiable.
Even the first day we got access to GPT-4, Winston went into his room for 14 hours and just redid a bunch of his associate tasks. When I looked at the work he was doing, it was essentially this hacky agentic workflow. He said, “I would need to go look up this case law, summarize it, take that summary, and use it to draft.”
Seeing him do that gave us the intuition very early on that this is the direction things are going. You can think of associates as agents. They get a task from a partner: “Hey, I have this high-level case strategy. I want to see if I can find a bunch of case law that supports it. Can you go research that, look it up, cite it, and write me a memo?”
A lot of the systems we’re starting to build look like that. One interesting direction the coding labs and research labs are going in is building these reinforcement-learning environments where you deploy agents, they interact with a codebase, and you see if they can pass unit tests.
In legal, that reinforcement-learning environment is a client matter. You have all the context of a fund formation, an acquisition, or litigation, and the models are starting to learn: “Let me go into the document-management system and see if I can find this. Let me go into the data room or do case-law research and get feedback from the partner.” I think that research direction is super interesting.
It’s really interesting that you make the associate analogy, because I remember when I led your Series B, which I think was maybe 2 years ago now—it was a while ago. I called a lot of your big customers and talked to the head of the law firm or the head of some of these institutions.
One thing I thought was really striking was, number 1, that they were adopting legal software, which had previously been really hard to sell into them. Because what you were doing was so striking and important, they were adopting you really quickly.
The second thing is that they weren’t threatened by it. I thought, “They’d be threatened because it may augment or eventually replace certain aspects of law, or help change that dramatically.” One insight they kept bringing up was really interesting. They said, “As we think ahead, as this sort of AI tooling and agentic workflow spreads through Harvey and companies like you, how do you think about the future of a law firm?”
Instead of hiring 100 associates, of whom you assume 10 will eventually become partners, maybe you only need 50. Maybe you only need 20. Are you even hiring enough people to know who would be a great partner? You’re going to shrink the set of people who are needed to do certain tasks over time, right? Right now, that isn’t true—it’s augmentation, and it’s expanding the business—but that could happen in the long run.
How do you think about the future of law, what law firms will look like, and the evolution of all that?
Yeah, this is a great question. I think it’s changed a lot in the past couple of years. Something we’re starting to talk with law firms about a lot is: How do we think about training the future generation of partners? To your point, these law firms have leverage ratios where you have a lot of associates but much fewer partners, and there is value to that because not everyone is going to become a partner. Part of going through that process is finding the person you would trust to do a very complex acquisition because they’ve gone through that experience.
I think the part I’m optimistic about is that, if I think back to over 10 years ago, when I learned to program, it was super painful. You had to go on Stack Overflow, and it was hard to learn multiple languages because you were like, “Okay, I’m just going to learn Python. I’m going to learn TensorFlow.” It was hard to even learn that, and it was hard to ask questions. When I was at Google, you didn’t want to ask a bunch of questions because people would be like, “Oh, you don’t know that?”
Stuck all the time.
Yeah, exactly. Now, with the models, programming is so fun to learn because you can just be like, “Here’s how to write this in Python. Translate it. Why is it written this way?” You can learn this so much more quickly.
We see lawyers doing that with Harvey, where they’ll say, “Generate this merger agreement. Why did we structure it that way?” We’re already starting to see some of that. But I think the really big opportunity for law firms is: How do they take all of the internal partner feedback and data that they’ve created and use that to start training? I think that’s one big piece.
I think another conversation we’re having is, to your point, how do you generally start restructuring firms? This is one where we have some intuitions, but a lot of it is going to depend on the firm, the region, the size, their specialty, and the types of clients they serve. One of the things that’s very challenging with law firms is that they are really a collection of all these practice areas. The firms that specialize in litigation look different from the firms that specialize in large transactions versus midsize transactions.
Usually, the big firms do a collection of these. A lot of what we’re spending time on is practice area by practice area. Can we go and sit with the fund-formation group and their private-equity clients and start thinking about what that would look like in terms of the workflows, the staffing, and the pricing?
Mhm.
I think it is a really interesting problem where a lot of the value in the product and the platform is not just the product itself, but how we help enable these firms to transform. When you think about it from that perspective, our goal is: How do we make these law firms more profitable? It’s not just a product problem. It’s thinking about their holistic business and where we fit in that bigger picture.
Mhm.
Yeah. It’s really interesting because when you look at the set of functions that a partner fills—and I’m thinking in particular of consulting firms, less about law firms, simply because I’m a little bit more familiar with consultancies—some of it is pattern recognition, high-level thinking, and strategy, and then part of it is sales.
Yeah.
Really being able to make that client connection. To your point, it’s interesting to think more broadly about how AI can augment all parts of their business versus just the legal workflows.
Yeah. And to your point, I don’t think that part changes. We’re now larger consumers of legal services, and when we think of the best partners we’ve worked with, I don’t think the models are doing—
Yeah.
—what they do anytime soon. I think what’s interesting is that the role of law-firm partners actually doesn’t change that much, in the same way that I don’t think the role of very senior engineers changes with this. You’re largely delegating work, and what you’re getting paid to do is: Here’s the high-level strategy, here are the right abstractions, go write the code or do the legal research to help me do it, and I will interface with the client.
My guess is that doesn’t change too much, but some of the lower-level functions do change because of this technology.
One of the things that you said in another conversation we were having was that there’s an analogy you could make between a great senior partner, like a Gordon Moody type, and a distinguished engineer working on systems at Google. I think for a more technical or general business audience that doesn’t really know what Gordon Davidson does, they might assume what Elad said, which is, “Isn’t 50% of that his network or his reputation?”
What you were pointing out is that there’s expertise in the ability to predict a sequence of arguments that is going to get you to the answer you want or manage risk. How does that translate to an RL environment or a task for you?
Yeah, this is a good question. For background context for the audience, Gordon Moody was a partner at Wachtell, which is one of the top transactional firms in the world. He joined us early on and is now an adviser.
The analogy I was giving is: Why is a senior, distinguished distributed-systems engineer at Google so valuable? A lot of it is the experience they have architecting these systems. None of this is public, so it won’t go into the models for a long time. If you’re building search at Google, these people can just point out, “Hey, if you build this system this way at this scale, it’s going to collapse for some reason that’s super unintuitive.”
One of the examples that Gordon talked about early on was that he was part of the process when Michael Dell took Dell private, then restructured it and took it public again. This was a multiyear, super-complex financial and legal restructuring of an incredibly large business. When you talk with him, what he’s incredibly good at is the same thing you see when you talk with a very senior engineer: He has the whole picture of this legal entity in his head.
At the time, they had to do the largest debt offering of all time. They had to invent a new financial instrument. It’s just understanding that, if I need to raise this much money to do this part of the transaction, this is how I would structure it. A lot of the value he brings is not just the relationship; it’s the technical understanding of how you architect these things, in the same way that you architect very large software projects.
I think when that translates to an RL environment, part of what is missing from the public models is the process of looking at one of these entities and figuring out, given all of the context, “I want to do this merger. This is the right way to structure it.” Just that process—and a lot—
It’s a reasoning trace, right, for an expert, just like it would be in code.
Yeah. If you looked at that dataset for one of those transactions, it would be: The client comes to Gordon and says, “I want to do this large merger or acquisition.” Then there would be meetings and emails talking about, “Okay, this is the background of the 2 companies. This is roughly how we would structure them. These are all the things we need to look into.”
A lot of the data would be Gordon giving tasks to associates, saying, “Okay, look into these risk factors of similar transactions we’ve done.” They would do research and say, “Okay, maybe we could structure it this way.” Then he would point out this really subtle thing: “Hey, actually, in this case, if you structured it this way, this thing’s going to happen.”
But none of that shows up. All you get from these public mergers is an SEC filing. You see the final result, but most of the value—or what you need, I think, to eventually improve these models—is the decision-making process, in the same way that you need these reasoning traces to train these models to do any of these reasoning tasks.
One of the things you mentioned is that the labs are all very focused on RL scaling in coding and math domains. I think of those as highly verifiable—not perfectly so—but how do you think about the appropriateness of law for RL, given that it’s not as easily verifiable?
Yeah, this is one of the biggest problems. I remember we had conversations early on when we were trying to figure out what the right evaluation structure was. I think the hardest thing about legal work is that most of these tasks are very long-form text generation.
There are definitely subsets of legal work that are super verifiable, such as going into a data room and finding all the change-of-control provisions. You can build these traditional datasets. But for something like “Generate this merger agreement”—
It’s really hard to just give some binary result: This is good or this is bad.
And I think this has been a big research problem with all the labs we work with and internally. There’s just this open question of how you build that reward function. If you think about what that reward function is at the law firms, it’s the partners. At the end of the day, there’s no way to verify this besides having a senior partner who’s done a bunch of these say, “Yeah, this looks pretty good.”
Internally, these law firms have a bunch of data: all the edits that went into a document and the feedback. We’re starting to think about how to use that to train these reward functions. I would say that is one of the really big problems, but I think one of the interesting things is that you actually have the same problem in programming.
In the short term, programming is verifiable because you can look at unit tests, but once you get into real software engineering, like the unit—
There is no unit test. It's like I deployed a system design.
Yeah. It's like I deployed this and a million users used it for 6 months and it didn't crash. Mergers are the same: you can make sure the filing is correct, but at the end of the day, 3 years later, the companies are still merged and they didn't take on litigation they didn't expect or something like that. That is eventually the really valuable human experience, right?
That's what you pay really good software engineers or really good lawyers for: they have that decade-long track record of building these systems, and they haven't fallen apart. A lot of this stuff is the same way—you can't unit-test it; it's hard to verify. So it is, I think, this really interesting open research problem.
One thing that you guys are doing on sort of the other end from pushing the bounds of what Harvey products and the models can do is just getting them deployed, and you recently started this deployed engineering force. This is confusing to me because I'm like, you're not necessarily an application-building company, which is how people have traditionally thought of FDE. Why are you doing this?
I would say this is closer to Sierra's agent engineering program. But what we're starting to run into a lot is that, early on, I think we did a really good job of building a horizontal platform. We didn't do that much customization for customers, in the sense of building specific things for specific customers. The nice thing about legal was that we could build things like Workflow Builder into the product that would let customers customize the product. For very large customers like PwC, we did some customization.
But now we're getting to the point where, when we're starting to talk with law firms about, "Hey, we want to take a bunch of this data and help you build a model or build agents," there is some amount of, "We need to go into your environment and figure out how to connect all the data." We're starting to connect to a lot of their business systems—their billing systems, governance systems, and so on.
Especially when we start working with the Walmarts, the very large banks, and the Fortune 500, they're much less standardized than these law firms. There is just this massive amount of work where we go to a large bank and they say, "We don't have any document-management system for our legal department. Can you just build us one?"
Mhm.
There is a massive amount of demand from people who just want smart technical people to sit here and help them think about their business and their operations, and how they should start mapping that into GenAI systems. For us, it's a really good way to figure out the roadmap.
For example, Blue Owl is one of the fastest-growing private equity firms that we recently started working with, and we meet with them all the time. They're just like, "There are all these things that we feel like we could map into GenAI. We don't quite know what it's going to look like, but let's just sit together and figure it out."
I would say that's a lot of the genesis of the program: how do we get more people who can work with all these customers and start paving the way for some of these new roadmaps in different verticals?
Yeah. I think what you're describing, too, is a very standard enterprise playbook. In Silicon Valley, people almost forgot because of the SaaS era that if you're Oracle, Dell, IBM, or any of these larger organizations, this is how you sell software.
Exactly.
Right? You have a platform, you have a bunch of customization around it, and people have bespoke data sets.
Right?
This is the standard way to do it. As you do it over and over again, you start repeatedly turning that into part of the platform. A lot of these started with doing something that resembled FDE, and then you get big enough that you get this implementation ecosystem. There are all these third parties that will come in and implement.
It'll be like the certified vendor. I think the interesting thing we're actually starting to see is that law firms are starting to do this for their in-house clients. They're starting to go and take Harvey to their clients and say, "Hey, buy Harvey and we'll help you build all the workflows and implement it," because they have the scale and the expertise to build this, whereas typically these in-house teams—the smaller ones—don't have the budget or the in-house team to build this.
That could be a good revenue driver for the law firms that you work with, in terms of a new line of business that they can offer.
Yeah.
I was really struck by—I don't know if it was Day 0; you can correct me—but it was within the 1st year where the very 1st version of Harvey was really an individual lawyer productivity tool, right? I'm an associate or a more senior person at a law firm. I want to get a piece of work done. Can you just make it less painful?
But the transition quickly to, "How do we transform the business, make the business more profitable, with organized teams in the ecosystem?" I think happened pretty quickly. Anything that is a business transformation just requires a lot of engagement.
Yeah.
Given how much you guys have invested in customer success and how that's driven adoption, I feel like a big piece of it is just how quickly AI has happened.
Yeah.
I would not necessarily have predicted that all the customers you're working with would be like, "Yes, in years 1 and 2 of this company selling, we're adopting." But part of it is you guys are helping them, right?
Yeah. No, and I think this was still surprising. When I look back, it was surprising how quickly some of these law firms adopted this. Our 1st customer we actually met through you, and he introduced us to David Wakeling at A&O. That was in our 1st year, and they went from a small pilot to firm-wide and investing in this.
I think you're seeing this in a couple of verticals with Cursor and OpenEvidence, where this technology is so transformative for industries that are so tech-heavy and knowledge-based. They just haven't had tools like this. Early on, we did find these customers that were like, "Oh, this is worth really betting on," but I think the pace has still been pretty surprising.
I asked the internet through X what questions we should ask you, and a popular one was, "Why aren't you guys building a law firm? Are you going to build a law firm and compete with all your customers?"
Yeah. No, we get this question. When we first started Harvey and were doing research, we actually talked to 30 people from Atrium. I think, interestingly, we also talked to Sam and Jason. Jason was the GC of OpenAI at the time and was the GC at Y Combinator when they did the Atrium investment.
What struck us was that the people who worked there said it was a really good idea, and they were super excited about the prospects. Then there were some challenges around the legal and the execution. But when we dug more into it, the big challenge that they ran into was that you're essentially just building 2 different companies, right? You're building a law firm and you're building a tech company.
It's already really hard to build product and engineering, do AI, and scale sales. I think the big issue you run into if you try to do both of these is that you can only do 1 thing well, and doing a law firm well is very different from building a software company well. I think that's 1 point.
The bigger point is that, for us, it feels like the best outcome is if we can figure out how to help every law firm become an AI-first law firm—not how to build 1 ourselves. The real problem we're trying to solve is: can we make every law firm more profitable? A part of that is how they work with their clients. Can you make their clients get better, faster, cheaper legal services?
I think solving that equation at scale is a much bigger opportunity than building a single law firm, because you get conflicted out. You can't scale this. So I think this is something we don't do. We've gotten this question, but I think it's not the focus for the company.
Analogous to other markets in software, law feels like an area where I've been very surprised personally about how large the scope of the problem is if you're really ambitious about what you can do. I didn't realize—you were telling me that if you do a really large M&A, let's say of 2 global companies, Microsoft and Activision or something, there are 100 outside counsel firms here.
You know why? Because in New Zealand, where both companies have customers, you have a tax implication, and the person who understands that lives in New Zealand, right?
Yeah. Oh, it's crazy.
And so I think, like other markets, the SMB version of this looks really different from the high-end enterprise version of this.
Yeah. And so I do think it’s hard. It just seems hard to imagine coalescing all of that expertise in a law firm and a software company at the same time, versus, well, Harvey now has, what, 40 customers in New Zealand.
Exactly. And if you think about those transactions, it’s also not just law firms. There are investment banks, and you maybe have PwC or a tax advisor, and there could be an HR consultancy that helps you think about how you’re merging headcount.
So for us, the bigger opportunity seems to be: How do we build the platform that lets professional service providers and their clients collaborate? I think a lot of the problems you need to solve there are—the biggest is secure collaboration across many of these entities, secure data sharing. How do you build and deploy AI systems across these very complex projects?
And I think, to your point, the scope of this—legal is $1 trillion, professional services is something like $3 trillion to $5 trillion. There’s just this massive amount of room to grow. We think our expertise is going to be in building the product, the technical systems, the AI systems that enable that. We want to give that infrastructure to all of the different law firms rather than compete with them, because I just don’t think you can.
I think one of the things that’s really striking about this sort of wave or era of AI is that there are deeply technical people building giant companies in really different industries.
Yeah.
And you come from a research background. You worked at one of the major labs in terms of foundation models and other areas—RL environments, reinforcement learning. What has been your biggest surprise in terms of transitioning into being a founder and running a company and building something from the ground up like that?
I think maybe—not a surprise, but the biggest mental model shift—is that I think the 10 years before Harvey, I was doing a mix of mainly AI research and trying to start companies, but always largely as an IC.
Mhm.
And I think the shift from when this started working and scaling—just how much I had to change my mental model of the type of company we’re building, how you do this at scale, how you operate—I think that was the biggest surprise, or thing that I’ve had to change.
But it’s been a crazy experience, going from Winston and me in an Airbnb to 500 people in about 3.5 years. Then I think also how you build these products at scale and the complexity of this industry—that has been a really hard but interesting experience.
It’s been amazing to see what you all have accomplished. It’s such a short period of time.
I was thinking back to when we pitched to both of you guys 3.5 years ago, and we were like, “Hey, AI plus legal,” and you guys were like, “Sounds good.”
I think a really important aspect of that, too, is you all started this company before GPT-4 came out and before a lot of the shifts in the models happened. I remember you showing side by side GPT-3.5 versus GPT-4, and what you were doing worked on 4 but not on 3.5.
You were part of that very early wave that had conviction this was so important as a trend. Was that because of your experience in the labs? Was it something else? What drove you? Not many people were actually starting AI companies when you all got started. It was, to your point, AI plus legal—nobody was doing that.
Yeah. Yeah. It’s something where now everyone’s like, “Oh, this is such an obvious idea.”
Yeah. Now it’s text in, text out.
But at the time, no one was thinking about this. I think it was a combination of a couple of things. A lot of the best people I had worked with at the time had gone to OpenAI, and so I was working on large language models at Meta. You saw GPT-1, GPT-2, GPT-3.
If you were working in AI for the past 10 years, one of the big problems was: How do you pull all this together? You built systems where, okay, this is really good at vision, this is really good at specific things, but no one really had the general solution. And you saw things like LaMDA.
I think with that trend, what I’d seen is, anytime you make that initial—okay, this is how you do it—you can usually just scale, and this stuff keeps working. With GPT-3.5, you were like, “This is getting really interesting, but it’s not quite there.”
The bet was, okay, OpenAI may be one of the people to crack it. I know a lot of the people there.
That was part of it, and then I think the other big part was just Winston was a lawyer, and I think we had become super close. I never thought we’d start a company together, but just the way I heard him talk about the legal industry—he, even though he was a first-year associate, just had this intuition not just of the work he was doing, but the structure of the firm.
I would hear him talk about the firm and be like, “Here’s what all the different partners are doing. Here’s why our firm strategy is this way.” He was in the process of convincing some partners to leave to start a law firm with him, which is insane for a first-year associate. And so it was just like, okay, this will be really fun.
He showed me a bunch of his legal tech. It seemed like the perfect application. Then when we saw GPT-4, I was just like, “Oh, the time is now. This is the perfect application.”
I think it’s really noteworthy that even 6 months into working with you guys, I was saying, “Will our capabilities really advance that quickly?” And both you and Winston were like, “Absolutely. We should have the ambition to take on the full complexity of any type of legal work that’s possible, because the models will keep getting better.”
That seems like a super-obvious mainstream point of view today, but in, I don’t know, the middle of 2022, I think it was a strong, unique intuition to have.
Yeah. Yeah. I think that was something we did really well. We just had this belief, and I think it’s the same thing that you see with the programming products: If you had built something where all this does is check that your Python code doesn’t have bugs—which you could have done better with 3.5—you wouldn’t have built something like Cursor.
The intuition was just these models can help you do any programming task in any programming language. I think we felt that same way in legal. I had a bit of intuition; I did a bit of investment banking and private equity, and it was the same workflows where you could just do any of them with these models.
I think keeping the product open-ended enough that it gave us the room—now we can build into all these things and other professional services—I think that was super important.
It’s a really interesting analogy, because for code it took an extra 2 years, I think, for the main coding companies to really emerge as the ones that are likely to win.
Yeah, that sounds right.
Right. And so you folks started, I think, 3.5 years ago, and you had a product almost immediately. You were up and running really fast. I think Cursor didn’t really launch its IDE until 24 months ago, something like that.
Then Cognition was slightly in that era, and then obviously Claude Code 6 months later. So everything came in a time-delayed way for code, even though GitHub Copilot was one of the first products and everybody knew that was really important.
I always think that’s really interesting, because there were so many coding companies that got started under the premise, but somehow it’s these ones that started a little bit later that really were the ones who took off. I always wonder why that is. What caused that?
My guess—part of my intuition here—was just you guys were very capability-focused from the beginning. AGI is less trendy than it is now, but both you and Winston—as a former investment banker turned AI researcher—were like, “It’s going to be able to do so much.”
The coding people thought that, too.
Yeah, they were.
People were very ambitious.
I just think that maybe you folks immediately focused on product, and that was part of the difference.
I think it was finding the right form factor. In legal, it was maybe a bit more obvious, where the initial form factor was essentially—the initial feature we built that none of the products had at the time was: Upload a document and do something with it, right? And that is a lot of legal tasks.
It was that and then really accurate citations. When you showed people that, they were like, “Oh, this is crazy, because that’s so much of my job.”
I think with coding—
The initial models were also not quite as good. You needed maybe a bit more capability from the base models, and then you needed, I think, to figure out the right way to integrate this into the IDE.
Yeah.
But I mean, I remember the first version of the product that I built was mainly—I used GPT-4, because most of my background was distributed systems and AI research, and I still don’t know React.
I just knew JavaScript and kind of put this together, but I'd be like, “Hey, GPT-4, help me make this,” and you could kind of already see it at the time with programming. That was part of what gave me the intuition that I could analogize to what Winston was doing.
You mentioned that you folks have gone from basically 2 founders to 500 people over the last 3 and a half years or so. You're obviously growing really quickly. The business is working, and you have tons of customer demand.
What are you hiring for? What are you looking for in terms of the next set of employees, or what types of roles are you hiring for right now?
Yeah. On the technical side, we mentioned FTEs. I think, in general, across roles, just strong engineers, and then I would say maybe specific callouts: We just hired a site lead for New York, so we're starting to scale up that office. More folks on the front end and scaling product in general, and then more AI folks as well.
So, yeah, any strong engineer, please apply.
Okay, last question for you.
How many pull-ups can you do? Just kidding.
We did find out. Yeah. Well, I don't know if that was the max, but these guys can both do 15 pull-ups—guys, with a wink in the middle.
Okay, okay, okay, guys. We get it.
What do you mean—in what set?
In 1 set.
We've got to do the 24-hour challenge.
Yeah. Oh, what's that?
Just how many can you do in a day?
No, really?
Yeah.
It's a lot.
Yeah. You can upload it to your TikTok.
Put it on the No Priors TikTok.
Does anyone have a TikTok?
No.
Okay. I don't have a TikTok.
Do you know what TikTok is very good for?
arowana videos.
Oh, this is good.
Yeah, that's really good. Super good.
Yeah, that's some really funny arowana videos.
Yeah, like people—there's one where the Miami girl visits.
Oh my God.
She's like, “Are you fun?” [laughter] So it's very good. I highly recommend it.
I'll have a couple sent.
Yeah. No, that and Twitter—I feel like that's where all my time goes.
Yeah. TikTok—everyone's videos. Okay, you can pick. We'll pick one of these two. I asked some other people involved in the company, “What questions should I ask Gabe?”
Oh boy. We covered some of them, but one of them was, “Why do you still sleep on an air mattress?”
Okay, so I don't sleep on an air mattress. I have a good mattress. I don't have a bed frame, which is where that's coming from.
When we moved from LA to San Francisco, my bed frame broke. The first year and a half of the startup, things were so crazy that I was like, “This is what a startup founder should do.” At some point, I was like, “I need to get a bed frame,” and I ordered one.
It came, and I got a call from the apartment. They were like, “Hey, you didn't sign out, you didn't fill out the insurance—the movers' insurance—so we can't let them bring this up.” I was like, “Okay.” I called them and was like, “Hey, do you guys have renters' insurance?” They're UPS. They were just like, “We don't do that.”
Then I was just like, “I don't have time to deal with it,” and I haven't dealt with it. We have other problems to solve.
Okay, yeah.
So I physically can't do anything except the company right now.
And pull-ups.
Exactly.
The other question was: There's a bunch of foresight in starting Harvey when you guys did. When you look forward, do you have a prediction that you think others don't necessarily agree with you on right now that is not mainstream?
One comment I'll definitely make on the foresight is, I think we've gotten comments like, “Oh, overnight success,” and, “Oh, you saw this coming.” I would say I actually just spent the decade before Harvey trying to start a company like Harvey. I think I was just super early, and then eventually it was like, “Oh, now is the right time,” and you were kind of in the right position.
My guess is that people are now catching up to how capability-building—as you called it—Winston and I were. I think people in Silicon Valley have a good sense of where these models are going, but I think generally people don't appreciate how much better they're going to continue getting.
It's hard to internalize.
It's really weird. Yeah, it's really weird.
I build things and I'm like, “Oh my God, code generation works. It just really works now.”
It's crazy. Yeah. And to me, I think the interesting thing will be the transition from these models being really smart individually. If you think about a lot of what we've done in the past 20 years with SaaS, it's: How do we use software to make these massive organizations?
I think that will be the continued trend, where a lot of what we're starting to think about is that law firms have 10×ed in size compared to before computers and the internet. I think that's going to happen again, but in maybe a different way than in the past 20 years.
A lot of people still talk about copilots and individual productivity, but a lot of the things we're starting to think about are organizational productivity. How do you build these systems at scale where—for our internal engineering team, for example—a really interesting question for Cursor and Codex is that making someone program 20% faster doesn't make you build a product 20% faster?
So we're starting to think about what broader infrastructure you need so these companies can develop software and products faster. Then, kind of the same analogy applies to legal. I think that's one of the things we're thinking about that I maybe don't hear people talk about as much.
Kind of collaborative AI, in some sense. It's sort of like the Figma transition: You're an individual-contributor designer versus working collaboratively with the design team.
Exactly.
What you're talking about is doing that for law, doing that for code, doing that for different verticals, and having AI as a layer on top of that. So it's super interesting.
Yeah. And I think, to that point, it's like, how are humans and AIs going to work super effectively? Because even at these large companies, you have huge teams of different specialized people who have different functions.
When I hear a lot of people talk about these models, they kind of talk about them as like, “Oh, AI will just get smart and do all of this.” I don't think that's the way this evolves, the same way it's not just like hiring 100,000 people and now you've built Walmart. It's like so much of it is how you organize all of these—
3 million, actually.
Yeah, 3 million, actually. Yeah, how you organize all of these. I think that will be one of the really interesting problems for—
Interesting. Yeah.
I'm seeing that a lot in the context of both AI-driven roll-ups as well as this company BrainCo that I helped get up and running, where a lot of the AI implementation issues are around people management and workflow optimization. It's much less about whether you can build the AI and much more about how you actually change the organization to be able to adopt it properly.
Yeah, no, and we're starting to work with a lot of private equity firms. I think it's interesting starting to see how they're thinking about that, because I think that will be a really interesting space.
Awesome. Thanks, Gabe.