Jev 如何把 AI 变成真正能办事的软件
Ben HorowitzMartin CasadoDiogo Almeida
Jev 的核心押注,是 AI 缺的不是智能,而是嵌入软件、能够可靠执行自动化的那一层。 Diogo Almeida 追问的是:“这些自动化到底都去哪儿了?”模型已经聪明得惊人,但默认解决方案仍然是外挂一个聊天机器人,而不是赋予产品新的能力。
编程代理让传统软件开发提速;Jev 要做的,则是让软件本身变得更有能力。 Claude Code、Codex 和 Cursor 生成的仍然是“和10年前代码一样的东西”,而 Jev 增加了一种自然语言原语:表达意图、配合状态机运行,并以一定置信度决定下一步做什么。
真正释放经济价值的关键不是原始基准表现,而是可靠性。 Almeida 将其定义为“每一次都是同样的智能”,这不同于系统可用性或输出完全一致:输入等价时,即便混入 UUID 这类无关变化,也应稳定得到智能响应,让开发者能够把结果纳入代码。
Almeida 曾认为泛化后的 RLHF 可能通向 SHGI,随后却正面撞上了惊人智能与广泛自动化之间的鸿沟。 Martin Casado 以 GPQA 为例,称其中的问题甚至连 Google 都不知道答案;Almeida 的判断是,行业优化的是人类对输出的评价,而不是任务能否可靠完成。OpenAI 自 2020 年起就在尝试自动化客服,但企业里的普通工作依然基本没有实现自动化。
长尾确实存在,但 Almeida 不接受它成为几乎什么都不自动化的借口。 他的务实答案是先自动化回报高、流程固定的工作,再处理困难的边缘案例;他暗示,声称能响应95%请求的支持团队,覆盖的可能只是重复性的密码重置,而面对独特个案时,覆盖率约为50%。
这套机会论可能带来一场“反向 SaaS 末日”。 既有 SaaS 厂商知道哪些工作流最重要,也拥有覆盖大批客户的分发能力;如果 Jev 式原语能够升级产品,而不只是加一个聊天入口,这些老牌公司可能成为“整个 AI 游戏中最大的赢家之一”。Almeida 开玩笑地把这一结果称为“SaaS-appaloosa”。
终局不是又一个面向用户的模型封装,而是一层新的系统基础设施。 Almeida 想象,智能会从顶层交互深入软件内部——起初“更像数据库,而不是标准库”,最终则像“里面装着一点大脑”的逻辑门;至于状态一致性、安全性、成本、速度和强保证等问题,嘉宾仍未给出答案。
1. Jev 让 AI 从编写软件转为软件的一部分
Almeida 的电梯式介绍是:“这些自动化到底都去哪儿了?”AI 是一颗“未经雕琢的钻石”:智能程度惊人,但除了聊天机器人和编程代理之外,几乎派不上用场。TypeSafe 的答案是让 AI 进入软件,而 Jev 是其首个旨在显著扩张自动化能力的模型。
一位主持人将 Claude Code 和 Codex 称为“即时生成软件”:它们通过自然语言按需创建程序,但仍受限于普通软件的表达能力。Almeida 想要的是“智能软件”——能够表达意图,并自动化传统代码难以直接规定的事情。
关键区别在于,Jev 并不是用一个更快的工程师替代人类工程师。人类程序员和编程代理都获得了一种新的原语:用自然语言描述目标行为,将其与状态机组合,再由系统以一定置信度决定怎么做。
Almeida 接受“Jev 确实就是一个分类器”这种看似过于简单的说法,因为分类器本来就是为了产生实用价值而设计的。更尖锐的比较是:在一个狭窄工作流上,Jev 可能已经胜过2019年一支机器学习团队在收集数据、训练模型、搭建定制基础设施后所能做出的成果。
2. 务实的系统思维,而非“神模型”,塑造了产品
Jev 位于 Almeida 入职介绍中的维恩图交集:AI 擅长的事情,与代码内部有价值的事情。它不应去外推模型处理不好的浮点值,而应把智能放在概率和分类能够产生有用行为的地方。
Almeida 猜测,设计空间可能是一根连接“语言输入、语言输出”系统与命令式程序的滑杆。他目前的指引是“每美元对应的智能”,但也承认短期内“每秒对应的智能”可能更重要;因此,Jev 在成为标准库之前,可能会更像一个数据库。
这段经历解释了他的系统思维:数学竞赛把他带入计算机科学,一次 Kaggle 获胜来自“自动化一切”,而 Isabelle Guillon——Almeida 认定她是 SVM 的合著者,并不太确定地回忆说她是第一作者——把他带进了研究社群。此后他又与 Jeremy Howard、Google Brain 和 OpenAI 共事,但 Almeida 仍认为自己更像计算机科学家,而不是 AI 研究者。
Horowitz 给出了乐观的框架:“我们构建的是产品,不是神。”他期待的是一个拥有更多、更好工作的世界,而不是更少工作的世界,并将 Safe Superintelligence 视为更广泛的积极 AI 运动的一部分。Almeida 则拒绝“一个统治一切的大脑”这种悲观叙事:简单任务依然顽固地停留在人工处理阶段,所以他眼下看到的悲剧不是自动化失控,而是强大的智能几乎还没有被释放到生产系统中。
3. 基准测试中的智能,与可依赖的经济活动出现分化
Almeida 的看法在2021年底前后发生了变化,当时 RLHF 的泛化能力远超他的预期。他的团队专门测试一些古怪问题——比如“为什么在冥想前吃袜子很重要?”——并先确认答案没有出现在网上;模型给出的合理、近似人类的回答让他们相信,这种能力不是骗人的把戏。
他起初认为模型有相当大的机会成为 SHGI。结果并非如此,“我的整个世界都崩塌了”,这迫使他面对一个如今驱动 Jev 的问题:如果这种智能是真实存在的,为什么把它转化为有用的自动化依然如此困难?
他的答案是评估错配。由于人类负责判断模型质量,行业优化的是取悦人类评审者,而不是完成任务;这会让模型看起来能力非凡,却让软件层无法信任或组合它们的输出。
Casado 提到数学题和 GPQA 这类看似已经被攻克的问题来强调这种错配;GPQA 被描述为连 Google 都不知道答案的问题,但模型在普通生产性任务上依然会失败。Almeida 承认现实工作存在厚尾式例外,却拒绝接受由此推导出的标准:自动化不需要解决每一个边缘案例。真正的“煤矿里的金丝雀”是,显而易见且回报高的工作究竟能不能自动化——OpenAI 自 2020 年起就在推进客服自动化,但企业内部真正实现自动化的工作依然很少。
4. 可靠性意味着始终如一的智能,而非完全相同的 token
Almeida 表示,产品发布后的使用广度和用例数量都极其出人意料;Casado 说,他几乎每小时都会听到有人用 Claude 处理新的任务。Almeida 强调,用户可能看不到的是为可靠性付出的多年“血汗与泪水”。他相信,可靠性每提升1个百分点,就可能解锁一批开发者目前不敢自动化的工作流。
他区分了3个概念:可用性是正常运行时间或 SLA;确定性更接近于复现同一个结果;可靠性则是“每一次都是同样的智能”。加入一个无关的 UUID 不应改变功能答案,尽管输出的 token 不必完全一致。
终极测试在于,开发者能否不构造示例查询,直接用 Jev 编程;他们之所以能始终“处于持续的流动状态”,是因为相信每次调用都会返回代码能够理解和处理的结果。
编程代理仍然是互补工具,但能力有边界。Almeida 认为它们擅长语法、拙于语义,并且“极其不擅长架构”;架构是软件中最具人性、最具创造力的一层。Casado 补充说,模型在架构上的水平可能只有第50百分位,但当速度是决定性杠杆、Codex 可以运行一整晚时,第50百分位的架构仍可能胜过第60百分位的架构。
5. SaaS 与系统都可能围绕这一新原语重建
针对“SaaS 末日”的说法,Almeida 认为,软件可能便宜、容易复制,却并不容易复现:大量行为都藏在底层。既有 SaaS 公司理解用户和工作流,也拥有分发能力,因此天然适合把能力升级扩散到现有客户群。
设想中的界面变化远不只是增加聊天功能。表单和选项菜单可能消失,因为它们看起来只是把自然语言映射成类似 Java 的结构;一个未经核实的开发者应用案例涉及语音交互:系统持续判断语音究竟是命令还是文本输入,以及应当把输入放在哪里。
系统边界仍未确定。Casado 提出状态一致性、可靠性和系统层面的强保证;Almeida 认为日志分析、邮件分析、用户界面设计和人机交互都会改变,也提到空中交通管制系统可能成为应用场景。Horowitz 的原则是,先自动化容易的工作,再处理困难的工作。
Almeida 最深层的类比是 TCP:把不可靠的东西转化为可靠的一层。如今模型甚至可能无法遵守 JSON schema,开发者只能把输出交给人或另一个语言模型,于是形成“人在回路中的”聊天机器人,或一个不断运行的 agent “while” 循环。Jev 要做的,是把自然语言变成可用于生产的状态机。Horowitz 称它为“反挫败机器”,而 Almeida 心中的乌托邦,是技术终于能够“按我的意图行事”(“do as I mean”)。
完整逐字稿
Where the hell is all that automation? AI is incredibly smart, but so useless at everything else. It doesn't matter how many AI-based coding agents you use; the software doesn't actually get better. Maybe you're just running it faster. It seems like you're getting worse. OpenAI has been trying to automate customer service since 2020.
Instead, I want smart software. I want to expand the capabilities of the software itself so that what should be automated finally becomes automated.
What I like most about your slogan is, “We build a product, not a god.” This is so cool, because if we had any other leader of a big lab, even if they were happy, they would hide it. Your view is so different. You say, “No, we will create a much better world.” Due to certain nuances, I don't think we're on the path to an AI apocalypse or anything like that.
1. Meet Diogo and Type Safe
Today we have with us the founder and leader of TypeSafe, Diogo Almeida, who is a true hero to Martin and me. Not only is he creating a really interesting product, but he's also starting what we think is a very important movement. We are very happy about today's meeting. Welcome to the show.
Thank you.
Maybe you could briefly explain what Jev is, what TypeSafe is, and why it is important?
Is this a Chris-friendly format or not?
Yes, yes, yes. Okay, cool.
I was once asked to do an “elevator pitch,” and I tend to get carried away and I'm not good at it, but I realized that my favorite pitch for Jev is, “Where the hell is all that automation?”
It's incredibly tragic. There's so much intelligence. AI is incredibly smart, but it's not that I hate chatbots or coding agents. I adore them myself, but they are so useless at everything else, and it's just a tragedy.
It's tragic because we have such a huge diamond in the rough that's not yet ready for work. That's why TypeSafe builds AI for software. We want to make AI powerful not only for people, but also for creating real software, and Jev is, for us, the first model in this area that will significantly improve automation.
Yes, it's interesting because it's already become a real hit in the programming world. One of the things that made us ask, “What the hell is going on here?” is that every developer we know calls us and says, “Oh, this is incredibly cool. It's fast. This is great. Everything is getting better.”
How is that possible? Everyone thinks, “Well, we have Cloud Code, we have Codex. Don't we already have that?” What's the difference, and how does it lead to true automation?
I wish I had some visuals, because I have a favorite picture for this. I like Cloud Code and Codex. I like Gary Tanenbaum's description of it as just-in-time software. It's an incredible way to describe what they do. It's like creating software on the fly, and you can program it in natural language, but it has the same expressive power as regular software.
I want smart software instead. Instead of automating software development, I want to expand the capabilities of the software itself so that what can and should be automated becomes automated. In more refined language, I want to express things like intention. I want to expand the vocabulary of what we can do.
I can talk about all sorts of weird sci-fi stuff I want, but programming is about hyper-detailing valuable things and then endlessly reproducing them. It's so cool, and I just want more of it.
2. Smart software, not just faster code
One way to think about it is this: instead of a tool that, to some extent, replaces the software engineer with a faster, maybe not even as good, engineer, you say, “No, no, no, we're going to give existing engineers superpowers to write much, much better and more interesting things.”
Yes, I actually mean that. By the way, I think a lot of people miss this point, and it's so subtle and so important to get it out. If you're using something like Claude Code or Codex, which is great, or Cursor, which is also great, they write code, but that code is the same as what a human would write. Maybe it's better, maybe it's worse, but it's basically still code, the same as code from 10 years ago.
The thing about Jev is that, whether you're a Cloud Coder or a human, you get this new primitive—a new thing that you put into the code that actually extends the capabilities of the software. So instead of writing code, it's something you add to your code.
Which one? Continue.
No, you continue.
What's interesting, by the way, is that it's a very powerful primitive, and it would be great if you could explain it, but it's also a little different from how programmers think. For example, there's the concept of probabilities, and therefore, perhaps—
That is, the intelligence level within the software. Like a library where you can describe what you want in natural language, add a state machine, and it chooses what to do with a certain level of confidence, which has never been done before on this scale.
There's a lot of nuance there. I'll first jump into one thing I like about what you said—the question of where the hell all the automation is. I love software so much. I would like to write it all day. I don't recommend people to be CEOs, but I don't care.
It's strange that AI is so cool and the software hasn't changed in 10 years. For me, it doesn't make sense, and the most we can do is add a chatbot on the side that can sometimes perform actions, but not all of them, because some actions are not reliable enough. I just wanted to make that little digression. I like this moment, so I'll go back to the idea that it's a slightly different way of thinking.
Yes, I think “machine-native” doesn't exactly match the bits perfectly. This is, in fact, the art we are working on. On the first day of onboarding, I draw a Venn diagram of what AI is good at and what is valuable in code, and we are in the middle.
We don't output, for example, extrapolated floating-point numbers, because AI simply doesn't do a good job of that. But things like probabilities are nothing new, and this is like the argument over whether Jev is just a classifier.
Jev is definitely a classifier. Classifiers are cool. They were created to be useful. They are designed to be useful. In general, this is the same interface as some natural-language-processing concepts, because they came from practitioners who were trying to make systems work.
What I see now is that Jev, in my opinion, is probably better than a machine-learning team from 2019 that would do it for you, because you can program on the fly. Who knows what can be built? In 2019, there weren't many good machine-learning teams to fill narrow niches, collect datasets, measure them, and do all that. This is just the beginning. I feel like there's still a lot ahead.
Do you think there's a certain slider here, where at one end there's “language at the input, language at the output,” as we have today, and on the other side there's a regular imperative program? You can move between them, or do you think this is a point in the design space—language input, machine-state output—that will become a universal tool for programmers?
That's a difficult question. I'll say what's on my heart. Deep down, I think it's a slider. When I was designing our primitives, I might have made mistakes due to my own preferences, but right now my main guideline is intelligence per dollar. To be clear, this could be a mistake. Intelligence per second may be more valuable in the short term.
Even our interface—we call the input a “state”—was designed that way intentionally. It's the same as with a code patch. This should be an internal part of programs. Deep down, we are truly optimizing for that. Most of my work is about even more complex internal state structures in programs. Can intelligence be added there? I think it will be a constant struggle.
We take design very seriously. Pragmatically, I think certain things happen by themselves. For example, it's easier to create AI that works at millisecond scales, so for a while it will be more like a database than a standard library. But I would like to see this become a standard library as well.
Can I step back for a second? What alchemy creates someone like Diogo? When a man and a woman love each other—
But listen, you're talking like an AI researcher. You're still talking like a systems engineer, and you're talking like a programmer. Usually these things don't overlap much, and you take the AI that we've developed as an entity and turn it into a programmer's tool. Tell me a little about your personal path, which is somewhat unconventional.
My path in AI is somewhat unconventional. I was a mathematician. I was a mathlete. I describe it this way: I was pretty good at math—it's a little embarrassing. I was good enough at math to attract girls, so that's pretty good.
No, that was the point.
You must be pretty cool. What kind of girls do you attract when you're so good at math?
This is interesting. This is something our audience needs to know.
We have to inspire young people.
Don't do this. Young people, don't do this. It's not worth it. Just be cool, calm, and interesting, and don't try to seem better than you are. I can't believe I said that.
I was a mathlete, but I've never actually liked math. I never made any effort. I was just a big fish in a small pond. For me, mathematics was the path I was directed toward, but I hated it because it was all about winning competitions.
Computer science is actually very similar to mathematics. It's essentially math, but cool and useful. It's fun and interesting, and I still love conducting algorithmic interviews.
Is this the best I can do? I don't know, but do I like it? Yes. And does that allow me to judge people very well? Yes, it does. So I love computer science. I consider myself a computer scientist much more than an AI researcher, despite my background.
What really got me into this was that I also won a competition on Kaggle. Not through complex math, but by automating everything with more and more nested loops. It was a systems problem, and that event ultimately led me here. For example, I was forced to speak at NeurIPS. It's usually an honor, but I hated it because I just wanted to be in the thick of things.
Was this from that Kaggle competition?
Yes.
Wow, wow.
Actually, the presenter on Kaggle was Isabelle Guillon, co-author of SVM. In fact, I think she was the first author of the SVM paper, although I'm not 100% sure who the first author was. She saw that I was someone who didn't fit into the research community at all, took me under her wing, introduced me to all the AI experts, and my career went in that direction.
And then OpenAI?
No, it was a startup with Jeremy Howard.
Are you kidding? I adore Jeremy.
Yes. Fantastic. And then Google Brain for a while. Eventually, I got tired of doing nothing. I thought, “You know what? AI is a hell of a lot of fun.” I joined OpenAI for that very reason, and it worked out really well—surprisingly well.
So you said something that's quite unusual in today's world: AI is actually a lot of fun. Your company also has a completely different behavior and view of AI than everyone else. My favorite thing you say is, “We create a product, not a god.” It's so apt. If it were any other leader of a large laboratory, they would try to hide their joy, even if they felt it.
Your approach is so different. You say, “No, we're going to create a much better world, and it's going to be great. There won't be fewer jobs; there will be more jobs, and they'll be much better jobs. Everyone will have a great time.” Just being around you shows that you truly believe it.
So tell us about it, because for us, Safe Superintelligence is more than a company. It's a whole movement toward a positive future that most people in the AI world don't seem to like, or at least don't share.
I think they don't understand it. You know, it's like complaining about a classifier at the level of machine-learning problems while everyone else is having a party with Jeff. It's like, hell, we can do whatever we want, and if you're not a developer, it's hard to understand what's really going on.
I agree with that 100%.
I really think they're painting a pretty bleak picture, which I obviously disagree with. I think it comes from this blind belief in the monolithic model that everyone believes in.
Exactly. One big brain to rule them all.
This is what we did. It sounds much more sinister.
Yes, that's what people hear. Of course, that's what people hear.
But is this one brain really on its way to ruling us all? We still haven't automated very simple things that, in my opinion, people shouldn't have to do. There are a lot of really basic things.
It hurts me when the world doesn't match reality. Part of the pain is, where the hell is all this automation? How can AI be so incredibly intelligent, and why are there so many financial incentives to automate everything? You can find an excuse for diffusion, but I don't believe it at all. It's not worth mentioning names, but this is obviously not true.
Part of the problem is that the mismatch with reality, combined with the fact that AI has enormous potential, makes it tragic to me that we haven't released it. So now it's like a little holiday for me, but I was afraid. And all the Jeff users, it's like—there's a happy AI, the people on Jeff, and there's a gloomy AI, the people who aren't there.
Yes, yes, yes. It's actually quite a fascinating dichotomy.
I'll tell you, on your point about automation, I had an interesting conversation this morning with David George, who runs our growth fund, because we were talking about new tools. I asked, “Have you tried the Muse thing?” He's like, “Oh, she's great.” I asked, “What are you doing with her?” He said, “I finally canceled my subscription to The New York Times.” And I said, “That's hard to do.”
3. Where's all the automation?
This is just the tip of the iceberg of the terrible things we need to automate.
I think if we want to be intellectually honest and truly strive for the “north star” of automation, we can't fall into the same anti-patterns that AI has fallen into—namely, focusing on individual use cases.
Many people ask me, “What are your favorite uses?” And I say, “I'm not sure if they work.” I want them to run in the background so that someone can trust them to run without having to check, and so that people can build something new on top of them.
And, you know, Layout, layout, right? Layout, but also security, right? This is a different type of security where if you want to run something with associated resources and access to things, you need guarantees to do so, or at least statistical guarantees that nothing will go wrong. Break it into pieces, make it fancy, something like that.
I don't think our models will do that anytime soon unless someone writes software around them. That would be a very cool achievement, and we should give credit for it. But not in a way that makes us lose responsibility.
I'm just curious: how long did this intuition take to mature? I remember talking to you, maybe in 2017. It was a long time ago. We really talked about it, and you already had a lot of these ideas. You said that data is important, and you said that you wanted to focus on the task.
Did you know that this would all end up being a classifier, or was it just an intuition that there was a different perspective on this whole AI movement?
There's a funny story about that conversation. During a talk in 2017, my topic was very similar. I think it was called something like “AI: Modular in Theory and Flexible in Practice,” which is very typical of software systems, so I'm sticking with that idea a bit.
I think this all actually started right before ChatGPT. When we released these things, I had no premonition. To be honest, I was very, very pleasantly surprised by the generalization capabilities of reinforcement learning from human feedback.
When was that?
Probably at the end of 2021—maybe the fourth quarter of 2021. It was really very general.
If you read the paper, unlike others who are trying to prove themselves right, we actually tried to disprove it using the scientific method: Is this a scam? My favorite question was, “Why is it important to eat socks before meditation?” We made sure that this was not already on the internet, and the models were able to provide plausible, human-like responses.
For us on the team, it was a moment of truth: this is not a scam. In machine learning, you should always be wary of fraud.
What really blew me away was how we released it. I'm obviously a big fan of the possibilities of AI, and I did a lot to release this model. I really thought the model had a good chance of becoming SHGI. When that didn't happen, my whole world just collapsed.
So you were a little on the other side for a while, like that “Crazy Train” or...
Well, no. I'm just like, no, no, no, no, no. RL generalizes. Maybe we have SHGI. Reinforcement learning from human feedback generalizes quite well. Reinforcement learning from verifiable rewards is something that doesn't generalize that well, from what I saw.
You were behind the scenes at ChatGPT. You were behind those early GPTs. It was a completely different goal: creating a chatbot that would communicate with a human, not a programmer's tool, and so on. So I'm just curious—
Actually, at the very beginning, sometime in 2020 at OpenAI, when we were talking about SHI, people were describing it as Elijah and every statement he made. It's not exactly like that, but we were talking about the definition of AGI at OpenAI.
Part of it is that it's intentionally blurry. It's a wide tent so everyone can be inside. But due to certain nuances, I don't think we're on the path to self-improving AI. I didn't think so before, and I don't think so now.
I really think that what OpenAI has defined as general-purpose AI is quite achievable. Automating most of the world's economically valuable labor sounds like—oh, I don't know. There is so much different work there, but most of it is very routine and simple. In terms of volume, to be able to outsource the work, you need simple instructions that anyone can follow.
As far as I can tell, the level of intelligence required for this has been in these models for quite some time. My main complaint is, “Why isn't this available yet?”
Since the advent of reinforcement learning from human feedback, the industry has been divided: big promises but weak results. I think GPT-3 was pretty balanced at the time. But when humans judge the quality of the models, they seem to be very accurate, because humans are the judges and we optimized for that judge, not for automation. That was the missing link.
I would say that's when it dawned on me: Why isn't this thing more useful?
So you think the measure should be how much we can automate real, productive tasks? When you say “many promises and few results,” is that the dimension you're referring to?
The ability to automate tasks. In my heart, it feels like cool science fiction, and I believe this is the first warning sign for such “cool” fiction.
Are you seriously claiming that the math is solved, or even that 2 years ago GPQA—the questions that even Google doesn't know the answers to—was solved, but we still can't handle ordering a self-driving car? This is very difficult to grasp at the same time.
4. Is it just a data problem?
And I think a lot of people don't have a good answer to that. Can I test one thing? It may not make sense, but isn't there an argument that the real world is different from the digital world? It has heavy tails in its distributions; there are many outliers, and we don't have all the data.
Couldn't it be that the reason we don't do productive things in the real world is simply that we don't have data for this distribution? We don't train on it, and that's why it's all reduced to these lower-dimensional manifolds, like math or code?
I don't quite agree with the data argument, as far as I'm concerned. I believe there is a long tail. Certainly, it would be kind of silly to deny it. But I don't think, in my situation—with the canary in the coal mine—we need to automate this long tail.
I think we need to be extremely pragmatic about everything. Creating reliable software is always an investment, right? Do you know what the three great virtues of a programmer were? Laziness, so as not to do it again, arrogance, and there was a third.
No—yes, I remember. This has been around since the days of Perl.
Yes, yes. There's a third one, dude. I would like to remember it. But laziness means spending 10 hours automating a 5-minute task once and never coming back to it again.
This should be a solution based on return on investment for people who automate things. I just wish it could be automated.
Yes. And I think people will just create new kinds of work, hence the Jevons paradox, when that becomes possible. But as a benchmark, I feel it's useful to see if we can actually automate what AI seems like it should be able to automate.
OpenAI has been trying to automate customer service since 2020. It's not that simple. It's pretty amazing. It's just crazy.
Well, within companies, very little is automated right now. The projects didn't work.
Yes, apart from programming, which works great. Can you perhaps categorize the types of problems that you think are easier to automate? This is quite interesting.
We were actually looking at support even before the current wave of generative AI. It was interesting to meet the company, and the company said, “We respond to 95% of all support requests.”
And I'm like, “That's so much.” But then you really look at the data and understand that it's all about resetting passwords. If you were judging by uniqueness, it was 50% or something like that.
So it seems that when you deal with people and natural systems, there is a very long tail of exceptions.
To what extent did you even foresee such a wide range of use cases? I probably get a message every hour that someone is using Claude for a new task. I'm like, “I had no idea.” Did you expect this to happen? Are you surprised by it?
Extremely surprised. I didn't expect this to happen. This launch wasn't something that anyone would have expected. They'd probably have to be crazy.
I don't think anyone would have expected ChatGPT for developers, because ChatGPT was for regular users, and that's weird. I don't even really know what percentage of people who joined Jeff's party are developers. I can't imagine anyone other than developers using it. I don't know how they would use it.
But even my friends who aren't developers are just part of this Twitter party, joking around and all that. So, number 1, it's phenomenal.
5. Reliability and "do what I mean"
Number 2, this is going to be hard to convey in a short message because it was blood, sweat, and tears over many years. The extent to which I care about reliability is a lot. Reliability is the essence of this thing. If you don't understand this, it will be very difficult to create an analogue that meets the criteria.
It seems to me that every percent of reliability will be valuable to everyone, even if it's not the most valuable thing in terms of market capitalization, because it will simply open up new opportunities. We're fighting for all sorts of strange levels of reliability that we don't even fully understand, because we're just trying to introduce this “electric motor” of AI—of intelligence—into people's workplaces, and they'll decide what to do with it.
What does reliability mean in this context? Is it the availability of the model, or is it that I call the model and it returns the same result? How should I even think about reliability for something that is inherently stochastic?
It's not so much the first; the second is closer to the truth. I would describe the first as uptime or an SLA. I would call the second one closer to determinism, and the third one I would consider more like reliability.
I would describe reliability as the same intelligence every time.
Oh, interesting.
Yes. It's not exactly determinism, because I believe determinism is useful for unit tests, but not for real systems. Think about it: if you add a UUID to the query, the result should be the same because, functionally, it's the same.
Yes, but it is not entirely deterministic.
Right. I think there is another level, the name of which I don't know yet. Maybe it's what I would call a certain form of intelligence, where the result doesn't have to be identical, but it has to be intelligent every time.
If you were in that situation, would that thought be understandable to a human being? The developer can take this into account in the code.
Really, for me, the ultimate measure of reliability would be to reach a level where people can program with Jev without creating example queries. When you just trust it, you're in a constant state of flow, creating great, incredible software.
And many programs are at this stage right now, aren't they?
I don't think it's entirely unrealistic, and I'm perfectly fine with it.
By the way, this is a bit of a strange question, so feel free—if it's too strange, just don't worry about it. But it occurred to me that the value of things like coding agents actually diminishes if you have a primitive like that.
You could say, “Whatever. Codex builds all the software for me, but it doesn't use Jev, so the software it builds is somewhat limited.” Or you could say, “As a human, I'll write a program without using a coding agent, but I have this very universal primitive that makes writing easier.”
Do you see a future where coding agents use Jev and you manage them, and do you have any redundancy? Or do you think people are implementing it themselves? This is more of a question about coding agents than it is about Claude.
Oh, yes, definitely. I feel like I'm not as immersed in the intricacies of programming as I would like. You two may know better than me about this, and that's unfortunate, but my experience is that coding agents are very good at syntax and really bad at semantics.
I would say they're incredibly bad at architecture. I think architecture is the most human, creative part of software development. That's why I like using AI agents for coding.
I think this is almost certainly not part of the data distribution. It would be creepy if they were learning from our users' data. So most likely, that's not the case. But I don't see a problem with instructing them on syntax when it's part of the distribution.
The thing is, the models may not just be terrible at architecture; they may be at the 50th-percentile level. And if you yourself don't know anything about architecture, then this will be enough.
These are all gray areas and compromises that you will have to navigate. Sometimes speed is exactly the leverage your company or project needs. You're willing to choose a 50th-percentile architecture over a 60th-percentile one because you want to move faster, so Codex can run all night or something.
Yes, by the way, there's an interesting phenomenon in the market. When AI agents appeared, a “SaaS apocalypse” occurred and their value fell sharply. And when Claude came along, every SaaS company said, “This is the best thing in the world.”
So explain this.
6. SaaS apocalypse, reversed
I don't even know what else to say. It seems quite natural to me. In this story of the SaaS apocalypse, the idea that, in my opinion, has not been fully realized is that software is very cheap and easy to copy.
I can believe the first part, but not the second, because a lot of things happen under the hood. Maybe I'm too fanatical about software.
We all are.
Yes, okay, okay. I didn't know where this code could be. We have a lot of heritage in this regard.
So I don't think it really worked. The markets seem to disagree, but in my opinion, SaaS provides the same value as before. Maybe the markets are just scared.
But I think SaaS is going to be one of the biggest winners in this whole AI game. And I want to work really, really well with all those big, boring, user-aware SaaS companies because, in my opinion, they know best what workflows are worth automating.
What do people need? This is their main bread and butter. They spend big money because software is always a capital investment, but you invest it up front to make this experience even better, and then it spreads to the entire huge user base.
True.
So I think it's going to be—I’m not going to make any predictions about the financial markets—but in terms of opportunity, it's going to be like a reverse SaaS apocalypse, and I'm just excited about it. I need to come up with a name.
Yes, you should have a name.
SaaS-appaloosa.
That sounds a little too much fun.
All SaaS applications will suddenly become much more useful. And by the way, a significant portion of a SaaS company's capital investment goes into actually reaching all its customers.
So if you've already covered all your customers, and then you've not just made a chatbot in your SaaS product but have really, significantly improved the software itself...
Hmm, this is incredible. I don’t know if this is a realistic dream, but I think there’s a world where forms with options to choose from simply disappear. It seems to me that they just map natural language, which the software already understands, into Java-like output.
It’s literally been around since the 1980s. We used to call it 4GL. Remember the MR2? Fourth generation language.
Yes. I also think it’s from the ’80s. It might even be an insult. I wasn’t born yet.
I think the concept of “do what I mean” will reach a whole new level. If I can mention one developer app, I don’t know how reliable it is, so I can’t promise anything, but it was so cool. Someone was using their voice to control a computer, and they were constantly making decisions: Is this a command, or is this text input? Where is this text input?
This sounds incredibly cool. I think the interfaces might just change completely, and we might have to do it cheaper and faster.
Yes, then you’re already in Star Trek.
There’s a deep intuition here that, if you’re using AI to build software today, you’re still creating the same software as before. But if you look at the average pull request for a large company, it’s about 10 lines, right? Seriously, it really is. I was at Google, and we did some research. It’s about 10 lines.
You automate 10 lines. And, by the way, those 10 lines are part of the learning from the client or something, so you’ve optimized something that’s actually pretty minimal. What it doesn’t do is give the software new capabilities. It’s like automating something that ultimately turns out to be relatively insignificant.
Now truly new opportunities have emerged, and it’s entirely possible that the software will simply get better. It never even occurred to me before that no matter how many AI agents you use to code, the software doesn’t actually get better. Maybe you write it faster, but are you really getting worse simply because there’s less supervision?
So I think it’s often less safe.
Yes, that’s right, without a doubt. But you can now argue that applications will have new functionality because of this, because you’re providing this new primitive. In a sense, it speaks natural language—it can reason—but it combines that with a state machine.
7. New capabilities, not more code
If people took that as the main takeaway, it would be the biggest compliment to what we do. I even think it’s almost too grand a vision: going beyond the 3 logic gates that we have into something where our types are like the same logic gate, but with a little brain inside.
That would be the best compliment to the legacy of type safety, because it’s a huge, nontrivial thing for the world. I’m not going to promise more than I can deliver, but I will fight for it.
I think there are still open questions about how deep this can go in serious areas like state consistency, reliability, or a real system level where you need to provide strong guarantees. This will 100% change things like log analysis, email analysis, user-interface design, and human interaction.
That’s true, but you could argue that, over time, it will become something like a smart database. And also air-traffic-control systems—anything, right? What we really need is a little scary.
I think my philosophy is to automate the easy work before the hard work.
I also think that a whole era of probabilistic programming will open up. By the way, there’s a huge history of probabilistic programming, which actually died in the ’70s. I’m familiar with this.
I actually think it’s similar to what you could call neurosymbolic AI. Your co-founder Eric came from this very field. He told me.
Oh, cool. Yes, he studied biology a lot.
What I mean is, I’m not a fan. My brand is pragmatism—extreme pragmatism. I’ve never been a fan of biologically inspired things at all. Have you ever noticed this? I mean, I don’t think they ever worked.
It’s useful for motivating crazy people to work on something for decades until it works, and then they refine it into an engineered version. That’s the history of AI today: neural networks, for sure. But a lot of the stories about how it worked were inaccurate, right?
The hierarchical features of the applications didn’t really work, because otherwise ResNet wouldn’t work. Long story. I really think it’s promising, but from a systems perspective, I’m not excited about that part. I’m excited for the world—not that I’m going to program it—because it sounds really complicated.
But I think that when we have a lot of intelligence with different trade-offs in cost and speed, system specialists will make choices where developers will be 1,000 times smarter than them. They just need a rough reference to make a rough guess and optimistically push the data here and there.
It’s going to be incredible, these capabilities that will be available in extreme systems.
The good news is that we’re going to be able to rebuild systems again, and that’s great, right? We have a new primitive. No, seriously, we have a new primitive—a new way of thinking about building software.
We’ve done this before. We did it with the internet. We went from mainframes to client-server architectures. We do it periodically. And, by the way, because of cybersecurity issues, we’ll probably have to rebuild almost all systems to make them secure.
I think so. I think it’s pretty obvious that there’s no—or at least no critical infrastructure here, for sure.
8. Apps vs the guts of systems
Do you think of it more as applications, SaaS, and analytics, or more as systems fundamentals—or all of that together?
Do you mean what I’m thinking about, or a general application for this thing?
When you think, for example, about working on Jev and imagining how people would adapt it, do you even have any idea?
I have a little bit of an idea. I think of it a little bit as diving into the very depths of TCP. You know how TCP is a transition from unreliable to reliable? That’s my analogy.
When I think about AI—and that’s why I’m interested in intelligence per dollar, just to be clear—I came to this conclusion by going from the opposite direction, from an AI-based economic revolution. AI is everywhere in science fiction and all that. It’s as if all software has AI everywhere.
I ask myself: What percentage of calls to AI—imagine that’s a function—are intended for human consumption, where style is needed and all that? That will be a lot of nines. And, from the same question, how many will be at the first level, and how many will be somewhere deep in the jungle?
I think there will be a lot of nines in the slums, but it will all start from level 1. If you’re not aiming for the jungle—wait, that’s kind of weird. If you don’t aim for the very essence, it will take you a long time to reach your goal.
People don’t realize how much AI has been built into software in secret. Even if you tried to build AI into the software, it behaved unpredictably, because software doesn’t understand natural language very well. You do all these strange things: You put in the prompt, “Here is the JSON output I need. Here is the schema,” and it never listens to you.
As a result, you just had to take that output and give it to the person. You say, “Well, to hell with it, right?” Or you give it to another language model. That’s the “while” loop—the agent’s “while” loop, right?
Going from the basics, it should be either a human in the loop, which is a chat, or an agent, which is a “while” loop, because natural language needs to be routed back into the next loop.
I’ll put it this way: I watched it, and it was like the 5 stages of grief. People take AI and say, “I’m going to use this in my software.” Then there’s denial, trying to make it work, anger, and then they move to acceptance: “Okay, never mind. I’ll just give it to another language model or a person.”
So it was all very chaotic. I think this is the first time I’ve seen someone say, “You can take a language model, take AI, and actually turn it into a state machine,” and do it productively.
I hope so. I also don’t want to overpromise and underdeliver on this. I don’t know if it’s ready for all the uses it’s been so loudly promised for. I really, really want this, and my team will, of course, fight for it. We really care about reliability.
We could have released this much sooner. I don’t think people understand this, and honestly, I don’t think they will. Judging by the Twitter discussions, people may never understand it intellectually, but they’ll just have a feeling of, “Oh, I can trust this.”
This is the anti-frustration machine.
I hope so. I hope “do as I mean”—for me, it’s about fluidity in the world, so that everything moves more smoothly and connects like gears. I have my own utopia with artificial intelligence in various directions, and “do as I say” is a huge part of it.
Imagine if all technology just did what you meant. This isn’t science fiction. Look how smart artificial intelligence is, right?
Yes, it’s amazing, and perhaps that’s where we should end: “Do as I say.”
Yes, I like it.
Thank you, Diogo. It was a great conversation. I really liked it.
Mm-hmm.