[BidClub_]
The a16z Show · · 42 min

How Jev Turns AI Into Software That Gets Things Done

Ben HorowitzMartin CasadoDiogo Almeida

AI & SoftwareTechnicalCompany Building
YouTube ↗
TL;DR
  • Jev’s core bet is that AI’s missing layer is not intelligence but reliable automation embedded inside software. Diogo Almeida’s question is “where the hell is all that automation?”: models are extraordinarily smart, yet the default fix is still to bolt on a chatbot rather than give products new capabilities.

  • Coding agents make conventional software faster to produce; Jev is meant to make the software itself more capable. Claude Code, Codex, and Cursor still generate “the same as code from 10 years ago,” while Jev adds a natural-language primitive that expresses intent, works with a state machine, and chooses what to do with a confidence level.

  • Reliability—not raw benchmark performance—is the economic unlock. Almeida defines it as “the same intelligence every time,” distinct from uptime or identical outputs: equivalent inputs, even with irrelevant changes such as a UUID, should reliably receive an intelligent response that developers can incorporate into code.

  • Almeida once thought generalized RLHF might lead toward SHGI, then confronted the gap between impressive intelligence and broad automation. Martin Casado highlights the mismatch by invoking GPQA questions described as ones even Google does not know the answers to, while Almeida’s diagnosis is that the industry optimized human judgments of outputs rather than dependable task completion. OpenAI has been trying to automate customer service since 2020, yet ordinary company work remains largely unautomated.

  • The long tail is real, but Almeida rejects it as an excuse for automating almost nothing. His pragmatic answer is to automate high-return, routine work before hard edge cases; even support operations claiming to respond to 95% of requests may cover only repetitive password resets, while unique cases are around 50%, he suggests.

  • The opportunity thesis is a potential “reverse SaaS apocalypse.” Established SaaS vendors know which workflows matter and have distribution across large customer bases; if Jev-like primitives upgrade the product instead of merely adding chat, those incumbents could become “one of the biggest winners in this whole AI game.” Almeida jokingly calls the outcome “SaaS-appaloosa.”

  • The end state is a new systems layer, not another user-facing model wrapper. Almeida imagines intelligence moving from top-level interactions deep into software—initially “more like a database than a standard library,” eventually functioning like logic gates “with a little brain inside”—while the speakers leave open questions around state consistency, security, cost, speed, and strong guarantees.

Digest · the substance, structured for research

1. Jev shifts AI from writing software to becoming part of software

  • Almeida’s elevator pitch is “where the hell is all that automation?” AI is a “diamond in the rough”: remarkably intelligent, but largely useless outside chatbots and coding agents. TypeSafe’s answer is AI for software, with Jev as its first model aimed at materially expanding automation.

  • One host frames Claude Code and Codex as “just-in-time software”: they create programs on demand through natural language, but retain ordinary software’s expressive limits. Almeida wants “smart software” instead—software that can represent intention and automate things conventional code cannot readily specify.

  • The crucial distinction is that Jev does not replace an engineer with a faster engineer. Human programmers and coding agents both receive a new primitive: describe a desired behavior in natural language, combine it with a state machine, and have the system choose what to do with some level of confidence.

  • Almeida accepts the apparently reductive description—“Jev is definitely a classifier”—because classifiers were designed to be useful. His sharper comparison is that Jev may already outperform what a 2019 machine-learning team could build for a narrow workflow after collecting data, training models, and constructing bespoke infrastructure.

2. Pragmatic systems thinking—not a “god model”—shaped the product

  • Jev occupies the overlap in Almeida’s onboarding Venn diagram: what AI does well and what is valuable inside code. It should not extrapolate floating-point values that models handle poorly; it should introduce intelligence where probabilities and classification create useful behavior.

  • Almeida suspects the design space is a slider between language-in/language-out systems and imperative programs. His current guide is “intelligence per dollar,” though he concedes intelligence per second might matter more near term; Jev may therefore resemble a database before it becomes a standard library.

  • His path explains the systems bias: mathematics competitions led to computer science, a Kaggle win came from “automating everything,” and Isabelle Guillon—whom he identifies as an SVM co-author and tentatively recalls as its first author—pulled him into the research community. Work with Jeremy Howard, Google Brain, and OpenAI followed, but Almeida still identifies more as a computer scientist than an AI researcher.

  • Horowitz supplies the optimistic framing: “we build a product, not a god,” a better world with more and better jobs rather than fewer. He treats Safe Superintelligence as part of a broader positive-AI movement. Almeida rejects the bleak “one big brain to rule them all” framing: simple tasks remain stubbornly manual, so he sees today’s tragedy not as runaway automation but as powerful intelligence that has barely been released into productive systems.

3. Benchmark intelligence diverged from dependable economic work

  • Almeida’s view changed around the end of 2021, when RLHF generalized far better than he expected. His team tested deliberately strange questions—“why is it important to eat socks before meditation?”—after checking that answers were not online; plausible human-like responses convinced them the capability was not a scam.

  • He initially believed the model had a meaningful chance of becoming SHGI. When that did not happen, “my whole world just collapsed,” forcing the question that now drives Jev: if the intelligence is real, why does it remain so hard to convert into useful automation?

  • His proposed answer is evaluation mismatch. Because humans judge model quality, the industry optimized the human judge rather than task completion; that can make models look extraordinarily capable while leaving the software layer unable to trust or compose their outputs.

  • Casado highlights the mismatch by invoking the apparent solution of math and GPQA—described as questions even Google does not know the answers to—despite failures at ordinary productive tasks. Almeida concedes that real-world work has heavy-tailed exceptions but rejects the implied standard: automation need not solve every edge case. The “canary in the coal mine” is whether obvious, high-return work can be automated at all—OpenAI has pursued customer-service automation since 2020, yet very little is automated within companies.

4. Reliability means consistent intelligence, not identical tokens

  • Almeida says the launch and breadth of use cases were extremely surprising; Casado reports hearing hourly about people using Claude for new tasks. Almeida stresses what users may miss: years of “blood, sweat, and tears” spent on reliability. He believes every additional percentage point will unlock workflows developers currently cannot risk automating.

  • He separates three concepts: availability is uptime or an SLA; determinism is closer to reproducing the same result; reliability is “the same intelligence every time.” Adding an irrelevant UUID should not change the functional answer, even though the tokens need not be identical.

  • The ultimate test is whether developers can program with Jev without constructing example queries—remaining “in a constant state of flow” because they trust each call to produce an intelligible result their code can handle.

  • Coding agents remain complementary but bounded. Almeida finds them excellent at syntax, weak at semantics, and “incredibly bad at architecture,” which he calls software’s most human and creative layer. Casado adds that models may be only at the 50th-percentile level in architecture, yet a 50th-percentile architecture may beat a 60th-percentile one when speed is the decisive leverage and Codex can run all night.

5. SaaS and systems may both be rebuilt around the new primitive

  • Against the “SaaS apocalypse,” Almeida argues that software may be cheap and easy to copy but is not easy to reproduce: substantial behavior lives under the hood. Incumbent SaaS companies understand users, workflows, and distribution, making them natural vehicles for spreading a capability upgrade across an existing customer base.

  • The envisioned interface change is larger than adding chat. Forms and option menus might disappear because they seem to map natural language into Java-like structures; one unverified developer-app example involved voice while continuously deciding whether speech was a command or text input and where that input belonged.

  • The speakers leave system boundaries open. Casado flags state consistency, reliability, and strong system-level guarantees; Almeida says log analysis, email analysis, user-interface design, and human interaction will change, and he also mentions air-traffic-control systems as a possible application. Horowitz’s rule is to automate easy work before hard work.

  • Almeida’s deepest analogy is TCP: a layer that turns something unreliable into something dependable. Today, models can fail to honor JSON schemas, so developers route output to a person or another language model, creating a human-in-the-loop chatbot or an agent “while” loop. Jev aims to turn natural language into a productive state machine. Horowitz calls it an “anti-frustration machine,” while Almeida’s utopia is technology that can finally “do as I mean.”

Full transcript
Ben Horowitz

Where the hell is all that automation? AI is incredibly smart, but so useless at everything else. It doesn't matter how many AI-based coding agents you use; the software doesn't actually get better. Maybe you're just running it faster. It seems like you're getting worse. OpenAI has been trying to automate customer service since 2020.

Instead, I want smart software. I want to expand the capabilities of the software itself so that what should be automated finally becomes automated.

What I like most about your slogan is, “We build a product, not a god.” This is so cool, because if we had any other leader of a big lab, even if they were happy, they would hide it. Your view is so different. You say, “No, we will create a much better world.” Due to certain nuances, I don't think we're on the path to an AI apocalypse or anything like that.

1. Meet Diogo and Type Safe

Today we have with us the founder and leader of TypeSafe, Diogo Almeida, who is a true hero to Martin and me. Not only is he creating a really interesting product, but he's also starting what we think is a very important movement. We are very happy about today's meeting. Welcome to the show.

Diogo Almeida

Thank you.

Ben Horowitz

Maybe you could briefly explain what Jev is, what TypeSafe is, and why it is important?

Diogo Almeida

Is this a Chris-friendly format or not?

Ben Horowitz

Yes, yes, yes. Okay, cool.

Diogo Almeida

I was once asked to do an “elevator pitch,” and I tend to get carried away and I'm not good at it, but I realized that my favorite pitch for Jev is, “Where the hell is all that automation?”

It's incredibly tragic. There's so much intelligence. AI is incredibly smart, but it's not that I hate chatbots or coding agents. I adore them myself, but they are so useless at everything else, and it's just a tragedy.

It's tragic because we have such a huge diamond in the rough that's not yet ready for work. That's why TypeSafe builds AI for software. We want to make AI powerful not only for people, but also for creating real software, and Jev is, for us, the first model in this area that will significantly improve automation.

Martin Casado

Yes, it's interesting because it's already become a real hit in the programming world. One of the things that made us ask, “What the hell is going on here?” is that every developer we know calls us and says, “Oh, this is incredibly cool. It's fast. This is great. Everything is getting better.”

How is that possible? Everyone thinks, “Well, we have Cloud Code, we have Codex. Don't we already have that?” What's the difference, and how does it lead to true automation?

Diogo Almeida

I wish I had some visuals, because I have a favorite picture for this. I like Cloud Code and Codex. I like Gary Tanenbaum's description of it as just-in-time software. It's an incredible way to describe what they do. It's like creating software on the fly, and you can program it in natural language, but it has the same expressive power as regular software.

I want smart software instead. Instead of automating software development, I want to expand the capabilities of the software itself so that what can and should be automated becomes automated. In more refined language, I want to express things like intention. I want to expand the vocabulary of what we can do.

I can talk about all sorts of weird sci-fi stuff I want, but programming is about hyper-detailing valuable things and then endlessly reproducing them. It's so cool, and I just want more of it.

2. Smart software, not just faster code

Martin Casado

One way to think about it is this: instead of a tool that, to some extent, replaces the software engineer with a faster, maybe not even as good, engineer, you say, “No, no, no, we're going to give existing engineers superpowers to write much, much better and more interesting things.”

Diogo Almeida

Yes, I actually mean that. By the way, I think a lot of people miss this point, and it's so subtle and so important to get it out. If you're using something like Claude Code or Codex, which is great, or Cursor, which is also great, they write code, but that code is the same as what a human would write. Maybe it's better, maybe it's worse, but it's basically still code, the same as code from 10 years ago.

The thing about Jev is that, whether you're a Cloud Coder or a human, you get this new primitive—a new thing that you put into the code that actually extends the capabilities of the software. So instead of writing code, it's something you add to your code.

Martin Casado

Which one? Continue.

Ben Horowitz

No, you continue.

Martin Casado

What's interesting, by the way, is that it's a very powerful primitive, and it would be great if you could explain it, but it's also a little different from how programmers think. For example, there's the concept of probabilities, and therefore, perhaps—

Diogo Almeida

That is, the intelligence level within the software. Like a library where you can describe what you want in natural language, add a state machine, and it chooses what to do with a certain level of confidence, which has never been done before on this scale.

There's a lot of nuance there. I'll first jump into one thing I like about what you said—the question of where the hell all the automation is. I love software so much. I would like to write it all day. I don't recommend people to be CEOs, but I don't care.

It's strange that AI is so cool and the software hasn't changed in 10 years. For me, it doesn't make sense, and the most we can do is add a chatbot on the side that can sometimes perform actions, but not all of them, because some actions are not reliable enough. I just wanted to make that little digression. I like this moment, so I'll go back to the idea that it's a slightly different way of thinking.

Yes, I think “machine-native” doesn't exactly match the bits perfectly. This is, in fact, the art we are working on. On the first day of onboarding, I draw a Venn diagram of what AI is good at and what is valuable in code, and we are in the middle.

We don't output, for example, extrapolated floating-point numbers, because AI simply doesn't do a good job of that. But things like probabilities are nothing new, and this is like the argument over whether Jev is just a classifier.

Jev is definitely a classifier. Classifiers are cool. They were created to be useful. They are designed to be useful. In general, this is the same interface as some natural-language-processing concepts, because they came from practitioners who were trying to make systems work.

What I see now is that Jev, in my opinion, is probably better than a machine-learning team from 2019 that would do it for you, because you can program on the fly. Who knows what can be built? In 2019, there weren't many good machine-learning teams to fill narrow niches, collect datasets, measure them, and do all that. This is just the beginning. I feel like there's still a lot ahead.

Martin Casado

Do you think there's a certain slider here, where at one end there's “language at the input, language at the output,” as we have today, and on the other side there's a regular imperative program? You can move between them, or do you think this is a point in the design space—language input, machine-state output—that will become a universal tool for programmers?

Diogo Almeida

That's a difficult question. I'll say what's on my heart. Deep down, I think it's a slider. When I was designing our primitives, I might have made mistakes due to my own preferences, but right now my main guideline is intelligence per dollar. To be clear, this could be a mistake. Intelligence per second may be more valuable in the short term.

Even our interface—we call the input a “state”—was designed that way intentionally. It's the same as with a code patch. This should be an internal part of programs. Deep down, we are truly optimizing for that. Most of my work is about even more complex internal state structures in programs. Can intelligence be added there? I think it will be a constant struggle.

We take design very seriously. Pragmatically, I think certain things happen by themselves. For example, it's easier to create AI that works at millisecond scales, so for a while it will be more like a database than a standard library. But I would like to see this become a standard library as well.

Ben Horowitz

Can I step back for a second? What alchemy creates someone like Diogo? When a man and a woman love each other—

But listen, you're talking like an AI researcher. You're still talking like a systems engineer, and you're talking like a programmer. Usually these things don't overlap much, and you take the AI that we've developed as an entity and turn it into a programmer's tool. Tell me a little about your personal path, which is somewhat unconventional.

Diogo Almeida

My path in AI is somewhat unconventional. I was a mathematician. I was a mathlete. I describe it this way: I was pretty good at math—it's a little embarrassing. I was good enough at math to attract girls, so that's pretty good.

Martin Casado

No, that was the point.

Ben Horowitz

You must be pretty cool. What kind of girls do you attract when you're so good at math?

Martin Casado

This is interesting. This is something our audience needs to know.

Ben Horowitz

We have to inspire young people.

Diogo Almeida

Don't do this. Young people, don't do this. It's not worth it. Just be cool, calm, and interesting, and don't try to seem better than you are. I can't believe I said that.

I was a mathlete, but I've never actually liked math. I never made any effort. I was just a big fish in a small pond. For me, mathematics was the path I was directed toward, but I hated it because it was all about winning competitions.

Computer science is actually very similar to mathematics. It's essentially math, but cool and useful. It's fun and interesting, and I still love conducting algorithmic interviews.

Is this the best I can do? I don't know, but do I like it? Yes. And does that allow me to judge people very well? Yes, it does. So I love computer science. I consider myself a computer scientist much more than an AI researcher, despite my background.

What really got me into this was that I also won a competition on Kaggle. Not through complex math, but by automating everything with more and more nested loops. It was a systems problem, and that event ultimately led me here. For example, I was forced to speak at NeurIPS. It's usually an honor, but I hated it because I just wanted to be in the thick of things.

Ben Horowitz

Was this from that Kaggle competition?

Diogo Almeida

Yes.

Ben Horowitz

Wow, wow.

Diogo Almeida

Actually, the presenter on Kaggle was Isabelle Guillon, co-author of SVM. In fact, I think she was the first author of the SVM paper, although I'm not 100% sure who the first author was. She saw that I was someone who didn't fit into the research community at all, took me under her wing, introduced me to all the AI experts, and my career went in that direction.

Ben Horowitz

And then OpenAI?

Diogo Almeida

No, it was a startup with Jeremy Howard.

Ben Horowitz

Are you kidding? I adore Jeremy.

Diogo Almeida

Yes. Fantastic. And then Google Brain for a while. Eventually, I got tired of doing nothing. I thought, “You know what? AI is a hell of a lot of fun.” I joined OpenAI for that very reason, and it worked out really well—surprisingly well.

Ben Horowitz

So you said something that's quite unusual in today's world: AI is actually a lot of fun. Your company also has a completely different behavior and view of AI than everyone else. My favorite thing you say is, “We create a product, not a god.” It's so apt. If it were any other leader of a large laboratory, they would try to hide their joy, even if they felt it.

Your approach is so different. You say, “No, we're going to create a much better world, and it's going to be great. There won't be fewer jobs; there will be more jobs, and they'll be much better jobs. Everyone will have a great time.” Just being around you shows that you truly believe it.

So tell us about it, because for us, Safe Superintelligence is more than a company. It's a whole movement toward a positive future that most people in the AI world don't seem to like, or at least don't share.

Diogo Almeida

I think they don't understand it. You know, it's like complaining about a classifier at the level of machine-learning problems while everyone else is having a party with Jeff. It's like, hell, we can do whatever we want, and if you're not a developer, it's hard to understand what's really going on.

Ben Horowitz

I agree with that 100%.

Diogo Almeida

I really think they're painting a pretty bleak picture, which I obviously disagree with. I think it comes from this blind belief in the monolithic model that everyone believes in.

Ben Horowitz

Exactly. One big brain to rule them all.

Diogo Almeida

This is what we did. It sounds much more sinister.

Ben Horowitz

Yes, that's what people hear. Of course, that's what people hear.

Diogo Almeida

But is this one brain really on its way to ruling us all? We still haven't automated very simple things that, in my opinion, people shouldn't have to do. There are a lot of really basic things.

It hurts me when the world doesn't match reality. Part of the pain is, where the hell is all this automation? How can AI be so incredibly intelligent, and why are there so many financial incentives to automate everything? You can find an excuse for diffusion, but I don't believe it at all. It's not worth mentioning names, but this is obviously not true.

Part of the problem is that the mismatch with reality, combined with the fact that AI has enormous potential, makes it tragic to me that we haven't released it. So now it's like a little holiday for me, but I was afraid. And all the Jeff users, it's like—there's a happy AI, the people on Jeff, and there's a gloomy AI, the people who aren't there.

Ben Horowitz

Yes, yes, yes. It's actually quite a fascinating dichotomy.

I'll tell you, on your point about automation, I had an interesting conversation this morning with David George, who runs our growth fund, because we were talking about new tools. I asked, “Have you tried the Muse thing?” He's like, “Oh, she's great.” I asked, “What are you doing with her?” He said, “I finally canceled my subscription to The New York Times.” And I said, “That's hard to do.”

3. Where's all the automation?

This is just the tip of the iceberg of the terrible things we need to automate.

Diogo Almeida

I think if we want to be intellectually honest and truly strive for the “north star” of automation, we can't fall into the same anti-patterns that AI has fallen into—namely, focusing on individual use cases.

Many people ask me, “What are your favorite uses?” And I say, “I'm not sure if they work.” I want them to run in the background so that someone can trust them to run without having to check, and so that people can build something new on top of them.

And, you know, Layout, layout, right? Layout, but also security, right? This is a different type of security where if you want to run something with associated resources and access to things, you need guarantees to do so, or at least statistical guarantees that nothing will go wrong. Break it into pieces, make it fancy, something like that.

I don't think our models will do that anytime soon unless someone writes software around them. That would be a very cool achievement, and we should give credit for it. But not in a way that makes us lose responsibility.

Ben Horowitz

I'm just curious: how long did this intuition take to mature? I remember talking to you, maybe in 2017. It was a long time ago. We really talked about it, and you already had a lot of these ideas. You said that data is important, and you said that you wanted to focus on the task.

Did you know that this would all end up being a classifier, or was it just an intuition that there was a different perspective on this whole AI movement?

Diogo Almeida

There's a funny story about that conversation. During a talk in 2017, my topic was very similar. I think it was called something like “AI: Modular in Theory and Flexible in Practice,” which is very typical of software systems, so I'm sticking with that idea a bit.

I think this all actually started right before ChatGPT. When we released these things, I had no premonition. To be honest, I was very, very pleasantly surprised by the generalization capabilities of reinforcement learning from human feedback.

Ben Horowitz

When was that?

Diogo Almeida

Probably at the end of 2021—maybe the fourth quarter of 2021. It was really very general.

If you read the paper, unlike others who are trying to prove themselves right, we actually tried to disprove it using the scientific method: Is this a scam? My favorite question was, “Why is it important to eat socks before meditation?” We made sure that this was not already on the internet, and the models were able to provide plausible, human-like responses.

For us on the team, it was a moment of truth: this is not a scam. In machine learning, you should always be wary of fraud.

What really blew me away was how we released it. I'm obviously a big fan of the possibilities of AI, and I did a lot to release this model. I really thought the model had a good chance of becoming SHGI. When that didn't happen, my whole world just collapsed.

Ben Horowitz

So you were a little on the other side for a while, like that “Crazy Train” or...

Diogo Almeida

Well, no. I'm just like, no, no, no, no, no. RL generalizes. Maybe we have SHGI. Reinforcement learning from human feedback generalizes quite well. Reinforcement learning from verifiable rewards is something that doesn't generalize that well, from what I saw.

Ben Horowitz

You were behind the scenes at ChatGPT. You were behind those early GPTs. It was a completely different goal: creating a chatbot that would communicate with a human, not a programmer's tool, and so on. So I'm just curious—

Diogo Almeida

Actually, at the very beginning, sometime in 2020 at OpenAI, when we were talking about SHI, people were describing it as Elijah and every statement he made. It's not exactly like that, but we were talking about the definition of AGI at OpenAI.

Part of it is that it's intentionally blurry. It's a wide tent so everyone can be inside. But due to certain nuances, I don't think we're on the path to self-improving AI. I didn't think so before, and I don't think so now.

I really think that what OpenAI has defined as general-purpose AI is quite achievable. Automating most of the world's economically valuable labor sounds like—oh, I don't know. There is so much different work there, but most of it is very routine and simple. In terms of volume, to be able to outsource the work, you need simple instructions that anyone can follow.

As far as I can tell, the level of intelligence required for this has been in these models for quite some time. My main complaint is, “Why isn't this available yet?”

Since the advent of reinforcement learning from human feedback, the industry has been divided: big promises but weak results. I think GPT-3 was pretty balanced at the time. But when humans judge the quality of the models, they seem to be very accurate, because humans are the judges and we optimized for that judge, not for automation. That was the missing link.

I would say that's when it dawned on me: Why isn't this thing more useful?

Ben Horowitz

So you think the measure should be how much we can automate real, productive tasks? When you say “many promises and few results,” is that the dimension you're referring to?

Diogo Almeida

The ability to automate tasks. In my heart, it feels like cool science fiction, and I believe this is the first warning sign for such “cool” fiction.

Are you seriously claiming that the math is solved, or even that 2 years ago GPQA—the questions that even Google doesn't know the answers to—was solved, but we still can't handle ordering a self-driving car? This is very difficult to grasp at the same time.

4. Is it just a data problem?

Martin Casado

And I think a lot of people don't have a good answer to that. Can I test one thing? It may not make sense, but isn't there an argument that the real world is different from the digital world? It has heavy tails in its distributions; there are many outliers, and we don't have all the data.

Couldn't it be that the reason we don't do productive things in the real world is simply that we don't have data for this distribution? We don't train on it, and that's why it's all reduced to these lower-dimensional manifolds, like math or code?

Diogo Almeida

I don't quite agree with the data argument, as far as I'm concerned. I believe there is a long tail. Certainly, it would be kind of silly to deny it. But I don't think, in my situation—with the canary in the coal mine—we need to automate this long tail.

I think we need to be extremely pragmatic about everything. Creating reliable software is always an investment, right? Do you know what the three great virtues of a programmer were? Laziness, so as not to do it again, arrogance, and there was a third.

Martin Casado

No—yes, I remember. This has been around since the days of Perl.

Diogo Almeida

Yes, yes. There's a third one, dude. I would like to remember it. But laziness means spending 10 hours automating a 5-minute task once and never coming back to it again.

This should be a solution based on return on investment for people who automate things. I just wish it could be automated.

Martin Casado

Yes. And I think people will just create new kinds of work, hence the Jevons paradox, when that becomes possible. But as a benchmark, I feel it's useful to see if we can actually automate what AI seems like it should be able to automate.

OpenAI has been trying to automate customer service since 2020. It's not that simple. It's pretty amazing. It's just crazy.

Diogo Almeida

Well, within companies, very little is automated right now. The projects didn't work.

Martin Casado

Yes, apart from programming, which works great. Can you perhaps categorize the types of problems that you think are easier to automate? This is quite interesting.

Diogo Almeida

We were actually looking at support even before the current wave of generative AI. It was interesting to meet the company, and the company said, “We respond to 95% of all support requests.”

And I'm like, “That's so much.” But then you really look at the data and understand that it's all about resetting passwords. If you were judging by uniqueness, it was 50% or something like that.

So it seems that when you deal with people and natural systems, there is a very long tail of exceptions.

Martin Casado

To what extent did you even foresee such a wide range of use cases? I probably get a message every hour that someone is using Claude for a new task. I'm like, “I had no idea.” Did you expect this to happen? Are you surprised by it?

Diogo Almeida

Extremely surprised. I didn't expect this to happen. This launch wasn't something that anyone would have expected. They'd probably have to be crazy.

I don't think anyone would have expected ChatGPT for developers, because ChatGPT was for regular users, and that's weird. I don't even really know what percentage of people who joined Jeff's party are developers. I can't imagine anyone other than developers using it. I don't know how they would use it.

But even my friends who aren't developers are just part of this Twitter party, joking around and all that. So, number 1, it's phenomenal.

5. Reliability and "do what I mean"

Number 2, this is going to be hard to convey in a short message because it was blood, sweat, and tears over many years. The extent to which I care about reliability is a lot. Reliability is the essence of this thing. If you don't understand this, it will be very difficult to create an analogue that meets the criteria.

It seems to me that every percent of reliability will be valuable to everyone, even if it's not the most valuable thing in terms of market capitalization, because it will simply open up new opportunities. We're fighting for all sorts of strange levels of reliability that we don't even fully understand, because we're just trying to introduce this “electric motor” of AI—of intelligence—into people's workplaces, and they'll decide what to do with it.

Martin Casado

What does reliability mean in this context? Is it the availability of the model, or is it that I call the model and it returns the same result? How should I even think about reliability for something that is inherently stochastic?

Diogo Almeida

It's not so much the first; the second is closer to the truth. I would describe the first as uptime or an SLA. I would call the second one closer to determinism, and the third one I would consider more like reliability.

I would describe reliability as the same intelligence every time.

Martin Casado

Oh, interesting.

Diogo Almeida

Yes. It's not exactly determinism, because I believe determinism is useful for unit tests, but not for real systems. Think about it: if you add a UUID to the query, the result should be the same because, functionally, it's the same.

Martin Casado

Yes, but it is not entirely deterministic.

Diogo Almeida

Right. I think there is another level, the name of which I don't know yet. Maybe it's what I would call a certain form of intelligence, where the result doesn't have to be identical, but it has to be intelligent every time.

If you were in that situation, would that thought be understandable to a human being? The developer can take this into account in the code.

Really, for me, the ultimate measure of reliability would be to reach a level where people can program with Jev without creating example queries. When you just trust it, you're in a constant state of flow, creating great, incredible software.

Martin Casado

And many programs are at this stage right now, aren't they?

Diogo Almeida

I don't think it's entirely unrealistic, and I'm perfectly fine with it.

Martin Casado

By the way, this is a bit of a strange question, so feel free—if it's too strange, just don't worry about it. But it occurred to me that the value of things like coding agents actually diminishes if you have a primitive like that.

You could say, “Whatever. Codex builds all the software for me, but it doesn't use Jev, so the software it builds is somewhat limited.” Or you could say, “As a human, I'll write a program without using a coding agent, but I have this very universal primitive that makes writing easier.”

Do you see a future where coding agents use Jev and you manage them, and do you have any redundancy? Or do you think people are implementing it themselves? This is more of a question about coding agents than it is about Claude.

Diogo Almeida

Oh, yes, definitely. I feel like I'm not as immersed in the intricacies of programming as I would like. You two may know better than me about this, and that's unfortunate, but my experience is that coding agents are very good at syntax and really bad at semantics.

I would say they're incredibly bad at architecture. I think architecture is the most human, creative part of software development. That's why I like using AI agents for coding.

I think this is almost certainly not part of the data distribution. It would be creepy if they were learning from our users' data. So most likely, that's not the case. But I don't see a problem with instructing them on syntax when it's part of the distribution.

Martin Casado

The thing is, the models may not just be terrible at architecture; they may be at the 50th-percentile level. And if you yourself don't know anything about architecture, then this will be enough.

These are all gray areas and compromises that you will have to navigate. Sometimes speed is exactly the leverage your company or project needs. You're willing to choose a 50th-percentile architecture over a 60th-percentile one because you want to move faster, so Codex can run all night or something.

Ben Horowitz

Yes, by the way, there's an interesting phenomenon in the market. When AI agents appeared, a “SaaS apocalypse” occurred and their value fell sharply. And when Claude came along, every SaaS company said, “This is the best thing in the world.”

So explain this.

6. SaaS apocalypse, reversed

Diogo Almeida

I don't even know what else to say. It seems quite natural to me. In this story of the SaaS apocalypse, the idea that, in my opinion, has not been fully realized is that software is very cheap and easy to copy.

I can believe the first part, but not the second, because a lot of things happen under the hood. Maybe I'm too fanatical about software.

Ben Horowitz

We all are.

Diogo Almeida

Yes, okay, okay. I didn't know where this code could be. We have a lot of heritage in this regard.

So I don't think it really worked. The markets seem to disagree, but in my opinion, SaaS provides the same value as before. Maybe the markets are just scared.

But I think SaaS is going to be one of the biggest winners in this whole AI game. And I want to work really, really well with all those big, boring, user-aware SaaS companies because, in my opinion, they know best what workflows are worth automating.

What do people need? This is their main bread and butter. They spend big money because software is always a capital investment, but you invest it up front to make this experience even better, and then it spreads to the entire huge user base.

Ben Horowitz

True.

Diogo Almeida

So I think it's going to be—I’m not going to make any predictions about the financial markets—but in terms of opportunity, it's going to be like a reverse SaaS apocalypse, and I'm just excited about it. I need to come up with a name.

Ben Horowitz

Yes, you should have a name.

Diogo Almeida

SaaS-appaloosa.

Ben Horowitz

That sounds a little too much fun.

Diogo Almeida

All SaaS applications will suddenly become much more useful. And by the way, a significant portion of a SaaS company's capital investment goes into actually reaching all its customers.

So if you've already covered all your customers, and then you've not just made a chatbot in your SaaS product but have really, significantly improved the software itself...

Hmm, this is incredible. I don’t know if this is a realistic dream, but I think there’s a world where forms with options to choose from simply disappear. It seems to me that they just map natural language, which the software already understands, into Java-like output.

Ben Horowitz

It’s literally been around since the 1980s. We used to call it 4GL. Remember the MR2? Fourth generation language.

Martin Casado

Yes. I also think it’s from the ’80s. It might even be an insult. I wasn’t born yet.

Diogo Almeida

I think the concept of “do what I mean” will reach a whole new level. If I can mention one developer app, I don’t know how reliable it is, so I can’t promise anything, but it was so cool. Someone was using their voice to control a computer, and they were constantly making decisions: Is this a command, or is this text input? Where is this text input?

This sounds incredibly cool. I think the interfaces might just change completely, and we might have to do it cheaper and faster.

Ben Horowitz

Yes, then you’re already in Star Trek.

Diogo Almeida

There’s a deep intuition here that, if you’re using AI to build software today, you’re still creating the same software as before. But if you look at the average pull request for a large company, it’s about 10 lines, right? Seriously, it really is. I was at Google, and we did some research. It’s about 10 lines.

You automate 10 lines. And, by the way, those 10 lines are part of the learning from the client or something, so you’ve optimized something that’s actually pretty minimal. What it doesn’t do is give the software new capabilities. It’s like automating something that ultimately turns out to be relatively insignificant.

Now truly new opportunities have emerged, and it’s entirely possible that the software will simply get better. It never even occurred to me before that no matter how many AI agents you use to code, the software doesn’t actually get better. Maybe you write it faster, but are you really getting worse simply because there’s less supervision?

Martin Casado

So I think it’s often less safe.

Diogo Almeida

Yes, that’s right, without a doubt. But you can now argue that applications will have new functionality because of this, because you’re providing this new primitive. In a sense, it speaks natural language—it can reason—but it combines that with a state machine.

7. New capabilities, not more code

If people took that as the main takeaway, it would be the biggest compliment to what we do. I even think it’s almost too grand a vision: going beyond the 3 logic gates that we have into something where our types are like the same logic gate, but with a little brain inside.

That would be the best compliment to the legacy of type safety, because it’s a huge, nontrivial thing for the world. I’m not going to promise more than I can deliver, but I will fight for it.

Martin Casado

I think there are still open questions about how deep this can go in serious areas like state consistency, reliability, or a real system level where you need to provide strong guarantees. This will 100% change things like log analysis, email analysis, user-interface design, and human interaction.

Diogo Almeida

That’s true, but you could argue that, over time, it will become something like a smart database. And also air-traffic-control systems—anything, right? What we really need is a little scary.

Ben Horowitz

I think my philosophy is to automate the easy work before the hard work.

Diogo Almeida

I also think that a whole era of probabilistic programming will open up. By the way, there’s a huge history of probabilistic programming, which actually died in the ’70s. I’m familiar with this.

Martin Casado

I actually think it’s similar to what you could call neurosymbolic AI. Your co-founder Eric came from this very field. He told me.

Diogo Almeida

Oh, cool. Yes, he studied biology a lot.

Ben Horowitz

What I mean is, I’m not a fan. My brand is pragmatism—extreme pragmatism. I’ve never been a fan of biologically inspired things at all. Have you ever noticed this? I mean, I don’t think they ever worked.

It’s useful for motivating crazy people to work on something for decades until it works, and then they refine it into an engineered version. That’s the history of AI today: neural networks, for sure. But a lot of the stories about how it worked were inaccurate, right?

The hierarchical features of the applications didn’t really work, because otherwise ResNet wouldn’t work. Long story. I really think it’s promising, but from a systems perspective, I’m not excited about that part. I’m excited for the world—not that I’m going to program it—because it sounds really complicated.

But I think that when we have a lot of intelligence with different trade-offs in cost and speed, system specialists will make choices where developers will be 1,000 times smarter than them. They just need a rough reference to make a rough guess and optimistically push the data here and there.

It’s going to be incredible, these capabilities that will be available in extreme systems.

Martin Casado

The good news is that we’re going to be able to rebuild systems again, and that’s great, right? We have a new primitive. No, seriously, we have a new primitive—a new way of thinking about building software.

Ben Horowitz

We’ve done this before. We did it with the internet. We went from mainframes to client-server architectures. We do it periodically. And, by the way, because of cybersecurity issues, we’ll probably have to rebuild almost all systems to make them secure.

Diogo Almeida

I think so. I think it’s pretty obvious that there’s no—or at least no critical infrastructure here, for sure.

8. Apps vs the guts of systems

Martin Casado

Do you think of it more as applications, SaaS, and analytics, or more as systems fundamentals—or all of that together?

Diogo Almeida

Do you mean what I’m thinking about, or a general application for this thing?

Martin Casado

When you think, for example, about working on Jev and imagining how people would adapt it, do you even have any idea?

Diogo Almeida

I have a little bit of an idea. I think of it a little bit as diving into the very depths of TCP. You know how TCP is a transition from unreliable to reliable? That’s my analogy.

When I think about AI—and that’s why I’m interested in intelligence per dollar, just to be clear—I came to this conclusion by going from the opposite direction, from an AI-based economic revolution. AI is everywhere in science fiction and all that. It’s as if all software has AI everywhere.

I ask myself: What percentage of calls to AI—imagine that’s a function—are intended for human consumption, where style is needed and all that? That will be a lot of nines. And, from the same question, how many will be at the first level, and how many will be somewhere deep in the jungle?

I think there will be a lot of nines in the slums, but it will all start from level 1. If you’re not aiming for the jungle—wait, that’s kind of weird. If you don’t aim for the very essence, it will take you a long time to reach your goal.

People don’t realize how much AI has been built into software in secret. Even if you tried to build AI into the software, it behaved unpredictably, because software doesn’t understand natural language very well. You do all these strange things: You put in the prompt, “Here is the JSON output I need. Here is the schema,” and it never listens to you.

As a result, you just had to take that output and give it to the person. You say, “Well, to hell with it, right?” Or you give it to another language model. That’s the “while” loop—the agent’s “while” loop, right?

Going from the basics, it should be either a human in the loop, which is a chat, or an agent, which is a “while” loop, because natural language needs to be routed back into the next loop.

Martin Casado

I’ll put it this way: I watched it, and it was like the 5 stages of grief. People take AI and say, “I’m going to use this in my software.” Then there’s denial, trying to make it work, anger, and then they move to acceptance: “Okay, never mind. I’ll just give it to another language model or a person.”

So it was all very chaotic. I think this is the first time I’ve seen someone say, “You can take a language model, take AI, and actually turn it into a state machine,” and do it productively.

Diogo Almeida

I hope so. I also don’t want to overpromise and underdeliver on this. I don’t know if it’s ready for all the uses it’s been so loudly promised for. I really, really want this, and my team will, of course, fight for it. We really care about reliability.

We could have released this much sooner. I don’t think people understand this, and honestly, I don’t think they will. Judging by the Twitter discussions, people may never understand it intellectually, but they’ll just have a feeling of, “Oh, I can trust this.”

Ben Horowitz

This is the anti-frustration machine.

Diogo Almeida

I hope so. I hope “do as I mean”—for me, it’s about fluidity in the world, so that everything moves more smoothly and connects like gears. I have my own utopia with artificial intelligence in various directions, and “do as I say” is a huge part of it.

Imagine if all technology just did what you meant. This isn’t science fiction. Look how smart artificial intelligence is, right?

Ben Horowitz

Yes, it’s amazing, and perhaps that’s where we should end: “Do as I say.”

Diogo Almeida

Yes, I like it.

Ben Horowitz

Thank you, Diogo. It was a great conversation. I really liked it.

Diogo Almeida

Mm-hmm.