[BidClub_]
Latent Space · · 62 min

Unsupervised Learning x Latent Space Crossover Special

JacobswyxAlessio FanelliJordan

Podcast
TL;DR
  • The episode’s strongest investor call is that AI value is shifting above the model layer, where applications can charge for utility instead of cost-plus infrastructure. swyx says the industry went from mocking GPT wrappers to treating wrappers as “the only thing that’s interesting,” with coding, support, and deep research as leading proven forms; conversation summarization is a strong use case but not necessarily an agent. Product experience, distribution, network effects, and execution speed now matter more than owning a bespoke model.

  • Reasoning models reopened scaling just as pre-training seemed tapped out, but DeepSeek showed how quickly proprietary model advantages can compress. swyx argued that R1’s full reasoning traces were net-unique in open source, yet the panel separates DeepSeek’s achievement from a durable “team open source”: enterprise open-model usage was estimated at roughly 5% and going down, while one DeepSeek-driven narrative coincided with about a 15% one-day decline in NVIDIA. “There’s just different companies, and they choose to open source or not.”

  • Agent frameworks look overbuilt for workloads still in flux; protocols and state are the more durable bets. The panel calls today the “jQuery era” and argues MCP may matter more as an XMLHttpRequest-like protocol than another all-encompassing framework. Memory is the underappreciated layer: without persistent knowledge beyond context windows, agents are less able to become smarter or learn on the job.

  • Current product-market fit is concentrated in coding, customer support, and deep research, with the next wave aimed at revenue creation. swyx estimates OpenAI’s $20-to-$200 deep-research upgrade could represent billions in ARR; when Jacob asked whether access had quickly dropped to $20, swyx clarified that access had been expanded. Voice products can create value under imperfect reliability: if home-service businesses miss 50% of calls, an AI effective 75% of the time can still create substantial revenue.

  • Coding is becoming the first full-stack battle among model vendors, IDEs, and autonomous agents. Cursor’s roughly $9 billion-$10 billion valuation makes acquisition harder, while Anthropic, Cognition, and others may converge across “inner-loop” IDE work and “outer-loop” cloud agents. The remaining moat is operational schlep—codebase integrations, security approvals, infrastructure, and enterprise trust—not merely model quality.

  • App defensibility is conventional but compounding: network effects, brand, integrations, and relentless velocity beat “train your own model” stories. Chai’s marketplace between users and model providers is swyx’s network-effect example; elsewhere, category leadership has reportedly doubled ACVs even where investors expected pricing compression. Jacob’s framing is “the thousand small things” that make a product delightful and let it exploit every new model within three to six months.

  • The largest technical fork is whether reinforcement learning can move beyond verifiable domains. If RL works for coding and math but not law, sales, or marketing, the result could be fully autonomous technical agents alongside human-dependent creative copilots: “to write the most basic sales email,” a person may still be the tastemaker. Bob McGrew’s reported reliability rule also leaves an infrastructure question: each move from 90% to 99% to 99.9% was described as requiring another order of magnitude of compute.

  • The explicit public-market pair was long Google and short Apple, although Jordan sees Apple’s Private Cloud Compute as an important exception. Google’s models are gaining practical usage—Flash reportedly wins most days in swyx’s daily frontier-model bake-off—but fragmented developer surfaces still impede adoption. Apple Intelligence disappointed badly, while PCC may help bring on-device-style privacy to cloud workloads; Jacob argues that AI will need single-tenant guarantees in multitenant environments.

Digest · the substance, structured for research

1. Reasoning rescued scaling, but DeepSeek compressed the moat

  • The opening surprise was the abrupt transition from Ilya Sutskever’s NeurIPS-era “scaling is dead” mood to reasoning models: “It’s so over” became “we’re so back” almost immediately, with inference-time compute appearing to replace pre-training as the scaling frontier.

  • swyx found the timing “suspiciously neat.” Strawberry had reportedly been underway for roughly two years, and Noam Brown’s arrival at OpenAI was presented as evidence of a major strategic bet: the work was ready just as pre-training appeared tapped out. “If there were an Illuminati, this would be what they planned.”

  • Open models present a paradox. Ankur from Braintrust estimated enterprise usage at about 5% and declining. Jacob’s explanation was that enterprises remain in use-case-discovery mode and keep choosing the strongest available model; swyx said DeepSeek showed him how short a model company’s exclusive product-building window could be.

  • The disagreement on DeepSeek is worth preserving: Alessio called replication “an order of magnitude cheaper” than invention and therefore overhyped, while swyx argued R1’s full reasoning traces were net-unique in open source. His harder correction: “There’s no team open source,” especially if DeepSeek stops releasing models and everyone else merely distills from it.

2. Legacy product surfaces caused incumbents to miss the AI-native turn

  • swyx’s surprise was that Zapier, Airtable, Retool, and Notion failed to capture AI software creation while Bolt and Lovable did. The incumbents had distribution and low-code DNA, but used natural language to improve existing actions, documents, and tables instead of rebuilding software creation “from whole cloth.”

  • Timing may have mattered as much as imagination: Bolt and Lovable reportedly went from $0 to $20 million in three months as models crossed a usability threshold. Their numbers transformed an architectural curiosity into a market investors could no longer ignore.

  • Jordan’s negative incumbent case was Apple Intelligence. Apple appeared ideally positioned to build a personal assistant, yet produced viral-but-bad message summaries and a BBC-linked false summary saying a man had shot himself when he had not: the consumer product fell far below the panel’s expectations.

  • His Apple exception was Private Cloud Compute, which carries on-device-style privacy into larger cloud workloads. Jacob’s mechanism: GPUs make dedicated VPC deployments impractical, so AI needs “single-tenant guarantees in multi-tenant environments”—an architectural opening as consumer AI handles more sensitive data.

3. Protocols and memory look more durable than agent frameworks

  • The framework critique is not that LangChain lacked velocity—it shipped relentlessly as abstractions changed—but that workloads remain too unstable for developers to choose a framework they can trust for years. The panel’s analogy: agents are still in a “jQuery era,” while the market keeps building more jQuery before discovering its React.

  • swyx’s modification was that even React may be premature: “The thing is the protocol and not the framework.” MCP resembles XMLHttpRequest, the enabling primitive behind AJAX, because a protocol can survive shifting implementations without freezing today’s agent assumptions into a monolith.

  • Memory is the missing agent primitive. OpenAI’s Agents SDK and Cloudflare’s agent work omitted persistent “memory memory”—facts, preferences, and knowledge extending beyond a context window—even though MCP’s initial servers included an implementation swyx recommends as a starting point.

  • Alessio pushed back that memory may be difficult rather than ignored, citing active work around LangMem, Letta, and Zep. swyx acknowledged the difficulty and cited those efforts, but kept the investment call: stateful abstractions should become standard, and “anything stateful should be interesting to VCs” because the business model resembles databases.

4. General models are squeezing new labs toward proprietary data

  • Jacob remains surprised by the number of new companies training broad models; swyx had thought that category died at the end of the prior year. The basic objection is strategic: every entrant wants AGI and benchmark leadership, but “they can’t all do it,” and no missing general capability clearly demands another lab.

  • Alessio’s reading of Noam Shazeer’s answer was that test-time compute may eventually let models perform AI engineering and discover the next algorithmic breakthrough themselves. That leaves little room for undifferentiated challengers unless they possess a genuinely different method.

  • The plausible exception is unique data generation: robotics, biology, and materials companies can gather specialized data through physical systems, wet labs, and experiments. Even there, the unresolved “bigger is better” question persists—smaller vertical models may be cheaper, while the next general model may simply outperform them.

  • BloombergGPT supplied the cautionary example. Bloomberg concluded closed models were better and stopped the model effort, but retained the valuable assets around it: its data pipeline, assembled team, fine-tuning capability, and roughly 12 or 13 teams applying generative AI across the company. “Everything but the model survived.”

5. Coding is becoming the first full-stack model-versus-app battle

  • Model companies are moving into products partly because expensive research creates near-term monetization pressure, but every move creates a new enemy. Search threatens Perplexity; deep research crowds application vendors; coding products raise the question of whether Anthropic remains Cursor’s supplier or becomes its competitor.

  • Cursor’s roughly $9 billion-$10 billion valuation has made it “hard to acquire,” in swyx’s view; at $5 billion-$6 billion, joining OpenAI and contributing coding data might have been plausible. Now it likely must remain independent while diversifying away from any single model provider.

  • Alessio argued Devin may be more exposed than Cursor because Anthropic could pursue autonomous agents without building an IDE. swyx’s pushback was Claude Code, while Jacob expected both sides to converge: IDEs will add agents, and agents will expose interactive coding environments.

  • The moat differs by loop. Inner-loop tools operate inside the IDE and a Git commit; outer-loop agents work between commits and must integrate codebases, satisfy security requirements, and endure enterprise “schlep.” That gives Cognition a possible moat—and perhaps favors Sourcegraph, whose decade of contracts and codebase trust is difficult for two-year-old startups to reproduce.

6. Product-market fit is real, narrow, and expanding

  • swyx originally set a $100 million bar to rebut claims that generative AI products were unreliable toys. His examples were Copilot, Jasper’s writing category—with the caveat “Jasper—no longer”—and Cursor as a coding agent.

  • Deep research is the newest validated form: long-running agents produce reasoned reports across OpenAI, Gemini, Grok, Perplexity, and vertical products such as Brightwave for finance. swyx estimates OpenAI’s $20-to-$200 upgrade could mean billions in ARR; when Jacob asked whether access had been dropped to the $20 tier, swyx clarified that access had been expanded.

  • Customer support joins coding and deep research as a “killer agent” category. Alessio reads Brett Taylor’s choice to build Sierra—despite his enthusiasm for developer tools and ability to fund almost anything—as strong evidence that support contains durable budgets, alongside Decagon and similar entrants.

  • Google is becoming the usage counterweight. Gemini handles swyx’s YouTube summaries, native image generation, and thinking workloads; Flash “wins most days” in his manually run daily model bake-off. Google Cloud, Vertex, and AI Studio fragmentation still hurts developers, but both swyx and Jacob chose the same trade: long Google, short Apple.

7. The next application wave must move from savings to revenue

  • The first wave sold cost reduction into work companies already outsourced to BPOs, especially support. Alessio sees the narrative shifting with Jensen Huang’s GTC wording: last year, “the more you buy, the more you save”; this year, “the more you buy, the more you make.”

  • Jacob’s concern is that outsourced cost centers may also face the fiercest price competition because buyers already accept lower performance for savings. Revenue-producing tools could prove more defensible: a go-to-market product that reliably creates incremental sales can charge more because its ROI is visible.

  • Voice AI illustrates why imperfect systems can already expand revenue. Home-service businesses such as electricians reportedly miss 50% of incoming calls; an agent effective only 75% of the time still captures substantial demand. The next bottleneck then appears: can the electrician hire, train, and deploy enough people to fulfill twice as much work?

  • The proposed next forms include screen-sharing assistants, outbound sales, hiring, finance, personal AI, and focused education products. swyx’s education lesson from trying to make Python a Singaporean national language was institutional resistance—teachers were unprepared—so he favors narrow products such as Speak over attempts to rebuild education all at once.

8. Durable app moats are networks, brands, and compounding execution

  • swyx’s clearest answer on app defensibility was “network effects.” Chai, a Character.AI competitor, lacks a proprietary-model advantage but connects users with outside model providers; that marketplace may hurt near-term differentiation while making the company robust to future model changes.

  • Brand is already producing unexpected economics in enterprise categories. Where investment committees predicted commoditization, pricing compression, and shrinking ACVs, Jordan says some ACVs have doubled because customers pay a premium for the category-defining company already invited into every buying process.

  • Jacob called unique datasets and self-trained models an early “head fake.” The real defense sounds less magical: “the thousand small things” behind delightful UX, broad product coverage, and organizational velocity. Every model release creates an existential three-to-six-month race to exploit new capability before a competitor does.

9. Agent infrastructure matters, but applications capture the richer margin

  • The attractive “LLM OS” layer sits around models: E2B-style code execution, memory, search, security, and other tools agents require. It is distinct from bare-metal serving; the application layer has nevertheless been “way more interesting” because vendors charge for customer utility while infrastructure often collapses toward cost plus.

  • Alessio’s security thesis is symmetrical: wherever AI improves offense, defense must respond in email, identity, red teaming, and binary inspection. Models add semantic understanding to historically syntactic rules—reasoning about what code is trying to accomplish, not merely matching isolated instructions.

  • Research labs will absorb some infrastructure. swyx’s heuristic is that every checkbox in ChatGPT’s custom-GPT builder represents a startup, and those checkboxes are becoming APIs; OpenAI Search was compared with Axon at roughly five times the price for the same amount of research, leaving room for multi-provider strategies.

  • swyx struggles with standalone fine-tuning, AI DevOps/SRE, and real-time voice infrastructure as venture-scale categories. Jordan disputed the SRE skepticism: full autonomy is distant, but even a 10% reduction in mean time to recovery can support meaningful revenue. swyx acknowledged the incremental case while maintaining that basic anomaly detection predates LLMs.

10. Reward, reliability, silicon, and identity remain unresolved

  • swyx’s largest question is whether RL can succeed outside verifiable domains. Coding and math have checkable rewards; law, contracts, marketing, and sales require taste. Failure would divide the market between autonomous agents in verifiable work and copilots elsewhere, producing the odd world of automated software but human-written sales emails.

  • Bob McGrew’s reported “rule of nines” makes reliability a compute problem: moving from 90% to 99%, then 99% to 99.9%, costs an additional order of magnitude each time, with the progression described as taking roughly two to three years. That raises both hardware availability and NVIDIA’s continued dominance as ecosystem-level questions.

  • swyx sees a theoretical opening for transformer-specific silicon because GPUs remain general enough for gaming, crypto, and AI. The bet requires stable workloads—and therefore confidence that transformers persist—but post-2019 or 2020 entrants can specialize more aggressively than chip companies designed before transformers had clearly won.

  • The emergent protocol problem is agent authentication: when Operator or another agent visits a site, it must indicate that it is an agent acting for a user rather than the user themself. swyx calls this a necessary “new SSO effectively for agents,” then allows the uncomfortable possibility that Sam Altman’s answer was early: “Maybe you have to scan your eyeballs.”

Jacob

Today, my partner Jordan and I have a special episode of Unsupervised Learning, a crossover with one of our favorite AI podcasts, Latent Space. If you're not already a listener, Latent Space is a technical newsletter and podcast by and for AI engineers. It had over two million downloads in 2024, and it's become a go-to resource for anyone who wants to understand the cutting edge of AI infrastructure, tooling, and product. If you like this show, it's definitely worth checking out. Given we've all spent a lot of time talking to some of the sharpest minds in AI, we thought it'd be fun to interview each other. In this episode, we dig into the questions we're constantly thinking about, what surprised us most last year, what we're paying most attention to right now, how we think about defensibility at the app layer, and which public companies we're long or short on. It's a different kind of episode, and I think you'll really enjoy it. Now here's my conversation with Swyx and Alessio from Latent Space.

Well, thanks so much for doing this, guys. I feel like we've been excited to do a collaboration for a while.

swyx

I love crossovers.

Alessio Fanelli

Yeah.

swyx

Yeah, this is great. The ultimate meta about podcasters talking to other podcasters.

Jacob

Yeah, it's a lot. Podcast all the way up. I figured we'd have a pretty free-ranging conversation today, but I brought a few conversation starters to kick us off.

1. Reasoning And Open Models

One interesting place to start is that, obviously, it feels like this world is changing every few months. As you guys reflect on the past year, what surprised you the most?

Alessio Fanelli

I think definitely reasoning models. We're kind of on the right here. We're causing—

swyx

Oh, open source.

Alessio Fanelli

Oh, that... I think there's two. There's what surprised us in a good way and maybe in a bad way, I would say. In a good way, reasoning models, and I think the release of them right after the NeurIPS “scaling is dead” talk by Ilya. I think there was maybe a little “it's so over,” and then “we're so back” in such a short period of time.

Jacob

It was really fortuitous timing: right as pre-training died. I mean, obviously, I'm sure within the labs they knew pre-training was dying and had to find something, but from the outside, it felt like one right into the other.

Alessio Fanelli

Right.

Jacob

It felt like one thing immediately followed the other.

Alessio Fanelli

Yeah, exactly. That was a good surprise.

swyx

If you want to make that comment about timing, I think it's suspiciously neat. We know that Strawberry was being worked on for 2 years or so, and we know exactly when Noam Brown joined OpenAI. That was obviously a big strategic bet by OpenAI.

For it to transition so nicely, when pre-training is kind of tapped out, into “Oh, now inference time is the new scaling law,” is very convenient. If there were an Illuminati, this would be what they planned.

Jacob

Or if we're living in a simulation or something. Then you said open source as well?

Alessio Fanelli

Yeah. Well, no, I think open source. We were discussing this on the negative. I would say—

swyx

Specifically, open models.

Alessio Fanelli

Yeah, I was surprised by the lack—

swyx

Like the Llamas of the world.

Alessio Fanelli

I was surprised by the lack of adoption. I mean, people use it, obviously, but I would say nobody's really a huge fanboy. I think the local Llama community and some of the more obvious use cases really like it, but when we talk to enterprise folks, it's like, “It's cool.”

I think people love to argue about licenses and all of that, but the reality is that it doesn't really change the adoption path of AI.

swyx

The specific stat that I got from Ankur, from Braintrust, in one of the episodes that we did, was—I think he estimated that open-source model usage in enterprises is at 5% and going down.

Jacob

It feels like basically all these enterprises are in use-case discovery mode. It's like, “Let's just take what we think is the most powerful model and figure out if we can find anything that works.”

So much of it feels like discovery, and then right as you've discovered something, a new generation of models is out, and so you have to go do discovery with those. I think, obviously, we're probably optimistic that open-source models increase in uptake.

My biggest surprise last year was open-source related, but it was just how fast open source caught up on reasoning models. It was kind of unclear to me, over time, whether there would be a compounding advantage for some of the closed-source models. In the early days of scaling, there was a tight time loop, but over time, would the gap increase? If anything, it feels like it shrunk.

DeepSeek specifically was just really surprising. In many ways, if the value of these model companies is—

swyx

You have a model for a period of time, and you're the only one that can build products on top of that model while you have it. God, that time period is much shorter than I thought it was going to be a year ago.

Yeah. I don't like this label of how fast open source caught up, because it's really how fast DeepSeek caught up. Now we have some evidence that DeepSeek is basically going to stop open-sourcing models.

Jacob

Yeah.

swyx

There's no team called Open Source. There's just different companies, and they choose to open-source or not. We got lucky with DeepSeek releasing something, and then everyone else is basically distilling from DeepSeek.

Distillations catching up is such an easier, lower bar than actually catching up. You're training something from scratch that is competitive. On that front, I don't know if that's happening. Basically, the only player right now is—we're waiting for Llama 4.

Alessio Fanelli

I mean, it's always an order of magnitude cheaper to replicate what's already been done than to create something fundamentally new, and so that's why I think DeepSeek overall was overhyped. Obviously, it's a good open-source new entrant, but at the same time, there's nothing fundamentally new there other than executing what's already been done really well.

Jacob

Yeah.

Alessio Fanelli

Right? So—

swyx

Well, but I think the traces are maybe the biggest thing. I think most previous open models were the same model, just a little worse and cheaper.

Alessio Fanelli

Yeah.

swyx

R1 is the first model that had the full traces, so I think that's a net-unique thing in open source. But, yeah, I think we talked about DeepSeek in our end-of-year 2024 recap, and we were mostly focused on cheaper inference. We didn't really have DeepSeek—

Jacob

DeepSeek-V3 was out then—

swyx

Yeah.

Jacob

—and we were like, “That was already talking about fine-grained mixture-of-experts and all that.” So, very—

swyx

That's a great receipt to have, to be like, “Yeah, end of year '24.”

Alessio Fanelli

We were there. Let's go.

Jacob

That's an impressive one.

swyx

You follow the right-whale believers on Twitter. It's pretty obvious.

I actually had a lot of my hedge fund and private-equity friends call me up. They were like, “Why didn't you tip us off about DeepSeek?” And I'm like, “Well, it's been there.” It's actually kind of surprising that NVIDIA fell, what, 15% in one day because of DeepSeek.

I think it's just whatever the public-market narrative decides is a story becomes the story, but really, the technical movements are usually 1–2 years in the making before that.

Jacob

Basically, these people were telling on themselves that they didn't listen to your podcast, because it was in the end-of-year 2024 recap.

swyx

No, no, we weren't banging the drum. It's also on us to be like, “No, this is an actual tipping point.” Our function as podcasters and industry analysts is to raise the bar or focus attention on things that you think matter. Sometimes we're too passive about it, and I think I was too passive there.

I'd be happy to own up to that.

Jacob

No, I feel like over time, you guys have moved into this more extreme role of taking stances on things that are or aren't important. I feel like you've done that with MCP of late—

And a bunch of things.

2. AI Engineering Moves Above Models

swyx

Yeah. So, the general push is AI engineering—you’ve got to rep the shirt. MCP is part of that. But the general movement is: what can engineers do above the model layer to augment model capabilities? It turns out it’s a lot. And it turns out we went from making fun of GPT wrappers to now, I think, the overwhelming consensus is that GPT wrappers are the only thing that’s interesting.

Jacob

Yeah. I remember Arvind from Perplexity came on our podcast, and he was like, “I’m proudly a wrapper.”

swyx

Yeah.

Jacob

Anyone who’s talking about differentiation pre-product-market fit is a ridiculous thing to say. Build something people want, and then over time you can worry about that.

swyx

Yeah. I interviewed him in 2023, and I think he may have been the first person on our podcast to proudly be a GPT wrapper.

Jacob

Yeah.

swyx

And obviously, he’s built a huge business on that.

Jacob

Well, now we all can’t get enough of it.

Yeah, on that.

swyx

Yeah. So that was Alessio’s one, and we prepped individual answers just to be interesting.

Jacob

In the same Uber on the way up.

swyx

Yeah.

Jordan

Oh, I was driving too.

Jacob

Oh, you were driving.

Oh, wow.

swyx

I mean, it was a Tesla.

Jacob

Glad you made it here okay.

3. AI Builders Escape Low Code

swyx

Mine was actually—it’s interesting that low-code builders did not capture the AI builder market, right? AI builders being Bolt and Lovable; low-code builders being Zapier, Airtable, Retool, Notion—any of those. When you’re not technical, you can build software.

Somehow, all of them missed it. Why? It’s bizarre. They should have the DNA. They already have the reach and the distribution. Why? I have no idea.

Jacob

They also had the ability to fast-follow. I’m surprised there wasn’t more.

swyx

Yeah, there’s just nothing.

Jacob

Yeah. What do you make of that?

swyx

It seems—and not to come back to the AI engineering PhD—it takes a certain kind of founder mindset or AI engineer mindset to be like, “We will build this from whole cloth and not be tied to existing paradigms,” I think.

When Zapier decided to do Zapier AI, they were like, “Oh, you can use natural language to make Zapier actions,” right? When Notion decided to do Notion AI, they were like, “Oh, you can write documents or fill in tables with AI.” They didn’t do the next step because they already had their base, and they were like, “Let’s improve our baseline.”

The other people who actually tried to create it from whole cloth were like, “We’ve got no prior preconceptions. Let’s see what kind of software people can build from scratch,” basically. I don’t know. That’s my explanation. I don’t know if you guys have any retros on the AI builders.

Jacob

Or did they get lucky starting that product journey right as the models were reaching the inflection point where they started to work?

swyx

There’s the timing issue. Yeah. I don’t know. To some extent, I think the only reason you and I are talking about it is that both of them have reported ridiculous numbers—basically, 0 to 20 million in 3 months.

Jacob

Jordan, did you have a big surprise?

Jordan

Yeah. Some of what’s already been discussed. I guess the only other thing would be on the Apple side in particular.

swyx

Oof.

Jacob

Those text-message summaries—phew.

swyx

But they’re funny.

Jacob

They’re funny, and how bad they are.

swyx

They’re great.

Jacob

Yeah, and how often.

Jordan

Very viral. For the last couple of years, we’ve seen so many companies trying to do personal assistants—all these various consumer things—and one of the things we’ve always asked is, “Well, Apple is in prime position to do all this.”

Then with Apple Intelligence, they just totally messed up in so many different ways. Then the whole BBC thing—saying that the guy shot himself when he didn’t—and just so many things at this point. I would have thought that they would’ve ironed out their AI products better, but they just didn’t really catch on.

Jacob

Second on this list of generally overly broad opening questions would be: anything that you guys think is overhyped or underhyped in the AI world right now?

Jordan

Overhyped agent frameworks. Sorry.

swyx

Not naming any particular ones.

Jordan

I’m sorry. Not naming any particular ones. I would say there’s just an overall chase to try to be the framework when the workloads are in such flux that I think it’s so hard to reconcile the 2.

I think what Harrison and LangChain have done so amazingly is product velocity. The initial abstractions were maybe not the ending abstractions, but they were just releasing stuff every day, trying to be on top of it. But I think now we’re past that. What people are looking for now is something that they can actually build on and stay on for the next couple of years.

We talked about this with Brett Taylor on our episode, and it feels like it’s the jQuery era of agents and LLMs. It’s kind of like single-file, big frameworks—kind of like a lot of developers—but maybe we need React. I think people are still just trying to build jQuery. I don’t really see a lot of people doing React-like stuff.

swyx

Yeah. Maybe the only modification I’d make about that is that maybe it’s too early even for frameworks at all.

Jacob

Yeah. Do you think there’s enough stability in the underlying model layer and patterns to add this?

swyx

Right. The thing is the protocol and not the framework. Frameworks inherently embed protocols, but if you just focus on a protocol, maybe that works. Obviously, MCP is currently leading in that area.

The comparison there would be: instead of just jQuery, it is XMLHttpRequest, which is the thing that enabled AJAX. That was the sort of inciting incident for JavaScript becoming popular as a language.

Jordan

I would largely agree with that. On the React side of things, I think we’re starting to see more frameworks go after more of that. I guess Mastra is sort of on the TypeScript side and more of a sort of—

swyx

Mastra? Yeah.

Jordan

Yeah, yeah. The traction is really impressive there. So I think we’re starting to see more surface area there, but I think there’s still a big opportunity.

Jacob

What do you have for an over- or underhyped thing?

Jordan

On the underhyped side, I know I mentioned Apple already, but I think the Private Cloud Compute side, with PCC, could be really big. It’s under the radar right now, but in terms of basically bringing the on-device sort of security to the cloud, they’ve done a lot of architecturally interesting things there.

swyx

Who’s “they”?

Jordan

Apple.

swyx

Oh, okay.

Jordan

On the PCC side. I actually think that—

swyx

So negative on Apple Intelligence, but positive—

Jordan

But on the—

swyx

On Apple Cloud.

Jordan

On the more local-device side, I think there will be a lot of workloads still on device. But when you need to speak to the cloud for larger LLMs, I think Apple has done really interesting things on the privacy side.

Jacob

Yeah. We seeded a company that does that, so—

Jordan

Yeah. Especially as things become more consumerized—

Jacob

Did you just set him up on purpose? That felt like a perfect—Yeah, no, I was like, “Let’s go, Jordan.” Come on.

swyx

Did you have a meeting before this episode?

Jordan

Tell me about that company after.

Jacob

Yeah. We’ll chat after. But yes, I think the unique thing about LLM workloads is that you just cannot have everything be single-tenant, because you just cannot get enough GPUs.

Even large enterprises are used to having VPCs, where everything runs privately, but now you just cannot get enough GPUs to run in a VPC. So I think you’re going to need to be in a multitenant architecture, and you need, like you said, single-tenant guarantees in multitenant environments.

Jacob

It's an interesting space. What about you, swyx?

swyx

Underhyped, I want to say memory—stateful AI. As part of my keynote—for every conference I do, I give a keynote, and I try to do the task of defining an agent. It's evergreen content for a keynote.

Alessio Fanelli

Evergreen content.

swyx

Evergreen content for a keynote.

But I did it in a way that was, I think, what a researcher would do. You survey what people say, and then you categorize it and go, “Okay, this is what everyone calls agents, and here are the groups of definitions. Pick and choose.” Right?

It was very interesting that the week after that, OpenAI launched their Agents SDK and formalized what they think agents are. Cloudflare also did the same with us. None of them had memory.

Alessio Fanelli

Yeah.

swyx

It's very strange.

Alessio Fanelli

Mm.

swyx

Obviously, there's conversation memory, but there's not memory-memory, like a knowledge graph of facts about you that can exceed the context length.

Curiously, if you look closely enough, there's a really good implementation of memory inside MCP. When they launched with the initial set of servers, they had a memory server in there, which I would recommend as where you start with memory.

But I think if there were a better memory abstraction, a lot of our agents would be smarter and could learn on the job, which is something that we all want. For some reason, we've all just ignored that because it's convenient to—

Alessio Fanelli

Do you feel like it's being ignored, or is it just a really hard problem and lots of—

swyx

Yeah.

Alessio Fanelli

I feel like lots of people are working on it. It just feels like it's proven more challenging—

swyx

Yeah. Harrison has LangMem, which I think he's relaunched again now. Then we had Letta come speak at our conference. I don't know Zep. I think there's a bunch of other memory guys.

But something like this should be normal in the stack. Basically, I think anything stateful should be interesting to VCs because it's databases, and you know how those things make money.

4. New Model Companies Face Pressure

Alessio Fanelli

I think on the overhyped side, the only thing I'd add is that I'm still surprised how many net-new companies there are training models. I thought we were past that.

swyx

I would say they died at the end of last year, and now they've resurfaced.

Alessio Fanelli

Yeah.

swyx

That's one of the questions you had down there: Is there an opportunity for net-new model players? I would've said no. I don't know what you guys think.

Alessio Fanelli

Yeah, sorry.

swyx

That's fine.

Jacob

I don't have a reason to say no, but I also don't have a reason to say this is what's missing and you should have a new model company do it. But again, I'm an engineer.

swyx

All these guys want to pursue AGI. They all want to be like, “Oh, we'll hit all the benchmarks,” and they can't all do it.

Jacob

Yeah. I mean, look, I don't know if Ilya has this secret approach up his sleeve—something beyond test-time compute.

swyx

Mm-hmm.

Alessio Fanelli

But it was funny. We had Noam Shazeer on the podcast last week, and I was asking him, “Is there some other algorithmic breakthrough? What do you make of Ilya?”

And he's like, “Look.” I think what he implicitly said was that test-time compute will get to the point where these models are doing AI engineering for us. At that point, they'll figure out the next algorithmic breakthrough, which I thought was pretty interesting.

Jordan

I agree with you folks. I think that we're most interested, at least from our side, in foundation models for specific use cases and more specialized use cases.

I guess the broader point is, if there's something that these companies can latch onto and be known for being the best at, maybe there's a case for that. Largely, though, I do agree with you that I don't think there should be more model companies at this point. That's just—

Jacob

I think it's these unique datasets, right? Obviously, robotics has been an area we've been really interested in—

Jordan

Mm.

Jacob

—which is an entirely different set of data that's required on top of a good LLM, and then biology, materials science—

Jacob

More about the specific use cases, basically.

Jacob

Yeah, but also specific, like—

Jordan

Yeah. Specific markets.

Jacob

I think a lot of these models are super-generalizable, but finding opportunities where, for a lot of these bio companies, they have wet labs. They're literally running a ton of experiments. The same is true on the materials science side.

I still feel like there are some opportunities there. But the core LLM agent space is tough to compete with the big ones.

swyx

Yeah. Agreed.

Jacob

Yeah. But they're moving more into product. So I think—

swyx

Yeah.

Alessio Fanelli

—that's the question: If they could do better vertical models, why not do that instead of trying to do Deep Research and Operator and these different things?

swyx

Mm-hmm.

Alessio Fanelli

In my mind, it's financial pressure. They need to monetize in a much shorter timeframe because the costs are so high.

Alessio Fanelli

But maybe it's not that easy to do.

Jacob

You think it would be a better business model to do a vertical—

Alessio Fanelli

Well, it's more like, why wouldn't they? You make fewer enemies if you're a model builder, right?

Now, with Deep Research and search, Perplexity is an enemy, and Gemini Deep Research is more of an enemy, versus if they were doing a finance model, you know—

Jacob

Mm-hmm.

Alessio Fanelli

—or whatever. They would just enable so many more companies. They had Habia as one of the customer case studies for ChatGPT Search, but they're not building a finance-based model for them.

So is it because it's super hard and somebody should do it, or is it because the new models are going to be so much better that the vertical models are useless anyway? Like, the better, less—

Jacob

This is the better, less—

Jordan

Exactly.

Jacob

It still seems to be a somewhat outstanding question. I'd say all the signs of the last few years seem to be that a general-purpose model is the way to go.

Training a hyper-specific model in a domain is maybe cheaper and faster, but it's not going to be higher quality. I mean, we were talking to Noam and Jack Rae from Google last week, and they were like, “Yeah, this is still an outstanding—”

We check this every time we have a new model, and that still seems to be holding. I remember a few years ago it felt like all the rage was the BloombergGPT model, and everyone was like, “You gotta take massive data and—”

swyx

Yeah, I had the CPO at Bloomberg present on that.

Jacob

Yeah. That must be a really interesting episode to go back to, because I feel like—

swyx

Yeah.

Jacob

—very shortly thereafter, the next OpenAI model came out and beat it on all sorts of—

swyx

No, it was a talk. We haven't released it yet. But, yeah, basically, they concluded that the closed models were better, so they just—

Jacob

Okay, yeah.

swyx

They stopped.

Jacob

So they thought it out. Interesting.

swyx

Exactly.

Jacob

Yeah, because I feel like that's been the—

swyx

But he's very insistent that the work that he did, the team he assembled, and the data that he collected are useful for more than just the model. Basically, everything but the model survived.

Jacob

What are the other things?

swyx

The data pipeline—

Jacob

Okay.

swyx

—the team that they assembled for fine-tuning and implementing whatever models they ended up picking. It seems like they are happy with that, and they're running with that. He runs 12 or 13 teams at Bloomberg, just—

Working GenAI across the company.

Jacob

I mean, I guess we’ve all been alluding to it, but because it’s a natural transition, the other broad opening I have is: What are we paying most attention to right now? I think the model companies coming into the product area is going to be fascinating to see play out over the next year, with these frenemy dynamics—

swyx

Yeah.

Jacob

It feels like it’s going to first boil up around Cursor and Anthropic.

Alessio Fanelli

Mm-hmm.

Jacob

The way that plays out over the next 6 months, I think, will be what we’re—

swyx

What is Cursor-Anthropic? Do you mean Cursor versus Anthropic?

Jacob

Yeah, I mean, I assume over time Anthropic wants to get more into the application side of coding. I assume over time Cursor will want to diversify away from just using the Anthropic model.

swyx

Uh-huh.

It’s interesting that Cursor is now worth, like, $9 billion to $10 billion.

Jacob

Yeah.

Alessio Fanelli

Mm-hmm.

swyx

They’ve made themselves hard to acquire. I would have said: You should just—

Alessio Fanelli

Yeah.

swyx

Get yourself to $5 billion or $6 billion and join OpenAI. All the training data goes to OpenAI, and that’s how they train their coding model. Now it’s not as complicated. Now they need to be an independent company.

Jacob

Increasingly, it seems the model companies want to get into the product layer. Seeing, over the next 6 to 12 months, whether having the best model lets you start from a cold start on the product side and get something in market—or whether the companies with the best products, even if they eventually have to switch to a somewhat worse, tiny bit worse model, are where developers ultimately choose to go—I think that’ll be super interesting.

swyx

Yeah.

Alessio Fanelli

But don’t you think that Devin is more in trouble than Cursor? I feel like Anthropic, if anything, wants to move more toward—I don’t think they want to build the IDE. If I think about coding, it’s kind of like you look at it like a cube. The IDE is one way to get the code, and then the agent is the other side.

I feel like Anthropic wants to be more on the agent side and then hand you off to Cursor when you want to go in depth, versus trying to build a cloud IDE. I don’t think that’s—

swyx

I would say—

Alessio Fanelli

I don’t know how you think about it.

swyx

The existence of Claude Code doesn’t support what you say. Maybe they would, but they haven’t shown that yet.

Jacob

I assume both just converge eventually, where—

Alessio Fanelli

Yeah.

Jacob

You’ll be able to do both.

swyx

So, in order to be an outer-loop coding agent, we’re talking about the distinction between the inner loop and the outer loop, right?

Jacob

Yeah.

swyx

The inner loop is inside Cursor, inside your IDE, between Git commits, and the outer loop is between Git commits in the cloud. To be an outer-loop coding agent, you have to be more of a, “We will integrate with your codebase. We’ll sign whatever security thing you need us to sign,” that kind of schlep.

I don’t think the model labs want to do that schlep. They just want to provide models. That would be my argument for why Cognition should still have some moat against Anthropic, simply because Cognition would do the schlep, the business development, and the infrastructure that Anthropic doesn’t really care about.

Jacob

I know the schlep is pretty sticky, though, once you do it.

swyx

It’s very sticky. Yeah. I think the natural winner of that should be Sourcegraph, but—

Jacob

Ah, another unprompted Metronome portfolio point.

swyx

Metronome portfolio point. Nice.

Jacob

Thank you. Nice.

swyx

I mean, they’re big supporters. I’m very friendly with both Quinn and Beyang, and they’ve done a lot of work with Cody, but not much work on the outer-loop stuff yet.

Any company where they’ve already had, “We’ve been around for 10 years. We have all the enterprise contracts. You already trust us with your codebase. Why would you go trust Factory or Cognition, two-year-old startups that just came out of MIT?” I don’t know.

5. Applications Find Product Market Fit

Jacob

I guess switching gears to the application side, I’m curious, for both of you, how do you characterize what has genuine product-market fit in AI today? I guess Alessio, more for you, and you’re sort of on the investing side: Is it more interesting to invest in the category of stuff that works today, or in where the capabilities are going long term? That’s hard.

swyx

I’m just going to do my job for you.

Jacob

Yeah, you’re like, “Man, that’s an easy layup.”

swyx

Tell us all your investment theses.

Jacob

I would say we mostly do seed investing, so it’s hard to invest in things that already work because it means they’re already late.

We try to be at the cusp. Usually, the investments we like to make have really not that much market risk. If this works, obviously people are going to use it, but it’s unclear whether or not it’s going to work, so that’s more what we skew toward.

We try not to chase as many trends. I was a founder myself, and sometimes I feel like it’s easy to just jump in and do the thing that’s hot. But becoming a founder to do something that’s underappreciated or doesn’t yet work shows some level of grit. You actually really believe in the thing. That alone, for me, makes me skew more toward that.

And you do a lot of angel investing too—

swyx

Yeah.

Jacob

So I’m curious how.

swyx

Yeah, but I don’t use that in my mental framework of things. I come at this much more as a content creator or market analyst. It really does matter to me what has product-market fit because I have to answer the question of what’s working now when people ask me.

Jacob

Yeah. Do you feel like, relative to the hype and discourse out there, there are a lot of things that have product-market fit, or a few things? Where—

swyx

A few things.

Jacob

Yeah.

swyx

So I have a list. 2 years ago, I wrote the “Anatomy of Autonomy” post, which was a first look at what’s going on in agents and what is actually making money. I think there are a lot of GenAI skeptics out there who are like, “These things are toys. They’re not reliable. Why are you dedicating your life to these things?”

For me, the product-market-fit bar at the time was $100 million. What use cases can reasonably fit $100 million? At the time, it was Copilot. It was Jasper—no longer.

Jacob

Mm-hmm.

swyx

But in that category of helping you write, which I think was helpful. And then Cursor, I think, was on there as a coding agent, plus-plus.

I think that list will just grow over time: the form factors that we know to work, and then we can adapt the form factors to a bunch of other things. The one that was most recently added to this is deep research.

Jacob

Yeah.

swyx

Anything that looks like deep research, whether it’s a Grok version, Gemini version, Perplexity version, whatever. Jordan has an investment that he likes called Brightwave, which is basically deep research—

Jacob

Mm-hmm.

swyx

For finance.

Jacob

Yeah.

swyx

Anything where it’s long-term agentic reporting, and it’s starting to take more and more of the job away from you and just give you a much more reasoned report, I think is going to work. That has some product-market fit, I think. Obviously, it has product-market fit.

I went through this exercise of trying to handicap how much money OpenAI made from launching OpenAI Deep Research. I think it’s billions. The sheer upgrade from $20 to $200 has to be billions in ARR. Maybe not all of them will stick around, but that is some amount of product-market fit—

Jacob

Didn’t they have to immediately drop it down to the $20 tier?

swyx

They expanded access. I wouldn’t say that.

Jacob

Which I thought was really telling of the market, right?

Jacob

I think it's going to be so interesting to see what they're actually able to get in that $200 or $2,000 tier, which we all think has a ton of potential.

swyx

Yeah.

Jacob

I thought it was fascinating. I don't know whether it was just to get more people exposure to it or the fact that Google had a similar product, obviously, and other folks did too. But it's really interesting how quickly they dropped it down.

swyx

I think that's just a more general policy: no matter what they have at the top tier, they always want to have smaller versions of that in the lower tiers.

Jacob

Yeah, just get people exposure to it.

swyx

Just get exposure. The brand of being first to market and the default choice is paramount to OpenAI.

Jacob

Yeah.

Though I thought that whole thing was fascinating because Google had the first product, right?

swyx

Yeah.

Jacob

Totally.

swyx

We interviewed them. I straight-up said to their faces, “OpenAI mocked you.”

Jacob

Totally.

swyx

And they were like, “Yeah.”

Jacob

Well, actually, I'm curious. This is totally off topic, but what is it going to take for Google? Google just released some great models a few weeks ago. I feel like—

swyx

It's happening.

Jacob

The stuff they're shipping is really cool.

swyx

It's happening.

Jacob

But I also feel like, at least in the broader discourse, it's still a drop in the bucket relative to—

swyx

Yeah. Again, I'm very bullish on this. I think it's happening. I think it takes some time. My Gemini usage is up. I use it a lot more for anything from summarizing YouTube videos to the native image generation that they just launched to Flash Thinking.

Jacob

Yeah.

swyx

I run a daily news recap called AI News that is 99% generated by models, and I do a bake-off between all the frontier models every day.

Jacob

Every day?

swyx

Yes.

Jacob

Does it switch?

swyx

Yes, it does switch. I manually do it.

Jacob

Yeah.

swyx

Flash wins most days. I think it's happening. I was thinking about tracking myself—the number of opens of ChatGPT versus Gemini—and at some point it will cross. I think Gemini will be my main. That will slowly happen for a bunch of people, and then that'll shift. I think that's really interesting. For developers, this is a different question.

Jacob

Yeah.

swyx

It's Google getting over itself—having Google Cloud versus Vertex versus AI Studio. They're all these 5 different brands slowly consolidating. It'll happen, just slowly, I guess.

Jacob

Yeah.

swyx

Yeah.

Jacob

I mean, another good example is that you cannot use the thinking models in Cursor. I know Logan Kilpatrick said they're working on it, but I think there are all these small things where, if I cannot easily use it, I'm really not going to go out of my way to do it. But I do agree that when you do use them, their models are great. So—

swyx

Yeah.

Jacob

They just need better bridges.

swyx

You had one of the questions in the prep: what public company are you long and short? And mine is Google—

Jacob

Yeah.

swyx

Google versus Apple. Long Google, short Apple.

Jacob

That was also my combo.

swyx

I think it does feel like Google's really cooking right now.

Yeah. So, okay, coming back to what product-market fit is—

Jacob

Yeah, now that we've come back from my complete, total sidetrack.

swyx

There's also customer support. We were talking in the car about Decagon and Sierra, obviously. Brett Taylor is the founder of Sierra. And, yeah, it seems like there are these layers of agents. I think you just look at the income statement or the org chart of any large-scale company, and you start picking them off one by one—what is interesting knowledge work—and they would just eat things slowly from the outside in.

Jacob

Yeah.

swyx

If that makes sense.

Alessio Fanelli

I mean, the episode with Brett—he's so passionate about developer tools, and yet he did not do a developer tools company. We spent 2 hours talking about developer tools and all of that stuff, and he's like, “I did a customer support company.” I'm like, man, that says something. You know what I mean?

swyx

Yeah.

Alessio Fanelli

When you have somebody like him who can raise any amount of money from anybody to do anything, to pick customer support as the market to go after, while also being the chairman of OpenAI, that shows you that these things have moats. They're going to stick around, you know? Otherwise, he's smarter than that.

So, yeah, that's a space where maybe initially I would've said, “I don't know if it's the most exciting thing to jump into.” But then, if you really look at the shape of how the workforces are structured and how the cost centers of the business really end up, especially for more consumer-facing businesses, a lot of it goes into customer support. All the AI story of the last 2 years has been cost-cutting.

swyx

Yeah.

Alessio Fanelli

I think now we're going to switch more toward growth.

swyx

Revenue.

Alessio Fanelli

You've seen Jensen: last year at GTC, he was saying, “The more you buy, the more you save.” This year he said, “The more you buy, the more you make.”

Jacob

Oh.

Alessio Fanelli

We were there.

swyx

Hot off the press.

Alessio Fanelli

We were there.

Jacob

I do think that's one of the most interesting things about this first wave of apps: almost the easiest thing that you could get real traction with was stuff that, for lack of a better way to frame it, people had already been comfortable outsourcing to BPOs or something.

swyx

Yeah.

Jacob

They had implicitly said, “Hey, this is a call center. We are willing to take some performance cut for lower cost in the past.” The irony of that—or what I'm really curious to see play out—is that you could imagine that is the area where price competition is going to be most fierce, because it's already stuff that people have said, “Hey, we don't need the 100% best version of that.”

I wonder whether this next wave of apps may prove even more defensible as you get these capabilities that actually increase the top line or whatnot. Take AI go-to-market, for example. You'd pay twice as much for it, because there's just a very clean ROI story to it. And so I wonder ultimately whether this next set of apps actually ends up being more interesting than the first wave.

swyx

Yeah.

Jordan

I think a lot of the voice AI ones are interesting too, because you don't need 100% precision and recall to actually have a great product. For example, we looked into a bunch of scheduling-intake companies for home services, for electricians and stuff like that. Today, they miss 50% of their calls. So even if the AI is only effective, say, 75% of the time—yeah, it's crazy, right?—that's totally fine, because that's still a ton of increased revenue for the customer. You don't need 100% accuracy.

As the models and the reliability of these agents get better, it's totally fine because you're still getting a ton of value in the meantime.

swyx

Yeah. I don't know how related this is, but one of my favorite meetings at—

Jordan

Yeah.

swyx

It is related. One of my favorite meetings at AI Engineer Summit—this was our first one in New York—and I just met a different crew than you meet here. Everyone here loves developer tools and infra. Over there, they're actually more interested in applications, so it's kind of cool.

I met this bootstrap team that only does appointment scheduling for vets. They're like, “This is an anomaly. We don't usually come to engineering summits because we usually go to vet summits.”

Jordan

I'm sure it's a massive pain point. They're willing to pay a lot of money.

Jacob

Yeah.

Alessio Fanelli

But this is my point about saving versus making more: if an electrician takes 2X more calls, do they have the bandwidth to actually do 2X more in-house?

swyx

Then they get hired.

Alessio Fanelli

Well, yeah, exactly. That's the thing: I don't think most businesses today are structured to overnight 2–3× their bandwidth. I think that's a startup thing. Most businesses can't do it, so—

swyx

Then you make an electrician agent.

Alessio Fanelli

Well, no, totally. How do you do a recruiting agent for electricians? Or electrician training?

Jordan

That's a good point.

Jacob

Yeah.

Alessio Fanelli

How do you do Lambda School for electricians? I don't know. It's like—

Jacob

It's going to be whack-a-mole for the bottlenecks in these businesses—

Speaker 3

Yeah, exactly.

Speaker 0

—you're like, "Oh, now we have a ton of demand. Cool, where do we go?"

swyx

Yeah.

Speaker 2

Yeah.

swyx

So, just to round out this PMF thing, I think this is relevant in a sense: it's pretty obvious that the killer agents are coding agents, support agents, and deep research, right? Roughly, right? We've covered all those 3 already. Then you have to turn to offense and go, "Okay, what's next?"

Jacob

I mean, I also just like summarization of voice and conversation—

swyx

Yep.

Speaker 0

Absolutely.

swyx

We actually had that on there.

Speaker 0

Yeah.

swyx

I just didn't put it as an agent because it seems less agentic, you know? But yes.

Jacob

Still a good AI use case.

swyx

That one I've seen—I would mention Granola. And what's the other one?

Speaker 0

Um—

swyx

Monterey?

Jacob

I think Abridge was the one you wanted to mention.

Alessio Fanelli

I was going to say Abridge, yeah.

swyx

Abridge? Abridge? Okay. I'll just call out what I had on my slides for the agent engineering thing.

Speaker 0

Yeah.

swyx

It was screen sharing, which I think is actually kind of underrated: an AI watching you as you do your work and just offering assistance. Outbound sales, so instead of support, just more outbound sales. Hiring—

Speaker 0

You say outbound sales has product-market fit?

swyx

No, it will. It's coming.

Speaker 0

Oh, on the commercial side—yeah, I totally agree with that.

swyx

Yeah. Hiring, like the recruiting side. Education, personalized teaching, I think.

Jacob

I'm kind of shocked we haven't seen more there.

Jordan

Mm.

swyx

Yeah.

Jacob

I don't know if that's—

swyx

It's like Duolingo is the thing.

Jacob

Yeah, I mean, Speak and some of these practice—

swyx

Speak.

Jacob

Practice.

swyx

Yeah.

Jacob

Yeah.

swyx

Yeah, interesting. And then finance—there are a ton of finance use cases that we can talk about. And then personal AI, which we also had a little bit of that. But I think personal AI is harder to monetize. I think those would be what I would say is up-and-coming in terms of what I'm currently focusing on.

Jacob

I feel like this question has been asked a few different ways, but I'm curious what you guys think. If we just froze model capabilities today, are there trillions of dollars of application value to be unlocked? Like AI education—if we just stopped all model development today, with this current generation of models, we could probably build some pretty amazing education apps.

Or how much of this is contingent upon people having had 2 years with GPT-4 and, I don't know, 6 months with the reasoning models? How much is contingent upon it just being more time with these things versus the models actually having to get better? I don't know. It's a hard question, so I'm going to just throw it to you.

Jordan

Yeah. Well, I think the societal thing is maybe harder, especially in education. Can you basically dodge the education system? Probably you should, but can you? I think it's more of a human—

Jacob

But people pay for all sorts of get-ahead things outside of class, and certainly in other countries there's a ton of consumer spending on education. It feels like the market opportunity is there.

swyx

Yeah, and in private education, I think. Public is very different.

Jacob

Yeah.

swyx

One of my most interesting quests from last year was kind of reforming Singapore's education system to be more AI-native.

Jacob

Just what you were doing on the side while you were—

swyx

Yes.

Jordan

While waiting for our actual, yeah.

Jacob

That's a great side quest.

swyx

My stated goal is for Singapore to be the first country that has Python as a first language, as a national language. Anyway, the pushback I got from the Ministry of Education was that the teachers would be unprepared to do it.

Jacob

Hmm.

swyx

It was really interesting. The immediate pushback was the de facto teachers' union being resistant to change. And I'm like, "Okay."

Jacob

Yeah.

swyx

That's par for the course. Anyway, not to dwell too much on that, but I think education is one of those things that everyone has strong opinions on because we all have kids, and we've all been through the education system.

But I think it's going to be like the domain-specific Speak. Such an amazing example of top-down: we will go through the idea maze, and we'll go to Korea and teach them English. It's like, what the hell? I would love to see more examples of that—really focused, where no one tried to solve everything. Just do your thing really, really well.

Jacob

On this trend of difficult questions that come up, I'm going to ask you the one that my partners like to ask me every single Monday, which is: how do you think about defensibility at the app layer?

swyx

Oh, yeah. That's great.

Jacob

Just give me an answer I can copy and paste and have an auto-response.

swyx

Network effects.

Jacob

Yeah.

swyx

Honestly, network effects. I think people don't prioritize those enough because they're trying to make the single-player experience good, but then they neglect the multiplayer experience.

I always think about load-bearing episodes. As podcasts, you do 1 a week, and some of those you don't really talk about ever again, while others you keep mentioning every single podcast. I think the recap episodes for us are pretty load-bearing. We refer to them every 3 months or so.

One of them, I think, for us is Chai—for me, it's Chai Research. Even though that wasn't a super popular one among the broader community outside of Chai, the Chai community. For those who don't know, Chai Research is basically a Character.AI competitor.

Jacob

Yeah.

swyx

Right? They were bootstrapped, they were founded at the same time, and they have outlasted Character.AI, de facto.

Jacob

Yeah.

swyx

It's funny. I would love to ask Milosz Kerzner a bit more about the whole Character.AI thing, but—

Jacob

Good luck getting past the Google cops.

swyx

He doesn't have his own models, basically. He has his own network of people submitting models to be run, and I think that, in the short term, is going to hurt him because he doesn't have proprietary IP. But in the long term, he has the network effect to make him robust to any changes in the future.

I want to see more of that, where he's basically looking at himself as kind of a marketplace. He's identified the choke point, which is to build the app—or the protocol layer—that interfaces between the users and the model providers, and then make sure that the money flows through. And that works.

I wish more AI builders or AI founders emphasized network effects, because that's the only thing that you're going to have at the end of the day.

Jacob

Yeah.

swyx

And brand leads into network effects as well.

Jacob

Yeah. I guess it's harder in the enterprise context, right? But I feel—it's funny, we do this exercise, and I feel like we talk a lot about the velocity and the breadth you're able to build of product surface area. There's also just the ability to become a brand in the space.

I'm shocked that even in 6–9 months, an individual company can become synonymous with an entire category, and then they're in every room for customers—

swyx

Yeah—

Jacob

—and some—

swyx

Well—

Jacob

—of the other startups are clawing their way to try and get into 1/20th of those rooms.

Jordan

There's a bunch of categories that we talk about in IC, and it's like, "Oh, pricing compression is going to happen. It's not as defensible, so ACVs are going to go down over time." In actuality, some of these ACVs have doubled, we've seen, and the reason for that is just people go to them and pay for that premium of being that brand.

Jacob

Yeah. I mean, what I'm struck by is there was such a head fake in the early days of AI apps where people were like, “We want this amazing defensibility story.” And then what's the easiest defensibility story? It's like, “Oh, total unique data set or train your own model or something.” I feel like that was just a total head fake. I don't think that's actually useful at all.

swyx

Mm-hmm.

Jacob

You sound much less articulate when you're like, “Well, the defensibility here is the thousand small things that this company does to make the user experience, design, everything, just delightful,” and just the speed at which they move to both create a really broad product, but then also, every 3–6 months, when a new model comes out, it's kind of an existential event for any company, because if you're not the first to—

swyx

Yeah.

Jacob

—figure out how to use it, someone else will.

swyx

Yeah.

Jacob

And so velocity really matters there. And it's funny, in our internal discussions, we've been like, “Man, that sounds pretty similar to how we thought about application SaaS companies.” There isn't some revolutionary reason—you don't sound like a genius when you're like, “Here's applications: application SaaS company A is so much better than B.” But it's a lot of little things that compound over time.

6. AI Infrastructure Finds Its Role

What about the infrastructure space, guys? I'm curious: how do you think about where the interesting categories are here today? Where do you want to see more startups, or where do you think there are too many?

Jordan

Yeah. Yeah, we call it kind of the LLM OS, but I would say—

swyx

Not we. I mean, Andre—Andre calls it LLM OS.

Jacob

Mm-hmm.

Jordan

Well, but, yeah, we—

Jacob

You and Andre, the three of you, call it the LLM OS.

Jordan

Well, we have this Four Horsemen of AI framework that we use, and LLM OS is one of them. But, yeah, I mean, code execution is one. We've been banging the drum. Everybody now knows we're investors in E2B.

Jacob

Mm-hmm.

Jordan

Memory is one that we kind of touched on before. Super interesting. Search, we talked about. I think those are more—

Alessio Fanelli

Not traditional infra, not like the bare-metal infra. It's more like the infra around the model—

swyx

Tools for agents. Yeah.

Alessio Fanelli

—you know? Which I think is where a lot of the value is going to be.

swyx

The security ones? Yeah.

Alessio Fanelli

Yeah, yeah. In cybersecurity, I mean, there's so much to be done there. Basically, in any area where AI is being used by the offense, AI needs to be applied on the defense side. Email security, identity, all these different things. So we've been doing a lot there.

As well as having you rethink things that used to be costly, like red teaming, and maybe used to be a checkbox in the past. Today, they can actually be helpful to make your app secure.

And there's this whole idea of semantics, right, that now the models can be good at. In the past, everything was about syntax. It's kind of like very basic constraint rules. I think now you can start to infer semantics from things that are beyond just simple recognition, to understanding why certain things are happening a certain way.

So, in the security space, we're seeing that with binary inspection, for example. There's kind of the syntax, but then there are the semantics of understanding what this code overall is really trying to do, even though this individual syntax is saying something specific.

Not to get too technical, but I think infrastructure overall is a super interesting place if you're making use of the model. If you're just serving the models, I'm less bullish—not that it's not a great business, but I think it's a very capital-intensive business.

swyx

Mm-hmm. Yeah.

Alessio Fanelli

I think infrastructure is great; people will make money. But I don't think there's as much interest from us, at least.

Jordan

How do you guys think about what OpenAI and the big research labs will encompass as part of the developer and infrastructure category?

Alessio Fanelli

Yeah. That's why I would say Search is the first example of one of the things we used to mention. We had Axon on the podcast, and Perplexity obviously as an API—

swyx

Mm-hmm.

Alessio Fanelli

—as an API.

swyx

The basic idea is, if you go into the ChatGPT custom GPT builder, what are the checkboxes? Each of them is a startup.

Alessio Fanelli

Yeah.

swyx

Yeah. Yeah.

Alessio Fanelli

And now they're also APIs.

swyx

Yes. Yeah.

Alessio Fanelli

So now Search is also an API. We'll see what the adoption is. In traditional infra, everybody wants to be multicloud. So maybe we'll see the same, where ChatGPT Search or the OpenAI Search API is great with the OpenAI models because you get it all bundled in, but their price is very high.

If you compare it to Axon, I think it's 5 times the price for the same amount of research, which makes sense if you have a big OpenAI contract. But maybe if you're just picking and choosing, you want to compare different ones.

swyx

Yeah.

Alessio Fanelli

They don't have a code execution one. I'm sure they'll release one soon. So they want to own that too. But the same question we were talking about before, right? Do they want to be an API company or a product company? Do you make more money building ChatGPT Search or selling the Search API?

swyx

The broader lesson, instead of going, “We did applications just now, and then what do you think is interesting infrastructure?”—it's not 50/50. It's not equal-weighted.

Alessio Fanelli

Yeah.

swyx

It's just very clearly that the application layer has been way more interesting.

Alessio Fanelli

Yeah.

swyx

There's interesting infrastructure plays, and I even want to push back on the whole GPU-serving thing, because Together AI is doing well, and Fireworks is doing well.

Jordan

No, that's how it works. It's like data centers and inference providers for the—

Alessio Fanelli

Oh, yeah, yeah, yeah.

Jordan

—you know.

Alessio Fanelli

I think it's all the capital.

swyx

I see. I see. Yeah.

Alessio Fanelli

Again, maybe you guys have—

swyx

Capital efficiency. Yeah.

Alessio Fanelli

—a much larger fund, so, you know, I'm sure you have GPU pods, but—

swyx

Yeah. Yeah. So, that's one thing I have been learning. I think I have historically had a dev tools and infra bias, and so has he, and we've had to learn that applications actually are very interesting and also maybe kind of the killer application of models, in the sense that you can charge for utility and not for cost, right? Whereas most infrastructure reduces to cost-plus. That's not where you want to be for AI.

I thought it'd be interesting for me to be the only non-VC in the room, saying what's not investable, because then I won't be canceled for saying, “Your whole category is not investable.”

Jordan

The interesting thing is that this thing's not investable, and then 3 months later, we're desperately chasing people.

swyx

Exactly. So you don't want to be on the record—

Jordan

This space changes so fast. It's like every opinion you hold, you have to hold it quite loosely.

swyx

I'm happy to be wrong in public. I think that's how you learn the most, right?

Jordan

Yeah.

swyx

Fine-tuning companies is something I struggled with, and still, I don't see how this becomes a big thing. You kind of have to wrap it up in a broader enterprise AI company, like a services company, like Writer AI, where they will fine-tune as part of the overall offering, but that's not where you spike. It's kind of interesting.

And then I'll just group AI DevOps, and there's a lot of AI SRE out there. It seems like there's a lot of data out there that should be able to be plugged into your codebase or your app to sort of self-heal or whatever. It's just, I don't know if that's been a thing yet, and you guys can correct me if I'm wrong.

And the last thing I'll mention is voice real-time infra. Again, very interesting, very hot, but again, how big is it? Those are the main 3 that I'm thinking about for things I'm struggling with.

Jordan

Yeah, I guess a couple of comments. On the AI SRE side, I actually disagree with that one.

swyx

Yeah.

Jordan

I think the reason they haven't sort of taken off yet is because the tech is just not there quite yet. And so it goes back to the earlier question: do we think about investing toward where the companies will be when the models improve versus now? I think in the short term we'll get there, but it's just not there just yet.

Jordan

But I think it’s an interesting opportunity overall.

swyx

Yeah. My pushback to you is: while it’s monitoring a lot of logs, it’s basically anomaly detection rather than—there’s a whole bunch of stuff that can happen after you detect the anomaly, but it’s really just anomaly detection, and we’ve always had that. This is not a transformers or LLM use case. This is just regular anomaly detection.

Alessio Fanelli

It’s more in terms of: it’s not going to be an autonomous SRE for a while. The question is, how much can the latest AI advancements increase the efficacy of bringing your MTTR down?

swyx

Yeah.

Alessio Fanelli

Even if it’s a 10% improvement over beforehand, that’s still potentially a lot of revenue.

swyx

Okay.

Alessio Fanelli

Right. That’s the way I would think about it now. A few years from now, if it’s actually an autonomous SRE just replacing it altogether, then that’s a totally different thing.

swyx

Hmm. Cool. I’ll look out for it.

Alessio Fanelli

Yeah.

swyx

Hmm.

7. The Biggest Unanswered AI Questions

Jacob

Switching back to overly broad questions, what do you feel is the biggest unanswered question in AI today that has large implications for the ecosystem?

swyx

Yeah. I’ve been banging the drum on RL, and I think it’s clear that you can do RL successfully on verifiable domains. I would say the question is whether or not we can figure out how to do that in non-verifiable domains. Law is a great example.

Jacob

Totally.

swyx

Can you do RL on contracts and documents? Marketing and sales, going back to outbound sales: can you do RL to simulate what an outbound interaction leads to? It’s unclear. If not, then I think we’ll be stuck with agents in the more verifiable domains, and we’ll just have copilots in the non-verifiable ones because you’ll still need a person to be the tastemaker.

Jacob

I had the exact same thought, and I feel like it’s the question. I’m trying to think of the implications: if it doesn’t work, the world could be weird, where you have fully autonomous AI coders and no one does any software or math, or even some areas of science, but then, to write the most basic sales email, it’s still—

swyx

Yeah.

Jacob

It’s always so hard to predict how the world will turn out. Of all the sci-fi that was written 50 years ago, I don’t think anybody foresaw that future. That is—

swyx

It is.

Jacob

—a really weird future.

swyx

Yeah, yeah.

Speaker 0

Do either of you have a different one?

swyx

Well, yeah. No, go. We’ll go back and forth.

Alessio Fanelli

Biggest unanswered question. I don’t know if this is a good answer, but Bob McGrew, who we had on the podcast, was talking about the rule of 9s they have at OpenAI: to go from 90% reliability to 99% is an order-of-magnitude increase in compute. Then 99% to 99.9% is another order-of-magnitude increase, and that happens every 2 to 3 years.

How are we going to scale accordingly through this next part? I think there are a lot of unanswered questions from a hardware perspective. As part of that, from an availability perspective, is NVIDIA just going to continue to be dominant? Obviously, AWS is going hard into their Trainium chips.

Jacob

Mm-hmm.

Alessio Fanelli

I’m blanking on it. Thank you. I think there’s a big ecosystem around CUDA that has obviously allowed NVIDIA to remain dominant, but what’s going to happen? Is there anyone who’s going to combat that to increase the availability of GPUs, or are we just going to be constrained when we actually need way more compute?

swyx

Yeah. My quick thoughts: I’ve been an individual investor in Madex, and I’m the only individual named as an investor. It’s kind of really funny because everyone else was funds, and then it was just me.

There are all these dedicated-silicon startups that are coming up and trying to challenge NVIDIA. The simple answer is that these GPUs are the most general things possible by design. That’s why they do gaming, crypto, and AI.

As long as the architecture seems stable, there seems to be a case to be made for that. The only question is who will win. Obviously, there are a whole bunch of competitors, including AMD, which is trying to make a play for it.

Speaker 0

Yeah.

swyx

But so will AWS, and so will every other company. Microsoft has a chip, and Facebook has a chip. Who knows who will win? It’s very interesting that this seems to be such a valuable prize. It’s NVIDIA that you’re competing with.

Speaker 0

Yeah.

swyx

No one has really made a real dent there yet. I kind of agree with you, but I think it basically is all about workload stability. It’s a bet on the death of transformers, basically. Even the state-space-model people would agree that it wouldn’t really change that much.

The overall consensus is that you don’t even use state-space models individually. You would use them in a mixture with transformers anyway.

Speaker 0

Yeah.

swyx

Just go bet on transformers. Bake it into the chip, and you will have way more ASICs for transformers, and that’s fine. Prima facie, there should be a company that wins that.

Speaker 0

Yeah.

swyx

I don’t know who will win that.

Speaker 0

Yeah.

Alessio Fanelli

I wish we knew.

swyx

I think anyone has to start after 2019 or 2020, because anyone who started before that will still be too general.

Jacob

Mm-hmm.

Alessio Fanelli

Yeah.

swyx

Transformers hadn’t won yet at the time. I have one more. The most emergent one that came out of the New York conference I did was agent authentication.

Jacob

Hmm.

swyx

The Information just published something that they’re worried about: when Operator, or whoever, accesses your website on your behalf, how does it indicate that it’s not you, but an agent of you?

My general philosophy on agent experience, or any of the reinvention of every part of the stack for agents, is that it’s all unnecessary except for this agent-authentication thing. We really need new SSO, effectively, for agents.

Jacob

Yeah. Is it going to be crypto? Both crypto people are really amped about the—

swyx

You know, it’s really frustrating when Sam Altman is right, but maybe you have to scan your eyeballs. Maybe you just have to. Maybe he saw this five years ago and was like, “You gotta scan your eyeballs,” and the rest of us are just behind him, as usual.

Jacob

Oh, I love it. Now I’ll move to the quick-fire round, where we’ll go around the horn and get quick takes on things. The first is dream podcast guest.

Alessio Fanelli

John Carmack.

swyx

Yeah. John is six steps away from solving AGI, apparently, so we can just ask him how far along he is. For me, it’s Andre. He’s a listener and supporter of the pod, and when I launched the whole AI engineer push that we have, he was the first one to legitimize it. He was like, “You know, there will be more AI engineers than ML engineers.”

I think that made everyone else pay attention. Latent Space only exists because he and other people helped promote it.

Jacob

Yeah.

Jordan

I also had Andre, so I guess we’re thinking the same thing there.

Jacob

Yeah, me too. Mine’s a little bit of a cheat. Clearly, they’re writing a book about OpenAI now, and at some point we’ll get to do the Acquired: OpenAI episode. There are probably just so many amazing stories from the last 5 or 6 years.

swyx

Do you know about Doomers? It’s a play. I’m actually going to it this Saturday. Someone made a play about the board drama from last year.

Jacob

Really?

swyx

Yeah, from 2 years ago.

Jordan

Wow.

swyx

Yeah.

Jacob

Wow. That’s wild.

Let us know how it is.

Alessio Fanelli

Yeah, let us know.

Jacob

Maybe we should have the director of that on the pod.

swyx

Yeah. No, I think it’s a lot of fan fiction, basically, but someone will write the accounts, and it’ll be interesting and fascinating. A lot of it will be fake because it’s a complex beast, right? You’re just getting an oral history of what happened.

Jacob

Yeah.

swyx

Yeah.

Jacob

All right. For the next one, I figured you could shout out either a new source you use to stay up to date or a startup that you’re not invested in that you’re excited about.

swyx

Oh.

Jacob

Or you can give it to him.

Alessio Fanelli

My news source is Sean.

Jacob

That’s what I was going to say. I literally wrote swyx’s Twitter.

Alessio Fanelli

So, in our Discord—we have a Latent Space Discord. Any link that ever matters on the internet, Sean’s going to post it into Discord. So all I do is open Discord.

swyx

Wow.

Alessio Fanelli

We have 40 or 50 different channels by topic.

swyx

Yeah, yeah.

Alessio Fanelli

I open Discord, and I’m like, “Okay, AI,” and then I go to developer tools, then I go to creator economy, then I go to stock and macro. They’re all there, so thank you.

swyx

Yeah, that’s very true. That is very true.

We actually met because of the Discord.

Alessio Fanelli

Yeah.

swyx

It was a COVID thing, because everyone was at home and just started a Discord. And yeah, that was the origin of Latent Space, just chatting on the Discord.

Alessio Fanelli

It used to be called Dev/Invest.

swyx

Yeah.

Alessio Fanelli

So it was all about developer tools investing.

swyx

Yeah.

Jacob

Yeah, yeah.

swyx

Yeah. Yeah. I was not prepared for the news sources thing. It’s hard. It’s really shitty to say, but just in-person conversations.

Jacob

Yeah.

swyx

I think the reason I have to be here in SF is because I make friends with people who know things and are smarter than me, and we go for chats, and they’re nice enough to share some stuff. Sometimes I worry that I am being used in order to put things out there that are maybe not true. So I have to exercise my own judgment as to—

Jacob

Yeah.

swyx

—what that is.

Jacob

I think one of the cool things about the podcast in general is just the opportunity to take these conversations that happen in closed rooms and try to bring them onto the airwaves. I’m curious: how much do you feel like the private discourse is similar to the public discourse?

swyx

In many ways, it is surprisingly similar. As in, people at OpenAI learn things about OpenAI from us, which is interesting. And then there are some ways in which it is drastically dissimilar, and those are the things I just cannot repeat until it’s public.

Jacob

This has been super fun. We were looking forward to this for a while.

Alessio Fanelli

Yeah, this was great.

Jacob

We want to make sure everyone around the horn gets an opportunity to plug whatever they want to plug. So we'll leave the last word to all of us, I guess. Where can folks go to learn more about Latent Space and all the exciting things you do? I want to make sure our listeners have a good sense of everything.

Alessio Fanelli

Yes. So we have a Substack. Latent.space is the website. And then please subscribe on YouTube. We're doing a lot of YouTube. We're trying to do better video and all that, so—

swyx

He set our OKRs—

Alessio Fanelli

Come—

swyx

—and it's basically all YouTube.

Alessio Fanelli

Come watch us on YouTube. It's very important for me personally. So even if you don't care, just open it.

Jacob

You gotta hit those OKRs.

Alessio Fanelli

Just mute it.

swyx

Well, we have to increase our production value. Look at this.

Alessio Fanelli

I know. I know. We only have three cameras. Yeah, and then Sean does a lot of the writing outside of the podcast—

swyx

Yeah.

Alessio Fanelli

—on the newsletter.

swyx

Yeah.

Alessio Fanelli

So—

swyx

Yeah, so it's trying to be a newsletter and community and podcast and whatever else that we do. Um, yeah, so, I guess our— for me, I guess there's Latent Space, but then there's also the other big piece, which is the conference that I run.

Jacob

Yeah.

swyx

And the idea is that I think sometimes you just get the good stuff from people if you just put them in front of a lot of people. And that's really—I'm mining people for content, and sometimes you put a mic in front of them and they yap for an hour. Other times you have to put them in front of a prestigious conference, and then they drop some alpha. And so the next one for us is going to be June. It's the AI Engineer World's Fair. And it should be the largest technical conference for AI.

Jacob

And ours is simple. Just subscribe to Unsupervised Learning on YouTube. Also, Swyx, thanks so much. This was awesome.

swyx

Thanks for having us.

Alessio Fanelli

Yeah, it was fun.

Jacob

It was good to see you guys. Thanks for coming on.

Jacob

Hey, guys, this is Jacob. Just one more thing before you take off. If you enjoyed that conversation, please consider leaving a five-star rating on the show. Doing so helps the podcast reach more listeners and helps us bring on the best guests. This has been an episode of Unsupervised Learning, an AI podcast by Redpoint Ventures where we probe the sharpest minds in AI about what's real today, what's going to be real in the future, and what it means for businesses in the world. With the fast-moving pace of AI, we aim to help you deconstruct and understand the most important breakthroughs and see a clearer picture of reality. Thank you for listening, and see you next episode.

Unsupervised Learning x Latent Space Crossover Special | BidClub