[BidClub_]
The a16z Show · · 44 min

What DeepSeek Means For The Future Of AI | Tech Veterans Weigh In

Steph SmithMartin CasadoSteven Sinofsky

YouTube
TL;DR
  • DeepSeek R1 is a genuine Chinese research achievement, but neither a “$6 million model” nor proof that frontier AI suddenly became trivial. The team had roughly 18 months of buildup, with relevant contributions already appearing in public literature, though little was announced; it also released the arguably more impressive V3 base model about two months earlier. The cited spend concerned a particular chain-of-thought training effort. The market nevertheless spent a weekend preparing to “trade away a trillion dollars of market cap,” which Martin Casado called a complete overreaction.
  • R1’s most consequential features are its permissive MIT-like license and its released reasoning traces. Martin called it “free as in free beer, for real”: applications can adopt it broadly, while developers can use its chain of thought to distill capable student models that run on smaller devices. That pushes AI toward “AGI in your pocket” and makes distribution—not merely benchmark leadership—the strategic variable.
  • The investment question is not whether models or apps win, but how their value changes over time. Casado allowed that valuable applications might need vertically integrated models, limiting DeepSeek’s direct impact on OpenAI and Anthropic; Sinofsky answered that both views can be right because “the variable is time.” Models attract users through raw magic, competitors catch up through distillation, and durable value can migrate into stateful workflows and configuration.
  • DeepSeek looks more like a scale-out catalyst than a reason to short NVIDIA. Sinofsky contrasted scarce, liquid-cooled data-center compute with effectively free MIPS on phones and potentially seven billion endpoints; smaller specialized models expand where inference can happen without eliminating hyperscale workloads. Casado called it another step toward “AGI in your pocket,” while Sinofsky argued that the TAM had expanded.
  • AI infrastructure resembles the internet’s fiber buildout, but its financial foundation is materially stronger. Investors again seek exposure through physical infrastructure because private software winners are hard to identify, creating some risk of excess capacity. Yet the primary builders are cloud companies with hundreds of billions of dollars on their balance sheets, while NVIDIA can take a price dip—unlike the leveraged WorldCom-era structure that helped turn fiber oversupply into a crisis.
  • Benchmark leadership will matter less than application-specific reliability and enterprise adoption. Research products will be judged on truth, sources and footnotes rather than parameter counts; productive applications will combine multiple models, fine-tuning and persistent workflow. Enterprise controls such as single sign-on, filtering and disabling features by user may sound mundane, but the speakers see them as sticky, monetizable moats.
  • The panel’s real wake-up call is for US policy, not for OpenAI, Anthropic or NVIDIA. Casado argued that restrictions on open source, chips, software and model weights failed to prevent capable Chinese researchers from building and releasing R1; Sinofsky sharpened the analogy: “The lesson is not Sputnik. The lesson is the internet.” Their prescription is faster domestic research and permissionless diffusion, with frontier labs urged to build applications in a market Sinofsky expects to reach “100x” today’s TAM.
Digest · the substance, structured for research

1. DeepSeek was an accumulated engineering result, not a weekend miracle

  • Casado’s baseline: DeepSeek came from a small Chinese quant hedge fund functioning as a quasi-research organization. Roughly a year and a half of buildup preceded R1, and the model landed within a competitive range of leading systems. What surprised the broader market was not an invisible breakthrough, but the team’s ability to aggregate contributions that specialists could already find in public literature.

  • Casado argued that V3, released about two months earlier, was probably the more impressive feat: a near-GPT-4-class base model and the prerequisite for building a reasoning system. He treated the reported $6 million chain-of-thought spend as plausible beside figures discussed by Anthropic and OpenAI—not as the cost of creating DeepSeek’s entire capability from scratch.

  • Casado called the headline number “irrelevant”: DeepSeek’s paper described training innovations, but financial markets turned a Friday release into a weekend frenzy and arrived Monday ready to “trade away a trillion dollars of market cap.” The magnitude of that reaction said more about positioning and narrative uncertainty than about the paper’s actual claims.

  • Steph Smith situated the hype in the o1 cycle: non-reasoning models seemed to be asymptoting around GPT-4, then OpenAI’s o1 revived excitement around reasoning, compute costs and NVIDIA. R1’s results made people question whether the next wave would require the expected level of spending.

  • The panel preserved uncertainty around possible advantages. Rumors ranged from a CCP “scoop” and hidden spending to deliberate Chinese New Year timing; none was endorsed. A more substantive but still unproven theory was that V3 could use both the broader web and China’s relatively isolated internet, plus lower-cost access to educated experts capable of annotating reasoning steps.

2. An MIT license and visible reasoning traces accelerate diffusion

  • Casado identified an undercovered distinction: R1 arrived under a license that was basically MIT-like and unusually permissive for a recent open-source model. For application companies already combining several proprietary and open models, licenses “really matter”; this one is approximately a page long and “free as in free beer, for real,” making broad proliferation far easier.

  • OpenAI’s o1 exposed a reasoning product without exposing its chain of thought; DeepSeek released the reasoning traces. Those traces let a powerful teacher model train much smaller student models through distillation, producing high-quality systems more quickly and cheaply—and, crucially, models capable of running across far smaller devices.

  • Sinofsky connected licensing to where standardization eventually forms, recalling the GPL, BSD and MIT battles over software distributions. Selling the model layer may resemble charging for HTTP servers: early internet companies focused on monetizing HTML and HTTP, while much larger businesses emerged in shopping, travel and television. Once distribution costs approach zero, selective corporate “openness” is not equivalent to a license that lets an ecosystem build freely.

3. Model magic attracts users, but workflows retain them

  • Casado’s countercase was deliberately unresolved. Models might commoditize, pushing value toward applications; alternatively, the best applications—ChatGPT or the newly launched research products, for example—might require ownership of the underlying model. Under that second path, DeepSeek matters less competitively because it is not yet building the vertically integrated applications.

  • Sinofsky replied, “You’re right both times,” because the missing variable is time. Windows and Office once appeared inseparable from application success; then search emerged without owning either, and other applications later emerged and integrated what they needed. New platforms reorganize the spectrum rather than simply replacing one incumbent category.

  • The important applications may not exist yet, just as the defining internet and mobile products did not exist before their platforms. An AI-native image product may initially resemble Canva, but an AI-native presentation tool should not reproduce PowerPoint’s “3,000 different formatting commands”; generation removes much of the kerning, nudging and coloring that defined the old product.

  • Casado’s portfolio observation: a large model can feel magical when first exposed directly, yet rival models catch up rapidly because distillation works so well. Defensible companies use that initial engagement to build stateful, configurable and retentive applications around the model. “The last two years have been the story of the large model”; the next era is workflow powered by many models.

4. Scale-out multiplies endpoints without abolishing hyperscale compute

  • Sinofsky framed DeepSeek as a change in trajectory from scale-up to scale-out. IBM might lease 100 or 500 new mainframes annually, Sun could sell 500,000 workstations, and PCs reached 10 million units per quarter. Smaller models similarly trade maximum compute per machine for vastly more endpoints—potentially seven billion—with local MIPS that feel free beside liquid-cooled data centers.

  • Sinofsky’s best analogy was the telephone industry’s quality-of-service obsession. In 1994, he showed AT&T visitors a postage-stamp Lillehammer Olympics video running at roughly five to 15 frames per second, with audio supplied by a phone call; they laughed at its deficiencies. Today Netflix streams everywhere. Sinofsky substituted hallucination for QoS, while Casado mapped the same lesson onto models: imperfect systems already unlock coding and creativity that the old model could not.

  • Casado called DeepSeek “another step to basically AGI in your pocket” and said his reaction was not to short NVIDIA. Sinofsky likewise rejected that trade, arguing that the scale-out step expands TAM and enables specialized models for phones while some inference remains hyperscale. He compared the moment to browsers gaining JavaScript and suddenly enabling broad application invention.

  • Casado accepted the fiber-buildout parallel: institutions again buy infrastructure because they cannot easily access private software winners. Sinofsky’s WorldCom example supplied the contrast: WorldCom had roughly $40 billion in debt, whereas today’s principal builders are cash-rich cloud companies. Martin said NVIDIA can take a price dip and remain fine. Excess capacity might occur, yet he does not expect the same balance-sheet chain reaction or crash structure.

5. Useful metrics and enterprise controls replace generic model rankings

  • Asked whether parameters and broad coding tests should give way to scale-out benchmarks, Sinofsky said the old rankings were “silly to begin with.” Browser magazines once timed progressive image rendering because it was measurable, not important. Metrics will now follow applications: research systems need truth, footnote links and sources, pulling retrieval and vector databases back into prominence.

  • Casado expects sophisticated applications to orchestrate many models and fine-tune them for particular jobs. The raw model is not the application; persistence, configuration and workflow make the system more defensible and harder to substitute.

  • Sinofsky’s enterprise moat is readiness: single sign-on, filtering, feature controls and even disabling AI for individual users. Scott Belsky’s Adobe Firefly example carried the point—consumers may not care much that images are licensed, but enterprises do. Meeting those constraints early makes products sticky and creates a natural basis for pricing.

6. DeepSeek indicts restrictive policy more than incumbent AI labs

  • Casado’s “biggest takeaway” was that US policy had aimed at the wrong bottleneck: restricting open source, frontier labs, chips, software and model weights in the name of safety or containment. China produced and openly released strong work despite chip controls. His conclusion was categorical: fund domestic research, move faster and treat AI as a race the US can win unless regulation obstructs it.

  • Sinofsky pushed back on the metaphor while agreeing with the prescription: “The lesson is not Sputnik. The lesson is the internet.” The relevant Al Gore contribution, in his telling, was regulation that allowed the internet to flourish despite AT&T, WorldCom and AOL preferring more controlled structures. DeepSeek reflects worldwide technical diffusion, not merely a geopolitical surprise.

  • Encryption export controls had already demonstrated the problem: “You can’t outlaw math.” Attempts to export-control PlayStations because of their potential for weapons simulations likewise collided with global markets. In an already connected world, Sinofsky said the calendar time required for technical diffusion is “zero.”

  • Neither speaker treated DeepSeek as a crisis for NVIDIA, OpenAI or Anthropic. Sinofsky mentioned an estimate he had seen putting DeepSeek near 35% of OpenAI’s daily active users, but read the spike partly as frictionless experimentation. Casado called it a wake-up call for frontier labs to move quickly while remaining bullish on them, and argued that they should build applications and accept “coopetition.” Sinofsky said TAM would become “100x”—every endpoint, with revenue on both the app and developer sides.

Steven Sinofsky

The lesson is not Sputnik. The lesson is the internet.

Martin Casado

This is another step to basically AGI in your pocket.

Steven Sinofsky

There are always pockets of people innovating. WorldCom and AT&T did not predict that the internet was going to come out of universities.

Martin Casado

It really is the AI race, just like we went through the space race, and we need to win.

Steven Sinofsky

There’s no way this doesn’t play out like the internet.

Steph Smith

It’s been a busy few weeks. I don’t know about you guys, but my Twitter feed, all the podcasts—everything—has been DeepSeek everywhere. Maybe unsurprisingly, but what’s your TL;DR? Let’s just start there in terms of what came out, and maybe also your take on why it blew up in the way it did, because we’ve seen lots of releases in the last, let’s say, 2 years since ChatGPT.

Martin Casado

The quick overview, of course, is that, out of essentially nowhere, a small quant hedge fund—a quasi computer science research organization—in China released a whole model. Now, those in the know know that it didn’t just appear. There’s been a year and a half or so of buildup, and they’re really good researchers. Nothing was announced, but it appeared to take the whole rest of the world by surprise.

I think there were 2 big things about it that really caught everybody’s attention. One was, “How did they go from nothing to this thing?” It seems to be a comparable level of capabilities to everybody else. Then this number got thrown around that it only cost $5 million.

Steph Smith

Yeah, the number was $6 million.

Martin Casado

The number is irrelevant, but it turns out they wrote a paper and said, “Hey, we innovated in this particular set of things on training.” Even here, everybody’s like, “Oh, well, that was pretty clever.”

Then, because of the weirdness that we don’t need to get into around the financial public markets and how this whole thing happened on a Friday, the whole weekend was everybody whipping themselves into a frenzy so they could wake up Monday morning and trade away $1 trillion of market cap. That seems to be a complete overreaction and craziness, but that’s not what we’re here to talk about.

Steph Smith

To your point, there are a lot of moving parts here and a lot to consider. It’s actually a fairly complicated situation. There has been this view that the traditional, non-reasoning LLMs were starting to maybe asymptote around GPT-4, and there hadn’t been a big advancement. Then there was going to be this new breath of life: OpenAI released a reasoning model, which is o1, and everybody was very excited about that.

In this grand tapestry, you had all this excitement about o1 and how that was going to drive compute costs and NVIDIA. Then R1 comes out and looks pretty good, and all of a sudden people are saying, “Well, if you can do it just as cheaply, is this actually going to drive the next wave?” There was a lot of buildup to o1, which lent to the R1 hype, and then people didn’t really know what to think about it.

I agree with you that this was a total market overcorrection. It’s also worth pointing out that, in addition to people saying, “Wow, this is a great model,” there were a lot of theories and rumors: “Maybe this is the CCP doing a scoop. Maybe it cost a lot more. Maybe this was very intentional. It was right by Chinese New Year.” There are just a ton of rumors, so maybe we’ll do our best to dissect everything going on.

Yeah, maybe let’s just do that, because to both of your points there was a lot here: the performance element, quotes around cost, the China element, virality—it hit number 1 in the App Store—and shipping speed. Martin, you shared that they released an image model shortly after, and you mentioned it was released on a Friday, so there’s a huge mixture of people reacting, some who know what they’re talking about and some who don’t, quite frankly. We’re about 10 days or so out from this release, which, by the way, as both of you said, was the R1 release. There was the V3 release 2 months ago or so, which was the base model. So now that we’re a little bit further out, what’s the signal from the noise?

Martin Casado

So maybe I’ll give you the lens of, “Chinese people are smart.” There’s one lens that I hold, which is that China has great researchers. DeepSeek has actually released a number of state-of-the-art models, including V3, which is probably a more impressive feat. It’s almost like a ChatGPT-4 model.

By the way, to create one of these chain-of-thought reasoning models, you need to have a model like that, which they had done and we had known about. All of the contributions they’ve made have been in the public literature somewhere; nobody had just aggregated them.

This is a very smart team that has been executing very, very well for a long time in AI. They are some of the top researchers. The fact that they spent $6 million just doing the chain of thought is not out of whack with what Anthropic has now said they’ve spent and what OpenAI has said they’ve spent.

This is a meaningful contribution from a good team in China, so it means something and we should respond to it. Some of the outcry is warranted. I do think that we should respond to it, but I don’t think for the reasons a lot of people are saying.

Steven Sinofsky

I completely agree with that. In fact, you also saw people outside of that team in China piling on to try to make it more intergalactic than it was. My favorite example is my old friend Kai-Fu Lee coming out on X and saying something about how this is why he said 2 years ago that Chinese engineers are better than American engineers.

To your point about people viewing this as reaching some cosmic level of progress, the previous base models—the GPT lineage—seemed to asymptote around GPT-4. But I think what’s super interesting about that is that the asymptote was true if you looked at it through the lens of the function that everybody was optimizing.

To my view, this was a crazy, hyperscaler view of the world: We need more compute and more data, more compute and more data, and we’re just on that loop. A lot of people from the outside were saying, “Well, you’re going to run out of data.” As a microcomputer person, I would say that, at some point, you’re going to end up breaking the problem up to the 7 billion endpoints of the world, which will have vastly more compute than you can ever squeeze into 1 giant, nuclear-powered data center.

A lot of what they did was a step-function change—not necessarily an improvement, just a change in the trajectory. That, to me, is the part where the hyperscalers needed to take a deep breath and say, “Okay, why did we get to where we were?” You were Google, Meta, or OpenAI, funded by Microsoft, and you all had billions and billions of dollars. You saw the problem through the lens of capital and data. Of course, you had English-language data, which there’s more of than anybody else’s, so you could keep going.

When Microsoft was small, we used to decide whether things were small problems, medium problems, or large problems. At one point, we started joking that we had lost the ability to understand small and medium problems and solutions. We only had large, which was trivial, and then huge and ginormous. Our default was ginormous because we thought, “We could do it and no one else could, and that’s a strategic advantage.” I feel like that’s where the AI community in the US, particularly—or the West, if you will—got just a little carried away. It was like every startup that has too much money: the snacks get a little too good.

Steph Smith

I’ve heard 2 theories about why they were able to do this. One of them is this constraint theory you mentioned, which I think is actually very true: We’ve just been using the blunt instrument of compute and the blunt instrument of all data, and we haven’t thought about a lot of engineering under constraints.

There’s a second theory I heard. I don’t know if it’s true, but it’s tantalizing. The reason people have said that V3 is so good is that it has access to the Chinese internet as well as the public internet, which is somewhat isolated. We don’t really have access to the internal Chinese internet, and we certainly don’t train from it, as far as I know. They do, so it could be the case that both things are true: They could have had a data advantage, and they definitely had the engineering-constraint advantage you mentioned.

Martin Casado

Even on the data, their starting point is the Chinese internet per se. That has much more structure to it; it’s a much better training set. Insofar as human-annotated data is important here—and for chain of thought, you do want experts saying, “Here’s how I would reason about a problem”—this whole chain of thought is basically, “What are the reasoning steps about something?”

If you want to look at a place to arbitrage really smart, educated people at a relatively low cost, it’s hard to beat China globally. They definitely have access to a bunch of potentially highly educated, annotated data, which is very relevant here.

I happen to be of the belief—it sounds like you are, too—that this did not come out of nowhere. It’s not a scoop. It’s a great team taking advantage of what it has. But there are still things that are very significant about it that are worth talking about. For example, the license is very significant, and the fact that they decided to release the reasoning steps is very significant. Maybe it’s worth talking through those, too.

Steph Smith

Yeah, please, because I do feel like those are 2 things you’re not seeing headlines about. You’re seeing headlines about all the other things we talked about. You mentioned the reasoning traces: Those were released, whereas the comparable o1 traces were previously not. Then there’s the open-source license. Let’s talk about this.

Martin Casado

There are 2 things that are pretty remarkable about DeepSeek R1 that have implications for adoption. We haven’t seen a license this permissive recently for an open-source model. It’s basically an MIT license—1 page, and you can do anything. It’s free as in free beer, for real.

At a16z, we have one of the largest portfolios of AI companies, both at the model layer and at the app layer. I will say that any company at the app layer is using many models. It’s not just 1 model. I have yet to see the GPT wrapper; they’re all using a lot of models. They do use open-source models, and licenses really matter. This is definitely going to result in a lot of proliferation.

The second thing is that a reasoning model actually thinks through the steps of the problem and uses that chain of reasoning, or chain of thought, to come up with deeper answers. When OpenAI released o1, they did not release that chain of thought. We don’t know why they didn’t do it, but it turns out that, if you have access to that chain of thought, it allows you to train smaller models very quickly and very cheaply. That’s called distilling.

The general term “distilling” in the LLM world means that you have a teacher model train a student model that’s much smaller. It turns out that you can get very, very high-quality smaller models by distilling these public models. The implications are both that this is more useful for somebody using R1 and that you get a lot more models that can run on much smaller devices. You just get more proliferation that way. It’s actually a very big step when it comes to the proliferation of this model.

Steven Sinofsky

There’s a tendency to either peg yourself at “It should just be open” without really defining it, which I think is important in this case. Because of where they came from and because they don’t have a business model, part of what was unique about this was that it was a hedge fund—almost like a side project, but not really a side project. It was a whole thing, and it had the effect of, “We’re just going to give the whole thing away.”

The rest of the companies are still trying to figure out their revenue models, which I would argue was probably premature. It starts to look a little like, “Let’s charge for a web server,” and it turns out that the business of selling HTTP service is not a great business.

Everybody got focused on the first breakthrough, which was the LLM. In a lot of ways, if you look back at the internet, that’s exactly what happened. Everybody got very focused on monetizing the first part of the internet, which was HTML and HTTP. Then along came Microsoft and a bunch of other companies to say, “That’s not the best layer to monetize. In fact, there might not be any money in that layer. The real money is going to be in shopping, plane tickets, television, and a bunch of other things.”

Even other companies, like AT&T, got wound up trying to monetize an even lower layer. That’s not how you’re going to get to 7 billion endpoints.

The licensing model really matters because what’s going to happen is that there’s going to end up being some level of standardization. I don’t know where in the stack or at what level, but there is going to be some level of standardization. The licensing model for the different layers is going to start to matter a lot.

Anyone who was around during the internet remembers the battles over the different GPL, BSD, and MIT licenses. You remember this very well, right? When you were doing a distribution, you were going to have to release it, and it turned out that even your distribution—what part of it you released and how you released it—was a huge issue because it could make or break a whole approach.

I think the US industry lost sight of the importance of that because it got so used to this model where “open” just means we’re a business and we pick and choose what we throw out there as evidence that we’re an open company.

Martin Casado

I totally agree, and I don’t think that view is aligned with how technology has been shown to evolve in an era where there’s no cost for distribution. Before, when there was a cost for distribution, all those things were sort of irrelevant because, even if it was free, you still couldn’t figure out how to get anybody to use it.

I do want to take the other side of this because I actually tend to agree with you. What you just said could be the case: The model may be the wrong place to focus, and everybody thinks there’s a lot of value there, so they’re playing all these cute games with openness instead of focusing on distribution.

But there’s another view, which is that the models really are pretty valuable. In particular, the model itself isn’t an app, but it could be that, if you’re building an app, you need to vertically integrate into the model.

If I’m building the next version of ChatGPT—or, as we saw today, Deep Research launched—it could be that the apps actually require you to own the model. In that case, DeepSeek is less relevant because they’re not building apps, and the impact on OpenAI or Anthropic isn’t as great.

There’s a fork that we don’t know the answer to. Fork number 1 is that maybe the models get commoditized and you need to focus at the app layer, in which case the license doesn’t matter. Or the models really matter up the stack, in which case the whole DeepSeek phenomenon really isn’t as impactful an event as people are making it out to be.

Steven Sinofsky

I’m going to build on that just because I want to say you’re right both times. The variable is T, and the variable is time.

There’s no way this doesn’t play out like the internet. It just has to. What we saw was that, for a while, building 1 app seemed like a crazy thing because you had to own Windows and Office. Then a new app came along that didn’t own any of those things: search.

That’s why a lot of people, because of their age and what they lived through, immediately jump to, “These LLMs are going to replace search.” It turns out that’s actually going to be really, really hard because there are a lot of things search does that models are bad at—really bad at.

What’s going to happen is that a new app is going to emerge. When the new app emerges, it’s going to get vertically integrated. The research app is a great example of that. Then other apps are going to spring up, and you’ll have Google Maps, search, and Chrome. It goes back and eats the things it couldn’t do before.

I really feel like that’s the trajectory we’re on now. It’s still a matter of where and what integrates, but the apps that ended up mattering on the internet literally didn’t exist before the internet. I think that’s what people are losing sight of.

The same thing happened with mobile. There were no social apps. Fine, I get it—there was GeoCities and a bunch of other things—but people get so caught up in thinking that the new thing is going to replace something. The problem with zero-sum thinking is that it’s so dangerous.

You can think of everything as the spectrum. When something new comes along, the whole spectrum gets divided up differently. That’s what Google said when it bought Writely: What are people going to do on the internet? They’re going to type stuff. What are they going to type? They’re going to type with other people.

We’re seeing this happen now. Someone will come up with a model that does something in a consumer space, like text-to-image, and over time people will say, “That’s kind of like Canva.” It’s the AI-native version of an existing app.

It looks like Canva, Word, PowerPoint, or Excel, but what’s important is that they’re actually different. Nothing is ever going to be PowerPoint again. The whole reason PowerPoint existed was to render something that couldn’t be rendered before.

The whole product is 3,000 different formatting commands. That’s not a number I made up. It’s 3,000 ways to kern, nudge, squiggle, color, and style things. It turns out you don’t need to do any of that in AI, so the whole product isn’t going to have any of those things.

Then it turns out all those things make it really hard to make it multiuser. When Google comes along and bundles its competitor that’s going to replace it, it’s totally focused on sharing.

Steph Smith

Steven, let me ask you this. You said something really interesting, which is that this has to pan out like the internet. You’ve used examples of different companies and the waves—the mobile wave and the cloud era. All of those are things we can learn from.

I just want to probe you: Is there something different here? To bring it back to DeepSeek, this means it’s very important to realize that China is a very credible player. But I don’t think that R1 itself, as a standalone, is going to have that deep of an impact.

Martin Casado

There are actually parallels when it comes to capital buildout that you see in AI. There’s a special parallel that Marc Andreessen reminded me of, which people don’t tend to see as well.

In the early days of the internet, in the mid-to-late 1990s, a lot of investors and big pools of money—banks or sovereign wealth funds—wanted exposure to the internet, but they had no idea how to invest in software companies. What were these new software companies? Who were these people? They were all private companies.

What did all of them do? They invested in fiber infrastructure. There was a whole dot-com-era fiber buildout. We’re starting to see this happen again: A lot of banks and big investors are saying, “We want to build data centers,” because they don’t know how to invest in startups. We know how to invest in startups.

On the one hand, you could say that we’re going to see all of this capital expenditure go into physical infrastructure, and therefore we’re going to have another fiber-glut equivalent—a data-center glut.

The counterpoint, where I think things are different, is that at the time of the fiber buildout, you had one company that was cooking its numbers and had a ton of debt to build all of this out. When the price of fiber dropped, that company went out of business, which caused a huge issue.

We have a much better foundation for the AI wave—much, much better. The primary investors are the big 3 cloud companies, which have hundreds of billions of dollars on their balance sheets. Even if all of this goes away, they’ll be fine. NVIDIA can take a price dip. NVIDIA will be fine.

I don’t think we’re heading toward the same kind of glut and crash that other people are predicting. It’s appealing to draw parallels to the internet, but I don’t think that parallel is there.

Steven Sinofsky

I’m completely with you on that. This part is going to look like the amount Google invested in the early 2000s, the amount Facebook invested 5 years later, or the amount that people forget Microsoft poured into Bing. I don’t know—$30 billion or $40 billion?

It’s still number 3 or whatever, but it doesn’t matter. I would bet—and I don’t know this as a fact—that Meta is spending more money on VR than it is on AI right now, just to show you. Maybe Apple is, too.

Apple is probably spending some amount that’s bigger than gargantuan. It really isn’t about the investment profile or who is going to win or lose financially, and I think that’s a super important point to hammer home.

There’s a certainty that nobody’s going to come out of this unscathed, but the scathing is not going to be at all what anybody thinks. It’s not going to be like what it was before. The structure that people remember—WorldCom, for example—had something like $40 billion in debt. There were companies we’ve all forgotten about that went bankrupt during that era. There was one in Seattle whose name I’m forgetting, but it had $20 billion just disappear.

Steph Smith

To your point, these companies have had so much cash on their balance sheets. They’ve been waiting for a moment to invest in the next generation, which also contributes to their willingness to scale up as much as they did.

Let’s talk about that. In your article, you talk about the difference between scale up and scale out, and the natural tendency in these early parts of the wave is to scale up, when really there tends to be a shift toward software going to zero cost.

Steven Sinofsky

When you’re big, you want to double down on being big, so you start building bigger and bigger computers that don’t distribute the computation elsewhere. If you’re IBM, you say, “The next mainframe is another mainframe that’s even bigger.” If you’re Sun Microsystems, you keep building bigger and bigger workstations. If you’re Digital Equipment, you build bigger and bigger minicomputers.

All along, you’re just doing more MIPS, in the acronym sense, than the previous maker, for less money. Then the microcomputer comes along, and not only does it do fewer MIPS, it costs nothing. There are going to be gazillions of them.

You went from an era when IBM would lease 100 or 500 new mainframes in a year, and Sun might sell 500,000 workstations, to selling 10 million computers in a quarter. That’s scale out: less computing, but at many more endpoints. It’s also a deep architectural win because it gives more people more control over what happens and reduces the cost.

Today, the most expensive MIPS you can get are in a nuclear-powered data center with liquid cooling and all the rest. The MIPS on my phone are free and readily available for use. That has been a blind spot for the model developers.

They all do it now. I run Llama on my Mac, and the first time you do it, your mind is blown. Then you start to think, “That’s just how it should happen.”

You look at Apple and its strategy—the execution hasn’t been great, but the idea that all these things will surface as features popping up all over my phone, and that they’re not going to cost anything and my data isn’t going to go anywhere—that has to be the way this evolves.

There will be some set of features that are only hyperscale cloud inference, just as most data operations happen in the cloud now. But most databases are still on my device.

Martin Casado

I’m smiling because this is the story from a microcomputer guy. I’ll tell the story from an internet guy. Do you remember the switching wars?

Steven Sinofsky

Absolutely.

Martin Casado

For the longest time, you had the telephone networks. They were perfect: They converged in milliseconds, never dropped anything, and gave you guaranteed quality of service. Then came the internet. You had none of those things. Convergence took minutes, it dropped packets all the time, and you couldn’t enforce quality of service.

There were crazy wars at the time. People said, “Why are you doing this internet stuff? It’s silly. We know how to do networking.”

What the switching people—the telephone people—didn’t understand was what happens when you have best-effort delivery and then enable the endpoints. They just didn’t understand that the value could be at the endpoints. They couldn’t think that way; they needed the value to be in the network. That really brought us the internet.

I think the exact same thing is playing out. I see it all the time. People look at these models and say, “They hallucinate,” or, “They’re not correct at these things.” But they enable an entirely new set of things, like creativity and coding. It’s an entirely white space, and it’s going to grow very quickly. To assume that they somehow need to fit the old model is irrelevant to where they’re going to go.

Steven Sinofsky

What I do is substitute QoS for hallucination. I was going to meetings in the 1990s with all these pocket-protector AT&T people who would show up and yell at Bill Gates, “QoS! QoS!” We had to go look up what QoS was because not only were we not using TCP/IP, the network we were using never worked. It was a PC-based network, and IBM was being stuffy.

I was talking to a networking genius, so I should explain the ping of death, but it was hilarious. They would just sit there and talk about QoS. I literally did this: They were telling me about QoS and wouldn’t shut up about it, and I didn’t know what it was, so I walked them over to my office.

This was in the winter of 1994. I said, “Look, here is a video of the Lillehammer Winter Olympics playing on my Mac.” It was a postage stamp—not the size of an iPhone icon, but close. They said, “That’s 15 frames a second.” I said, “I know. It’s usually 5.”

They asked, “Where’s the audio?” I said, “If I want audio, I just call up this phone number on your system.” They laughed at me. Here we are, of course, using Netflix on every device all over the world.

They couldn’t understand that these paradigms—where the liabilities either don’t matter or just become features—could work. That’s what gave birth to Cisco.

Martin Casado

Exactly. They just said, “This is how we’ve been doing it, and it all works.” It only worked in crazy, weird universities and in the Defense Department, for people who cared. Now that’s all we use.

I want to tie this back to DeepSeek because the reason we’re getting so excited about it is that we’ve seen these types of things before. We’ve seen things like DeepSeek come out before, and it’s not zero-sum. It doesn’t replace the old thing; it’s a component of the new thing.

We still haven’t even envisioned the new thing. It’s like the internet is just coming right now. Our excitement is about the new thing to come. When I saw DeepSeek, I thought, “Amazing. This is another step to basically AGI in your pocket.”

These can run on small models, and it shows that we’re moving forward. My reaction was not, “I need to short NVIDIA,” or whatever. I think that’s the wrong answer.

Steven Sinofsky

I read the “Let’s short NVIDIA” blog post that flew around that whole weekend, and I thought, “Are you kidding me? Jensen is a genius, and his company is filled with geniuses.” What about the TAM that just expanded? Don’t you like that?

The scale-out step just happened, so now you can see everybody doubling down. The point you made earlier is super insightful and really important: the enabling of specialized models. That’s what’s going to end up being on your phone, and that’s what’s going to enable the app layer to really exist.

To me, this is the equivalent of the browser getting JavaScript. Once the browser got JavaScript, you could do anything you needed without going to some standards body or building your own browser. I think that’s where we are right now.

Steph Smith

One follow-up: If you think about how this has progressed to date, I feel like the benchmarks have always been, “Which model has the most parameters?” and, “How is it doing on this coding test?” That isn’t necessarily representative.

Should we expect a different set of benchmarks or things that we judge these models by? Should we be looking at what device the model can fit on and how much it costs? Or should we just be looking at the app layer? Does there need to be some kind of shift that moves us away from “bigger is better,” as you’re saying, and toward something that represents scale out?

Steven Sinofsky

I thought all those benchmarks were silly to begin with. To me, they seemed like the benchmark we used to do with browsers: how fast they could finish rendering a whole picture.

Marc Andreessen invented the image tag in the browser, and the neat thing they did in their implementation was progressively render the image. What that did was empower stopwatches all over the world—and magazines—to write about who finished rendering a picture faster.

You can measure that today, but it doesn’t matter. I think those benchmarks will all go away, and we’re quickly going to get to, “What does it actually do?”

The measure that’s going to start to matter will depend on the application people are going after. Take the research apps that just appeared this week. Everybody has a research app now. When you’re doing research, the metric that matters is truth.

All of a sudden, we’re back to hallucinations, but now you’re giving footnote links and sources. What’s really happening under the covers is that it’s a little less generative and a little more information retrieval. Vector databases, looking things up, and reproducing information matter.

We’re going to get to a point where people develop something probably along the lines of ImageNet, and they’re going to generate thousands and thousands of routine tests that ask, “Is this true?”

Martin Casado

This is a total aside, but you reminded me of a weird historical erratum: Andreessen made the image tag, so in a way he’s also the grandfather of some AI.

CLIP, which is an AI model, will take an image and describe it. The way it does that is by using the metadata and image tags.

Steven Sinofsky

That’s right. He created the metadata to do this.

Martin Casado

Back on the topic of images, here’s one thing I’ve noticed working with these companies: These models are actually pretty magical by themselves. If you have a big model and just expose it, people use it. That’s very different from computers. You just put the model out there, and people use it.

The problem is that all the other models catch up very quickly because they distill so well. That’s not defensible in a way. The defensible companies I’ve seen put out a compelling model, and then, once users are engaging with it, they find ways to build an app around it that’s actually retentive.

It starts converging on something like PowerPoint, where it becomes more stateful and requires configuration. That tends to be very defensible. The applications that use models use lots of models, and they do a lot of fine-tuning in those models.

I would say that the last 2 years have been the story of the large model. They’ve been magical. People use them, people really like them, and the first time you’re in ChatGPT, you think, “This is amazing.”

Now I think we’re in the era of the workflow around models. These are stateful, complex systems, with many models powering apps in more sophisticated ways.

Steven Sinofsky

The many-model point is a great one. This is what happened with user interfaces. The whole notion of the user interface that IBM put forward was derived from its green screens and 3270s.

They made a set of rules about how the UI should work: 40 characters by 40 characters, this is the F10 button, and so on. Then people started building all sorts of UI frameworks. It looks exactly like the browser today, where there are a zillion frameworks on the endpoint. You pick and choose, do what you want, and invent a new calendar dropdown if you want. You don’t have to waste your time; it’s really up to you.

That aspect of creativity is extremely important to applications. What else is going to matter a lot for apps to be differentiable and to have a moat? The apps are also going to embrace the enterprise.

For better or worse, one of the lessons we keep learning is that, if you want to get adoption in the enterprise, you have to do a bunch of work to turn off parts of your app, filter parts of your app, disable things, or whatever it is. The smartest entrepreneurs are going to recognize the need for single sign-on at the beginning, role-based access control, and all of those things.

Every time, it turns into a way to price the product. It’s not super hard, and I think a lot of dumb stuff has been done around AI alignment, censorship, whose point of view it is, and all this other stuff. There’s now a whole industry that wants to show up and tell you all the things they don’t want out of AI.

The smartest entrepreneurs are going to get ahead of that, and they’ll be there to sell it, because that’s actually enormously sticky in the enterprise. We’re going to see smart productivity tools embrace that immediately. It could be at the most granular level, like turning it off for specific users.

Steph Smith

We had Scott Belsky at Speedrun recently, and to your point, he talked about Adobe. Someone said, “You have all these licensed images for Firefly. Do consumers really care about that?” He said, “Honestly, not really. But you know who does care? The enterprise.”

Those are 2 different modalities, and founders are going to have to figure that out. I do want to touch on DeepSeek as a Sputnik moment. That can be viewed through the lens of geopolitics—the US and China—but if you think about Sputnik, it wouldn’t have been a moment if Kennedy hadn’t made his moon-landing speech and if we hadn’t actually gotten there. In other words, changes had to be made.

Let’s say you’re in a boardroom and you’re an adviser. I don’t want to talk to the board; I want to talk to the US government.

Martin Casado

For me, the biggest takeaway from DeepSeek is none of what we’ve talked about so far. The biggest takeaway is how blind our policies around AI have been.

Our previous policies around AI have been, “We can’t open-source because it will enable China. We have to limit our big labs. We have to put all this regulation on top of it.” The reason is safety and all this other stuff. Then there are export controls: We can’t enable other countries, so we put export controls on chips. We’ve talked about putting export controls on software and weight limits.

That was our entire policy. For me, the biggest takeaway from the whole DeepSeek thing is that this is the wrong way to do policy. China has a lot of very smart people. They’re incredibly capable, they’re great researchers, and they can build things as well as we can. They can open-source it.

We did not enable them. They did this even with export controls on chips. Basically, all of our activity has been for naught. What we should be doing is funding and investing in our research labs, and going as fast as we can.

It really is the AI race, just like we went through the space race, and we need to win. We have everything we need to win. The only thing in our way is our own regulatory environment.

Steven Sinofsky

The lesson is not Sputnik. The lesson is the internet. Al Gore famously claimed to have invented the internet, but what he really did was invent the regulation that allowed the internet to flourish.

It could have been a Sputnik moment, and they could have looked at the internet and said, “Oh my God, this is Sputnik,” and then tried to turn it into what AT&T and WorldCom wanted. They were lobbying to make that happen. AOL wanted it to happen that way, too.

Instead, they ignored that and went with what made the internet strong to begin with. What gave us this DeepSeek moment was the strength of the worldwide technology community. As much as people want to own it and be the singular provider, that’s not going to work.

I think it’s a Sputnik moment in the sense that it’s a wake-up call for half the world. It isn’t a geopolitical wake-up call. It’s not about war; it’s about technology diffusion.

We’ve had so many misfires since then. We had the whole encryption war, where we tried to put export controls on encryption. People thought we were being silly as an industry when many of us championed this. You can’t outlaw math. It turns out that it is outlawing math.

The fact that DeepSeek used those chips also shows that it’s very hard to put export controls on things in the global economy. We’ve seen this before. We were going to export-control PlayStations because you could do weapons simulations on them. The PlayStation was the first to use SGI technology, and we were going to export-control that because we couldn’t let it get into Saddam Hussein’s hands.

It was a total failure because global markets are global markets. We’re much better off investing in our own infrastructure, which we did at the time. We did a great job of that.

It’s a great analogy with the internet, and we should be doing exactly that again. Some politician needs to stand up and be the Al Gore of this moment. I think we will get that.

There is now a wake-up call. The futility of the past 4 or 5 years of this kind of thing is very, very clear. I mean that even more broadly than you were saying: the people who wanted to control this technology at a very granular level, all these think tanks and institutes that were aligned, the number of books written, the number of academic departments started, and the number of assaults on technology companies to align them.

There were whole meetings in Switzerland about aligning with world leaders. That’s just not how anything evolves. The biggest lesson for computing, starting in 1981 with the IBM PC—or frankly in 1977 with Apple—has been the creativity at the edge and simply enabling it.

The problem regulators had was that they had never faced regulating a connected world before. The other lesson from DeepSeek is that the world is already connected and already native in all of this stuff. The amount of actual calendar time it takes for something to diffuse technically is zero.

DeepSeek was, I think, 35% of the daily active users of OpenAI—that’s the number I saw this morning—which is a giant spike because all the same people are just trying it out. There’s no friction. It takes no time.

It’s unbelievably exciting to be part of what’s going on right now. We don’t need to throw water on it or be party poopers.

Martin Casado

I personally don’t think this is a crisis moment for OpenAI or Anthropic. Apps are hard to build. Right now, the apps they’ve put out are very complex, and they actually know their users and have very specific use cases.

For them, it’s a wake-up call that they can’t slouch and have to move very quickly, but I’m still very bullish on our AI labs. I think they can stay ahead, too.

There’s a view of DeepSeek as a crisis moment for NVIDIA, OpenAI, and Anthropic. I don’t buy any of that. I think it’s more of a wake-up call for the regulatory environment. We should all acknowledge that there’s going to be global competition, and we need to stay ahead.

The right reaction from all of these frontier companies is that they should start building apps. The best feedback loop for building a great platform for other people to use is to build apps yourself.

There’s this whole conversation about competing with your partners. Our industry is coopetition through and through. That’s Andy Grove’s lesson. Everybody should be prepared for the big players to compete with them, but history has shown that’s no surefire success.

If the TAM is 10 times bigger, there’s just a lot of room for a lot of people.

Steven Sinofsky

People forget that Microsoft spent more than 10 years as a distant number 3 in the applications business. It was a platform shift that all the other players ignored that caused it to win.

The TAM is going to be 100 times bigger. It’s going to be every endpoint. The revenue is going to come from the app side, and then there will be a developer side. It will just be a different pricing model for different sets of scenarios, but it’s going to be there.

Steph Smith

Everything is rising right now, since this is this positive, growing world. Do you have any thoughts, just quickly, on the fact that this came from an algorithmic hedge fund, a quant? Is that different from what you expected, or does it signal that more people can participate?

Martin Casado

It’s a good reminder that there are always pockets of people innovating. WorldCom and AT&T did not predict that the internet was going to come out of universities. They did not think that CERN was going to invent the protocols that became foundational.

Steven Sinofsky

They also didn’t expect a failed corporate lab to develop TCP/IP, which became the standard. It’s not just any old lab. It wasn’t IBM’s lab; it was literally a lab they had all but shut down because it had failed, just down the street at PARC.

Remember, SRI was one of these places that you don’t even think about anymore. Most of this isn’t going to be in any history that’s written 5 years from now. That’s the excitement.

What DeepSeek Means For The Future Of AI | Tech Veterans Weigh In | BidClub