[BidClub_]
No Priors · · 23 min

No Priors Ep. 100 | With Sarah and Elad

Sarah GuoElad Gil

YouTube
TL;DR
  • DeepSeek was a real open-source reasoning advance, not a $5.5 million repudiation of frontier compute. Elad said comparable final runs already cost roughly $5–10 million and suspected hundreds of millions preceded DeepSeek’s distilled result, making NVIDIA’s roughly 20% selloff “a bit unwarranted.”
  • The larger investment signal is a 180x decline in GPT-4-equivalent inference cost per token over 18 months. Benchmark gaps are narrowing, yet Sarah stressed that even a multiple of DeepSeek’s reported cost remains far below a multibillion-dollar or Stargate-sized entry price—a genuine “narrative violation.”
  • Frontier leadership still buys distribution, workflow lock-in and a possible recursive research advantage. Better models can generate synthetic data, label data and write code for successors; whether that produces “liftoff” is uncertain, while widely available base models could instead become “a big leveler.”
  • OpenAI’s Deep Research immediately raises the analyst bar, but its outputs cannot be treated as authority. Elad said he would compare median analyst or intern work against it because “the comp is hard”; Sarah found it especially useful for surveying unfamiliar domains, but in domains she knows well, users are “really going to have to audit the outputs.”
  • AI’s emerging control over knowledge makes model plurality and open source strategically important beyond economics. Elad compared AI’s blind spots to Gell-Mann Amnesia: people spot errors in familiar fields, then trust the same system elsewhere; Sarah warned that an AI combining search, social networks and media into “one single device that you interrogate” creates censorship and propaganda risk.
  • Stargate reflects uncertainty about scaling returns, not evidence that capital has stopped mattering. Elad said he could not imagine an AGI lab not wanting “the biggest cluster they could have” if it were free or financeable, although he considered continued pretraining gains likely to become less efficient.
  • Their 2025 map favors foundation-model consolidation, vertical AI, agents and autonomy, with robotics still at the proof stage. Harvey, Decagon, Sierra, Cognition, Tesla, Waymo and Applied Intuition illustrate the opportunity set; consumer resurgence may come from cheap low-latency models, while biology, materials and health could benefit from smarter domain-specific data generation.
Digest · the substance, structured for research

1. DeepSeek advanced the frontier without rewriting compute economics

  • Elad’s read was that DeepSeek mattered but remained “roughly on trend”: it delivered a state-of-the-art Chinese, open-source reasoning model and genuinely novel reinforcement-learning techniques that other laboratories were beginning to adapt.

  • The reported $5.5 million covered a final run, not the entire program. Elad said knowledgeable practitioners put comparable runs around $5–10 million and believed DeepSeek likely spent hundreds of millions on tooling, data, experimentation, pretraining and post-training; NVIDIA’s roughly 20% decline therefore looked “a bit unwarranted.”

  • Sarah emphasized the release sequence: V3 appeared in December without crashing NVIDIA, while R1—a reasoning counterpart to OpenAI o1—produced the “narrative violation.” Post-training made the model useful, and a Chinese laboratory’s rapid catch-up challenged a 20-year US-versus-China technology-dominance story. She also noted that DeepSeek’s mobile app briefly became a top App Store contender; Elad thought the attention was more likely curiosity about the leading Chinese model than proof that consumers choose the cheapest capable model.

2. Falling costs commoditize capability but do not erase frontier value

  • Sarah’s pushback—worth keeping—was that even a sizable multiple of $6 million is not a multibillion-dollar or Stargate-sized entry price. Elad answered that the direction was already obvious: GPT-4-equivalent inference cost per token had fallen “180x, not 180%” in 18 months.

  • Elad pointed to Artificial Analysis’s independently rerun benchmarks: across reasoning, knowledge, math, coding, multilinguality and cost, leading models were getting “closer and closer” rather than separating. A new breakthrough might temporarily leapfrog the field, but convergence was the prevailing trend.

  • Frontier leadership still offers market share, optimized-prompt and tooling stickiness, plus better synthetic data, labeling and coding for the next model. Elad hedged the stronger “liftoff” claim; Sarah added that broadly available high-quality base models could instead be “a big leveler” for self-improvement.

3. Deep Research raises the analyst bar while weakening epistemic visibility

  • Elad said Deep Research immediately raises the bar for knowledge work and that he would compare the work of a median analyst or intern against it: “the comp is hard.” Sarah found it especially useful for surveying unfamiliar domains, building a comprehensive view and identifying experts.

  • Her reservation was its implicit authority ranking: when deciding which web claims and ideas were good, it often required users to “audit the outputs.” It could orient you, but its conclusions could not simply be taken as given.

  • Elad invoked Gell-Mann Amnesia: readers recognize errors when a newspaper covers their specialty, then trust its next page on an unfamiliar subject. He said AI could become a definitive information source while making its sources less evident. Sarah warned that a system combining search, social networks and media into “one single device that you interrogate” creates propaganda and censorship risk; she viewed a multi-AI, multi-company world as an offset and open source as especially important for civil liberties.

4. Stargate exposes uncertainty around scaling returns

  • Elad framed infrastructure demand as “uncertainty rather than risk.” Algorithmic efficiency, synthetic data and test-time scaling make capability gains difficult to forecast, but he could not imagine a serious AGI laboratory declining “the biggest cluster that they could have if it was free.”

  • That revealed his underlying call: pretraining should continue producing gains, though probably less efficiently. He offered no strong conclusion on capital-market depth or sovereign involvement, preserving those as open questions rather than treating Stargate’s scale as self-validating.

5. Vertical AI and agents headline the 2025 map

  • Elad expected partial foundation-model consolidation—especially in image, video, voice and secondary models—if the FTC became friendlier, alongside new races in physics, biology and materials. His central application call was an “era of vertical ops,” citing Harvey, Decagon, Sierra and increasingly agentic products such as Cognition.

  • Sarah defined agents pragmatically as systems that “do multi-step tasks successfully,” manage state and act beyond content generation. Security, support, SRE and coding already supplied examples; copilots should naturally expand by taking on more work and getting better at handling failures.

  • Elad expected self-driving to command attention through Tesla and Waymo, with Applied Intuition as his “dark horse.” Sarah expected technical proof of breakthroughs in robotics and generalization this year, “though not deployments”; Elad similarly anticipated only a “glimmer of how this thing will work versus the whole thing.”

  • Elad predicted a consumer resurgence, while Sarah said small, low-latency models could unlock free experiences, whether local or web-based in the browser; edge compute mattered only when it was transparent to the user. Elad also argued that reasoning may improve reliability as much as task complexity, so apparent technical failures deserve repeated reassessment. As innovation diffuses beyond the tip of the spear, he expects smarter domain-specific data-generation and data-capture strategies—potentially including biotech innovation—in biology, materials and perhaps health.

Sarah Guo

Hey listeners, welcome back to No Priors. This episode marks a special milestone: today is our 100th show. Thank you so much for tuning in each week with me and Elad, and it's been an exciting last couple of weeks in AI, so we have lots to talk about. Why don't we start with the news of the hour—or really, the last month at this point? DeepSeek: what's your overall reaction?

Elad Gil

DeepSeek is one of those things that's really important in some ways, and then also kind of what you'd expect would happen from a trend-line perspective. I think there was a lot of interest around DeepSeek for 3 reasons.

Number 1, it was a state-of-the-art Chinese model that seemed to have really caught up with a number of things on the reasoning side and in other areas relative to some of the Western models, and it was open source. Number 2, there was a claim that it was done very cheaply, so I think the paper talked about a $5.5 million run as sort of the end. Lastly, I think there's this broader narrative of who's really behind it and what's going on, and some perception of mystery, which may or may not be real.

As you walk through each one of those things, I think on the first one, having a state-of-the-art open-source model with some recent capabilities built in, they actually did some really nice work. You read through the paper, and there's some novel techniques in reinforcement learning that they worked on, which I know some other labs are starting to adapt. I think some other labs had also come up with similar things over time, but it was clear they'd done some real work there.

On the cost side, everybody that I've at least talked to who's savvy to it basically views every final run for a model of this type as roughly in that kind of dollar range—$5 million to $10 million, something like that. The real question is how much work went in behind that before they distilled down this smaller model. My sense is that everybody thinks they were spending hundreds of millions of dollars on compute leading up to this.

From that perspective, it wasn't really novel, and I think that sort of 20% drop in NVIDIA stock and everything else that happened as news of this model spread was a bit unwarranted. The last one was just speculation about what's going on: Is it really a hedge fund? Is something else happening? It felt a little bit speculative. There are all sorts of reasons that it is exactly what they say it is, and then there are some circumstances in which you could interpret things more broadly. That's kind of my read on it. What do you think?

Sarah Guo

Yeah, I think it's interesting, especially the delayed reaction to it. To your point, it's also what you might expect, especially given the historical precedent with GPT-3.5 and then ChatGPT. DeepSeek V3, the base model—the big AI model pretrained on a lot of internet data to predict the next token—was out in December, right? NVIDIA stock did not crash based on that news.

I think it's interesting to recognize that people obviously do not just want the raw likelihood of the next word in a streaming way. The work of post-training and making it more useful for human feedback or more specific data, like high-quality examples of prompts and responses, just like we've seen with the chat models such as ChatGPT—the instruction fine-tuning that made this such a breakthrough experience—really mattered.

Then, as you said, the narrative-violation release of the R1 reasoning model as a parallel model to OpenAI o1, I think that was also the breakthrough moment in terms of people's understanding of this.

Elad Gil

Well, it's also 20 years of a China-America technology-dominance narrative.

Sarah Guo

Right. I think it was also kind of this zeitgeist around us versus China, and where the West is far ahead. Will they ever catch up, et cetera? This showed that Chinese models can get there really fast.

I do think the cost thing was a huge component of it, and again, I think cost may have been misstated or misunderstood. At least, it's not clear to me that final model runs at scale are in this price range. But I completely agree with what you were saying before: experimentation tends to be a multiple. You need tooling and data work, experimentation, the pretraining run, data-generation costs, post-training, and inference. I'm sure I'm missing something here, but it seems very unlikely that there hasn't been a large multiple of $6 million spent in total.

I think there was also a narrative violation in that even at a multiple of $6 million, it's not a multibillion-dollar entry price or a Stargate-sized entry price to go compete. I think that's something that really shook the market.

Elad Gil

That should be expected, because if you look at the cost of training a GPT-4-level model today versus 2 years ago, it's a massive drop in cost. If you look at inference costs for a GPT-4-level model, somebody on my team worked it out, and in the last 18 months we saw an 180x decrease in cost per token for equivalent-level models—180x, not 180%. The cost collapse on these things is already quite clear.

That's true in terms of training equivalent models, and it's true in terms of inference. I view this as roughly on trend. Maybe it's a little bit better, and they've come up with some advanced techniques, which they absolutely have, but it does feel a little bit overstated from the perspective of how radical it is.

Sarah Guo

I do think it's striking what they did, and it kind of pushes U.S. open source forward as well, which I think will be really important. But I think people need to look at the broader picture of these curves that are already happening.

Do you think it's proof that models are commoditizing—the fact that they are so much cheaper for a given level of capability over the last 18 months?

Elad Gil

There's this really great website called Artificial Analysis that allows you to look at the various models and their relative performance across a variety of different benchmarks. The people who run this actually do the benchmarks themselves. They'll retest them rather than just take the paper at face value.

You see that, for a variety of different areas, these models have been growing closer and closer in performance. There are different aspects of reasoning and knowledge, scientific reasoning and knowledge, quantitative reasoning and math, coding, multilinguality, and cost per token relative to performance. They graph this all out for you and show you, by provider and by state-of-the-art model, how things compare.

Things are getting closer over time versus more dispersed over time. I think, in general, the trend line is already in this direction, where it seems like a lot of people have moved closer and closer to parity than they were, say, 18 months ago, when I think there were enormous disparities. Obviously, there are certain areas where different models are still quite a bit ahead, but on average things are starting to net out a little bit more.

That may change. Maybe somebody comes out with an amazing breakthrough model and leapfrogs everybody else for a while, but it does seem like the market has gotten closer than it was even just a year ago.

Sarah Guo

What do you think is the value of being the leader at the frontier?

Elad Gil

I think there are 3 or 4 different types of value. One is capturing market share. Do you just get more people using you, and then do they stick because they're used to it, or because they've optimized prompts or other things for what you're doing, or because of other tooling?

The second thing is that, if you're actually using the model to help advance the next model, having something that's dramatically better can make a difference. That could be data labeling, artificial-data generation, or other aspects of post-training. There are lots of things you could start doing when you have a really good model to help you. It could be coding and coding tools; it could be all sorts of things.

There is an argument that some people make that, at some point, as you move closer and closer to some form of liftoff, the more state-of-the-art the model is, the more it bootstraps into the next model faster, and then it just accelerates for you and you stay ahead. I don't know if that's true or not. I'm just saying that's something that some people speculate about sometimes. Are there other things that you can think of?

Sarah Guo

No. I think one thing you mentioned—maybe if I just extend it—is underpriced, or not yet understood enough: the idea that if you have a high-quality-enough base model to do synthetic-data generation for the next generation of model, that's actually a big leveler.

If you believe there will be continued availability of more and more powerful base models, that's a big leveler of the playing field in terms of having self-improving models. That's an interesting thing that people have not really talked about.

There are different ways to have value from being at the frontier. One of the things that was really interesting to me was that the DeepSeek mobile app became a top contender in the App Store for a little bit. I think there's one belief that the cheapest, most capable model in the market actually matters to consumers, and they can tell, and that will drive consumer adoption. That's what happened, and that's why you need to have the state-of-the-art model to create these new experiences.

There's a competing view, which is that this whole drama is quite interesting, and people are trying it because they want to see what the leading Chinese AI model is like—whether it's as good as OpenAI, Anthropic, and such.

Elad Gil

I definitely believe that leading capability can lead to a novel product that draws consumer attention. But I think in this case it's more the latter.

Sarah Guo

A couple of other things that happened this past week: On the OpenAI side, they released Deep Research, speaking of really interesting advancements and capabilities. Secondly, they announced Stargate, which was a massive series of investments across AI infrastructure that was announced with Trump at the White House. What are your views on those 2 things? In some sense, they overlap in terms of OpenAI really advancing different aspects of what's state-of-the-art right now.

Elad Gil

Deep Research is a really cool product. I encourage everybody to try it. The biggest deal to me is that it immediately raises the bar for a number of different types of knowledge work. Where I might have hired a median intern or analyst before—I mean, we don't do that here, but where one could hire a median analyst or intern—I'm going to immediately compare a bunch of their work to what you could do with Deep Research. Your ability to do better with Deep Research, and the comp, is hard.

I'd say it is a really valuable product. I expect other people to adopt this pattern too, but I think it's a really novel innovation. Kudos to the team.

Sarah Guo

I would say I think it is more useful, at least to me upon first blush, in domains I understand less, to do surveying and make sure I have a comprehensive view and understand who the experts are. In an area where I feel like I have a lot of depth, I take issue with its implicit authority ranking and its ability to determine what ideas out there, what on the web, is good and what isn't.

From my initial prompting and experimentation in a domain I'm familiar with, I'm like, “Oh, man, you're really going to have to audit the outputs here.” It will orient you, but you can't take many of the claims here as given.

Elad Gil

It's the AI form of the Murray Gell-Mann amnesia effect, which was coined by the guy who wrote Jurassic Park. I can never remember how his name is pronounced—Gell-Mann or Gilman. Murray Gell-Mann was a physicist who came up with quarks and a few other things. He was a Nobel Prize winner and was considered widely brilliant, and the effect was named after him by Michael Crichton.

The idea is that if you're reading a page in The New York Times about something you really understand, you're like, “This is so dumb. How could they write this? I don't believe it.” Then you just turn the page and look at something you don't know anything about, and you assume they got it all right. Why would you do that? You instantly forgot that they got everything you know about wrong. Why would they get the other thing right? Maybe they also got that wrong.

It's this really interesting kind of cognitive dissonance around what this thing actually knows or doesn't know. If it's getting expertise wrong in a domain I understand, does that mean it's also getting it wrong in domains I don't understand? Of course, we never apply that as people. We just assume it's right in the domains we don't understand, which I think is really interesting psychologically.

But it also has real implications for how people will use AI in the future, because these things will become the definitive source of a lot of people's primary information. In some senses, it's really overlapping with some of the search use cases in deep ways, and you have something where the sources traditionally have been less evident.

I know people are working on different ways to surface what the primary sources are for some of these things, but it does have really interesting implications for how you think about knowledge in the modern era as you're using AI, and especially as you're using agents that then just go and do things and report back, and you don't even know what they did.

Sarah Guo

I think it's a very interesting topic. I'm not sure how you solve that from a UX perspective, or maybe it's somewhat unsolvable, given that it also reflects what knowledge on the web is.

It really does feel like a dangerous thing from a propaganda and censorship perspective. Social networks were kind of V1 of that—or maybe certain aspects of the web were V1, and social networks are V2—and this is kind of the big version, because it's a mix of search, Twitter, Facebook, everything else that you're using, and all the media outlets, all into one single device that you interrogate.

That's kind of where these AIs are going, and so the ability to control the output of these things is extremely powerful, but also very dangerous. That's why I'm happy that we're in a multi-AI, multi-company world. There's a way to offset that, and that's where open source becomes incredibly important if you worry about civil liberties.

What do you think about Stargate?

Elad Gil

Maybe there are a couple of different implied questions in Stargate. One is: How much does it matter in the race to continue to have access to the largest infrastructure? I'm going to skip the question about whether or not it's real—there's a lot of money involved here.

I think another question is how deep the capital markets are to continue funding this stuff. Maybe a final one is the involvement of different sovereigns or quasi-sovereigns in this. I don't know if I have a strong opinion on the latter 2.

The way I think about the dynamic of how much the capital matters, and the implied question of whether we continue to see scaling in pretraining as a dominant factor, is really as uncertainty rather than risk. If you think about capabilities as emergent, and people not being sure what algorithmic efficiencies counteract the improvements that will come from more scale, the things you can do to generate new data to improve in other vectors, and what we're going to get out of test-time scaling, I just think it's very hard to predict.

But I fail to see a scenario where anybody trying to build AGI—any of the large research labs—wouldn't want the biggest cluster they could have if it were free, or if the capital were available to them. That, to me, says more than anything else: We are going to get more out of pretraining. Is it going to be as efficient? I think that's unlikely.

Sarah Guo

We're a little bit delayed on this, but we'll just give ourselves a free pass given that it's Episode 100. Predictions for 2025. Happy New Year.

Elad Gil

It's February, but I'm going to say Happy New Year. It's like the Larry David episode.

Sarah Guo

Yeah, basically. There's some statute of limitations on how late into the year you can say Happy New Year. We're now a month in, so of course we're way over that. We should probably say Happy Valentine's Day, even though we're 2 weeks early.

Elad Gil

No, I don't like that. What was the vibe for 2025?

Sarah Guo

You can just do things. You can just say Happy New Year.

Elad Gil

Happy New Year a lot.

Sarah Guo

Yeah, I'm going to do whatever the fuck I want for a whole year. It's going to be amazing.

Elad Gil

On 2025, I think there are a few things that are likely to happen. First, the foundation-model market should at least partially consolidate, and it may be in the ancillary areas. That's image generation, video, voice, and a few other areas like that. Maybe some secondary LLMs or foundation models will also consolidate.

I do think we're going to see a lot of consolidation, particularly if the FTC is a little bit friendlier than the prior regime. We'll also see the expansion of new races in physics, biology, materials, and the like. That will happen alongside the general scaling of foundation models, which will continue. That includes reasoning and other things.

That's one big area. The second area is that we're going to see vertical AI apps continue to work at scale. It's Harvey for legal, Decagon and Sierra for customer success, and a variety of folks, like Abridge, for medical scribing, et cetera. I think it'll be the era of vertical ops, and a subset of those will start adding more and more agentic things to them. Some folks, like Cognition, are already doing that.

Third, self-driving will get a lot of attention. Obviously, Tesla and Waymo are starting to see really interesting adoption of full self-driving and robotaxis. Applied Intuition is a dark horse to watch more generally on the automotive stack.

Fourth, I think some consumer things will get large-scale experiments happening in a way that hasn't happened until now. I'm starting to see consumer startups, and I'm starting to see more consumer applications from incumbents. I actually think we're going to see a bit of a resurgence in consumer. It may take a while, but I think that'll happen.

Lastly, there are things that we all know will happen and that are really early, but we may start to see some interesting behavior with agents and some early robot stuff. It'll be one of those things where it's more going to be a glimmer of how this thing will work versus the whole thing, but I think some of those developments will be very exciting.

Those would be my 5 predictions for 2025. How about you? What do you have?

Sarah Guo

We agree on a number of different things. I think the whole definition of “agent” is super fuzzy, but if we just think of it as doing multistep tasks successfully in some sort of end-user environment and taking action beyond just generating content, we're already seeing that. I think we're going to see that more broadly as reasoning models get better and product companies, or vertically integrated companies, get better at handling failure cases and managing state intelligently.

We're already seeing that in security, support, and SRE, and I think that will continue to happen. This already happened in coding, as you were alluding to, but I think companies doing copilot products will naturally extend to agents. They'll just try to do more and take more on.

One of the inputs to broader consumer-experience automation, as you describe it, is just way more capable, small, low-latency models. I don't think we have any monotonic movement toward compute at the edge. When people say “edge compute” for the sake of edge compute, I'm like, “Nobody cared.” But if you can make that transparent to the user and it's free, then your ability to ship things that are free is obviously unlocked, and I think that's cool.

I also think there will be a lot of web apps. I don't think it necessarily has to be on-device consumer products. Undoubtedly, there will be some, but I also think there will be things running on the internet that just become part of your application stack, on your browser, that will do really interesting things over time.

Elad Gil

Yeah, well, stuff in the browser can also use the GPU.

Sarah Guo

I just think the ability to run locally might be a big unlock for them. I don't know if you and I disagree on the timeline, but I think we're going to see technical proof of breakthroughs in robotics and generalization this year, though not deployments.

Elad Gil

One thing that's maybe mispriced, just because it's very new, is that people don't really know how to think about reasoning. I would claim that one thing it's as much an improvement in reliability as it is in the complexity of the task.

One mistake that entrepreneurs and investors make, and that I have made, is that you look at something and it's not working, and the issue is a technical issue, and then you assume it's not going to work. But I think in AI you have to keep looking again and again, because things can begin to work really quickly.

Maybe one last one is something I've seen small examples of with our EMBED program and also broadly in the portfolio. Because you have this diffusion of innovation—not just with customers, but with the types of entrepreneurs who go take something on—we're beyond just the tip of the spear now. More and more people are thinking, “I can do stuff with AI.”

I think we're going to get smarter data-generation strategies for different domains where you need domain knowledge as well as an understanding of AI. Examples here could be biology and materials science. You need a set of scientists who are capable of innovating on data capture, which might literally be a biotech innovation versus a computer-science innovation, to understand the potential of deep learning—and that the bottleneck was data, and then the type of data you were looking for.

I think that is happening, and I think that's really exciting. This may be the year where we see something really interesting happen on the health side, as an example, where you need specialized data, but it's not as hard as the atomic world of biomolecular design or something.

Sarah Guo

Anything else we should talk about? Your facial hair.

Elad Gil

We could. Should I bring it back?

Sarah Guo

I liked the beard. I like the beard-and-hat era.

Elad Gil

Oh, interesting. Yeah, maybe I should go back to that.

Sarah Guo

The last question for today: We're on Episode 100. What do you think the state of the world will be relative to AI when we're at Episode 200?

Elad Gil

I don't think we're part of this anymore. I think it's just 2 agents going back and forth, teaching us stuff, and you and I are no longer the hosts or the choosers of topics. We're just nodes into the network.

Sarah Guo

Will they be as good-looking as us?

Elad Gil

They'll be better computers. We'll see. I still like some art more than some Midjourney art.

Sarah Guo

There are some beautiful things on there.

Elad Gil

Okay, Episode 200. That's, like, what—

Sarah Guo

It's almost 2 years if it's weekly.

Elad Gil

I think we're either in the R-Chef Farm[?], or we're sitting on a beach in Ibiza, post-abundance. That's a prediction you heard here first.

Sarah Guo

Hopefully I'll see you at Episode 200—or in Ibiza. I think the third alternative is not this great.

No Priors Ep. 100 | With Sarah and Elad | BidClub