《DeepSeek 恐慌指南:紧急特别节目》
- DeepSeek 所宣称的经济性,挑战了“前沿 AI 必须由最富有的实验室、最好的芯片和数百亿美元基础设施支撑”这一假设。 这家成立约1年的中国公司称,V3 使用受出口限制的芯片训练,成本为550万美元——约为可比 OpenAI 模型的百分之一;但 Kevin Roose 提醒,这个数字需要“打个折扣”看待。如果数据经得起核验,行业的进入壁垒、模型定价和利润率都将下移。
- DeepSeek 发布 R1 推理模型、随后将模型带给消费者后,技术故事演变成了一场市场事件。 NVIDIA 下跌约18%,市值蒸发数千亿美元;与此同时,DeepSeek 的免费、无广告应用冲上美国 App Store 第一名。R1 被定位为 OpenAI o1 和 o3 的对应产品,其可见的思考过程让提示词不再像“往喷泉里扔一枚硬币”。
- Casey Newton 的反方判断是,AI 成本快速下降本来就是既有趋势,而不是对整个行业的突然证伪。 Ethan Mollick 分享的一张图显示,在某些情况下,GPT-4 级别推理的成本在几年内下降了1,000倍;而此前美国的投资也为 DeepSeek 提供了可继续搭建的技术基础:“之所以他们能做到便宜,原因之一是其他人先付出了昂贵代价。”
- 即便 Big Tech 的数据中心不会变成闲置资产,定价压力也已立即出现。 一家大型 AI 公司的一名员工告诉 Kevin,客户已经在询问:如果从 OpenAI API 切换到 DeepSeek,能否节省80%;与此同时,同一批训练服务器和芯片可以转用于推理,随着使用量增长,具备规模优势的既有厂商反而可能受益。
- DeepSeek 表明,开放模型可以更快缩短与封闭实验室的领先差距,但也放大了地缘政治和安全风险。 R1 的权重可以下载和修改;它会拒绝回答一些政治敏感问题,例如天安门广场相关问题;而 DeepSeek 没有披露清晰的安全计划。Casey 仍警告,不要仅仅因为中国的存在就本能地加速 AI 竞赛,尤其是当相关倡议者可能从冲突中获利时。
- Jevons 悖论是既有厂商看多叙事的出口,但几位主持人并未把它视为定论。 Satya Nadella 认为,更高的效率会让 AI 变成“我们怎么都用不够的一种商品”,即使单位成本下降,也能维持需求。Kevin 指出,这也正是 Microsoft CEO 在市场抛售期间会说的话;Casey 的结论仍是,DeepSeek 意义重大,但一些反应已经“冲过了滑雪板前端”。
1. DeepSeek 把一篇效率论文变成了消费者冲击
Kevin 追溯了 DeepSeek 的起点:High-Flyer,这家成立约1年的中国 AI 公司正是从这家对冲基金中脱胎而出。它第一次获得大规模关注是在圣诞节前后,当时发布了 V3,并称其可以与美国领先系统竞争。
真正令人震惊的不只是性能。DeepSeek 称,V3 使用的是受出口限制的“二流”芯片,而非顶级 H100,单次训练的原始成本为550万美元。Kevin 保留了必要的谨慎,提醒读者对这个数字“打个折扣”看待;但即便如此,这也意味着训练成本大约低至原来的百分之一。
随后发布的 R1,是 DeepSeek 对 OpenAI o1、o3 这类推理模型的回应;几天后,面向大众的应用上线。数百万人下载了这款免费、无广告的聊天机器人,认为它“和 ChatGPT 一样好,甚至更好”,推动这款罕见的中国消费应用登上美国 App Store 第一名。
Casey 认为,DeepSeek 的产品设计突破之一,是在回答问题时展示它如何理解查询。提示词不再像“往喷泉里扔一枚硬币”,用户可以看懂模型的思路,并据此提出更好的追问。
2. 抛售重估了 AI 护城河
NVIDIA 下跌约18%,市值蒸发数千亿美元,这让 DeepSeek 的意义超越了又一次模型发布。Kevin 认为,投资者担心的是“利润率下滑和商品化”:一个名不见经传的后来者,可能用更弱的芯片和其一小部分投入,就追平了行业领先者。
受到威胁的核心前提是“越大越好”——前沿模型需要数十亿乃至数千亿美元、庞大数据中心和最好的 GPU。如果 DeepSeek 的说法成立,领先模型的成本可能只需数百万乃至数千万美元,这会吸引更多竞争者,并限制服务商的收费空间。
Casey 的反驳值得保留:DeepSeek 并非脱离昂贵的美国创新独立出现。“之所以他们能做到便宜,原因之一是其他人先付出了昂贵代价”;早期实验室承担了摸索相关技术的成本,DeepSeek 得以在这些技术之上继续搭建。
3. 更便宜的模型未必会让算力资产闲置
Casey 表示自己“没那么慌了”,因为成本下降早已极其剧烈,也是市场普遍预期的方向。他引用 Ethan Mollick 分享的一张图指出,GPT-4 级别推理成本在几年内下降了1,000倍;DeepSeek 只是再次凸显了一条已经启动的曲线。
数据中心投入同样可以复用。为大规模训练购买的同一批芯片和服务器,也能用于推理;随着 AI 采用率上升,手握产能的既有厂商可以承接查询,并围绕这些流量建立业务。Casey 认为,DeepSeek 自己的需求就说明它仍然“恨不得拥有所有这些芯片”,所以它的成功并不能证明出口管制失效。
Kevin 保留了短期看空逻辑:他认识的一家大型 AI 公司员工说,客户已经在询问,用 DeepSeek 替代 OpenAI API 能否节省80%。即便算力需求能够延续,模型供应商仍面临压力,因为买家希望推理服务“尽可能便宜”。
Satya Nadella 给出的答案是 Jevons 悖论:效率提升可能带来足够大的消费增长,从而抵消节省的成本。他所说的 AI 使用量将“飙升”,或许能够维持 Microsoft 的经济模型;但 Kevin 指出,这正是投资者预期一家处于恐慌中的公司 CEO 会传递的信息。
4. 开放权重正在缩短每一家封闭模型公司的领先期
Casey 认为,DeepSeek 的成就有一部分是“精巧地照搬美国开创的技术”。真正的变化在于速度:过去,开放模型通常会落后 OpenAI 或 Google 的系统一截;DeepSeek 则表明,它们如今可以快得多地追上来。
Kevin 表示,DeepSeek 展示了如何对 Llama 3 或 Llama 4 这类模型进行蒸馏——在基本不牺牲性能的情况下,让模型变得更小、更便宜。一旦技术路径变得清晰,Casey 说,维持可防守的护城河就会更难。
在 AI 上投入数十亿美元、却持续免费开放模型的 Meta,可能最直接地感受到这种压力。The Information 报道称,Meta 已设立4个内部战情室应对 DeepSeek;Casey 说自己没有独立报道,但相信这些消息。他的辛辣概括是:Meta“本来应该是最擅长照搬别人东西的公司”,现在却遇到了一个更快的模仿者。
5. 开放扩散同时放大地缘政治和安全风险
DeepSeek 的模型仍然明显带有中国烙印:据报道,R1 会拒绝回答有关天安门广场的提示词。Kevin 担心,如果中国企业取得领先,其审查法律和价值观可能嵌入广泛使用的 AI,并且很难再被移除。
Casey 接受这一担忧,但不相信仅仅因为中国存在,就应当加速 AI 竞赛。当那些拥有可能从旷日持久冲突中获益的投资的人,用地缘竞争为更快的 AI 竞赛辩护时,他会“扬起眉毛”。
AI 安全专家担心,能力不断增强的系统可能很快达到通用智能或超级智能,并以人类不愿看到的方式收场。R1 的开放权重意味着任何人都可以下载、运行或修改它,这进一步加剧了这样的担忧:危险的未来系统将“非常难以控制”,也很难被限制在一个国家或一家公司之内。
Casey 看不到 DeepSeek 披露过任何安全理念,甚至找不到一名明确的安全研究员:其显然采取的路线是“构建 AGI,把它交给尽可能多的人”,然后看看会发生什么。
Well, Casey, the last time we recorded an emergency podcast, you were at Gate E8 of the San Francisco airport, and we were talking about OpenAI and how Sam Altman had just been fired. Are you at the airport today? And if not—
No.
Would you like me to mail you an Auntie Anne’s pretzel so you feel more comfortable?
Yeah, that’d be—no. All things being equal, Kevin, it’s actually much more comfortable to record here in my home studio and not have to compete with the PA system announcing flights to Houston.
Casey, we are here today to talk about a little company called DeepSeek, which probably most people had not heard of, but that is causing a major series of events in the U.S. stock market and around the U.S. tech industry this week.
That’s right. By now, our listeners have probably seen that the stock market dipped on Monday and that some companies whose fortunes are closely tied to AI dipped quite dramatically. But they also might have just noticed it in the App Store, where DeepSeek has hit number 1 this week, which is a rarity for a Chinese consumer app to do in the United States. Suddenly, everywhere you look, there are signs of DeepSeek affecting the world.
And we should talk directly about what some listeners may be thinking: Why are we interrupting our normal production schedule to do a special emergency episode about DeepSeek? It is not unusual for people in the AI world to start freaking out about some new development or breakthrough, or some new model that was released, but I believe that this is the real deal. I think this is a big moment in the history of AI development, and it is really taking a toll on stock markets in ways that I think are really interesting.
I mean, you said “dip,” but NVIDIA stock, one of the highest-performing stocks on the market over the past few years and certainly the one that is most closely correlated with people’s feelings about AI, is down about 18% today. That represents hundreds of billions of dollars wiped off the market cap of just 1 company by this announcement from DeepSeek. So I think this is a broader story than just the stock market. I think it has tons of implications for other companies developing AI and also for concerns that a lot of people working on AI safety have about how this technology could get out of hand.
Yeah, I’m excited to get into it too, but I will signal that I think there are also some reasons not to freak out, and so I’m going to be trying to bring some of those to the discussion. But today, Kevin, I think we really want to do 3 things. 1, we want to tell you what DeepSeek is. 2, we want to give you some insight into why people think this is such a big deal. And then 3, I think we want to debate a little bit back and forth just how big a deal this really is.
So, yeah, let’s get into it.
All right, so let’s start with what DeepSeek is. Kevin, we have mentioned it on the show before, but tell us a little bit about this new model and why it has taken the world by storm.
1. DeepSeek Cuts Training Costs
Well, let’s talk first about DeepSeek itself. You may remember, if you listened to our episode a couple weeks ago about this, that DeepSeek is a Chinese AI company. It is about 1 year old, and it spun out of a hedge fund called High-Flyer. It was something that, I think, outside of China, most people were not paying attention to until late last year, when they released something called V3.
That was an AI model that they said was competitive with some of the leading AI models created by American AI companies. It really caught people’s attention, not just because it came out of this little-known Chinese AI startup, but because of what DeepSeek said about how it was trained and how much it cost to train.
Yeah, so tell us about what was so interesting about how they did it and what it cost.
The first interesting thing about DeepSeek that caught people’s attention was that they had managed to make a good AI model at all from China. For several years now, the availability of the best and most powerful AI chips has been limited in China by Chinese export controls. If you are NVIDIA or another American company, you are not allowed to export your most powerful AI chips to China.
So DeepSeek came out with this paper and said, “Well, we actually didn’t use your fancy AI chips. We used a kind of second-rate AI chip that was artificially limited in order to be able to export them to China. We have a bunch of those, and using just those lesser AI chips, we were able to get a model to perform as well as you American tech companies with all your fancy H100s.”
The second thing that really caught people’s attention was the cost. DeepSeek claimed that it had spent just $5.5 million training V3, and I think there are some reasons to take that number with a grain of salt. But just in terms of the raw cost of the training run for that model, $5.5 million is peanuts relative to what American AI companies spend training their leading models—something on the order of 100 times cheaper than what something like an OpenAI model of equivalent performance would cost to train.
Right, and this comes against a backdrop of all the U.S. tech giants saying, “We’re going to spend tens of billions of dollars this year to increase our capacity and data centers, and the amount of compute power that we’ll have.” So this really stood in stark contrast to that.
That tells us a little bit about what V3 is. The training and the cost were perhaps more interesting than the model itself, which is just a chatbot like a lot of us have already used. But V3 came out around Christmas, Kevin, so why is the market reacting so strongly now?
2. The App Sparks Panic
A couple of things happened in the past week or so that have led to the freak-out we’re seeing now. The first is that last week, DeepSeek released another model, R1, which was its attempt at a so-called reasoning model.
Basically, V3, the last model, was kind of similar to things like Claude or Gemini. It was a basic language model. But R1 was more like OpenAI’s o1 and o3, which are its newest reasoning models.
That happened early last week, and then a couple days later, DeepSeek did something else: It released an app where anyone could download DeepSeek and use its model in a very easy way. This is when people really started to go from being interested and fascinated by DeepSeek to panicking about it, because all of a sudden, millions of Americans were downloading this app, using DeepSeek’s models, and realizing, “Oh, wait, this is as good as or better than ChatGPT. It’s free. It doesn’t have any ads. It seems to be quite smart.”
And it does something that OpenAI’s models don’t do, which is show you the internal thought process that it is going through as it is producing these answers.
Yes.
That is something that OpenAI’s models do not show the user, but DeepSeek’s models do, and I think people found that really compelling.
Yes, and that last point that you mentioned is really important, because I suspect all the AI companies are going to copy this now. The process of using a chatbot today is that you ask a question. I’ve likened it before to throwing a penny in a fountain, right? You’re just making a wish, and you see what you get back.
What’s really interesting about DeepSeek is that as it’s answering your question, you’re seeing how the computer understood your query. So if you want to ask a follow-up question, you now have a much better sense of how the computer understood you, and this actually does seem to be a conceptual breakthrough in product design, just as much as the underlying science.
All right, so that gives us a sense of what DeepSeek is and what its latest models are. Let’s talk about why everyone is freaking out, and maybe more specifically, take a look at who is freaking out. As we mentioned at the top, 1 of the big ways people are noticing this is through the decline in the stock market. Kevin, give us a sense of the industry reaction to what the DeepSeek models might mean.
3. Investors Fear Commoditization
I would say the people who are freaking out the most are investors in the biggest American AI companies, as evidenced by all of the tech stocks selling off today. I think you could categorize that as a fear of declining margins and commoditization.
That sounds very boring, but basically, what they’re saying is, “Look, if a Chinese AI company that no one had ever heard of until a few weeks ago can come along and, for a fraction of our costs, develop a model that is as good as or better than the leading models on the market, with substandard chips, by the way, then the barrier to entry in this market is just not nearly as high as we thought it was.”
One of the fundamental assumptions over the past few years when it came to AI was that bigger was better, right? In order to build the most powerful models, you needed billions of dollars, maybe tens or hundreds of billions of dollars, and huge data centers, and all of the leading chips.
And that assumption may no longer be true. If what DeepSeek claims is true and checks out, then it may mean that it only costs single-digit millions or double-digit millions of dollars to build a leading model, which would radically shift what these companies are able to charge for their models, as well as the number of competitors in the market.
Yes, I definitely agree it changes what companies might be able to charge, but I would also note that nothing that DeepSeek did is possible without American innovation. One of the reasons that it was cheap for them is because it was expensive for everyone else, and other companies did spend hundreds of billions of dollars figuring this out. So, worth saying.
All right, let's talk about a second group of people who have been really rattled by this series of announcements, Kevin, and that is folks who are paying attention to geopolitics.
4. China Raises New AI Concerns
Yeah, so a lot of people who worry about China in general are worried about this DeepSeek announcement because DeepSeek is obviously a Chinese company. If you're a person who's worried about Chinese tech dominance or the possibility that Chinese firms could eventually get to something like AGI first, I think you are especially worried about what DeepSeek is doing.
And I think we should also say that the models themselves are recognizably Chinese. Over the weekend, I saw people testing out various queries on DeepSeek-R1, including things like, “Tell me about what happened at Tiananmen Square.” The model just refuses to answer them. And so I think there is a worry that if Chinese companies do take the lead in AI, then Chinese values and censorship laws may become embedded into this technology in a way that is very hard to extract.
Yeah. I think that's true. I also always urge caution when people try to use the existence of China as a reason to dramatically accelerate the AI race. A lot of those people have made investments that will pay off handsomely if we find ourselves in some sort of protracted and awful conflict with China. So whenever anyone starts talking about China in the context of AI, my eyebrows arch up a little bit.
All right, now, Kevin, there is one more group of folks that I think is quite justly nervous about what they're seeing out there with DeepSeek. Who is that?
5. Open Source Raises Safety Risks
So the third group of people that I would say are freaking out about DeepSeek are AI safety experts: people who worry about the growing capabilities of AI systems and the potential that they could very soon achieve something like general intelligence or possibly superintelligence, and that that could end badly for all of humanity.
And the reason that they're spooked about DeepSeek is that this technology is open source. DeepSeek released R1 to the public. It's an open-weights model, meaning that anyone can download it and run their own versions of it or tweak it to suit their own purposes.
And that goes to one of the main fears that AI safety experts have been sounding the alarms on for years, which is that once this technology is invented, it is very hard to control. It is not as easy as stopping something like nuclear weapons from proliferating, and if future versions of this are quite dangerous, it suggests that it's going to be very hard to keep that contained to one country or one set of companies.
Yeah, I mean, say what you will about the American AI labs, but they do have safety researchers. They do at least have an ethos around how they're going to try to make these models safe. It's not clear to me that DeepSeek has a safety researcher. Certainly, they have not said anything about their approach to safety, right?
As far as we can tell, their approach is, “Yeah, let's just build AGI, give it to as many people as possible, maybe for free, and see what happens.” And that is not a very safety-forward way of thinking.
So, Casey, that is a lot of information that we just dumped on our listeners, but really what I want to know is, are you freaked out about this?
Mm-hmm.
Do you think that this is as big a deal as some people seem to think it is?
6. The Case Against Panic
Mm. I think as I'm doing my reading and having conversations with folks this morning, my sense is I am freaking out a bit less than some other folks that I'm talking to. I think this is a big deal and merits discussion, but I also think that people may be getting a bit over their skis when it comes to thinking through the implications here.
So make that case, because all I'm seeing all over my timelines is people saying, “This is the Sputnik moment for AI. This is the biggest moment in AI since the release of ChatGPT. Everyone needs to stop what they're doing and pay attention.” So what is the case that you are seeing out there that people are hyperventilating a bit over nothing?
Sure. So let's take a few different points. One reason why people are really nervous here is that DeepSeek was able to train this model very cheaply, and I want to be clear: this is a significant technical achievement.
At the same time, the cost of training and inference has been falling rapidly in AI for a long time now. Ethan Mollick, who we've had on the show before, posted a chart on X that showed this decline, and in some cases, for example, running inference on a GPT-4-level model, the cost of that has fallen a thousandfold over the past couple of years.
So things have already been moving in this direction, and I think most people who work in AI expected that it would continue to go there. So that is the first point that I would make.
Got it. And if you are Satya Nadella at Microsoft, Sam Altman at OpenAI, or Sundar Pichai at Google, are you worried that you are going to spend tens or hundreds of billions of dollars building out new data centers and filling them with all the fanciest GPUs, and that some Chinese startup is going to just take everything that you do and copy it 3 months later for pennies on the dollar?
So this is a great question, which leads me to a second reason why I think at least some folks may be overreacting here, and that is that in most cases, the money that is being spent to build out the data centers that will handle these giant training runs can be repurposed.
The same servers and chips that you would use to do that can also be used to serve what is called inference, so basically actually answering the questions. As more and more people start to use AI, it will be those giants that actually have the capacity to serve those queries. They will be able to build businesses around that.
And by the way, that is another reason why I don't think that DeepSeek is evidence that the export controls failed. The folks over at DeepSeek would love to have all of these chips, not just to do the big training runs, but also so that they could serve all of the demand that they are currently generating, right? So I think that's another important thing to keep in mind as this discussion moves forward.
Yeah, that makes a lot of sense to me. I do think the cost dynamics here are very important, because I talked to a person I know who works at one of these big companies, and he said that a lot of their customers are already starting to ask, “Well, could we shift over from using the OpenAI APIs and their models to using DeepSeek if it saves us 80% of our costs?”
And so I think in the short term, there is reason for the American AI companies to worry because people want these things to be as cheap as possible.
Yeah, and let me just say, as somebody who spent $200 to upgrade to GPT Pro last week so I could try Operator, I'm looking forward to that price going down.
7. Open Source Closes The Gap
But that leads to, I think, maybe the third reason that I think people might be overreacting a little bit here, which is that a lot of what we are seeing is essentially a fancy ripping-off of techniques that were pioneered here in the United States, right?
It has long been the case that open-source models were just a little bit behind the models that were made by the big labs. You look at Meta's Llama models, which until DeepSeek were seen as the best open-weights models that were out there. They weren't as good as what OpenAI or Google or others were doing, right?
Where I do think that this gets super interesting is that DeepSeek is showing us open source can now catch up faster than it used to, right? The labs used to have a little bit longer lead, but now people are just getting cleverer and cleverer about these techniques. And so it is getting harder to build that defensible moat because this is just one of those technologies where once you figure out basically how people are doing it, you could just get in there and do it too.
Yeah. Now, Casey, I'm curious what, if anything, you are hearing from inside Meta specifically, because I think this is one of the most fascinating angles.
Meta is a company that has spent billions of dollars developing AI models and, unlike most of its competitors, has chosen to release those models freely. And part of what DeepSeek has shown is that you can take a model like Llama 3 or Llama 4, and you can distill it. You can make it smaller and cheaper. You can do that without sacrificing a lot of the performance.
And so there were some reports in recent days that Meta is basically at DEFCON 1 right now, that they are—The Information reported that they have 4 war rooms at Meta headquarters devoted to trying to figure out how to respond to the DeepSeek threat.
Yeah.
And—
And by the way, I hope they were the same war rooms that Meta used to use to protect America from election interference.
They say, “Hey, get out of here. We got something else we gotta figure out.”
So are you hearing anything from Meta? Because I think that is the company that I would say has the most to worry about when it comes to DeepSeek, because DeepSeek is doing essentially what they do but at a fraction of the cost.
Yeah, so I do not have my own original reporting to share on this yet, but I do trust the information that they are freaking out. And the reason is that Meta is supposed to be the best company at ripping other people off. And so when they find out that some Chinese Johnny-come-latelys are gonna be better than they are at ripping things off, they're gonna have something to say about it.
Nothing could be more poetic now that DeepSeek has ripped off all the American companies. Meta is coming back, and they say, “Oh, you think you're good at ripping people off? Just wait until we have plumbed the guts of DeepSeek-V3 and DeepSeek-R1. We're gonna be back on top sooner or later, bucko.”
Yes. Now, I wanna ask you about one other reaction that I saw on social media, which was from Satya Nadella, the CEO of Microsoft. He went on his X account late last night and posted the following: “Jevons paradox strikes again. As AI gets more efficient and accessible, we will see its use skyrocket, turning into a commodity we just can't get enough of.” And then he linked to a Wikipedia article about Jevons paradox.
So, Casey, did you understand this, and if so, what did you make of it?
Well, I did, because we had just discussed Jevons paradox on this very show, Kevin.
It's true.
When Hugging Face's Sasha Luccioni came on and explained Jevons paradox, which is essentially as stuff becomes more efficient, you simply increase demand for it, thereby canceling out a lot of the efficiency gains. So when I saw Satya tweet “Jevons paradox,” I said, “Once again, Hard Fork has set the national news agenda—and if you're not listening, fix that.”
Yeah, many people are talking about Jevons paradox. I predict that this is going to be something I'm gonna hear about at every single party I go to for the next 6 months.
And, just to connect the dots a little bit, I think what Satya is trying to say here is that DeepSeek is not actually a threat to companies like Microsoft because as the cost of building and using AI models comes way down, people are just gonna wanna use them more and more, and so the overall demand and Microsoft's overall profitability will not change. Which could be true, but I would also just say it is exactly what you would expect the CEO of Microsoft to say on a day where investors were panicking and selling their stock.
It is. It is. All right. Well, Kevin, I think that's a pretty good overview of what DeepSeek is doing, why people are freaking out, and at least some thoughts about exactly how freaked out you should be.
There is a lot more to say about this subject, and if you are starving for even more discussion of DeepSeek, I can promise you that we'll have more to say on our regularly scheduled episode of Hard Fork this Friday.
Yes, Casey, I love doing these emergency podcasts. They fill me with a sense of danger—
Mm-hmm.
—and excitement, living on the edge.
I love 'em for that reason. I love them for a second reason, Kevin, which is that I get paid by the episode. So here's to many more emergencies in 2025.