投资人 Meng Xing 谈开源影响、中国 AI 人物与 RL 数据的未来
- Kimi K3 让 Xing Meng 相信,中国开源实验室已经贴近前沿,不再只是快速追赶者。 Kimi 仍是一个不到400人的理想主义实验室,专注于 AGI。它在算力受限的情况下,将包括 KDA 在内的底层架构创新整合进一款2.8万亿参数模型,其中一些创新此前只在300亿至400亿参数的“玩具模型”上验证过。Xing Meng 回忆,约一年前发布的 K2 若放在闭源模型榜单中,大概排名第6或第7;K3 已接近第2或第3。他预计,Kimi 与 DeepSeek 将在开源模型排名中轮流占据第1。
- 据称,中国开源模型真正可交易的影响在 Token 收入,而不在于吸引前沿研究人员的注意力。 前沿实验室的研究人员并不太在意这些模型,尝试过的人也很少;但大量 Agent 开发者已经转向 K3。Meng 听到的说法是,这显著影响了前沿实验室的营收,尤其是 Anthropic:7月 Token 消耗增速低于预期,有人甚至称已经停滞;如果总消耗持平,而用户用更便宜的模型替代昂贵模型,收入就会下降。用户可能会组合使用不同模型,例如用 K3 或 DeepSeek 写代码,再用 Claude 审查。即使在 Anthropic 的 Fable 5、Opus 5、Opus 4.8,有时还有 Opus 4.6 之间,用户也未必总能判断哪一个更适合某项任务。
- ByteDance 据报道决定不蒸馏前沿模型,这一选择既勇敢又意味深长,不能简单理解为承认落后。 Meng 特别在意那个据称附带的条件——接受 Seed 可能落后于中国同行:如果一家公司拒绝蒸馏,可能已经落后,也可能面临落后的风险。他区分了直接蒸馏与通过合成数据进行的间接蒸馏,认为所有人或多或少都在蒸馏某些东西。他表示,Seed 和 DeepSeek 可以承受落后几年,因为不需要融资,从而保留了用“正确的方式”冲到顶尖的野心。
- ByteDance 的 Seedance 源于一次高风险的规模下注,当时几乎没有足够先例证明这条路行得通。 Meng 称,相比此前的多媒体模型,Seedance 的规模扩大了一个数量级,需要的 GPU 从数千张增至数万张;与此同时,LLM 和编码任务对算力的需求也在上升。他将成果归功于 ByteDance 一名非常年轻的女性研究员。收入似乎主要来自专业消费者:Topview、LiblibAI 等视频画布服务据称在 Seedance 发布后,从年初的数百万美元或中个位数百万美元收入增长到接近1亿美元甚至更高,部分买家还需要预付 API Token。主持人称,今年中国短片的影响力已经超过电影院;Meng 在确认其中既包括真人出演也包括 AI 短片后,接受了这一比较。
- “新实验室”的定义正在变化,Meng 认为 Cursor 的路径变得越来越重要。 一个拥有独特、几乎“不可触碰”用户轨迹数据的垂直应用,可以通过后训练打造 Composer 这样的模型,并有机会超过许多预训练模型。优势不在于新方法,而在于获得前沿实验室没有的数据。法律、医疗等垂直工具,可能比通用应用更有机会训练出有辨识度的模型。更深层的问题在于,研究人员对数据的选择往往反映他们对编程和数学的熟悉程度;真正的测试环境应当是人们每天使用的工作软件——Slack、Zoom、Salesforce 以及类似系统。
- 训练数据的未来属于环境,难度则从稳定的优化任务一路延伸到部分可观测和动态变化的世界。 Meng 描述了静态 Kernel 优化环境、类似 Slack 的部分观测环境,以及金融或社交媒体环境——在后两者中,每一次行动都会改变下一状态。Vending-Bench 将这一思路推向物理世界,让模型运营虚拟自动售货机;其创建者还在附近搭建了真实商店或市场,以更不可预测的条件测试 LLM。
- 数据供给正在沿两种模式快速增长。 Scale、Mercor 等按人头计价的供应商,把招募来的劳动力像咨询服务一样出售;新一批供应商则销售环境,价格从几百美元起,最高约10万美元甚至更高。一个 Booking.com 任务可能只需约1000美元或更少,而重建大部分 AWS 式 DevOps 环境,成本可能接近50万美元。Meng 认为,前一种模式目前收入更高,后一种由研究人员主导、提前布局的模式增速更快。在韩国最近一场会议上,年轻研究人员告诉他,自己最想做的是创办数据供应商。
- AI for science 的真正瓶颈是验证,而不只是建模。 AI 科研 Agent 已能阅读文献、设计实验、在部分情况下控制实验设备、分析结果并撰写论文,但科学数据库、工具、工作流和湿实验执行仍然困难重重。蛋白质设计模型面临严重的数据与多模态问题,虚拟细胞模型则试图在无需昂贵实验的情况下测试生物相互作用。真正的验证标准,是药物能否在人体内发挥作用并治愈疾病;完整反馈可能要等到研发流程约第5或第6年、进入 FDA 批准的临床试验阶段后才会出现。
- Meng 认为,硅谷最大的误解是把中国 AI 的进展完全归因于政府支持。 他表示,美国生态没有充分认可那些在艰难约束下推动进展的个人、创业者和研究人员,并相信只要他们能在中国做到,就能在任何地方做到。
1. Kimi 是不到400人的理想主义实验室,刚刚交付接近前沿的模型
- Meng 对 Kimi 的画像是一个“汇聚了真正聪明的年轻人才、去实现非常理想主义目标的枢纽”——目标是 AGI。这个团队经历了3年媒体关注度最高的时期、“糟糕的日子”,以及重返巅峰的过程,始终没有转向应用或多媒体。几乎还在上高中的实习生也能获得真正的算力,并参与构建 K3 的关键组件。
- 团队规模是刻意控制的:Xing 表示,Kimi 目前仍不到400人;主持人估计,3年下来团队大约只有300人左右。这并不是因为资源不足,而是为了保持团队完整,避免上下文和知识在扩张中流失。
- K3 是否让他感到意外?答案是肯定的。Xing 回忆,约1年前发布的 K2 若放在闭源模型榜单中,大概排第6或第7;如今仍处于开源状态的 K3 已接近第2或第3。将底层架构创新——其中一些只在300亿至400亿参数的“玩具模型”上验证过——整合进一场2.8万亿参数的训练,通常会拖累短期表现;Kimi 却在算力有限、周期很短的情况下完成了这件事。“如果我说自己完全不意外,那就是在撒谎。”
2. DeepSeek 创始人、稀缺催生的创新,以及中国实验室为何各有性格
- 谈到从未见过面的 DeepSeek 创始人,Meng 说:“他似乎很早就看透了这个世界”——知道如何变得富有、赚钱、交易和招募人才;但他对财富或影响力似乎并不感兴趣,也没有利用这些能力去积累财富或扩大影响力。这个实验室从中国高校招募未经雕琢的年轻人才,而不是从美国顶级实验室挖来高阶、知名研究员。几位 VC 朋友晚上9–10点去办公室拍照,发现按仍在工作的员工占比计算,DeepSeek 和 Kimi 是最拼的两家公司。“很多努力其实都是自发的。”
- 他反复遇到一种地位倒置:OpenAI、Anthropic 和 Gemini 的基础设施人员,竟然“仰望 DeepSeek 的年轻人”,其中包括1998年及以后出生的工程师。背后的机制就是约束:“只要把资源压缩到足够稀缺,足够多的优秀人才聚在一起,魔法自然会发生。” DeepSeek 还会发布异常详尽的技术论文及配套工作;Google 等机构的研究人员会说:“我以前不知道这件事还能这样做。”
- Kimi Delta Attention 与 DeepSeek Sparse Attention 在技术上不同,但共享一些原则。不同模型版本之间的能力似乎在不断跃迁。与美国资深研究人员频繁跳槽不同,中国研究人员较少在国内实验室之间流动,因此每家实验室都会形成独特性格,由创始人的信念和被其吸引来的人共同塑造。
3. Z.ai 与 MiniMax:两家上市实验室,两种不同性格
- 主持人介绍 Z.ai 时,提到 GLM-5.2 已获得强劲的编码采用率。Xing 认为,Z.ai 的工作质量不错,团队规模更大,核心围绕 Tang Jie 教授及其学生构建,并已拓展到电脑操控、多模态及其他产品套件。他问道,Z.ai 的市值是否已经达到1.23万亿元人民币、接近2000亿美元——即便2025年收入仍然非常小。相比 Moonshot 和 DeepSeek,Z.ai 的聚焦程度较低,更像一家按 OpenAI 路线发展的完整 AI 公司。
- MiniMax 的股价经历了“过山车”:它曾与 Z.ai 一起上涨,随后大幅回落,目前接近或略高于 IPO 发行价,约为峰值的1/5至1/6。公司最初打造了类似 Character.AI 的知名陪伴产品,也帮助定义了中国公司如何构建这类产品;后来转向多模态模型,包括 MiniMax M1 和 Hailuo 系列。但 Xing 认为,MiniMax 在纯 LLM 的推理和编码方面相对不够成功。
- 2024年初的一个人才信号是:许多原本可能加入 ByteDance 或 TikTok 的“酷人”都去了 MiniMax,其中包括产品经理、开发者、研究员和市场团队人员。“整体氛围有点像 TikTok。”
4. ByteDance 拒绝蒸馏:勇敢、意味深长,但只有部分公司负担得起
- 据报道,Zhang Yiming 曾告诉员工,Seed 将拒绝蒸馏前沿模型,“即使这意味着我们会落后于中国同行”。Meng 抓住了这句话中的条件:“如果你不做蒸馏,就意味着你会落后,或者已经落后了。这对我来说才是更大的新闻。”
- 他将蒸馏分成两类:直接蒸馏是把一套 Prompt 分布发送给前沿模型,再用返回结果训练,包括在可获得时使用 Claude 的思维链;间接蒸馏则通过合成数据完成。由于大量合成数据由模型生成,“默认情况下,合成数据蒸馏了用于生成它的模型”。因此,所有人或多或少都在做某种蒸馏。Meng 表示,合成数据和直接蒸馏都已证明有效,但他没有证据判断谁在做什么。
- 为什么拒绝蒸馏?因为这条捷径“会给我们一种虚假的满足感,因为我们是用错误的方式实现了目标”。当主持人问是否有人能承受落后几年时,Meng 回答:“如果你不需要融资,可以。”他认为 Seed 和 DeepSeek 就拥有这种余裕。
5. 据称,开源真正冲击的是 Token 经济,而非前沿研究人员的注意力
- 这次出行让 Meng 感到意外的是:尽管围绕 K3 的讨论很多,开源也受到广泛关注,但前沿实验室的研究人员“并不太在意这件事”,真正试用过这些模型的人很少。相比之下,Agent 开发者已经有相当一部分转向 K3,尤其是新闻报道呈现出的情况。
- Xing 听到的说法是,这已经显著影响了前沿实验室的营收,尤其是 Anthropic。7月 Token 总消耗增速低于预期,有人称已经停滞。如果总消耗持平,用户从昂贵模型切换到便宜模型,就会压低收入。一个典型的混合用法是用 K3 或 DeepSeek 写代码,同时保留 Claude 做审查。
- 对普通任务而言,模型间的差异正在收窄:“大多数时候,我不认为你能那么容易分辨顶级模型之间的区别。”即使在 Anthropic 的 Fable 5、Opus 5、Opus 4.8,有时还有 Opus 4.6 之间,用户也无法总是判断哪一个更适合某项任务。K3 并不算特别便宜,但比可比模型便宜。
- 为什么非编码 Token 的增长会停滞?企业推动 AI 落地可能需要2–3年,因为合规、安全和隐私工作都需要时间。白领任务也缺少容易预先定义的目标——“在看到 PowerPoint 之前,很难把目标设定好”——因此人类仍在流程中,吞吐量受到限制。Meng 提出的杠杆是提速:如果把制作一份 PowerPoint 的时间从10分钟缩短到10秒,就能进行更多轮迭代。“这还没有发生。”
6. Seedance 的高风险规模下注,构建了专业消费者驱动的短片经济
- 回到2023年,Meng 认为,如果要猜谁最可能做出未来最强的多媒体模型,答案显然会是 ByteDance、Kuaishou 或 Meta,因为它们拥有数据、分发、平台和 GPU。Kuaishou 的 Kling 曾登上第1,随后被 ByteDance 的 Seedance 超越。ByteDance 将 Seedance 的规模在此前多媒体模型基础上扩大了一个数量级,需要的 GPU 从数千张增至数万张;当时几乎没有证据表明这套方法会奏效。他将成果归功于 ByteDance 一名非常年轻的女性研究员。尽管 Meta 拥有相关数据和能力,但至今仍未展现出他预期的进展。
- 收入主要由专业消费者驱动,而非个人消费者。Topview、LiblibAI 等视频画布服务据称在年初的收入只有数百万美元或中个位数百万美元,Seedance 发布后却增长到接近1亿美元甚至更高。一度有买家必须预付款,才能锁定 Seedance API Token。
- 主持人称,今年中国短片已经超过电影院。Xing 先确认这其中既包括真人出演的短片,也包括 AI 短片,随后表示这一比较是合理的。他说,AI 生成短片此前最擅长科幻、传统中国科幻和古装武侠,因为这些题材对特效的要求高于对细腻人类情感的要求。Seedance 的演示展示了更流畅的情绪转折,例如从大笑转为哭泣,以及更准确的高动态动作,例如武打场景中手臂不再穿模。因此,它几乎已经成为许多短片的独家供应商。
- 对介于电视和游戏之间的互动媒体,Xing 修正了此前的乐观判断:“我得说,我有点失望。”主持人以 Detroit: Become Human 为例,介绍这种一半电影、一半游戏的形式;Xing 表示,这类内容已经出现过个别爆款,但质量和普及度还没有高到足以全面起飞。他仍希望这一幕最终发生,并说:“我认为它会比游戏本身更大。”
7. 新实验室正在被重新定义:Cursor 路径与研究员品味问题
- 过去对新实验室的定义,核心是年轻研究人员或拥有学术背景的资深研究人员。如今,Xing 更多想到 Cursor:一家专注应用的公司,拥有超级应用和用户轨迹数据,通过后训练而非预训练训练出 Composer,使其能够超过许多预训练模型。“这不是一种新方法,而是一套前沿实验室没有的新数据。”
- 判断标准是:数据必须独特且“不可触碰”,而垂直领域比通用市场更适合。法律或医疗应用有机会针对独特任务调优模型,从而超过通用模型;相比之下,通用 Agent 的 Prompt 与训练 Claude、OpenAI 模型时已经使用过的数据相似。Xing 提到了 Harvey 及美国其他垂直应用 Agent,并预计中国也会出现类似发展。
- 更深层的观点是,数据选择“取决于核心研究人员自己的品味”。刚毕业的 PhD 可能熟悉编码和数学,却“可能从未做过多少 PowerPoint 或会计工作”。衡量模型的真正标准,应当是人们实际使用的工作环境——Slack、Zoom、Salesforce 以及其他 SaaS 系统,而不是由研究人员偏好想象出来的环境。Cursor 之所以较早出现,部分原因就在于编码是研究人员熟悉的领域。
8. 训练数据的未来是环境:看世界会在多大程度上反过来施加约束
- 数据已经从静态材料——无论是爬取数据还是合成数据——转向环境:模型在其中尝试完成任务,并从奖励中学习。Xing 的分类从静态 Kernel/CUDA 优化开始,这类任务“可以重复100万次而不会改变”;再到部分可观测环境,例如新员工在 Slack 里“试图搞清楚老板到底想要什么”;最后是交易或社交媒体等动态环境,在那里每次行动都会改变下一状态。一条社交媒体帖子会改变粉丝的认知,因此同一个实验不能简单重复1000次。
- Vending-Bench 通过让 LLM 运营3台虚拟自动售货机,测试从虚拟世界向物理世界过渡的能力:模型需要选择供应商、设定价格并决定营销方式。创建者告诉 Xing,他们还在附近搭建了一个市场或商店,用来测试模型能否在不可预测的条件下经营现实世界的生意。“这些数据不在互联网上……必须在现场生成。”
9. 数据供应商正在两种完全不同的商业模式上快速增长
- 过去6个月,数据供应商的收入大幅增长,许多供应商的 ARR 已达到20亿至30亿美元。中国供应商尚未达到这一规模,但增长很快。第一种模式是按人头计价,Scale 和 Mercor 都在采用:它们像咨询公司一样出售招募来的劳动力,属于高度依赖关系的 B2B 生意,还要投入大量资源宴请决定数据采购的研究人员。
- 第二种模式销售环境,价格从几百美元起。一个 Booking.com 式的预订任务可能只需约1000美元或更少;如果要为一个高难度 DevOps Agent 几乎完整复刻 AWS 体验,成本可能接近50万美元。数据公司的研究员创始人会提前3个月预判实验室需求:创建一个 benchmark,将其做出影响力,再销售能帮助模型在该 benchmark 上取得高排名的数据。Xing 认为,按人头计价的模式目前收入更高,而提前布局环境的模式收入增速更快。
- 在上个月韩国的一场学术会议上,年轻研究人员告诉他,自己最想做的是创办销售数据的创业公司:“他们不做模型公司,也不做 Agent 公司——他们要做数据公司。”主持人描述了从人工标注转向环境和 benchmark 构建的更广泛趋势;Xing 表示,中国供应商同样是从人工标注业务起步。他还提到 HLE,即 Humanity's Last Exam,作为高端专家标注的例子。
10. AI for science:易于构建,难于验证;以及硅谷对中国的盲区
- AI 科研 Agent 可以阅读文献、设计实验、在部分情况下控制实验设备、分析结果,以及撰写或发表论文。Xing 说,这些系统“如今已经不算难做”,但要让它真正优秀且有意义仍然困难,部分衡量标准是它能否产出一篇受到认可的论文。障碍包括作为底层模型的 OpenAI 和 Anthropic 模型并未针对科学数据库和工具进行专门训练;此外,湿实验的执行、数据采集和验证都需要时间与人工介入。
- 在模型侧,自 AlphaFold、尤其是 AlphaFold 3 以来,DeepMind 以及 Chai、Boltz 等公司一直在推进最先进的蛋白质设计和复合结构预测。Xing 还提到他们投资过的蛋白质设计模型公司。这些模型可以尝试从头设计,创造自然界不存在的蛋白质或抗体;虚拟细胞世界模型则可能在不进行昂贵湿实验的情况下,测试设计出的结合体是否会按预期发生相互作用。最大的问题是细胞数据有限,以及多模态能力不足:酶、抗体、抗原和肽必须被放在同一系统中表示,而不能只交给彼此割裂的小模型处理。
- 结构性的验证缺口依然严重:虽然存在中间验证器,但“真正的验证器,是这种药物在人体内是否有效,以及能否治愈疾病”。完整反馈可能要等到研发流程约第5或第6年、进入 FDA 批准的临床试验阶段后才会出现。“你无法想象 LLM 会有这么长的反馈周期……但这就是药物设计模型今天面对的现实。”
- Xing 认为,硅谷的一个误解是把中国 AI 的进展完全绑定在政府通过资金、算力等方式提供的支持上。他认为,美国生态没有充分认可那些在艰难约束下取得进展的个人、创业者和研究人员;只要他们能在中国做到,就有能力在任何地方做到。
1. Kimi, Moonshot AI and China’s young AI researchers
Kimi is a very idealistic company and a hub of really smart young talent gathered together to achieve a very idealistic goal: AGI. DeepSeek is also made up of hardworking, driven people who have been able to develop great open-source models. I think they'll be able to alternate the number-one spot among open-weight models over time.
Let's talk about this debate over distillation in general and ByteDance's decision not to go down that route.
I think it's a very brave decision by the head of ByteDance to say that this is not the route we're taking. If we go down that route, it'll give us a false sense of satisfaction because we achieved this in the wrong way.
How has open source changed AI?
If we're talking about frontier labs, researchers, and so forth, they haven't cared much about this.
Xing, welcome to Valley 101.
Thank you for having me.
2. Silicon Valley trip recap
You split your time between Beijing and Silicon Valley, and you come here once every quarter. What was your biggest takeaway from this trip to the Bay Area?
Before we came here, about a week ago, we had a few themes in mind, and we wanted to explore how they played out in the Valley. We wondered about models like Kimi and DeepSeek. We saw a lot of buzz and discussion on Twitter and even in Congress, but what is their real impact on frontier labs, agent companies, and researchers in general? Are they really that concerned? Are they really paying that much attention to them?
How does that change the economics of the world, especially leading up to the IPOs of OpenAI and Anthropic? Does that have an impact? We thought it would, but we wanted to see some proof.
The second theme is that, in the past 3 years since the launch of GPT-3.5, I think the key theme has always been pushing out the best next model and reaching the ultimate performance beyond imagination—being at the top of every benchmark possible. Now we're reaching a point where the model companies want to reach more users, and those users have to go beyond coders and other specialized users.
They want adoption rates to go up and to reach more diverse users. So you're looking at people who care a lot about efficiency and serving speed rather than only the ultimate level of intelligence. Token efficiency is important, and so is the cost of routing the right model to serve a specific task.
The third theme is AI for science. That has been a big factor and one of the big focuses for me and my team for almost a year now. We're looking at breakthroughs across biology, materials science, and even traditional drilling for minerals and applying AI in those directions.
Every time we come here, we explore those things because Silicon Valley is not only a hub for AI but also a hub for general science. Those are the things I'm most looking forward to and will spend a lot of time on.
Open source, this transition from capability to speed and inference, and then AI for science. We'll dig into each of those.
Let's start with the foundation models. One of the things I find really interesting in the US is that we're getting into this two-tier ecosystem. We have Anthropic and OpenAI, which are the duopoly in the startup world, and then we have the public companies and bigger companies, like Google and Meta.
Things look similar in China as well. We have the so-called four dragons of AI startups that focus on foundation models, and then we have the bigger companies like ByteDance and Alibaba. Would you agree with that characterization? Is that a good way of looking at the foundation-model labs in China?
Sure. I think those are the main ones.
I wanted to talk a little bit about the specific companies, and I'd love to know how you think of them from an investor perspective. Let's start with Moonshot, since we've been talking about Kimi K3 for a while. What kind of company is that?
I think Kimi is a very idealistic company. They're a hub of really smart young talent gathered together to achieve a very idealistic goal, which is AGI. At the very beginning, 3 years ago, that sounded unusual. Today it sounds normal—everybody is saying that—but 3 years ago, people would have thought you were too ambitious for your credentials in that way. They held on to that ambition and aspiration throughout the past 3 years.
They've been at the peak of media attention, they've gone through horrible times, and now they're back at the peak again. But I don't think they've changed much in terms of what they think or what they do. They're very focused. They're not an application company, and they don't build a lot of multimodal models. They're laser-focused on building AGI, in a sense.
They're one of the companies that probably has the best atmosphere for young researchers to join. They give a lot of attention and resources to interns and really young interns. They have interns who are almost in high school, which is not very usual, and those interns are able to access a lot of compute power and resources so that they can build important components of the latest model, Kimi K3, which was just launched.
I think they're a team-first or culture-first company. They establish a very attractive goal, attract talent, and remain laser-focused on that goal. If you talk to people in that organization, I think they'll tell you the same thing. They're a very small team today. I think they're still fewer than 400 people.
More or less. If we talk to the companies here in the US, and if I were to ask them to guess, after seeing Kimi K3, the buzz, and the performance, how many people they would guess this company has, they would be really surprised to learn that it only has 300 or so people after 3 years of development.
It's not that they don't have the resources to hire more people. They decided to keep the team small so that they could remain intact enough for everybody to know what everyone else was working on. They're close enough that there's no context loss when they're building things together.
Knowing that team, did the release of Kimi K3 surprise you? Do you think it proved that they've been on the right track all along?
I think so. I've always expected that from them, and from DeepSeek as well. I think those 2 teams are, in a way, 80% similar and 20% different. But on the similar side, they're really hardworking, driven people who have been able to develop great open-source models.
I think they'll alternate the number-one spot in the open-source model rankings over time. But being so close to the closed-source models, even if you're number 1 in the open-source-model rankings, is still amazing.
I think they were able to hit that level. If I remember correctly, when they launched Kimi K2 about a year ago, that was a big hit. It was one of the models that made the world aware of them, especially the US. At the time, it was ranked number 6 or number 7 on the closed-model list. If it had been a closed-source model, it would have been number 6 or 7.
Now they're getting close to number 2 or number 3 on that list while being an open-source model. That really surprised me as well, given how limited their resources are in terms of compute and people.
Importantly, as we discussed, they made a lot of fundamental innovations in the model architecture they just released. Although they tested many of those innovations on toy models with 40 billion or 30 billion parameters, it's still a big feat to integrate those research efforts into the actual 2.8-trillion-parameter model. That's a huge feat.
Normally, you want to stay on the safe side. When you introduce something as dramatic as those types of architectures, it usually brings down model performance rather than improving it. Over time, the net effect will probably be positive, but the short-term effect will actually drag the model down.
They were able to integrate everything together, bring the model size to almost 3 trillion parameters, and push it out with limited compute and a very short timetable. I would be lying if I said I wasn't surprised.
With those constraints in terms of compute, and then being able to fit those architectural innovations into the model, is that something that's being shared among the Chinese frontier labs?
The compute?
Yeah, I guess the—
The techniques for compute?
I mean the infrastructure-optimization techniques.
I think they've been shared across the world, more or less.
It's a little surprising to me. When I visited a lot of the frontier labs here—OpenAI, Anthropic, and Gemini—the infrastructure people especially looked up to the young engineers at DeepSeek and sometimes at Kimi as well.
A lot of the world’s innovation comes from those labs and from the really young guys who were born in 1998 or later. It’s in the hands of those people. They’re not just innovating at the best level in the country; they’re innovating at the best level in the world. That’s happening, willingly or less willingly, because they don’t have the resources, so they have to do it this way. If you make the resources scarce enough, magic will automatically happen, more or less, if you put enough great people together in that sense. So, yeah, I think that’s what we’re seeing happen.
But then, once it’s released, if you look at the latest architectures of Kimi and DeepSeek—
There are a lot of similarities, but also differences, because DeepSeek Sparse Attention, or DSA, is very different from Kimi Delta Attention, or KDA, which is Kimi’s sort of architecture. In terms of principle and philosophy, there are a lot of things that are similar. Then you see this iteration: when I release one model, they launch this feature, and then the next competitor’s model launches this feature. They go side by side, so, yeah, I think there’s definitely a lot of sharing in that.
But I don’t see the type of sharing that happens a lot in the U.S. scene, where there’s so much liquidity, or mobility, across the different firms for individual researchers. You see somebody at OpenAI one day, and 2 weeks later they’re at another firm. This happens almost every day, even at the very senior level. But in most of the labs in China, this is less frequent. They’re either in one Chinese lab and work their way up, or they may start startups, or they may come to U.S. labs. Moving across Chinese labs in a very known fashion isn’t as common as it is in the U.S.
And because of that, would you say that each Chinese lab has its own character in some way?
Yeah, I would say it’s shaped by the people who are there, but it’s also shaped mostly by the founder’s beliefs. You attract similar people to you once you set up the company as a founder. Liang is a very unique founder, and Yang is a very unique founder, but they’re unique in very different ways. Tang, obviously, is a professor, and he has his prodigies across his field. They’re all very different, and as a result, yes, I agree with you: different companies have different characters.
3. DeepSeek and unusual Liang Wenfeng
Let’s talk about DeepSeek. What’s unique about the founder, as you just mentioned, and what makes DeepSeek as a lab unique?
The guy is so unusual. I didn’t know about this guy until DeepSeek, although I should have, because his hedge fund is so important in the Chinese finance world. I never met him in person, but I have talked to people at his firm. In general, the perception is that he seems to have figured out the world at a very early stage: how to become rich, how to make money, how to trade, and how to hire talent. But, on the contrary, I don’t think he has much interest in wealth or influence—the typical things that people at this age would want to maximize.
It’s sort of even enviable to a lot of people, because he has the capability to build that, but he doesn’t capitalize on it. He has a very pure and idealistic view, almost the same words I used to describe Moonshot, in a way. They share that. They believe in hiring raw talent from China, and they’re not big fans of recruiting high-level, well-known researchers from the top labs in the U.S. or established researchers. They believe in young, fresh talent coming out of schools and building great work there. People work really, really hard.
A few weeks ago, some friends of ours in the venture-capital industry tried to take photos of companies’ floors at 9 p.m. and then at 10 p.m. to see what percentage of people were still working after 9. Those 2 are the hardest-working firms across many of the different firms out there, in terms of the percentage of people still working by then. That’s probably partially because they come to work late in the morning, but a lot of this is really voluntary. They really love working there.
The innovation they’ve done in terms of infrastructure, parallel computing, MLA, and all those types of things—DeepSeek Sparse Attention, in particular—is very innovative work. A lot of our friends here in the labs at Google, when they saw the release, were impressed because DeepSeek does a very detailed release of its technology. They open-source not only the model, but also a lot of the supporting work. They’re still one of the labs that writes thorough enough technical papers so that people can actually read and replicate what they do, rather than having to go through the whole process of figuring it out on their own.
A lot of the researchers here in the labs, when they read those papers, say, “Wow, I never knew this could be done this way.” There are really some crazy people who try to spend a lot of time grinding out the details and so forth. It’s surprising, because you would think that whatever DeepSeek and Moonshot have figured out should already have been figured out by the labs here, which are training models that are 10 times larger and dealing with problems that are 10 times harder than those of the Chinese labs. But that’s not the case.
4. Z.ai, MiniMax and what makes a great AI lab
What about Z.ai and MiniMax? Both of them also launched models recently. I think Z.ai has enjoyed a lot of great reputation since the launch of GLM-5.2, which is a model that’s doing really well in coding, and adoption of that model has been wild across China and the U.S.
They’re also very lucky to be listed early on. MiniMax and Z.ai are the 2 companies that went public in Hong Kong 8 months ago or so. The market value of Z.ai obviously skyrocketed to, was it, 1.23 trillion RMB? So almost $200 billion, which is a lot larger than many of the Chinese legacy internet companies out there. People think that’s crazy because they had very, very small revenue that year, in 2025.
But I think they’re doing some really nice work. They have a larger team, and they were built around Professor Tang Jie and his students, who have worked together for many years. They’ve expanded into a lot of different directions and product suites. They have computer use, multimodality, and different types of things. Relatively, I think they’re less focused, but nonetheless, they’re able to push out some good work. The culture, taste, or feel is a little bit different from the 2 companies we mentioned, Moonshot and DeepSeek. Z.ai is a full-fledged AI company, sort of like OpenAI in that way.
MiniMax, on the other hand, is also listed. Its stock price has enjoyed a roller-coaster ride. It went up alongside Z.ai and then dropped significantly. Today, they’re probably around their IPO price, a little higher, but probably 5 or 6 times lower than their peak.
When they started, they were known for pushing out not just models very early on, but also really good products, similar to Character.AI and so forth. They almost defined how Chinese companies should build companion AI products at the beginning. They built some very famous products and then pivoted toward building multimodality models. MiniMax M1 and the Hailuo series were very popular once they launched. Recently, they have quite a great model again, but I think they haven’t been very successful on the pure LLM side, in terms of reasoning, coding, and so forth.
Their highlights are usually around the periphery of intelligence, rather than coding or LLMs, right? It’s the multimodality, it’s the products, that sort of stuff.
I wouldn’t say this is a bad thing. I think it’s their unique trait in that way. There was a time, I think in early 2024, when I realized that all the really talented young people who would otherwise have joined ByteDance or TikTok—the cool people—all joined MiniMax at the time. They were able to recruit a lot of not just researchers, but also product managers, developers, and people in go-to-market roles. They were able to build a really attractive firm. The vibe is sort of like TikTok in a way, and they actually got a lot of people talking about joining as well. They’re unique, and I think they’re quite different from the other labs, as I mentioned.
5. ByteDance, Zhang Yiming’s bold decision, distillation debate
There has been a lot of discussion about distillation from the Chinese labs recently. ByteDance recently reportedly said in an all-hands meeting that Zhang Yiming told the staff they were going to refrain from distilling frontier models in order to preserve their technological capacity—the ability to develop their own technology. I’m really curious about your take on this debate over distillation in general and ByteDance’s decision not to do so.
I think it’s a very brave decision. Alongside that, I think I read the original comment as saying that even if it means we’re going to be behind our Chinese peers in terms of large language models, we will still prohibit distillation from the other labs and so forth. I think the key part is not prohibiting distillation. The condition for that is that people don’t think Seed or ByteDance’s large language model is a product that’s behind other people.
So it actually made me alert in that way: if you don’t distill, that means you’re going to be behind, or you’re already behind. That’s actually bigger news for me.
I think it’s quite brave because, when you think about distillation, there are 2 types of distillation. There’s one type that is blatant distillation, meaning that you have a set of prompts that you think covers the distribution of what users usually want. Then you send those to the frontier models and extract the results. In the case of Claude, you actually extract the chain of thought in between. In the case of GPT, you don’t have that, but you get a certain return back and train on top of that.
But there’s also this other type of distillation. If you think about it, a lot of the data we use today is synthetic. It’s not purely made by humans by hand, right? You wouldn’t call synthetic data distillation if you use it to train on that. But how did you create synthetic data in the first place? It’s a human using a model to create synthetic data.
So, by default, the synthetic data distills the model that you used to create it to begin with. It’s indirect distillation in that way. From that perspective, I think everybody distills something to a certain degree. It’s probably not blatant, but they distill some intelligence there.
So far, both have been proving effective. Synthetic data is obviously a substrate of all training today. It’s very important. It’s probably the key to scaling up training. But blatant distillation is very effective, as we’re probably seeing.
I have no evidence or proof of that. I don’t want to say who’s doing what in that way. But from what I see, it’s a very effective way to do so.
Having the head of ByteDance say that this is not the route they’re taking—if you go down that route, it will give us a false sense of satisfaction, because we achieved this in the wrong way and then thought we would be good, but we’re not—is a very brave decision. It’s also very suggestive of who Zhang Yiming is and what he wants to build Seed to become. We can afford to be behind for a few years, but we ultimately want to be the top of the world. In order to preserve that ambition, you have to do things the right way and in a unique way.
Can anyone afford to be behind for a few years in today’s AI competition landscape?
If you don’t have to raise funding? Yes.
Mhm.
So, in the case of Seed and in the case of DeepSeek, I think they have the luxury to do that.
Interesting. The reason why this distillation has obviously become such a hot and controversial topic in the US is because of the open-source models in China, and they have become so capable. Kimi K3 kind of proved that. But even before that, there was the DeepSeek moment last year. Then GLM-5.2 was released, and I think the US labs were alarmed at how good the Chinese models had become.
6. Are Chinese open-source models actually changing Silicon Valley?
This is when we went back to your purpose of coming to the Bay Area this time around. You wanted to see the impact of open source—I guess, the industry—on applications and the entire agentic AI ecosystem. What have you seen so far? How has open source changed AI?
The interesting observation I had is that, despite the heated debate and the great attention that has been placed on open source, Kimi K2, and so forth, if you’re talking to frontier-lab researchers and so forth, they haven’t cared much about this. Few of them have actually tried those models, and that was a surprise to me. I thought this would be your direct competition, not just from a competition standpoint, but also from a technical-report and paper-sharing perspective. You should be learning from that, but it gets less of their attention.
From an agent-developer standpoint, a lot of them, especially as seen in the news, have switched to Kimi K3. As a result, what I heard was that this has had a significant impact on the top-line revenue of the frontier labs, especially Anthropic.
We’re at the point where total token consumption in July has been growing slower than expected. Some would say it’s stagnated. If you look closer, that’s only consumption. If you look at token revenue, a lot of people are swapping from the frontier models to the cheaper models. So, if total consumption is flat—let’s say, I don’t know if it is, but let’s assume it is—total revenue will be going down because people are swapping more expensive models for cheaper models.
There’s a large impact, and the reason behind that is that you’re seeing agent companies, users, and individual coders making that change as well. There are innovative ways to do that. You might not swap all the way: you might use the likes of Kimi K3 or DeepSeek to write your code and then still use Claude to review your code, using a combination of them to maintain a certain level of accuracy and performance along the way.
For most people, you can hardly tell the difference from that perspective, unless you’re a very intensive user or you’re using it for highly complicated problems. But if you look back, even among the Anthropic models—between Fable 5, Opus 5, and Opus 4.8, sometimes even Opus 4.6—people can’t decide which one is better for their particular task. Sometimes the 4.6 is better. Sometimes Fable is better, which is supposed to be because it’s a larger model, uses more training cost, and came out later.
So, you’re seeing that all of them are more expensive than the cheaper models. Kimi K3 is actually not that cheap.
Compared to DeepSeek, which is usually the—
The cheapest option, but still cheaper, obviously, than the other models of the same size and so forth. So this is happening and hitting the token economy a lot.
I would say the agent companies really care about this, but the top frontier models care less. As a result, this discussion hasn’t been that helpful so far. You’re only seeing migration, with adoption rates going up. That’s sort of expected and what we learned, but there’s less discussion of what we can learn from those models and what we should take from them. The researchers here are commenting on the models less because they haven’t spent much time on them.
Yeah.
7. Are frontier AI models becoming commoditized?
So, are we at a stage where it has become kind of hard to differentiate between the different models for a normal company or regular users? If so, what are the implications for the frontier labs in both the US and China?
I think if your main task is not to build the next frontier model, build the next great database, build complicated coding systems, or solve the hardest math problem on Earth, then for the most part, I don’t think you’ll be able to tell the difference among the top models that easily.
The penetration of coding agents among coders has been really, really high. It would be very hard to find a coder who doesn’t use a coding agent today.
But they hope to bring non-coders up in terms of using the likes of Claude Cowork to solve, I guess, their PowerPoint or report problems.
The penetration is not that high, and the growth is not as high as expected. It’s partially due to the fact that there are 2 types of adoption. There’s adoption from the company, which is pushing from the top down. You have a business deal with your company, the company adopts it, and this pushes every employee to use it. Those are very slow—slower than expected.
Usually, their AI strategy lasts 2 to 3 years, and it takes them 2 to 3 years to fully get compliance, security, and privacy right. That’s just too long. So you’re not seeing that adoption going really fast.
On the individual level, first of all, with a lot of those tasks, you can’t really tell the difference in terms of capability among those models. Second, a lot of tasks—as you know, probably, Fable and a lot of those top models—the advantage of those is that they can run for a very long time.
If you set the right goals, set the right evaluations, and let it run, it’ll be able to figure out a complex problem while you sleep, right? The next morning you get up, and it’ll be done.
But for a lot of white-collar work, if you’re writing a PowerPoint or writing a report, it’s very hard to predefine the goal. If you make this color right in this way, then this is a great PowerPoint. It’s very hard to set this before you see the PowerPoint.
Humans are always in the loop in that way. You have to be in front of your computer, and even if the work is done by Codex or Claude, you still have to be in front of a computer to check the result every time.
And then when you see the result, you look at it, think about it, brainstorm, and probably set up some new goals. That means the token throughput cannot be very high, because it cannot run continuously for a very long time. You always have to be in the loop to do some work, check, and enhance the result.
The reason why I mentioned that I think token throughput should be a very important issue here is that if you build a PowerPoint in 10 seconds rather than 10 minutes, you can have more iterations of that loop for non-coding work and bring up the total token consumption in that area. But this is not yet happening. So far, you're seeing stagnation in those areas beyond coding.
The third area where token consumption can go up will be multimodal, multimedia models. You're seeing Seedance obviously going up, and we talk a lot about LLMs, but ByteDance has built this really powerful tool, Seedance, which is the best sort of multimodal, multimedia model coming out and making a lot of revenue—probably the best in the world by far. My personal belief is that it's the most promising tool for the next level of token consumption. But it has to be improved and become more interactive. The serving speed has to be faster for this to work.
8. Why China is leading in AI video generation
Yeah. Let's talk about multimodal models and video in particular, because you mentioned Seedance. I actually thought this was a great segue into the applications layer of AI in the US and China, because video models are one of the things where we're seeing a sharp contrast between the US and China. In China, we had Seedance and all of those startups working on video generation. In the US, those attempts have kind of faded over the past year. OpenAI kind of stopped the Sora 2 release, and I thought that was an interesting comparison. Why are the Chinese labs good at making video models?
I think, first of all, there's the resource-allocation part, and there's also the priority part, right? Coding has been so important that everything else should be less important, following that narrative. This narrative has been more and more central to the scene since, I guess, January of this year, the early part.
So everything has been cut. Everything that consumes compute is cut to preserve enough compute to make sure that coding is as good as possible. This is not to say that this is not happening in China as well. I think it's also happening there.
But if you travel back in time and say you're in 2023, and you would have to guess who would have the best multimodal model in the world 3 years later, the easy guess would be ByteDance, Kuaishou, and Meta. They have the most training data, the most multimedia platforms, the best distribution platforms, and enough GPUs and so forth. They have the data, they know what the product should train toward, and they have the distribution channel for that.
But it seems that Kling, which is the model for Kuaishou, worked out and was number 1, and then was surpassed by Seedance from ByteDance. Those 2 were expected in that way.
ByteDance's Seedance was really great because they took a very aggressive but risky shot, which was to train a much larger-scale model than all the multimedia models before that. I think they scaled it up by 1 order of magnitude, and they did it when there was very little evidence that it would work.
Looking at it retrospectively, it seems like the obvious choice, and everybody's doing it now. But back then, I think half a year ago or almost a year ago, it wasn't that obvious. They took the reward for taking that risk, I think.
Why was it a difficult bet back then?
Yeah, because most of the models are relatively small, and you use thousands of cards or GPUs for training compute rather than tens of thousands for that type of model. Being able to pull those resources out from training LLMs, and coding especially this year, is a more and more difficult decision to make, especially when you're resource-constrained.
So, yeah, that was not simple. I think that decision was led by a very, very young female researcher at ByteDance, and she deserves the credit for building this. Kudos to her. I think that's a very great move.
But what is disappointing is Meta. In a way, I suppose they should make a really great model. They have the data, the capability, and the resources to do that, and it is in their area of expertise. I think it's something they should do really well, but so far I haven't seen the progress yet.
Yeah. I guess they did make Dream, and then Muse came out a couple of months ago.
Yeah.
Would you also say there's consumer demand for this in China? Because in the beginning we talked about consumer-facing applications. MiniMax was doing that—obviously, it was a lot more companion work—but in China it seems like there are a lot of consumer-facing AI applications, and Seedance has obviously been used by everyday consumers. Everyone is making TikTok videos about it. AI dramas and short dramas are very popular on social media. Is that also what constitutes a different market in China versus the US?
I think Seedance's adoption goes up, and especially revenue goes up, but my view is that it's not yet mainly contributed by individual consumer users. It's mainly by prosumer users.
The reason you can say that is because you look at the applications that are built for prosumers: their revenue went up significantly after the release of Seedance. What are some examples?
Topview and LiblibAI, for example. Those are the ones. The 2 products are similar, but essentially they're canvases for building videos, and advertisers or short-video builders use the canvas to set up a workflow to generate certain videos.
At the beginning of the year, those companies were looking at revenue in the small millions or the mid-single digits. After Seedance's release, they quickly ramped up to almost 100 million or more. That's super-fast growth.
If you look at the deals they're making, it was even hard to get Seedance tokens at some point. You had to prepay for a certain amount in order to allocate them. Those are the buyers of Seedance's API, and the users behind that are prosumers who are making videos for AI short films and advertisements and all that stuff. As an individual, I don't think spending has gone up that high yet.
But the short-film market—the short-video market—has been going crazy, and short films have, I think, surpassed movie theaters. I think they already surpassed movie theaters this year.
Oh, this includes, I guess, both human-acted and AI short films, right?
Yes. Yes.
That makes sense.
But then, because of short films, in the past we were looking at this as an investor. We looked at a few companies doing that. My view was that AI-generated short films are great at science-fiction themes or traditional Chinese sci-fi, or ancient swordsmen, because they're really great at creating special effects.
They're not very good at creating detailed human emotional reactions. So if you're filming a high school drama with a lot of dating and small, intricate microexpressions, it's not very good for that. In the past, they were good at certain genres in Chinese film, which was good enough because those genres were the most popular to begin with.
But now, after Seedance, it's already very good at all of that. A lot of the Seedance demos are focused on 2 types of content. One is making very intricate emotional expressions: you have to laugh and then cry, with a very smooth transition across that. Sometimes it's better than humans, and sometimes better than actors.
The second is that Seedance specifically trained and spent a lot of effort making sure that highly dynamic motions are captured fairly accurately. For example, if you're engaged in a kung fu fight between 2 actors, in the past, sometimes if you hit someone, the hand would go through their arm and to the other side, or it would just look unnatural in some way.
It's very hard because it's happening very fast. The frame rate is high, and the action is compressed into a few frames. Sometimes it blurs out, and sometimes it's wrong. So I think the Seedance 2.0 version has spent a lot of effort improving that, which is one of the most important features for AI short films.
As a result, it's become almost the exclusive provider for a lot of short films in that way.
Yeah. Yeah. I actually thought it was interesting because in one of your previous interviews, you mentioned that one of the applications you look forward to is this middle ground between video and games—this interactive form of media.
Would you say that AI-generated videos are kind of a step toward that? One example you gave in that podcast was Fable Studio’s remake of South Park with AI. Are you seeing more examples like that with the advancing capabilities of AI video?
I have to say that I’m a little disappointed. Since I made that interview, I still had high hopes that it would become a mainstream type of media, but I don’t think it has played out as I expected so far.
I think naturally there should be something in between watching TV, which is leaning back and doing nothing but allowing a minimum of interaction, and playing games, which is leaning in and doing intense interaction. There should be something in between with medium interaction but endless content—a type of generated content. So it’s sort of between games and media in that way. I don’t think we’re there yet.
I think there’s this thing called Yingyou, which is sort of half movie, half game. You get to play; there’s a theme, and there’s a narrative.
I think it’s sort of a recreation of the famous game Detroit: Become Human, which is a video game.
Probably 10 years ago or so. There’s a lot of that content coming out, but I don’t think it’s of high enough quality yet to become a very popular genre. There are single hits that are making a lot of money, but it’s not taking off yet or becoming as popular as AI short films in that way. Yeah.
Yeah. But you still have hopes for that?
Yes. I would love to see that happen. I think this will be bigger than games themselves.
Nice. Yeah, I guess world models would be kind of a form of that, which we’ll get into, because we’re going to talk about new labs. That’s one of the verticals I really wanted to get into more, because in the US it has become pretty popular to invest in labs founded by young researchers, or often by experienced researchers who spend a lot of their time in an academic environment.
In China, you mentioned that you have been looking at a lot of AI-for-science companies and startups as well. Tell us a little bit more about that. What does funding for new labs look like in China compared to what you’ve done in the US?
9. Neo Labs, AI agents and the future of model companies
In the US, new labs have been the favorite of a lot of venture capital over the past year or so. There’s obviously much more funding for this to happen in the US than in China. I think China is more practical. There are new labs coming out. For example, Lin Junyang, the head of Tencent’s Hunyuan model, founded his lab. A famous researcher, Dai Cong, founded his lab.
Overall, my view is that the concept of new labs has changed, or is changing, along the way. In the past, new labs exclusively meant what I mentioned and defined just a moment ago. But now, when we think of a new lab—or what a new lab should look like—I’m thinking more like Cursor.
They’re an application-focused company to begin with. They built a super app, gathered user data, had user traces, and decided to train a model. They were able to do that precisely because they had a super app and user traces, and they were able to build a great model, which is Composer, by just doing post-training rather than pretraining. That model was able to beat a lot of the pretrained models out there in the world.
I think that’s the right recipe for a new lab these days. It’s not a new methodology; it’s that you have a new set of data that the frontier labs don’t have, and you’ll be able to capitalize on that data to build something different.
When I think about new labs, I’m thinking about the most popular apps out there today. At some point, they will decide to train their own model, either through post-training or maybe someday even pretraining. They will have to do that because they’ll want to, number 1, lower their costs, and number 2, possibly build something customized for their particular application.
You’re seeing this happen at Harvey and some of the vertical application agents here in the US, and I think you’ll be seeing this happen in China as well.
10. The next frontier of AI training data
What kind of companies would you say are most suitable for doing that? You said it was a company that has the user data and has an app.
Number 1, you have to have enough user data that is untouchable, that is unique to you. Number 2, the data is better suited if you’re in a vertical rather than a general market, because then you’re not only training a model to get a cheaper model and lower your costs, but you also actually have a shot at training a model that is unique.
If your application is general, then it’s hard to train a unique model because, essentially, users use general agents or search engines for general purposes. Their prompts are similar to what Claude and OpenAI are trained for. But if you’re looking at just a legal vertical or a medical vertical, then you’re solving a particular problem. Potentially, you can tune a model specifically for that purpose, and it’s easier to build a model that beats the general models in that way, on top of your unique data.
So I think the combination is a vertical application plus a super app. The question behind this—and I probably want to touch on this a little bit—is: What should a model be trained toward? What is the goal? What’s the evaluation for that?
Today, when you look at frontier models, everybody knows that the key to improving them is data. But who decides what data to use and purchase? It’s the researchers who own the training of the models. A lot of the data they use is attributed to the key researchers’ own tastes.
There’s some bias toward that because, if you’re a researcher, obviously you’re very familiar with coding. You’re familiar with probably math as well, but you’re probably not going to be very familiar with office work, because you personally have never done it—especially if you’re a PhD just out of your PhD program, directly entering a frontier lab and being put into a key role. You probably have never done much PowerPoint work, report work, or accounting work.
But then they’re acquiring human-labeled data and new environments to train the models on. The true environment that the model should be measured against in the future should be the exact environment that people are using, right? The Slack, Zoom, SaaS, and Salesforce software that people are using. There’s some goal in there, and then people make some effort, doing trajectories to solve certain processes within that environment.
Cursor is one of those examples. It’s one of the most popular pieces of SaaS software that people use. It just so happens that it’s coding software, so this happened really early because it’s something that researchers are very familiar with. But in the future, I think those will be the true environments—those working environments that truly measure how good the model can work. It’s not some imaginary environment imagined by the researchers or according to their research tastes.
That’s interesting, because I feel like a lot of discussions before happened around how, when you were developing foundation models, we had already used the world’s data. We had really drained the internet, and that’s why people were looking into synthetic data. But it seems like the future of data lies in those specific verticals. Would you say that these are still unmined territories, and that there’s still a lot of data when it comes to specific use cases and verticals?
Absolutely. I think there’s a long way to go. In the past, data was static. You gathered it by crawling the internet or synthesizing certain data, but those were static data.
Today, the most popular data are environments. You don’t build a piece of data; you build an environment, like SaaS software, for example. You can let your model attempt a bunch of tasks within that environment, and it will try on its own. It will do reinforcement learning along the way: once it’s successful, it will be rewarded; if it’s unsuccessful, it won’t be rewarded, and it will learn on its own.
If we’re talking about environments, then there are a lot of environments to be built. We start with the most static environments. For example, if you want to build something called kernel optimization, which is very popular in reinforcement learning today, you try to write certain CUDA operators so that your model runs faster. But the environment is set. You can rerun this a million times, and it won’t change, and you’ll be able to roll out the test equally and stably. So that’s the easiest.
Next is called partial observation. What if you can’t tell what all the parameters of the environment are? For example, if you’re in Slack, if you’re a new employee and you join a new company and you’re in the employee group chat, you’re trying to figure out what the boss wants. You don’t know what the boss wants.
You have to observe, and then gradually, over time, you probably know, because you're only given a partial view. Somebody probably tells you what the boss wants, but it's not specified; that information is untold. Over time, you learn this, maybe by exploring, interacting with other people, and so forth. So this is called partial-observation optimization, and that's sort of true of most environments. You obviously don't know the intricacy of everything to begin with.
And then there's static versus dynamic. For example, in a finance or trading environment, every action you take changes the environment itself. It will never be able to roll back to the original environment again. So every test you do changes it, and then, if it's successful, it might not be successful in the next step because the environment changed over time.
How do you adapt to that? Those are sort of new environments that you have to build, as this is new data, and you have to create them. Social media is that type of thing. Every post you make changes everybody around you. It changes your fans' perception of you, so you cannot do the same experiment 1,000 times and just see how it works, because the accumulation of the previous 999 times is already in the mindset of your fans, and it changes their behavior.
And from the virtual world to the physical world is also a direction where you can build. For example, I think there's one benchmark I love. There's one environment, or sort of data, that I love to share. It's called Vending-Bench. I don't know if you've heard of it.
We just went to their office right before this. They created a benchmark that is virtually setting up 3 different vending machines and asking an LLM to run those vending machines. The LLM can decide who to order the products from, set the price, decide on the marketing strategy, and so forth, just to see how it runs them.
That's an attempt to go from a virtual world to the real world. Although the benchmark was set up for virtual vending machines, they sort of created this in the minds of the LLMs—not to really set it up—but then they set up a market, I think at the Anton market close to here, which is really a shop, to let an LLM run the shop.
The goal is to test whether, if your LLM is so good at reasoning and doing all that work, it can run a shop in the physical world, where so many unpredictable things could happen, and still run it as well as you think it would. They change different models to do the same thing. Those data obviously aren't on the internet and cannot be crawled beforehand; they have to be made on the spot. There's a lot that can be done.
11. How US and China approaches data differently
Speaking of data, there are also a lot of data providers and data-labeling companies, with Scale AI being one of them in Silicon Valley. But there's still a lot of demand for data. I know that you've visited some of those companies as well. Tell me a little bit about the state of those companies in the US, and how different they are from the Chinese data providers.
I think we're seeing a boom in revenue over the past 6 months. I know they're making a lot of revenue already in the past year, but the past few months have seen a significant increase for a lot of them. Many of them have hit $2–3 billion in ARR at that scale. Some of the Chinese companies have been catching up—not to this scale yet, but very fast as well in terms of growth.
If you dive deep into those companies, they're different types of companies, although they all claim to be AI and LLM data providers. Some of them, like Scale and Mercor, sell their data in terms of headcount. They sort of price according to the old model: how many people they recruited to do how many hours, like a consulting or lawyer-type model.
Some of the new labs price based on environments. One environment is sold at a price ranging from a few hundred dollars to maybe $100,000, depending on how difficult it is. If it's a small Booking.com environment—trying to book a ticket—then it was probably $1,000 or less.
If you're trying to build an environment where you want to test the agent's capability—say, the DevOps capability—to turn on a cluster, get a cluster somewhere else, merge those 2 clusters, and do some crazy work in terms of DevOps, then you sort of have to recreate almost the entire AWS experience. That will probably cost close to half a million dollars. They're pricing based on that, and those are sort of new lab business models.
For the headcount models, they're largely relationship-driven, because today, as we talked about a little bit, a lot of the data is decided by researchers and their tastes. It's somewhat a traditional B2B sales problem, whereas you have to make great friends with the researchers who can make the decisions. There's a lot of wining and dining in that process to make sure it happens.
A lot of the traditional data providers are building their businesses off the typical business-development and sales model. A lot of new labs are building this off the fact that the founders of the data providers are researchers themselves. They could have gotten really high-paying jobs at a frontier lab, but they decided to become founder-led data providers.
Their unique trait is anticipating the needs of researchers in the labs 3 months ahead of time. What type of benchmark do you want to rank high on? We'll create that benchmark for you. What type of data would you need? We'll create that data for you. We'll sell you a combination of the benchmark, make it popular, and everybody will want to rank high on that benchmark. Then we'll sell you the coding data that will help you get really high on that benchmark.
That's a different way of competing. Can you be an academic opinion leader and push up your benchmark? If so, then you'll rightfully be able to sell all the data or code that will help with that benchmark.
Those are 2 different models: the prior one, where human labelers are priced based on headcount, and the latter one, which anticipates what researchers need and builds data specifically for that. I think the revenue is higher for the former, but the growth in revenue is higher for the latter, in terms of anticipating what they need and building data specifically for that. So that's really interesting.
The Chinese data providers are similar in that they started by building human-labeling businesses. In the past, one of the key benchmarks built by Scale AI was called HLE, or Humanity's Last Exam, and that was mostly a high-end expert-labeling effort.
People started with that. I think almost every data company started with that. Now they're transforming into a lot of environment builders and benchmark builders, so I think that's a trend.
But despite which side you are on, I think we're seeing revenue going crazy. Even at ICML, the academic conference, last month in—
South Korea. Yeah, South Korea.
Korea. We're seeing a lot of young researchers, and we asked them, “What's your number-one aspiration? If you want to get a job in academia, or do you want to build a startup?”
“Definitely want to build a startup, but we want to sell data. We want to be data providers.”
Oh, wow. Yeah, a lot of those aren't building model companies. They're not building agent companies; they're building data companies. So, yeah.
Let's move on to AI for science, which seems like another thing that you've been looking at during this trip. What kind of companies have you been chatting with when it comes to building AI for science?
Mostly. We're talking a lot about 2 groups. One is frontier labs, which usually have a division looking at AI for science. The other is a lot of fresh PhDs who are scientists themselves or are at the intersection between computer science and, say, biostatistics or biology—I don't know the name for that—people who work on genes, or people researching physics and materials science, and so forth.
I think they're mainly coming from these 2 groups. The third group is professors who specialize in certain fields and are looking to get some help moving into AI for science. The barrier to entering AI for science is lower nowadays with certain agents and so forth, so they're building their own capabilities in that area.
So, frontier-lab researchers who specialize in this area, PhDs at the intersection between computer science and those fields, and then some professors who are in the sciences.
12. AI for science
What's the biggest challenge in building AI for science?
I think it depends on what you're building. There's a lot of things thrown into this bucket called AI for science nowadays. A lot of them are building AI scientist agents, which are essentially agents that can read literature, design experiments, sometimes control lab equipment, analyze results, and sort of do that scientific discovery loop, and maybe, at the end, write papers and even publish papers.
That’s one area that I think isn’t that hard to build these days. Every lab is building its own AI research scientist nowadays. It’s just hard to build something that’s really good and meaningful in that way, essentially measured by whether you’re able to publish a paper that’s well received, right?
It’s not that easy to do that yet because there are a few steps. You have to analyze the literature, plug into enough databases, and plug into a certain number of tools. But the backbone models you’re using are still OpenAI and Anthropic models, and usually they’re not trained specifically on those types of scientific databases and tools. They’re not really good at using those things, and sometimes you have to build certain workflows or guardrails to make sure they do the right thing, rather than just let the model freestyle and run its own course. So that’s one of the difficulties.
Second, a lot of these experiments are done in wet labs, meaning that they have to have lab equipment and lab scientists to do the experiment, get the data back, and verify it. It’s not like math or coding, where the whole environment and all the verifiers are online, accurate, and stable. So it usually takes a long time, or sometimes it’s just disconnected. Certain agents can go as far as experiment design, but executing the experiment needs a human to intervene in that way.
There are companies like BioMap and Lila Sciences, and they’re all doing very well and building useful tools for researchers. There’s also the other side, where you’re building a model specifically for that. Since AlphaFold, DeepMind has obviously been building great models out there. Since AlphaFold 3, a lot of companies have been coming out and building frontier, state-of-the-art models, like Chai, like Boltz, and like DeepMind’s own Isomorphic Labs. We’ve invested in companies along the way that are building protein-design models as well as complex-structure-prediction models.
These models are trained on public and private datasets to be able to design things, and nowadays, de novo design—meaning design from nothing. You’re just creating something that doesn’t exist in nature, such as antibodies or proteins, and enabling them to bind to certain known targets so they can solve certain diseases.
We’re also looking to build world models for cells—virtual cells. This is essentially the flip side of the protein-design models: you can design a protein binder, but how would you test whether it can bind well or not? You have to design a cell in order to see the interaction that way. Otherwise, you have to do this experiment in a real lab, which is very costly and time-consuming.
That’s also very complicated because, unlike robotics or the real world, where you have cameras and all those sensors, for a cell you have very, very minimal data to build anything or sustain a large model to train on. So the data shortage is probably the biggest problem here. Multimodality is also a problem: enzymes, antibodies, antigens, and peptides are very different things. How would you encode them all into the same model and be able to represent them together? Otherwise, you have to build small models for each, and then that defeats the purpose of building one unified large language model, or a large model of any sort, to be able to do the prediction.
Those are some of the problems we’re seeing today. Essentially, all the things we’re doing in biology—I gave you a lot of examples in biology—but the true verifier is whether this drug works on a human body and cures the disease, right? We can have intermediate verifiers today: whether it binds well, or whether you do it in a wet lab. But the true verifier is whether this drug works on a human body and cures the disease. You won’t even be able to do that experiment unless you’re 5 or 6 years into the process, until you get to a clinical trial approved by the FDA. You don’t get any real, full feedback from a clinical trial until you’re a few years into this process. You can’t imagine that happening with an LLM or a physical-world model, but this is a reality for drug-design models today.
Yeah, very interesting. A really high-stakes drug-discovery problem with a lot to expect there.
Yeah.
Great. So we’ve talked a lot about foundation models, the difference between the AI ecosystems in the U.S. and China, and different applications. My last question is: what would you say is the 1 biggest misconception that Silicon Valley has about China’s AI ecosystem, China’s AI development, and tech in general?
13. Silicon Valley’s biggest misconception about China’s AI ecosystem
I won’t attempt to say that I understand the whole perception of the U.S. ecosystem of China. But I think, in my opinion, the U.S. ecosystem thinks of China’s AI ecosystem, the labs, or the technology as the result of improvement that’s entirely tied to government support, either through funding, compute, or all that stuff. They give a lot of credit to how the government supports this.
I think they haven’t given enough credit to the individuals, the entrepreneurs, and the researchers who are actually making this happen under very difficult constraints in the world. I think they would be able to do this anywhere in the world if they’re able to do this in China today. So that’s probably the biggest misconception.
Yeah, totally. Great. I think that’s all the time that we have today. Thank you so much for joining us. Any last words?
No, thank you. Thank you. Really glad to be here. Thank you.