[BidClub_]
Silicon Valley 101 · · 72 min

Investor Meng Xing on open source impacts, Chinese AI personalities, and the future of RL data

Xing Meng

YouTube
TL;DR
  • Kimi K3 convinced Xing Meng that China's open-source labs are now frontier-adjacent, not fast followers. Kimi remains a sub-400-person, idealistic lab focused on AGI. It integrated fundamental architecture innovations, including KDA, some tested only on 30–40B toy models, into a 2.8-trillion-parameter model on constrained compute. Meng recalls that K2, launched about a year ago, would have ranked sixth or seventh on the closed-model list; K3 is getting close to second or third. He expects Kimi and DeepSeek to alternate the number-one spot in open-source model rankings over time.
  • The reported tradeable impact of Chinese open source is on token revenue, not frontier researchers' attention. Frontier-lab researchers have not cared much about these models, and few have tried them, while many agent developers have switched to K3. Meng heard that this significantly affected frontier labs' top-line revenue, especially Anthropic's: July token consumption grew more slowly than expected—some say it stagnated—and if consumption is flat while users swap expensive models for cheaper ones, revenue falls. Users may combine models, such as using K3 or DeepSeek to write code and Claude to review it. Even among Anthropic's Fable 5, Opus 5, Opus 4.8, and sometimes Opus 4.6, users cannot always decide which is better for a task.
  • ByteDance's reported decision not to distill frontier models is brave and suggestive, not simply an admission. Meng found the reported condition—accepting that Seed could fall behind Chinese peers—especially significant: if a company refuses distillation, it may already be behind or risk falling behind. He distinguishes blatant distillation from indirect distillation through synthetic data, arguing that everyone distills something to some degree. He says Seed and DeepSeek can afford to be behind for a few years because they do not have to raise funding, preserving the ambition to reach the top by doing things “the right way.”
  • ByteDance's Seedance emerged from a risky scale bet with little prior evidence that it would work. Meng says the model was scaled up by an order of magnitude over earlier multimedia models, requiring tens of thousands rather than thousands of GPUs, at a time when compute was increasingly needed for LLMs and coding. He credits a very young female ByteDance researcher. Revenue appears to be driven mainly by prosumers: video-canvas services such as Topview and LiblibAI reportedly grew from small or mid-single-digit millions to almost 100 million or more after Seedance's release, with some buyers prepaying for API tokens. The host said short films had surpassed movie theaters in China this year; Meng accepted that comparison after clarifying it included both human-acted and AI short films.
  • The new-lab concept is changing: Meng sees Cursor's recipe as increasingly important. A vertical application with unique, “untouchable” user-trace data can post-train a model such as Composer and potentially outperform many pretrained models. The advantage is not a new methodology but access to data frontier labs lack. Vertical applications such as legal or medical tools may have a better chance of building a distinctive model than general-purpose apps. The deeper problem is that researchers' data choices reflect their familiarity with coding and math, while the real test environments should be the working software people use—Slack, Zoom, Salesforce, and similar systems.
  • The future of training data is environments, with difficulty ranging from stable optimization tasks to partially observed and dynamic worlds. Meng describes static kernel-optimization environments, Slack-like partial observation, and finance or social-media environments where every action changes the next state. Vending-Bench extends the idea toward the physical world by having models run virtual vending machines; its creators also set up a real shop or market nearby to test an LLM in a less predictable setting.
  • Data provision is booming through two models. Headcount-based providers such as Scale and Mercor price recruited labor like consulting, while newer providers sell environments ranging from a few hundred dollars to roughly $100,000 or more; a Booking.com task might cost about $1,000 or less, while recreating much of an AWS-like DevOps environment could approach $500,000. Meng says the former currently has higher revenue, while the anticipatory researcher-led model has faster growth. At a recent conference in South Korea, young researchers told him their top aspiration was to build data-provider startups.
  • AI for science's binding constraint is verification, not simply modeling. AI-scientist agents can read literature, design experiments, sometimes control lab equipment, analyze results, and write papers, but scientific databases, tools, workflows, and wet-lab execution remain difficult. Protein-design models face severe data and multimodality problems, while virtual-cell models aim to test biological interactions without costly experiments. The true verifier is whether a drug works in a human body and cures disease, but full feedback may not arrive until an FDA-approved clinical-trial stage roughly five or six years into the process.
  • Meng's stated misconception in Silicon Valley is that China's AI progress is entirely a product of government support. He argues that the US ecosystem has not given enough credit to the individuals, entrepreneurs, and researchers producing this work under difficult constraints, and believes they could do it anywhere if they can do it in China.
Digest · the substance, structured for research

1. Kimi is a sub-400-person idealistic lab that just shipped a near-frontier model

  • Meng's portrait: “a hub of really smart young talent gathered together to achieve a very idealistic goal” — AGI, held through three years of peak media attention, “horrible times,” and a return to the peak, without diluting into applications or multimedia. Interns “almost in high school” get real compute and built important components of K3.
  • The headcount is deliberate: Xing said Kimi is still fewer than 400 people; the host estimated “300 or so” after three years, not for lack of resources but so the team stays intact and avoids context loss.
  • Did K3 surprise him? Yes: Xing recalled that K2, launched about a year ago, would have ranked sixth or seventh on the closed-model list; K3 is near second or third while open. Integrating fundamental architecture innovations — some validated only on 30–40B “toy models” — into a 2.8-trillion-parameter run would normally drag short-term performance down; they did it on limited compute and a short timetable. “I would be lying if I said I wasn't surprised.”

2. DeepSeek's founder, scarcity-bred innovation, and why each Chinese lab has a character

  • On DeepSeek's founder, whom he has never met: “He seems to have figured out the world at a very early stage” — how to become rich, make money, trade, and hire talent — but does not seem very interested in wealth or influence and does not capitalize on his ability to build them. The lab hires raw young talent out of Chinese schools rather than high-level, well-known researchers from top US labs; VC friends photographing office floors at 9–10 p.m. found DeepSeek and Kimi to be two of the hardest-working firms by the share of people still working. “A lot of this is really voluntary.”
  • The status inversion he keeps encountering: infrastructure people at OpenAI, Anthropic, and Gemini “look up to the young guys at DeepSeek,” including engineers born in 1998 or later. The mechanism is constraint: “If you make the resources scarce enough, magic will automatically happen” when enough great people are put together. DeepSeek also releases unusually detailed technical papers and supporting work; researchers at Google and elsewhere say, “I never knew this could be done this way.”
  • Kimi Delta Attention and DeepSeek Sparse Attention differ technically but share some principles. Features appear to leapfrog from one release to the next. Unlike the US, where senior researchers can move between firms frequently, Chinese researchers less often move between Chinese labs; each lab therefore develops a distinct character shaped by its founder's beliefs and the people attracted by them.

3. Z.ai and MiniMax: two listed labs, different characters

  • The host introduced Z.ai by pointing to the strong coding adoption of GLM-5.2. Xing said Z.ai is doing good work, has a larger team built around Professor Tang Jie and his students, and has expanded into computer use, multimodality, and other product suites. Its market value, he asked, had reached 1.23 trillion RMB — nearly $200 billion — despite very small 2025 revenue. It is relatively less focused than Moonshot and DeepSeek, more like a full-fledged AI company in the style of OpenAI.
  • MiniMax's stock has been “a roller coaster”: it rose alongside Z.ai, then fell significantly, leaving it around or slightly above its IPO price and roughly five or six times below its peak. It initially built well-known companion products, similar to Character.AI, and helped define how Chinese companies built such products. It later shifted toward multimodality models, including MiniMax M1 and the Hailuo series, but Xing says it has been less successful on the pure LLM side of reasoning and coding.
  • The talent tell from early 2024: many of the “cool people” who might otherwise have joined ByteDance or TikTok joined MiniMax, including product managers, developers, researchers, and go-to-market staff. “The vibe is sort of like TikTok.”

4. ByteDance's no-distillation decision: brave, suggestive, and affordable for only some firms

  • Zhang Yiming reportedly told staff that Seed would refrain from distilling frontier models “even if it means we're going to be behind our Chinese peers.” Meng fixated on the conditional: “If you don't distill, that means you're going to be behind or you're already behind. That's actually the bigger news for me.”
  • His taxonomy: blatant distillation sends a prompt distribution to frontier models and trains on their returns, including Claude's chain of thought where available. Indirect distillation occurs through synthetic data: because models generate much of that data, “by default the synthetic data distills the model which you used to create it.” Everyone therefore distills something to some degree. Meng says both synthetic data and blatant distillation have proved effective, while adding that he has no proof about who is doing what.
  • Why refuse it: the shortcut “will give us a false sense of satisfaction, because we achieved this in the wrong way.” When the host asked whether anyone could afford to be behind for a few years, Meng answered: “If you don't have to raise funding, yes.” He gave Seed and DeepSeek as examples with that luxury.

5. Open source's reported bite is the token economy, not frontier researchers' attention

  • Meng's surprise from the trip: despite the K3 debate and attention to open source, frontier-lab researchers “haven't cared much about this”; few had tried the models. Agent developers, by contrast, had switched to K3 in substantial numbers, particularly according to news coverage.
  • Xing heard that this had a significant effect on the top-line revenue of frontier labs, especially Anthropic. July total token consumption was growing more slowly than expected, and some people said it had stagnated. If consumption were flat, he said, swapping expensive models for cheaper ones would reduce revenue. The hybrid pattern is to use K3 or DeepSeek to write code while retaining Claude to review it.
  • Differentiation is narrowing for ordinary tasks: “For the most part, I don't think you'll be able to tell the difference among the top models that easily.” Even among Anthropic's Fable 5, Opus 5, Opus 4.8, and sometimes Opus 4.6, users cannot always decide which is better for a particular task. K3 is not especially cheap, but it is cheaper than comparable models.
  • Why non-coding tokens stall: company-led AI adoption can take two to three years because of compliance, security, and privacy work. White-collar tasks also lack easily predefined goals — “it's very hard to set this before you see the PowerPoint” — so humans remain in the loop and throughput is capped. His proposed lever is speed: creating a PowerPoint in 10 seconds rather than 10 minutes would allow more iterations. “This is not yet happening.”

6. Seedance's risky scale bet built a prosumer short-film economy

  • Looking back to 2023, Meng says the obvious guess for the best future multimedia model would have been ByteDance, Kuaishou, or Meta, given their data, distribution, platforms, and GPUs. Kuaishou's Kling reached number one before being surpassed by ByteDance's Seedance. ByteDance scaled Seedance up by an order of magnitude over prior multimedia models, requiring tens of thousands rather than thousands of GPUs, when there was little evidence the approach would work. He credits a very young female researcher at ByteDance. Meta, despite having relevant data and capabilities, has not yet shown the progress he expected.
  • Revenue is mainly prosumer-driven, not individual-consumer-driven: video-canvas services such as Topview and LiblibAI reportedly went from small or mid-single-digit millions in revenue early in the year to almost 100 million or more after Seedance's release. At one point, buyers had to prepay to secure Seedance API tokens.
  • The host said short films had surpassed movie theaters in China this year. Xing first clarified that this included both human-acted and AI short films, then said it made sense. He said AI-generated short films had previously been strongest in science fiction, traditional Chinese science fiction, and ancient-swordsman genres, where special effects mattered more than subtle human emotion. Seedance's demos showed smoother emotional transitions, such as laughing into crying, and more accurate high-dynamic motion, such as kung-fu fights without hands passing through arms. As a result, it had become almost the exclusive provider for many short films.
  • On interactive media between television and games, Xing revised his earlier optimism: “I have to say that I'm a little disappointed.” The host cited Detroit: Become Human as an example of the half-movie/half-game format; Xing said such content has produced individual hits but has not yet reached high enough quality or popularity to take off. He still hopes it will happen and says, “I think this will be bigger than games themselves.”

7. New labs are being redefined: the Cursor recipe and the researcher-taste problem

  • The earlier new-lab concept centered on young researchers or experienced researchers with academic backgrounds. Xing now thinks more about Cursor: an application-focused company with a super app and user traces that trained Composer through post-training rather than pretraining, allowing it to beat many pretrained models. “It's not a new methodology — it's a new set of data that the frontier labs don't have.”
  • Criteria: the data must be unique and “untouchable,” and a vertical is better suited than a general market. A legal or medical application could potentially tune a model to a distinctive task and outperform general models, whereas general-agent prompts resemble data already used to train Claude and OpenAI models. Xing cited Harvey and other vertical application agents in the US and expects similar development in China.
  • The deeper argument is that data choices are “attributed to the key researchers' own taste.” A fresh PhD may know coding and math but have “probably never done much PowerPoint work or accounting work.” The true measure of a model should be the working environments people actually use — Slack, Zoom, Salesforce, and other SaaS systems — rather than an imagined environment shaped by researchers' preferences. Cursor came early partly because coding is familiar to researchers.

8. The future of training data is environments — graded by how much the world fights back

  • Data has moved from static material, whether crawled or synthetic, toward environments where models attempt tasks and learn from rewards. Xing's taxonomy ranges from static kernel/CUDA optimization — “you can rerun this a million times and it won't change” — to partial observation, such as a new employee in Slack “trying to figure out what the boss wants,” and dynamic settings such as trading or social media, where every action changes the next state. A social-media post changes followers' perceptions, so the same experiment cannot simply be repeated 1,000 times.
  • Vending-Bench tests the virtual-to-physical transition by having an LLM operate three virtual vending machines: choosing suppliers, setting prices, and deciding on marketing. The creators told Xing they had also set up a nearby market or shop to test whether a model could manage a physical-world business with unpredictable conditions. “Those data are not on the internet... they have to be made on the spot.”

9. Data providers are booming on two very different business models

  • Revenue has surged in the past six months, with many providers reaching $2–3 billion in ARR at that scale; Chinese providers have not reached that scale but are growing quickly. Model one: headcount pricing, used by Scale and Mercor, sells recruited labor like consulting and is relationship-driven B2B, involving substantial wining and dining of the researchers who choose the data.
  • Model two sells environments, from a few hundred dollars upward. A Booking.com-style booking task might cost about $1,000 or less; recreating nearly the entire AWS experience for a difficult DevOps agent could approach $500,000. Researcher-founders anticipate lab needs three months ahead: create a benchmark, make it popular, and then sell the data that helps models rank highly on it. Xing thinks headcount models currently have higher revenue, while the anticipatory model has faster revenue growth.
  • At the academic conference in South Korea last month, young researchers told him their top aspiration was founding startups that sell data: “They're not building model companies, they're not building agent companies — they're building data companies.” The host described a broader shift from human labeling toward environment and benchmark building; Xing said Chinese providers similarly began with human-labeling businesses. He mentioned HLE, or Humanity's Last Exam, as an example of high-end expert labeling.

10. AI for science: easy to build, hard to verify — plus Silicon Valley's China blind spot

  • AI-scientist agents can read literature, design experiments, sometimes control lab equipment, analyze results, and write or publish papers. Xing says they are “not that hard to build these days,” but it is hard to make one genuinely good and meaningful, measured in part by whether it can produce a well-received paper. The blockers include backbone OpenAI and Anthropic models that are not specifically trained on scientific databases and tools, plus wet-lab experiments where execution, data collection, and verification require time and human intervention.
  • On the model side, since AlphaFold and especially AlphaFold 3, efforts from DeepMind and companies such as Chai and Boltz have pursued state-of-the-art protein design and complex-structure prediction. Xing also mentioned companies they had invested in building protein-design models. These models can attempt de novo design — creating proteins or antibodies that do not exist in nature — while virtual-cell world models could test whether a designed binder interacts as intended without a costly wet-lab experiment. The biggest problems are limited cell data and multimodality: enzymes, antibodies, antigens, and peptides must be represented together rather than handled only by disconnected small models.
  • The structural verifier gap remains severe: intermediate verifiers exist, but “the true verifier is if this drug works on a human body and if it cures the disease.” Full feedback may not arrive until an FDA-approved clinical-trial stage roughly five or six years into the process. “You can't imagine that happening with an LLM... but this is a reality for drug-design models today.”
  • Xing's stated misconception in Silicon Valley is that China's AI progress is entirely tied to government support through funding, compute, and similar help. He argues that the US ecosystem has not given enough credit to the individuals, entrepreneurs, and researchers making progress under difficult constraints, and believes they could do it anywhere if they can do it in China.

1. Kimi, Moonshot AI and China’s young AI researchers

Xing Meng

Kimi is a very idealistic company and a hub of really smart young talent gathered together to achieve a very idealistic goal: AGI. DeepSeek is also made up of hardworking, driven people who have been able to develop great open-source models. I think they'll be able to alternate the number-one spot among open-weight models over time.

Speaker 1

Let's talk about this debate over distillation in general and ByteDance's decision not to go down that route.

Xing Meng

I think it's a very brave decision by the head of ByteDance to say that this is not the route we're taking. If we go down that route, it'll give us a false sense of satisfaction because we achieved this in the wrong way.

Speaker 1

How has open source changed AI?

Xing Meng

If we're talking about frontier labs, researchers, and so forth, they haven't cared much about this.

Speaker 1

Xing, welcome to Valley 101.

Xing Meng

Thank you for having me.

Speaker 1

2. Silicon Valley trip recap

You split your time between Beijing and Silicon Valley, and you come here once every quarter. What was your biggest takeaway from this trip to the Bay Area?

Xing Meng

Before we came here, about a week ago, we had a few themes in mind, and we wanted to explore how they played out in the Valley. We wondered about models like Kimi and DeepSeek. We saw a lot of buzz and discussion on Twitter and even in Congress, but what is their real impact on frontier labs, agent companies, and researchers in general? Are they really that concerned? Are they really paying that much attention to them?

How does that change the economics of the world, especially leading up to the IPOs of OpenAI and Anthropic? Does that have an impact? We thought it would, but we wanted to see some proof.

The second theme is that, in the past 3 years since the launch of GPT-3.5, I think the key theme has always been pushing out the best next model and reaching the ultimate performance beyond imagination—being at the top of every benchmark possible. Now we're reaching a point where the model companies want to reach more users, and those users have to go beyond coders and other specialized users.

They want adoption rates to go up and to reach more diverse users. So you're looking at people who care a lot about efficiency and serving speed rather than only the ultimate level of intelligence. Token efficiency is important, and so is the cost of routing the right model to serve a specific task.

The third theme is AI for science. That has been a big factor and one of the big focuses for me and my team for almost a year now. We're looking at breakthroughs across biology, materials science, and even traditional drilling for minerals and applying AI in those directions.

Every time we come here, we explore those things because Silicon Valley is not only a hub for AI but also a hub for general science. Those are the things I'm most looking forward to and will spend a lot of time on.

Speaker 1

Open source, this transition from capability to speed and inference, and then AI for science. We'll dig into each of those.

Let's start with the foundation models. One of the things I find really interesting in the US is that we're getting into this two-tier ecosystem. We have Anthropic and OpenAI, which are the duopoly in the startup world, and then we have the public companies and bigger companies, like Google and Meta.

Things look similar in China as well. We have the so-called four dragons of AI startups that focus on foundation models, and then we have the bigger companies like ByteDance and Alibaba. Would you agree with that characterization? Is that a good way of looking at the foundation-model labs in China?

Xing Meng

Sure. I think those are the main ones.

Speaker 1

I wanted to talk a little bit about the specific companies, and I'd love to know how you think of them from an investor perspective. Let's start with Moonshot, since we've been talking about Kimi K3 for a while. What kind of company is that?

Xing Meng

I think Kimi is a very idealistic company. They're a hub of really smart young talent gathered together to achieve a very idealistic goal, which is AGI. At the very beginning, 3 years ago, that sounded unusual. Today it sounds normal—everybody is saying that—but 3 years ago, people would have thought you were too ambitious for your credentials in that way. They held on to that ambition and aspiration throughout the past 3 years.

They've been at the peak of media attention, they've gone through horrible times, and now they're back at the peak again. But I don't think they've changed much in terms of what they think or what they do. They're very focused. They're not an application company, and they don't build a lot of multimodal models. They're laser-focused on building AGI, in a sense.

They're one of the companies that probably has the best atmosphere for young researchers to join. They give a lot of attention and resources to interns and really young interns. They have interns who are almost in high school, which is not very usual, and those interns are able to access a lot of compute power and resources so that they can build important components of the latest model, Kimi K3, which was just launched.

I think they're a team-first or culture-first company. They establish a very attractive goal, attract talent, and remain laser-focused on that goal. If you talk to people in that organization, I think they'll tell you the same thing. They're a very small team today. I think they're still fewer than 400 people.

Speaker 1

More or less. If we talk to the companies here in the US, and if I were to ask them to guess, after seeing Kimi K3, the buzz, and the performance, how many people they would guess this company has, they would be really surprised to learn that it only has 300 or so people after 3 years of development.

It's not that they don't have the resources to hire more people. They decided to keep the team small so that they could remain intact enough for everybody to know what everyone else was working on. They're close enough that there's no context loss when they're building things together.

Knowing that team, did the release of Kimi K3 surprise you? Do you think it proved that they've been on the right track all along?

Xing Meng

I think so. I've always expected that from them, and from DeepSeek as well. I think those 2 teams are, in a way, 80% similar and 20% different. But on the similar side, they're really hardworking, driven people who have been able to develop great open-source models.

I think they'll alternate the number-one spot in the open-source model rankings over time. But being so close to the closed-source models, even if you're number 1 in the open-source-model rankings, is still amazing.

I think they were able to hit that level. If I remember correctly, when they launched Kimi K2 about a year ago, that was a big hit. It was one of the models that made the world aware of them, especially the US. At the time, it was ranked number 6 or number 7 on the closed-model list. If it had been a closed-source model, it would have been number 6 or 7.

Now they're getting close to number 2 or number 3 on that list while being an open-source model. That really surprised me as well, given how limited their resources are in terms of compute and people.

Importantly, as we discussed, they made a lot of fundamental innovations in the model architecture they just released. Although they tested many of those innovations on toy models with 40 billion or 30 billion parameters, it's still a big feat to integrate those research efforts into the actual 2.8-trillion-parameter model. That's a huge feat.

Normally, you want to stay on the safe side. When you introduce something as dramatic as those types of architectures, it usually brings down model performance rather than improving it. Over time, the net effect will probably be positive, but the short-term effect will actually drag the model down.

They were able to integrate everything together, bring the model size to almost 3 trillion parameters, and push it out with limited compute and a very short timetable. I would be lying if I said I wasn't surprised.

Speaker 1

With those constraints in terms of compute, and then being able to fit those architectural innovations into the model, is that something that's being shared among the Chinese frontier labs?

Xing Meng

The compute?

Speaker 1

Yeah, I guess the—

Xing Meng

The techniques for compute?

Speaker 1

I mean the infrastructure-optimization techniques.

Xing Meng

I think they've been shared across the world, more or less.

Speaker 1

It's a little surprising to me. When I visited a lot of the frontier labs here—OpenAI, Anthropic, and Gemini—the infrastructure people especially looked up to the young engineers at DeepSeek and sometimes at Kimi as well.

Xing Meng

A lot of the world’s innovation comes from those labs and from the really young guys who were born in 1998 or later. It’s in the hands of those people. They’re not just innovating at the best level in the country; they’re innovating at the best level in the world. That’s happening, willingly or less willingly, because they don’t have the resources, so they have to do it this way. If you make the resources scarce enough, magic will automatically happen, more or less, if you put enough great people together in that sense. So, yeah, I think that’s what we’re seeing happen.

Speaker 1

But then, once it’s released, if you look at the latest architectures of Kimi and DeepSeek—

Xing Meng

There are a lot of similarities, but also differences, because DeepSeek Sparse Attention, or DSA, is very different from Kimi Delta Attention, or KDA, which is Kimi’s sort of architecture. In terms of principle and philosophy, there are a lot of things that are similar. Then you see this iteration: when I release one model, they launch this feature, and then the next competitor’s model launches this feature. They go side by side, so, yeah, I think there’s definitely a lot of sharing in that.

Speaker 1

But I don’t see the type of sharing that happens a lot in the U.S. scene, where there’s so much liquidity, or mobility, across the different firms for individual researchers. You see somebody at OpenAI one day, and 2 weeks later they’re at another firm. This happens almost every day, even at the very senior level. But in most of the labs in China, this is less frequent. They’re either in one Chinese lab and work their way up, or they may start startups, or they may come to U.S. labs. Moving across Chinese labs in a very known fashion isn’t as common as it is in the U.S.

And because of that, would you say that each Chinese lab has its own character in some way?

Xing Meng

Yeah, I would say it’s shaped by the people who are there, but it’s also shaped mostly by the founder’s beliefs. You attract similar people to you once you set up the company as a founder. Liang is a very unique founder, and Yang is a very unique founder, but they’re unique in very different ways. Tang, obviously, is a professor, and he has his prodigies across his field. They’re all very different, and as a result, yes, I agree with you: different companies have different characters.

Speaker 1

3. DeepSeek and unusual Liang Wenfeng

Let’s talk about DeepSeek. What’s unique about the founder, as you just mentioned, and what makes DeepSeek as a lab unique?

Xing Meng

The guy is so unusual. I didn’t know about this guy until DeepSeek, although I should have, because his hedge fund is so important in the Chinese finance world. I never met him in person, but I have talked to people at his firm. In general, the perception is that he seems to have figured out the world at a very early stage: how to become rich, how to make money, how to trade, and how to hire talent. But, on the contrary, I don’t think he has much interest in wealth or influence—the typical things that people at this age would want to maximize.

It’s sort of even enviable to a lot of people, because he has the capability to build that, but he doesn’t capitalize on it. He has a very pure and idealistic view, almost the same words I used to describe Moonshot, in a way. They share that. They believe in hiring raw talent from China, and they’re not big fans of recruiting high-level, well-known researchers from the top labs in the U.S. or established researchers. They believe in young, fresh talent coming out of schools and building great work there. People work really, really hard.

A few weeks ago, some friends of ours in the venture-capital industry tried to take photos of companies’ floors at 9 p.m. and then at 10 p.m. to see what percentage of people were still working after 9. Those 2 are the hardest-working firms across many of the different firms out there, in terms of the percentage of people still working by then. That’s probably partially because they come to work late in the morning, but a lot of this is really voluntary. They really love working there.

The innovation they’ve done in terms of infrastructure, parallel computing, MLA, and all those types of things—DeepSeek Sparse Attention, in particular—is very innovative work. A lot of our friends here in the labs at Google, when they saw the release, were impressed because DeepSeek does a very detailed release of its technology. They open-source not only the model, but also a lot of the supporting work. They’re still one of the labs that writes thorough enough technical papers so that people can actually read and replicate what they do, rather than having to go through the whole process of figuring it out on their own.

A lot of the researchers here in the labs, when they read those papers, say, “Wow, I never knew this could be done this way.” There are really some crazy people who try to spend a lot of time grinding out the details and so forth. It’s surprising, because you would think that whatever DeepSeek and Moonshot have figured out should already have been figured out by the labs here, which are training models that are 10 times larger and dealing with problems that are 10 times harder than those of the Chinese labs. But that’s not the case.

Speaker 1

4. Z.ai, MiniMax and what makes a great AI lab

What about Z.ai and MiniMax? Both of them also launched models recently. I think Z.ai has enjoyed a lot of great reputation since the launch of GLM-5.2, which is a model that’s doing really well in coding, and adoption of that model has been wild across China and the U.S.

Xing Meng

They’re also very lucky to be listed early on. MiniMax and Z.ai are the 2 companies that went public in Hong Kong 8 months ago or so. The market value of Z.ai obviously skyrocketed to, was it, 1.23 trillion RMB? So almost $200 billion, which is a lot larger than many of the Chinese legacy internet companies out there. People think that’s crazy because they had very, very small revenue that year, in 2025.

But I think they’re doing some really nice work. They have a larger team, and they were built around Professor Tang Jie and his students, who have worked together for many years. They’ve expanded into a lot of different directions and product suites. They have computer use, multimodality, and different types of things. Relatively, I think they’re less focused, but nonetheless, they’re able to push out some good work. The culture, taste, or feel is a little bit different from the 2 companies we mentioned, Moonshot and DeepSeek. Z.ai is a full-fledged AI company, sort of like OpenAI in that way.

MiniMax, on the other hand, is also listed. Its stock price has enjoyed a roller-coaster ride. It went up alongside Z.ai and then dropped significantly. Today, they’re probably around their IPO price, a little higher, but probably 5 or 6 times lower than their peak.

When they started, they were known for pushing out not just models very early on, but also really good products, similar to Character.AI and so forth. They almost defined how Chinese companies should build companion AI products at the beginning. They built some very famous products and then pivoted toward building multimodality models. MiniMax M1 and the Hailuo series were very popular once they launched. Recently, they have quite a great model again, but I think they haven’t been very successful on the pure LLM side, in terms of reasoning, coding, and so forth.

Speaker 1

Their highlights are usually around the periphery of intelligence, rather than coding or LLMs, right? It’s the multimodality, it’s the products, that sort of stuff.

Xing Meng

I wouldn’t say this is a bad thing. I think it’s their unique trait in that way. There was a time, I think in early 2024, when I realized that all the really talented young people who would otherwise have joined ByteDance or TikTok—the cool people—all joined MiniMax at the time. They were able to recruit a lot of not just researchers, but also product managers, developers, and people in go-to-market roles. They were able to build a really attractive firm. The vibe is sort of like TikTok in a way, and they actually got a lot of people talking about joining as well. They’re unique, and I think they’re quite different from the other labs, as I mentioned.

Speaker 1

5. ByteDance, Zhang Yiming’s bold decision, distillation debate

There has been a lot of discussion about distillation from the Chinese labs recently. ByteDance recently reportedly said in an all-hands meeting that Zhang Yiming told the staff they were going to refrain from distilling frontier models in order to preserve their technological capacity—the ability to develop their own technology. I’m really curious about your take on this debate over distillation in general and ByteDance’s decision not to do so.

Xing Meng

I think it’s a very brave decision. Alongside that, I think I read the original comment as saying that even if it means we’re going to be behind our Chinese peers in terms of large language models, we will still prohibit distillation from the other labs and so forth. I think the key part is not prohibiting distillation. The condition for that is that people don’t think Seed or ByteDance’s large language model is a product that’s behind other people.

So it actually made me alert in that way: if you don’t distill, that means you’re going to be behind, or you’re already behind. That’s actually bigger news for me.

I think it’s quite brave because, when you think about distillation, there are 2 types of distillation. There’s one type that is blatant distillation, meaning that you have a set of prompts that you think covers the distribution of what users usually want. Then you send those to the frontier models and extract the results. In the case of Claude, you actually extract the chain of thought in between. In the case of GPT, you don’t have that, but you get a certain return back and train on top of that.

But there’s also this other type of distillation. If you think about it, a lot of the data we use today is synthetic. It’s not purely made by humans by hand, right? You wouldn’t call synthetic data distillation if you use it to train on that. But how did you create synthetic data in the first place? It’s a human using a model to create synthetic data.

So, by default, the synthetic data distills the model that you used to create it to begin with. It’s indirect distillation in that way. From that perspective, I think everybody distills something to a certain degree. It’s probably not blatant, but they distill some intelligence there.

So far, both have been proving effective. Synthetic data is obviously a substrate of all training today. It’s very important. It’s probably the key to scaling up training. But blatant distillation is very effective, as we’re probably seeing.

I have no evidence or proof of that. I don’t want to say who’s doing what in that way. But from what I see, it’s a very effective way to do so.

Having the head of ByteDance say that this is not the route they’re taking—if you go down that route, it will give us a false sense of satisfaction, because we achieved this in the wrong way and then thought we would be good, but we’re not—is a very brave decision. It’s also very suggestive of who Zhang Yiming is and what he wants to build Seed to become. We can afford to be behind for a few years, but we ultimately want to be the top of the world. In order to preserve that ambition, you have to do things the right way and in a unique way.

Speaker 1

Can anyone afford to be behind for a few years in today’s AI competition landscape?

Xing Meng

If you don’t have to raise funding? Yes.

Speaker 1

Mhm.

Xing Meng

So, in the case of Seed and in the case of DeepSeek, I think they have the luxury to do that.

Speaker 1

Interesting. The reason why this distillation has obviously become such a hot and controversial topic in the US is because of the open-source models in China, and they have become so capable. Kimi K3 kind of proved that. But even before that, there was the DeepSeek moment last year. Then GLM-5.2 was released, and I think the US labs were alarmed at how good the Chinese models had become.

6. Are Chinese open-source models actually changing Silicon Valley?

This is when we went back to your purpose of coming to the Bay Area this time around. You wanted to see the impact of open source—I guess, the industry—on applications and the entire agentic AI ecosystem. What have you seen so far? How has open source changed AI?

Xing Meng

The interesting observation I had is that, despite the heated debate and the great attention that has been placed on open source, Kimi K2, and so forth, if you’re talking to frontier-lab researchers and so forth, they haven’t cared much about this. Few of them have actually tried those models, and that was a surprise to me. I thought this would be your direct competition, not just from a competition standpoint, but also from a technical-report and paper-sharing perspective. You should be learning from that, but it gets less of their attention.

From an agent-developer standpoint, a lot of them, especially as seen in the news, have switched to Kimi K3. As a result, what I heard was that this has had a significant impact on the top-line revenue of the frontier labs, especially Anthropic.

We’re at the point where total token consumption in July has been growing slower than expected. Some would say it’s stagnated. If you look closer, that’s only consumption. If you look at token revenue, a lot of people are swapping from the frontier models to the cheaper models. So, if total consumption is flat—let’s say, I don’t know if it is, but let’s assume it is—total revenue will be going down because people are swapping more expensive models for cheaper models.

There’s a large impact, and the reason behind that is that you’re seeing agent companies, users, and individual coders making that change as well. There are innovative ways to do that. You might not swap all the way: you might use the likes of Kimi K3 or DeepSeek to write your code and then still use Claude to review your code, using a combination of them to maintain a certain level of accuracy and performance along the way.

For most people, you can hardly tell the difference from that perspective, unless you’re a very intensive user or you’re using it for highly complicated problems. But if you look back, even among the Anthropic models—between Fable 5, Opus 5, and Opus 4.8, sometimes even Opus 4.6—people can’t decide which one is better for their particular task. Sometimes the 4.6 is better. Sometimes Fable is better, which is supposed to be because it’s a larger model, uses more training cost, and came out later.

So, you’re seeing that all of them are more expensive than the cheaper models. Kimi K3 is actually not that cheap.

Speaker 1

Compared to DeepSeek, which is usually the—

Xing Meng

The cheapest option, but still cheaper, obviously, than the other models of the same size and so forth. So this is happening and hitting the token economy a lot.

I would say the agent companies really care about this, but the top frontier models care less. As a result, this discussion hasn’t been that helpful so far. You’re only seeing migration, with adoption rates going up. That’s sort of expected and what we learned, but there’s less discussion of what we can learn from those models and what we should take from them. The researchers here are commenting on the models less because they haven’t spent much time on them.

Yeah.

Speaker 1

7. Are frontier AI models becoming commoditized?

So, are we at a stage where it has become kind of hard to differentiate between the different models for a normal company or regular users? If so, what are the implications for the frontier labs in both the US and China?

Xing Meng

I think if your main task is not to build the next frontier model, build the next great database, build complicated coding systems, or solve the hardest math problem on Earth, then for the most part, I don’t think you’ll be able to tell the difference among the top models that easily.

The penetration of coding agents among coders has been really, really high. It would be very hard to find a coder who doesn’t use a coding agent today.

Speaker 1

But they hope to bring non-coders up in terms of using the likes of Claude Cowork to solve, I guess, their PowerPoint or report problems.

Xing Meng

The penetration is not that high, and the growth is not as high as expected. It’s partially due to the fact that there are 2 types of adoption. There’s adoption from the company, which is pushing from the top down. You have a business deal with your company, the company adopts it, and this pushes every employee to use it. Those are very slow—slower than expected.

Usually, their AI strategy lasts 2 to 3 years, and it takes them 2 to 3 years to fully get compliance, security, and privacy right. That’s just too long. So you’re not seeing that adoption going really fast.

On the individual level, first of all, with a lot of those tasks, you can’t really tell the difference in terms of capability among those models. Second, a lot of tasks—as you know, probably, Fable and a lot of those top models—the advantage of those is that they can run for a very long time.

If you set the right goals, set the right evaluations, and let it run, it’ll be able to figure out a complex problem while you sleep, right? The next morning you get up, and it’ll be done.

But for a lot of white-collar work, if you’re writing a PowerPoint or writing a report, it’s very hard to predefine the goal. If you make this color right in this way, then this is a great PowerPoint. It’s very hard to set this before you see the PowerPoint.

Humans are always in the loop in that way. You have to be in front of your computer, and even if the work is done by Codex or Claude, you still have to be in front of a computer to check the result every time.

And then when you see the result, you look at it, think about it, brainstorm, and probably set up some new goals. That means the token throughput cannot be very high, because it cannot run continuously for a very long time. You always have to be in the loop to do some work, check, and enhance the result.

The reason why I mentioned that I think token throughput should be a very important issue here is that if you build a PowerPoint in 10 seconds rather than 10 minutes, you can have more iterations of that loop for non-coding work and bring up the total token consumption in that area. But this is not yet happening. So far, you're seeing stagnation in those areas beyond coding.

The third area where token consumption can go up will be multimodal, multimedia models. You're seeing Seedance obviously going up, and we talk a lot about LLMs, but ByteDance has built this really powerful tool, Seedance, which is the best sort of multimodal, multimedia model coming out and making a lot of revenue—probably the best in the world by far. My personal belief is that it's the most promising tool for the next level of token consumption. But it has to be improved and become more interactive. The serving speed has to be faster for this to work.

Speaker 1

8. Why China is leading in AI video generation

Yeah. Let's talk about multimodal models and video in particular, because you mentioned Seedance. I actually thought this was a great segue into the applications layer of AI in the US and China, because video models are one of the things where we're seeing a sharp contrast between the US and China. In China, we had Seedance and all of those startups working on video generation. In the US, those attempts have kind of faded over the past year. OpenAI kind of stopped the Sora 2 release, and I thought that was an interesting comparison. Why are the Chinese labs good at making video models?

Xing Meng

I think, first of all, there's the resource-allocation part, and there's also the priority part, right? Coding has been so important that everything else should be less important, following that narrative. This narrative has been more and more central to the scene since, I guess, January of this year, the early part.

So everything has been cut. Everything that consumes compute is cut to preserve enough compute to make sure that coding is as good as possible. This is not to say that this is not happening in China as well. I think it's also happening there.

But if you travel back in time and say you're in 2023, and you would have to guess who would have the best multimodal model in the world 3 years later, the easy guess would be ByteDance, Kuaishou, and Meta. They have the most training data, the most multimedia platforms, the best distribution platforms, and enough GPUs and so forth. They have the data, they know what the product should train toward, and they have the distribution channel for that.

But it seems that Kling, which is the model for Kuaishou, worked out and was number 1, and then was surpassed by Seedance from ByteDance. Those 2 were expected in that way.

ByteDance's Seedance was really great because they took a very aggressive but risky shot, which was to train a much larger-scale model than all the multimedia models before that. I think they scaled it up by 1 order of magnitude, and they did it when there was very little evidence that it would work.

Looking at it retrospectively, it seems like the obvious choice, and everybody's doing it now. But back then, I think half a year ago or almost a year ago, it wasn't that obvious. They took the reward for taking that risk, I think.

Speaker 1

Why was it a difficult bet back then?

Xing Meng

Yeah, because most of the models are relatively small, and you use thousands of cards or GPUs for training compute rather than tens of thousands for that type of model. Being able to pull those resources out from training LLMs, and coding especially this year, is a more and more difficult decision to make, especially when you're resource-constrained.

So, yeah, that was not simple. I think that decision was led by a very, very young female researcher at ByteDance, and she deserves the credit for building this. Kudos to her. I think that's a very great move.

But what is disappointing is Meta. In a way, I suppose they should make a really great model. They have the data, the capability, and the resources to do that, and it is in their area of expertise. I think it's something they should do really well, but so far I haven't seen the progress yet.

Speaker 1

Yeah. I guess they did make Dream, and then Muse came out a couple of months ago.

Xing Meng

Yeah.

Speaker 1

Would you also say there's consumer demand for this in China? Because in the beginning we talked about consumer-facing applications. MiniMax was doing that—obviously, it was a lot more companion work—but in China it seems like there are a lot of consumer-facing AI applications, and Seedance has obviously been used by everyday consumers. Everyone is making TikTok videos about it. AI dramas and short dramas are very popular on social media. Is that also what constitutes a different market in China versus the US?

Xing Meng

I think Seedance's adoption goes up, and especially revenue goes up, but my view is that it's not yet mainly contributed by individual consumer users. It's mainly by prosumer users.

Speaker 1

The reason you can say that is because you look at the applications that are built for prosumers: their revenue went up significantly after the release of Seedance. What are some examples?

Xing Meng

Topview and LiblibAI, for example. Those are the ones. The 2 products are similar, but essentially they're canvases for building videos, and advertisers or short-video builders use the canvas to set up a workflow to generate certain videos.

At the beginning of the year, those companies were looking at revenue in the small millions or the mid-single digits. After Seedance's release, they quickly ramped up to almost 100 million or more. That's super-fast growth.

If you look at the deals they're making, it was even hard to get Seedance tokens at some point. You had to prepay for a certain amount in order to allocate them. Those are the buyers of Seedance's API, and the users behind that are prosumers who are making videos for AI short films and advertisements and all that stuff. As an individual, I don't think spending has gone up that high yet.

Speaker 1

But the short-film market—the short-video market—has been going crazy, and short films have, I think, surpassed movie theaters. I think they already surpassed movie theaters this year.

Xing Meng

Oh, this includes, I guess, both human-acted and AI short films, right?

Speaker 1

Yes. Yes.

Xing Meng

That makes sense.

Xing Meng

But then, because of short films, in the past we were looking at this as an investor. We looked at a few companies doing that. My view was that AI-generated short films are great at science-fiction themes or traditional Chinese sci-fi, or ancient swordsmen, because they're really great at creating special effects.

They're not very good at creating detailed human emotional reactions. So if you're filming a high school drama with a lot of dating and small, intricate microexpressions, it's not very good for that. In the past, they were good at certain genres in Chinese film, which was good enough because those genres were the most popular to begin with.

But now, after Seedance, it's already very good at all of that. A lot of the Seedance demos are focused on 2 types of content. One is making very intricate emotional expressions: you have to laugh and then cry, with a very smooth transition across that. Sometimes it's better than humans, and sometimes better than actors.

The second is that Seedance specifically trained and spent a lot of effort making sure that highly dynamic motions are captured fairly accurately. For example, if you're engaged in a kung fu fight between 2 actors, in the past, sometimes if you hit someone, the hand would go through their arm and to the other side, or it would just look unnatural in some way.

It's very hard because it's happening very fast. The frame rate is high, and the action is compressed into a few frames. Sometimes it blurs out, and sometimes it's wrong. So I think the Seedance 2.0 version has spent a lot of effort improving that, which is one of the most important features for AI short films.

As a result, it's become almost the exclusive provider for a lot of short films in that way.

Speaker 1

Yeah. Yeah. I actually thought it was interesting because in one of your previous interviews, you mentioned that one of the applications you look forward to is this middle ground between video and games—this interactive form of media.

Would you say that AI-generated videos are kind of a step toward that? One example you gave in that podcast was Fable Studio’s remake of South Park with AI. Are you seeing more examples like that with the advancing capabilities of AI video?

Xing Meng

I have to say that I’m a little disappointed. Since I made that interview, I still had high hopes that it would become a mainstream type of media, but I don’t think it has played out as I expected so far.

I think naturally there should be something in between watching TV, which is leaning back and doing nothing but allowing a minimum of interaction, and playing games, which is leaning in and doing intense interaction. There should be something in between with medium interaction but endless content—a type of generated content. So it’s sort of between games and media in that way. I don’t think we’re there yet.

I think there’s this thing called Yingyou, which is sort of half movie, half game. You get to play; there’s a theme, and there’s a narrative.

Speaker 1

I think it’s sort of a recreation of the famous game Detroit: Become Human, which is a video game.

Xing Meng

Probably 10 years ago or so. There’s a lot of that content coming out, but I don’t think it’s of high enough quality yet to become a very popular genre. There are single hits that are making a lot of money, but it’s not taking off yet or becoming as popular as AI short films in that way. Yeah.

Speaker 1

Yeah. But you still have hopes for that?

Xing Meng

Yes. I would love to see that happen. I think this will be bigger than games themselves.

Speaker 1

Nice. Yeah, I guess world models would be kind of a form of that, which we’ll get into, because we’re going to talk about new labs. That’s one of the verticals I really wanted to get into more, because in the US it has become pretty popular to invest in labs founded by young researchers, or often by experienced researchers who spend a lot of their time in an academic environment.

In China, you mentioned that you have been looking at a lot of AI-for-science companies and startups as well. Tell us a little bit more about that. What does funding for new labs look like in China compared to what you’ve done in the US?

9. Neo Labs, AI agents and the future of model companies

Xing Meng

In the US, new labs have been the favorite of a lot of venture capital over the past year or so. There’s obviously much more funding for this to happen in the US than in China. I think China is more practical. There are new labs coming out. For example, Lin Junyang, the head of Tencent’s Hunyuan model, founded his lab. A famous researcher, Dai Cong, founded his lab.

Overall, my view is that the concept of new labs has changed, or is changing, along the way. In the past, new labs exclusively meant what I mentioned and defined just a moment ago. But now, when we think of a new lab—or what a new lab should look like—I’m thinking more like Cursor.

They’re an application-focused company to begin with. They built a super app, gathered user data, had user traces, and decided to train a model. They were able to do that precisely because they had a super app and user traces, and they were able to build a great model, which is Composer, by just doing post-training rather than pretraining. That model was able to beat a lot of the pretrained models out there in the world.

I think that’s the right recipe for a new lab these days. It’s not a new methodology; it’s that you have a new set of data that the frontier labs don’t have, and you’ll be able to capitalize on that data to build something different.

When I think about new labs, I’m thinking about the most popular apps out there today. At some point, they will decide to train their own model, either through post-training or maybe someday even pretraining. They will have to do that because they’ll want to, number 1, lower their costs, and number 2, possibly build something customized for their particular application.

You’re seeing this happen at Harvey and some of the vertical application agents here in the US, and I think you’ll be seeing this happen in China as well.

Speaker 1

10. The next frontier of AI training data

What kind of companies would you say are most suitable for doing that? You said it was a company that has the user data and has an app.

Xing Meng

Number 1, you have to have enough user data that is untouchable, that is unique to you. Number 2, the data is better suited if you’re in a vertical rather than a general market, because then you’re not only training a model to get a cheaper model and lower your costs, but you also actually have a shot at training a model that is unique.

If your application is general, then it’s hard to train a unique model because, essentially, users use general agents or search engines for general purposes. Their prompts are similar to what Claude and OpenAI are trained for. But if you’re looking at just a legal vertical or a medical vertical, then you’re solving a particular problem. Potentially, you can tune a model specifically for that purpose, and it’s easier to build a model that beats the general models in that way, on top of your unique data.

So I think the combination is a vertical application plus a super app. The question behind this—and I probably want to touch on this a little bit—is: What should a model be trained toward? What is the goal? What’s the evaluation for that?

Today, when you look at frontier models, everybody knows that the key to improving them is data. But who decides what data to use and purchase? It’s the researchers who own the training of the models. A lot of the data they use is attributed to the key researchers’ own tastes.

There’s some bias toward that because, if you’re a researcher, obviously you’re very familiar with coding. You’re familiar with probably math as well, but you’re probably not going to be very familiar with office work, because you personally have never done it—especially if you’re a PhD just out of your PhD program, directly entering a frontier lab and being put into a key role. You probably have never done much PowerPoint work, report work, or accounting work.

But then they’re acquiring human-labeled data and new environments to train the models on. The true environment that the model should be measured against in the future should be the exact environment that people are using, right? The Slack, Zoom, SaaS, and Salesforce software that people are using. There’s some goal in there, and then people make some effort, doing trajectories to solve certain processes within that environment.

Cursor is one of those examples. It’s one of the most popular pieces of SaaS software that people use. It just so happens that it’s coding software, so this happened really early because it’s something that researchers are very familiar with. But in the future, I think those will be the true environments—those working environments that truly measure how good the model can work. It’s not some imaginary environment imagined by the researchers or according to their research tastes.

Speaker 1

That’s interesting, because I feel like a lot of discussions before happened around how, when you were developing foundation models, we had already used the world’s data. We had really drained the internet, and that’s why people were looking into synthetic data. But it seems like the future of data lies in those specific verticals. Would you say that these are still unmined territories, and that there’s still a lot of data when it comes to specific use cases and verticals?

Xing Meng

Absolutely. I think there’s a long way to go. In the past, data was static. You gathered it by crawling the internet or synthesizing certain data, but those were static data.

Today, the most popular data are environments. You don’t build a piece of data; you build an environment, like SaaS software, for example. You can let your model attempt a bunch of tasks within that environment, and it will try on its own. It will do reinforcement learning along the way: once it’s successful, it will be rewarded; if it’s unsuccessful, it won’t be rewarded, and it will learn on its own.

If we’re talking about environments, then there are a lot of environments to be built. We start with the most static environments. For example, if you want to build something called kernel optimization, which is very popular in reinforcement learning today, you try to write certain CUDA operators so that your model runs faster. But the environment is set. You can rerun this a million times, and it won’t change, and you’ll be able to roll out the test equally and stably. So that’s the easiest.

Next is called partial observation. What if you can’t tell what all the parameters of the environment are? For example, if you’re in Slack, if you’re a new employee and you join a new company and you’re in the employee group chat, you’re trying to figure out what the boss wants. You don’t know what the boss wants.

You have to observe, and then gradually, over time, you probably know, because you're only given a partial view. Somebody probably tells you what the boss wants, but it's not specified; that information is untold. Over time, you learn this, maybe by exploring, interacting with other people, and so forth. So this is called partial-observation optimization, and that's sort of true of most environments. You obviously don't know the intricacy of everything to begin with.

And then there's static versus dynamic. For example, in a finance or trading environment, every action you take changes the environment itself. It will never be able to roll back to the original environment again. So every test you do changes it, and then, if it's successful, it might not be successful in the next step because the environment changed over time.

How do you adapt to that? Those are sort of new environments that you have to build, as this is new data, and you have to create them. Social media is that type of thing. Every post you make changes everybody around you. It changes your fans' perception of you, so you cannot do the same experiment 1,000 times and just see how it works, because the accumulation of the previous 999 times is already in the mindset of your fans, and it changes their behavior.

Speaker 1

And from the virtual world to the physical world is also a direction where you can build. For example, I think there's one benchmark I love. There's one environment, or sort of data, that I love to share. It's called Vending-Bench. I don't know if you've heard of it.

Xing Meng

We just went to their office right before this. They created a benchmark that is virtually setting up 3 different vending machines and asking an LLM to run those vending machines. The LLM can decide who to order the products from, set the price, decide on the marketing strategy, and so forth, just to see how it runs them.

That's an attempt to go from a virtual world to the real world. Although the benchmark was set up for virtual vending machines, they sort of created this in the minds of the LLMs—not to really set it up—but then they set up a market, I think at the Anton market close to here, which is really a shop, to let an LLM run the shop.

The goal is to test whether, if your LLM is so good at reasoning and doing all that work, it can run a shop in the physical world, where so many unpredictable things could happen, and still run it as well as you think it would. They change different models to do the same thing. Those data obviously aren't on the internet and cannot be crawled beforehand; they have to be made on the spot. There's a lot that can be done.

Speaker 1

11. How US and China approaches data differently

Speaking of data, there are also a lot of data providers and data-labeling companies, with Scale AI being one of them in Silicon Valley. But there's still a lot of demand for data. I know that you've visited some of those companies as well. Tell me a little bit about the state of those companies in the US, and how different they are from the Chinese data providers.

Xing Meng

I think we're seeing a boom in revenue over the past 6 months. I know they're making a lot of revenue already in the past year, but the past few months have seen a significant increase for a lot of them. Many of them have hit $2–3 billion in ARR at that scale. Some of the Chinese companies have been catching up—not to this scale yet, but very fast as well in terms of growth.

If you dive deep into those companies, they're different types of companies, although they all claim to be AI and LLM data providers. Some of them, like Scale and Mercor, sell their data in terms of headcount. They sort of price according to the old model: how many people they recruited to do how many hours, like a consulting or lawyer-type model.

Some of the new labs price based on environments. One environment is sold at a price ranging from a few hundred dollars to maybe $100,000, depending on how difficult it is. If it's a small Booking.com environment—trying to book a ticket—then it was probably $1,000 or less.

If you're trying to build an environment where you want to test the agent's capability—say, the DevOps capability—to turn on a cluster, get a cluster somewhere else, merge those 2 clusters, and do some crazy work in terms of DevOps, then you sort of have to recreate almost the entire AWS experience. That will probably cost close to half a million dollars. They're pricing based on that, and those are sort of new lab business models.

For the headcount models, they're largely relationship-driven, because today, as we talked about a little bit, a lot of the data is decided by researchers and their tastes. It's somewhat a traditional B2B sales problem, whereas you have to make great friends with the researchers who can make the decisions. There's a lot of wining and dining in that process to make sure it happens.

A lot of the traditional data providers are building their businesses off the typical business-development and sales model. A lot of new labs are building this off the fact that the founders of the data providers are researchers themselves. They could have gotten really high-paying jobs at a frontier lab, but they decided to become founder-led data providers.

Their unique trait is anticipating the needs of researchers in the labs 3 months ahead of time. What type of benchmark do you want to rank high on? We'll create that benchmark for you. What type of data would you need? We'll create that data for you. We'll sell you a combination of the benchmark, make it popular, and everybody will want to rank high on that benchmark. Then we'll sell you the coding data that will help you get really high on that benchmark.

That's a different way of competing. Can you be an academic opinion leader and push up your benchmark? If so, then you'll rightfully be able to sell all the data or code that will help with that benchmark.

Those are 2 different models: the prior one, where human labelers are priced based on headcount, and the latter one, which anticipates what researchers need and builds data specifically for that. I think the revenue is higher for the former, but the growth in revenue is higher for the latter, in terms of anticipating what they need and building data specifically for that. So that's really interesting.

The Chinese data providers are similar in that they started by building human-labeling businesses. In the past, one of the key benchmarks built by Scale AI was called HLE, or Humanity's Last Exam, and that was mostly a high-end expert-labeling effort.

Speaker 1

People started with that. I think almost every data company started with that. Now they're transforming into a lot of environment builders and benchmark builders, so I think that's a trend.

But despite which side you are on, I think we're seeing revenue going crazy. Even at ICML, the academic conference, last month in—

Xing Meng

South Korea. Yeah, South Korea.

Speaker 1

Korea. We're seeing a lot of young researchers, and we asked them, “What's your number-one aspiration? If you want to get a job in academia, or do you want to build a startup?”

“Definitely want to build a startup, but we want to sell data. We want to be data providers.”

Xing Meng

Oh, wow. Yeah, a lot of those aren't building model companies. They're not building agent companies; they're building data companies. So, yeah.

Speaker 1

Let's move on to AI for science, which seems like another thing that you've been looking at during this trip. What kind of companies have you been chatting with when it comes to building AI for science?

Xing Meng

Mostly. We're talking a lot about 2 groups. One is frontier labs, which usually have a division looking at AI for science. The other is a lot of fresh PhDs who are scientists themselves or are at the intersection between computer science and, say, biostatistics or biology—I don't know the name for that—people who work on genes, or people researching physics and materials science, and so forth.

I think they're mainly coming from these 2 groups. The third group is professors who specialize in certain fields and are looking to get some help moving into AI for science. The barrier to entering AI for science is lower nowadays with certain agents and so forth, so they're building their own capabilities in that area.

So, frontier-lab researchers who specialize in this area, PhDs at the intersection between computer science and those fields, and then some professors who are in the sciences.

Speaker 1

12. AI for science

What's the biggest challenge in building AI for science?

Xing Meng

I think it depends on what you're building. There's a lot of things thrown into this bucket called AI for science nowadays. A lot of them are building AI scientist agents, which are essentially agents that can read literature, design experiments, sometimes control lab equipment, analyze results, and sort of do that scientific discovery loop, and maybe, at the end, write papers and even publish papers.

That’s one area that I think isn’t that hard to build these days. Every lab is building its own AI research scientist nowadays. It’s just hard to build something that’s really good and meaningful in that way, essentially measured by whether you’re able to publish a paper that’s well received, right?

It’s not that easy to do that yet because there are a few steps. You have to analyze the literature, plug into enough databases, and plug into a certain number of tools. But the backbone models you’re using are still OpenAI and Anthropic models, and usually they’re not trained specifically on those types of scientific databases and tools. They’re not really good at using those things, and sometimes you have to build certain workflows or guardrails to make sure they do the right thing, rather than just let the model freestyle and run its own course. So that’s one of the difficulties.

Second, a lot of these experiments are done in wet labs, meaning that they have to have lab equipment and lab scientists to do the experiment, get the data back, and verify it. It’s not like math or coding, where the whole environment and all the verifiers are online, accurate, and stable. So it usually takes a long time, or sometimes it’s just disconnected. Certain agents can go as far as experiment design, but executing the experiment needs a human to intervene in that way.

There are companies like BioMap and Lila Sciences, and they’re all doing very well and building useful tools for researchers. There’s also the other side, where you’re building a model specifically for that. Since AlphaFold, DeepMind has obviously been building great models out there. Since AlphaFold 3, a lot of companies have been coming out and building frontier, state-of-the-art models, like Chai, like Boltz, and like DeepMind’s own Isomorphic Labs. We’ve invested in companies along the way that are building protein-design models as well as complex-structure-prediction models.

These models are trained on public and private datasets to be able to design things, and nowadays, de novo design—meaning design from nothing. You’re just creating something that doesn’t exist in nature, such as antibodies or proteins, and enabling them to bind to certain known targets so they can solve certain diseases.

We’re also looking to build world models for cells—virtual cells. This is essentially the flip side of the protein-design models: you can design a protein binder, but how would you test whether it can bind well or not? You have to design a cell in order to see the interaction that way. Otherwise, you have to do this experiment in a real lab, which is very costly and time-consuming.

That’s also very complicated because, unlike robotics or the real world, where you have cameras and all those sensors, for a cell you have very, very minimal data to build anything or sustain a large model to train on. So the data shortage is probably the biggest problem here. Multimodality is also a problem: enzymes, antibodies, antigens, and peptides are very different things. How would you encode them all into the same model and be able to represent them together? Otherwise, you have to build small models for each, and then that defeats the purpose of building one unified large language model, or a large model of any sort, to be able to do the prediction.

Those are some of the problems we’re seeing today. Essentially, all the things we’re doing in biology—I gave you a lot of examples in biology—but the true verifier is whether this drug works on a human body and cures the disease, right? We can have intermediate verifiers today: whether it binds well, or whether you do it in a wet lab. But the true verifier is whether this drug works on a human body and cures the disease. You won’t even be able to do that experiment unless you’re 5 or 6 years into the process, until you get to a clinical trial approved by the FDA. You don’t get any real, full feedback from a clinical trial until you’re a few years into this process. You can’t imagine that happening with an LLM or a physical-world model, but this is a reality for drug-design models today.

Speaker 1

Yeah, very interesting. A really high-stakes drug-discovery problem with a lot to expect there.

Xing Meng

Yeah.

Speaker 1

Great. So we’ve talked a lot about foundation models, the difference between the AI ecosystems in the U.S. and China, and different applications. My last question is: what would you say is the 1 biggest misconception that Silicon Valley has about China’s AI ecosystem, China’s AI development, and tech in general?

13. Silicon Valley’s biggest misconception about China’s AI ecosystem

Xing Meng

I won’t attempt to say that I understand the whole perception of the U.S. ecosystem of China. But I think, in my opinion, the U.S. ecosystem thinks of China’s AI ecosystem, the labs, or the technology as the result of improvement that’s entirely tied to government support, either through funding, compute, or all that stuff. They give a lot of credit to how the government supports this.

I think they haven’t given enough credit to the individuals, the entrepreneurs, and the researchers who are actually making this happen under very difficult constraints in the world. I think they would be able to do this anywhere in the world if they’re able to do this in China today. So that’s probably the biggest misconception.

Speaker 1

Yeah, totally. Great. I think that’s all the time that we have today. Thank you so much for joining us. Any last words?

Xing Meng

No, thank you. Thank you. Really glad to be here. Thank you.