Nathan Labenz
Hello and welcome back to The Cognitive Revolution. Today, I'm honored to share a special cross-post from the China Talk podcast, hosted by Jordan Schneider, China Talk analyst Irene Jiang, and Nathan Lambert of AI2 and the Interconnects Substack, featuring a conversation with Zixuan Li, director of product and generative AI strategy at Z.AI, also known as Zhipu AI, about the culture, incentives, and constraints shaping Chinese AI development.
I imagine that many, even in our AI-obsessed audience, will not be familiar with Z.AI, but their model releases demonstrate that they are a significant player worthy of our attention. As of today, their latest GLM-4.6 model holds the number 19 spot on the LMArena text leaderboard. Its Elo rating is roughly 65 points behind the current leaders, which means that it still wins 2 out of 5 head-to-head comparisons with the leaders, and it happens to sit right next to Qwen 3 Max, Kimi K2 Thinking, and DeepSeek V3.2.
Together, these 4 models, all from China, are indeed the top 4 open-weight models available today, though it should be noted that Mistral is not too far behind. On the WebDev Arena leaderboard, GLM-4.6 does even better, coming in at number 9, making it competitive with GPT-5.1 and meaningfully behind only Gemini 3 and the new Claude Opus 4.5.
All that said, this conversation goes way beyond benchmarks and touches on a number of important topics, including why we should understand Chinese companies' open-weight strategy not so much as an ideological commitment, but as a practical marketing tactic used by companies that are seeking to gain global mindshare while recognizing that Western enterprises simply can't and won't use their APIs. We discuss culturally distinct AI use cases, such as role-play, that matter in China and how these drive different fine-tuning priorities than what we typically see from Western companies.
We also discuss the role that Silicon Valley thought leaders play in establishing credibility for Chinese companies, even in their home market; the market for AI talent in China and why very few people at Z.AI even know that Zixuan studied at MIT; Zixuan's view that there is a wall and that further architectural breakthroughs will be needed; the extreme velocity with which Z.AI releases models, which often involves shipping within hours of completing training; what, if anything, people in China fear about the AI future; and how Chinese companies generally still view themselves not as peers or rivals to leading American companies, but as upstarts that are just trying to keep pace and will be quite happy if they can secure for themselves a meaningful niche in the global AI marketplace.
This is a fascinating, detail-rich conversation, and regardless of your attitude on U.S.-China competition, it's clear to me that we in the West don't hear nearly enough from the actual builders inside Chinese AI labs. So, I appreciate Jordan for allowing me to cross-post this episode, and of course, I encourage everyone to subscribe to China Talk, online at chinatalk.media. With that, I hope you find as much value as I did in this behind-the-scenes look at frontier AI development in China, with Zixuan Li of Z.AI from the China Talk podcast.
Speaker 1
Zixuan Li, who studied in the U.S. before moving back to China, works at Zhipu AI, or Z.AI. We're going to let him introduce himself and his role. Co-hosting today are Irene, longtime China Talk analyst, and Nathan Lambert of AI2 and the Interconnects Substack. Welcome to China Talk, everyone.
Speaker 2
Personally, I've known of Zhipu AI, or Z.AI, for at least a year. Then, in December, there was kind of a mind meld, and I was like, "Well, there's another DeepSeek moment," when they released GLM-4.5. You'll have to correct me on whether or not it was before or after the Kimi K2 model.
Zixuan Li
After.
Speaker 2
I guess it was after, and that was a matter of weeks or days after Kimi K2. It's just like, "Wow, okay, there are a lot of people building great models." It's fun to get to learn about some of them and how it compares between U.S. labs and Chinese labs, and I think a lot of it is more in common than different.
Zixuan Li
Hi, everyone. I'm Zixuan Li from Z.AI, and I manage a lot of things, like global partnerships, Z.AI Chat, model evaluation, and our API services. If you have heard of the GLM coding plan, I'm actually in charge of this thing, too. Nice to meet you, everyone.
Speaker 1
Thank you for introducing yourself. I'd love to hear more about how you ended up working in AI after you moved back to China, and why AI specifically.
Zixuan Li
Actually, I applied for multiple roles at companies like Moonshot and MiniMax, but got rejected or neglected because there are so many resumes going to them every day. I studied AI for science and AI safety at MIT, so I did a lot of research on AI for applications and AI alignment.
That's not very relevant to what we are doing right now, but it actually gives me a sense of what's going on in the frontier area. So it helps me a lot in understanding what OpenAI and Anthropic were doing at that time. I think it's very innovative to have this sort of idea and experience.
Speaker 1
Was it ever a debate for you whether to stay in the U.S., or did you always know you were going back to China after grad school?
Zixuan Li
I already knew I was going back because my family is here. But I got the job after graduation, because it's hard to get a job. I continuously applied for jobs, but finally, after 1 month, I got the opportunity to interview with Z.AI.
At that time, I wasn't in charge of the overseas department because our focus was on the domestic area. It was a domestic chatbot, so I was responsible for the strategy of developing a domestic chatbot. It's called ChatGLM.
Speaker 1
Got it. So maybe let's do a little bit of Zhipu AI backstory. When was it founded, and how would you place it within the broader landscape of teams developing models in China?
Zixuan Li
Great. Zhipu AI, and also Z.AI, was founded in 2019. We were chasing AGI at that time, but not with LLMs, with some graph networks or graph computing. We did something like Google Scholar called AMiner.
We used that type of thing to connect all the data resources from journals and research papers into a database, and people could easily search for and map scholars and their contributions. So it was very popular at that time.
But we shifted to exploring large language models in 2020, and we launched our paper, GLM, in 2021. So that's, I believe, 1 year ahead of the launch of GPT-3.5. It was a very early stage, and we were one of the first companies to explore large language models.
After that, we continuously improved the performance of our models and tried new architectures. GLM is a new architecture, actually, but we are going to explore more in the future. I believe that we got famous by the launch of GLM-4.5 and GLM-4.6, because I think they are very capable in coding, reasoning, and agentic tool use.
That's more useful compared to the previous version. People may know us through Claude Code, Kilo Code, and other tools. So we need to combine with these coding products, and that got us famous.
Speaker 1
Let's talk a little bit more about the evolution of GLM-4.5. I don't know, Nathan, this is your question. Why am I asking this question?
Speaker 2
Well, let's think. What does it take to transition from the models that you were early with to things that get international recognition? I have known of Z.AI and your work for years, and then it's like, snap of the fingers, and now this model is on everybody's radar who's paying attention.
Does this feel like something that was just going to happen for you overnight in developing the models? What does that feel like when you go through it? How do you get to that moment? Because there are a lot of people who want to do that at their companies.
Zixuan Li
That's a very interesting point, because in 2024, everyone was interested in Chatbot Arena, right? We saw GPT-4, and we saw Gemini performing very well on Chatbot Arena, so that was our interest, because we paid attention to end users' experience. When there were 2 answers, which one did you prefer?
We did a lot of things on that, and we performed very well on Chatbot Arena, ranking maybe 6th, or between 6th and 9th.
But in 2025, with the launch of Manus and Claude Code, we realized that coding and agentic stuff are more useful, or they can contribute more economically and also in terms of efficiency for people. So I think chat is no longer our top priority.
Instead, we do more exploration on the coding side and the agent side. We observe the trend, and we do a lot of experiments on it. We need to follow the trend and also predict the future.
Speaker 2
Do you feel like you're better at executing code versus Chatbot Arena? Because I like GLM-4.5, and I think GLM-4.5-Air and GLM-4.6 are extremely renowned for this. I think that when you train these models, the process can look very similar depending on what your target is.
I'm just wondering if it was a shift or if it's just that sometimes things work out better than others.
Zixuan Li
It’s a shift, actually. We pay more attention to the coding stuff. On Z.ai Chat, it’s free, right? Nobody’s paying for the chat.
People pay for cloud API use and for agentic stuff, but we just let users chat with the chatbot freely. So, that’s a shift. But we need to continuously improve performance in normal chat and maybe role-playing, but that’s not our top priority.
Nathan Labenz
Let’s talk a little bit about the talent and the internal culture that allowed you to put out GLM-4.5. What do you think is different about Zhipu AI, or what distinguishes Zhipu AI from other labs, both in the US and China?
Zixuan Li
First of all, we’re more collaborative inside the company. Everyone is working toward this single goal. Maybe we have separate team heads, like a pre-training team and a post-training team, but they’re working very closely. They just sit next to each other, working toward a single goal: trying to build a unified reasoning, agentic, and coding model.
We’ve built 3 separate models, as we illustrated in our tech report. We then distilled these 3 teacher models into one single model, GLM-4.5. That’s our goal, and that’s how I believe we built GLM-4.5 more efficiently compared to other companies.
And we’re super young, right? Another point is that, as you mentioned, talent is important. I believe that nowadays you need to do the research yourself. You need to do the training yourself as the head of the team, so you cannot let others do this stuff for you.
Nathan Labenz
Why is that?
Zixuan Li
Because things change really fast. Maybe during your training, there comes GLM-4; there comes GPT-5—anything can happen. You need to feel the trend yourself. You need to combine the results from experiments, the trends, and what’s going on within your competitors’ teams to feel the move yourself.
It’s super important. Even our founder did the experiments himself. He looked at the papers. You need to do things simultaneously, not just set goals for people and let others do the stuff for you.
Nathan Labenz
Yeah, it’s very fast-paced. I think before we started recording, you were also mentioning that there are a lot of PhD students involved. Are these people actively pursuing their PhDs, kind of new grads, or a mix of all of them?
I work at a research institute that is very open-source, and we have a lot of full-time students who are part of it. But when you look at other closed labs in the US, there’s not nearly as much intermingling with academic institutions. I think that could be a really powerful thing if you have this, because there’s a lot of experimental stuff going on there. Do you feel like it’s a kind of open door between some academic institutions and your work?
Zixuan Li
Definitely. There are a lot of PhD students currently here. I believe they’re both pursuing their academic work and working on GLM simultaneously. But they can combine them together, right?
If you’re doing a really innovative job, like training a unified agentic coding model, it’s one of your greatest achievements ever. People won’t say, “Okay, I need to do another research project. Let me finish this first, and then we’ll go back to GLM.” They’ll try to treat GLM as their single biggest achievement.
Everyone is really devoted to this stuff. We hardly see anyone not devoting themselves to training GLM.
Nathan Labenz
Could you talk a little more broadly about the talent market? You mentioned earlier that you had to put your résumé in a lot of places. What does it look like right now? What’s the kind of hierarchy, and what are folks looking for? What are employers looking for, and what is the talent looking for?
Zixuan Li
On the research and engineering side, I think they’re looking for papers, GitHub code, and competition experience. They’re also looking for your experience using GPUs and your experience training models.
For the non-technical side, they’re looking for how you’re going to grow model performance, expand your branding, and a lot of other things. If you’re going to be a product manager, they’re looking for your coding skills, your vision in this area, and how you do the stuff yourself. Those are very important.
I think it’s pretty similar, but you mentioned hierarchy. In terms of hierarchy, large companies choose the people first because they have more money. They can pay more, like Baidu and Alibaba. But for startups, we need people to fight together. You need to fight against other competitors. You need to drive yourself to finish the goals because you don’t get paid that much.
You need ambition. You truly need to enjoy working with really young, talented people and trying to build something like GLM. It seems to come from nowhere and tries to beat other competitors’ models.
Nathan Labenz
How big would you say the team is—the number of people who are actually training the model? I think in the US, it’s generally accepted that the core research and engineering staff normally doesn’t get to be more than 100 to 200 people at OpenAI or somewhere like that, and then there’s a lot of support around them in terms of product and distribution. Do you feel like this is similar, or is there a small core research team?
Zixuan Li
It’s similar—100 to 200 people. I think that’s enough.
Nathan Labenz
Yeah, because—
Zixuan Li
You need to be focused, right? There are people preparing data, and there are people doing the product stuff. But for the core team, you don’t need that many people because you need to stay focused. These people need to be really talented. They cannot make many mistakes, right?
Nathan Labenz
Do you know if that’s different at bigger companies?
Zixuan Li
I think for bigger companies, there might be different groups. They have more GPUs, and they can do more exploration. For example, at ByteDance, they’re chasing top performance not only in text generation but also in video generation, speech, and other areas, so they can allocate resources to multiple teams.
But inside these teams, I think the core members are still the same—maybe 10 to 20, and another 80 or 100 doing the training or data preparation.
Nathan Labenz
There was a lot made in Chinese and Western media about how DeepSeek was biased against people who had studied abroad. I’m curious about any other broader dynamics you see with relation to returnees versus people who did their whole education in China.
Zixuan Li
I think there’s no bias. They want the best people, and generally, the interviewees who only stayed in China performed the best in their interviews. They have no bias. But maybe people coming from the US or other countries just did worse in their interviews.
I believe that’s not a bias, because they’re judging very well. They have their standards. Maybe their standard is different from what you do around the world. Actually, I believe that’s another issue.
Even inside China or inside our team, it’s the same standard. I joined Zhipu AI after coming back from the US, but I think nobody actually knows.
People will never ask, “Are you studying abroad?” or “Do you have a master’s degree from MIT?” I believe maybe only 10 people in Z.ai know about this. So there’s no bias because people don’t care.
Nathan Labenz
Yeah.
Let’s talk a little bit about open versus closed source and Chinese model developers broadly, and Z.ai in particular. How do you think people—what’s the thought process behind so many models going open source in recent years?
Zixuan Li
First, I think, generally, we need to devote more to the research area. Llama is doing this, Qwen is doing this, and Kimi is doing this. We’re also doing this. We want to contribute more to academia and to the exploration of all possibilities. I think that’s our top priority.
But beyond that, as a Chinese company, we need to be open to get accepted by some companies. People will not use the Z.ai API to try your models. Maybe they deploy them on Fireworks, maybe they use them on Groq, and maybe they download them to their own chips. I think it’s not easy to get famous in the United States because people just don’t accept your API. The models need to be hosted in the U.S., so I think it’s necessary to be open right now for people to use GLM.
Nathan Labenz
I mean, this is what our company does. Where I work, I wouldn’t be able to sign up for the API service as an enterprise, but I can still use multiple Chinese models when I’m training. I’m using multiple models and might come across this. So it’s not surprising, but it’s a good time to articulate it.
Zixuan Li
Yeah, we also learned from DeepSeek. We had a closed-source version in 2024. Our flagship model was closed source back then. But when DeepSeek-R1 launched, we realized that you can do this thing simultaneously. You can be really famous for open-sourcing your model while getting some business return through the API or other things, like collaboration. You need to expand the cake first and then take a bite of it.
Nathan Labenz
Maybe taking one step back, why is it so important for Chinese model makers to get famous in the U.S., or to get global adoption more broadly?
Zixuan Li
Because I think there’s a better ecosystem for developers and research in the United States. You need to get accepted by the top researchers, right? If we don’t open-source our models, we’ll never have the opportunity to join this conversation. That’s also important.
We learn from X, YouTube, and Reddit every day, and all the Chinese tech media are also paying attention to U.S. KOLs and influencers.
Speaker 3
This was very surprising, I think, to both Nathan and me—the way the Chinese media covers the models, especially the Chinese models that Americans are talking about. It’s a very curious trend.
Zixuan Li
Yeah, because you have people like Andrej Karpathy, Sam Altman, and Elon Musk. They not only talk about their own models, but also about what’s going on elsewhere, so everyone knows. If they post a tweet, everyone knows what’s going on, what models they’re picking, and what preferences they have. Their views on, maybe, Qwen versus Claude—all of social media will try to grasp their ideas immediately.
That’s very important. We also learned this from DeepSeek. Frankly speaking, we used to neglect the importance of the global economy. We thought we needed to sell our products and APIs directly to Chinese enterprises. But nowadays, Chinese enterprises are still paying attention to your global brand and your global performance.
Nathan Labenz
Yeah, this reminded me of something I’ve been curious about. We know the conversation is recursive. We know that the Chinese tech pace is a lot faster than what American Silicon Valley is looking at. But is there anything about the AI debate or discourse in China that Western media tends to miss, in your opinion? Are there any issues, debates, or things that people are really interested in that people in the English-speaking discourse tend not to understand?
Zixuan Li
I just talked to a professor from Germany yesterday, and he mentioned some models that he knew people were talking about these days, like Llama, Qwen, and even Mistral, but not GLM. So there are many people still missing out on that.
Nathan Labenz
That’s personal. In the San Francisco circles, more people are talking about GLM than Mistral and arguably Llama these days. So you’ve made a lot of progress.
Zixuan Li
Yeah, we’ve made a lot of progress. But we also track the discussions on Reddit and other social media, and we still see a lot of people asking, “What is GLM? Is it a good model? Where does it come from? Did it come from nowhere?” There are still a lot of similar questions. We only have 20,000 followers on X, so that’s quite few. Nobody actually has a very deep understanding of GLM compared to other models.
Nathan Labenz
I think DeepSeek has around 1 million. It’s crazy.
Zixuan Li
Yeah, and that’s even big for a whole lot of American companies. For a new American tech company, that would even be big. It’s remarkable. Mistral and Cohere also get much more attention compared to Kimi and Z.ai, so we still need to do better with our branding and our engagement in the technical community.
Nathan Labenz
You mentioned selling API access to Chinese companies. Tell us a little bit about adoption in China and what the sales process is like. Do they all just have VPNs and use Claude anyway? What’s it like trying to do enterprise sales in China?
Zixuan Li
You have 2 types of enterprises. One is companies that can use an API. There are also companies that need to deploy the model on their own chips, and they cannot accept sending data to other companies, even Z.ai or Alibaba. That’s a requirement.
For those companies, they require DeepSeek. There are teams deploying DeepSeek for them—not from DeepSeek. Any company can deploy DeepSeek, right? They usually build on top of the DeepSeek model with RAG, data storage, workflows, and other things.
The other type uses APIs, maybe from tech companies or media companies. These companies accept APIs because they need to standardize their workflows. For API companies, I think they choose based on the balance between performance and price. ByteDance is doing great in that area. I believe ByteDance dominates API services.
Qwen is still trying to sell its APIs because Qwen3-Max is a closed-source version, right? If you have—yeah, I’ve heard of it.
They have open-sourced some models but also keep some things closed source for selling. For us, we have open-sourced our foundation models. We are frequently asked, “What’s different about your service from the open-source version?” because we can't deploy the open-source version ourselves. So we need a better engineering team and faster decoding speed. We need to do more on top of just a good model.
That might be our unique selling point. We need to do searches. We need to build our MCP. We’re trying to get a competitive advantage over other GLM providers.
Nathan Labenz
Is that annoying? Or fun?
Zixuan Li
It’s fun because I think it’s necessary to open-source your models. So how you get a bite in that case is really important. We’ve been figuring this out for a long time, but recently we found that a subscription is a good idea with the GLM Coding Plan.
With a subscription, your users will become stickier. They love this area because you don’t have to worry about how much one prompt consumes in your dialogue. Maybe inside Claude Code, a lot of interaction will consume a million tokens, but you don’t have to worry about it. So we will figure it out for our users.
Nathan Labenz
Do you think you have meaningful adoption there? In the US market, I could start using Claude Code, Codex, Gemini, and whatever all for free, along with some basic Cursor. That’s why I was wondering: are people in the US actively using this? Is this a growing market that you think you’re going to eat into? GLM has one, and I might have tried it, but I’ve always thought, “Oh, I have my own ChatGPT subscription.” I’m just wondering whether, on the ground, it feels optimistic—whether it feels like something that’s really shifting the needle.
Zixuan Li
Definitely, I’m very optimistic, because we don’t have to persuade 50% of people to do this. Maybe you only need 5%. But 5% is a huge market. If 5% of Claude Code users shift their model to GLM, it’s a huge market.
Nathan Labenz
Yeah, and it’s growing so fast.
Zixuan Li
But not just for Claude Code, because we’re trying new ideas like role-playing. Many people are still curious and are using GLM on JanitorAI because we did very well in role-playing. We’re trying to reach more markets: coding markets and agentic markets. Maybe one day Meta will be using our model.
Nathan Labenz
All right, we’ve got to take a step back and explain role-playing. What is it? How do you make a model that’s good at it? What are people using it for?
Zixuan Li
Before GLM-4.6, with models like GLM-4.5, we were relatively weak in role-playing because we had to train on that data. So we needed to create some data and let the model follow the instructions. For role-playing, there’s often a very long system prompt. If you don’t train on that kind of material, the model will forget who it is, forget all the instructions, and just use its general performance to carry on the conversation.
But for a role-playing task, if you give the model very long instructions, it will strictly follow those instructions and show more emotion or more behavior in accordance with them.
Nathan Labenz
Just to be clear, this is people having a conversation saying, “I’m a Japanese pirate. I’m raiding the coast of Taiwan in 1570, and I want to plan an attack to defend the fort.” People write out five pages of background, and then these are chatbots, right? You’re having conversations where you’re playing a character.
It’s like playing a text-based RPG from the 1980s, except it’s AI and it just generates the content. To be clear, I’m not sure everyone knows what role-playing means when it comes to AI.
Zixuan Li
We also try something very interesting, like Family Guy. We have our own Stewie. You just give a description of what Stewie does and a history, and then you can create your own Stewie. We perform very well in text generation, but if we have a speech model, we can recreate a Stewie.
Nathan Labenz
Was there a specific kind of pre-training data or reinforcement learning that you needed to do to get this? Or did it just suddenly become clear that this was really good at pretending to be cartoon characters?
Zixuan Li
I believe it’s mainly post-training data.
Nathan Labenz
There’s been a big discussion lately in the US about people being worried that folks are falling in love with AI. There’s also this whole discussion about AI psychosis, where ChatGPT convinces people who trust it too much to harm themselves.
I’m curious about your broad sense of that type of discussion, in China broadly and then internally in your firm, regarding these questions about people using AI for play or emotional support.
Zixuan Li
I just read the post from OpenAI yesterday because they invited a lot of experts to try to frame a model that, I think, is not addictive. They train the model to say it’s an AI instead of saying it’s a human being, not letting people attach to ChatGPT anymore.
I’ve read this, and a bunch of people have read this, so we can discuss it. I think it’s discussed internally when people find relevant material. We have the news. It’s a hot topic, and we can discuss it.
But from a broader audience perspective, I think not many people are looking into this because we’re not there yet. If we have a model that can perform like GPT-5, then we can move on to removing the addiction. But performance is still not on par with these top closed-source models. We need to chase these models first.
While we chase these models, we’ll shift our focus to data collection and data preparation, and sometimes the model behavior will change dramatically. If we do similar things with our previous model, it will be outdated in the next version. So performance is still very relevant currently.
Nathan Labenz
I’m guessing this is somewhere in the rundown, but how is the balance of optimism versus fear about AI as a long-term trajectory in your lab versus China generally? In the US, there’s a very large concentration of people who worry deeply about the long-term potential of AI, whether it’s a powerful entity, a concentration of power, or other things.
There are people who think this is the most important technology that has ever been invented and that we have to be really serious about it. I’m wondering where on this spectrum you think the lab has a culture of, or whether it’s not really something that’s debated and you’re just building a useful thing and going to keep making it better.
Zixuan Li
I think developers fear it the most. When you use Claude Code or Codex, you get that fear in a very concrete way. They can do all the tasks for you, especially for junior developers.
But for writers and managers, I think it’s simpler because we have SaaS. We have other technologies helping them already, so large language models like ChatGPT are just another helper for them. I cannot feel fear coming from the general public.
But specifically for software developers and data analysts, they fear it the most because they try out the new models and products more frequently than the general public. They can feel the power.
Many people use DeepSeek and other chatbots. DeepSeek can help them brainstorm ideas, polish their writing, and do translations for them. But they don’t believe that this work can replace them. For developers, it’s a different story.
Nathan Labenz
What are the main fears? Is it just people’s jobs getting taken away, or AI taking over the world? For the people who are worried, what are they worried about?
Zixuan Li
Maybe jobs. Jobs being taken away.
Nathan Labenz
That’s pretty different from the US. There’s definitely a huge culture—not a majority in terms of the number of people, but a very vocal minority—that influences a lot of the thinking about the risks of AI well beyond just job loss.
Job loss is almost an assumption for many people in the US. Then there are added fears on top of that, and I think that’s a very different media ecosystem and thought ecosystem.
Zixuan Li
I definitely know about this because I did the research.
Nathan Labenz
You lived here. You lived through some of this, obviously.
Zixuan Li
Everyone at MIT was talking about how AI would change the world—not on the positive side, but on the negative side.
Nathan Labenz
Why do you think this is? Is it that Chinese society is a little more practical, or does job loss feel more imminent? Is it because it’s less of a market-driven economy?
Zixuan Li
I believe that people just know about DeepSeek because maybe only 1 million people follow the latest trends, while there are 1 billion people doing their daily work who aren’t impacted by AI. The more you learn, the more fear you will have.
Nathan Labenz
What’s the vibe among these younger engineers you’re talking about—the junior folks who are a little scared? I’m generally curious what gets them into this work in the first place and what makes them want to work at places like Z.ai.
Zixuan Li
At Z.ai, I think we lack people, so there’s no fear about losing jobs here. We have a lot of things to do. But for other companies, especially large enterprises, they may have 10,000 people doing similar things, like data analytics.
And also in back-end engineering. So they might think that if other people are using our code or a generic tool, maybe they just need 50% of the people. Yes, but they can do nothing. They need to wait for their bosses or the founders to make the decision. Like what's happening at Amazon, right? For layoffs, you can do nothing; you just wait for the results.
Nathan Labenz
I wanted to jump in here and also ask about translation. Z.ai's models are very strong at making very contextually rich translations from Chinese to English, and they point that out on social media. Could you talk a bit more about the process behind that, if you know? And what's the secret sauce to translating memes?
Zixuan Li
Yeah, exactly. We're doing very well in translation, especially translation between Chinese and English. I think we are on par with Gemini 2.5 Pro. But you mentioned memes, and memes are also one of our weapons, because we just prepared the data and we understand the culture. We can even translate emoji.
Nathan Labenz
What do you mean? How does that work?
Nathan Labenz
You mean, like, Tencent emojis to Apple emojis?
Zixuan Li
No, you can—if you enter a sentence talking about AI and you use a whale to replace DeepSeek, we might translate this back to DeepSeek.
Zixuan Li
And if you give us a sentence about animals, we will translate it into a whale. You understand the context.
Nathan Labenz
Is it because Chinese internet talk is just so cryptic?
Zixuan Li
Yeah, because people are very naughty. They're naughty, and they sometimes use emoji. There are a lot of companies that include animal names in their brands or logos, and we're trying to use those to replace what people actually use. People also use abbreviations, right? So all those things need to be translated correctly.
Nathan Labenz
I remember a few years ago there was all this discussion like, “Oh, it's going to be really hard to train Chinese models to speak colloquially because all the data is behind walled gardens.” Tencent has the Tencent data, Xiaohongshu has the Xiaohongshu data, and Alibaba has—I don't even know what data they have. Was that a problem for you guys when you were doing more colloquial internet speech, or is there enough out there that you can just scrape stuff and figure it out?
Zixuan Li
We need statistical data, right? We don't have the actual data. We cannot scrape anything from other WeChat user profiles.
Nathan Labenz
Yeah.
Zixuan Li
But we know people are talking, especially in public areas. In the open areas, we can observe what's going on on Xiaohongshu, TikTok, and other platforms. We especially pay attention to their comment areas, because people are really naughty there in their comments. When the TikTok refugees thing happened, we benefited from it because more people, or more software, needed automatic translation. We're trying to acquire some large customers through our translation capabilities.
Nathan Labenz
Does anyone train on danmu data?
Zixuan Li
Definitely. We're trying to collect memes from everywhere, especially for our vision model, because memes are always in image format. I'm trying to understand them with our vision model. I think it's very interesting, and it's also very necessary, because if you cannot translate the comments in a very accurate way, they will not purchase your model.
Unlike YouTube, if you use YouTube's auto-translation, it won't grasp the exact meaning. People just need to understand, “This English version is about this, and I can read it in Chinese; 80% is enough for me.” But for apps like X, Xiaohongshu, and WeChat, you need to understand 100% of the comment area.
Nathan Labenz
Is it a challenge to balance data across markets—not to mention culture? You're marketing to Western users as well as your domestic market. Is that a technical challenge, to feel like you have to do both excellently?
Zixuan Li
I think it's a challenge, but we can do very well in Chinese and English, and we're trying to explore more in French and even Hindi. So we have data on Hindi. We can perform very well in, I believe, 20 languages. But beyond that, we're still exploring the data and software, so we need to register on their platforms to see what people are doing out there. Sometimes it's hard to figure out. I'm trying to learn from Gemini and GPT-5: Why do they do so great at translation?
Nathan Labenz
Can we talk a little bit about compute? There are all these rumors—we're recording this October 29, in the evening U.S. time. Are you excited to buy some Blackwells if they come on the market in the next few weeks?
Zixuan Li
Blackwell is great because it's not only the chip, but also FP4, right? FP4 can reduce a lot of cost. We're trying to use the best we can get. That's a strategy. I think it's pretty clear.
On the model-training side, for the architecture, we use the best. For the chips, not the best, but I think we do the best trade-off between performance and cost.
Nathan Labenz
Do you guys train outside of China as well, or only on the domestic clouds?
Zixuan Li
Yeah, we do inference from outside China. But all the training is going on here.
Nathan Labenz
How do you feel about Huawei chips and software? Are they going to make it?
Zixuan Li
Yeah, we are going to use them, because we have multiple models, like GLM-4.6, the upcoming GLM-4.6 Air, and our previous version. So we need to find the best use case for all sorts of chips, domestic chips and NVIDIA chips. We need to classify the use case, because for one customer maybe it needs 30 tokens per second, and for another customer it needs 80 tokens per second. So maybe for one customer or one use case, some chips are enough, and for others, we need better chips and better inference techniques.
Nathan Labenz
Do you try to do any API sales, or just enterprise sales in general, outside the U.S. or China? Since we mentioned having a lot of languages and whatnot, do you see any use cases coming from other places?
Zixuan Li
We have 2 platforms. In China, our platform is called BigModel. BigModel is like a large language model; it's a simple translation. BigModel.cn. We also have Z.ai. It's called api.z.ai, and it's our overseas platform. I'm actually in charge of api.z.ai. All of our services are hosted in Singapore. So, actually, I'm an employee of a Singaporean company.
Nathan Labenz
Oh, sorry. I wasn't clear. I meant, do you see much demand coming from non-U.S. countries for Z.ai? Like, other countries?
Zixuan Li
A lot of countries—India, Indonesia, even Norway, and also Brazil. But it depends on who's using Reddit and who's using X, because we basically build our growth on X and Reddit. Maybe someone on YouTube—if people are watching these materials or videos, they will purchase it. But we're trying to do Telegram or other things, so it might shift the proportion of our users.
India and Indonesia are huge markets. But more revenue is coming from the U.S. compared to other countries because they pay more. They buy the Pro plan and Max plan instead of the Lite plan. In terms of users, I think India has the most users. But the U.S. market generates 50% of overseas revenue.
Speaker 1
Jordan, are you on the Chinese plans yet? What's your AI bill? How do we diversify this internationally? I'm on about $500 a month. It's not good.
Speaker 2
I don't know. Just charge it to the firm. Charge it to the Allen estate, Nathan. Come on. We've got to save you. Irene, what was your question earlier?
Speaker 3
Building off what we were talking about earlier with third-party walled gardens, does Z.ai have any thoughts about doing AI search on the Chinese internet, and what that would look like in China, where there increasingly is no unified open internet?
Zixuan Li
I think that's a challenge also for U.S. product builders, because Google doesn't have a search API, and Bing is trying to stop its search API. So there are other third-party providers like SerpAPI, and they basically just scrape the data. They quickly send a request to Google and scrape the page, right? So that's also very challenging for builders like Perplexity and even ChatGPT.
Nowadays, using our own technology or trying to grasp multiple resources from different platforms, I think that's very reasonable. There are other technologies like Manus; they just browse the internet themselves without using an API. I think that's more doable these days when you want to see multiple resources and distinguish the best use case and the best resources. You need to really log into an account and see the data yourself, read the page yourself, instead of just using whatever API gives you.
Speaker 3
Nathan, maybe you want to ask some broader research-direction-type stuff, or whatever else is on your mind.
Nathan Labenz
Where are you planning to take your models next? I think less in-domain, but how do you make models better given that everybody has limited compute and data resources, that we're changing from chat to agent, and it's just—how far out do you think? Or do you think about the very short-term problems? There are just so many directions that you can take it.
Zixuan Li
I don't know. Yeah, I can give some names of the ideas we're exploring right now: on-policy training, on-policy reinforcement learning, because we are quite mature in off-policy reinforcement learning. But for on-policy learning, we still need to explore more. And also multi-agent systems.
When you look at Z.ai Chat, it actually acts like a single agent. One model does the search itself, comes back and does another round of search, then comes back and can generate slides, a presentation, or a poster. Things like that, but it's all performed by a single actor: GLM-4.6.
Nathan Labenz
For our models, do we think you have to change your models a lot in order to do this? I think so much of 2025 has been changing the training stack away from, “We are a chatbot,” to, “Now we are an agent.” What do you think we should change the most about our models, given that? I mean, it's almost like the Air model, the faster model, might be more useful because you can have more of them and things like this.
Zixuan Li
That's the reason why we need to do a very solid evaluation. We have different product solutions, and currently the single agent works very well on our platform. But we need to do more—to try different ideas and see whether we can improve the speed and performance with a multi-agent architecture, and explore other possibilities.
For single agents, they have better context management because you have the best model that can see all the context ahead of the current conversation and follow the instructions maybe better.
Nathan Labenz
That might lose some context. Or, for orchestration, it's hard: if you give 4 agents the same context, they might all try the same thing, and they might not work together well, and stuff like this.
Zixuan Li
Yeah, and maybe even if 1 agent has a hallucination, it will ruin all of the research. But we are also trying to make a longer context window and a longer effective context, because we all know that you said your model can do a 1-million-token context window, but actually it just performs very well inside 60K or 100K. You can release whatever size of context window you want, but it's whether or not it actually works.
Nathan Labenz
How much do you think it's going to be scaling the type of transformers that we have, which is making the long context better, like just improving the data, versus there being fundamental walls that this is approaching? It's the low-hanging-fruit question. Do you think there's a ton to keep improving? Is it easy to find the things to do and you just don't have time?
Zixuan Li
It's not easy. We believe it's an architecture thing. Data can improve, but it cannot cross the wall. There is a wall. So we need better architecture, pre-training data, and post-training data.
Nathan Labenz
Do you think you're starting to hit this wall, or do you just see it coming already? Is this something you're forecasting, or are you saying, “This specific thing—data alone is not solving it for us”? People in the U.S. who are training these models just don't talk about it. They're like, “I don't know. I can't say it.” The models I train are smaller. I think our biggest models are around 30B scale, so when you scale up, you start to see very different limits. What's happening?
Zixuan Li
We need to do some experiments. GLM is a 355-billion-parameter model, right? But we cannot do experiments with this large model. We need to do experiments with some smaller models, maybe 9 billion parameters or 30 billion parameters, and test our hypotheses.
Ninety percent of the time, we just fail. Experiments—you cannot win every time. But you need to do a lot of scientific work to finally get the right answer. So, if you're talking about whether the GLM-4.6 architecture will hit the wall, actually, there's a wall. But we need to shift our focus and start from maybe a new architecture or a new framework for doing this stuff.
Nathan Labenz
So it sounds like doing one of these bigger runs where—I don't know, I don't know if that's necessarily barely making it, but definitely stressful for you.
Zixuan Li
Yes, it's stressful, but we are going to use some engineering techniques to try to compress the context windows to make our users happy, because you don't normally need that much. You don't normally need 1 million tokens. If it cannot perform very well, you can compress the context window to 60K or 30K to make it work.
Nathan Labenz
You mentioned earlier that your inference is abroad, but training is at home. What's behind the rationale for that decision?
Zixuan Li
I think the rationale is very simple, because we provide services to overseas customers, so I think it's a requirement to store the data overseas, right? It's a very strict policy for our Z.ai endpoint. We change that privacy policy every month to make it stricter and more coherent with people's expectations.
But for inference and training, I think it's simpler because we don't have many resources. We only have these resources, and we need to utilize them.
Nathan Labenz
But doing it on Nebius or AWS in Malaysia or Singapore—it's too expensive, it's too slow? You guys already have enough chips at home. What's the thinking there?
Zixuan Li
I think it's not very slow. It's fast, because we don't only change the location of the GPU, but also the CPU and the database. If they're all in Singapore, it's very, very fast. But if you have to go back from Singapore to mainland China and then go back to Singapore, it will be slow.
Nathan Labenz
Okay, but on the training side—in the training side in particular.
Zixuan Li
On the training side, I think it's very simple because we're not open there yet. We're not at peak. We don't have to choose between Amazon, Google, and their own infrastructure. They're doing very complicated stuff, but for us, I think we're still in the initial stage.
We don't have many complicated structures with these large inference providers, so things are just very simple here.
Nathan Labenz
Yeah, for now.
Zixuan Li
For now. For now.
Speaker 1
Irene or Nathan, any more training questions before Irene wraps this up? Only sensitive questions that I don't expect to have an answer to. Like, how big is your next model? How many GPUs do you have? It's like, I don't know. It's not a real question; it's just a curiosity.
Zixuan Li
For our next generation, we're going to launch GLM-4.6 Air, and I don't know what they are called. Maybe Mini. It's a 30-billion-parameter model, so it becomes smaller in a couple of weeks.
I think that's all for 2025. For 2026, we're still doing experiments. Like I said, I try to explore more, but we're doing experiments on smaller models. So they will not be put into practice in 2026, but they give us a lot of ideas about how we're going to train the next generation. We'll see.
When this podcast launches, I believe we will already have GLM-4.6 Air, GLM-4.6 Mini, and also the next GLM-4.6 Vision model.
Nathan Labenz
How long does it take from when the model is done training until you release it? What is your thought process in getting it out fast versus—
Zixuan Li
Get it fast—several hours. Yeah, several hours.
Nathan Labenz
So, would you just open-source it? I love it.
Zixuan Li
When we finish the training, we do some evaluation. After the evaluation, we just release it. We don't have arrangements like sending the endpoint to LMArena or Artificial Analysis and trying to let them evaluate it first and then release the model. We don't have this.
We also don't have a media-buzz thing that tries to make it famous before it's launched. Because we are very transparent, and we believe that if you want to open-source the model, open-sourcing itself is the biggest event.
Nathan Labenz
So this is why you time it to something like that or anything?
Zixuan Li
Yeah, because we're trying to do some marketing. From my side, I want to make it longer. We want a week for me to collaborate with inference providers, benchmark companies, and coding agents, and let everyone try the model before it's released.
But from the company's perspective, if open-source is the most important thing, you only need to prepare the materials for open-source. You need benchmarks and maybe a technical blog. And it's very stressful for me because I need to negotiate with multiple partners within several hours.
We have a new model coming in 2 hours, maybe 3 hours, and maybe you're sleeping. This is huge. Sorry, we don't give you enough time to connect to the model or integrate it, but we're trying to post it with your tweet afterward.
Nathan Labenz
Can you talk a little bit about hours? I mean, in America, we've got our own thing: 0-0-2.
Zixuan Li
What is 0-0-2?
Nathan Labenz
Midnight to midnight with a 2-hour break. So dumb.
Zixuan Li
I think hours vary a lot, even inside the company. Someone would just leave the company at 7:00 p.m.; someone will never leave the company. For me, I work 18 hours a day because I need to negotiate with U.S. large-firm CEOs or the founder of a coding agent. I need to discuss with Fireworks, with Marina[?], and maybe with Kilo Code.
They’re CEOs, so I need to follow their time and do the meetings maybe at 2:00 a.m. or 3:00 a.m. It’s all possible. Oof. Yeah, but for our researchers or engineering team, I think your brains can only work maybe 8 hours a day. If you feel tired and need to get some rest, you should.
I think it’s impossible to ask a top researcher to work 40 hours a week, because that means you’re working really inefficiently, or you’re just attending meetings. You can join meetings for 20 hours a day; you just sit here and listen to other people talking. But you want to read papers, do experiments, and write code. I think 8 hours is enough.
Nathan Labenz
It’s very sensible.
My PhD advisor always said that you can totally change the world if you do 4 hours a day of top technical work. You just go walk in the sun after that. You did a good job.
There are a couple of final questions, then. I’ve always wanted to ask Chinese AI folks this because I feel like the conversation on value propositions can be really Western-centric. How do you explain the value of your work to, let’s say, kids in high school in Beijing or your grandmother?
Zixuan Li
My work?
Nathan Labenz
Yeah, or Z.ai’s work, or data science work. How do you explain the value of that to other people—to kids or older people in China?
Zixuan Li
It’s hard. I can only say I do a similar thing to DeepSeek. We’re just a company like DeepSeek; we do a similar thing because DeepSeek is so famous. Everyone in high school and in kindergarten knows about DeepSeek. For other companies, even Qwen, you cannot explain them to kindergarten kids or high school students.
The value proposition is simple: we are one of the best coding models you can find, especially in China. But high school students always ask, “We have DeepSeek, so what are you doing? Why do we need you? Are you doing the same thing? Are you better? Are you faster? If I’m not using DeepSeek, I’ve got other apps. Why do I need your app?”
That’s where it comes down to. We still need to improve the model performance. I think that’s the top priority. The product and the user experience come second. Without a solid model, nobody will pay attention to you, because we’re all at the same level. Only the most famous one gets all the attention.
Nathan Labenz
So you think the sudden attention to AI models in society came straight out of DeepSeek and the kind of nationalism associated with that?
Zixuan Li
Yeah, I think there’s a hype. It became so famous, even in China, so we remain unknown even here. I believe that a lot of students in Chinese universities haven’t tried GLM or even heard of this company. Everybody reads the news, but not everyone goes to this building to visit GLM.
DeepSeek is all over the news and social media, so it’s really tough to explain our contribution or our value. We say we have a gigantic model or a gigantic tool-use model, but what is tool use? What is search? We’re trying to do more in the future.
Nathan Labenz
Do you think Chinese society is starting to find AI more valuable or scary? Do you have a sense?
Zixuan Li
Valuable. Yeah, because we’re not there yet. AI is not strong enough to make people fear it. There are still hallucinations and problems with not following instructions, so all of those issues make people feel, “Oh, it’s still very silly for me.” Or there’s an agent, but it has hallucinations. How can I use it?
There are still a lot of issues to solve before it feels more real or terrifying to people.
Nathan Labenz
We end every episode with a song. Does Z.ai have a theme song? What do people listen to when they code around the office?
Zixuan Li
No, actually, because our founder loves running. He’s a pro marathoner.
Nathan Labenz
What’s his marathon time?
Zixuan Li
His marathon time is below 3 hours.
Nathan Labenz
That’s solid running.
Zixuan Li
Yeah, because the founder of Moonshot really loves songs, but our founder doesn’t have much interest in songs. For our anniversary, we have a half marathon to celebrate the anniversary. It’s crazy.
Nathan Labenz
I’ve got to go do this. I’m going to go run the Z.ai half marathon next year.
Zixuan Li
I have an intern who finished the half marathon in 3 hours and 15 minutes. She’s a girl. It was crazy. I just waited for her at the finish line. She was almost dead. It was super crazy, but we need to work very long hours, so energy is very important.
Nathan Labenz
So, no music, just sports.
I don’t know if this makes you a good boss for waiting or a terrible boss for making her do it in the first place. Those interns, man—they’ve got to earn their slot and show their dedication.
Zixuan Li
She’s actually the product manager of Z.ai Chat. She built this.
Nathan Labenz
Good. She earned it. After making her do the half marathon, I’m glad you guys gave her a job at the end.
Zixuan Li
She ate 2 hamburgers after that.
Nathan Labenz
Okay, good. All right. Well, this was really fun. Thank you so much for joining the show.