[BidClub_]
The Cognitive Revolution · · 83 分钟

中国 AI 新贵:Z.ai 如何在数小时内构建、基准测试并发布模型,来自 ChinaTalk

Nathan LabenzZixuan Li

YouTube
TL;DR
  • Z.AI 有意从通用聊天转向编程和智能体,由此进入前沿模型的讨论中心。 Nathan Labenz 的基准快照显示,GLM-4.6 在 LMArena 文本榜排名第19,较头部模型落后约65个 Elo 分,但与头部模型对比时仍能赢下约2/5;在网页开发榜排名第9,表现可与 GPT-5.1 竞争。Zixuan Li 表示,公司只是顺着用户认为价值更高的方向走:“Coding and agentic stuff are more useful”,而普通版 Z.AI Chat 仍然免费。

  • 开放权重既是对研究社区的贡献,也是 Z.AI 绕开西方企业可能无法使用中国托管 API 这一现实约束的实用路径。 即使企业不能使用 Z.AI 的 API,也可以通过 Fireworks、其他服务商或自有芯片运行 GLM;Li 表示,如果不开放模型,“we’ll never have the opportunity to join this conversation”。DeepSeek-R1 已证明,扩大影响力与获得商业回报可以并存:“You need to expand the cake first and then take a bite of it.”

  • 这套变现逻辑并不要求市场领先;在高速增长的编程市场中拿下一个狭窄切片,可能就已足够。 GLM Coding Plan 通过订阅制消除用户对智能体循环可能消耗100万 tokens 的焦虑,使用黏性也高于按量计费的 API。Li 的说法很直接:“You don’t have to persuade 50% of people… Maybe you only need 5%”;而 Claude Code 用户的5%就已经是“a huge market”。

  • Z.AI 把组织速度当作竞争资产,训练和评估结束后仅数小时便发布模型。 GLM-4.5 将推理、智能体和编程3个专业教师模型合并,由一支100至200人的高度协同核心团队完成;团队负责人和创始人仍亲自做实验。组织内部的指令是“Get it fast”,有时只提前两三个小时通知集成合作方,因为“the open-source itself is the biggest event”。

  • Li 不认为单靠更多数据就能让当前架构无限延续:“There is a wall.” GLM-4.6 拥有3550亿参数,但由于全规模实验不现实,假设必须在90亿或300亿参数模型上验证;他说这些实验约90%会失败。名义上100万 tokens 的上下文窗口,实际可能只有在6万或10万 tokens 内才有效,因此 Z.AI 正在探索新架构、on-policy 强化学习、更长的有效上下文和多智能体系统,并在必要时压缩生产环境上下文。

  • 中国用户场景带来了差异化的后训练重点,尤其是角色扮演、具备文化语感的翻译和社交媒体解读。 Z.AI 的角色扮演训练会使用长篇角色指令,让模型保留身份、情绪和规定行为。翻译工作覆盖缩写、弹幕、梗和依赖语境的 emoji——例如,当上下文谈论 AI 时,要把鲸鱼符号还原为“DeepSeek”。Li 称,中英翻译已与 Gemini 2.5 Pro 持平;社交应用要求接近完整理解,而不是普通视频翻译中“80% is enough”的标准。

  • Li 认为,中国近期对 AI 的焦虑更多集中在就业,而非生存风险;与此同时,Z.AI 的海外收入仍高度依赖美国。 开发者通过 Claude Code 和 Codex 直接感受到被替代的风险,但 Li 表示,更广泛的公众仍将会产生幻觉的模型视为有用工具,而非可怕威胁:“We are not there yet.” 印度拥有最多海外用户,但美国贡献了约50%的海外收入,因为当地客户购买价格更高的 Pro 和 Max 套餐;训练仍在中国完成,海外服务和数据则托管在新加坡。

摘要 · 为研究而整理的核心内容

1. 让 GLM 登上全球舞台的是编程,而非聊天榜单

  • Nathan Labenz 的基准快照显示,GLM-4.6 在 LMArena 文本榜排名第19,位于 Qwen 3 Max、Kimi K2 Thinking 和 DeepSeek V3.2 旁边——这4个是中国领先的开放权重模型,Mistral 也并未落后太远。GLM-4.6 虽然落后约65个 Elo 分,但这仍意味着它面对头部模型时能赢下5场中的2场。

  • 网页开发才是 GLM-4.6 更强的表现:排名第9,可与 GPT-5.1 竞争,明显只落后于 Gemini 3 和 Claude Opus 4.5。正是这一表现,让许多西方听众此前并不熟悉的公司,突然通过编程产品和集成进入视野。

  • Li 介绍,Z.AI 即 Zhipu AI 的起点是2019年,当时公司以图计算探索 AGI,并打造了学者图谱产品 AMiner。公司在2020年转向大语言模型,2021年发布 GLM,比 GPT-3.5 早1年;GLM-4.5 和 GLM-4.6 终于让这项持续多年的工作获得国际层面的辨识度。

  • 战略转折发生在 Z.AI 于2024年在 Chatbot Arena 排名约第6至第9之后。看到 Manus 和 Claude Code 在2025年获得关注,公司将最高优先级转向具备经济价值的编程和工具使用:“We need to follow the trend and also predict the future.”

2. 小而亲力亲为的组织,围绕统一模型运转

  • GLM-4.5 源自推理、智能体和编程3个独立教师模型,之后蒸馏为一个统一模型。Li 认为关键不在组织形式有多新,而在协作方式:预训练和后训练团队“just sit next to each other”,围绕同一个目标工作,而不是各自优化彼此割裂的任务指标。

  • 前沿变化发生在训练进行期间,管理层不能退回到只定目标的位置。Li 认为,“You need to feel the trend yourself”;Z.AI 的创始人会读论文、跑实验,把实时结果与竞争对手动向结合起来,而不是让下属替自己过滤证据。

  • 核心研发团队约100至200人,Li 认为这个规模已经够用,也更容易保持专注。在更大的公司里,一个领域团队可能仍只有10至20名核心成员,另外由约80或100人负责训练或数据准备。

  • 在读博士生将模型训练与学术研究视为兼容而非竞争的两条轨道:构建统一的智能体编程模型,可能是“one of your greatest achievements ever”。招聘看重论文、GitHub 工作、竞赛、GPU 经验和实际训练经历;创业公司的薪酬较低,则通过寻找有抱负、愿意“fight together”的人来弥补。

3. 海外求学经历不如实际成果重要

  • 面对中国实验室歧视海外毕业生的说法,Li 否认这一前提:公司要的是最好的人,海归可能只是在某种面试标准下表现更差。他在 MIT 学成后加入 Zhipu AI,但估计 Z.AI 只有约10人知道他在哪里读过书,因为“people don’t care”。

  • 中国的大公司——包括 Baidu 和 Alibaba——通常会优先吸纳人才,因为它们能提供更高薪酬。创业公司则需要自驱型员工,他们被年轻团队以及让某个“seems to come from nowhere”的东西与成熟模型竞争的可能性所吸引。

  • Li 自己的经历正体现了这一市场现实:在简历数量激增的情况下,其他知名 AI 公司拒绝或忽略了他的申请。他在 MIT 做的 AI for science 和 alignment 研究,与最初负责国内聊天机器人战略的岗位并不直接相关,但让他形成了对 OpenAI 和 Anthropic 如何追逐前沿的实际理解。

4. 开放权重先是分发渠道,其次才是意识形态

  • Li 首先将开放定义为对研究的贡献,把 Z.AI 与 Llama、Qwen 和 Kimi 并列。但商业约束同样关键:美国机构可能不能使用 Z.AI 的 API,却仍可自行部署 GLM、通过 Fireworks 或 Groq 获取 GLM,或者在自有芯片上运行。

  • Z.AI 在2024年将旗舰模型保持闭源后改变路线。DeepSeek-R1 证明,公司可以通过开源获得全球知名度,同时保留 API、合作和协作收入:“You need to expand the cake first and then take a bite of it.”

  • 全球技术认可也会反哺国内信誉。中国媒体会迅速转发 Andrej Karpathy、Sam Altman、Elon Musk 等美国科技人物的偏好;即使是中国企业,也会评估供应商的“global brand and global performance”,这纠正了 Z.AI 早期认为只靠国内 API 销售就足够的判断。

  • 但这种认可仍不完整。Z.AI 在 X 上仅有约20,000名粉丝,而 DeepSeek 约有100万;Reddit 用户仍会问 GLM 是“come from”哪里。Nathan Labenz 反驳称,如今旧金山圈子讨论 GLM 的程度已超过 Mistral,甚至可能超过 Llama,但 Li 认为,品牌和技术社区参与度仍落后于模型质量。

5. 在开放模型之上叠加服务与订阅

  • 中国企业需求分为自托管和 API 客户两类。数据敏感型组织会要求集成商把 DeepSeek 等模型部署在私有芯片上,再叠加 RAG、数据存储和工作流;科技和媒体公司则接受 API,主要按性能和价格选择,Li 认为这一市场目前由 ByteDance 主导。

  • Qwen3-Max 展示了一种混合策略:部分模型开放,而 Qwen3-Max 仍保持闭源以服务 API 销售。由于 Z.AI 开放了基础模型,买家会持续追问为何不自行托管;因此答案必须来自工程能力而非排他性,包括更快的解码、搜索、MCP 能力以及围绕同一套权重的其他改进。

  • GLM Coding Plan 把这套基础设施变成订阅服务。用户不必计算每次智能体交互消耗多少 tokens——一次 Claude Code 对话可能就会用掉100万——订阅制消除了这种计量焦虑。Li 认为,可预测性会带来更强黏性,而不只是换一种计费方式。

  • Labenz 质疑,在 Claude Code、Codex、Gemini 和基础版 Cursor 都可以试用的情况下,GLM 是否能获得付费用户。Li 给出的答案是市场份额算术,而非市场统治:即便转化 Claude Code 用户的5%,也足以形成有意义的规模;此外,智能体、角色扮演以及 Meta 等大型平台都可能提供进一步的细分市场。

6. 角色扮演与文化翻译需要独立的数据策略

  • 在 GLM-4.6 之前,GLM-4.5 的角色扮演能力相对较弱,因为后训练没有覆盖相关数据。训练中没有长篇角色设定,模型就会忘记自己是谁,退回到泛化对话;加入这类数据后,模型能够在互动过程中保留指令、情绪和规定行为。

  • Nathan Labenz 给出的具体类比是文字 RPG:用户提供5页身份和背景设定。Li 则举了一个流行文化样本:描述 Family Guy 中 Stewie 的性格和故事背景,Z.AI 就能生成对应的文字人格;再加入合适的语音模型,这种表现可以延伸到声音交互。

  • 翻译是另一项有意打造的专长。Li 称,GLM 的中英翻译表现已与 Gemini 2.5 Pro 持平,能够处理缩写、梗和 emoji:在 AI 语境中,鲸鱼可能指 DeepSeek;在动物语境中,同一符号则应继续译为鲸鱼。“We understand the culture”是其声称的优势。

  • Z.AI 无法抓取私密的 WeChat 对话,因此研究公开的 Xiaohongshu、TikTok 和评论区——那里用户“very naughty”——并使用弹幕和图片梗训练视觉能力。“TikTok refugees”事件推高了翻译需求;与 YouTube 上80%的理解度可能已经够用不同,X、Xiaohongshu 和 WeChat 等应用要求用户基本理解每一条评论。

7. 安全争论随能力而来,而失业最具现实感

  • Li 读过 OpenAI 关于降低不健康依恋的研究——训练 ChatGPT 识别自己是 AI 而非人类——并表示相关问题在适用时会被内部讨论。但在模型仍落后于最强闭源系统时,Z.AI 优先追求能力:“If we have a model that can perform like GPT-5, then we can move on to remove the addiction.”

  • 他的理由部分来自技术现实:Z.AI 为追求能力而改变数据采集和后训练方式,不同版本的行为可能剧烈变化,为旧模型设计的干预措施也会随之失效。这并不是说依恋风险不存在,而是公司认为当前真正的瓶颈仍是性能。

  • 软件开发者和数据分析师最直接地感受到恐惧,因为编程智能体已经能够完成具体任务,尤其会影响初级开发者。作家和管理者长期使用 SaaS 及其他辅助工具,因此又一个头脑风暴或润色工具暂时还不像替代品;在 Li 的叙述中,他听到最多、也最明确的担忧是失业。

  • Labenz 强调了这一点与美国少数高声讨论就业之外风险的人群之间的反差,Li 在 MIT 期间就接触过这种文化。他估计,约100万人密切跟踪前沿进展,而约10亿人仍在基本不受影响地日常工作:“The more you learn, the more fear you will have.”

8. 全球推理、混合芯片与本地搜索共同塑造部署

  • Z.AI 在中国训练模型,但通过新加坡向海外用户提供服务,国际服务也托管在那里。Li 表示,海外数据驻留是硬性要求,隐私政策几乎每月修改一次;将 GPU、CPU 和数据库放在新加坡,可以避免请求绕回中国大陆所产生的延迟。

  • Blackwell 的吸引力不只在芯片本身,也在于 FP4,后者可能大幅降低成本。不过,Z.AI 仍会根据客户需求,让 GLM-4.6、即将推出的 GLM-4.6 Air 及旧模型适配国产硬件和 NVIDIA 硬件:一个客户可能需要每秒30 tokens,另一个则需要80。

  • 印度拥有最多海外用户,印度尼西亚、挪威和巴西也有显著需求。用户主要通过 X、Reddit 和部分 YouTube 发现 Z.AI,这使用户构成带有平台偏差;但美国仍贡献约50%的海外收入,因为当地客户选择 Pro 和 Max 套餐,而非 Lite 套餐。

  • Irene Jiang 问到中国平台相互封闭时搜索如何进行。Li 认为美国也面临类似问题:Google 没有搜索 API,Bing 也在试图停止其 API。Z.AI 可以聚合多个资源,也可以让智能体登录平台自行浏览页面,从而保留通用 API 可能遗漏的源数据访问能力。

9. 下一轮增益需要架构实验,而大多数实验都会失败

  • 在 off-policy RL 已相对成熟后,Z.AI 正探索 on-policy 强化学习以及多智能体系统。当前产品本质上是一个 GLM-4.6 actor,能够反复搜索,并在保留此前上下文的同时生成幻灯片、演示文稿或海报。

  • 多智能体设计可能提升速度或性能,但编排会引入明显的权衡。如果多个智能体接收相同上下文,它们可能重复劳动或无法协同;一个产生幻觉的智能体,还可能污染整个研究结果。

  • 宣称的上下文长度不等于有效上下文长度:模型可能声称支持100万 tokens,但实际只有在6万或10万 tokens 内表现良好。Z.AI 可以将客户工作负载压缩到6万或3万 tokens,因为大多数用户并不需要100万;但 Li 不认为这一工程折中就是根本解决方案。

  • Li 的判断是绝对的:数据单独无法跨过“there is a wall”。GLM-4.6 拥有3550亿参数,因此 Z.AI 会在90亿或300亿参数模型上验证假设;“90% of the time we just fail”。继续进步需要更好的架构、预训练和后训练数据,以及可能全新的框架。

10. 数小时内发布,让上线变成协调难题

  • Li 预告了一款300亿参数的下一代模型,称其为 GLM-4.6 Air,但也表示名称可能改为 Mini。他预计 GLM-4.6 Air、GLM-4.6 Mini 和 GLM-4.6 Vision 在播客上线时已经可用。他还表示,为下一代准备的小模型实验不会在2026年投入实践,但会为未来训练提供思路。

  • 发布节奏快到几乎按字面执行:训练结束、评估跑完,模型几小时后就可以上线。Z.AI 不会先把 endpoint 发给 LMArena 或 Artificial Analysis 评估,也不会先做媒体造势;“if you want to open-source the model, the open-sourcing itself is the biggest event.”

  • Li 希望有约1周时间协调推理服务商、基准测试方和编程智能体合作伙伴。但实际情况可能是通知对方:“We have a new model coming in two hours, maybe three hours, maybe you are sleeping”,然后在模型发布后放大合作方的集成成果。

  • 由于美国合作方会议往往安排在凌晨2点或3点,Li 的全球化工作日可能长达18小时,但他不接受研究人员采用同样的作息:一个大脑每天认真读论文、做实验和写代码的时间可能只有8小时。Z.AI 仍需要模型质量和行业认知度才能脱颖而出;没有扎实的模型,“only the most famous one gets all the attention.”

Nathan Labenz

Hello and welcome back to The Cognitive Revolution. Today, I'm honored to share a special cross-post from the China Talk podcast, hosted by Jordan Schneider, China Talk analyst Irene Jiang, and Nathan Lambert of AI2 and the Interconnects Substack, featuring a conversation with Zixuan Li, director of product and generative AI strategy at Z.AI, also known as Zhipu AI, about the culture, incentives, and constraints shaping Chinese AI development.

I imagine that many, even in our AI-obsessed audience, will not be familiar with Z.AI, but their model releases demonstrate that they are a significant player worthy of our attention. As of today, their latest GLM-4.6 model holds the number 19 spot on the LMArena text leaderboard. Its Elo rating is roughly 65 points behind the current leaders, which means that it still wins 2 out of 5 head-to-head comparisons with the leaders, and it happens to sit right next to Qwen 3 Max, Kimi K2 Thinking, and DeepSeek V3.2.

Together, these 4 models, all from China, are indeed the top 4 open-weight models available today, though it should be noted that Mistral is not too far behind. On the WebDev Arena leaderboard, GLM-4.6 does even better, coming in at number 9, making it competitive with GPT-5.1 and meaningfully behind only Gemini 3 and the new Claude Opus 4.5.

All that said, this conversation goes way beyond benchmarks and touches on a number of important topics, including why we should understand Chinese companies' open-weight strategy not so much as an ideological commitment, but as a practical marketing tactic used by companies that are seeking to gain global mindshare while recognizing that Western enterprises simply can't and won't use their APIs. We discuss culturally distinct AI use cases, such as role-play, that matter in China and how these drive different fine-tuning priorities than what we typically see from Western companies.

We also discuss the role that Silicon Valley thought leaders play in establishing credibility for Chinese companies, even in their home market; the market for AI talent in China and why very few people at Z.AI even know that Zixuan studied at MIT; Zixuan's view that there is a wall and that further architectural breakthroughs will be needed; the extreme velocity with which Z.AI releases models, which often involves shipping within hours of completing training; what, if anything, people in China fear about the AI future; and how Chinese companies generally still view themselves not as peers or rivals to leading American companies, but as upstarts that are just trying to keep pace and will be quite happy if they can secure for themselves a meaningful niche in the global AI marketplace.

This is a fascinating, detail-rich conversation, and regardless of your attitude on U.S.-China competition, it's clear to me that we in the West don't hear nearly enough from the actual builders inside Chinese AI labs. So, I appreciate Jordan for allowing me to cross-post this episode, and of course, I encourage everyone to subscribe to China Talk, online at chinatalk.media. With that, I hope you find as much value as I did in this behind-the-scenes look at frontier AI development in China, with Zixuan Li of Z.AI from the China Talk podcast.

Speaker 1

Zixuan Li, who studied in the U.S. before moving back to China, works at Zhipu AI, or Z.AI. We're going to let him introduce himself and his role. Co-hosting today are Irene, longtime China Talk analyst, and Nathan Lambert of AI2 and the Interconnects Substack. Welcome to China Talk, everyone.

Speaker 2

Personally, I've known of Zhipu AI, or Z.AI, for at least a year. Then, in December, there was kind of a mind meld, and I was like, "Well, there's another DeepSeek moment," when they released GLM-4.5. You'll have to correct me on whether or not it was before or after the Kimi K2 model.

Zixuan Li

After.

Speaker 2

I guess it was after, and that was a matter of weeks or days after Kimi K2. It's just like, "Wow, okay, there are a lot of people building great models." It's fun to get to learn about some of them and how it compares between U.S. labs and Chinese labs, and I think a lot of it is more in common than different.

Zixuan Li

Hi, everyone. I'm Zixuan Li from Z.AI, and I manage a lot of things, like global partnerships, Z.AI Chat, model evaluation, and our API services. If you have heard of the GLM coding plan, I'm actually in charge of this thing, too. Nice to meet you, everyone.

Speaker 1

Thank you for introducing yourself. I'd love to hear more about how you ended up working in AI after you moved back to China, and why AI specifically.

Zixuan Li

Actually, I applied for multiple roles at companies like Moonshot and MiniMax, but got rejected or neglected because there are so many resumes going to them every day. I studied AI for science and AI safety at MIT, so I did a lot of research on AI for applications and AI alignment.

That's not very relevant to what we are doing right now, but it actually gives me a sense of what's going on in the frontier area. So it helps me a lot in understanding what OpenAI and Anthropic were doing at that time. I think it's very innovative to have this sort of idea and experience.

Speaker 1

Was it ever a debate for you whether to stay in the U.S., or did you always know you were going back to China after grad school?

Zixuan Li

I already knew I was going back because my family is here. But I got the job after graduation, because it's hard to get a job. I continuously applied for jobs, but finally, after 1 month, I got the opportunity to interview with Z.AI.

At that time, I wasn't in charge of the overseas department because our focus was on the domestic area. It was a domestic chatbot, so I was responsible for the strategy of developing a domestic chatbot. It's called ChatGLM.

Speaker 1

Got it. So maybe let's do a little bit of Zhipu AI backstory. When was it founded, and how would you place it within the broader landscape of teams developing models in China?

Zixuan Li

Great. Zhipu AI, and also Z.AI, was founded in 2019. We were chasing AGI at that time, but not with LLMs, with some graph networks or graph computing. We did something like Google Scholar called AMiner.

We used that type of thing to connect all the data resources from journals and research papers into a database, and people could easily search for and map scholars and their contributions. So it was very popular at that time.

But we shifted to exploring large language models in 2020, and we launched our paper, GLM, in 2021. So that's, I believe, 1 year ahead of the launch of GPT-3.5. It was a very early stage, and we were one of the first companies to explore large language models.

After that, we continuously improved the performance of our models and tried new architectures. GLM is a new architecture, actually, but we are going to explore more in the future. I believe that we got famous by the launch of GLM-4.5 and GLM-4.6, because I think they are very capable in coding, reasoning, and agentic tool use.

That's more useful compared to the previous version. People may know us through Claude Code, Kilo Code, and other tools. So we need to combine with these coding products, and that got us famous.

Speaker 1

Let's talk a little bit more about the evolution of GLM-4.5. I don't know, Nathan, this is your question. Why am I asking this question?

Speaker 2

Well, let's think. What does it take to transition from the models that you were early with to things that get international recognition? I have known of Z.AI and your work for years, and then it's like, snap of the fingers, and now this model is on everybody's radar who's paying attention.

Does this feel like something that was just going to happen for you overnight in developing the models? What does that feel like when you go through it? How do you get to that moment? Because there are a lot of people who want to do that at their companies.

Zixuan Li

That's a very interesting point, because in 2024, everyone was interested in Chatbot Arena, right? We saw GPT-4, and we saw Gemini performing very well on Chatbot Arena, so that was our interest, because we paid attention to end users' experience. When there were 2 answers, which one did you prefer?

We did a lot of things on that, and we performed very well on Chatbot Arena, ranking maybe 6th, or between 6th and 9th.

But in 2025, with the launch of Manus and Claude Code, we realized that coding and agentic stuff are more useful, or they can contribute more economically and also in terms of efficiency for people. So I think chat is no longer our top priority.

Instead, we do more exploration on the coding side and the agent side. We observe the trend, and we do a lot of experiments on it. We need to follow the trend and also predict the future.

Speaker 2

Do you feel like you're better at executing code versus Chatbot Arena? Because I like GLM-4.5, and I think GLM-4.5-Air and GLM-4.6 are extremely renowned for this. I think that when you train these models, the process can look very similar depending on what your target is.

I'm just wondering if it was a shift or if it's just that sometimes things work out better than others.

Zixuan Li

It’s a shift, actually. We pay more attention to the coding stuff. On Z.ai Chat, it’s free, right? Nobody’s paying for the chat.

People pay for cloud API use and for agentic stuff, but we just let users chat with the chatbot freely. So, that’s a shift. But we need to continuously improve performance in normal chat and maybe role-playing, but that’s not our top priority.

Nathan Labenz

Let’s talk a little bit about the talent and the internal culture that allowed you to put out GLM-4.5. What do you think is different about Zhipu AI, or what distinguishes Zhipu AI from other labs, both in the US and China?

Zixuan Li

First of all, we’re more collaborative inside the company. Everyone is working toward this single goal. Maybe we have separate team heads, like a pre-training team and a post-training team, but they’re working very closely. They just sit next to each other, working toward a single goal: trying to build a unified reasoning, agentic, and coding model.

We’ve built 3 separate models, as we illustrated in our tech report. We then distilled these 3 teacher models into one single model, GLM-4.5. That’s our goal, and that’s how I believe we built GLM-4.5 more efficiently compared to other companies.

And we’re super young, right? Another point is that, as you mentioned, talent is important. I believe that nowadays you need to do the research yourself. You need to do the training yourself as the head of the team, so you cannot let others do this stuff for you.

Nathan Labenz

Why is that?

Zixuan Li

Because things change really fast. Maybe during your training, there comes GLM-4; there comes GPT-5—anything can happen. You need to feel the trend yourself. You need to combine the results from experiments, the trends, and what’s going on within your competitors’ teams to feel the move yourself.

It’s super important. Even our founder did the experiments himself. He looked at the papers. You need to do things simultaneously, not just set goals for people and let others do the stuff for you.

Nathan Labenz

Yeah, it’s very fast-paced. I think before we started recording, you were also mentioning that there are a lot of PhD students involved. Are these people actively pursuing their PhDs, kind of new grads, or a mix of all of them?

I work at a research institute that is very open-source, and we have a lot of full-time students who are part of it. But when you look at other closed labs in the US, there’s not nearly as much intermingling with academic institutions. I think that could be a really powerful thing if you have this, because there’s a lot of experimental stuff going on there. Do you feel like it’s a kind of open door between some academic institutions and your work?

Zixuan Li

Definitely. There are a lot of PhD students currently here. I believe they’re both pursuing their academic work and working on GLM simultaneously. But they can combine them together, right?

If you’re doing a really innovative job, like training a unified agentic coding model, it’s one of your greatest achievements ever. People won’t say, “Okay, I need to do another research project. Let me finish this first, and then we’ll go back to GLM.” They’ll try to treat GLM as their single biggest achievement.

Everyone is really devoted to this stuff. We hardly see anyone not devoting themselves to training GLM.

Nathan Labenz

Could you talk a little more broadly about the talent market? You mentioned earlier that you had to put your résumé in a lot of places. What does it look like right now? What’s the kind of hierarchy, and what are folks looking for? What are employers looking for, and what is the talent looking for?

Zixuan Li

On the research and engineering side, I think they’re looking for papers, GitHub code, and competition experience. They’re also looking for your experience using GPUs and your experience training models.

For the non-technical side, they’re looking for how you’re going to grow model performance, expand your branding, and a lot of other things. If you’re going to be a product manager, they’re looking for your coding skills, your vision in this area, and how you do the stuff yourself. Those are very important.

I think it’s pretty similar, but you mentioned hierarchy. In terms of hierarchy, large companies choose the people first because they have more money. They can pay more, like Baidu and Alibaba. But for startups, we need people to fight together. You need to fight against other competitors. You need to drive yourself to finish the goals because you don’t get paid that much.

You need ambition. You truly need to enjoy working with really young, talented people and trying to build something like GLM. It seems to come from nowhere and tries to beat other competitors’ models.

Nathan Labenz

How big would you say the team is—the number of people who are actually training the model? I think in the US, it’s generally accepted that the core research and engineering staff normally doesn’t get to be more than 100 to 200 people at OpenAI or somewhere like that, and then there’s a lot of support around them in terms of product and distribution. Do you feel like this is similar, or is there a small core research team?

Zixuan Li

It’s similar—100 to 200 people. I think that’s enough.

Nathan Labenz

Yeah, because—

Zixuan Li

You need to be focused, right? There are people preparing data, and there are people doing the product stuff. But for the core team, you don’t need that many people because you need to stay focused. These people need to be really talented. They cannot make many mistakes, right?

Nathan Labenz

Do you know if that’s different at bigger companies?

Zixuan Li

I think for bigger companies, there might be different groups. They have more GPUs, and they can do more exploration. For example, at ByteDance, they’re chasing top performance not only in text generation but also in video generation, speech, and other areas, so they can allocate resources to multiple teams.

But inside these teams, I think the core members are still the same—maybe 10 to 20, and another 80 or 100 doing the training or data preparation.

Nathan Labenz

There was a lot made in Chinese and Western media about how DeepSeek was biased against people who had studied abroad. I’m curious about any other broader dynamics you see with relation to returnees versus people who did their whole education in China.

Zixuan Li

I think there’s no bias. They want the best people, and generally, the interviewees who only stayed in China performed the best in their interviews. They have no bias. But maybe people coming from the US or other countries just did worse in their interviews.

I believe that’s not a bias, because they’re judging very well. They have their standards. Maybe their standard is different from what you do around the world. Actually, I believe that’s another issue.

Even inside China or inside our team, it’s the same standard. I joined Zhipu AI after coming back from the US, but I think nobody actually knows.

People will never ask, “Are you studying abroad?” or “Do you have a master’s degree from MIT?” I believe maybe only 10 people in Z.ai know about this. So there’s no bias because people don’t care.

Nathan Labenz

Yeah.

Let’s talk a little bit about open versus closed source and Chinese model developers broadly, and Z.ai in particular. How do you think people—what’s the thought process behind so many models going open source in recent years?

Zixuan Li

First, I think, generally, we need to devote more to the research area. Llama is doing this, Qwen is doing this, and Kimi is doing this. We’re also doing this. We want to contribute more to academia and to the exploration of all possibilities. I think that’s our top priority.

But beyond that, as a Chinese company, we need to be open to get accepted by some companies. People will not use the Z.ai API to try your models. Maybe they deploy them on Fireworks, maybe they use them on Groq, and maybe they download them to their own chips. I think it’s not easy to get famous in the United States because people just don’t accept your API. The models need to be hosted in the U.S., so I think it’s necessary to be open right now for people to use GLM.

Nathan Labenz

I mean, this is what our company does. Where I work, I wouldn’t be able to sign up for the API service as an enterprise, but I can still use multiple Chinese models when I’m training. I’m using multiple models and might come across this. So it’s not surprising, but it’s a good time to articulate it.

Zixuan Li

Yeah, we also learned from DeepSeek. We had a closed-source version in 2024. Our flagship model was closed source back then. But when DeepSeek-R1 launched, we realized that you can do this thing simultaneously. You can be really famous for open-sourcing your model while getting some business return through the API or other things, like collaboration. You need to expand the cake first and then take a bite of it.

Nathan Labenz

Maybe taking one step back, why is it so important for Chinese model makers to get famous in the U.S., or to get global adoption more broadly?

Zixuan Li

Because I think there’s a better ecosystem for developers and research in the United States. You need to get accepted by the top researchers, right? If we don’t open-source our models, we’ll never have the opportunity to join this conversation. That’s also important.

We learn from X, YouTube, and Reddit every day, and all the Chinese tech media are also paying attention to U.S. KOLs and influencers.

Speaker 3

This was very surprising, I think, to both Nathan and me—the way the Chinese media covers the models, especially the Chinese models that Americans are talking about. It’s a very curious trend.

Zixuan Li

Yeah, because you have people like Andrej Karpathy, Sam Altman, and Elon Musk. They not only talk about their own models, but also about what’s going on elsewhere, so everyone knows. If they post a tweet, everyone knows what’s going on, what models they’re picking, and what preferences they have. Their views on, maybe, Qwen versus Claude—all of social media will try to grasp their ideas immediately.

That’s very important. We also learned this from DeepSeek. Frankly speaking, we used to neglect the importance of the global economy. We thought we needed to sell our products and APIs directly to Chinese enterprises. But nowadays, Chinese enterprises are still paying attention to your global brand and your global performance.

Nathan Labenz

Yeah, this reminded me of something I’ve been curious about. We know the conversation is recursive. We know that the Chinese tech pace is a lot faster than what American Silicon Valley is looking at. But is there anything about the AI debate or discourse in China that Western media tends to miss, in your opinion? Are there any issues, debates, or things that people are really interested in that people in the English-speaking discourse tend not to understand?

Zixuan Li

I just talked to a professor from Germany yesterday, and he mentioned some models that he knew people were talking about these days, like Llama, Qwen, and even Mistral, but not GLM. So there are many people still missing out on that.

Nathan Labenz

That’s personal. In the San Francisco circles, more people are talking about GLM than Mistral and arguably Llama these days. So you’ve made a lot of progress.

Zixuan Li

Yeah, we’ve made a lot of progress. But we also track the discussions on Reddit and other social media, and we still see a lot of people asking, “What is GLM? Is it a good model? Where does it come from? Did it come from nowhere?” There are still a lot of similar questions. We only have 20,000 followers on X, so that’s quite few. Nobody actually has a very deep understanding of GLM compared to other models.

Nathan Labenz

I think DeepSeek has around 1 million. It’s crazy.

Zixuan Li

Yeah, and that’s even big for a whole lot of American companies. For a new American tech company, that would even be big. It’s remarkable. Mistral and Cohere also get much more attention compared to Kimi and Z.ai, so we still need to do better with our branding and our engagement in the technical community.

Nathan Labenz

You mentioned selling API access to Chinese companies. Tell us a little bit about adoption in China and what the sales process is like. Do they all just have VPNs and use Claude anyway? What’s it like trying to do enterprise sales in China?

Zixuan Li

You have 2 types of enterprises. One is companies that can use an API. There are also companies that need to deploy the model on their own chips, and they cannot accept sending data to other companies, even Z.ai or Alibaba. That’s a requirement.

For those companies, they require DeepSeek. There are teams deploying DeepSeek for them—not from DeepSeek. Any company can deploy DeepSeek, right? They usually build on top of the DeepSeek model with RAG, data storage, workflows, and other things.

The other type uses APIs, maybe from tech companies or media companies. These companies accept APIs because they need to standardize their workflows. For API companies, I think they choose based on the balance between performance and price. ByteDance is doing great in that area. I believe ByteDance dominates API services.

Qwen is still trying to sell its APIs because Qwen3-Max is a closed-source version, right? If you have—yeah, I’ve heard of it.

They have open-sourced some models but also keep some things closed source for selling. For us, we have open-sourced our foundation models. We are frequently asked, “What’s different about your service from the open-source version?” because we can't deploy the open-source version ourselves. So we need a better engineering team and faster decoding speed. We need to do more on top of just a good model.

That might be our unique selling point. We need to do searches. We need to build our MCP. We’re trying to get a competitive advantage over other GLM providers.

Nathan Labenz

Is that annoying? Or fun?

Zixuan Li

It’s fun because I think it’s necessary to open-source your models. So how you get a bite in that case is really important. We’ve been figuring this out for a long time, but recently we found that a subscription is a good idea with the GLM Coding Plan.

With a subscription, your users will become stickier. They love this area because you don’t have to worry about how much one prompt consumes in your dialogue. Maybe inside Claude Code, a lot of interaction will consume a million tokens, but you don’t have to worry about it. So we will figure it out for our users.

Nathan Labenz

Do you think you have meaningful adoption there? In the US market, I could start using Claude Code, Codex, Gemini, and whatever all for free, along with some basic Cursor. That’s why I was wondering: are people in the US actively using this? Is this a growing market that you think you’re going to eat into? GLM has one, and I might have tried it, but I’ve always thought, “Oh, I have my own ChatGPT subscription.” I’m just wondering whether, on the ground, it feels optimistic—whether it feels like something that’s really shifting the needle.

Zixuan Li

Definitely, I’m very optimistic, because we don’t have to persuade 50% of people to do this. Maybe you only need 5%. But 5% is a huge market. If 5% of Claude Code users shift their model to GLM, it’s a huge market.

Nathan Labenz

Yeah, and it’s growing so fast.

Zixuan Li

But not just for Claude Code, because we’re trying new ideas like role-playing. Many people are still curious and are using GLM on JanitorAI because we did very well in role-playing. We’re trying to reach more markets: coding markets and agentic markets. Maybe one day Meta will be using our model.

Nathan Labenz

All right, we’ve got to take a step back and explain role-playing. What is it? How do you make a model that’s good at it? What are people using it for?

Zixuan Li

Before GLM-4.6, with models like GLM-4.5, we were relatively weak in role-playing because we had to train on that data. So we needed to create some data and let the model follow the instructions. For role-playing, there’s often a very long system prompt. If you don’t train on that kind of material, the model will forget who it is, forget all the instructions, and just use its general performance to carry on the conversation.

But for a role-playing task, if you give the model very long instructions, it will strictly follow those instructions and show more emotion or more behavior in accordance with them.

Nathan Labenz

Just to be clear, this is people having a conversation saying, “I’m a Japanese pirate. I’m raiding the coast of Taiwan in 1570, and I want to plan an attack to defend the fort.” People write out five pages of background, and then these are chatbots, right? You’re having conversations where you’re playing a character.

It’s like playing a text-based RPG from the 1980s, except it’s AI and it just generates the content. To be clear, I’m not sure everyone knows what role-playing means when it comes to AI.

Zixuan Li

We also try something very interesting, like Family Guy. We have our own Stewie. You just give a description of what Stewie does and a history, and then you can create your own Stewie. We perform very well in text generation, but if we have a speech model, we can recreate a Stewie.

Nathan Labenz

Was there a specific kind of pre-training data or reinforcement learning that you needed to do to get this? Or did it just suddenly become clear that this was really good at pretending to be cartoon characters?

Zixuan Li

I believe it’s mainly post-training data.

Nathan Labenz

There’s been a big discussion lately in the US about people being worried that folks are falling in love with AI. There’s also this whole discussion about AI psychosis, where ChatGPT convinces people who trust it too much to harm themselves.

I’m curious about your broad sense of that type of discussion, in China broadly and then internally in your firm, regarding these questions about people using AI for play or emotional support.

Zixuan Li

I just read the post from OpenAI yesterday because they invited a lot of experts to try to frame a model that, I think, is not addictive. They train the model to say it’s an AI instead of saying it’s a human being, not letting people attach to ChatGPT anymore.

I’ve read this, and a bunch of people have read this, so we can discuss it. I think it’s discussed internally when people find relevant material. We have the news. It’s a hot topic, and we can discuss it.

But from a broader audience perspective, I think not many people are looking into this because we’re not there yet. If we have a model that can perform like GPT-5, then we can move on to removing the addiction. But performance is still not on par with these top closed-source models. We need to chase these models first.

While we chase these models, we’ll shift our focus to data collection and data preparation, and sometimes the model behavior will change dramatically. If we do similar things with our previous model, it will be outdated in the next version. So performance is still very relevant currently.

Nathan Labenz

I’m guessing this is somewhere in the rundown, but how is the balance of optimism versus fear about AI as a long-term trajectory in your lab versus China generally? In the US, there’s a very large concentration of people who worry deeply about the long-term potential of AI, whether it’s a powerful entity, a concentration of power, or other things.

There are people who think this is the most important technology that has ever been invented and that we have to be really serious about it. I’m wondering where on this spectrum you think the lab has a culture of, or whether it’s not really something that’s debated and you’re just building a useful thing and going to keep making it better.

Zixuan Li

I think developers fear it the most. When you use Claude Code or Codex, you get that fear in a very concrete way. They can do all the tasks for you, especially for junior developers.

But for writers and managers, I think it’s simpler because we have SaaS. We have other technologies helping them already, so large language models like ChatGPT are just another helper for them. I cannot feel fear coming from the general public.

But specifically for software developers and data analysts, they fear it the most because they try out the new models and products more frequently than the general public. They can feel the power.

Many people use DeepSeek and other chatbots. DeepSeek can help them brainstorm ideas, polish their writing, and do translations for them. But they don’t believe that this work can replace them. For developers, it’s a different story.

Nathan Labenz

What are the main fears? Is it just people’s jobs getting taken away, or AI taking over the world? For the people who are worried, what are they worried about?

Zixuan Li

Maybe jobs. Jobs being taken away.

Nathan Labenz

That’s pretty different from the US. There’s definitely a huge culture—not a majority in terms of the number of people, but a very vocal minority—that influences a lot of the thinking about the risks of AI well beyond just job loss.

Job loss is almost an assumption for many people in the US. Then there are added fears on top of that, and I think that’s a very different media ecosystem and thought ecosystem.

Zixuan Li

I definitely know about this because I did the research.

Nathan Labenz

You lived here. You lived through some of this, obviously.

Zixuan Li

Everyone at MIT was talking about how AI would change the world—not on the positive side, but on the negative side.

Nathan Labenz

Why do you think this is? Is it that Chinese society is a little more practical, or does job loss feel more imminent? Is it because it’s less of a market-driven economy?

Zixuan Li

I believe that people just know about DeepSeek because maybe only 1 million people follow the latest trends, while there are 1 billion people doing their daily work who aren’t impacted by AI. The more you learn, the more fear you will have.

Nathan Labenz

What’s the vibe among these younger engineers you’re talking about—the junior folks who are a little scared? I’m generally curious what gets them into this work in the first place and what makes them want to work at places like Z.ai.

Zixuan Li

At Z.ai, I think we lack people, so there’s no fear about losing jobs here. We have a lot of things to do. But for other companies, especially large enterprises, they may have 10,000 people doing similar things, like data analytics.

And also in back-end engineering. So they might think that if other people are using our code or a generic tool, maybe they just need 50% of the people. Yes, but they can do nothing. They need to wait for their bosses or the founders to make the decision. Like what's happening at Amazon, right? For layoffs, you can do nothing; you just wait for the results.

Nathan Labenz

I wanted to jump in here and also ask about translation. Z.ai's models are very strong at making very contextually rich translations from Chinese to English, and they point that out on social media. Could you talk a bit more about the process behind that, if you know? And what's the secret sauce to translating memes?

Zixuan Li

Yeah, exactly. We're doing very well in translation, especially translation between Chinese and English. I think we are on par with Gemini 2.5 Pro. But you mentioned memes, and memes are also one of our weapons, because we just prepared the data and we understand the culture. We can even translate emoji.

Nathan Labenz

What do you mean? How does that work?

Nathan Labenz

You mean, like, Tencent emojis to Apple emojis?

Zixuan Li

No, you can—if you enter a sentence talking about AI and you use a whale to replace DeepSeek, we might translate this back to DeepSeek.

Zixuan Li

And if you give us a sentence about animals, we will translate it into a whale. You understand the context.

Nathan Labenz

Is it because Chinese internet talk is just so cryptic?

Zixuan Li

Yeah, because people are very naughty. They're naughty, and they sometimes use emoji. There are a lot of companies that include animal names in their brands or logos, and we're trying to use those to replace what people actually use. People also use abbreviations, right? So all those things need to be translated correctly.

Nathan Labenz

I remember a few years ago there was all this discussion like, “Oh, it's going to be really hard to train Chinese models to speak colloquially because all the data is behind walled gardens.” Tencent has the Tencent data, Xiaohongshu has the Xiaohongshu data, and Alibaba has—I don't even know what data they have. Was that a problem for you guys when you were doing more colloquial internet speech, or is there enough out there that you can just scrape stuff and figure it out?

Zixuan Li

We need statistical data, right? We don't have the actual data. We cannot scrape anything from other WeChat user profiles.

Nathan Labenz

Yeah.

Zixuan Li

But we know people are talking, especially in public areas. In the open areas, we can observe what's going on on Xiaohongshu, TikTok, and other platforms. We especially pay attention to their comment areas, because people are really naughty there in their comments. When the TikTok refugees thing happened, we benefited from it because more people, or more software, needed automatic translation. We're trying to acquire some large customers through our translation capabilities.

Nathan Labenz

Does anyone train on danmu data?

Zixuan Li

Definitely. We're trying to collect memes from everywhere, especially for our vision model, because memes are always in image format. I'm trying to understand them with our vision model. I think it's very interesting, and it's also very necessary, because if you cannot translate the comments in a very accurate way, they will not purchase your model.

Unlike YouTube, if you use YouTube's auto-translation, it won't grasp the exact meaning. People just need to understand, “This English version is about this, and I can read it in Chinese; 80% is enough for me.” But for apps like X, Xiaohongshu, and WeChat, you need to understand 100% of the comment area.

Nathan Labenz

Is it a challenge to balance data across markets—not to mention culture? You're marketing to Western users as well as your domestic market. Is that a technical challenge, to feel like you have to do both excellently?

Zixuan Li

I think it's a challenge, but we can do very well in Chinese and English, and we're trying to explore more in French and even Hindi. So we have data on Hindi. We can perform very well in, I believe, 20 languages. But beyond that, we're still exploring the data and software, so we need to register on their platforms to see what people are doing out there. Sometimes it's hard to figure out. I'm trying to learn from Gemini and GPT-5: Why do they do so great at translation?

Nathan Labenz

Can we talk a little bit about compute? There are all these rumors—we're recording this October 29, in the evening U.S. time. Are you excited to buy some Blackwells if they come on the market in the next few weeks?

Zixuan Li

Blackwell is great because it's not only the chip, but also FP4, right? FP4 can reduce a lot of cost. We're trying to use the best we can get. That's a strategy. I think it's pretty clear.

On the model-training side, for the architecture, we use the best. For the chips, not the best, but I think we do the best trade-off between performance and cost.

Nathan Labenz

Do you guys train outside of China as well, or only on the domestic clouds?

Zixuan Li

Yeah, we do inference from outside China. But all the training is going on here.

Nathan Labenz

How do you feel about Huawei chips and software? Are they going to make it?

Zixuan Li

Yeah, we are going to use them, because we have multiple models, like GLM-4.6, the upcoming GLM-4.6 Air, and our previous version. So we need to find the best use case for all sorts of chips, domestic chips and NVIDIA chips. We need to classify the use case, because for one customer maybe it needs 30 tokens per second, and for another customer it needs 80 tokens per second. So maybe for one customer or one use case, some chips are enough, and for others, we need better chips and better inference techniques.

Nathan Labenz

Do you try to do any API sales, or just enterprise sales in general, outside the U.S. or China? Since we mentioned having a lot of languages and whatnot, do you see any use cases coming from other places?

Zixuan Li

We have 2 platforms. In China, our platform is called BigModel. BigModel is like a large language model; it's a simple translation. BigModel.cn. We also have Z.ai. It's called api.z.ai, and it's our overseas platform. I'm actually in charge of api.z.ai. All of our services are hosted in Singapore. So, actually, I'm an employee of a Singaporean company.

Nathan Labenz

Oh, sorry. I wasn't clear. I meant, do you see much demand coming from non-U.S. countries for Z.ai? Like, other countries?

Zixuan Li

A lot of countries—India, Indonesia, even Norway, and also Brazil. But it depends on who's using Reddit and who's using X, because we basically build our growth on X and Reddit. Maybe someone on YouTube—if people are watching these materials or videos, they will purchase it. But we're trying to do Telegram or other things, so it might shift the proportion of our users.

India and Indonesia are huge markets. But more revenue is coming from the U.S. compared to other countries because they pay more. They buy the Pro plan and Max plan instead of the Lite plan. In terms of users, I think India has the most users. But the U.S. market generates 50% of overseas revenue.

Speaker 1

Jordan, are you on the Chinese plans yet? What's your AI bill? How do we diversify this internationally? I'm on about $500 a month. It's not good.

Speaker 2

I don't know. Just charge it to the firm. Charge it to the Allen estate, Nathan. Come on. We've got to save you. Irene, what was your question earlier?

Speaker 3

Building off what we were talking about earlier with third-party walled gardens, does Z.ai have any thoughts about doing AI search on the Chinese internet, and what that would look like in China, where there increasingly is no unified open internet?

Zixuan Li

I think that's a challenge also for U.S. product builders, because Google doesn't have a search API, and Bing is trying to stop its search API. So there are other third-party providers like SerpAPI, and they basically just scrape the data. They quickly send a request to Google and scrape the page, right? So that's also very challenging for builders like Perplexity and even ChatGPT.

Nowadays, using our own technology or trying to grasp multiple resources from different platforms, I think that's very reasonable. There are other technologies like Manus; they just browse the internet themselves without using an API. I think that's more doable these days when you want to see multiple resources and distinguish the best use case and the best resources. You need to really log into an account and see the data yourself, read the page yourself, instead of just using whatever API gives you.

Speaker 3

Nathan, maybe you want to ask some broader research-direction-type stuff, or whatever else is on your mind.

Nathan Labenz

Where are you planning to take your models next? I think less in-domain, but how do you make models better given that everybody has limited compute and data resources, that we're changing from chat to agent, and it's just—how far out do you think? Or do you think about the very short-term problems? There are just so many directions that you can take it.

Zixuan Li

I don't know. Yeah, I can give some names of the ideas we're exploring right now: on-policy training, on-policy reinforcement learning, because we are quite mature in off-policy reinforcement learning. But for on-policy learning, we still need to explore more. And also multi-agent systems.

When you look at Z.ai Chat, it actually acts like a single agent. One model does the search itself, comes back and does another round of search, then comes back and can generate slides, a presentation, or a poster. Things like that, but it's all performed by a single actor: GLM-4.6.

Nathan Labenz

For our models, do we think you have to change your models a lot in order to do this? I think so much of 2025 has been changing the training stack away from, “We are a chatbot,” to, “Now we are an agent.” What do you think we should change the most about our models, given that? I mean, it's almost like the Air model, the faster model, might be more useful because you can have more of them and things like this.

Zixuan Li

That's the reason why we need to do a very solid evaluation. We have different product solutions, and currently the single agent works very well on our platform. But we need to do more—to try different ideas and see whether we can improve the speed and performance with a multi-agent architecture, and explore other possibilities.

For single agents, they have better context management because you have the best model that can see all the context ahead of the current conversation and follow the instructions maybe better.

Nathan Labenz

That might lose some context. Or, for orchestration, it's hard: if you give 4 agents the same context, they might all try the same thing, and they might not work together well, and stuff like this.

Zixuan Li

Yeah, and maybe even if 1 agent has a hallucination, it will ruin all of the research. But we are also trying to make a longer context window and a longer effective context, because we all know that you said your model can do a 1-million-token context window, but actually it just performs very well inside 60K or 100K. You can release whatever size of context window you want, but it's whether or not it actually works.

Nathan Labenz

How much do you think it's going to be scaling the type of transformers that we have, which is making the long context better, like just improving the data, versus there being fundamental walls that this is approaching? It's the low-hanging-fruit question. Do you think there's a ton to keep improving? Is it easy to find the things to do and you just don't have time?

Zixuan Li

It's not easy. We believe it's an architecture thing. Data can improve, but it cannot cross the wall. There is a wall. So we need better architecture, pre-training data, and post-training data.

Nathan Labenz

Do you think you're starting to hit this wall, or do you just see it coming already? Is this something you're forecasting, or are you saying, “This specific thing—data alone is not solving it for us”? People in the U.S. who are training these models just don't talk about it. They're like, “I don't know. I can't say it.” The models I train are smaller. I think our biggest models are around 30B scale, so when you scale up, you start to see very different limits. What's happening?

Zixuan Li

We need to do some experiments. GLM is a 355-billion-parameter model, right? But we cannot do experiments with this large model. We need to do experiments with some smaller models, maybe 9 billion parameters or 30 billion parameters, and test our hypotheses.

Ninety percent of the time, we just fail. Experiments—you cannot win every time. But you need to do a lot of scientific work to finally get the right answer. So, if you're talking about whether the GLM-4.6 architecture will hit the wall, actually, there's a wall. But we need to shift our focus and start from maybe a new architecture or a new framework for doing this stuff.

Nathan Labenz

So it sounds like doing one of these bigger runs where—I don't know, I don't know if that's necessarily barely making it, but definitely stressful for you.

Zixuan Li

Yes, it's stressful, but we are going to use some engineering techniques to try to compress the context windows to make our users happy, because you don't normally need that much. You don't normally need 1 million tokens. If it cannot perform very well, you can compress the context window to 60K or 30K to make it work.

Nathan Labenz

You mentioned earlier that your inference is abroad, but training is at home. What's behind the rationale for that decision?

Zixuan Li

I think the rationale is very simple, because we provide services to overseas customers, so I think it's a requirement to store the data overseas, right? It's a very strict policy for our Z.ai endpoint. We change that privacy policy every month to make it stricter and more coherent with people's expectations.

But for inference and training, I think it's simpler because we don't have many resources. We only have these resources, and we need to utilize them.

Nathan Labenz

But doing it on Nebius or AWS in Malaysia or Singapore—it's too expensive, it's too slow? You guys already have enough chips at home. What's the thinking there?

Zixuan Li

I think it's not very slow. It's fast, because we don't only change the location of the GPU, but also the CPU and the database. If they're all in Singapore, it's very, very fast. But if you have to go back from Singapore to mainland China and then go back to Singapore, it will be slow.

Nathan Labenz

Okay, but on the training side—in the training side in particular.

Zixuan Li

On the training side, I think it's very simple because we're not open there yet. We're not at peak. We don't have to choose between Amazon, Google, and their own infrastructure. They're doing very complicated stuff, but for us, I think we're still in the initial stage.

We don't have many complicated structures with these large inference providers, so things are just very simple here.

Nathan Labenz

Yeah, for now.

Zixuan Li

For now. For now.

Speaker 1

Irene or Nathan, any more training questions before Irene wraps this up? Only sensitive questions that I don't expect to have an answer to. Like, how big is your next model? How many GPUs do you have? It's like, I don't know. It's not a real question; it's just a curiosity.

Zixuan Li

For our next generation, we're going to launch GLM-4.6 Air, and I don't know what they are called. Maybe Mini. It's a 30-billion-parameter model, so it becomes smaller in a couple of weeks.

I think that's all for 2025. For 2026, we're still doing experiments. Like I said, I try to explore more, but we're doing experiments on smaller models. So they will not be put into practice in 2026, but they give us a lot of ideas about how we're going to train the next generation. We'll see.

When this podcast launches, I believe we will already have GLM-4.6 Air, GLM-4.6 Mini, and also the next GLM-4.6 Vision model.

Nathan Labenz

How long does it take from when the model is done training until you release it? What is your thought process in getting it out fast versus—

Zixuan Li

Get it fast—several hours. Yeah, several hours.

Nathan Labenz

So, would you just open-source it? I love it.

Zixuan Li

When we finish the training, we do some evaluation. After the evaluation, we just release it. We don't have arrangements like sending the endpoint to LMArena or Artificial Analysis and trying to let them evaluate it first and then release the model. We don't have this.

We also don't have a media-buzz thing that tries to make it famous before it's launched. Because we are very transparent, and we believe that if you want to open-source the model, open-sourcing itself is the biggest event.

Nathan Labenz

So this is why you time it to something like that or anything?

Zixuan Li

Yeah, because we're trying to do some marketing. From my side, I want to make it longer. We want a week for me to collaborate with inference providers, benchmark companies, and coding agents, and let everyone try the model before it's released.

But from the company's perspective, if open-source is the most important thing, you only need to prepare the materials for open-source. You need benchmarks and maybe a technical blog. And it's very stressful for me because I need to negotiate with multiple partners within several hours.

We have a new model coming in 2 hours, maybe 3 hours, and maybe you're sleeping. This is huge. Sorry, we don't give you enough time to connect to the model or integrate it, but we're trying to post it with your tweet afterward.

Nathan Labenz

Can you talk a little bit about hours? I mean, in America, we've got our own thing: 0-0-2.

Zixuan Li

What is 0-0-2?

Nathan Labenz

Midnight to midnight with a 2-hour break. So dumb.

Zixuan Li

I think hours vary a lot, even inside the company. Someone would just leave the company at 7:00 p.m.; someone will never leave the company. For me, I work 18 hours a day because I need to negotiate with U.S. large-firm CEOs or the founder of a coding agent. I need to discuss with Fireworks, with Marina[?], and maybe with Kilo Code.

They’re CEOs, so I need to follow their time and do the meetings maybe at 2:00 a.m. or 3:00 a.m. It’s all possible. Oof. Yeah, but for our researchers or engineering team, I think your brains can only work maybe 8 hours a day. If you feel tired and need to get some rest, you should.

I think it’s impossible to ask a top researcher to work 40 hours a week, because that means you’re working really inefficiently, or you’re just attending meetings. You can join meetings for 20 hours a day; you just sit here and listen to other people talking. But you want to read papers, do experiments, and write code. I think 8 hours is enough.

Nathan Labenz

It’s very sensible.

My PhD advisor always said that you can totally change the world if you do 4 hours a day of top technical work. You just go walk in the sun after that. You did a good job.

There are a couple of final questions, then. I’ve always wanted to ask Chinese AI folks this because I feel like the conversation on value propositions can be really Western-centric. How do you explain the value of your work to, let’s say, kids in high school in Beijing or your grandmother?

Zixuan Li

My work?

Nathan Labenz

Yeah, or Z.ai’s work, or data science work. How do you explain the value of that to other people—to kids or older people in China?

Zixuan Li

It’s hard. I can only say I do a similar thing to DeepSeek. We’re just a company like DeepSeek; we do a similar thing because DeepSeek is so famous. Everyone in high school and in kindergarten knows about DeepSeek. For other companies, even Qwen, you cannot explain them to kindergarten kids or high school students.

The value proposition is simple: we are one of the best coding models you can find, especially in China. But high school students always ask, “We have DeepSeek, so what are you doing? Why do we need you? Are you doing the same thing? Are you better? Are you faster? If I’m not using DeepSeek, I’ve got other apps. Why do I need your app?”

That’s where it comes down to. We still need to improve the model performance. I think that’s the top priority. The product and the user experience come second. Without a solid model, nobody will pay attention to you, because we’re all at the same level. Only the most famous one gets all the attention.

Nathan Labenz

So you think the sudden attention to AI models in society came straight out of DeepSeek and the kind of nationalism associated with that?

Zixuan Li

Yeah, I think there’s a hype. It became so famous, even in China, so we remain unknown even here. I believe that a lot of students in Chinese universities haven’t tried GLM or even heard of this company. Everybody reads the news, but not everyone goes to this building to visit GLM.

DeepSeek is all over the news and social media, so it’s really tough to explain our contribution or our value. We say we have a gigantic model or a gigantic tool-use model, but what is tool use? What is search? We’re trying to do more in the future.

Nathan Labenz

Do you think Chinese society is starting to find AI more valuable or scary? Do you have a sense?

Zixuan Li

Valuable. Yeah, because we’re not there yet. AI is not strong enough to make people fear it. There are still hallucinations and problems with not following instructions, so all of those issues make people feel, “Oh, it’s still very silly for me.” Or there’s an agent, but it has hallucinations. How can I use it?

There are still a lot of issues to solve before it feels more real or terrifying to people.

Nathan Labenz

We end every episode with a song. Does Z.ai have a theme song? What do people listen to when they code around the office?

Zixuan Li

No, actually, because our founder loves running. He’s a pro marathoner.

Nathan Labenz

What’s his marathon time?

Zixuan Li

His marathon time is below 3 hours.

Nathan Labenz

That’s solid running.

Zixuan Li

Yeah, because the founder of Moonshot really loves songs, but our founder doesn’t have much interest in songs. For our anniversary, we have a half marathon to celebrate the anniversary. It’s crazy.

Nathan Labenz

I’ve got to go do this. I’m going to go run the Z.ai half marathon next year.

Zixuan Li

I have an intern who finished the half marathon in 3 hours and 15 minutes. She’s a girl. It was crazy. I just waited for her at the finish line. She was almost dead. It was super crazy, but we need to work very long hours, so energy is very important.

Nathan Labenz

So, no music, just sports.

I don’t know if this makes you a good boss for waiting or a terrible boss for making her do it in the first place. Those interns, man—they’ve got to earn their slot and show their dedication.

Zixuan Li

She’s actually the product manager of Z.ai Chat. She built this.

Nathan Labenz

Good. She earned it. After making her do the half marathon, I’m glad you guys gave her a job at the end.

Zixuan Li

She ate 2 hamburgers after that.

Nathan Labenz

Okay, good. All right. Well, this was really fun. Thank you so much for joining the show.