“AI 洗白”式裁员?+ 为什么 LLM 写不好 + Tokenmaxxing
这轮裁员潮是一个真实的劳动力市场信号,但本期并未把每一种 AI 解释都视为因果关系。 Atlassian 裁员 10%,约 1,600 个岗位;Block 裁员约 40%,即 4,000 人;Reuters 报道称,Meta 可能裁掉 20% 或更多员工,最多涉及 16,000 人,但 Meta 称这是“推测性报道”,截至录音时尚未宣布裁员。Casey Newton 的总结是:公司一直在说 AI 很重要,“迟早,我确实认为我们不得不相信他们”,即便过度招聘、股价疲弱或组织失灵也提供了其他解释。
对投资者而言,AI 同时充当经营模式和估值叙事。 Block 宣布裁员次日上涨 17%,而 Meta 正将可能的裁员与计划中的 1,350 亿美元资本开支并置;Kevin Roose 的框架是,公司未必在降低总成本,而是在“把成本从人力转向 AI”。这场押注是:薪酬支出可以变成算力支出,市场会在生产率提升得到验证之前,就奖励那些先讲出这个故事的管理层。
Meta 提出的“人力换算力”替代仍然高度不确定,因为其自身的 AI 执行表现一直不稳定。 Zuckerberg 称,过去需要大团队的项目如今可以由“一个非常有才华的人”完成;但主持人指出,Meta 已放弃 Behemoth,据报道因未达标而推迟 Avocado,并再次重组 AI 团队。前沿实验室 OpenAI、Anthropic 和 Google 本身都没有进行可比规模的批量裁员,不过 Kevin 指出,OpenAI 和 Anthropic 的规模小得多,一些落后公司也可能在借助 AI 追赶。
AI 的应用已经变成一个让员工怎么选都输的信号困境,也可能成为管理层管束员工的工具。 员工担心,大量使用 AI 会证明自己适应力强,但也会证明自己的工作可以被自动化;Casey 谨慎地观察到,反复裁员已经让 Meta 员工变得更沉默,即便控制员工并非其声称的目的。Kevin 认为这可能为科技行业工会化打开空间,而 Casey 给出的更尖锐的组织动员试金石是:“我想不出有什么事会比 Meta 软件工程师组成工会更让 Mark Zuckerberg 恼火。”
现代聊天机器人在许多任务上比 GPT-2 或 GPT-3 更有用,但后训练牺牲了意外性、声音和风格跨度。 Jasmine Sun 认为,早期模型“疯疯癫癫”且不可靠,却没有如今常见的破折号、三段式列表和“不是这个,而是那个”的句式;在她的测试中,GPT-3 模仿作家的说服力也强于 ChatGPT 5.4 Thinking。RLHF、脚本化对话、词语限制和人类偏好评分,把“疯癫又撞懵的模型”改造成了安全的企业助手。
更深层的创意写作瓶颈在于,艺术质量既无法被清晰验证,也无法建立在模型自身的人生经验之上。 一位写作评估员描述的评分标准会因为 3 个感叹号扣分,甚至按事实性给同人小说评分,说明适用于代码——代码要么能运行、要么不能——的方法,在主观艺术上会失效。LLM 可以生成打磨得很好的隐喻,但 Jasmine 认为,人类作者的作品若来自经历、观察或社群,就带有切身的利害关系;Casey 的反驳是,模型从未听过音乐,却仍能富有感染力地讨论音乐。
Token 使用量正在变成一种昂贵的地位指标,其激励可能跑在经济价值前面。 据称 OpenAI 过去 7 天用量榜首的员工消耗了 2100 亿 Token,约等于 33 个 Wikipedia;Kevin 听说 Anthropic 个人 Claude Code 用量最高的用户单月花费超过 15 万美元;一些公司如今甚至把消耗量纳入绩效评估。主持人援引 Goodhart 定律,以及那句把按代码行数衡量编程比作“按重量衡量飞机制造进度”的老话:一部分 tokenmaxxing 确实带来真实产出,但排行榜也会诱发浪费、个人项目、预算失控和人为制造的员工绑定。
1. AI 相关裁员是早期预警信号,并非清晰的因果实验
眼下的数字已相当可观:Atlassian 裁掉 10% 的员工,约 1,600 个岗位;Block 裁掉约 40%,即 4,000 人。Reuters 报道称,Meta 正准备裁掉 20% 或更多员工,潜在涉及 16,000 个岗位;但 Meta 称该报道为“推测性报道”,主持人强调,截至录音时并未发生裁员。
Atlassian 表示,裁员将为 AI 和企业销售提供资金;Block 称公司正转向更小、更扁平的团队;Meta 也公开拥抱一种更高强度使用 AI 的工作方式。Casey 的总体判断是:具体情况各不相同,但高管们一再指出 AI 很重要,“迟早,我确实认为我们不得不相信他们”。
Kevin 认为,科技从业者可能会成为最早一批受冲击者,因为他们的雇主既在开发这些工具,也在快速采用这些工具。但现有证据无法隔离出劳动力替代效应:股价下跌、疫情时期的招聘、战略重置和组织失灵,都与 AI 应用交织在一起。
Casey 从劳动者角度的反问切中了归因争论的要害:“如果对劳动者的影响是一样的,具体原因真的重要吗?”无论直接原因是自动化、人员过剩还是投资者压力,最终仍有数千人失去工作。
2. Atlassian 获得 AI 免检;Block 更像管理层在收拾残局
Atlassian CEO Mike Cannon-Brookes 表示,AI 并未取代员工,但如果否认技能结构或所需岗位数量发生变化,就“不诚实”。Casey 认为这相对坦率,在没有更清楚地知道哪些职能被裁掉之前,不愿把它称为 AI 洗白。
Atlassian 面临的更难问题,是主持人所说的“SaaSpocalypse”:它的商业工具把结构化工作流编码进产品,客户最终可能以低廉成本自行复现这些流程。股价已经遭受重创,裁员提供了一个新的市场叙事——员工更少、生产率更高——即便客户继续购买 Atlassian 产品时支付的价格更低。
Block 的历史让它对 AI 的解释更难令人信服。员工人数从 2019 年约 3,800 人增至 3 倍,而裁员前 5 个月,公司还花费 6,800 万美元,带 8,000 人飞去参加 Jay-Z 参与的活动;Casey 的结论是,AI “如果眯着眼看”或许可以为这次清理提供理由,但长期管理不善已经足以解释相当多的问题。
但市场仍然奖励了这个说法:Block 股价次日跳涨 17%。Kevin 将 AI 的叙事力量比作加密货币热潮——当时仅仅采用时髦话术就能推高股价;Casey 直白地总结:“公开市场确实可以这么容易就被欺骗。”
3. Meta 正以资本开支押注未经验证的“人力换算力”替代模式
Meta 可能裁员的消息,与今年计划中的 1,350 亿美元资本开支并列出现。Casey 认为,这一组合是在给投资者吃定心丸:管理层可以推进“公司历史上最大的押注”,同时释放出自己还没有“完全”失去费用控制的信号。
Zuckerberg 给出了生产率前提:“过去需要大团队的项目,如今一个非常有才华的人就能完成。”Kevin 更犀利的会计解读是,这些公司未必在总量上省钱——它们只是把支出从工资挪到了数据中心、模型、工具和 Token 上。
Kevin 听一位风险投资人说,一些高度 AI 原生的初创公司在 AI 工具上的支出已经超过工资支出。Kevin 称这可能是离群值,但也代表了高管们想象中的终点:大部分运营费用最终购买的将是机器劳动力,而非人类劳动力。
Casey 的保留意见是,Meta 尚未在全公司范围内证明这一生产率说法。它放弃了 Behemoth,据报道 Avocado 因仅略胜 Gemini 2.5 而被推迟,并再次重组 AI 团队;与此同时,OpenAI、Anthropic 和 Google——这些前沿模型开发者——并未大规模裁员。Kevin 还指出,OpenAI 和 Anthropic 的规模小得多,而一些正在裁员的公司可能是在试图借助 AI 追赶竞争对手。
4. AI 应用焦虑可能先于 AI 替代员工,起到管束作用
一位科技从业者描述了一个无解的选择:要么积极使用 AI,向管理层证明自己与方向一致;要么避免证明这份工作可以被自动化。Kevin 听到的是“角力、恐惧和焦虑”,而且在知道高管正主动规划裁员后,这些情绪被进一步放大。
Casey 不会断言,反复裁员是有意用来管住 Meta 员工,但他观察到裁员确实产生了这种效果。员工担心自己真的会丢掉工作后,内部抗议减少了;这支曾经愿意挑战管理层的队伍变得“安静得多”。
Kevin 猜测,这种压力是否最终会引发科技行业的大规模工会化。与没有工会的软件从业者不同,制造业工会历来会在岗位被自动化后协商调岗和再培训;Casey 的挑衅是:没有什么会比“Meta 软件工程师组成工会”更让 Zuckerberg 恼火。
5. 后训练让聊天机器人变得有用,却磨平了它们的声音
Jasmine 的论点比“人类写得更好”更窄:大多数文字本来就很差,而 LLM 在普通语言任务上胜过大多数人。她真正困惑的是,为什么行业领袖承诺模型能在编程和科学发现上达到超人水平,Sam Altman 却只谨慎地设想模型能写出“一个真正诗人的还算可以的诗”。
回看不同模型世代,Jasmine 更喜欢 GPT-2、尤其 GPT-3 的文风。这些系统可能会撒谎、跑题——主持人把 GPT-2 比作一个从楼梯上摔下来的人——但它们的语气更多变,也会给读者惊喜;GPT-3 还能模仿 Paul Graham 或 Richard Dawkins 等人的风格。
Jasmine 沿用旧的 GPT-3 风格提示词,让 ChatGPT 5.4 Thinking 重做测试,得到的结果被她称为“糟糕透顶”。相反,现代输出有一套可识别的口头禅:破折号、三段式列表、“不是这个,而是那个”,以及一种为企业助手场景优化的轻快能干腔。
她认为,机制在于后训练:实验室向基础模型喂入脚本化对话、行为限制、指定词汇,以及由人工评分者给出的 RLHF 评分。这些层把她所谓“疯狂、不可预测”和“疯癫又撞懵的模型”驯化了,却也把它们困在一种通用的“有求必应助手”人格里。
6. 创意质量让可验证奖励机制失灵
实验室承认,AI 研究者可能更懂什么是好代码,而不懂什么是好文章,因此 Mercor 和 xAI 等公司以每小时约 45 美元的价格招聘“创意写作专家”,有时还要求候选人拥有一本登上 New York Times 畅销榜的书或获得 Kirkus 星级书评。可专业能力并没有挽救荒谬的评估体系。
一名为 Scale AI 做写作评估员的承包商回忆称,回答中出现 3 个感叹号就会被扣分;他还被要求按事实性给同人小说评分——Jasmine 以此说明,资源充足的公司正试图把审美判断变成机械清单。
Kevin 将失败归因于可验证奖励:生成的代码可以测试,因为它要么能运行、要么不能;但没有评估员能稳定证明莎士比亚为何是莎士比亚,或 Neruda 的某首诗为何成功。用户需求又强化了这一结果,因为主流需求不是文学,而是“帮我写这封邮件”,在这种场景里,平庸反而有效。
Jasmine 的第二个解释是缺乏现实根基:模型可以写出“尝起来像差一点就是周五的临界日”这样的惊艳句子,但没有自己的人生为其提供切身利害、观察或视角;Casey 反驳说,模型从未听过音乐,却仍能出人意料地谈论音乐的感官质地,也许只是通过匹配听过音乐的人写下的文字模式。
7. 文本生成可以自动化;写作的其余环节仍然难以自动化
盲测可能显示读者更喜欢 AI 文章,直到得知其来源;Jasmine 也承认,人们会排斥明显的机器写作。但她质疑的是任务定义:她估计文本生成只占自己工作时间的 25%,采访、寻找选题、选择来源、报道,以及判断什么值得写,构成其余部分。
Kevin 提出“自我安慰”的反驳:写作者可能在重复工程师犯过的错误,把价值定义在模型尚未做到的事情上。Jasmine 的回答有实证依据,也保留限定:她已经尝试用 Claude 自动化自己 3 年但失败了,不过风格可能改善,微调可能有帮助,而她并没有说“永远不可能”。
类型小说既显示出进步,也暴露出限制。Sudowrite 联合创始人 James Yu 和其他从业者描述了大量工程工作,用来纠正后训练留下的轻快、谄媚、PG-13 倾向;Jasmine 把这种作者—模型协作称为“半人马模型”,因为人类必须不断提示、逼迫系统走向古怪与感官性。
她的判断带有明确条件:如果提供采访文字稿,且实验室像投资编程代理那样投入,模型最终可能写出强有力的特稿或文学小说。与“自动化 23 岁软件工程师”相比,她怀疑这在经济上并不划算;对于模型独立从现实世界采写,她仍然没那么乐观。
8. 个性化 Claude 编辑器有效,是因为它学会了一个作者的品味
Jasmine 的高效工作流并不是让 Claude 替她写作。她在一个 Claude 项目中载入自己的 Substack 档案、自由撰稿作品、发表后的复盘、受众、报道领域和目标,然后基于自己的写作追求而非“好写作”的通用标准,共同制定评判标准。
形成的评分框架识别出她的特征,例如她在硅谷的“圈内人类学家”位置、在创业公司行话和互联网俚语之间切换,以及从政策分析转向个人场景。她把审阅分成构思、结构、文风和最终事实核查几个阶段。
Claude 不替她编造素材,而是会指出总结式结尾很无聊,提醒她早先一篇文章以一个场景收尾得更有力,并追问飞机起飞时她有什么感受,或一段对话能否让干巴巴的政策变得有生命。判断权在她手里,但工具把她推向“作为写作者最好的自己”。
9. Tokenmaxxing 把 AI 采用与可量化生产率混为一谈
Token 是单词的片段,也是模型供应商计量消耗量的单位;大约 10,000 个 Token 可以生成 7,500 个单词。如今智能体编程一场会话就消耗数十万乃至数百万 Token,因为开发者运行的是更长、并发的进程,而不是交换一次提示和一次回复。
据称 OpenAI 过去 7 天用量榜首的员工消耗了 2100 亿 Token,约等于 33 个 Wikipedia,不过其中一部分是缓存 Token。Kevin 听说 Anthropic 个人 Claude Code 用量最高的用户单月花费超过 15 万美元;一位瑞典工程师告诉 Kevin,他在 Claude 上的花费超过自己的工资。
公司用排行榜激励实验、监测工程师是否采用了智能体编程;有些公司如今把 Token 消耗纳入绩效评估。实验室员工可能免费获得访问权限,但其他地方的员工可能把雇主预算用超。可 Goodhart 定律立刻生效:一旦把使用量设成目标,员工就能用无价值项目把数字做大;或者,正如有人猜测的那样,拿公司的算力做个人项目。
Casey 的历史类比很直接:用 Token 数量衡量编程,类似于用重量衡量飞机制造进度。Kevin 仍不愿把所有 tokenmaxxing 都视为作秀——一些重度用户可能确实更高效——但管理者应追问这些支出产出了什么,尤其是在至少一些营销绩效评估已开始计入 AI 使用评分、工程师开始向潜在雇主询问“我的 Token 预算是多少?”之际。
I just read the most heartwarming news this morning that I wanted to share with you, Kevin.
What's that?
The U.K. government has withdrawn a proposal to let A.I. companies train on copyrighted works after a backlash from artists like Dua Lipa. Did you see this?
No.
Dua Lipa said, “Don’t Start Now with this A.I.” My sugar boo, she’s litigating, Kevin. She’s making some new rules, and she’s saying, “We’re not gonna train on my copyrighted works.”
Wow.
And that’s why she is a queen. And so, Dua Lipa, if you’re listening, we salute you.
Yeah. Dua Lipa, you’re a Dua Keepa.
Yep. Period. Dua Lipa said artist rights.
Wow.
I'm Kevin Roose, a tech columnist at The New York Times.
I'm Casey Newton from Platformer.
And this is Hard Fork.
This week, a big wave of tech layoffs is raising the question: Has A.I. job loss truly begun? Then, writer Jasmine Sun is here to help us answer the question: Why are chatbots bad at writing? And finally, it’s token maxing time. Why are tech companies building leaderboards to measure who is spending the most on A.I.?
1. The AI Layoff Warning
Well, Casey, for years now, we’ve been monitoring for signs of an A.I. job apocalypse.
Yeah, we’ve been monitoring the situation.
It’s true. And over the past few weeks, I think we’ve gotten some early indications that something is happening in the labor market, especially for tech workers.
Yeah, we have certainly heard CEOs of companies announcing layoffs and invoking A.I. as a reason that it is happening, and so that has gotten our attention.
Yeah, so just a couple of examples from the last few weeks. Last week, Atlassian announced a 10 percent reduction in its staff, about 1,600 jobs, that they said were going to help them fund further investment in A.I. and enterprise sales. That came on the heels of a big round of layoffs at Block, the financial tech company formerly known as Square, which said that it was cutting its staff by about 40 percent, or about 4,000 jobs, saying that they were shifting the way that they were working to use smaller and flatter teams.
And then the big one that folks are expecting, maybe as soon as this week, is that Meta is reportedly poised to lay off 20 percent or more of the entire company. This was reported by Reuters last Friday, who said that their sources had told them that Meta was preparing to cut as many as 16,000 jobs, the largest layoffs at that company since late 2022 or early 2023, when they laid off 20,000 people. So as of this recording, that hasn’t happened yet, that we know of, but I know that people at Meta are very on edge and are awaiting further news about their jobs.
Meta, after this story came out, told Reuters that it was, quote, “speculative reporting.”
Which, if you’re not familiar with the language deployed by Meta communications staffers, means this is happening—but we don’t want to tell you it’s happening yet.
Correct.
So, Casey, I want to hear what you make of these layoffs, but first we should do our disclosures. I work for The New York Times, which is suing OpenAI, Microsoft and Perplexity.
And my fiancée works at Anthropic.
So, okay, Casey, what do you make of the fact that all these companies are referencing A.I. in some way as a reason for their layoffs?
Well, I think it’s a little different at each company, Kevin, and I think we can make a decent case for and against the idea that A.I. is really driving the show at each of them, so maybe we should get into that. But at the highest level, I would say companies do continue to tell us now that A.I. is a significant factor in the reduction of these workforces, and sooner or later, I do think we’re going to have to believe them.
Yeah, I think this is the early warning sign for a lot of people, especially in the tech industry, who are, I think it’s fair to say, going to be some of the first people to see their jobs change or disappear because of these new A.I. tools. But let’s get into some of the specifics here.
Yeah.
So, Casey, let’s start with Atlassian, the first company I mentioned. Their CEO, Mike Cannon-Brookes, said in a company blog post that the bar for what great looks like for software companies on growth, on profitability, on speed, on value creation, has gone up. He said, “We are choosing to adapt thoughtfully, decisively and quickly to drive durable, profitable growth.” He claimed that A.I. was not replacing people, but he said it would be disingenuous to pretend that A.I. doesn’t change the mix of skills we need or the number of roles required in certain areas.
Yeah, so I take him at his word. It seems like he himself is trying to walk a middle path there, right? And sort of not denying that A.I. is a factor here, but also not saying, “This is the only reason this is happening.”
I think some other context that is worth having is that Atlassian is one of the companies that could be part of what we’ve been calling the SaaSpocalypse around here, right? This is a company that makes tools for businesses. A lot of its products are essentially structured workflows, and there are those who believe that sooner or later, you’re just going to be able to code your own pretty cheaply.
Now, maybe you will still choose to buy a product from a company like Atlassian, but maybe you’re not going to be willing to pay nearly as much as you would have before. And so the company’s stock price has just been battered over the past year, and I think that has left them, one, hurting for cash a little bit, but two, and probably more importantly, looking for a different story that they can tell the stock market about what they’re doing. And so today that story is, “We’re gonna get rid of some of these workers, and we’re gonna figure out how to make our remaining workers more productive.”
Hmm. So there’s this term that’s been floating around called A.I. washing.
I thought it was when a software engineer finally took a shower.
And basically, the thesis is: These aren’t really layoffs about A.I. This is just sort of a convenient excuse that these companies are using.
Yeah.
Do you think Atlassian qualifies as A.I. washing?
I would like to get a little bit more detail on exactly who they are laying off here, which is a detail that we do have about some of these other companies that helps us answer that question. So I don’t know exactly how it is happening inside of Atlassian, but I think that their CEO was relatively straightforward, as these things go, in saying, “It’s a little bit about A.I., it’s not entirely about A.I.,” but, “Yes, keep your eye on A.I.”
So to me, that just reads as honest, and so I’m gonna give them a pass.
2. Block Shrinks Its Workforce
Okay. Let’s talk about Block. Jack Dorsey, the CEO of Block, gave an explanation about their layoffs. He said, quote, “We’re not making this decision because we’re in trouble. Our business is strong, but something has changed. I had two options: cut gradually over months or years as this shift plays out, or be honest about where we are and act on it now. I chose the latter.” Casey, your take.
So something to know about me and Jack Dorsey is I have a bit of a bias against him as a former Twitter user who misses that website dearly. At this point in 2026, I would not hire Jack Dorsey to run a lemonade stand. But if you want to talk about Block specifically, this is a company that tripled its headcount from about 3,800 people in 2019, in what seems like just classic inattention to what was happening in the business during pandemic-era boom times, right?
And I wonder if you saw this detail, because it truly took me out, Kevin. Five months before they announced the layoffs, Block spent $68 million to fly 8,000 people to an in-person event with Jay-Z.
Come on.
Yeah. So that’s the kind of famous attention to detail that has turned Jack Dorsey into one of the greatest visionaries in tech.
So look, is this about A.I.? Again, what does Block really do? They have those little iPads at the coffee shop—
Yeah.
—and then they have Cash App.
Mm-hmm.
Okay? How many people do you really need to run those products? Probably fewer than 10,000.
Hmm.
Is that about A.I.? I don’t know. Maybe if you squint. But again, this is a company whose stock price was cratering. They needed a different story to tell the market, and I do think you can make a case that A.I. will make the remaining workers more productive. So again, this is another one where it’s like, you could use A.I. to justify what’s happening, but you also could just say, “This company has been mismanaged for a while now.”
Yeah, you could use A.I. washing or Jay-Z washing, which seems to be what they are doing here.
Mm-hmm. Yes.
So this did seem to have an effect on their stock price. In fact, the day after Jack Dorsey announced the layoffs, Block’s stock shot up 17 percent. It’s gone down a little bit since then, but they’re still up from where they were before these layoffs.
And I think we should just say: This is also a part of the equation here, right? These are companies, largely public ones, that have investors’ attention. Right now, there’s this narrative power around AI: If you seem like a company that is investing heavily in AI tools and the AI way of working, your investors say, “Oh, that company is really forward-looking. They must have a plan for how to navigate this transition.” And so I think they’re seeing the power in telling the story that all this is related to AI.
Yeah, which, by the way, reminds me of the peak of crypto mania, when some publicly traded companies would just add a crypto term to their name, and their stock price would shoot up by about 40,000%.
Yes.
It turns out that the public markets actually can just be tricked that easily.
Yes.
That would give me some relief if I were a CEO, just knowing that I could fool people like that.
3. Meta Cuts For AI Infrastructure
So let’s talk about the third large tech company that is reportedly conducting layoffs: Meta. We don’t know exactly who or what teams are being affected by these layoffs, but this is a significant part of their workforce. They seem to be saying in their communications with the public what all of these other companies are saying: “We are going all in on the new way of working, and we are going to have to make some cuts to make that work.”
Yeah. On a recent earnings call, Mark Zuckerberg said, quote, “Projects that used to require big teams now can be accomplished by a single, very talented person.” We should also say that this cut is coming alongside this massive AI infrastructure investment, right? They’re going to spend $135 billion on capital expenditures this year. And even for a company of Meta’s size, that is real money. I know they’re trying to be careful not to spook the stock markets too much. This is obviously the biggest bet in the company’s history, and I think making some substantial cuts is going to signal to the market, “Hey, don’t worry. We’re not completely losing our minds here. We’re going to keep some of these expenses under control.”
Yeah, I think that’s a really important point, because what we’re seeing here at some of these companies is that they are not actually cutting costs in the aggregate by using these tools. They are just shifting the cost from human labor to AI.
Right.
They are plowing this money that they are going to save by laying off these thousands of people into the building of data centers and other AI infrastructure. Basically, the bet they’re making is that these new AI workers are going to be faster, more efficient, maybe cheaper in the long run, maybe not, but they are going to be able to do the work that used to require many thousands of people. And that is a profound shift in the way that companies are talking about their workers.
I recently talked to a venture capitalist who said that a lot of the AI startups that he sees, the most AI-native companies, are spending more on AI tools than they are on payroll. That may be an outlier, but I think that is where these companies believe that we are headed, where the majority of your expenses will not go to paying the salaries of human workers. It will go toward buying the AI tools and the tokens that your company runs on.
Yes, I think that’s absolutely the bet that they’re making. I also think it is worth noting that this is still mostly speculative, right? In the case of Meta specifically, this is a company that has arguably been struggling when it comes to AI. They had to abandon their last model, Behemoth, because it wasn’t very good. The Times reported last week that it’s delaying the release of its latest model, Avocado, because it hasn’t been hitting its performance targets. It’s apparently barely outperformed Gemini 2.5. What is this, last March?
Yeah, that model is really the pits.
That’s an Avocado joke.
That’s very good. Thank you.
So, again, this is not as simple as saying they’re able to cut 20% of their workforce because they’ve just made these massive gains. I’m sure there are individuals there who have made massive gains, but as a company, it still seems like it is somewhat mired in dysfunction. They just did yet another partial reorganization of their AI teams, and that always makes me raise my eyebrows.
Yeah. I will say, one thing that’s been surprising to me about this recent round of layoffs is that the companies that are making them are not the ones on the frontier, right? It is not the OpenAIs, the Anthropics, or the Googles. Those companies are not laying off people en masse because of these AI tools, which they are building and presumably have even better models than the ones they’re releasing to the public. So you have to think that part of this is just companies that are lagging behind their competition saying, “Well, maybe if we just use a bunch of AI, it’ll help us catch up.”
Yes, but also OpenAI and Anthropic are much smaller companies than some of the ones that we’ve been talking about today, at least in number of workers, right? I think it is interesting to think that Atlassian is bigger than OpenAI in terms of the number of people who work there, when you look at the relative value of what they’re generating.
DocuSign has 7,000 employees.
There’s no funnier sentence that is true in all of tech journalism. As somebody who has a paid subscription for DocuSign that I truly resent paying for, get to work over there, people.
Or get not to work.
Get not to work.
4. Workers Consider Union Power
Here’s another question that I would ask, Kevin. We’re seeing a bunch of layoffs. Are these AI-related or not? Does it actually matter if the effect on workers is the same, right? If you’re the worker, whether it’s about AI or not, you’re still out of a job.
Yeah, and it’s not clear to me what workers can or should be doing to protect themselves against these layoffs. One person I talked to said they work at one of these big tech companies, and they’re like, “Well, there’s just a lot of jostling and fear and anxiety right now. People don’t know if they should be using the AI tools a ton because then it shows that they’re getting with the program, or whether that just means that they’re proving that their work can be automated.”
I think there’s a lot of fear, suspicion, and mistrust inside these companies right now, and for good reason. Their executives are planning to lay them off.
Yes, and by the way, I think at least at some of these companies, that may not be an explicit reason for these layoffs, but some of the executives there would see that as a positive byproduct, right? Because if you’re Mark Zuckerberg, you lived through the 2020 era. You had these restive employees who wanted a lot of things from you, and they wanted to have a lot of control over what the company could and could not do and how it did it.
I know that executives over there really resented that sort of thing. And once Meta entered this new era of massive layoffs, employees over there did get really scared for all of the reasons that you would assume. They were like, “Oh, God, maybe I actually am going to lose my job.” All of a sudden, they got a lot quieter, and you started to see a lot fewer protests over there.
So I’m not going to say that these occasional mass layoffs are a way of keeping the workforce in line, but I have noticed that it seems to be having that effect.
Totally. And it makes me wonder whether something that I predicted was going to happen a year or two ago, but did not happen—the sudden and mass unionization of workers at these companies—may actually start to happen in the next year or two.
I think one major difference between what’s happening now at these tech companies and what has been happening for decades at manufacturing companies and car companies, among factory workers, is that those workers were by and large unionized. And so when the employers said, “Hey, we’re going to lay a bunch of you off,” they were able to negotiate. They were able to say, “Hey, maybe instead of laying us all off, maybe you could find other jobs for us. If our jobs are being automated, maybe we should be allowed to retrain to do something else.”
And that was largely successful. There were still layoffs, of course, but not the number that we’re seeing today at these tech companies. So do you think there’s any possibility of that, or is that just a union fever dream?
Here’s what I will say: I cannot think of anything that would make Mark Zuckerberg more mad than a union of software engineers at Meta. And I think the software engineers at Meta should use that information how they will.
You think that would make him more mad than getting booed at a UFC fight?
Absolutely. I think that probably just made him really sad.
Well, there you have it. If you want to make Mark Zuckerberg mad, Meta employees, sign your union card.
When we come back, why aren’t chatbots as good at writing as I am?
We’ll ask Jasmine Sun.
5. LLMs Still Struggle With Writing
Well, Casey, over the last couple of years, we've talked on this show about how AI models are getting better at so many things. They are getting better at coding, at competition math, at solving novel physics problems.
Mass domestic surveillance—autonomous weapons.
Yes. And I think the story of the last few years in AI has been one of sort of rapid, steady progress, but these systems are still sort of jagged, and they have flaws and weaknesses. And one place where they arguably haven't improved that much is in writing.
Now, that's our domain.
Yes. At least that is the argument that Jasmine Sun made in The Atlantic this week. She is a freelance journalist. Her piece was called “The Human Skill That Eludes AI,” and it's her attempt to understand why, despite so much progress in all these different areas, the models of today don't seem to be writing anything particularly good or compelling.
Yeah. And while I think the question of whether LLMs are good at writing is highly subjective and dependent on the use case, I do think Jasmine makes a really interesting technical case for why these models write the way they do.
Yes. And we should say, before we bring her in, Jasmine is a friend of mine. She has also been my researcher on the upcoming book that I'm working on, and I just think she's one of the best people writing about AI today. She writes on her Substack, which is called Jasmine News. It's J-A-S-M-I dot news, and you can read much more of her writing there.
All right. I'll allow it, but I do want to balance it out. By next week, bring me on one of your enemies.
Okay, let's bring her in. Jasmine Sun, welcome to Hard Fork.
Thanks for having me. I'm excited.
Hi, Jasmine.
So you wrote this great piece in The Atlantic this week about the human skill that eludes AI, and I want to start by challenging the subtitle of your piece. Why can't language models write well? Can't language models write well?
So I do say in the piece that most writing, period, is very bad, and so I think that language models are definitely better at writing and language than most humans are. But the question that I was really curious about is, why can't they write at a sort of literary, creative-fiction level?
Because the thing is, if you listen to these AI leaders talk about their aspirations, they say, “We're gonna cure cancer. We're gonna solve physics. We're gonna build a superhuman coder.” They are not shy about saying, “Oh, our AI models are gonna be better than 75 percent of human coders.” They're saying, “No, we will literally build a self-replicating factory tomorrow.”
And then Tyler Cowen asked Sam Altman in an interview from last October, “When do you think GPT will be able to write a Neruda poem?” And Sam Altman says, “Maybe in the future, ChatGPT will be able to write, quote, ‘a real poet's okay poem.’” So that was the thing that fascinated me: Even these guys who are more bullish than anybody else about the capabilities of their technology, they are very reserved about how much literary writing their models can do.
Mm.
And so that was the gap that I was really interested in.
Hmm.
Hmm.
6. GPT Two Had More Voice
And you start your piece with this interesting provocation, which is that, in some ways, GPT-2 was the peak of AI when it comes to creative writing. So explain that.
Part of what got me interested in this piece was I was actually doing research for your book. I was going through all of these previous generations of models and reading the outputs, and the thing that really shocked me was that, in a way, the writing style of GPT-2 and GPT-3 was so much more compelling to me than ChatGPT today.
It doesn't have any of the annoying tics. It doesn't have the em dashes, the tripartite lists, the “it's not this but that.” The tone was much more variable. It would actually surprise you. It would be funny. It would be poetic. And that shocked me, to go back a few generations and realize that maybe they were also lying all the time and all sorts of other things. But from a writing-style perspective, I preferred it, and I wanted to investigate that.
They were weird.
That shocks me. To me, talking to GPT-2 was like talking to somebody who had just fallen down the stairs. You know what I mean? It was like, “Do I need to get you to the hospital? Do you smell toast?”
Yeah, there are these amazing prompts from this early OpenAI prompt library where they would say, “I just won $175,000 in Las Vegas. What do I need to know about taxes?” And GPT-2 would start just writing some short story about an orphanage.
But, yeah, they were surprising.
Yes.
They were nutty. They were weird. They would absolutely be a terrible corporate assistant, a horrible coding intern. It can't do any of the things that modern LLMs can do that I'm very grateful for. But from a pure writing-style perspective, they were very good—GPT-3 in particular.
I found this set of samples that some guy did where it was, “Oh, write in the style of Paul Graham. Write in the style of Richard Dawkins,” whatever, and it could style-match much better than modern LLMs can. And particularly because so much of literary writing comes from voice and style, one of the things I was really interested in was: What did we lose? The LLMs can no longer emulate Paul Graham's style or whoever's style.
Because I would put in the same exact prompt that this guy gave GPT-3 into ChatGPT 5.4 Thinking or whatever, and it would be god-awful.
Hmm.
And I was like, “That's really weird.”
So tell us about what you learned about what happened after the GPT-2 and GPT-3 era that changed the way that these models respond to us.
Yeah, I think the answer is post-training, basically. So they started adding a post-training layer, which is basically saying: We have these crazy, unpredictable, nut-job, concussed models, and they need to learn how to behave because a model that can't behave is a very bad corporate assistant.
And so the AI researchers give them example dialogues and scripts to learn from. They give them words that they can and can't say. They do RLHF, which is a process by which human graders will rate which response is the most helpful-sounding or something like this.
And so now these post-trained models have been trapped, in a way, or trained or guided toward a very particular character or persona that is a very helpful assistant, but might be very bad at writing in creative and surprising ways.
Mm.
I mean, the way that you described it was that there is a phase within the post-training phase where these AI models are evaluated by humans. And that's part of what they call RLHF, or reinforcement learning from human feedback.
And what struck me in your reporting is that you actually talked to some people who have done this kind of feedback, who say that they're just being asked to grade things in ways that don't make sense. Right? Tell us about that.
Yeah, this is super interesting because these job listings you'll see on places like Mercor or xAI, Elon’s company, will list them directly. It'll be like, “Creative writing expert, $45 an hour. Must be a New York Times bestseller,” and have a starred Kirkus review or something like this.
Have you ever gotten a starred Kirkus review, Roos?
I think so.
Okay, good job.
Not sure.
All right.
You might qualify to help Elon—
Yeah.
—to help Annie from Grok write a little bit better.
Yeah, we're gonna get on that job listing. But okay, you were saying.
Yeah, so these companies realize that these AI researchers are really good at knowing what good coding is, but they don't actually know what good writing is, so they're like, “Why don't we hire some humans to find out?”
And so they'll commission MFAs, published authors, and sometimes just random guys with a blog or whatever.
And one of the people I talked to, who was a contractor for Scale AI as a writing evaluator and was doing this for one of the bigger labs, said that the rubric just didn’t make any sense. He would be told things like, “You have to grade them based on the number of exclamation marks there are. If something has 3 exclamation marks, that’s too many, and so you have to ding that one.”
Yeah, and I have to say, generally not bad writing advice.
Yeah.
I guess it depends on the length of the text, but 3 feels like a lot for many scenarios.
This is what they tell women in business communications. It’s like, “Take all those exclamation marks, replace them with periods. We’re just gonna remove all of the exclamation points.”
We teach women to shrink themselves.
Exactly.
Yeah.
He was being asked to grade these things. In another case, he got a bunch of fan fiction, and he was supposed to grade it on its factuality, since that was one of the criteria. I do imagine that one could devise better rubrics than this particular evaluator was given, but I think it does show, at least, that some of these very big companies that are very well-resourced simply do not know how to think about what good writing is.
Briefly, I want to underline that, because to me, that seems like the whole story. We are taking the entire internet, and we are grading it on factuality. So the LLM that you’re gonna get out of that is just probably not gonna be all that creative. And I wonder how much of it is related to this sort of verifiable reward—
Mm-hmm.
—system that a lot of these companies are using, where you have a system generate a bunch of code, and then you have another evaluator model check the code to see whether it’s good or not. That works in domains like programming, where the code either runs or it doesn’t, but creative writing doesn’t work that way. You can’t have an evaluator tell you, with any sort of consistency, whether something is good or not, and so it may just come down to preference. So I guess I’m curious: Do you see this as a technical problem that the labs are frustrated trying to solve, or is this just demand-related? Is this just what people want chatbots to sound like, and in every test where they pit different models against one another, the one that sounds like a bland corporate assistant wins, and so they go with that?
I think both are true. The majority of writing that we are asking the models to do is, “Write this email for me,” right? And they excel at that. They are truly great corporate email writers. They are much better at the whole passive-aggressive thing than I am.
At the same time, I do think, like you said, there is a technical challenge that has to do largely with verifiability. There are people who have spent decades of their lives attempting to articulate what makes Shakespeare Shakespeare, or what makes a Neruda poem a Neruda poem, and they will still not know in any kind of certain way. They will still get into debates with their fellow academics and literary critics about which writer is better than the other, because these things are subjective, because they are ineffable, because they are hard to put in a rubric, and that is the nature of art.
And to that point, you started this segment by talking about Sam Altman saying, “Hey, we just basically can’t write a great poem yet.” Sam Altman, a year ago, said the company had trained a good creative-writing model and posted a short story on X. Many people found it compelling. Is Sam Altman just not being consistently candid with us, Jasmine?
Ooh. Wouldn’t be the first time. But that short story, if you remember, had some great lines, like talking about the seams of mirrors or Thursday, the… What was it?
It was the liminal almost-Friday or something.
Yeah, the liminal day that tastes of almost-Friday.
Wait, I have to actually look this one up—
It’s so good.
—because it was so good.
While you’re looking it up, the thing about AI writing is that it comes up with all of these fun metaphors, and those metaphors are sometimes surprising, but the language is not grounded in life. That was my other thing: Aside from the verifiability, fundamentally, when I think about the writers who I really love—whether it’s journalists or poets or whatever—they are writing from life, right? A journalist goes out and talks to people, and they see stuff and observe the color of the sky in a particular way, or a poet is thinking about personal experiences that they’ve had.
Their writing has stakes. It comes from an emotional place. And the fact that LLMs, while being very talented and grammatically pristine or whatever, don’t have lives means that all of the metaphors they choose, all of the words they choose, and the examples they choose are just ungrounded, right? It’s not coming from a point of view, or a particular experience, or a particular community that makes the writing believable. I think part of what voice and style are is that they are very specific to the life that a person has had, and LLMs cannot get there in the same way a human who hasn’t really lived that life cannot get there.
I don’t know. I feel like it’s case-dependent. I’m a big music fan, and over the past few months, I have enjoyed putting questions about music, and in particular the sounds of certain bands, to an LLM, which sounds like a joke prompt because an LLM has never heard anything—
Mm.
—right? And yet I find that, in general, the models can have good conversations with me about the sound of music. Now, it may be that they are just pattern-matching based on a bunch of public writing on the internet by people who do have ears—
—and have heard, right? I’m very open to that.
Yeah.
But, again, I have just been struck by the way that it is able to write about sensory topics in an evocative way that, at least to me, surpasses what I would predict they would be able to do.
Yeah. I want to pose a couple objections that I think—
Okay.
—someone might make to—
Perfect.
—your article. One of them is: This is cope. This is Jasmine, a writer, a very talented writer, sort of finding the things that AI, in her view, is not good at yet and saying, “This is categorical proof that it will be very hard for AI to do these things.” This is the same reaction that software engineers had when models started getting really good at code. They would say, “Oh, well, it can’t do these other 10 things that I do,” and then, basically, just wait a few years, and the models will be better than all of us at everything, including writing.
I would love for it to be cope, because I try to automate myself away all the time. I have no deep attachment to having to do it. I like writing, but I have tried over and over and over for the past 3 years to automate my own job away and to get Claude to do my job for me. It cannot do it. This is very frustrating.
Mm-hmm.
It’s not out of a lack of trying. Again, I’m going back to the CEOs themselves and the things that they themselves are saying, right? It’s not just me, a writer; it’s Sam Altman saying, “This thing will cure cancer and solve physics, but it will not write better than a real poet’s okay poem.” And so I think that suggests that there is something that is at least perceived as a little bit different.
I think it’s very possible that the models will get much better at writing over the next few years. I don’t think it’s a never thing. I do think that reporting is hard to replicate. I think that having life experiences that are real and verifiable is hard to replicate. I think the style stuff can be improved, especially if you fine-tune the models. But I think what’s also interesting to me about this piece is that it shows how the market incentives, the demand incentives of these companies, do shape what we see as their abilities today.
Mm.
The other objection I’m imagining people might have, who are very AI-pilled, is—
Mm.
—that this is all in the eye of the beholder, right?
Mm-hmm.
There have been several studies now that have shown that if you give people a blind taste test of AI writing versus human writing, they prefer the AI writing until you tell them that it’s AI writing, and then the value in their eyes plummets. I did one of these in a New York Times quiz just recently. So is it possible that the models have already become superhuman at writing, but that the minute we learn that they are AI models generating text and not humans writing words with their fingers, we lose all interest in it just because of the source, not because of the quality of the writing?
I mean, I think it’s definitely interesting and true that people don’t want to like AI writing, and that is part of what bothers them when they see AI text that is obviously AI, even though, as you said, in these quizzes and tests, AI can outperform human writers in those narrow scenarios.
My quibble with a lot of these quizzes and tests is that, as a writer—and you guys are writers too—how much of your job is actually text generation? I think AI is a superhuman text generator, right?
Mm-hmm.
In my job, I am generating text probably 25% of the hours in my day. I spend a lot of time interviewing people. I spend a lot of time coming up with ideas. I spend a lot of time reading, and not just reading indiscriminately, but reading very particular sources that feel like the right ones.
Usually, at the point that you are doing one of these tests, you're saying, “Generate one paragraph very specifically about why Trump won the 2016 election, 500 words or less.” You've already given the prompt, which I think is a critical part of writing: What are you going to write about? You've often supplied some of the evidence and the guidance in the form of saying, “500 words or less,” and at that point, I do think that AI is probably a better text generator than almost all humans are.
But again, when I think about it, AI is still very bad at coming up with ideas for articles. It is still very bad at reporting. The non-text-generation parts of the role feel further away from automation. Again, I'm sort of a “never say never” person. Maybe it'll get there. I would be totally happy if Claude was able to give me good ideas for my next essays, but it's not there yet.
Well, we're already seeing LLMs make huge progress in genre fiction, right? Recently on the show, we talked to the author of a story in The Times about how authors of romance novels are now able to generate dozens of novels a year using LLMs. In fact, much of the discussion that we had was around how you just have to prompt them differently and relentlessly in order to get what you want.
Your piece, Jasmine, made me wonder: How much of getting a model to just write weird can be achieved by repeatedly telling it, in different ways, “Hey, be a little weirder”?
Some of it, but not all of it. I talked to, for example, James Yu, who is the co-founder of Sudowrite, which is one of the earliest creative-fiction AI writing assistants. I talked to some other folks who similarly were in the fiction-writing LLM space.
And like you said, to an extent, a lot of writers are already using these, already leaning on LLMs to generate large amounts of text, and it can be very successful, and it can meet readers' needs and whatever. But even these people who I was talking to were describing to me how freaking hard it is to undo all of the post-training that the labs have done.
They were applying immense amounts of engineering effort, which, in my conversations with them, clearly frustrated them, because it is so hard to get these models to stop being so chirpy, so sycophantic, so PG-13 and everything, in order to get them to this sort of base-model state where they're able to be weird again. So I think it's certainly possible, but I think the labs have made it quite challenging just because of the way that these models are trained.
The other thing that I think is important is that I tend to think that writing and a lot of creative work is actually the perfect use case for these centaur models, right? The idea that the human-plus-AI collaboration is where you can get the furthest. And when I listen to the interviews that you guys did about the fiction authors, I was thinking, “This is a centaur model,” right?
Without the human prompting and bullying the AI into getting weird and getting sensual and whatever, it was not going to do that on its own. I myself do use LLMs as a research assistant. I wrote about that inside The Atlantic piece about the way that Claude has now helped me edit my own work in a way that I found incredibly useful. But I do feel like the collaborative element is important for any domain where the personal perspective, lived experience, whatever, really matters.
Talk about that a little bit. You mentioned your editing process. How are you using AI to help you edit your work, and are you finding it useful?
Yeah, I feel like I really cracked this over the last couple months, which I'm very excited about. Because, again, I've tried to make these things write and edit for me over and over and over, and they've never really been able to do it.
So the thing that I realized was, if I make Claude into an editor that is not just trying to grade and give feedback on my work against some genericized standard of what good writing is, but actually does it against basically what my personal aspirations for writing are, it can give feedback that I find much, much more helpful.
So what I did was basically feed Claude my entire Substack archive of the writing that I've previously done, as well as some of my freelance work.
And just to get real specific, is this inside a Claude project, or how have you set this up? Because I know our listeners are going to want to try this.
Yes. I did it in a project—
Okay.
But on Claude's advice. I was like, “Do I need to Claude Code something?” And Claude was like, “No, that's overkill.”
Okay.
You don't need to code or anything. So, in a Claude project, I gave it my whole archive of writing. I also personally write retro notes to myself after everything I publish. So I have a notes app that's just me writing what was good and bad about everything I've ever written.
Hmm.
Just a few bullet points.
This is why Jasmine's gonna be our boss.
For sure. These are very low-quality bullet points, but I also gave it that because I wanted it to learn my taste. I wanted it to learn: What do I aspire to be? Where do I see myself falling short? And what am I proud of, right?
And so from those two things, plus a little bit more information about, “Here's my audience. This is my beat. These are my goals,” we were able to co-develop a rubric. Instead of asking, “How many exclamation marks does it have?” it would say things like, “Does this take advantage of your, quote-unquote, ‘insider anthropologist position in Silicon Valley?’” That's one of the things that Claude and I think distinguish my voice.
Or it'll also notice, “Oh, Jasmine, you tend to move between registers. You'll switch between startup jargon and internet slang and whatever. And I think the fact that you can do the hi-lo or move from policy to a personal scene is something that is characteristic of your writing.” And so again, we're co-developing these qualitative criteria.
Then I split it into phases: ideation phase, structure rubric, prose rubric, and final fact-checking. What I do now is put this all in a Claude project. I said, “Your job is to evaluate my drafts based on these criteria, but not to do the writing for me, and to make sure to prompt out of me what I can do better.”
I dump the draft into Claude. Claude will run phase-two structure on it. It'll say things like, “Your conclusion is just a summary, and this is really boring. In fact, in your piece about this and that, you actually ended on a scene, and I thought that was much more powerful, so why don't you try ending this one on a scene?”
And Claude will say, rather than inventing a scene, “What were you thinking when the plane took off? What were you feeling inside? Can you think of a scenario where you had a conversation with, say, a kid-safety advocate about AI that really resonated with you? Because right now it sounds like a dry policy explainer.” And that feedback I actually found incredibly useful.
Hmm. It is.
I'm still applying my own judgment to say, “Do I take it or not?” But this is about me becoming the best version of myself as a writer. It's about me self-improving and Claude pushing me to do that, which I found much, much more helpful.
Hmm.
Wow.
I want to ask you both a question as fellow writers. Do you feel the impulse to make your writing weirder because of AI to sort of stand out from the sea of slop? Because I find myself feeling this tug of, “Oh, that's a little weird aside that probably I should cut, but I think I'm gonna leave it in, because Claude would never do that,” right?
Mm.
It's like a marker that I am typing these words, and I feel like that's sort of my imprimatur that I'm leaving. My answer to you is yes, I absolutely feel that way, and I've gone back and tried to edit sentences to make them feel a little bit weirder or, in particular, to make them sound colloquial in a way that I know an LLM generally would not. And yes, it is for that reason.
I think that writing right now, we're all—not all, many of us are—on such high alert for the prospect that we might be reading slop that if you are a writer who does not want to be producing slop, you should be asking yourself that question.
Mm-hmm.
I think it makes me a lot more comfortable writing the way I want to write in the first place. I think maybe, unlike both of you, I didn't sort of come up through newsrooms where I was learning a very specific house style and all of these norms.
I can do news writing now. It’s something I’ve learned now, but I’m actually much more, quote-unquote, Internet- and blogging-native, which is a form that is voicey and irreverent and not as pristine, and will make inappropriate jokes. It’s just a looser form of writing. And so I think what it’s actually done is made me more comfortable doing the bloggy thing instead of always trying to write in a more professionalized journalistic tone.
Hmm. So I think we should leave this with a question for you, Jasmine, which is: Your piece makes the case very convincingly that today’s AIs are not very good at the kind of writing that I think we all value. Do you think they will get there, and what should the companies do to make their models better at writing?
I think that if we separate out text generation from reporting, which I’m not that bullish on the models doing, and we’re just talking about, say, literary fiction, or “Here’s a bunch of interview transcripts; write a magazine feature” or something, I think that if they applied as many resources toward that task as they do toward coding agents and things that actually make the money, I think that they could get there.
Will the companies ever find it financially advisable to spend all their resources on that instead of automating 23-year-old software engineers? Probably not. I would be grateful for that world. I don’t need them to take my job or these folks’ jobs, but I think it’s possible.
Look, they’re going to get around to it eventually. Okay? I hear what you’re saying—
Have you seen what writers make in this economy, Casey?
Eventually, like, they—
Those aren’t going to pay for a lot of data centers.
No, there is economic value in writing. And eventually, the AI companies will want that all to themselves.
You know what would be a very funny outcome of this, taking your point about the sort of guardrails of the models: Maybe the next great American novel will be written by Grok.
Oh, God.
And with that, Jasmine Sang, thank you for joining us. Thank you, Jasmine.
Thank you very much, Kevin and Casey.
Well, Kevin, you’ve recently returned from book leave and are once again writing in The New York Times. How does it feel to see your name in print again?
Feels great. It hasn’t happened yet, but when it does, it’ll be great.
Well, I got to take an early read at a story that you are publishing about the fact that tech companies have now created leaderboards to show which employees are using the most AI tokens in their work.
7. Tokenmaxxing Becomes A Workplace Metric
Yes, it’s a token frenzy out there, and the employees of these companies are competing among their colleagues, informally and for fun, but they’re taking it very seriously. They want to be the people at their company who are using the most AI tokens.
So let me just ask a basic question for listeners who may not be familiar. What is a token, and why is that something you might start keeping track of?
So a token is the basic atomic unit of AI labor. It’s basically a fragment of a word, and it is how AI model providers measure their consumption. So if you type in a prompt, “Help me write this essay,” an old model might have given you a couple hundred tokens in response. That would be a couple hundred words.
What has been happening over the past year or so, as these agentic coding tools have started taking off, is that the models are just much more token-hungry. You can use now hundreds of thousands or even millions of tokens in a single session, and so that is what is propelling these leaderboards: the idea that the more coding you’re doing, the more agentic tools you’re using, the more simultaneous processes you’re running, the higher your token count will be.
One measurement I found useful was that apparently it takes about 10,000 tokens to generate 7,500 words, if that helps to ground you at all. But as you just said, and I want to hear more about this, the more advanced systems are using way more tokens than that. So tell me about some of the numbers that some of the token all-stars are putting up on the boards.
So I don’t know all of the exact numbers, but I did learn that at OpenAI, where they do track this kind of leaderboard, the highest employee token count over a 7-day period recently was a guy who used 210 billion tokens. And this is, for rough scale, about 33 Wikipedias’ worth of text.
Hmm.
Now, all of that is not typing and receiving a response. Some of that is what they call cached tokens. So it’s not all being extruded from the model for the first time. But these are the kinds of numbers that I think even a year ago would have sounded completely insane.
Right. Now, is this guy working on a new mass domestic-surveillance program for the Department of Defense?
I don’t know, and OpenAI did not make him available for interviews.
Oh.
But what I wanted to do in writing this column was to try to call up a bunch of people or talk to a bunch of people who are in this billion-token club, the extreme power users, and just ask them, “Hey, how are you guys using all those tokens, and isn’t that very expensive, and how are you paying for it all?” And I learned a lot.
Yeah. Well, okay. So tell us, first of all, just how expensive it is.
Very expensive.
Yeah.
In fact, I heard that the top user of Claude Code, the top individual user of Claude Code, as measured by Anthropic, spent more than $150,000 on tokens last month. So extrapolate that. That is like an employee making more than $1 million a year.
They are burning that in a month, and I heard similar figures from some of these other extreme coders who are spending something on the order of thousands of dollars a day on tokens from these models. Now, we should also say the employees of these companies get their tokens for free, right?
Right.
So they are not shelling out; their companies are not shelling out. But at other companies, this is starting to become an issue because they are outstripping their budgets for these things.
So there are companies where there are engineers who legitimately are costing their employers maybe $150,000 a week because they’re getting tokens from one of the big providers.
Yeah, I talked to a software engineer in Sweden who said that he probably spends more than his salary on Claude. So this is essentially becoming a very expensive job perk for some of these coders.
8. Leaderboards Create Perverse Incentives
So talk to me about why employers want to create leaderboards to promote this to employees, because I could see other companies saying, “If you spent $150,000 on tokens last month, you actually don’t work at this company anymore, because we’re bankrupt.”
Right. So this was a big question that I had: Why is this going on? And it seems to be some combination of employee motivation and worker tracking, right? There are executives at these companies who think that the more tokens you use, the more productive you probably are.
And as we discussed in a previous segment on this show, these companies are very eager to have their workers start embracing the AI tools. And so at a number of these companies, I talked to people who said, “Yeah, this is just basically them trying to see who is really all in on the new way of programming.”
And you’ve talked to a number of people who are ranking high on these leaderboards. I realize you probably haven’t dug deep into their code, but what is your sense of how productive they actually are? What is the relationship between token usage and taking my company to the next level?
I mean, it’s very unclear, right? Some of these people may be just generating worthless projects.
I think the thing that worries a lot of the people I talk to about these leaderboards is that they just incentivize you to run up your token count, right?
Yes.
Because then you look like the special 10X engineer or 100X engineer who’s outperforming all your colleagues. So I think there are a number of companies that see this leaderboard business as a little strange and maybe counterproductive. But I do think that there is a feeling among the most heavy token users that they are being productive.
Yeah. I have to say, when I read your column, I thought, this just seems like it would create the worst incentives, right?
Yes.
There’s this idea of Goodhart’s law, right? When a measure becomes a target, it ceases to become a good measure. I can’t think of a better way to ensure that token usage becomes a bad measure than creating a leaderboard for it.
Totally.
What are the people inside the company saying about that?
Well, some of them are opposed to this whole leaderboard thing. I also talked with some folks who defended the leaderboards. They said, “Look, it’s never been all that easy to track the productivity of programmers.”
Some people have had their productivity measured by how many lines of code they generate or how many pull requests they made. These are imperfect proxies for how hard you’re working and how much you’re doing.
But the employees of these companies also see this, I think wisely, as a key to their own success. A number of these companies are now using A.I. token use and consumption as part of the performance review cycle. So you go in for your annual review, and your boss says, “Hey, it looks like you only used 70 million tokens last month. What’s going on?”
I think the engineers at these companies are getting wise to the fact that if they want to have a long, successful career, they better start using some tokens.
Yeah, but I imagine that some of them are really nervous about that, though, right? Because it seems clear to me that at least some of these companies want to incentivize token usage because the companies themselves suspect that the more we can get them using this stuff, the less long we will have to employ the humans.
Maybe, although I think it’s less about the A.I. systems replacing the humans and more about it being a radically different way of working, right?
Mm-hmm.
These are people who, most of them, have had long careers in software engineering. They grew up writing code by hand. They maybe grew up using some sort of A.I. assistant, like GitHub Copilot, and what people at these companies are saying is that these agentic engineering systems are just really different.
You have to approach them in a different way, and you have to spend a lot of time with them to understand what they’re good and not good at. To them, this is a way of motivating their employees to say, “Hey, go out and try the new thing.”
Yeah. I don’t know. I’ve been thinking a lot about this question of, if I were an engineer at one of these companies and I had this incentive to get on the leaderboard, how would I approach it? I do think that the instinct to waste a bunch of tokens to rise higher on the leaderboard could ultimately backfire. If you rise too high, people are going to ask you what you did with all the tokens.
Right.
If you’re number 1 at 10 billion tokens and you only managed to vibe-code a calculator or something, people are probably going to get mad at you.
Yeah, and I actually did talk to one person who speculated that the people at the top of the leaderboards are all doing side projects. They’re starting their—
They’re starting a new company.
—their side hustles.
They start a new company with the boss’s money. And if you’re doing that, I just want to say I salute you.
Yeah.
That is the right way to work in 2026.
Yeah. Maybe don’t be number 1 on the leaderboard if you’re doing that. Maybe try to stick around 6 or 7.
Yeah, like middle of the pack—
Yeah.
—is kind of where you want to aim yourself. I mean, let me ask: Is there any kind of token tracking that you think offers a reasonable signal? Do you think that if you’re a tech company, you should create a leaderboard?
No. I think that’s a bad idea for all the reasons that we just talked about, including—
Yeah.
—Goodhart’s law, which is that I think this is just going to lead to people wasting tokens and doing side projects. But if I’m the budget manager at a company and I’m seeing that people are spending multiples of their salary on A.I. tokens, I’m asking them some questions about what they’re doing with all that. If their answer is not, “I built an amazing new product that’s going to generate billions of dollars a year in revenue,” I’m trying to say, “Hey, could you maybe use a little less next month?”
Yeah. I have to say, I have been struck by how this idea of the token leaderboard represents a new incarnation of something that the software industry has been trying to figure out for a long time, which is: How can I figure out if my software engineers are productive?
I was talking recently to this very handsome software engineer who I’m engaged to about your column. He was telling me that he used to be evaluated on how many lines of code he contributed, and he told me about all the games that people used to play back in the day. “Oh, I wrote a quick algorithm to translate a bunch of stuff into some new languages, and it’s completely worthless, but it makes me look like I had a very productive week.”
And so I went back and looked into this, and they were doing this in the ’60s and ’70s. There’s this saying from the early days of computer programming that says, quote, “Measuring programming progress by lines of code is like measuring aircraft-building progress by weight.”
I have to say, I think the same thing kind of applies here, right? If you squint and look at it at the right level of abstraction, it’s probably true that some people who are using a lot of tokens are more productive than some people who aren’t. It just doesn’t quite seem like the right way to measure these things, and I just wonder how quickly the industry is going to figure that out.
Yeah. I think it’s going to be pretty soon, in part because the budgets are just getting very ridiculous. And especially the A.I. model providers are now seeing individual users consuming amounts of their services that entire companies would have consumed just a few months ago.
You know, maybe the last question I have for you about this is: What implications do you think it has for the broader economy, right? Because we know that in so many different sectors of the economy, managers are saying, “I want to incentivize my employees to use A.I., and I want to track how they’re using A.I.”
Do you think that as knowledge of these leaderboards spreads, we’re going to see people in nontechnical fields try to adopt their own version of them?
I hope not. I think it’s really a bad move, not just for tracking actual productivity and output, but just for morale, right?
I remember years ago when Gawker would have a traffic leaderboard at its office, so you could see how many clicks your stories were getting relative to other people. I don’t think anyone who worked there at the time thought that was incentivizing the right things or creating high morale among employees. Basically, everyone was just competing with each other all the time.
And I think in this case it’s even worse because it’s not necessarily even correlated with any success.
Mm-hmm.
It’s just pure, sort of, how many agents can you run in a parallel swarm to work 24/7 doing tasks of uncertain value?
Which is a great question to ask on a first date in San Francisco, too, by the way.
But anyway, I have to say, I worry that this idea of tokenmaxxing is going to spread into the broader economy. I was talking with somebody who works in marketing this week, and she was telling me that her job used to be evaluated solely on creativity. Then recently, the performance review got a new A.I. section, and everyone is being evaluated on how much A.I. they used.
From her perspective, she was like, “This was working fine. I didn’t need to use an A.I. tool to help me, but now my bonus might be based on how much of it I use.”
So I think this thing has already seeped out of the labs and is getting into the water elsewhere. I just hope that managers are really thoughtful about what they are incentivizing, and that maybe A.I. use for the sake of A.I. use is not going to be the boon to your company that you’re hoping it is.
Yeah. I think it’s going to be very case by case. I think there will be people who are tokenmaxxing who are way more productive than their colleagues and doing way more projects way more quickly. I think there will be other people whose managers look at their token budgets and say, “You spent this many tokens on what?” and will have to have some hard conversations.
But I think it’s very hard to draw with a broad brush and say, “All of this tokenmaxxing is pointless productivity theater.” It sounds to me, from my conversations, like some of it really is working for people.
Yeah. Well, on the flip side, I’ve also heard of people in my social circle who have gotten in trouble for spending too much on Claude in their life.
Wait, really?
Yeah. When I heard that, I was like, “Oh, your company’s not going to make it, bro. You’ve got to spend on this stuff.”
Well, what’s so interesting is that now it’s becoming part of job conversations for engineering jobs. People are going into new jobs and saying, “Well, what’s my token budget?” And for the employees of these big AI labs who have unlimited free access to the models, some of them are using so many tokens that they effectively can’t afford to quit their jobs, right? Because anywhere else they would work would have to pay for their tokens, and it would be completely unaffordable to employ them.
Yeah. I mean, those sound like real incentives, and better than the ones at Meta. Do you remember when Meta was spinning up superintelligence labs and they said, “You can sit really close to Mark Zuckerberg”? If I were them, I’d be like, “I’ll take the tokens, thanks.”
All right. Well, just to wrap this up, exactly how many tokens should a person use?
I think that’s something you have to look within yourself for.
Look within yourself?
Yeah.
Okay.
Yeah.
That’s between you and your God.
Yeah.
Yeah.
Do what Marc Andreessen will not: introspect.
Introspect.