[BidClub_]
Hard Fork · · 59 min

Why Tech C.E.O.s Are Blaming A.I. for Mass Layoffs

Ross DouthatAndrew Miller

YouTube
TL;DR
  • Tech’s latest cuts are an early warning, but not yet clean proof that AI has directly replaced workers. Atlassian is eliminating 10%, roughly 1,600 jobs; Block about 40%, or 4,000; and Meta was reportedly preparing cuts of 20% or more, potentially 16,000. Casey Newton’s calibration is the useful one: AI differs in causal importance at each company, yet “sooner or later, I do think we’re going to have to believe them.”

  • The hosts disclose relevant conflicts: Kevin Roose works for The New York Times, which is suing OpenAI, Microsoft, and Perplexity; Casey Newton’s fiancé works at Anthropic.

  • For public companies under pressure, AI offers both a productivity thesis and a marketable explanation for old-fashioned restructuring. Block had expanded from roughly 3,800 employees in 2019 to more than 10,000, then spent $68 million flying 8,000 people to an event with Jay-Z five months before its cuts; its shares rose 17% the next day. That makes “AI-washing” difficult to separate from correcting pandemic-era overhiring and management failures.

  • Kevin argues that Meta’s proposed labor cuts could shift costs from payroll to AI infrastructure rather than reduce aggregate spending. The company plans $135 billion of capital expenditure this year while Zuckerberg argues that projects once requiring large teams can now be completed by “a single very talented person.” Yet the productivity case remains speculative: Meta abandoned Behemoth, reportedly delayed Avocado for missing targets, and apparently achieved only a slight improvement over Gemini 2.5.

  • Employees are being placed in a no-win adoption test: use AI heavily to demonstrate alignment, or avoid proving that their work can be automated. The hosts argue that repeated layoffs also quiet internal dissent, whether or not workforce discipline is an explicit objective. Kevin Roose consequently revives a prediction he previously got wrong: fear could drive tech workers toward unionization and bargaining over retraining or reassignment.

  • Chatbots’ mediocre literary voice may be a consequence of product optimization, not an underlying inability to generate surprising prose. Jasmine Sun preferred aspects of GPT-2 and GPT-3 because they were variable, strange, and capable of style matching, whereas post-training and RLHF pushed later systems toward the “helpful assistant”: chirpy, sycophantic, safe, and repetitive. The commercial demand is for excellent corporate emails, while artistic quality lacks the verifiable rewards that accelerated coding.

  • The defensible near-term writing product is a personalized collaborator, not an autonomous author. Sun says text generation occupies only about 25% of her work; reporting, idea selection, reading, judgment, and lived experience remain harder to reproduce. Her productive Claude workflow instead learns her archive, retrospective notes, audience, and aspirations, then asks questions that push her toward “the best version of myself as a writer.”

  • Token consumption is becoming a new and potentially enormous component of technical labor costs. OpenAI’s top employee reportedly used 210 billion tokens in seven days—about “33 Wikipedias” of text—while Anthropic’s highest individual Claude Code user spent more than $150,000 in one month. One Swedish engineer said he probably spends more on Claude than his salary, turning unlimited access into both a job perk and a retention mechanism.

  • Token leaderboards confuse adoption with output and invite classic Goodhart’s-law gaming. Companies are incorporating usage into performance reviews—an employee might be challenged for using “only” 70 million tokens—despite no clear relationship between consumption and value. The investor-relevant question is therefore not who is token maxing, but whether escalating inference spend produces products, revenue, or merely “tasks of uncertain value.”

Digest · the substance, structured for research

1. AI has entered the layoff rationale before proving the substitution case

  • Kevin Roose opens with the scale: Atlassian is cutting about 1,600 jobs, Block roughly 4,000, and Meta was reportedly preparing its largest reduction since the 20,000 jobs eliminated in late 2022 and early 2023. Meta called the report “speculative,” and the cuts remained unconfirmed at recording.

  • The hosts disclose relevant conflicts: Kevin works for The New York Times, which is suing OpenAI, Microsoft, and Perplexity; Casey Newton’s fiancé works at Anthropic.

  • Casey resists one explanation for all three companies, but his highest-level conclusion is firmer: executives keep identifying AI as a significant workforce factor, “and sooner or later, I do think we’re going to have to believe them.”

  • Kevin treats tech as an early-warning market because its workers will be among the first to see jobs change or disappear. Casey’s sharper worker-level point: whether AI caused a dismissal “does it actually matter if the effect on workers is the same?”

2. Atlassian and Block expose two different versions of AI-washing

  • Atlassian CEO Mike Cannon-Brookes said AI was not simply replacing people, but that it would be “disingenuous” to deny changes in the skills mix or number of roles required. His stated objective is to adapt “thoughtfully, decisively, and quickly” for durable, profitable growth.

  • Casey’s framing places Atlassian inside the “SaaSpocalypse”: customers may eventually build structured workflows cheaply rather than pay established software vendors as much. With the stock battered, layoffs create a new market story—fewer workers, higher productivity—though Casey gives Cannon-Brookes a pass for acknowledging that AI is only part of the explanation.

  • Block looks less clean. Headcount rose from about 3,800 in 2019 to more than 10,000, and five months before eliminating roughly 40% of staff, the company spent $68 million flying 8,000 people to an event with Jay-Z. Casey’s verdict: AI may raise remaining-worker productivity, but “you also could just say this company has been mismanaged.”

  • The market nevertheless rewarded Jack Dorsey’s story: Block shares jumped 17% the day after the announcement. Kevin sees narrative power in presenting cuts as forward-looking AI adaptation; Casey compares it with peak crypto mania, while observing that “the public markets actually can’t just be tricked that easily.”

3. Meta is transferring spending from payroll to compute

  • Reuters reported that Meta could eliminate 20% or more of its workforce, potentially 16,000 jobs. The proposed cuts sit beside the company’s planned $135 billion of capital expenditure this year—“real money,” even at Meta’s scale—and may reassure investors that its largest-ever bet has some expense discipline.

  • Zuckerberg supplied the productivity premise: “Projects that used to require big teams now can be accomplished by a single very talented person.” Kevin’s interpretation is not that technology lowers total costs today, but that companies are shifting expenditure “from human labor to AI.”

  • A venture capitalist told Kevin that some especially AI-native startups already spend more on AI tools than payroll. The destination these companies imagine is one where salaries no longer dominate expenses and businesses instead purchase the models, infrastructure, and tokens on which their work runs.

  • Casey’s pushback is execution risk: Meta abandoned Behemoth because it “wasn’t very good,” reportedly delayed Avocado after missing performance targets, and apparently barely surpassed Gemini 2.5 before another partial AI reorganization. Frontier developers such as OpenAI and Anthropic are not making comparable mass cuts, although Casey notes that they employ far fewer people.

4. AI adoption has become an internal loyalty test

  • One big-tech employee described a genuine trap: heavy AI use might prove enthusiasm for the new program, but it might also demonstrate that the employee’s work can be automated. Kevin hears “fear and suspicion and mistrust” because workers know executives are planning reductions.

  • Casey observes that Meta’s earlier layoffs made employees quieter and reduced internal protests. He stops short of calling periodic cuts an intentional workforce-control mechanism, but says some executives would regard that effect as “a positive byproduct.”

  • Kevin revisits his failed prediction of sudden mass unionization and wonders whether it could happen in the next year or two. Unlike largely unionized manufacturing workers who bargained over retraining and reassignment, tech employees lack that union bargaining channel; Casey’s advice is pointed: nothing would make Zuckerberg angrier than “a union of software engineers at Meta.”

5. Post-training traded the old models’ weirdness for corporate usefulness

  • Jasmine Sun distinguishes competent language from literary writing: most human writing is bad, and models outperform many people, but even maximalist AI leaders remain cautious about art. Asked when GPT might write a Neruda poem, Sam Altman offered only the possibility of “a real poet’s okay poem.”

  • Sun found GPT-2 and GPT-3 more compelling stylistically than current ChatGPT. They lied, wandered, and could be unusable assistants—GPT-2 might answer a tax question with a story about an orphanage—but they were “surprising” and “nutty.” GPT-3 in particular was better at matching voices such as Paul Graham’s than ChatGPT 5.4 Thinking.

  • The loss, in her account, came through post-training. Example dialogues, prohibited language, and RLHF steered unpredictable base models toward a consistent helpful-assistant persona: excellent for office work, but constrained when creative prose depends on variable tone, surprise, and risk.

6. Literary quality does not fit the rewards that made coding improve

  • The evaluation machinery can become absurdly reductive. Listings might offer a creative-writing expert $45 an hour while requiring a New York Times bestseller and starred Kirkus review; one Scale AI contractor was instructed to penalize three exclamation marks and assess fan fiction for factuality.

  • Kevin calls the evaluator problem “the whole story”: “We are taking the entire internet and grading it on factuality.” Sun adds that code can be tested by whether it runs, while experts can debate for decades over what makes Shakespeare or Neruda good; subjective art cannot be reduced to a consistently verifiable reward.

  • Demand reinforces the technical bias. Most users want “write this email for me,” a task at which models excel, and preference tests reward the bland corporate assistant. Sun therefore thinks both propositions are true: labs face a hard evaluation problem, and the market actively selects the resulting voice.

  • Sun’s deeper objection is that model language is “not grounded in a life.” Journalists observe scenes and interview people; poets write from emotionally consequential experience. Casey pushes back with models’ evocative writing about music despite never hearing it, while conceding they may simply recombine criticism written by people with ears.

7. AI can generate prose, but writing remains a larger system of judgment

  • Kevin poses the “cope” objection: programmers once listed everything models could not do, only to watch the gap narrow. Sun says she has spent three years trying to automate herself with Claude and failed, though she explicitly allows that style and literary generation might improve substantially.

  • Blind tests complicate the claim because readers sometimes prefer AI prose until told its source. Sun’s answer is occupational: text generation takes perhaps 25% of her day; the rest includes finding ideas, interviewing, selecting particular sources, reporting, and deciding what deserves to be written.

  • Genre fiction shows both capability and constraint. Sudowrite co-founder James Yu and other practitioners described the engineering effort required to undo models’ chirpy, sycophantic, PG-13 post-training. Human authors must keep prompting and “bullying the AI into getting weird,” making the successful arrangement a Centaur rather than autonomous authorship.

  • Sun’s best workflow makes Claude a personalized editor. She loaded a project with her archive, freelance work, post-publication notes, audience, beat, and goals, then co-developed separate ideation, structure, prose, and fact-checking rubrics while instructing it to evaluate—not write—her drafts.

8. Personalized evaluation turns Claude into a demanding collaborator

  • The useful feedback is specific to Sun’s aspirations: Claude identifies her “insider anthropologist” position in Silicon Valley and her movement between startup jargon, internet slang, policy, and personal scenes. That is categorically different from counting exclamation marks against a generic standard.

  • Instead of fabricating an ending, Claude might say her conclusion merely summarizes, recall that another piece ended more powerfully on a scene, and ask: “What were you thinking when the plane took off?” Sun retains judgment over whether to act, describing the objective as becoming “the best version of myself as a writer.”

  • All three writers feel pressure to preserve odd, colloquial, or blog-native lines as proof of human presence amid “slop.” Sun says AI has made her more comfortable with her loose, irreverent internet voice rather than pushing her toward professionalized newsroom prose.

  • Her final forecast is conditional: if labs devoted writing-level resources comparable to coding agents, they could plausibly produce strong literary text or turn transcripts into features. Whether that beats “automating 23-year-old software engineers” financially is doubtful; the hosts’ darkly comic alternative is that Grok writes the next great American novel.

9. Token maxing is becoming a costly proxy for modern engineering

  • Kevin defines a token as “the basic atomic unit of AI labor,” roughly a word fragment. About 10,000 tokens can generate 7,500 words, but agentic coding sessions now consume hundreds of thousands or millions as engineers run longer and more numerous processes.

  • OpenAI’s highest employee total over one recent seven-day period was reportedly 210 billion tokens—roughly “33 Wikipedias” of text, though some were cached rather than newly generated. Kevin’s conversations focused on this emerging “billion-token club.”

  • The expense is already salary-scale. Anthropic’s top individual Claude Code user reportedly spent more than $150,000 in one month; other extreme users burn thousands of dollars daily, and a Swedish engineer said he probably spends more on Claude than his salary.

  • Unlimited internal access becomes a powerful perk: some AI-lab employees use so much that another employer could not afford their habits, effectively making them costly to recruit. Engineering candidates are consequently starting to ask, “What’s my token budget?”

10. Leaderboards turn an imperfect signal into a gameable target

  • Employers use leaderboards for motivation and tracking, assuming heavier token users are adopting agentic engineering more seriously. Some now incorporate consumption into performance reviews, potentially creating conversations like: “It looks like you only used, you know, 70 million tokens last month. What’s going on?”

  • Productivity remains unproven. Heavy users may complete far more projects, but they may also generate worthless work; one person speculated that people at the top could be building side companies with their employers’ tokens. Casey invokes Goodhart’s law, and Kevin says he would not create a leaderboard at all.

  • The historical analogy is lines of code, another proxy engineers learned to game. Casey quotes the old comparison: “Measuring programming progress by lines of code is like measuring aircraft building progress by weight.” Token volume may correlate loosely with output while still failing as an individual target.

  • The practice is already escaping engineering: a marketer said her performance review now has an AI-use section and that her bonus might be based on how much AI she uses, despite creativity previously being the objective. Kevin rejects calling all token maxing theater, but Casey’s warning stands: “AI use for the sake of AI use” may produce expense, rivalry, and 24/7 agent swarms doing “tasks of uncertain value.”

Kevin Roose

I'm Kevin Roose, a tech columnist at The New York Times.

Casey Newton

I'm Casey Newton from Platformer.

Kevin Roose

And this is Hard Fork. This week, a big wave of tech layoffs is raising the question: Has AI job loss truly begun? Then, writer Jasmine Sun is here to help us answer the question: Why are chatbots bad at writing? And finally, it's token maxing time. Why tech companies are building leaderboards to measure who is spending the most on AI?

Well, Casey, for years now, we've been monitoring for signs of an AI job apocalypse.

Casey Newton

Yeah, we've been monitoring the situation.

Kevin Roose

It's true. And over the past few weeks, I think we've gotten some early indications that something is happening in the labor market, especially for tech workers.

Casey Newton

Yeah, we have certainly heard CEOs of companies announcing layoffs and invoking AI as a reason that it's happening, and so that has gotten our attention.

Kevin Roose

Yeah. So, just a couple of examples from the last few weeks. Last week, Atlassian announced a 10% reduction in its staff, about 1,600 jobs, that they said would help them fund further investment in AI and enterprise sales. That came on the heels of a big round of layoffs at Block, the financial tech company formerly known as Square, which said that it was cutting its staff by about 40%, or about 4,000 jobs, saying that it was shifting the way it worked to use smaller and flatter teams.

And then the big one that folks are expecting, maybe as soon as this week, is that Meta is reportedly poised to lay off 20% or more of the entire company. This was reported by Reuters last Friday, which said that its sources had told them that Meta was preparing to cut as many as 16,000 jobs—the largest layoffs at that company since late 2022 or early 2023, when it laid off 20,000 people.

As of this recording, that hasn't happened yet, as far as we know, but I know that people at Meta are very on edge and are awaiting further news about their jobs.

Casey Newton

Meta, after this story came out, told Reuters that it was, quote, “speculative reporting.” Which, if you're not familiar with the language deployed by Meta communications staffers, means this is happening, but we don't want to tell you it's happening yet.

Kevin Roose

Correct.

So, Casey, I want to hear what you make of these layoffs, but first we should do our disclosures. I work for The New York Times, which is suing OpenAI, Microsoft, and Perplexity.

Casey Newton

And my fiancé works at Anthropic.

Kevin Roose

So, okay, Casey, what do you make of the fact that all these companies are referencing AI in some way as a reason for their layoffs?

Casey Newton

Well, I think it's a little different at each company, Kevin, and I think we can make a decent case for and against the idea that AI is really driving the show at each of them. So maybe we should get into that. But at the highest level, I would say companies continue to tell us that AI is a significant factor in the reduction of these workforces, and sooner or later, I do think we're going to have to believe them.

Kevin Roose

Yeah, I think there are probably some complications here, and we should get into them. But I think this is the early warning sign for a lot of people, especially in the tech industry, who are, I think it's fair to say, going to be some of the first people to see their jobs change or disappear because of these new AI tools.

But let's get into some of the specifics here. So let's start with Atlassian, the first company I mentioned. Their CEO, Mike Cannon-Brookes, said in a company blog post that the bar for what great looks like for software companies—on growth, on profitability, on speed, on value creation—has gone up. He said, “We are choosing to adapt thoughtfully, decisively, and quickly to drive durable, profitable growth.”

He claimed that AI was not replacing people, but he said it would be disingenuous to pretend that AI doesn't change the mix of skills we need or the number of roles required in certain areas.

Casey Newton

Yeah, so I take him at his word. It seems like he himself is trying to walk a middle path there, right? Not denying that AI is a factor here, but also not saying this is the only reason this is happening.

I think some other context that is worth having is that Atlassian is one of the companies that could be part of what we've been calling the SaaSpocalypse around here, right? This is a company that makes tools for businesses. A lot of its products are essentially structured workflows, and there are those who believe that sooner or later, you're just going to be able to code your own pretty cheaply.

Now, maybe you will still choose to buy a product from a company like Atlassian, but maybe you're not going to be willing to pay nearly as much as you would have before. And so the company's stock price has just been battered over the past year, and I think that has left them hurting for cash a little bit and, probably more importantly, looking for a different story that they can tell the stock market about what they're doing.

And so today that story is: “We're going to get rid of some of these workers, and we're going to figure out how to make our remaining workers more productive.”

Kevin Roose

Hmm. So there's this term that's been floating around called AI-washing, which is basically when a company wants to lay a bunch of people off, or maybe doesn't feel like it needs as many people.

Casey Newton

I thought it was when a software engineer finally took a shower.

[laughter]

Kevin Roose

And basically, the thesis is that these aren't really layoffs about AI. This is just a convenient excuse that these companies are using. Do you think Atlassian qualifies as AI-washing?

Casey Newton

I would like to get a little bit more detail on exactly who they are laying off here, which is a detail that we do have about some of these other companies that helps us answer that question. So I don't know exactly how it's happening inside Atlassian, but I think that their CEO was relatively straightforward, as these things go, in saying that it's a little bit about AI, it's not entirely about AI, but yes, keep your eye on AI.

So to me, that just reads as honest, and I'm going to give them a pass.

Kevin Roose

Okay, let's talk about Block. Jack Dorsey, the CEO of Block, gave an explanation about their layoffs. He said, quote, “We're not making this decision because we're in trouble. Our business is strong, but something has changed. I had 2 options: cut gradually over months or years as this shift plays out, or be honest about where we are and act on it now. I chose the latter.”

Casey, your take.

Casey Newton

Something to know about me and Jack Dorsey is that I have a bit of a bias against him as a former Twitter user who misses that website dearly. At this point in 2026, I would not hire Jack Dorsey to run a lemonade stand.

[laughter]

But if you want to talk about Block specifically, this is a company that tripled its headcount from about 3,800 people in 2019 to more than 10,000 today, in what seems like classic inattention to what was happening in the business during pandemic-era boom times, right?

And I wonder if you saw this detail, because it truly took me out, Kevin. 5 months before the layoffs, Block spent $68 million to fly 8,000 people to an in-person event with Jay-Z.

Kevin Roose

Come on.

Casey Newton

Yeah. So that's the kind of famous attention to detail that has turned Jack Dorsey into one of the greatest visionaries in tech.

[laughter]

So, look, is this about AI? Again, what does Block really do? They have those little iPads at the coffee shop, and then they have Cash App, okay? How many people do you really need to run those products? Probably fewer than 10,000.

Kevin Roose

Mhm.

Casey Newton

Is that about AI? I don't know. Maybe if you squint. But again, this is a company whose stock price was cratering. They needed a different story to tell the market, and I do think you can make a case that AI will make the remaining workers more productive.

So again, this is another one where you could use AI to justify what's happening, but you also could just say this company has been mismanaged for a while now.

Kevin Roose

Yeah, you could use AI-washing or Jay-Z-washing, which seems to be what they are doing here.

Casey Newton

Yes.

Kevin Roose

So this did seem to have an effect on their stock price. In fact, the day after Jack Dorsey announced the layoffs, Block's stock shot up 17%. It's gone down a little bit since then, but they're still up from where they were before these layoffs.

And I think we should just say this is also a part of the equation here, right? These are companies, largely public ones, that have investors' attention. And right now, there's sort of this narrative power around AI, where if you seem like a company that is investing heavily in AI tools and the AI way of working, your investors say, “Oh, that company is really forward-looking. They must have a plan for how to navigate this transition.”

And so I think they're seeing the power in telling the story that all of this is related to AI.

Casey Newton

Yeah, which, by the way, reminds me of the peak of cryptomania, when some publicly traded companies would just add a crypto term to their name and their stock price would shoot up by, like, 40,000%. It turns out that the public markets actually can't just be tricked that easily.

Kevin Roose

Yes. That would give me some relief if I were a CEO, just knowing that I could fool people like that. But anyways—

Casey Newton

Totally. And I don't think everyone is falling for it. Aaron Zamost, the former head of communications for what was then Square, wrote a guest essay in The New York Times recently, basically saying these job cuts aren't what they seem. This is just a company that wants to show that it is adapting to the new reality.

But then also, it’s just a vision that everyone has to get on board with or decide that they’re not comfortable with and leave.

Kevin Roose

Yeah, I have to say: What do we even think the future of Block is? Again, they make a little widget you can plug into your phone so that you can sell bracelets at the craft fair, and then it makes a PayPal competitor. What are we using AI to do here exactly? I don’t know what the roadmap is.

These people are not building a metaverse. They don’t appear to be building a machine god. So I truly do not know what they’re going to be up to over there. And if you know, don’t email me. [Laughter.] Don’t email Casey. Just Cash App him instead.

Casey Newton

Yeah, there you go.

Kevin Roose

So let’s talk about the third large tech company that’s reportedly conducting layoffs: Meta. We don’t know exactly who or what teams are being affected by these layoffs, but this is a significant part of their workforce. They seem to be saying in their communications with the public what all of these other companies are saying: “We are going all in on the new way of working, and we’re going to have to make some cuts to make that work.”

Casey Newton

Yeah, on a recent earnings call, Mark Zuckerberg said that, quote, “Projects that used to require big teams now can be accomplished by a single very talented person.” We should also say that this cut is coming alongside this massive AI infrastructure investment, right? They’re going to spend $135 billion on capital expenditures this year. And even for a company of Meta’s size, that is real money, right?

I know that they’re trying to be careful, again, trying not to spook the stock markets too much. This is obviously the biggest bet in the company’s history, and I think that making some substantial cuts is going to signal to the market, “Hey, don’t worry. We’re not completely losing our minds here. We’re going to keep some of these expenses under control.”

Kevin Roose

Yeah, I think that’s a really important point, because what we’re seeing here at some of these companies is that they’re not actually cutting costs in the aggregate by using these tools. They’re just shifting the cost from human labor to AI.

Casey Newton

Right.

Kevin Roose

They are plowing this money that they are going to save by laying off these thousands of people into the building of data centers and other AI infrastructure. Basically, the bet they’re making is that these new AI workers are going to be faster, more efficient, maybe cheaper in the long run, maybe not. But they are going to be able to do the work that used to require many thousands of people.

And that is a profound shift in the way that companies are talking about their workers. I recently talked to a venture capitalist who said that a lot of the AI startups that he sees, the most AI-native companies, are spending more on AI tools than they are on payroll. That may be an outlier, but I think that is where these companies believe we are headed: where the majority of your expenses will not go to paying the salaries of human workers. It will go toward buying the AI tools and the tokens that your company runs on.

Casey Newton

Yes, I think that’s absolutely the bet they’re making. I also just think it is worth noting that this is still mostly speculative, right? In the case of Meta specifically, this is a company that has arguably been struggling when it comes to AI. They had to abandon their last model, Behemoth, because it wasn’t very good.

The Times reported last week that it’s delaying the release of its latest model, Avocado, because it hasn’t been hitting its performance targets. It’s apparently barely outperformed Gemini 2.5. What is this, last March?

Kevin Roose

Yeah, that model is really the pits.

Casey Newton

So—

Kevin Roose

That’s an avocado joke.

Casey Newton

That’s very good. Thank you. So, again, this is not as simple as saying they’re able to cut 20% of their workforce because they’ve just made these massive gains. I’m sure there are individuals there who have made massive gains, but as a company, it still seems like it is somewhat mired in dysfunction.

They just did yet another partial reorg of their AI teams, and that always makes me raise my eyebrows.

Kevin Roose

Yeah, I will say one thing that’s been surprising to me about this recent round of layoffs is that the companies making them are not the ones on the frontier, right? It is not the OpenAIs, the Anthropics, or the Googles. Those companies are not laying off people en masse because of these AI tools, which they are building and presumably have even better models than the ones they’re releasing to the public.

So you have to think that part of this is just companies that are lagging behind their competition saying, “Well, maybe if we just use a bunch of AI, it’ll help us catch up.”

Casey Newton

Yes, but also, OpenAI and Anthropic are much smaller companies than some of the ones we’ve been talking about today, at least in number of workers, right? I think it is interesting to think that Atlassian is bigger than OpenAI in terms of the number of people who work there when you look at the relative value of what they’re generating.

Kevin Roose

DocuSign has 7,000 employees.

Casey Newton

There’s no funnier sentence that is true in all of tech journalism. [Laughter.] As somebody who has a paid subscription for DocuSign that I truly resent paying for: Get to work over there, people.

Kevin Roose

Or get not to work.

Casey Newton

Get not to work. Here’s another question that I would ask Kevin. So we’re seeing a bunch of layoffs. Are these AI-related or not? Does it actually matter if the effect on workers is the same, right? If you’re the worker, whether it’s about AI or not, you’re still out of a job.

Kevin Roose

Yeah, and it’s not clear to me what workers can or should be doing to protect themselves against these layoffs. One person I talked to—they work at one of these big tech companies—said, “Well, there’s just a lot of jostling and fear and anxiety right now. People don’t know if they should be using the AI tools a ton because then it shows that they’re getting with the program, or whether that just means that they’re proving that their work can be automated.”

I think there’s a lot of fear and suspicion and mistrust inside these companies right now, and for good reason. Their executives are planning to lay them off.

Casey Newton

Yes, and by the way, I think at least some of these companies, that is maybe not an explicit reason for these layoffs, but some of the executives there would see that as a positive byproduct, right? Because if you’re Mark Zuckerberg, you lived through the 2020 era. You had these restive employees who wanted a lot of things from you, and they wanted to have a lot of control over what the company could and could not do and how it did it. I just know that executives over there really resented that sort of thing.

Once Meta entered this new era of massive layoffs, employees over there did get really scared for all of the reasons that you would assume. They’re like, “Oh, God, maybe I actually am going to lose my job.” All of a sudden, they got a lot quieter, and you started to see a lot fewer protests over there.

I’m not going to say that these occasional mass layoffs are a way of keeping the workforce in line, but I have noticed that it seems to be having that effect.

Kevin Roose

Totally. And it makes me wonder whether something that I predicted would happen a year or 2 ago, but did not—the sudden and mass unionization of workers at these companies—may actually start to happen in the next year or 2.

I think one major difference between what’s happening now at these tech companies and what has been happening for decades at manufacturing companies, car companies, and factories is that those workers were by and large unionized. So when the employers said, “Hey, we’re going to lay a bunch of you off,” they were able to negotiate. They were able to say, “Hey, maybe instead of laying us all off, maybe you could find other jobs for us. If our jobs are being automated, maybe we should be allowed to retrain to do something else.”

And that was largely successful. There were still layoffs, of course, but not the number that we’re seeing today at these tech companies. So do you think there’s any possibility of that, or is that just a union fever dream?

Casey Newton

Here’s what I will say: I cannot think of anything that would make Mark Zuckerberg more mad than a union of software engineers at Meta, and I think the software engineers at Meta should use that information however they will.

Kevin Roose

You think that would make him more mad than getting booed at a UFC fight?

Casey Newton

Absolutely. [Laughter.] I think I probably just made him really sad.

Kevin Roose

Well, there you have it. If you want to make Mark Zuckerberg mad at employees, sign your union card.

Well, Casey, over the last couple of years, we’ve talked on this show about how AI models are getting better at so many things. They are getting better at coding, at competition math, and at solving novel physics problems—

Casey Newton

Mass domestic surveillance, autonomous weapons. [Laughter.]

Kevin Roose

And I think the story of the last few years in AI has been one of rapid, steady progress. But these systems are still jagged, and they have flaws and weaknesses. One place where they arguably haven’t improved that much is in writing.

Casey Newton

Now, that’s our domain.

Kevin Roose

Yes. At least that is the argument that Jasmine Sun made in The Atlantic this week. She is a freelance journalist.

Her piece was called “The Human Skill That Eludes AI,” and it’s her attempt to understand why, despite so much progress in all these different areas, the models of today don’t seem to be writing anything particularly good or compelling.

Casey Newton

Yeah. And while I think the question of whether LLMs are good at writing is highly subjective and dependent on the use case, I do think Jasmine makes a really interesting technical case for why these models write the way they do.

Kevin Roose

Yes. And we should say, before we bring her in, Jasmine is a friend of mine. She has also been my researcher on the upcoming book that I’m working on, and I just think she’s one of the best people writing about AI today. She writes on her Substack, which is called Jasmine News. It's j a s m i dot news and you can read much more of her writing there.

Casey Newton

All right, I’ll allow it, but I do want to balance it out by next week bringing on one of your enemies. Okay, let’s bring her in. Jasmine Sun, welcome to Hard Fork.

Jasmine Sun

Thanks for having me. I’m excited.

Kevin Roose

Hi, Jasmine. So, you wrote this great piece in The Atlantic this week about the human skill that eludes AI, and I want to start by challenging the subtitle of your piece.

Jasmine Sun

Mm-hmm.

Kevin Roose

Why can’t language models write well?

Jasmine Sun

[Laughter] Can’t language models write well? I do say in the piece that most writing, period, is very bad, and so I think that language models are definitely better at writing and language than most humans are. But the question that I was really curious about is: Why can’t they write at a sort of literary, creative-fiction level?

Because the thing is, if you listen to these AI leaders talk about their aspirations, they say, “We’re going to cure cancer. We’re going to solve physics. We’re going to build a superhuman coder.” They are not shy about saying, “Our AI models are going to be better than 75 percent of human coders.” They’re saying, “No, we will literally build a self-replicating factory tomorrow.”

And then Tyler Cowen asked Sam Altman in an interview from last October, “When do you think GPT will be able to write a Neruda poem?” And Sam Altman says, “Maybe in the future ChatGPT will be able to write a real poet’s okay poem.”

So, that was the thing that fascinated me. Even these guys who are more bullish than anybody else about the capabilities of their technology are very reserved about how much literary writing their models can do. That was the gap that I was really interested in.

Casey Newton

Mm-hmm. And you start your piece with this interesting provocation, which is that in some ways GPT-2 was the peak of AI when it comes to creative writing. So, explain that.

Jasmine Sun

Part of what got me interested in this piece was that I was doing research for your book, and I was going through all of these previous generations of models and reading the outputs. The thing that really shocked me is that, in a way, I found the writing style of GPT-2 and GPT-3 so much more compelling than ChatGPT today.

It doesn’t have any of the annoying tics. It doesn’t have the em dashes, the tripartite lists, or the “it’s not this but that” construction. The tone was much more variable. It would actually surprise you. It would be funny. It would be poetic.

That shocked me, to go back a few generations and realize that maybe they were also lying all the time and doing all sorts of other things, but from a writing-style perspective, I kind of preferred it. I wanted to investigate that.

Kevin Roose

Weird. That shocks me. To me, talking to GPT-2 was like talking to somebody who had just fallen down the stairs. [Laughter] You know what I mean? It was like, “I think we need to get you to the hospital. Do you smell toast?” [Laughter]

Jasmine Sun

There are these amazing prompts from the early OpenAI prompt library where they would say, “I just won $175,000 in Las Vegas. What do I need to know about taxes?” And GPT-2 would start writing some short story about an orphanage. [Laughter]

But they were surprising. They were nutty. They were weird. They would absolutely be a terrible corporate assistant and a horrible coding intern. They can’t do any of the things that modern LLMs can do, which I’m very grateful for. But from a pure writing-style perspective, they’re very good.

GPT-3 in particular—there’s this set of samples that some guy did where it was, “Write in the style of Paul Graham. Write in the style of Richard Dawkins,” and so on. It could style-match much better than modern LLMs can. Particularly because so much literary writing comes from voice and style, that was one of the things I was really interested in: What did we lose that means the LLMs can no longer emulate Paul Graham’s style, or whoever’s style?

I would put in the exact same prompt that this guy gave GPT-3 into ChatGPT 5.4 Thinking, or whatever, and it would be god-awful. I was like, “That’s really weird.”

Casey Newton

So, tell us about what you learned about what happened after the GPT-2 and GPT-3 era that changed the way these models respond to us.

Jasmine Sun

Yeah, I think the answer is post-training, basically. They started adding a post-training layer, which is basically saying, “We have these crazy, unpredictable, nutjob, concussed models.” And they need to learn how to behave, because a model that can’t behave is a very bad corporate assistant.

The AI researchers give them example dialogues and scripts to learn from. They give them words that they can and can’t say. They do RLHF, which is a process by which human graders rate which response sounds the most helpful, or something like that.

Now these post-trained models have been trapped, in a way—trained or guided toward a very particular character or persona that is a very helpful assistant, but might be very bad at writing in creative and surprising ways.

Kevin Roose

Mm-hmm. I mean, the way that you described it was that there’s a phase within the post-training phase where these AI models are evaluated by humans. That’s part of what they call RLHF, or reinforcement learning from human feedback.

What struck me in your reporting is that you actually talked to some people who have given this kind of feedback to the models and say that they’re just being asked to grade things in ways that don’t make sense.

Jasmine Sun

Yeah.

Kevin Roose

Right. Tell us about that.

Jasmine Sun

Yeah, this is super interesting, because these job listings on places like Mercor, or xAI—Elon’s company will list them directly—will say something like, “Creative-writing expert, $45 an hour, must be a New York Times best seller and have a starred Kirkus review.”

Casey Newton

Have you ever gotten a starred Kirkus review, Ruth? [Laughter]

Ruth

I think so.

Kevin Roose

Okay, good job. All right. You might qualify to help Annie from Grok write a little bit better.

Casey Newton

Yeah, we’re going to get her that job listing. But okay, you were saying.

Jasmine Sun

These companies realize that AI researchers are really good at knowing what good coding is, but they don’t actually know what good writing is. So they’re like, “Why don’t we hire some humans to find out?”

They’ll commission MFAs and published authors, and sometimes just random guys with a blog or whatever. One of the people I talked to was a contractor for Scale AI as a writing evaluator, and he was doing this for one of the bigger labs. He said that the rubrics just didn’t make any sense.

He would be told things like, “You have to grade them based on the number of exclamation marks.” If something has 3 exclamation marks, that’s too many, and so you have to ding that one.

Kevin Roose

Yeah, and I have to say, generally not bad writing advice. [Laughter] I guess it depends on the length of the text, but 3 feels like a lot for many scenarios.

Jasmine Sun

This is what they tell women in business communications. It’s like, “Take all those exclamation marks and replace them with periods.” We’re just going to remove all of the ideas.

Casey Newton

We teach women to shrink themselves.

Jasmine Sun

Exactly. Yeah. He was being asked to grade these things, and another one was that he got a bunch of fan fiction and was supposed to grade it on its factuality, since that was one of the criteria.

I do imagine that one could devise better rubrics than this particular evaluator was given, but I think it does show, at least, that some of these very big, very well-resourced companies simply do not know how to think about what good writing is.

Kevin Roose

Yeah, and is that what—just briefly, I want to underline that, because to me that seems like the whole story. We are taking the entire internet and grading it on factuality. The LLM that you’re going to get out of that is probably just not going to be all that creative.

Jasmine Sun

Well, I wonder how much of it is related to this verifiable reward system that a lot of these companies are using, where you have a system generate a bunch of code and then you have another evaluator model check the code to see whether it’s good or not.

That works in domains like programming, where the code either runs or it doesn’t, but creative writing doesn’t work that way. You can’t have an evaluator tell you, with any sort of consistency, whether something is good or not. It may just come down to preference.

So, I guess I’m curious: Do you see this as a technical problem that the labs are frustrated trying to solve, or is this just demand-related? Is this just what people want chatbots to sound like? In every test where they pit different models against one another, the one that sounds like a bland corporate assistant wins.

Kevin Roose

And so, they go with that.

Jasmine Sun

I think both are true. The majority of writing that we are asking the models to do is, “Write this email for me,” right? They excel at that. They are truly great corporate email writers. They are much better at the whole passive-aggressive thing than I am.

At the same time, I do think, as you said, there is a technical challenge that has to do largely with verifiability. There are people who have spent decades of their lives attempting to articulate what makes Shakespeare Shakespeare or what makes a Neruda poem a Neruda poem, and they will still not know in any certain way. They will still get into debates with their fellow academics and literary critics about which writer is better than the other.

Because these things are subjective, because they are ineffable, because they are hard to put in a rubric—that is the nature of art.

Kevin Roose

And to that point, you started this segment by talking about Sam Altman saying, “Hey, we just basically can’t write a great poem yet.” Sam Altman said a year ago that the company had trained a good creative-writing model and posted a short story on X. Many people found it compelling. Is Sam Altman just not being consistently candid with us, Jasmine?

Jasmine Sun

Ooh. [Laughter] Wouldn’t be the first time. But that short story, if you remember, had some great lines, like talking about the seams of mirrors, or Thursday, the—

Kevin Roose

The liminal, almost-Friday, or something.

Jasmine Sun

Yeah, the liminal day that was almost Friday.

Kevin Roose

Wait, I had to actually look this one up because it was so good.

Jasmine Sun

I think while you’re looking it up, the thing about AI writing is that it comes up with all these fun metaphors, and they are kind of surprising sometimes. But the language is not grounded in a life, and that was my other thing.

Aside from the verifiability, fundamentally, when I think about the writers who I really love, whether they’re journalists or poets or whatever, they are writing from life, right? A journalist goes out and talks to people, and they see stuff and observe the color of the sky in a particular way. A poet is thinking about personal experiences that they’ve had. Their writing has stakes. It comes from an emotional place.

The fact that LLMs, while being very talented and grammatically pristine, don’t have lives means that all of the metaphors they choose, all of the words they choose, and the examples they choose are ungrounded, right? They’re not coming from a point of view or a particular experience or community that makes the writing believable.

I think part of what voice and style are is that they’re very specific to the life that a person has had, and LLMs cannot get there in the same way a human who hasn’t really lived that life cannot get there.

Casey Newton

I don’t know. I feel like it’s case-dependent. I’m a big music fan, and over the past few months, I’ve enjoyed putting questions about music, and in particular the sounds of certain bands, to an LLM, which sounds like a joke prompt because an LLM has never heard anything, right? And yet, I find that in general, the models can have good conversations with me about the sound of music.

It may be that they’re just pattern-matching based on a bunch of public writing on the internet by people who do have ears and have heard. I’m very open to that. But, again, I’ve just been struck by the way that it is able to write about sensory topics in an evocative way that, at least to me, surpasses what I would have predicted they would be able to do.

Kevin Roose

I want to pose a couple of objections that I think someone might make to your article. One of them is: This is cope. This is Jasmine, a writer, a very talented writer, finding the things that AI, in her view, is not good at yet and saying this is categorical proof that it will be very hard for AI to do these things.

This is the same reaction that software engineers had when models started getting really good at code. They would say, “Oh, well, it can’t do these other 10 things that I do.” And that basically means, just wait a few years, and the models will be better than all of us at everything, including writing.

Jasmine Sun

I would love for it to be cope because I try to automate myself away all the time. I have no deep attachment to having to write. I like writing, but I have tried over and over and over for the past 3 years to automate my own job away and to get Claude to do my job for me. It cannot do it. This is very frustrating. It’s not out of a lack of trying.

I’m going back to the CEOs themselves and the things that they themselves are saying, right? It’s not just me, a writer. It’s Sam Altman saying this thing will cure cancer and solve physics, but it will not write better than a real poet’s okay poem. I think that suggests that there is something that is at least perceived as a little bit different.

I think it’s very possible that the models will get much better at writing over the next few years. I don’t think it’s a never thing. I do think that reporting is hard to replicate. I think having life experiences that are real and verifiable is hard to replicate. The style stuff can be improved, especially if you fine-tune the models. But I think what's also interesting to me about this piece is that it shows how the market incentives, the demand incentives of these companies do shape what we see their abilities are today.

I’m sort of a never-say-never person. Maybe it’ll get there. I would be totally happy if Claude were able to give me good ideas for my next essays, but it’s not there yet.

Kevin Roose

The other objection I imagine people who are very AI-pilled might have is that this is all in the eye of the beholder, right? There have been several studies now that have shown that if you give people a blind taste test of AI writing versus human writing, they prefer the AI writing until you tell them that it’s AI writing, and then the value in their eyes plummets. I did one of these in a New York Times quiz just recently.

Is it possible that the models have already become superhuman at writing, but that the minute we learn that they are AI models generating text and not humans writing words with their fingers, we lose all interest in it just because of the source, not because of the quality of the writing?

Jasmine Sun

I think it’s definitely interesting and true that people don’t want to like AI writing, and that is part of what bothers them when they see AI text that is obviously AI, even though, as you said, in these quizzes and tests, AI can outperform human writers in those narrow scenarios.

My quibble with a lot of these quizzes and tests is that, as a writer—and you guys are writers, too—how much of your job is actually text generation? I think AI is a superhuman text generator, right? I am generating text probably 25% of the hours in my day. I spend a lot of time interviewing people. I spend a lot of time coming up with ideas. I spend a lot of time reading, and not just reading indiscriminately, but reading very particular sources that feel like the right ones.

Usually, at the point that you are doing one of these tests, you’re saying, “Generate one paragraph very specifically about why Trump won the 2016 election, 500 words or less.” You’ve already given the prompt, which I think is a critical part of writing: What are you going to write about? You’ve often supplied some of the evidence, the guidance, and the form of it, saying, “500 words or less.”

At that point, I do think that AI is probably a better text generator than almost all humans are. But, again, AI is still very bad at coming up with ideas for articles. It is still very bad at reporting. The non-text-generation parts of the role feel further away from automation.

Again, I’m a never-say-never person. Maybe it’ll get there. I would be totally happy if Claude were able to give me good ideas for my next essays, but it’s not there yet.

Casey Newton

Well, we’re already seeing the LLMs make huge progress in genre fiction, right? Recently on the show, we talked to the author of a story in The New York Times about how authors of romance novels are now able to generate dozens of novels a year using LLMs. In fact, much of the discussion that we had was around how you just have to prompt them differently and relentlessly in order to get what you want.

Your piece, Jasmine, made me wonder how much of getting a model to just write weird can be achieved by repeatedly telling it, in different ways, “Hey, be a little weirder.”

Jasmine Sun

Some of it, but not all of it. I talked to, for example, James Yu, who is the co-founder of Sudowrite, one of the earliest creative-fiction AI writing assistants. I talked to some other folks who were similarly in the fiction-writing LLM space.

As you said, to an extent, a lot of writers are already using these and leaning on LLMs to generate large amounts of text. It can be very successful, and it can meet readers’ needs.

But even these people I was talking to were describing to me how freaking hard it is to undo all of the post-training that the labs have done. They’re applying immense amounts of engineering effort that, clearly, in my conversations with them, frustrates them. It is so hard to get these models to stop being so chirpy, so sycophantic, and so PG-13 in order to get them to this sort of base-model state where they’re able to be weird again.

So I think it's certainly possible, but I think the labs have made it quite challenging just because of the way that these models are trained.

The other thing that I think is important is I tend to think that writing and a lot of creative work is actually the perfect use case for these Centaur models, right? The idea that the human-plus-AI collaboration is where you can get the furthest. I listened to the interviews you guys did about the fiction authors, and I was thinking, this is a Centaur model, right? Because without the human prompting and bullying the AI into getting weird and getting sensual and whatever, it was not going to do that on its own.

I myself do use LLMs as a research assistant. I wrote about that inside the Atlantic piece about the way that Claude has now helped me edit my own work in a way that I found incredibly useful, but I do feel like the collaborative element is important for any domain where the personal perspective, lived experience, or whatever really matters.

Kevin Roose

Talk about that a little bit. You mentioned your editing process. How are you using AI to help you edit your work, and are you finding it useful?

Jasmine Sun

I feel like I really cracked this over the last couple months, which I'm very excited about, because again, I've tried to make these things write and edit for me over and over and over, and they've never really been able to do it.

The thing that I realized was, if I make Claude into an editor that is not just trying to grade and give feedback on my work against some genericized standard of what good writing is, but actually evaluating it against basically what my personal aspirations for writing are, it can give feedback that I find much, much more helpful.

What I did was basically feed Claude my entire Substack archive of the writing that I've previously done, as well as some of my freelance work.

Casey Newton

And just to get real specific, is this inside a Claude project, or have you set this up? Because I know our listeners are going to want to try this.

Jasmine Sun

Yes, I did it in a project, but on Claude's advice. I was like, “Do I need to Claude Code something?” And Claude was like, “No, that's overkill.” So you don't need to code or anything.

In a Claude project, I gave it my whole archive of writing. I also personally write retro notes to myself after everything I publish. So I have a notes app that's just me writing what was good and bad about everything I've ever written, just a few bullet points.

Casey Newton

This is why Jasmine's going to be our boss.

Jasmine Sun

I'm sorry. These are very low-quality bullet points, but I also gave it that because I wanted it to learn my taste. I wanted it to learn what I aspire to be, where I see myself falling short, and what I'm proud of, right?

From those two things, plus a little bit more information about, like, here's my audience, this is my beat, these are my goals, we were able to co-develop a rubric of—rather than “How many exclamation marks does it have?”—things like, “Does this take advantage of your ‘insider anthropologist’ position in Silicon Valley?” That's one of the things that Claude and I think distinguishes my voice.

It'll also notice, like, “Oh, Jasmine, you tend to move between registers. You'll switch between startup jargon and internet slang and whatever.” And I think the fact that you can do the high-low or move from policy to a personal scene is something that is characteristic of your writing. So again, we're co-developing these qualitative criteria.

Then I split it into phases: an ideation-phase rubric, a structure rubric, a prose rubric, and final fact-checking. What I do now is put this all in a Claude project. I said, “Your job is to evaluate my drafts based on these criteria, but not to do the writing for me, and to make sure to prompt out of me what I can do better.”

I dumped a draft into Claude. Claude will run phase 2, structure, on it. It'll say things like, “Your conclusion is just a summary, and this is really boring. In fact, in your piece about this and that, you actually ended on a scene, and I thought that was much more powerful. So why don't you try ending this one on a scene?”

And rather than inventing a scene, Claude will say, “What were you thinking when the plane took off? What were you feeling inside? Can you think of a scenario where you had a conversation with, say, a kid-safety advocate about AI that really resonated with you? Because right now it sounds like a dry policy explainer.”

That feedback I actually found incredibly useful. I'm still applying my own judgment to say, “Do I take it or not?” But I'm like, “You know, this is about me becoming the best version of myself as a writer. It's about me self-improving, AI pushing me to do that.” I found that much, much more helpful.

Casey Newton

I want to ask you both a question as fellow writers. Do you feel the impulse to make your writing weirder because of AI, to sort of stand out from the sea of slop? Because I find myself feeling this tug of, “Oh, that's a little weird aside that probably I should cut, but I think I'm going to leave it in because Claude would never do that,” right? It's a marker that I am typing these words, and I feel like that's my imprimatur that I'm leaving.

Kevin Roose

My answer to you is yes, I absolutely feel that way. I've gotten back and tried to edit sentences to make them feel a little bit weirder, or in particular to make them sound colloquial in a way that I know an LLM generally would not be. And yes, it is for that reason.

I think that writing right now—we're all, or many of us, on such high alert for the prospect that we might be reading slop—that if you are a writer who does not want to be producing slop, you should be asking yourself that question.

Jasmine Sun

I think it makes me a lot more comfortable writing the way I want to write in the first place. Maybe unlike both of you, I didn't come up through newsrooms where I was learning a very specific house style and all of these norms.

I can do news writing now. It's something I've learned, but I'm actually much more, quote-unquote, internet- and blogging-native, which is a form that is voicey and irreverent and not as pristine, and will make inappropriate jokes. It's just a looser form of writing.

So I think what it's actually done is made me more comfortable doing the bloggy thing instead of always trying to write in a more professionalized journalistic tone.

Kevin Roose

So I think we should leave this with a question for you, Jasmine, which is: Your piece makes the case very convincingly that today's AIs are not very good at the kind of writing that I think we all value. Do you think they will get there, and what should the companies do to make their models better at writing?

Jasmine Sun

I think that if we separate out text generation from reporting, which I'm not that bullish on the models doing, and we are just talking about, say, literary fiction, or here's a bunch of interview transcripts, write a magazine feature or something, I think that if they applied as many resources toward that task as they do toward coding agents and things that actually make them money, I think that they could get there.

Will the companies ever find it financially advisable to spend all their resources on that instead of automating 23-year-old software engineers? Probably not. I would be grateful for that world. I don't need them to take my job or these folks' jobs, but I think it's possible.

Casey Newton

Look, they're going to get around to it eventually.

Jasmine Sun

Okay.

Casey Newton

You know, I hear what you're—

Kevin Roose

What writers make in this economy, Casey?

Casey Newton

Eventually—

Kevin Roose

Those aren't going to pay for a lot of data centers.

Casey Newton

No, there is economic value in writing, and eventually the AI companies will want that all to themselves.

Kevin Roose

You know what would be a very funny outcome of this, taking your point about the sort of guardrails of the models? Maybe the next great American novel will be written by Grok.

Jasmine Sun

Oh, God.

Kevin Roose

And with that, Jasmine Sun, thank you for joining us.

Jasmine Sun

Thank you very much, Kevin and Casey.

Casey Newton

Well, Kevin, you've recently returned from book leave and are once again writing in The New York Times. How does it feel to see your name in print again?

Kevin Roose

Feels great. Hasn't happened yet, but when it does, it'll be great.

Casey Newton

Well, I got to take an early read at a story that you are publishing about the fact that tech companies have now created leaderboards to show which employees are using the most AI tokens in their work.

Kevin Roose

Yes, it's a token frenzy out there, and the employees of these companies are competing among their colleagues, sort of informally and for fun, but they're taking it very seriously. They want to be the people at their company who are using the most AI tokens.

Casey Newton

So let me just ask a basic question for listeners who may not be familiar: What is a token, and why is that something you might start keeping track of?

Kevin Roose

A token is the basic atomic unit of AI labor. It's basically a fragment of a word, and it is how AI model providers measure their consumption.

So if you type in a prompt—“Help me write this essay”—an old model might have given you a couple hundred tokens in response. That would be a couple hundred words. And what has been happening over the past year or so, as these agentic coding tools have started taking off, is that the models are just much more token-hungry.

You can now use hundreds of thousands or even millions of tokens in a single session. And so, what is propelling these leaderboards is the idea that the more coding you’re doing, the more agentic tools you’re using, and the more simultaneous processes you’re running, the higher your token count will be.

Casey Newton

One measurement I found useful was that, apparently, it takes about 10,000 tokens to generate 7,500 words. If that helps ground you at all. But as you just said, and I want to hear more about this, the more advanced systems are using way more tokens than that. So, tell me about some of the numbers that some of the token all-stars are putting up on the boards.

Kevin Roose

I don’t know all of the exact numbers, but I did learn that at OpenAI, where they do track this kind of leaderboard, the highest employee token count over a 7-day period recently was from a guy who used 210 billion tokens. And, for rough scale, this is about 33 Wikipedias’ worth of text.

Not all of that is typing and receiving a response. Some of that is what they call cached tokens, so it’s not all being extruded from the model for the first time. But these are the kinds of numbers that I think even a year ago would have sounded completely insane.

Casey Newton

Right. Now, was this guy working on a new mass domestic surveillance program for the Department of Defense?

Kevin Roose

I don’t know, and OpenAI did not make him available for interviews. But what I wanted to do in writing this column was to try to call or talk to a bunch of people who are in this sort of billion-token club—the sort of extreme power users—and just ask them, “Hey, how are you guys using all those tokens? Isn’t that very expensive, and how are you paying for it all?” And I learned a lot.

Casey Newton

Yeah. Okay, so tell us, first of all, just how expensive it is.

Kevin Roose

Very expensive. In fact, I heard that the top user of Claude Code, the top individual user of Claude Code as measured by Anthropic, spent more than $150,000 on tokens last month. If you extrapolate that, it’s like an employee making more than $1 million a year and burning that in a month. And I heard similar figures from some of these other extreme coders who are spending something on the order of thousands of dollars a day on tokens from these models.

Now, we should also say the employees of these companies get their tokens for free, right? So they’re not shelling out; their companies are not shelling out. But at other companies, this is starting to become an issue because they are outstripping their budgets for these things.

Casey Newton

So there are companies where there are engineers who are legitimately costing their employers maybe $150,000 a week because they’re getting tokens from one of the big providers.

Kevin Roose

Yeah, I talked to a software engineer in Sweden who said that he probably spends more than his salary on Claude. So, this is essentially becoming a very expensive job perk for some of these coders.

Casey Newton

So, talk to me about why employers want to create leaderboards to promote this to employees, because I could see other companies saying, “If you spend $150,000 on tokens last month, you actually don’t work at this company anymore because we’re bankrupt.”

Kevin Roose

Right. So, this was a big question I had: Why is this going on? And it seems to be some combination of employee motivation and worker tracking. There are executives at these companies who think that the more tokens you use, the more productive you probably are. And as we discussed in a previous segment on this show, these companies are very eager to have their workers start embracing the AI tools.

At a number of these companies, I talked to people who said, “Yeah, this is basically them trying to see who is really all in on the new way of programming.”

Casey Newton

And you’ve talked to a number of people who are ranking high on these leaderboards. I realize you probably haven’t dug deep into their code, but what is your sense of how productive they actually are? What is the relationship between token usage and taking my company to the next level?

Kevin Roose

It’s very unclear. Some of these people may just be generating worthless projects. I think the thing that worries a lot of the people I talked to about these leaderboards is that they incentivize you to run up your token count, because then you look like the special 10x engineer, 100x engineer who’s outperforming all your colleagues.

So, I think there are a number of companies that see this leaderboard business as a little strange and maybe counterproductive. But I do think that there is a feeling among the heaviest token users that they are being productive.

Casey Newton

Yeah, I have to say, when I read your column, I thought, “This just seems like it would create the worst incentives.” There’s this idea of Goodhart’s law: When a measure becomes a target, it ceases to be a good measure. I can’t think of a better way to ensure that token usage becomes a bad measure than creating a leaderboard for it.

Kevin Roose

Totally.

Casey Newton

What are the people inside the company saying about that?

Kevin Roose

Well, some of them are opposed to this whole leaderboard thing. I also talked with some folks who defended the leaderboards. They said, “Look, it’s never been all that easy to track the productivity of programmers. Some people have had their productivity measured by how many lines of code they generate or how many pull requests they made.”

These are imperfect proxies for how hard you’re working and how much you’re doing. But the employees of these companies also see this, I think wisely, as a key to their own success. A number of these companies are now using AI token use and consumption as part of the performance review cycle.

So, you go in for your annual review, and your boss says, “Hey, it looks like you only used 70 million tokens last month. What’s going on?” And so, I think the engineers at these companies are getting wise to the fact that if they want to have a long, successful career, they better start using some tokens.

Casey Newton

Yeah, but I imagine that some of them are really nervous about that, though. Because it seems clear to me that at least some of these companies want to incentivize token usage because the companies themselves suspect that the more we can get them using this stuff, the less time we will have to employ the humans.

Kevin Roose

Maybe, although I think it’s less about AI systems replacing humans and more about it being a radically different way of working. These are people who, for the most part, have had long careers in software engineering. They grew up writing code by hand, and they may have grown up using an AI assistant like GitHub Copilot.

What people at these companies are saying is that these agentic engineering systems are just really different. You have to approach them in a different way. You have to spend a lot of time with them to understand what they’re good at and not good at. And to them, this is a way of motivating their employees to say, “Hey, go out and try the new thing.”

Casey Newton

Yeah, I don’t know. I’ve been thinking a lot about this question: If I were an engineer at one of these companies and I had this incentive to get on the leaderboard, how would I approach it? And I do think the instinct to just waste a bunch of tokens to rise higher on the leaderboard is ultimately a bad one. If you rise too high, people are going to ask you what you did with all the tokens.

If you’re number one at 10 billion tokens and you only managed to vibe-code a calculator or something, people are probably going to get mad at you.

Kevin Roose

Yeah, and I actually did talk to one person who speculated that the people at the top of the leaderboards are all doing side projects. They’re starting their new companies.

Casey Newton

They started a new company with the bosses’ money. And if you’re doing that, I just want to say I salute you. That is the right way to work, in my opinion. Don’t be number one on the leaderboard if you’re doing that. Maybe try to stick around 6 or 7. Middle of the pack is where you want to aim yourself.

Let me ask: Is there any kind of token tracking that you think offers a reasonable signal? Do you think that, if you’re a tech company, you should create a leaderboard?

Kevin Roose

No, I think that’s a bad idea, for all the reasons that we just talked about, including Goodhart’s law. I think this is just going to lead to people wasting tokens and doing side projects.

But if I’m the budget manager at a company and I’m seeing that people are spending multiples of their salary on AI tokens, I’m asking them some questions about what they’re doing with all that. And if their answer is not, “I built an amazing new product that’s going to generate billions of dollars a year in revenue,” I’m trying to say, “Hey, could you maybe use a little less next month?”

Casey Newton

Yeah. I have to say, I have been struck by how this idea of the token leaderboard just represents a new incarnation of something that the software industry has been trying to figure out for a long time, which is: How can I figure out if my software engineers are productive?

I was talking recently to this very handsome software engineer who I’m engaged to about your column.

And he was telling me that he used to be evaluated on how many lines of code he contributed. And he told me about all the games that people used to play back in the day: “Oh, I wrote a quick algorithm to translate a bunch of stuff into some new languages, and it's completely worthless, but it makes me look like I had a very productive week.” So I went back and looked into this, and they were doing this in the ’60s and ’70s.

And there's this saying from the early days of computer programming that eventually arose: “Measuring programming progress by lines of code is like measuring aircraft building progress by weight.” And I have to say, I think that the same thing kind of applies here, right? If you squint and look at the right level of abstraction, it's probably true that some people who are using a lot of tokens are more productive than some people who aren't. It just doesn't quite seem like the right way to measure these things, and I just wonder how quickly the industry is going to figure that out.

Kevin Roose

Yeah, I think it's going to be pretty soon, in part because the budgets are just getting very ridiculous, and because the AI model providers are now seeing individual users consuming amounts of their services that entire companies would have consumed just a few months ago.

Casey Newton

You know, the last question I have for you about this is just what implications you think it has for the broader economy, right? Because we know that in so many different sectors of the economy, managers are saying, “I want to incentivize my employees to use AI, and I want to track how they're using AI.” So do you think that, as knowledge of these leaderboards spreads, we're going to see people in nontechnical fields try to adopt their own version of them?

Kevin Roose

I hope not. I think it's really a bad move, not just for tracking actual productivity and output, but just for morale, right? I remember years ago when Gawker would have a traffic leaderboard at their office, so you could see how many clicks your stories were getting relative to other people's. And I don't think anyone who worked there at the time thought that was incentivizing the right things or creating high morale among employees.

Basically, everyone was just competing with each other all the time. And I think in this case, it's even worse because it's not necessarily even correlated with any success. It's just pure, sort of, how many agents you can run in a parallel swarm to work 24/7 doing tasks of uncertain value.

Casey Newton

Which is a great question to ask on a first date in San Francisco, too, by the way. [laughter] But anyway, I have to say, I worry that this idea of token maxing is going to spread into the broader economy. I was talking with somebody who works in marketing this week, and she was telling me that her job used to be evaluated solely on creativity.

And then recently, the performance review got a new AI section, and everyone is being evaluated on how much AI they use. From her perspective, she was kind of like, “This was working fine. I didn't need to use an AI tool to help me, but now my bonus might be based on how much of it I use.”

So I think this thing has already seeped out of the labs and is getting into the water elsewhere. I just hope that managers are really thoughtful about what they are incentivizing, and that maybe AI use for the sake of AI use is not going to be the boon to your company that you're hoping it is.

Kevin Roose

Yeah, I think it's going to be very case-by-case. I think there will be people who are token maxing who are way more productive than their colleagues and doing way more projects way more quickly. I think there will be other people whose managers look at their token budgets and say, “You spent this many tokens on what?” and will have to have some hard conversations.

But I think it's very hard to draw with a broad brush and say all of this token maxing is pointless productivity theater. It sounds to me from my conversations like some of it really is working for people.

Casey Newton

Yeah, well, I will say, on the flip side, I've also heard of people in my social circle who have gotten in trouble for spending too much on Claude.

Kevin Roose

Really?

Casey Newton

Yeah, and [laughter] when I heard that, I was like, “Oh, your company's not going to make it, bro. You got to spend on this stuff.”

Kevin Roose

Well, what's so interesting is now it's becoming part of job conversations for engineering jobs. People are going into new jobs and saying, “What's my token budget?” And for the employees of these big AI labs who have unlimited free access to the models, some of them are using so many tokens that they effectively can't afford to quit their jobs, right? Because anywhere else they would work would have to pay for their tokens, and it would be completely unaffordable to employ them.

Casey Newton

Yeah, I mean, those sound like real incentives, and better than the ones that Meta offered. Do you remember when Meta was spinning up a superintelligence lab and they said, “You can sit really close to Mark Zuckerberg”? [laughter] If I were them, I'd be like, “I'll take the tokens, thanks.”

Kevin Roose

All right, well, just to wrap this up, exactly how many tokens should a person use?

Casey Newton

I think you have to look within yourself.

Kevin Roose

Yourself?

Casey Newton

Yeah.

Kevin Roose

Okay.

Casey Newton

Yeah. That's between you and your God.

Kevin Roose

Yeah.

Casey Newton

Do what Marc Andreessen will not. And introspect.

Why Tech C.E.O.s Are Blaming A.I. for Mass Layoffs | BidClub