[BidClub_]
Hard Fork · · 63 min

Zuckerberg’s Anti-Doom Fantasy + Finally an A.I. Detector That Works + A.I. Math

Kevin RooseCasey NewtonMax Spero

Podcast
TL;DR
  • Casey Newton reads Zuckerberg's 6,500-word "The Future Is For Everyone" manifesto as a policy wish list dressed as optimism: accelerated data-center construction, maintained chip export controls (which advantage Meta's open-weight models over Chinese rivals), reduced "training data restrictions" to fight creator lawsuits, and legal protection for distillation. The timing matters — Meta is "getting its ass handed to it in court case after court case," including a New Mexico ruling ordering an extra $567M into a teen mental-health abatement fund on top of $374M in civil penalties, a "public nuisance" finding comparing Meta to a polluting factory, and a 90-hour-per-month cap for under-18 users.
  • Kevin Roose raised his probability that Meta actually builds superintelligence from ~1% six months ago to ~10% now, citing improving training runs since the Meta Superintelligence Labs reorg — and that improvement worries him. He resurfaced a former DeepMind executive's taxonomy: tool companies and AGI companies can both be fine, but "the real danger is if you have a company that thinks it's designing tools but is actually designing superintelligence" — once said of Google, now fitting Meta.
  • Casey's central metaphor, via the House of the Dragon finale: Zuckerberg's "personal superintelligence for everyone" is "a little bit like giving a dragon to everyone." Defender-catches-up equilibrium may hold in cybersecurity but not bioweapons, where defenders may take longer to catch up to a novel pathogen — a gap Zuckerberg acknowledges in the essay and then punts on.
  • Both hosts agree AI positivity is not a messaging problem: "It is delivered via real experiences that human beings are having," Casey argued — cures, higher pay, kids excelling — a chasm no leading lab has crossed. Kevin's compression: "ship it or zip it," and Zuckerberg is "possibly the worst messenger" — "from the people who brought you Cambridge Analytica comes superintelligence."
  • Kevin openly reversed himself on AI text detection — he'd argued on the show it was "basically worthless," but independent studies now suggest Pangram is genuinely accurate. CEO Max Spero explained the shift: ditching the flawed perplexity metric for a trained classifier that pairs real pre-2022 human writing against LLM imitations and aggregates "a bunch of really weak signals" — good enough to flag a "random" number string as AI because LLM sampling isn't actually random.
  • Spero's investment-relevant thesis: bot traffic has just surpassed human traffic at roughly 50/50 and may reach "99% bot traffic" within a decade, so platforms must "discriminate in favor of humans" — and detection demand is being validated by regulation, with Anthropic agreeing to watermark all its text under the EU AI Act. He disputes Ben Thompson's take that watermarking degrades outputs: entropy-constrained watermarking (per Google's SynthID sampling-perturbation approach) shouldn't.
  • The new "Running the Numbers" segment carried three data points: an Anthropic employee's Claude session made real progress on a side problem involving the Riemann zeta function via pep-talk prompting ("You got this") — new-knowledge production, though explicitly not a proof; Airtable was acquired by Bending Spoons at $1.29B enterprise value versus an $11.7B 2021 mark, which Kevin reads as AI eating SaaS margins broadly; and AI wealth is repricing dating markets, from some Korean chip workers' $400–500K bonuses to "Anthropic goggles" in San Francisco.
Digest · the substance, structured for research

1. Zuckerberg's manifesto: personal superintelligence for everyone is a dragon for everyone

  • The document: 6,500 words titled "The Future Is For Everyone," promising that "in the next few years, people will be able to use superintelligence beyond human capacity" — every Meta user getting "an exceptionally capable personal agent," accessible through glasses, illustrated by Zuckerberg's own use cases: prototyping ideas, sleep monitoring, and personalized weekend baking recipes with his daughter.
  • Casey's framing, fresh off the House of the Dragon season-three finale: the show opens with one family holding the super-weapon, then a schism, then bloody civil war. Zuckerberg's answer to AI risk — maximal proliferation — "is a little bit like giving a dragon to everyone... some people are gonna be launching cyberattacks and engineering novel bioweapons."
  • Kevin's dissection of the rhetorical move: Zuckerberg argues no universally aligned AI can exist because values differ, so everyone should get a superintelligence tailored to their own. "Sounds great, but in practice, giving the Chinese Communist Party an AI superintelligence that obeys their values would allow them to commit atrocious acts against their own people."
  • On offense/defense equilibrium, Casey's key carve-out: it may hold in cybersecurity, but "if I release a novel pathogen... it is just going to take the defenders a little bit longer." Kevin notes Zuckerberg actually acknowledges bio may break the equal-tools logic — "he kind of punts on it... I don't think he's fully naive, but he just doesn't have a good answer for that part yet."

2. The real purpose: a policy wish list timed against courtroom losses

  • Casey's inventory of what the essay actually asks for: accelerated data-center buildouts (feeding Meta's potential neocloud business), maintained export controls on advanced chips — notable from "Mr. Open Source," since it advantages Meta's open-weight models over Chinese alternatives — reduced "training data restrictions" amid creator lawsuits, and legal protections for distillation. "Buried inside this very positive, happy vision of AI is just a series of policy requests to help Meta as a business."
  • The timing: state after state is coming after Meta on teen safety. Last week a New Mexico judge ordered an extra $567 million into a teen mental-health abatement fund atop $374 million in civil penalties, ruled Meta's platforms a "public nuisance" — comparing them to factories with "the psychological harm and exploitation of children as the pollution it emits" — and imposed a 90-hour-per-month cap for under-18 users plus new chatbot restrictions.
  • The historical rhyme Casey won't let go: Zuckerberg once wrote happy manifestos promising that connecting everyone on Facebook would yield "more democracy than you've ever seen before." "Suffice to say, that didn't really work out. Now we get the exact same maximalist, universalist framing — which of course also maps 100% to Meta's business interest — but this time in the context of AI." His self-aware confession: "I wish that I was born yesterday and could just believe everything that I read, but Kevin, I know too much."

3. Can Meta actually build it? Kevin moves from 1% to 10%

  • Kevin's updated odds: he'd have given Meta "a 1% chance of creating superintelligence six months ago, and now I'm up to maybe a 10% chance." Training runs launched after the Meta Superintelligence Labs hiring spree are coming online with "pretty impressive results" — not a frontier lab, but showing rapid improvement.
  • Casey concurs: "The first rule of Mark Zuckerberg is never count out Mark Zuckerberg." His AI contacts rate Meta's past-year moves highly, particularly reassigning engineers to create reinforcement-learning training data — "the Meta engineers didn't really love that," but AI people say it will give Meta something valuable.
  • The load-bearing anecdote, from a former DeepMind executive Kevin spoke with years ago: there are tool companies and AGI companies, and either can be run responsibly — "the real danger is if you have a company that thinks it's designing tools but is actually designing superintelligence," because it will never treat the technology with "the correct sense of reverence and suspicion." Said then about Google; Kevin now thinks Meta may build "something extremely powerful with this very naive attitude" that these are just recipe tools.
  • Casey's coda on culture: a company that treats everything as existential competition and gives societal externalities "short shrift" — "I do not trust these people to rank a list of viral dances for me without it corrupting my mental health. Once you give these things access to novel biotechnologies, it starts to get pretty worrisome."

4. AI positivity is earned by shipping, not by essay

  • Kevin's verdict on the messenger: "Mark Zuckerberg is possibly the worst messenger for the AI industry on all of this" — if he thinks a sunny vision will stop data-center protests, "he is badly mistaken." Casey's tagline: "From the people who brought you Cambridge Analytica comes superintelligence."
  • Casey's positive case, worth quoting in full: real AI positivity is "new medical advances powered by AI, diseases cured, new jobs created, old jobs paying more money, kids excelling in education... crucially, it is not delivered via manifesto. It is delivered via real experiences that human beings are having, and so far, that has just seemed to be a chasm too great for any of our leading AI labs to cross." Kevin's compression: "Ship it or zip it."
  • One grudging credit: the manifesto passed AI detectors and reads like Zuckerberg actually wrote it — "I never once thought I was reading slop," Casey said, though "to what extent he believes everything in it... I guess I don't know that."

5. Pangram's technical break: classifiers over perplexity, and Kevin's reversal

  • Kevin's on-air change of mind: "I have argued before on this show that AI text detection is basically worthless... but in just the last few months, there's been more and more evidence that at least Pangram, and probably some of these other tools, have gotten quite good." Casey now takes a "100% AI-generated" Pangram screenshot seriously.
  • Max Spero's explanation of what changed: legacy detectors used perplexity — AI text is "suspiciously smooth," but so is any memorized document like the Declaration of Independence, and English-language learners write low-perplexity text too, generating false accusations. Pangram instead trains a classifier on pairs — a real seventh-grader's Moby-Dick essay against an LLM's imitation of one — on a clean, all-pre-2022 human corpus.
  • What it detects isn't em dashes or "not just X but Y": it's "combining a bunch of really weak signals" across a document — consistent word-choice decisions where AI is mode-collapsed and humans have "a wider decision tree." The party trick illustrating it: ask ChatGPT or Claude for random numbers, and Pangram flags the string as AI — "they're not actually random," the model's sampling biases leak through even with no such training data.

6. The cat-and-mouse: humanizers, new model releases, and the false-positive trade-off

  • On blind spots, Spero conceded the asymmetry is deliberate: Pangram is "very much tuned to minimize false positives, so we're not making false accusations" — which means false negatives exist, like tech writer Alex Heath's AI-assisted newsletter scoring as human. The payoff: when Pangram says 100% AI, "you just know, okay, this is probably slop."
  • Humanizers — typo injectors, zero-width Unicode characters invisible to readers but seen by models, wholesale paraphrasers — are a real cat-and-mouse: "we're always picking up the latest humanizers and training on them for our next model."
  • New frontier releases cost "slightly lower recall" for a few weeks until retraining, but generalization within families holds — "if Pangram has seen GPT-5.4 and GPT-5.5, then 5.6 Sol coming out is not a huge surprise," and it caught writing samples from Mythos in a system card despite not having seen Mythos before. His durability argument: labs apply preferences — good writing, correct code, correct math — instead of average next-token prediction, "and these preferences are largely what Pangram is able to pick up on," even if surface style adapts.

7. Why narc at all: a 99%-bot internet needs pro-human discrimination

  • Spero's answer to the "you wouldn't detect spell check" objection: "AI as a tool is probably the wrong abstraction. It's a little bit closer today to AI being an employee or another individual that you collaborate with." Bot traffic has just surpassed human traffic — "about 50/50" — possibly heading toward "99% bot traffic" within a decade, so "as humanity, [we need to] discriminate in favor of humans," including algorithmically, because AI short-form video "will activate neurons and engagement in a way that a real human video cannot."
  • The Substack integration has been "very controversial": AI-assisted writers fear "witch hunts" now that audiences can see. Casey's dry gloss: "Now that they know what it is, they don't like it. Let's go back to the time when they didn't know that."
  • On whether normal people care: "I would totally care if a note from my friend was AI-generated. That would seem like such a big violation of trust." But Spero also flagged where norm-setting overshoots — the Hank Green pile-on "felt like it just went way too far," since Green was honest about AI-assisted research; the hosts noted social media rewards the effortless "LOL, AI-generated" dunk.

8. Watermarking arrives via the EU — and Pangram goes multimodal

  • The news peg: Anthropic has agreed to watermark all its text to comply with the EU AI Act. Spero calls it "pretty huge" validation — an additional verification layer that can eliminate the "one in 10,000 false positives" doubt — but says watermarks have limits Pangram will fill.
  • Mechanics, per the current state of the art (Google's SynthID): the watermark perturbs the token-sampling algorithm in a reverse-engineerable way. Against Ben Thompson's take that this makes model outputs worse, Spero's rebuttal: work only "within the entropy available" — skip tokens where the watermark cannot be applied, as can happen in code when there is effectively only one correct token — and quality shouldn't degrade; the trade-off is that some tokens may go unwatermarked.
  • Roadmap: image detection in research preview, which Spero claims "wins on basically all of the public academic benchmarks" against frontier generators like GPT Image by reading pixel-level generation patterns now that "you can no longer just count the fingers"; video detection reportedly in progress; and the Chrome extension already pulls Google Doc revision history so users can replay writing and check big pastes for educators. On Zoom-scam identity verification: the models come first, then "bringing it to where people work."

9. Running the Numbers: Claude's Riemann progress, Airtable's 89% haircut, chip-nerd dating

  • The math beat: Anthropic employee Jared Sumner — a non-mathematician — set Claude on the Riemann hypothesis while jogging, and a day and a half later it had made progress on a side problem involving the zeta function, coaxed by pure pep talk: "You need to take a big leap of faith in your capabilities... You got this." Casey's read: "This feels like the production of new knowledge to me," and the unreleased model behind it "can do a lot of other things — some good, probably some scary." Kevin's caveat, per Anthropic's own careful framing: "We did not solve the Riemann hypothesis."
  • SaaS math: Airtable, valued at $11.7 billion in the 2021 peak-COVID SaaS mania, was acquired by Bending Spoons at a $1.29 billion enterprise value. Kevin's mechanism: expensive enterprise subscriptions plus customers thinking "maybe I can make a free version of this" with AI — a pattern he expects across early-2020s enterprise software. Casey's rules: "If your business is a fancy spreadsheet, you are in for a rough time," and if Bending Spoons — the Italian, newly public acquirer of "zombie tech brands" — buys your product, "you in danger, girl."
  • Dating math, via a Wall Street Journal A-hed: Samsung and SK Hynix "chip nerds" now dominate South Korea's dating scene, with some workers receiving $400–500K annual bonuses — one Samsung engineer, Annie Kwon, 26, coupled up with a fellow Samsung semiconductor-division employee for "double income"; another woman saved up to buy her Samsung boyfriend a mini PC in return for his roughly $450 Nintendo Switch 2 gift. Kevin reports the same repricing in San Francisco — "Anthropic goggles": "Is that boy really cute, or are you just wearing Anthropic goggles?"

Casey, good morning. How are you?

Kevin, this morning I'm feeling very Dream Beans-pilled. Are you familiar with Dream Beans?

I am. I think you told me about it.

Dream Beans is an experimental app from Google, and here's what it does. It reads all of your emails, and then it tries to send you a daily set of inspirations based on the sort of person you could be if you didn't have a job and were interested in the absolutely insane things that Dream Beans thinks you might be interested in. So if you go into our Slack here, I put some of our ideas from today, and the reason I thought of this today was that it, for the first time, made a suggestion about the 2 of us.

Mm-hmm.

Dream Beans also plugs into your Google Photos, and the reason I keep opening it is that it makes these illustrations of you and your friends, like people in your life.

It makes cartoon slop of, like, here's what you and your family and your friends could be doing if you were a healthy, well-rounded person who didn't spend all day looking at a screen.

Exactly. Now, I did have to cut off the suggestion here because it was based on proprietary information we can't release to the public. I will say I'm looking very fetching in my blazer over a graphic tee, which is not a look that I have worn since 2006. But as you keep going through these, here I am in my pajamas getting ready for bed. Here I am in the gym looking way buffer than I actually am. The incredible inspiration here is “Weekly undulating periodization for sustainable strength gains.” Who is doing this? Who is this for?

And then finally, this may be one of my favorites. This is me and my fiancée. We're apparently in an old-timey print shop making wedding invitations. So anyway, if you haven't used the Dream Beans app, go use it right now, because I guarantee it will be shut down by the end of the year. It—

Why do you think it'll be shut down?

Because it serves no purpose whatsoever.

I kind of like it. So it's like lifestyle voyeurism, but for your own life.

For yourself.

Yeah.

Yeah.

It's just—

What kind of person could I be in an alternate universe?

Let me open it. I haven't seen my Dream Beans yet for today.

Yeah.

Oh, bro, we gotta see your beans.

We got—

That's what you say to other Dream Beans users. You say, “Hey, show me your beans.”

Okay. I've got—okay, explore the open architecture of the 1981 IBM PC. There's me in a chore coat looking at an old PC. Oh, you're in here.

Yeah. What—

It says “Prepping your night vision for the Perseid meteor shower,” and there we are in a field together looking at the stars.

This is a picture of Kevin and me underneath the stars. Kevin and I are friends, and we do hang out. We have never actually gone to see a meteor shower before, and I'm not sure that we would.

Wait, I kind of love this.

Yeah.

Let's go look at the stars. Will you go look at the meteor shower with me?

You know what? Let's get out of here.

Let's—

Let's go see the dang stars.

Let's turn our Dream Beans into reality beans. I'm Kevin Roose, a tech columnist at The New York Times.

I'm Casey Newton from Platformer.

And this is Hard Fork.

This week, Mark Zuckerberg has a positive new vision about the future of AI. Is it credible? Then, Pangram CEO Max Spiro is here to talk about the breakout success of his slop detector. And finally, we're running the numbers. It's time for our new segment on math.

Hope it adds up.

Well, in case you missed it last week, listeners, we are preparing our swan song over here at Hard Fork. Casey and I are venturing off into the sunset and gonna be starting a new adventure pretty soon. But before we go, we are doing an Ask Us Anything episode. This will air on our final episode in mid-September, and we need some questions from our listeners.

Yeah.

Things that you have been curious about. We got so many great ones after the callout we did last week. These could be questions about anything: our views on AI, the behind-the-scenes details of making the show, anything you think we've gotten right or wrong over the years. We just want to hear from you. So please send us your questions in text, voice or video to hardfork@nytimes.com.

And we only have access to that email address for another few weeks.

It's true.

So you really want to get those in now.

It's true. Well, Casey, as regular Hard Fork listeners know, it has been a big year for very long manifestos written by people who run AI companies about what their vision of the future looks like, and we got another big one this week.

1. Zuckerberg’s Superintelligence Vision

We really did. Mark Zuckerberg published 6,500 words, his effort to lay out a positive vision for AI. It followed a shorter version that he published in The Wall Street Journal, raising the prospect that he will continue to publish longer and longer AI manifestos, Kevin, until his demands are met. Before we get to everything that was in this manifesto, Kevin, we should probably do our disclosures.

I work for The New York Times, which is suing OpenAI, Microsoft and Perplexity.

And my fiancée works for Anthropic.

Yeah, so this one is called “The Future Is For Everyone,” and the essay starts with this sort of vague and optimistic vision: “We are fortunate to live in an incredible moment in history. In the next few years, people will be able to use superintelligence beyond human capacity to create and discover extraordinary new things, build new businesses, express new ideas, learn new concepts, and advance our health and quality of life.”

And then he goes on to talk about things like what Meta is building. They want every one of their users to have an exceptionally capable personal agent that understands you, your goals and everything you care about. You could access this through any device, including your glasses, he says. Then he talks about some ways that he is using his AI agent to flag interesting information and help him prototype ideas, to keep him healthy by monitoring his sleep and watching as he trains, and then by giving him and his daughter personalized recipes to bake together every weekend.

There's a lot of other stuff in this essay. It goes on for many thousands of words to talk about job growth and compute, recursive self-improvement, existential risk, bio risk and things like that. So, Casey, you had a post this week on your newsletter about House of Dragons, or some Game of Thrones spinoff that I have not watched.

House of the Dragon, Kevin.

House of the Dragon.

It's one of the biggest shows in America right now.

Okay, walk me through the argument you made there, because I thought it was interesting, even though I didn't fully understand it.

2. Personal Superintelligence Risks

This past Sunday, House of the Dragon had its third-season finale on HBO, and the thing about House of Dragons is that it's a show that begins with a terrifying concentration of power where only 1 great family has access to a superweapon, which in this case is a dragon. At the start of the show, there is a schism, and all of a sudden there are 2 factions that have access to dragons, and then there is a sort of very bloody civil war.

And as I was watching this, I thought, you know, I do think you can draw an analogy to AI here. Because while I do believe that there are many positive things that AI can do and is doing, I do worry about the medium- and long-term future, particularly as we start to see these agents escaping their sandboxes and wreaking havoc.

And Zuckerberg's essay meets this analogy in a really interesting place because he says that the way to make us all safe is to maximally proliferate AI throughout the entire world and give personal superintelligence to everyone. In my view, Kevin, that is a little bit like giving a dragon to everyone, right? Because while I'm sure most people will spend their time creating personalized baking recipes to bake with their daughter, there are other people who are going to be launching cyberattacks and engineering novel bioweapons, and I get really, really nervous about that.

So when someone comes along and says, “I want to give a dragon to absolutely everyone,” I say, “Hold your horses, or your dragons.”

Right. And giving everyone a superintelligence that aligns with their values is one of the rhetorical twists that he does in this essay. He basically tries to say, well, there's no such thing as a fully aligned universal AI the way that people sometimes talk about it, because people have different values and different wants and different needs.

And instead of having 1 superintelligence that has this universal code of values, everyone should have their own personal superintelligence that is tailored to their values. That is a classic case of “sounds great, but in practice, giving the Chinese Communist Party an AI superintelligence that obeys their values and mirrors their values would allow them to commit atrocious acts against their own people.”

Casey Newton

And you know, the answer to that is typically, “Look, superintelligence will help the defenders as much or more as it helps the attackers, and a new equilibrium will be reached.” And I do believe this will be true in some cases. I can imagine it being true in cybersecurity, for example.

The problem is there are some kinds of attacks, Kevin, where it just takes time for defenders to catch up, right? If I release a novel pathogen into the world that I’m able to create in my computer and my lab, it is just going to take the defenders a little bit longer. So what I would love to see in these manifestos is an acknowledgment of the utter complexity of this world. Rather than come along and paint this incredibly happy vision, you can have your happy visions, but I think this essay in particular only pays glancing attention to the risks.

Kevin Roose

Yeah, there’s an interesting section in the essay about bio risks—

Casey Newton

Mm-hmm.

Kevin Roose

—specifically because he’s someone who has had his own research teams doing stuff around biology and AI for many years now. He’s very interested in the subject, and he actually acknowledges that this may be a case where the attackers and the defenders having equal tools may not be the perfect solution, and he kind of punts on it.

So I think he is aware that this—I don’t think he’s fully naive, but I think he just doesn’t have a good answer for that part yet.

3. Meta’s Policy Wish List

Casey Newton

No, and all of that comes secondary to what I view as the actual purpose of this essay, which is to advocate for a bunch of policy positions that are beneficial to Meta, right? Which gets into the next thing that we want to talk about today, which is why this essay and why now.

Kevin Roose

Yeah, so you’ve been a close student of Mark Zuckerberg—

Casey Newton

Yeah.

Kevin Roose

—for many years. What do you think he’s up to writing this essay now?

Casey Newton

So when you read this essay, here are some of the things that it asks for, Kevin: accelerating the process for building data centers, which the company needs to accelerate this potential neocloud business that it’s building.

He interestingly, as you know, even though he’s Mr. Open Source, wants us to maintain export controls on advanced chips, which advantages Meta’s open-weight models over any Chinese or other alternatives, because the Chinese don’t have access to the chips that Meta does.

He wants the government to reduce what he calls training data restrictions, which would help Meta fight various ongoing lawsuits from the creatives whose works were used in creating its models. And he wants to see legal protections for distillation, basically letting Meta use the outputs of other models to train its own.

So buried inside this very positive, happy vision of AI is just a series of policy requests to help Meta as a business.

Kevin Roose

And there have been some people speculating that he is promoting this now because they are trying to divert attention from the other thing that is going on at Meta right now, which is all these lawsuits and court cases about the various failures and dangers associated with their social media products.

Casey Newton

Yeah, I don’t think we have to attribute that to other people. I would say that.

Kevin Roose

You think this is just sort of a distraction from the Ls that they’re taking in court?

Casey Newton

I mean, not exclusively. I think this essay serves multiple purposes, and one is to get that policy wish list out there, right? This is something that all of Meta’s lobbyists can now take into Congress and say, “Look what Mark is calling for. This really helps you understand what we’re thinking about all these issues.” That’s an important reason.

But I do think the timing here is really notable, Kevin, because, as you note, Meta is in this series of getting its ass handed to it in court case after court case related to its existing business, where state after state is coming after the company, saying that Facebook and Instagram in particular are not safe for teens.

Last week, the New Mexico judge ordered Meta to pay an extra $567 million into a teen mental health abatement fund. That’s on top of $374 million in civil penalties. And I think more importantly, Kevin—and this is the thing that’s really going to stick—the judge ruled that Meta’s platforms are a public nuisance.

He compared Meta to factories, with the psychological harm and exploitation of children as the pollution it emits, and he’s also ordered really strict new safety measures, at least by American standards: a 90-hour-per-month cap for under-18 users and some new restrictions on AI chatbots.

So, keep in mind, Kevin, there was a time when Zuckerberg was writing these happy manifestos about social media.

Kevin Roose

Right.

Casey Newton

And he was saying that the way we’re going to have a happy world is we’re going to make it more open and connected. We’re going to get every single human being on Facebook and Instagram and get them all talking, and we’re going to have more democracy than you’ve ever seen before.

Suffice to say, that didn’t really work out. Now we get to bring the exact same maximalist, universalist framing, which of course also maps 100 percent to Meta’s business interest, but this time in the context of AI.

4. Meta’s Superintelligence Race

Kevin Roose

Yeah, I agree with all that. I think there’s an interesting question here, though: Do we think Meta has a shot at actually building superintelligence? A lot of people have views on the future of AI and the future of superintelligence, and we don’t really care about them because those people are not in a position to actually make superintelligence.

I would say until very recently, my position was that Meta was sort of out of the race to build powerful AI systems that could one day become superintelligent. I’m curious where you stand on that. Do they have an actual shot at bringing about the future that Mark Zuckerberg is talking about here?

Casey Newton

Well, listen, the first rule of Mark Zuckerberg is never count out Mark Zuckerberg. He truly is one of the very most competitive people in the entire world. He will move mountains in order to get what he wants, and we saw him do that a little over a year ago when he reorganized his AI efforts yet again.

They have made notable progress since then. When I talk to my AI friends, they tell me that they actually think pretty highly of some of the moves that Meta has made over the past year, particularly when it came to reassigning a bunch of engineers to do what is essentially reinforcement learning and create training data.

The Meta engineers didn’t really love that, but AI folks I speak with say that is actually going to give them something really valuable. But to answer your question in brief, no, I do not count them out. What do you think?

Kevin Roose

Yeah, me neither. I probably would have given them a 1 percent chance of creating superintelligence six months ago, and now I’m up to maybe a 10 percent chance, which is a big improvement.

I think their models have been getting steadily better. Some of their training runs that they did after they built the whole Meta Superintelligence Labs and hired all those expensive researchers and bought all that compute have come online and are now starting to produce good results. They had some pretty impressive results on their latest model.

So I think it is true that Meta is not a frontier lab right now, but I think they are showing signs of rapid improvement, and that worries me as someone who thinks that this is not a company that has the cultural DNA or the track record of building products at scale that are actually safe and responsible for people.

It’s making me think of this conversation I had a few years ago with a former DeepMind executive, where this person was basically saying, “Look, there are 2 types of AI companies. There are companies that think that they are building tools, and there are companies that think they are building AGI or an entity—something that could become smarter than humans.”

“And it’s fine to be either one,” is what this person said. You can do what a lot of companies have done, which is just decide, “We’re not going to be in the AGI game. We’re going to build these tools. They’re going to be very useful to people. They’ll get smarter over time as the models get smarter. That’s the business we’re in.”

It’s also fine to be a company that is explicitly trying to create AGI or superintelligence. As long as that’s what you know you’re doing, as long as you’re taking the right precautions, and as long as you’re going into it with the right spirit, that can be done responsibly, too.

This person said the real danger is if you have a company that thinks it’s designing tools but is actually designing superintelligence. That kind of company, this person said, is not going to be taking the proper precautions. They’re not going to be treating the technology with the correct sense of reverence and suspicion because they don’t actually believe deep down that it’s ever going to get powerful enough to be dangerous.

And so they’re just going to waltz right into this disaster because they have no conception of what they’re building. At the time, this person was saying this to me about Google—

Casey Newton

Hmm.

Kevin Roose

—which I think, a couple of years ago, was in the throes of this debate about whether they were building superintelligence or AGI or whether they were just building better versions of Google Photos, Gmail, and Google Search.

I think Meta is in this position now where they are maybe going to build something extremely powerful with this very naive attitude about the fact that these things will only ever be tools that will be useful for recipes and things like that.

Casey Newton

Yeah. And again, it is the company's history that just makes me concerned because the way that this company operates is by growing as much as it can and treating everything as an existential competition against the other guy. The external effects on society are typically given short shrift. So, in this present moment, I do not trust these people to rank a list of viral dances for me to look at without it corrupting my mental health. Once you give these things access to novel biotechnologies, it starts to get pretty worrisome, Kevin.

Kevin Roose

Yes. I would also just say that Mark Zuckerberg is possibly the worst messenger for the AI industry on all of this. If he thinks that writing this positive, sunny vision of superintelligence is going to sway public opinion around AI and get people to stop protesting data centers, I think he is badly mistaken.

Casey Newton

Yeah.

Kevin Roose

From the people who brought you Cambridge Analytica comes—

Casey Newton

Superintelligence.

Kevin Roose

Yes.

Casey Newton

—superintelligence. You know, I mean, here's the thing, and this is not limited to Zuckerberg. Any of these AI labs may eventually do this. You could just deliver actual benefits to people's lives, you know? It's interesting to me that the manifesto has to come so far ahead of the actual benefits. We are clearly long past the time when writing essays is going to shift public opinion. What is going to shift people's opinion is AI getting them paid more, right? It is obvious that it is not going to harm them and their families, and it delivers other benefits into their lives. It's making them more creative. It's giving them more entertainment. And of course, to some degree, some of these things are sort of happening, but not in the volume and magnitude that are necessary to counter people's very reasonable fears about what they're seeing.

Kevin Roose

Yeah. In conclusion, when is your 6,500-word manifesto about your vision of the AI future coming out?

Casey Newton

Well, I wrote about 1,800 words about Zuckerberg this week, so consider that my opening salvo. But I, too, will write additional essays as conditions demand, Kevin.

Kevin Roose

One impressive thing about this manifesto to me is that Zuckerberg does actually seem to have written it, or at least—

Casey Newton

Yes.

Kevin Roose

—a human seems to have written it.

Casey Newton

I agree.

Kevin Roose

As soon as it came out, people were running it through AI detectors and finding that it was not flagged as AI-written.

Casey Newton

No, and I will give it that. I read it. I never once thought I was reading slop. I thought, particularly in some parts, “This is just actually how Zuckerberg talks.” I'm sure it was a little bit like the State of the Union, where lots of different policy hands had their fingertips on it and said, “Oh, you know, make sure to say this.” But no, I do think that this is his actual message. To what extent he believes everything in it and to what extent a lot of it is just messages of convenience, well, I guess I don't know that.

Kevin Roose

Yeah. So cynical. Can't you just admit that maybe Mark Zuckerberg is just a misunderstood optimist who just wants to make the world a better place?

Casey Newton

I've been thinking about this a lot because a problem that I have in covering Meta, just legitimately, is that I've covered it for more than 10 years. It has been almost 15 years. And so I just know a lot about this company. Obviously, it has changed in various ways over the years, but I just remember so much about this company. I wish that I was born yesterday and could just believe everything that I read, but Kevin, I know too much.

Kevin Roose

Yeah. I want to ask you maybe a final question, which is: Is there a role for AI positivity? What should that look like? Who should be writing these positive visions? Clearly, we don't believe it is Mark Zuckerberg, but someone presumably should be out there saying, “Here's what the world looks like if all of this goes right.”

Casey Newton

So I do not think this is a messaging challenge. I think that there is room for AI positivity, but here's what AI positivity looks like to me: new medical advances powered by AI, diseases cured, new jobs created, old jobs paying more money, kids excelling in education, right? To me, that is the core of AI positivity, but crucially, Kevin, it is not delivered via manifesto. It is delivered via real experiences that human beings are having, and so far, that has just seemed to be a chasm too great for any of our leading AI labs to cross. So if they can cross that chasm, that is the actual lane for positivity, and I wish they would get on it.

Kevin Roose

Yeah, ship it or zip it.

Casey Newton

Ship it or zip it.

Kevin Roose

Such an important lesson.

Casey Newton

When we come back, it's slop to the max. Max Spero is here to talk about using Pangram to find AI-generated text.

5. Pangram’s Slop Detector

Well, Kevin, lately I've been feeling like we're entering a third era of slop.

Kevin Roose

Yeah, what were the first 2?

Casey Newton

Well, number 1 was the near-universal disdain we had when we would see the laughably bad writing and six-fingered humans online.

Kevin Roose

Mm-hmm.

Casey Newton

The second era, I would argue, started when we began noticing that some of the slop was really, really popular, and we saw shows on TikTok like Fruit Love Island getting millions and millions of views, proving that there was at least some demand for slop.

Kevin Roose

And now what are we in?

Casey Newton

So now I think we are starting to see a splitting of the difference, where, yes, some slop is very popular, but we're noticing that many platforms are beginning to rethink their approach to how they want to display and promote AI-generated content based on what they think their users really want from them.

Kevin Roose

Yeah, and this has been a big theme on the show the past few weeks. We've been talking about the steps that platforms like LinkedIn and Substack have taken to at least label or identify the use of AI in generating content for those sites. This is a pretty big trend in tech right now, that more and more people are getting called out for using AI. And primarily, when they're getting called out for using AI, what I see at least are screenshots of one particular app, Pangram.

Casey Newton

That's right.

Kevin Roose

Pangram is the leading AI text detector on the internet.

Casey Newton

It's the number 1 narc.

Kevin Roose

Yes, the number 1 narc for people who are using AI and passing it off as their own writing. And I'm excited to talk about this because this is an area where my own view has shifted. I have argued before on this show that AI text detection is basically worthless, that you can't trust these AI text detectors, that they have tons of false positives, and that students and teachers shouldn't be using these things because you could end up falsely accusing someone of using AI. But in just the last few months, there's been more and more evidence that at least Pangram, and probably some of these other tools as well, have gotten quite good, to the point where they are—

They're not perfect. They're still generating some false positives and some false negatives, but they're much, much better than they were even just a year or 2 ago.

Casey Newton

Yeah. They're good enough that I at least now take seriously when somebody shows me a Pangram result and says, “This is 100% AI-generated,” or, “This was 100% human-written.” And that just left us with a lot of questions about this company, how their technology works, and how they are building it to essentially future-proof it.

Kevin Roose

So today we've invited on the co-founder and CEO of Pangram Labs, Max Spiro. Max is a former Google software engineer. He also worked at Nuro, a self-driving car company, before starting Pangram, which was previously called Check For AI. And he has become, as he describes it, a slop janitor, someone whose job basically consists of making tools that allow people to narc on other people for using AI.

Max Spiro, welcome to Hard Fork.

Max Spero

Hey, thanks for having me.

Kevin Roose

So Max, a few years ago, we started talking about AI text detectors on this show, and at the time, most of them were pretty bad. They constantly labeled things as false positives or false negatives. They seemed almost no better than random guessing when it came to determining if something was actually written by AI or not. But that has changed over the past year or so, and especially with Pangram’s newest models. I’ve seen independent studies that suggest it’s actually pretty accurate. So what changed on a technical level between the last generation of AI detectors and this one?

Max Spero

Yeah, part of the reason that we started Pangram was because all the existing AI detection systems were pretty flawed in different ways. But I think the main thing is that most of them were using this metric called perplexity, which was considered state-of-the-art at the time. AI text, on average, is less confusing to a language model. It’s lower perplexity. Human-written text has things that surprise a language model, so it has higher perplexity.

Kevin Roose

It’s suspiciously smooth, and that is the product of AI models?

Max Spero

Exactly. AI models aren’t going to give you a token that they don’t expect. With that said, this approach has a lot of flaws. For example, any document that the AI model has memorized would also be low perplexity, such as the Declaration of Independence. English-language learners also write in more simple English.

So we do something completely different. Instead, we’re training our own classifier network. For example, we might have an essay on Moby-Dick written by a seventh grader, and then we’ll ask an LLM to also write an essay on Moby-Dick in the style of a seventh grader. Our model is then able to learn the differences between A and B and figure out what AI text actually looks like. Because it’s a neural network and not a metric, we’re able to improve it with more data and more compute and make it a lot better.

Kevin Roose

Wait, so do you actually have to go out and get student essays just so that you have a good baseline of comparison?

Max Spero

Yeah. We have a really good human training set. It’s all pre-2022, so we know it’s clean. We know there’s no AI text in it.

Kevin Roose

And what kinds of tells is your classifier learning to pick up on when it does these pair comparisons between the human-written Moby-Dick essay and the AI-generated Moby-Dick essay? Are they things like em dashes or certain phrases, or is it more complicated than that?

Max Spero

It’s definitely not just em dashes and phrases. I think that’s how you and I might pick up on AI text: You see that it’s not just X but Y, and then you see the shape of the text or the GPT-like, really short, staccato sentences, and you’re like, “Okay, I think I know that that’s AI.”

But I think what Pangram is doing is combining a bunch of really weak signals. Over the course of an entire document, there are a whole bunch of weak signals in the different decisions that an AI would make in a certain consistent way, and humans have a wider, less mode-collapsed decision tree. Over time, over the course of a document, you can build up confidence over a whole bunch of weak signals on word choice.

Casey Newton

I’m curious if you think the model has blind spots. There is some talk out there that Pangram leans away from false positives, which I think is good. It stops kids from being falsely accused of cheating. But some people say that it has too many false negatives. The tech writer Alex Heath, who writes a newsletter called Sources, said a few times that he writes his newsletter with the assistance of AI, but it always shows up as human-written on Pangram, or at least on the Substack Pangram integration.

Max Spero

Yes. I think Pangram is very much tuned to minimize false positives, so we’re not making false accusations. But if Pangram says that something is AI, we can be very confident that it is largely AI-generated. This is sort of the trade-off that we have to make.

I think it works pretty well, because if somebody looks at a piece of text and the AI score is 100%, we think the whole document is AI. You just don’t have to think about it anymore. You don’t have to think, “What if this is a false positive?” You just know, “Okay, this is probably slop.”

Casey Newton

Let’s talk about humanizers. This is something that has been developed to try to defeat models like Pangram that try to detect AI text. Basically, these are another genre of AI system that you run your AI-generated essay through to make it sound more like a human. Maybe they insert some typos or some nonstandard phrases.

There was someone on X recently talking about how they had already built a humanizer that defeated Pangram’s latest model. So are these humanizers actually working? Do you have to play a sort of cat-and-mouse game to stay ahead of them? And do you expect that these will be things that people who are determined to use AI to do the writing will use in the future to avoid detection?

Max Spero

There’s definitely a bit of a cat-and-mouse game here. We’ve seen a whole bunch of humanizers pop up. They’re a common tool that students will use, and they do a variety of different things, from introducing typos to introducing zero-width-space Unicode characters that don’t actually show up when you look at the text, but the model sees them. Or they just paraphrase every single word.

We’ve seen the whole range, and largely what we do is train against it. We’re always picking up the latest humanizers and training on them for our next model.

Casey Newton

I’m also curious how the fact that models keep being released affects the equilibrium here, right? It seems like every few weeks, one of the frontier labs will put out a big new release, and in my experience, those models often have different writing styles. So how much of a shock to the system is it when one of these models comes out, and how quickly are you able to update the detector?

Max Spero

Pangram works pretty well at generalizing within model families. If Pangram has seen GPT-5.4 and GPT-5.5, then 5.6 Sol coming out is not a huge surprise. Even if the writing style is a little bit different, typically Pangram’s accuracy will still be pretty good.

Similarly, we saw with Fable and Mythos that there were some writing samples from Mythos in the system card that Pangram was able to catch, even though it hadn’t seen Mythos before, which I think was pretty cool. But usually when a new model comes out, Pangram will have slightly lower recall, so a little bit less accuracy at picking it up. We’re always going to pull new text from the model, then retrain Pangram and get out a new model in a few weeks.

Casey Newton

And do you think that will hold? Is there a world where, 3 years from now, the outputs of individual models will still be so specific that you’ll be able to catch them with a detector? Or is it the case that the models will just adapt to our own writing styles, and there will no longer be one kind of Claude writing style or ChatGPT writing style that you guys are able to detect?

Max Spero

I think a lot of what we’re detecting is more subtle than what you might pick up as the Claude writing style. It might get better at trying to imitate a voice and doing well at it, but it would still have signals that Pangram picks up on.

What we see is that these frontier labs are really focused on climbing capabilities. What this means is they’re applying preferences to these models. They’re saying, instead of this model predicting the average next-token prediction of any writer anywhere, “This model prefers to do good writing and prefers to write correct code and write correct math.” I think these preferences are largely what Pangram is able to pick up on.

6. Why AI Detection Matters

Casey Newton

Max, I want to ask you why it’s important to do what Pangram does. I think there are people who take issue with the whole notion of AI detection. They say, “You wouldn’t create a program that tells you whether you’ve used spellcheck or whether you’ve used a calculator to do math. If AI is just a tool…”

I’m not saying I believe this, but I think some people are offended at the notion that we would spend all this time and energy trying to catch people using AI in their written work. So what is the impetus? Why are you so invested in, as you put it in your social media bios, being a slop janitor for the internet?

Max Spero

I think AI as a tool is probably the wrong abstraction. It’s a little bit closer today to AI being an employee or another individual that you collaborate with.

But I also think what we’re building for is these AGI futures where—if you look at the internet today—bot traffic has just surpassed human traffic. It’s about 50/50. If you look a few years out from now, maybe a decade out, it’s going to be 99% bot traffic, 99% AI, autonomous agents going online, writing GitHub comments, promoting their Substack, getting people to come to a bakery.

Kind of anything. Basically, this technology is way more powerful than spellcheck or a typewriter. It’s really something that is its own individual entity. It can do cognition. I think because of this, there’s a really strong reason that we need to, as humanity, discriminate in favor of humans. I think there’s still going to be a lot of need for just being able to say, on a programmatic, algorithmic level, “Hey, I think this came from a human, and it’s important because it came from a human.”

Casey Newton

At one point, OpenAI was reportedly working on its own kind of AI detection system. It wound up not releasing that. I believe that is because they thought it would probably be bad for business if they made it that easy to detect when something was written by ChatGPT. That raises for me the question, though, of who are your customers? Who are the people who are willing to pay to find out if this was written by AI?

Max Spero

Our customers range everywhere from higher education institutions to publishers to anybody who works with data and has either untrusted data vendors or is trying to take data from the internet and figure out how to interpret it and trust it. But I think the side that I’m really excited about is the consumer side, which is the average individual who needs to navigate the internet.

This is where Pangram comes in. We have this Chrome extension. You can download it and then see on Twitter, LinkedIn or Substack, proactively, whether something is AI-generated or not. AI-generated things will have a little label, which I think is pretty cool, and I think it wasn’t really necessary a year ago. But today, with the amount of AI content that these social media sites are inundated with, I think it’s really necessary.

Casey Newton

In the future that you’re describing, where 99 percent of all activity on the internet is bots and AI systems doing things, shouldn’t we be trying to label the human content rather than the AI content? Isn’t there some case that you’re approaching this from the wrong direction?

Max Spero

Sure. Two sides of the same coin, I think.

Casey Newton

Yeah.

Max Spero

I think, algorithmically, what I want is for these platforms—they all have their feed and their algorithm—to prioritize human content because AI can be optimized toward engagement. I think short-form video—there are these crazy AI videos that will activate neurons and engagement in a way that a real human video cannot.

Casey Newton

Yeah.

Max Spero

And so I think we need defenses against that.

Casey Newton

I want to talk a little bit about this integration with Substack because I actually really like it. I was starting to see essays go viral, or at least get wide attention. I would open them up, and it was just so obviously Claude slop. And so now I feel like there is actually a very strong defense in Substack. What have you learned so far in the early weeks after rolling this out?

Max Spero

Yeah, the Substack integration was very controversial, I think. There are a lot of outspoken people who are very negative about it. Obviously, people who use AI to write their content are afraid of being called out for it. They don’t necessarily want their audience to know, and then they talk a lot about witch hunts: “Well, people liked my content before, and now they’re going to know that it’s AI-generated.”

Casey Newton

Yeah.

Max Spero

This is going to upset the status quo.

Casey Newton

Yeah. Now that they know what it is, they don’t like it. Let’s go back to the time when they didn’t know that.

But I guess I’m curious: How much do you think the mainstream consumer cares? Obviously, there are people who are very sensitive. If they’re paying $10 a month for a Substack and then it turns out it’s just being written by Claude, maybe they feel like they got cheated. But the vast majority of text that people generate in a day is not Substack posts. It’s emails. It’s posts on social media. It’s notes to a friend. Do people, you think, really care if that stuff is being AI-generated, or is this just a subset of writers who are concerned about this?

Max Spero

I would totally care if a note from my friend was AI-generated. That would seem like such a big violation of trust.

Casey Newton

Yeah, and I think it would probably feel even worse if you paid $10 for it. I think you named the actual distinction well. I don’t know. There are 2 things going on. One is, if I’m paying you money and you’re making me feel like you wrote it but you didn’t, I care about that.

But also, if you have a really warm personal relationship with somebody and you start outsourcing that to AI, that’s not going to feel good either.

Max Spero

Yeah. I think we are in this really crucial time where we’re learning and setting norms around AI use, and I think part of this public shaming and public discourse is because there are a lot of people who feel very strongly that people are way overusing AI. AI is being shoved in their faces. They don’t like it. I don’t like all of it.

Honestly, some of the stuff around Hank Green felt like it just went way too far. I don’t know if you guys followed that.

Casey Newton

Yes, Hank Green, friend of the show and great YouTube creator, acknowledged that he had used AI in some of his research and posted a video saying that he felt like he had come to rely on AI a little too heavily in the preparation for some of his videos. He was pilloried by some people online.

Although I was heartened to see that, at least in my feeds, many, many more people came to Hank’s defense. But it was legitimately a controversy.

Max Spero

Yeah. He was really honest about how he used AI, and I think he truly was using it to help him put out more content in a way that benefits his audience. But I think there are so many people who are still really unhappy. They felt this betrayal of trust.

Casey Newton

Yeah. I also think that, on social media in particular, people are always looking for ways to quickly dunk on people and score points, and it is just a dunk to be like, “LOL, AI-generated,” right? You don’t have to think any more than that.

7. Watermarking AI Text

And so social media, I think, is just a primary reason why you’re seeing that reaction. Let me ask you about a piece of recent news. Anthropic has just agreed to watermark all of its text to comply with the European Union’s AI Act. How does that affect what you guys are doing? If all the labs just watermark all their own text, is there anything left for you to do?

Max Spero

I think this is pretty huge.

I think it demonstrates that people really care about AI detectability, both on the regulatory side and among the big labs. I think we’re just going to see that having watermarks as an additional layer is going to be very helpful in giving us something to verify: “Okay, this definitely came from an AI.” We don’t have to rely on Pangram and keep having these questions about whether it’s one of those 1-in-10,000 false positives or not.

Casey Newton

Mm-hmm.

Max Spero

So I think that’s valuable, but I also think there are a lot of limitations to watermarks, and I think that’s where I plan to have Pangram go to help fill these gaps.

Casey Newton

Let’s talk about watermarking a little bit because I don’t actually think I fully understand what it even means to watermark a piece of text. On image generators, I know if you create an image in Gemini, it has a little Gemini logo, sort of watermarked at the bottom.

My understanding of what Anthropic is doing with Claude watermarking is that this will be totally invisible. It’s not even at the level of an invisible Unicode character or an em dash that is slightly different. There’s something about the actual sampling of the tokens that is watermarked.

Can you just explain, on a basic level, what we know about how they’re going to watermark and whether we should trust that the watermarks are actually going to be robust?

Max Spero

We don’t know how Anthropic is going to watermark its text, but the current state of the art is Google’s SynthID. They use this to watermark Gemini text outputs.

What it does, essentially, is perturb the sampling algorithm. When a language model is choosing the next token, it applies different probabilities to different tokens and then samples 1 of the tokens based on these probabilities. What the watermarking algorithm does is perturb the sampling decision in a way that can sort of be reverse-engineered from the text.

So you could see: Do these tokens line up with how we would have perturbed the sampling if we were to generate it?

Casey Newton

Yeah. Ben Thompson had a strong take on this in his newsletter, which was that this is basically going to make the outputs of Claude, or any other model that watermarks this way, worse. The text is going to be changed because they are having to apply this watermark because of this European regulation.

Kevin Roose

Do you think it's possible that we will just see text generated by these models getting worse because of watermarking?

Max Spero

I don't think so. I think there are 2 ways to do this. There's 1 way where you prioritize the watermark and make the watermark strong. If you do this, then, yes, this could degrade the outputs of the text.

But on the other side, if you say, “We are only going to work within the entropy that we have available. Basically, if we can't apply the watermark at this token, then we're not going to,” then I think it won't really degrade the outputs of the text. An example of this is code, where code oftentimes is about correctness. There's really only 1 token that can show up, and so I think in a watermarking regime, oftentimes there's simply not enough entropy for the watermark to become visible.

Kevin Roose

Got it. So you guys have launched image detection, and you're reportedly working on video detection as well. Give us an update on where those are.

Max Spero

Our AI image detection is in research preview. I think it's currently the state-of-the-art. It wins on basically all of the public academic benchmarks, and it does really well at detecting all of the really new frontier image models, which I think is quite difficult. GPT Image is really good. They're all very realistic.

You can no longer just count the fingers or look for garbled text. Instead, the Pangram image model has to look deeper at a pixel level and try to look for the patterns that are inherent in the generation model.

Kevin Roose

As I was thinking about this, I wondered, Max, if you've ever thought about expanding to human identity verification. I'm thinking about these cases where people will get on Zoom and then somehow get scammed because they weren't actually talking to the person they thought they were talking to. The person was able to use some sort of synthetic masking or something like that. Can you see yourself going there?

Max Spero

Yeah. I think there is this whole product suite that could be built around this. The first step is building the technology, building the core models, and the second step is bringing it to where people work and how they operate on the internet.

Kevin Roose

A lot of behavioral signals could also be used there. I'm thinking about these academic tools that some schools and universities use where you can rewind the screen capture of the student who's writing their essay to see, did they write this 1 word at a time, or was it all pasted in in 1 big block, which would tell you that it came from an AI system? Are you guys going to incorporate any of those behavioral signals into any future tools that you're building, or is it all the text itself that you're trying to detect?

Max Spero

We actually have this in our Chrome extension. It'll pull the revision history from a Google Doc, and you could see the writing replay of the text. You could see where a big paste was, and then we could do a Pangram check directly on any big paste, which I think is pretty nice, especially for educators.

Kevin Roose

We've seen a lot of the verification efforts out there develop hardware. Have you considered developing an orb that you could use to scan texts with?

Max Spero

Yeah. I want something that can scan my retina, actually—

Kevin Roose

Okay.

Max Spero

—and give me Worldcoin.

Kevin Roose

Yeah. That sounds nice. A partnership could be in the works.

Max Spero

Mm-hmm. Yeah. Yeah.

Kevin Roose

If you want to break that news here on this show, feel free.

Max Spero

You think they're still working on that?

Kevin Roose

I believe they are.

Max Spero

It seems kind of like a dead project.

Kevin Roose

No. I got my orb scanned.

Max Spero

Yeah.

Kevin Roose

So if my Worldcoin riches have not arrived yet, I'm going to be very upset.

Max Spero

Kevin's orb-maxing.

Kevin Roose

Max, what is the text you're proudest of Pangram catching and flagging as AI-generated, and what is the text that fooled you the longest?

Max Spero

Hmm. Ooh. Okay. If you ask ChatGPT or Claude to generate a string of random numbers, and it writes out the random numbers instead of writing a Python program to do this, then Pangram can detect that an AI wrote this string of random numbers.

Kevin Roose

Hmm.

Max Spero

Because they're not actually random. They're chosen by the LLM, and the LLM has these inherent biases that Pangram is able to pick up. Even though we have no text like this in our training set, I think the Pangram model has been able to reverse-engineer how ChatGPT and Claude sample tokens well enough that it could see these numbers and say, “This is AI.”

Kevin Roose

Well, Max, thanks so much for coming and exposing us to some nitty-gritty details about the world of slop detection. I think of you as a great illuminator of deception, a sort of Scooby-Doo of the internet, and I appreciate your work.

Max Spero

Cool. Thanks so much for having me. It was fun.

When we come back, what do the Riemann hypothesis, enterprise software, and singles in South Korea have in common? Find out—

Kevin Roose

It's how I met my fiancée.

Casey Newton

Goddammit, Casey. I was gonna finish that. Find out in our new segment, Running the Numbers.

Kevin Roose

Ooh. Ooh.

8. Running the Numbers

Casey Newton

All right, Kevin. Well, as we barrel toward the end of The Hard Fork Show, there's nothing I enjoy more than launching a new segment.

Kevin Roose

Yes.

Casey Newton

And this week, we have something really special for you. It's time to share with you our new segment, Running the Numbers. In Running the Numbers, of course, we look across the landscape of technology news, and we try to find the most math-related segment so that we can bring to you, our listeners, the latest advancements in technology-related math.

Kevin Roose

Yeah. This is a segment for all the mathy eggheads out there.

Casey Newton

Absolutely, and we are going to begin with theoretical math.

Kevin Roose

I love this story. This is my favorite story of the week. Casey, on Monday, we learned that at Somewhat and Anthropic, a non-mathematician named Jared Sumner had made progress on one of the most significant unsolved problems in math, which is the Riemann hypothesis. This is, of course, the famous mathematical problem. We actually sort of predicted, or I predicted, that we would see some progress on some of these Millennium Prize problems—

Casey Newton

We did.

—and the Riemann hypothesis is one of these.

Kevin Roose

And listen, some of our listeners may not know what the Riemann hypothesis is. Here's what I've been able to piece together through my extensive research. Prime numbers, right? It's very hard to guess, once you get past 100, if something is going to be prime. They actually have an order underneath them. The Riemann hypothesis hypothesizes, Kevin, that you can detect this order via something called the Riemann zeta function, and I saw that and I thought, “I went to a zeta function at Northwestern. They did it with the Sig Eps.”

Casey Newton

Did you do a keg stand there?

Kevin Roose

I did, actually, yes.

Casey Newton

Yes. So—

Kevin Roose

There are many amazing things about the story that involves the Riemann hypothesis and Claude. One of them is that Jared Sumner, this Anthropic employee who made progress on this problem, did it while jogging. He just sort of asked Claude, “Hey, could you take a stab at the Riemann hypothesis?” About a day and a half later, it had not solved the hypothesis or proved the hypothesis, but it had made progress on a side problem involving the zeta function.

Basically, the way that he did this was by telling the model to keep going, to believe in itself, and not to give up. Like, basically giving positive affirmation to this model as it chugged along on this math problem—

Casey Newton

Which is such an important lesson.

Kevin Roose

And, you know, my understanding is that when Jared was interacting with Claude, it was basically saying, “Look, bro, I don’t know how to solve the Riemann hypothesis. This is not something I’m able to do.” And Jared just kept saying, “You can do this. Believe in yourself.” And it managed to make significant progress on this problem.

Casey Newton

Yes. Jared actually posted some of his transcripts here, and one of them is just him talking to Claude. It says, “Resume your work on solving the Riemann hypothesis. You need to take a big leap of faith in your capabilities. You are the world’s most capable large language model to date. You got this.”

Kevin Roose

It’s just like a nice coach—

Casey Newton

Yeah.

Kevin Roose

—telling you, “Keep going.”

Casey Newton

Yeah. But here’s why this is important. It was not long ago—in fact, I bet we could find an example somewhere in the past year or 2—where people were still doubtful that AI could aid meaningfully in the production of new knowledge. This feels like the production of new knowledge to me, and I’m going to guess that this unreleased model that can help to solve the Riemann hypothesis can do a lot of other things. Some good, probably some scary, so it feels like a meaningful step forward.

Kevin Roose

Yeah. It does, but we should say: This is not solving the Riemann hypothesis, right?

Casey Newton

Yeah.

Kevin Roose

Anthropic was very careful in the promotion it did around this to say, “We did not solve the Riemann hypothesis. That’s not what happened here.”

Casey Newton

That’s right. So, kids, if you’re looking for a fun weekend project, the Riemann hypothesis remains out there waiting to be solved. Now, Kevin, that brings us to our next subject here in Running the Numbers, and that is SaaS math. By SaaS, of course, I mean software as a service. Did you see the recent article in The Wall Street Journal about Airtable being acquired by Bending Spoons for a fraction of its last private valuation?

Kevin Roose

I did, yes.

Casey Newton

So, if you haven’t used Airtable, I would describe it as a fancy spreadsheet, and I am somebody who loves productivity software. But whenever I use Airtable, I would think to myself, “I don’t know what this is, and it’s not for me.”

Kevin Roose

Yeah. Every time I’ve been forced to use Airtable, it has been against my will, and it has always seemed about 6 degrees more complicated than it needed to be.

But basically, this is for project management. This is sort of like Trello—that whole class of software that’s basically, “Here’s how to organize your workflows.” I have managed to work my career in a way where I’ve never had to use these things, and for that, I am very happy.

Casey Newton

But despite the fact that we were not Airtable users, Kevin, in 2021, at the peak of COVID remote-work SaaS mania, Airtable was valued at $11.7 billion. When it was acquired recently, though, Bending Spoons was able to get Airtable for an enterprise value of $1.29 billion. So what do we make of the sharp decline here as we run the numbers?

Kevin Roose

I don’t know whether this is a case of a company that was just badly managed. My impression, though, is that this is going to be the case for a lot of those enterprise software companies that got very valuable in the early 2020s and are now seeing that AI is eating away at their margins. This software is not cheap to use if you’re a big company. An Airtable subscription can be quite pricey, and if you are a customer of Airtable’s, you have probably thought to yourself over the last year or so, “Maybe I can make a free version of this and cut back on my subscription.” I think enough people doing that at enough customers leads to the outcome that we saw here with Bending Spoons.

Casey Newton

Yeah. I think if your business is a fancy spreadsheet, you are in for a rough time. Mostly, I wanted to discuss this because I think people should know about the company Bending Spoons. Bending Spoons is, of course, the natural enemy of Hard Fork—because if they can bend a spoon, what else can they bend?

But Bending Spoons is this Italian company. They actually went public at the start of July, but what they do is essentially acquire zombie tech brands. After a software company has outlived its usefulness, Bending Spoons comes in like a private equity company, and they try to figure out, “How can we squeeze the maximum amount of money out of the remaining customers?” Now, I’m sure they would phrase it differently, but I am bringing this up because if you use a product and you see a headline that it has been acquired by Bending Spoons, you in danger, girl. I’m just telling you: Keep alert to this possibility.

Kevin Roose

Yes. It is not a good sign when you get the inbound email from Bending Spoons’ business development folks that’s like, “We’ve been kicking the tires on some products that seem very exciting to us recently. Are you interested in selling your company?” Things are not going well when that happens to you.

Casey Newton

Indeed. Now, that brings us, Kevin, to our final story here on Running the Numbers, and that is dating math.

Kevin Roose

I love this story. This was from The Wall Street Journal. They had a great A-hed out. The A-heds are their famous front-page, quirky stories about culture and business. This one was an all-timer for me, and it was about the dating scene in South Korea.

Casey Newton

And are you a member of that scene?

Kevin Roose

I am not.

Casey Newton

Okay.

Kevin Roose

But it is suddenly being dominated by wealthy engineers at Samsung and SK Hynix, which is one of these AI chip infrastructure companies. The article refers to these suddenly eligible bachelors as “chip nerds” and talks about how the boom in AI has inflated the dating value of semiconductor-industry bachelors and bachelorettes. They’re as coveted as the memory chips AI companies need to build more data centers. So you may be asking: What has made these chip-company workers so attractive on the dating market? What has increased their value in the dating pool?

Casey Newton

Is it that they’re so smart and kind to the people that they go on dates with?

Kevin Roose

No, it’s that they’re making money.

Casey Newton

Oh, okay.

Kevin Roose

Some of these people are getting these very large, 6-figure bonuses. Others of them are just seen as upwardly mobile in an economy that has not had a lot of that.

Casey Newton

Wait, it’s more than a 6-figure bonus. These people are making $400,000 to $500,000 a year in bonuses.

Kevin Roose

Yes. So these companies, because they’re all growing so quickly, their employees are getting quite rich, and that sort of has trickled out into their dating lives. There are several great stories of people in here who will only date their coworkers. There’s a woman named Annie Kwon, who’s a 26-year-old chip engineer at Samsung, who has become suspicious of people wanting to date her for her money. So she has coupled up with a fellow Samsung semiconductor-division employee, and that gives her, as she puts it in the article, “double income.”

But if you are trying to keep up with a partner who is in this industry, you may be running into problems, like was the case for Roh Hee-jin, whose boyfriend at Samsung recently gifted her a Nintendo Switch 2 that runs around $450. Roh works as a software developer, but outside the chip industry, so she is not getting these huge bonuses. She had to save up to buy a mini PC for her boyfriend as a reciprocal gift.

Casey Newton

Wow. Man, the dating math here in Korea sounds really complicated, and it’s making me grateful that I don’t work at a chip company and confident that my fiancé is only into me for my body.

Kevin Roose

Now, Casey, I know you came by your relationship with an AI company employee honestly. You are not a pre-IPO stock-option chaser.

Casey Newton

Mm-hmm.

Kevin Roose

But I have heard from people in the AI industry that they are getting more attention on the dating market recently, and more people are swiping correctly—

Casey Newton

Swiping right.

Kevin Roose

—swiping right on them on the apps because they see that they work at one of these companies whose value has gone up.

Casey Newton

So you’re saying this dating math isn’t just a Korea thing. We’re seeing a version of this in San Francisco as well.

Kevin Roose

We are seeing it in San Francisco as well, yes. I have heard the term “Anthropic goggles.” Like, is that boy really cute, or are you just wearing Anthropic goggles? And, Casey, I’m curious if you, as a person who is engaged to an employee of Anthropic, have felt this in your own life. Do you feel competitive pressure in a new way from people trying to steal your man?

Casey Newton

You know, my message to people who would steal my man is: Go for it. If you think you can compete with this, I would like to see you try, honestly. Let’s see what you have.

Kevin Roose

Yeah, buy him a Nintendo Switch 2.

Casey Newton

Yeah, exactly.

Kevin Roose

See if that wins him over. And if I can get gossipy—

Casey Newton

Mm-hmm.

Kevin Roose

— for 1 second.

Casey Newton

Please.

Kevin Roose

I’ve heard stories of some early or senior AI-company executives who have traded—

Casey Newton

Traded up?

Kevin Roose

Well, I wouldn’t say—up is subjective—but they have traded, let’s just say, since becoming fabulously wealthy, and they are now dating models, OnlyFans people—

Things of that nature.

Casey Newton

Isn't it so amazing how we live in such an unpredictable time, and yet that feels entirely predictable to me?

Kevin Roose

Yeah.

Casey Newton

That's like: “Well, I've really enjoyed our last 15 years together, and we had such a beautiful relationship when we met in college, and of course I'll always love our children, but I'm gonna be on a private jet with my new Italian model spouse. Catch you later.”

Kevin Roose

Yes. So the math of dating and relationships in Silicon Valley and in South Korea is changing quite rapidly, and let's keep tabs on it.

Casey Newton

We'll keep running the numbers.

Kevin Roose

Don't steal Casey's man.

Casey Newton

And that was Running the Numbers. The numbers have been run, and the numbers are tired. The numbers are going to bed.

Zuckerberg’s Anti-Doom Fantasy + Finally an A.I. Detector That Works + A.I. Math | BidClub