[BidClub_]
Hard Fork · · 72 分钟

给互联网设年龄门槛 + Cloudflare 对阵 AI 爬虫 + HatGPT

Matthew Prince

播客
TL;DR
  • 年龄验证正成为互联网基础设施,但英国的落地显示,儿童安全政策很快就可能演变成隐私与言论监管体系。《网络安全法》要求服务评估未成年人是否可能接触色情、自杀鼓励或进食障碍相关内容,并在风险达到一定程度时部署“高质量年龄保障”。该制度覆盖 X、Reddit 社区,甚至可能波及 Wikipedia。成年人可能需要提交身份证件、信用卡或面部扫描,才能访问过去“基本完全私密”的内容。

  • 绕过手段可能削弱安全收益,却保留监控成本。超过400,000人签署废除请愿书,VPN成为显而易见的绕过方式,用户还用 Death Stranding 里表情丰富的游戏角色骗过面部识别。但这套模式仍在扩散:美国已有24个州通过年龄验证法案,最高法院也维持了得州针对色情内容占比超过三分之一网站的相关要求。

  • 真正关键的层面可能是保护隐私的年龄保障,而不是散落在全网的数千个身份数据库。Apple 提议的 API 会把家长提供的年龄转换为匿名令牌,再交给开发者;Meta 希望由 Apple 和 Google 承担更多法律责任,而 Apple 则希望提供基础设施、同时不承担这类责任。Kevin Roose 预计,年龄门槛互联网访问将“某种程度上不可避免”,因此决定性竞争在于实现架构。

  • Cloudflare 表示,AI 已严重侵蚀互联网以流量换内容的交易。Matthew Prince 的公司位于“超过20%的互联网”之前,他说,为内容带来流量的难度已接近10年前的10倍;对 OpenAI 来说是750倍,对 Anthropic 则是30,000倍。如果出版商无法通过名声、广告或订阅获得认可或收入,那么在750倍难度下已经“死了”,在30,000倍下则会被遗忘。

  • Cloudflare 7月1日默认拦截爬虫,是在制造 AI 内容市场所需的稀缺性。Prince 的核心判断非常明确:“AI 公司必须为内容付费”,因为它们如今复制原作,却不再返还可变现的流量。他预计,定价会从类似 iTunes 的“一首歌99美分”模式,转向类似 Spotify 的订阅或收入分成;能否接入独特内容,将成为模型差异化的关键。

  • Cloudflare 可能同时成为执行层和收费市场撮合者,在带来机会的同时增加集中度风险。出版商与 AI 的直接交易不会给 Cloudflare 带来收入;如果由 Cloudflare 谈判或促成交易,Prince 提出20%–30%的分成设想,而拦截和分析服务即使对小网站也保持免费。他承认一家覆盖20%互联网的公司同时扮演这些角色令人不安,但认为“任何市场的第一步都必须是制造稀缺性”。

  • HatGPT 的几则故事显示,AI 产品捕捉身份、言论和劳动的速度,已经超过社会规范的适应速度。LeBron James 的律师挑战未经授权的合成视频;Sam Altman 警告,ChatGPT 的治疗对话不受法律保密保护;Amazon 的收购目标 Bee 会记录用户及其周围所有人的声音;Meta 将允许部分编程候选人在面试中使用 AI。Casey Newton 希望建立的底线很简单:用户应当看见、控制并删除系统留存的关于他们的信息。

摘要 · 为研究而整理的核心内容

1. 年龄保障正在割裂曾经同质化的互联网

  • Casey Newton 将英国《网络安全法》描述为西方民主国家“监管在线言论最全面的尝试之一”。该法于2023年通过,要求服务评估未成年人是否可能接触色情、自杀鼓励、进食障碍或其他列明危害;如果评估发现相关风险,就必须部署“高质量年龄保障”,不能只依赖一个很容易核验的“我已满18岁”按钮。

  • 相关条款于25日星期五生效后,许多成年人毫无预警地发现,浏览内容可能需要驾驶证、信用卡或基于摄像头的年龄估算。Casey 担心的是,一个过去“基本完全私密”的访问,突然变成了与个人身份信息绑定的互动,而这些信息会如何存储、后续如何使用,都不清楚。

  • 监管范围远不止显而易见的成人网站。关于戒烟、苹果酒等看似与性无关主题的 X 和 Reddit 社区也面临限制;Wikipedia 警告称,隐私问题可能迫使它限制英国访问,而不是收集自己并不想要的信息。

  • Kevin 更大的判断是:过去大约40年里,一个13岁孩子和一个50岁成年人可以接触同一张信息网络;年龄保障可能按年龄把互联网拆分成实质不同的产品。真正的结构性变化不是色情网站被拦截,而是互联网出现了分裂。

2. 简单绕过手段削弱安全逻辑,法律却持续扩张

  • 公众反应“不是保持冷静、照常生活的局面”:超过400,000人签署了要求撤销政策的请愿书。VPN提供了最简单的逃生通道,用户可以伪装成身处英国境外,按另一司法辖区的规则访问服务。

  • 更尖锐的例子来自 Death Stranding。一些摄像头检测会要求用户微笑或皱眉,于是有人用游戏的拍照模式让主角做出这些表情,再把生成的图像交给验证系统。Casey 开玩笑说,年龄门槛引发的反弹可能会让它成为英国最畅销的游戏。

  • Ofcom 表示,这项法律“不是万能药”,但仍坚持年龄检查能阻止儿童随意撞见有害内容。整个交锋留下了核心错位:有决心的用户可以绕过门槛,普通成年人却仍要承担摩擦和隐私暴露。

  • 这股政策方向已经跨过大西洋。美国已有24个州通过年龄验证法,最高法院也维持了得州一项针对色情材料占比超过三分之一网站的规定。大法官 Clarence Thomas 的理由是,网站不像店员,“无法看着访客来估计他们的年龄”。

3. 设备级令牌可以保护隐私,但责任归属仍有争议

  • Kevin 为年龄门槛进行了最强辩护:酒吧会查验身份证,成人杂志过去也一直放在柜台后面。Casey 同意应当让未成年人远离成人内容,但反对让人反复披露敏感信息,甚至可能“余生都要”这样做。

  • Apple 提出的年龄保障 API 提供了节目主持人偏好的架构。家长在设备设置时输入孩子的年龄或生日;Apple 对其匿名化,再向应用发送一个令牌,例如表明用户是13岁。该功能当时尚未上线,但 Apple 告诉 Bloomberg,很快就会推出。

  • 这种设计能减少每个开发者掌握的信息,但没有解决谁来负责。Meta 推动州级法案,把验证义务更多转给 Apple 和 Google;Apple 则希望提供基础设施,却不愿在儿童触及受禁内容时承担法律责任。有人把 Apple 比作商场,把 Facebook 比作商场里的酒铺。

  • Kevin 反驳称,平台本来就会从行为数据推断用户年龄。YouTube 计划利用浏览和观看模式;此前报道也显示,Meta 知道 Instagram 上存在未满13岁的用户。澳大利亚转而考虑禁止16岁以下儿童使用 YouTube,说明“社交媒体”规则可以多快扩大。

4. Te 说明身份数据库如何放大泄露风险

  • Te 是一款让女性匿名分享与男性约会经历的应用,曾短暂登上 Apple App Store 榜首。据报道,注册需要扫描驾驶证,并通过自拍进行性别核验——这正是会让验证系统变成高价值攻击目标的集中式信息收集。

  • Kevin Roose 表示,据他了解,Te 在未能保护部分上传的验证材料后遭到黑客攻击,相关内容随后泄露。Matthew Prince 说,4chan 用户又拿到这些自拍,将其改造成“Hot or Not 风格的网站”和其他滥用项目。Casey 更广泛的教训是,不同服务商的安全能力参差不齐,因此敏感数据库越多,暴露面就会按比例扩大。

  • Kevin 的价值判断仍然相互拉扯。成年人应继续访问合法信息,但家长“已经被逼到绝境”,不知该如何让孩子接触智能手机和社交媒体;现有控制工具甚至可能要求家长变成“兼职 IT 人员”。不过,他仍认为更广泛的年龄验证不可避免,并希望在下一次“Te 式泄露”发生前先采用保护隐私的实现方式。

  • Casey 把这个问题放进他所谓的“持续6个月的言论权回收”中,与针对记者、学术界和广播媒体的压力并列。要求用户出示政府身份证件才能查看网站,可能并不会困扰所有人,但他认为这些限制“都是同一件事”,回头看时可能会显得重要得多。

5. AI 答案正在侵蚀搜索以流量换复制的交易

  • Matthew Prince 表示,Cloudflare 位于“超过20%的互联网”之前,因此能看到一项持续25年的隐性交易:出版商允许 Google 复制自己的内容,Google 则返还访客,出版商可以通过广告或订阅将访客变现。

  • Google 逐步改变了这笔交易。答案框开始直接展示事实,把用户留在 Google 内部,反转了创始人过去“让人迅速离开自己的网站”的自豪说法;如今 AI Overviews 会直接总结更多查询,用户无需访问底层来源。

  • Cloudflare 的对比非常刺眼:把流量吸引到一篇内容上的难度,已经接近10年前的10倍。Prince 认为,对 OpenAI 来说难度是750倍,对 Anthropic 是30,000倍,因为用户越来越多地消费衍生内容、跳过脚注,并开始信任不断进步的 AI 答案。

  • 在 Prince 的分类里,内容生产的动机包括自我满足或名声、广告、订阅,或几者的组合。既没有认可,也没有收入,生产就会停止:出版商在10倍难度下苦苦挣扎,在750倍下已经“死了”,在30,000倍下则会“被埋进地下、彻底遗忘”。他称这对互联网构成生存威胁。

6. Cloudflare 正在制造稀缺性,让内容拥有价格

  • Prince 没有声称知道最终机制,但他的起点判断非常明确:“AI 公司必须为内容付费。”搜索引擎用流量补偿复制;AI 系统复制内容,却几乎不返还流量,因此出版商已经没有理性理由继续免费开放访问。

  • 包括 Amazon 与 The New York Times 的协议、OpenAI 与出版商的安排在内,现有交易说明市场愿意付费,但双边合同留下了搭便车问题。“Sam 不可能在 OpenAI 当冤大头”:一家企业不能购买授权内容,同时看着竞争对手以零成本抓取同一批作品。

  • 因此,Cloudflare 在7月1日宣布“内容独立日”,将参与网站的 AI 爬虫拦截设为默认,除非创作者主动选择开放访问或接受补偿。Prince 的市场逻辑很简单:“没有一定程度的稀缺性,就不可能有市场。”

  • 定价还会不断迭代。他把初期固定费率比作 Apple 的 iTunes“一首歌99美分”模式,而后者最终让位于 Spotify 约“每月10美元”的无限量订阅。最终形态可能是固定费率、订阅或收入分成,而不一定真的是按爬取次数付费。

7. 强制执行的网络控制,让 robots.txt 请求产生后果

  • Prince 把 robots.txt 比作限速标志:它传达请求,却没有物理执行力。而且它通常过于粗糙,只能广泛适用规则,无法让出版商区分具体内容、爬虫和商业安排。

  • Cloudflare 追踪到 AI 公司通过 Bing 缓存、Internet Archive,甚至能够返回页面描述的广告网络,寻找同一批材料。有些公司据称使用住宅代理来隐藏身份;Prince 说,最糟糕的行为看起来不像普通企业,更像“朝鲜黑客”,同时多次强调 OpenAI 的表现相对规矩。

  • Cloudflare 能执行这些规则,是因为客户流量经过一套为网络安全而建的网络。用于识别黑客的同一套系统,也能拒绝伪装爬虫访问:行为恶劣的机器人可以被送进“暂停区”,而与 Cloudflare 达成协议的 AI 公司则能获得高效、结构化的内容传输。

  • 预期中的控制面将足够细致。创作者可以授权 OpenAI、拒绝另一家爬虫,或者最终直接发布标准“价目表”,即使规模不足以单独谈判也能如此。Cloudflare 会把身份识别、拦截、分析和授权传输整合起来,而不是继续依赖爬虫的礼仪。

8. 内容交易所可能惠及长尾出版商,也让 Cloudflare 获利

  • Prince 说,出版商的回应带着明显的“欣喜”,反复按下“禁止访问”。让他意外的是,AI 公司大体接受“内容是驱动我们引擎的燃料”,因此必须付费——前提是 Cloudflare 能提供一个“公平竞争的环境”,防止不那么守规矩的对手免费拿走内容。

  • Google 仍然是关键变量。Prince 把它不断变化的交易比作“温水煮青蛙”:一项有益的安排,在 Google 持续留住更多注意力后,慢慢变成了惩罚。他希望说服能够奏效,同时也指出,全球范围的调查最终可能迫使 Google 合作。

  • Casey 希望可信出版物能进入深度研究产品,或许通过订阅者凭证实现,但 Prince 不愿把市场限制在大型出版商之内。未知网站也能生产有价值的原创内容,初创公司需要负担得起的输入;健康的交易因此需要“很多卖家和很多买家”,而不是只有大型媒体与 AI 企业之间的合同。

  • 如果是出版商自行与 OpenAI 达成的直接协议,Cloudflare 不会抽成。如果由 Cloudflare 谈判或促成交易,Prince 设想收取类似 Spotify 的20%–30%分成,不过具体数字尚未确定。拦截和分析仍将免费;Casey 说,即使 Platformer 只能收到80美元,只要由其他人找到买家,也可能值得。

9. 更好的激励机制可能重振知识,也可能把知识集中到 AI 巨头内部

  • 当被问及特朗普总统说对每篇文章或每本书收费“不可行”时,Prince 区分了政府强制定价与私人市场。他的理解是,政府反对澳大利亚或加拿大式的强制要求,而 Cloudflare 的杠杆来自技术:创作者获得补偿,否则就失去访问。

  • AI Labyrinth 提供了惩罚性的一端。Cloudflare 不只是拦截行为不端的爬虫,还可以把它送进无穷无尽的 AI 生成页面,“大规模污染他们的数据”。Prince 对那些在 GPU 上投入100亿美元的公司说,至少应该拿出一部分资本,去资助系统正在消耗的内容。

  • 他有意夸张的诊断——“当今世界所有问题最终都归咎于 Google”——针对的是激励设计。Google 教会出版商把流量视为“神”,Facebook 和 TikTok 又强化了注意力经济,创作者于是学会触发皮质醇和愤怒,而不是生产经久耐用的知识。

  • Prince 希望 AI 的需求能够付费给那些填补世界“巨大奶酪块”空洞的创作者。他眼中的“黑暗镜像”是:记者和研究人员只能成为5家强大 AI 公司的雇员,知识则被分割到不同意识形态和国家阵营中。他承认,Cloudflare 因位于20%的互联网之前而拥有单方面权力,并提到公司此前在大规模枪击事件后撤销对8chan 的保护;但他认为出版商正在消亡,必须采取行动,而“任何市场的第一步”都是稀缺性。

10. HatGPT 展示 AI 如何撞上肖像、隐私、招聘与言论边界

  • 据报道,LeBron James 的律师向 InterlinkAI 发出停止侵权通知,原因是其社区制作了未经授权的视频,描绘他怀孕、无家可归以及其他侮辱性场景,其中一些“带有种族主义色彩,糟糕透顶”。Casey 认为,就他所知,这可能是首批已知的名人反对 AI 公司滥用其肖像的案例之一,也预示着围绕同意和人格权的争议将不断增加。

  • Sam Altman 警告,类似治疗咨询的 ChatGPT 对话不享有法律保密保护,可能在诉讼中被提交。Casey 欢迎这一披露,但希望建立有关留存、画像和广告的规则:用户应知道聊天机器人存储了什么,能够查看它“知道”的内容,并删除相关记录。

  • Meta 将允许部分编程候选人在面试中使用 AI,Kevin 称这是对传统 LeetCode 测试的“Roy Lee 胜利”。他们的实际结论是,评估应当接近 AI 辅助下的真实工作。另一方面,Substack 的推荐系统推送了一份纳粹通讯后,Casey 觉得自己得到了验证——这正是促使他离开的放大风险,而 Substack 随后将该功能下线。

  • Amazon 的收购目标 Bee 制作一款手环,把日常谈话转录成可搜索的记忆和任务。Joanna Stern 的评价是“实用得令人印象深刻”,却“真的他妈令人毛骨悚然”;主持人质疑,是否值得牺牲佩戴者和旁观者双方的隐私,只为提醒自己买牛奶。较轻松的插曲包括用于捕捉偷猎者的热成像机器鹿,以及一只八哥把图像转换为声音、再转换回图像,充当有损存储设备。

Speaker 1

People on social media, Gen Z people, have started referring to robots as clankers. Have you heard this?

Speaker 2

Oh, yeah, I believe I saw something about this.

Speaker 1

Yeah, so people are saying, “Oh, I hate when I call customer support and a clanker picks up,” and this is their new derogatory slang.

Speaker 2

Well, it pairs nicely with another new piece of slang that I wonder if you’ve heard, because people are talking about people who use tools like ChatGPT a lot, and they’re calling them sloppers.

Speaker 1

Really?

Speaker 2

Have you heard this?

Speaker 1

No.

Speaker 2

Yeah, so if you’re out on the internet and you’re publishing something, and it seems like it’s obviously just AI, somebody might call you a slopper.

Speaker 1

Hmm.

Speaker 2

So we have clankers and sloppers.

Speaker 1

Clankers and sloppers. I just think this makes me a little uncomfortable, even though clankers is not actually—I think it’s sort of a tongue-in-cheek thing. I just don’t think we should be calling the robots names. I don’t believe in slurring against anyone, human or not.

Speaker 2

You believe in an appeasement strategy with the robots.

Speaker 1

Yes.

Speaker 2

Just give them what they want.

Speaker 1

They are keeping score.

Speaker 2

I believe that.

This week, why age gates are suddenly popping up all over the internet, and how some could create more problems than they solve. Then, Cloudflare CEO Matthew Prince returns to the show to discuss his company’s new plan to help websites fight back against AI scrapers. And finally, we’re passing the hat for some HatGPT.

Kevin, how old are you?

Speaker 1

It’s none of your business. How old are you?

Speaker 2

That’s none of your business. But guess what? If we lived in the United Kingdom, it would be the government’s business, Kevin, because as many people found out over the last week, to use a lot of different websites, you now have to prove your age.

Speaker 1

Yes. This is the age-gating issue, which I’ve heard a lot of people talking about. I know you wrote about it in your newsletter this week. I’m excited to talk to you about it. Before we get into what is happening and some of the consequences and the reactions to this law, I wonder if you could just sell me on the stakes. Why talk about how the UK’s Online Safety Act is mandating age verification for websites? Why does it matter to me, an American?

1. Age Gates Reshape The Internet

Speaker 2

Sure. I would say a couple of things. Number 1, this truly is one of the most far-reaching attempts we have seen by a Western democracy to regulate speech online. We’ve talked on this show in the past about various ways that people want to protect children online, and requiring folks to verify their ages is one of the ways that has been discussed. But it’s extremely rare to see it roll out across an entire country the way it has in the UK.

So that’s thing 1. Thing 2 is this stuff is coming to the United States. In June, the Supreme Court upheld a law in Texas that requires residents of Texas to do something very similar if they want to access adult content online, and so it’s really not just a UK thing, Kevin. We are starting to see a gradual erosion of people’s freedom of expression online.

Speaker 1

Yeah, I think that’s well taken, and I might just even take a further step back and say I think the internet for the last, call it, 40 years has been a kind of informational free-for-all, right? You are going to experience the same internet whether you were 13 years old or 25 years old or 50 years old.

But what we are talking about now, and what is happening in the UK, is you have to establish your age in order to access a number of different kinds of services, including social media. One outcome here would just be that the internet actually fragments in this way, where if you are 12 or 13 years old, you are just going to have a very different experience of the internet than someone who is 25 or 50.

Speaker 2

Yeah, and I think in addition to all of that, Kevin, there is just the fact that the way that websites are collecting this information is putting people at risk in ways that we can talk about.

Speaker 1

Okay, so let’s just start with what is happening in the UK. What’s the backdrop here?

2. The UK Age Gate Rollout

Speaker 2

Yeah, so in 2023, the United Kingdom passed something called the Online Safety Act, and it includes provisions that require online services to try to spare minors from seeing what they would call harmful content online. Porn is a big part of that, but it also includes stuff about, like, if you have a pro-suicide website or a pro-eating-disorder website, you are now required in the United Kingdom to first do a sort of risk assessment of your website: Could children access this, and would they be exposed to this certain list of harms?

And if so, you then have to implement what they call high-quality age assurance, which means that, no, you can’t get away with just putting a box on your website that says, “Yes, I promise I’m 18.” All of this goes into effect last Friday, the 25th, and things start to go a little bit haywire.

Speaker 1

What happens?

Speaker 2

Well, first of all, I think a lot of people in the UK just didn’t realize that this was about to happen. And so they’re sitting down for their evening visit to Pornhub, and all of a sudden they find that they are being asked to upload their driver’s license to prove that they’re actually 18.

There are a number of different ways that people are allowed to prove their identity. You could also, for example, show your credit card, because you can’t have a credit card unless you’re at least 18. Or you can use your phone or laptop’s camera to take a picture of you, and they’ll do some sort of AI in the background to figure out if they think that you’re actually 18.

But for a lot of adults who didn’t see this coming, this produces some real anxiety because all of a sudden, an experience that had previously been basically completely private, just visiting a website, is now being linked to your personally identifying information, and who knows what’s happening to it once you actually submit it.

Speaker 1

Right. Now, is this just affecting porn sites, or are there other kinds of websites where people are being asked to prove that they’re of age?

Speaker 2

It is affecting many more kinds of sites. X is being affected. Many different subreddits are being affected. So, like, a stop-smoking subreddit, a cider subreddit, various things that do not immediately seem like adult content and could actually be quite beneficial to minors. All of a sudden, you need to prove who you are if you want to access them.

Also, Wikipedia has said that they may have to limit access to the site in the UK because of privacy concerns that are created by this law, which would require them to collect a lot of information that they don’t actually want to collect.

Speaker 1

Hmm. Now, is there porn on Wikipedia?

Speaker 2

I haven’t been able to find any. You?

Speaker 1

No.

Speaker 2

Okay. Well, I will say there are some Greek vases with some incredibly curvaceous men and women, and depending on what kind of night you’re having.

Speaker 1

Right.

Speaker 2

Yeah.

Speaker 1

So how are people in the UK reacting to this?

3. Britain Finds The Workarounds

Speaker 2

This is not a keep-calm-and-carry-on situation, Kevin, okay? More than 400,000 people have signed a petition saying that they want these changes to be reversed, and I will be curious to see where that goes.

In the meantime, though, people are just finding new ways to get around this, and I think, to the extent that this law actually sticks around, this will be the reason why: People have just found all sorts of relatively easy workarounds.

One thing you can do is use what they call a virtual private network, or VPN. All that does is lie to your internet service provider and say, “Hey, I’m not in the UK. I’m in the United States, where I can still watch most of what I want to watch.” And so we’re seeing a lot of that.

But we’re also seeing people get much more creative. Now, are you familiar with the video game Death Stranding?

Speaker 1

No, I’m not.

Speaker 2

Unfortunately, this is so funny if you know what Death Stranding is. It’s this kind of very strange and beautiful video game made by a Japanese auteur, and people are using photo mode in the game to take pictures of the protagonist in order to fool the UK’s age-verification technology. The reason this works is because, in the photo mode of this game, you can tell the character, “Smile,” or, “Frown.” And it turns out that a lot of the age gates in the UK are doing the same sort of thing—

Speaker 1

Huh.

Matthew Prince

—where, if you’re using the camera on your laptop, it’ll be like, “Okay, now smile. Okay, now frown.” People are now just doing this with the game Death Stranding. So if Death Stranding winds up being the best-selling video game in the UK this year, we’ll know why.

Speaker 1

That’s amazing. So, okay, has there been any response from the UK government or the safety authorities about people’s reactions to this? Are they saying, “We’ll keep improving the systems,” or something like that?

Speaker 2

There was a response from Ofcom, the UK’s media regulator. They have said that this act is not, quote, “a silver bullet,” but, quote, “Until now, kids could easily stumble across porn and other online content that’s harmful to them without even looking for it. Age checks will help prevent that.”

Speaker 1

All right, so that is what’s happening in the UK. Where are we in the US with these age-gating laws? I’ve read some headlines about people in some states who are having to go through age verification to get to an adult website, but where are these laws, and what is the regulatory picture here?

4. America Expands Age Verification

Speaker 2

Age-verification laws have been passed in 24 states. And last month, the Supreme Court upheld a law in Texas that requires websites where more than one-third of the content is sexual material to use one of these age-verification methods. Justice Clarence Thomas spoke for the majority when he said, “Unlike a store clerk, a website operator cannot look at its visitors and estimate their ages. Without a requirement to submit proof of age, even clearly underage minors would be able to access sexual content undetected.” So I suspect we’re going to start to see these laws in more and more places.

Speaker 1

All right, let me try to steelman the case for age verification here, because I think I am undecided about what I think about it. But I think you have decided that these laws are a bad idea, so I want to try to make the opposite case. We age-gate things all of the time—

Speaker 2

Mm-hmm.

Speaker 1

—to prevent minors from getting access to it or being exposed to it. So if I walk into a bar and the bouncer asks to see my ID, that is allowed. No one has contested whether that is a violation of the Constitution. If I walk into—I don’t know if this probably exists anymore, but when I was a kid, there were these things called nudie magazines. If you wanted to go into a newsstand and get them, they were behind the counter. They would have these special barriers on them, and you would have to prove that you were of age to be allowed to buy them. So why is what’s going on with age verification on adult websites on the internet any different?

Speaker 2

I share your concern here. I think that we should come up with ways to prevent minors from accessing this kind of adult material. I just object to the way in which they’re doing it, which requires that people share a lot of personal information for the rest of their lives.

Speaker 1

Okay, so you’re not opposed to age-gating in—

Speaker 2

No.

Speaker 1

—like, as a concept. You’re just opposed to the way that they’re doing it in the UK.

Matthew Prince

Yes. And this year, Apple suggested a way to do age verification that I think is a lot more elegant, Kevin. It’s not available quite yet. They said this week to Bloomberg that it’s coming soon. But basically, they’re going to offer what they call an age-assurance API.

If you’re a parent and you’re setting up your child’s device, you can just tell the device, “Hey, my kid is X years old,” or, “Here’s my kid’s birthday.” Apple can then essentially anonymize that and pass it through to a developer. So if you are Facebook, Instagram, or TikTok, you just get a little token that says, “This person is 13,” and you may want to show them different kinds of content, or you might want to restrict certain features. Essentially, it puts the onus on the parents and the device makers to do this. That way, if you’re just a normal adult using the internet, you don’t have to worry about uploading your driver’s license to visit a website. So I think that’s a very elegant solution.

Speaker 1

Yeah, I prefer on-device age verification rather than making every website operator go and do their own version of this and store all the driver’s-license photos and whatever else. I did not know that Apple was building that. So you’re saying that’s going to be released soon?

Speaker 2

Yes, they’re going to release that soon. Now, interestingly, Meta is trying to ensure that this does not become the way that this is handled, because that still puts the onus on Meta to do a lot of the age checking. They want Apple and Google, which also has a big app store, to have the legal liability in cases where a minor does access material that they’re not supposed to.

Meta has been leading a charge lobbying for a lot of bills around the country that would put the onus on Apple and Google to do all the verification instead of Meta having to play a role in it. And they’ve been having some success.

Speaker 1

Hmm. Now, this confuses me because, if I know one thing about Meta and other social media sites, it’s that they are very good at collecting information on users and using machine learning to sort of detect who is more interested in what and who is part of what consumer segment. I assume that these platforms already know, or have very good guesses about, how old all of their users are. So what is the issue with just having them do the detection?

I also saw that YouTube this week is looking at things like your browsing history, your consumption patterns, and the kinds of videos you’re searching for to make its best estimate of whether you are underage or not. So why shouldn’t the platforms have responsibility or a role here, too?

Speaker 2

I think the platforms do and should have a responsibility. What we have seen is a lot of reporting over the past couple of years that, at Meta in particular, they were writing reports about how many under-13 users they had on Instagram. Jeff Horwitz at The Wall Street Journal did a lot of this reporting. It’s both very disturbing and darkly comic to read.

Yes, they absolutely knew that they had all of these younger users, and so they’ve spent the last year trying to release a lot of features that essentially make it more difficult for under-13s to use the platform. So all the platforms are kind of belatedly coming around on this.

You mentioned the YouTube thing. Here’s something really interesting about YouTube: This week, the Australian government said that they were going to ban YouTube for kids under 16—

Speaker 1

Hmm.

Speaker 2

—because they consider it social media, which was a major reversal from what they were saying before. I sort of think this is the sort of thing that, if it was actually carried out, could cause the government of Australia to be toppled.

Speaker 1

A bunch of angry Minecraft preteens are going to storm the whatever the White House of Australia is.

Speaker 2

Yeah, yeah. But we’ll never know. If you know what the seat of government is in Australia, please email hardfork@nytimes.com. We were unable to find this information online due to age checks.

Anyway, look, I’ve been a little glib about this, maybe in the spirit of trying to record an entertaining segment about tech policy, but the truth is that this stuff is very complicated. I think everyone has a role to play, which is something that is very satisfying to say but does not actually solve the problem. Because ultimately, if you want to solve the problem, somebody has to be responsible. Somebody has to try to implement a solution.

Inevitably, when you implement a solution, people are going to be caught up in it. They’re going to be falsely tagged as underage when they’re overage, or vice versa. So it is really messy. What I’m trying to say is that there are more privacy-preserving ways of understanding a person’s age online, and I would like to see us focus on those rather than this sort of blunt-hammer approach that they’re taking in the UK, which in practice is mostly just going to annoy a lot of adults and potentially put their public information out there in a way that could be breached. That’s my way, Kevin, of introducing the conversation about Te.

Speaker 1

Oh, yes. Let’s spill the tea about Tea.

Matthew Prince

Let’s spill the tea about Tea.

Speaker 1

So this is a story—

Speaker 2

Kevin, what was Te? And, by the way, are there any red flags about you on Tea? I haven’t looked yet.

5. Te Exposes The Privacy Risk

Speaker 1

I have not been able to look, nor will I be looking, because of what happened over the past week. Te is an app that had a viral moment over the past week, in part because it briefly reached number 1 on the iOS App Store, ahead of ChatGPT, Instagram, and all these other apps.

To my understanding, although I’ve never been on it, Tea is an app that allows women to anonymously divulge experiences they’ve had with men.

Speaker 2

Yeah.

Speaker 1

So basically, dish the dirt about the guy you dated who was sort of a jerk to you or behaved in a way that you didn’t like.

And what got people's attention was that you could only register for this app if you were a woman. So they would do some kind of verification process when you signed up. I think you were asked to scan your driver's license—

Matthew Prince

Yes.

Speaker 1

And also submit a selfie so they could use AI to try to detect whether you are a woman or not and keep out all the men. This drove a lot of people on the internet insane. My understanding is that they essentially hacked Te. They found that Te had not secured some of these uploads that users had made, and they leaked a bunch of people's verification information and photos that they had submitted to Te as a kind of revenge by the men.

Speaker 2

Yeah. And this gets to my exact concern with some of these age-gating methods: it is just left up to the service provider to decide how you're going to verify people's ages. While most countries do have some laws that regulate data, in practice, we see breaches all the time, right? It feels like every week you see headlines about one app or another being breached, and in this case, you have material that is just ripe for someone to do stuff that is really abusive.

Indeed, a bunch of 4chan users got ahold of all of these Te selfies. They started creating Hot or Not-style websites, just doing all kinds of gross stuff with them. And that, again, I think is just a really predictable outcome of building these kinds of age-gating technologies.

Speaker 1

Yeah. I think that's a really good point and an example of why we should not be throwing this to the website operators and platforms, because inevitably some of them are going to use really secure, really well-designed age-verification services, but some of them are just going to do this in a very sloppy way and inevitably leak people's personal information.

Speaker 2

Yeah. And now every online service just has to have this database lying around. It just increases the risk across the entire internet for all of us.

Speaker 1

Yeah. So I think it makes a lot of sense to do it in a more centralized way through Apple and Google and their app stores. But I have also read that Apple has been lobbying against some of these state bills that would force them to do exactly that. You just told me that Apple is building an age-verification feature, but I've also read that Tim Cook at one point personally lobbied the governor of Texas not to do this kind of age-verification thing. So what is going on here?

Matthew Prince

So I think it's mostly just that they want to avoid the legal liability here. I read one story that said, “Look, Apple is a mall. Facebook is a liquor store in the mall. You need to hold the liquor store accountable for who they are selling liquor to.” This is the Apple metaphor. I think Apple does clearly want to play a role in the solution here. That's why they've crafted this whole API.

But they want it to be on the honor system, right? Where it's like, “Hey, we're going to do this. We can all agree this is a good thing. Now leave us alone.” And what the lawmakers are saying is, “Yeah, that's a good thing, but also we want to hold you legally liable if we find out that a bunch of 8-year-olds are watching eating disorder content on TikTok.”

Speaker 1

Yeah. I think that makes a lot of sense. And I'll also say, this is an issue, this age-gating thing, where I find that my values are in tension—

Speaker 2

Mm-hmm.

Speaker 1

Because I agree with you that the First Amendment is important, that adults should be able to access all manner of information on the internet, even if it is lewd, obscene, or inappropriate for children. I don't want a UK-style system where you have to upload your driver's license to some database that may or may not be secure just to go look at a subreddit.

At the same time, I totally understand that parents have legitimate concerns here. I talk to so many parents who are at their wit's end about how to safely introduce things like social media, smartphones, or the internet to their children without opening the floodgates. Parental controls are the thing that everyone talks about, but any parent can tell you these are not perfect. You practically have to become a part-time IT person to be able to understand and control what your underage kids are doing on the internet.

So I totally understand where the anger comes from at these companies that have made it very clear that they don't want to be doing any of this. They don't want this to be their responsibility. But I think that we are going to move to an internet with age verification in the United States and around the world. I just think that is going to happen. There's enough public pressure at this point that it is sort of inevitable in my mind, and so I would like to see it done in a way that preserves the privacy of those users, rather than having Te-style leaks every 6 months.

Speaker 2

I would like that, too. Here's the last thing I would say about this, Kevin. This may not be the free expression that you, the listener, care about the most. You may be happy to upload your ID if you want to look at adult content on the internet or even Wikipedia.

But what I am going to say is I view this very much as part and parcel with a widespread clawback of speech rights in the United States over the past 6 months. Look at the attacks on free speech that we've seen against journalists, against academia, against broadcast media, right? And now you are coming along and saying, “By the way, now you need to show government ID to look at a website.” These things are all of a piece, and I think it's important that we think about them as being of a piece, because it may be that 6 months, a year from now, we look back and we think, “Wow, we have really lost a lot of ground with free expression.” And the time to talk about that online is while you still can.

Speaker 1

I see. Well, thank you for explaining this to me. It was not immediately clear to me why I needed to care about a tech regulation in the United Kingdom, but I think I get it now.

Speaker 2

Yeah. I feel like just doing this segment has aged me about 10 years.

Speaker 1

Yeah. Also, I'm looking at a list of British expressions that mean “really angry,” and I have to read some of these to you.

Speaker 2

What are they?

Speaker 1

A bear with a sore head.

Throwing a wobbly.

Speaker 2

Throwing a wobbly?

Speaker 1

Yeah. Anyway, I'm throwing a wobbly over this regulation, and I think a lot of other Brits are in a right tizzy. When we come back, Matthew Prince from Cloudflare will explain how he's trying to keep the internet safe from AI crawlers.

Speaker 2

Not to mention the clankers and the sloppers.

Speaker 1

Well, Casey, we've got a very exciting guest with us today. Matthew Prince, the CEO of Cloudflare, is here to talk about something new that they've been working on.

Speaker 2

Yeah, and I am so excited about this, Kevin, because it actually is in the dead center of my own interests as somebody who has a website on the internet.

Speaker 1

And Casey, because we are talking about AI in this segment, we should talk about our disclosures.

Speaker 2

That's right, Kevin. My boyfriend works for the notorious AI-scraping company Anthropic.

Speaker 1

And my employer, The New York Times, is suing OpenAI and Microsoft over alleged copyright violations related to the training of large language models. So Cloudflare, for those who may remember our last podcast with Matthew, is the plumbers and bouncers of the internet, right? They make a lot of these security services, like DDoS protection, that many websites rely on to keep their services safe. They also do a host of other security and data-related things.

And Matthew is a diehard supporter of the internet. A couple of years ago, when all these AI bots started scraping the internet to feed data into their models, he and Cloudflare got really nervous about this. They have been doing a lot of interesting things to try to preserve what they see as the heart of the internet. Recently, they made a lot of waves by announcing that Cloudflare would start to introduce default blocking of these AI data scrapers, basically preventing AI companies from automatically being able to exploit and scrape data from the websites that they're visiting. This is a step that he said would help to protect content creators like you online.

Speaker 2

Yeah. I'm excited to talk to him about that, and also, Kevin, just about some of the research that Cloudflare has been doing into the state of the web as a whole. You and I have both been worried about what the rise of AI means for the internet in general. Cloudflare has taken a close look at that and has some interesting things to say.

Speaker 1

Yeah. It's one of the more underappreciated companies on the internet. Their decisions about even seemingly small things can have enormous ripple effects throughout the industry. I thought it was a really good time to bring on Matthew and ask him some questions about what they're doing and why.

Speaker 2

Well, let's bring him in.

Speaker 1

Matthew Prince, welcome back to Hard Fork.

Matthew Prince

Thanks for having me.

Speaker 1

So we're going to talk about AI and scraping and all the steps that you're taking to protect the internet. But before we describe the solution that you're proposing, which you announced on something called Content Independence Day, I want to ask you to describe the problem.

You had some pretty shocking numbers in your post about how much internet traffic patterns have already changed as a result of AI. Talk a little bit about those and why you felt the need to do something about it.

6. AI Traffic Starves The Web

Matthew Prince

Sure. At Cloudflare, we sit in front of north of 20% of the web, so we see a lot of what's happening online, and we understand a lot of how the web works and the business model that exists behind it. Twenty-five years ago, basically, Google struck an implicit deal with publishers: “You let us copy all of the content that you're creating, and in exchange, we'll send you traffic.”

The whole web—the economy of the entire web—was really built on that search-based interface. However, over the last 10 years, Google has changed what that interface looks like. It started subtly about 8 years ago, when they introduced the answer box. If you type in something like, “When did Cloudflare launch?” instead of you having to go to 10 blue links, it just says, “September 27, 2010.”

That was a pretty radical change because, back in the day, Larry and Sergey used to brag about how their job was to get people off Google.com as quickly as possible, right? They wanted—

Speaker 1

Ooh.

Matthew Prince

They even measured the time up in the corner to show you just how fast you were leaving their site, and now all of a sudden Google was keeping you on the site. More recently, they've introduced something called AI Overviews. If you do a search, most of the time now an overview shows up, which is an AI summary of what's going on, and that has made a big difference in how much traffic goes to content creators from search.

Speaker 1

How much of a difference? Put some numbers on that.

Matthew Prince

If you take 10 years ago as the standard and compare it with today, it's almost 10 times harder to get traffic today for a piece of content than it was 10 years ago.

Speaker 1

Ooh.

Matthew Prince

And it's gotten significantly worse just since the introduction of the AI answer box. That's just with Google, and that's the good news. If you look at the AI companies, OpenAI is 750 times harder. Anthropic is 30,000 times harder.

What's basically happening is that as the interface of the web is shifting from search to AI, people are consuming derivatives, not originals. They're not following the footnotes, and we've actually seen it get even worse as people are trusting the AI systems more because the AI systems are getting better.

What scares us about that is that there are really 3 reasons people create content. Maybe not you, Kevin, but one is ego—

Speaker 1

Vanity is one.

Matthew Prince

Vanity. Ego, right? There's a whole bunch of content that's just created for the ego and the fame of doing it.

Speaker 1

Yeah.

Matthew Prince

But the business model is you either sell subscriptions, or you sell ads, or both of those things. The value has to be either to make people famous or to make them rich through ads or subscriptions. That is what has built the web over the last 25 years.

The problem is that publishers are struggling already at 10 times harder. I worry that as more and more of the interface of the web looks like OpenAI or Anthropic, they're dead at 750 times, and they're buried in the ground and forgotten about at 30,000 times.

Speaker 1

Ooh.

Matthew Prince

And so if that's the case, if there's not an incentive for people to create content, I think people will stop creating content, and that really is an existential threat to the web.

Speaker 1

Okay, that's a pretty good and concise description of the problem. So what is the solution here? What are you all doing?

7. Cloudflare Blocks The Scrapers

Matthew Prince

I don't pretend to know exactly what the solution is, but I know some of the aspects of how we have to get there. The answer has to be that AI companies have to pay for content.

The deal is different than it used to be with search. Search copied your content, and in exchange, it sent you traffic that you could monetize. If now they're copying your content and they're not sending you anything, then why would you give them your content in the first place?

AI companies have to pay for content. I think what's encouraging is that we're actually seeing some AI companies doing that. Amazon just announced a deal with The New York Times, and OpenAI has done a number of content deals that are out there.

But the problem with a lot of those deals that we realized was that Sam at OpenAI can't be a sucker. He can't pay for content and then have all of his competitors get it for free. Another way to think of this is that you can't have a market unless you have some level of scarcity.

What we thought the first step was, and what we announced on July 1 in conjunction with the who's who of the world's publishers, was that we're going to, by default, block the ability for AI crawlers to get content unless they are compensating the content creators for getting that content. We think that that's super important.

Exactly how the compensation works, again, I think there are a lot of different models. I analogize it to—and at the risk of hubris—when Apple announced the introduction of iTunes, 99 cents a song, right? That wasn't what the final business model was. The final business model that we've come to is more like Spotify, where it's $10 a month and you get kind of all-you-can-eat from the Spotify catalog.

I think we're going to take some iterations to figure out where we land. We might start at some fixed price, and we might evolve to something that's more like a Spotify model over time. But the first step has to be actually saying, “If you're an AI company, you can't get content for free.”

Speaker 1

Now, if I'm a website operator and I want to prevent AI bots from scraping my site, I can already do that through robots.txt. I can just put a little file in the metadata, and I can say, “Hey, these 3 crawlers, you're allowed to scrape, but these other 6, you're not allowed to scrape.” So how is what you are building with Cloudflare different from that?

Matthew Prince

The first thing is that with robots.txt, the analogy would be like a speed-limit sign, where it says, “Okay, you can drive 55 miles an hour,” right? There's no law of physics that says you have to drive 55 miles an hour, and I think a lot of us maybe look at the speed-limit sign and go, “Ah, you know, probably 60 is okay,” right? It's the same thing. It is a recommendation. It's not actual enforcement. That's problem number 1.

Problem number 2 is that it's incredibly blunt in terms of what it does. You basically have to apply it across a significant portion of your site. You can't say, “Okay, let these things out, but don't let these things.”

The last problem is that a lot of the AI companies, because we see a huge amount of the internet and can track how this works, if they hit a robots.txt file, what they then do is find some other way to go out and scour the internet to find your same content.

They'll do a search against Bing and try to pull the cached content, or they'll go look at the Internet Archive, or they'll actually do things that are incredibly sneaky, like pinging ad server networks to get a description of the page that comes back to them and trying to find all kinds of things around it.

We've tracked some of the worst-performing AI companies, and there's a huge range. OpenAI, they're actually the good guys here.

They're doing things right. They're trying to do things the right way, and they have, by far, the best behavior. There are others that most closely resemble North Korean hackers, where they're literally using residential proxies to spoof who they are to try to get around the various blocks.

Speaker 2

Mm.

Matthew Prince

And so I think that in addition to a road sign for the badly behaving bots, we actually need something where we say, “We're gonna take away your car because you keep going 400 miles an hour down the road.”

Speaker 2

Tell us a bit about how technically the solution that you've built works. How are you able to build those kinds of fine-grain controls?

8. Cloudflare Enforces The Rules

Matthew Prince

Yeah. Cloudflare is essentially a giant network, and a ton of the web sits behind us, so it has to flow through our pipes. Our primary business for most of our history has been cybersecurity. Every day we go to war with the North Korean hackers, the Russian hackers, the Iranian hackers, and the Chinese hackers who are trying to get into our customers' systems one way or another. We're really good at identifying, no matter what they pretend to be, who they are, what they're doing, and stopping them, literally by just not letting their traffic get to our customers.

What we realized, though, was that actually gave us the perfect position to be able to help anyone who is a content creator also set enforceable rules of the road, where we can say that if this bot is behaving in a bad way, we're going to block it. For instance, as we have now studied this and we have very good evidence on what even some of the major AI companies are doing that's really sleazy, we're gonna put them in timeout. We're actually going to stop their ability to access a huge portion of the web, even if they pretend that they're doing it, and we've got the records to do that. Our security teams have actually investigated and figured that out.

On the other hand, if OpenAI has done a deal, we also want to make sure that they get that content as efficiently as possible and structured in a way that is possible. So the content creator can say, “I want to allow OpenAI to come to my page.” I think what we'll develop over time is a standard rack rate, where you as a content creator, even if you're small, can say, “Okay, I'm happy to let this content be scraped, but here's the price for it.” That's something that, you know, we'll see how that market develops. I think that's going to be the next step of this project.

Speaker 2

Now, you announced this new approach to AI crawlers a few weeks ago and flipped this switch that blocks all the crawlers by default. What has the reaction been from publishers, from AI companies, and from people who are inside Cloudflare looking at the data?

Matthew Prince

From publishers, not surprisingly, I'm on a lot more people's Christmas lists. Publishers were really struggling with how to deal with this. They saw the problem, and they didn't have a good technical solution to be able to lock it down. They're not cybersecurity companies, so they have a harder time being able to track this.

The glee that we've seen from publishers as they've turned the system on, seen what was going on, and then had the ability to push a button and say, “Disallow, disallow, disallow, disallow,” has been really palpable. I think that's why this resonated across such a wide swath of the publisher base.

What surprised me has been the reaction of the AI companies. I kind of thought that they were just gonna throw up all over this and hate it, but it hasn't been that. For the most part, with a couple of exceptions, they've said, “Listen, we get it. Ultimately, content is the fuel that runs our engine, and we need to pay for it, but the key is it needs to be a level playing field.”

What I am encouraged by is that in all the conversations that I've been a part of, the AI companies said, “If you can make it a level playing field, if you can make it something that's fair, then we're willing to pay for content.” I think that's a dramatic step that says we're headed in the right direction. Now, making it fair is gonna be a trick, and folks like Google, who can just believe that they've got a God-given right to be able to copy everything off the web, are gonna take some persuading to get there.

But I think as the industry lines up and says, “This is the right thing to do,” we'll be able to get Google to hopefully voluntarily support it. If not, there's certainly enough investigations going on around the world that, one way or another, I think they will be persuaded or compelled to get behind this effort.

Speaker 2

I mean, basically what you're describing is a revenue-share deal, right?

Matthew Prince

That's right.

Speaker 2

And it seems only logical that a company like Google, which is gonna be making a lot more revenue by crawling everyone's content because they're such a source of queries, should be paying more than the upstart AI company that just shows up on day 1 and—

Matthew Prince

That's right.

Speaker 2

I would hope that they would be open to that.

Matthew Prince

I think I am encouraged, and I do believe that Google really does believe in the ecosystem, and they get it. But it's a little bit of the frog boiling in water, where 10 years ago or 25 years ago, when Google started, it was a good deal. Get in the nice pot of water. You're happy. You're a frog, right?

But they've slowly turned the heat up, and that's made what used to be a good deal into a much, much worse deal. So the deal needs to get renegotiated, and that's a big piece of it.

The other thing that I think is important is that you need to have access to content, and that's the thing today that's cheap. But my prediction is that over time, the real differentiator between the different AI companies is who has access to the most interesting content. We've seen this play out with things like Netflix, where you can see that by getting original content, it actually drives subscribers.

I think that at the end of the day, it's exactly right that fundamentally what content producers should be arguing for is, “We should be getting a share of whatever the revenue that you're generating from users is,” because we're the fuel that runs the engine that's powering your business.

Speaker 2

I feel like you can already see this today, and Kevin, I'm sure you've seen this. Matthew, I'd be curious if you have, too. If you ever use one of these chatbots to run a deep-research report because you're trying to really bone up on some set of historical facts or current events that you want to write about, you get these results and half of them are from bleepbloop.com, like newswire.xyz—publishers that no one has ever heard of and that you're not entirely sure are on the level.

Often when I run those searches, I think, “God, I wish that the sources I trust actually did have deals with these chatbots,” or that I could log in with my Bloomberg credentials and then just have you read Bloomberg, too, as part of the report that you're making.

Matthew Prince

Yeah. And I think all of those things will come, but 2 things. One, bleepbloop.com actually sometimes is gonna have really interesting original content. Part of the key here is we don't wanna create a situation where the only people who get the content deals are the major, major publishers.

Speaker 2

Sure.

Matthew Prince

The New York Times—

Speaker 2

Yeah.

Matthew Prince

—can do a big deal with OpenAI. At the same time, you wanna make sure that the small AI companies, the new startups, are able to also get access to this. So a vibrant market here is lots of sellers and lots of buyers, and you wanna be able to do that.

Speaker 1

Right. You mentioned that publishers are very happy about the steps that Cloudflare has taken to block these AI crawlers, and that AI companies were not as excited, understandably, but maybe they would get on board, too.

I did talk to one AI executive about this, who basically accused you and Cloudflare of setting up a new tollbooth between the AI companies and the content companies because you are not only providing this technology, but you are also inserting yourselves as the merchant of record in these transactions. This person asked me to ask you what percentage of each payment Cloudflare plans to take. What percentage does Cloudflare plan to take from these transactions?

Matthew Prince

First of all, if you're The New York Times and you do a deal with OpenAI, that's your deal. We don't get any of that. If we facilitate that—if we're the ones who go out and negotiate it—then, yeah, we'll take some percentage of it. I have no idea what that will be, but I think it'll be something reminiscent of what Spotify takes of a subscriber's revenue versus what they pay out. Usually, that's sort of in the 20% to 30% range of what that is.

I think the only way that we should get something is if we're actually generating value from this. If not, then you should do the deals yourself between the AI companies and the publishers. In that case, sure, we'll provide you with the interface to be able to stop it, but that's not something that we would take any percentage of.

What's also important is that we provide this at no cost to even our free users because we think that this is fundamentally important to the long-term health of the internet. It shouldn't be something that only the big companies can get access to. No matter who you are, if you're signing up for Cloudflare, you get these tools, the analytics, the understanding, and the ability to block it. Again, I think we'll be less involved in the transactions for folks like The New York Times, but for bipbopbloop.com or whatever—

Speaker 2

Mm-hmm.

Matthew Prince

That's probably a porn site that we're pointing people to. You can sort of envision what it might be. But if that's the case, then I think that's a place where, if we're doing work, we should get compensated for that work.

Speaker 2

I have a website on the internet, Platformer.news. That's where my newsletter is, and I'm super interested in this because I frankly just do not have the time or energy to go out and try to strike deals with AI companies. I also have no idea what the value of Platformer is in that particular marketplace.

Matthew Prince

Yep.

Speaker 2

So if somebody wants to go make a market, and then I can just show up and you tell me, “Hey, we can make you 80 bucks this year,” or whatever it is, I'm interested. And if you want to take 20% of that, sure.

Matthew Prince

Yeah, and I think that feels fair.

Speaker 2

Yeah.

Matthew Prince

Hopefully, we're not the only ones doing this. From my perspective personally, my wife and I own a small newspaper in our hometown. We see how hard and how important it is to have local news, and there's got to be a business model for it. The reporters need to eat. It costs money to print papers. You have to have a business model there.

Speaker 2

Mm-hmm.

Matthew Prince

Personally, I think we've built a $60 billion company on the back of the internet. I feel an enormous responsibility to give back to the internet and actually protect it. Our mission is to help build a better internet, and if you talk to anyone at Cloudflare, that's why they work for us. With all due respect to the AI executive, they've built an entire company by stealing content creators' content and not compensating them for that. If they keep doing that, people will stop creating content, which not only kills the internet but kills them in the process.

Speaker 1

Yeah. I think there are a lot of people who would agree with you. One of them is not President Trump. He said just last week, in connection with some new AI policy rollouts that he was doing, that he basically took the side of the AI companies and said that they shouldn't have to pay for copyrighted material. He said, quote, “You can't be expected to have a successful AI program when every single article, book, or anything else that you've read or studied, you're supposed to pay for. It's not doable.” What is the incentive for an AI company to agree to something like a pay-per-crawl system if the current administration is signaling that they're not going to face any penalty for just breaking copyright the old way?

Matthew Prince

Yeah. Like many things that sometimes come out of the Trump administration, nuance is a little bit lost here. We've actually talked with a number of people in the Trump administration, and I think what they are concerned about is much more a government-instituted kind of “You must pay X,” in the way that we've seen come out of Australia and Canada. I think that's what they're trying to signal that they are not in favor of. They are very much in favor of, as far as we can tell, private-market solutions where that gets created.

The incentive for the AI company is that you need content in order to build your tools, and we've just stopped you from getting it. So regardless of what the law says, we're going to be able to technically stop the AI companies from being able to get the content unless they're compensating the content creators for it.

Speaker 1

Hmm.

Speaker 2

Hmm.

Speaker 1

I want to ask you, Matthew, about another release that you all did earlier this year related to this issue of AI crawling, which was something called the AI Labyrinth.

Matthew Prince

Yeah.

Speaker 1

This was a system designed to basically, instead of blocking AI crawlers, just redirect them to an endless series of AI-generated links and pages, basically trapping them in this labyrinth that they could not escape from.

Matthew Prince

Yes.

Speaker 1

So my first question is, do you worry that the robots will remember that you did this to them and take their revenge?

Matthew Prince

That is not a risk factor that we have currently included in our S-1, but I'll talk to our legal team about it. We now have data that there are some AI companies—big AI companies, well-funded—that are just behaving horribly. And frankly, if they're going to behave like hackers, then we're going to behave like trolls right back to them. We can feed enough garbage into their system that they will create garbage content.

When we come back to what's the incentive, you don't want to piss us off because, again, we believe in the future of the internet, we believe in supporting journalists, and we believe in supporting content creators. We're on the right side of history here. The right thing to do is say, “Yeah, you're spending $10 billion on GPUs, you're spending billions of dollars on employees, and you should be dedicating at least something to paying for content.”

Speaker 1

Did the labyrinth work? What were the results of that?

Matthew Prince

Yeah. For real, again, we try not to put OpenAI into the labyrinth. OpenAI is a good actor, doing all the things right. So we don't throw them into the labyrinth.

Speaker 1

This message was not endorsed by The New York Times legal team.

Matthew Prince

That's right. But for bad actors—and there are bad actors out there—we can pollute their data, and we can pollute it at scale.

Speaker 2

I'm having what I think is an entirely novel feeling during a Hard Fork recording: I'm feeling optimistic about the state of the media. Every single time we talk about the internet and the media on this show, I'm like, “Well, it's going to be really tough, everybody. Batten down the hatches. The storm is here.” But Matthew, you actually are giving me optimism that when someone with the right incentives—and I do feel like you have the right incentives—shows up with great technology and the willingness to go out and make a market, maybe by hook and by crook, we actually still will have a media industry 5 years from now.

Matthew Prince

Well, I think it—again, I might lose people because this is starting to get a little woo-woo—but everything that's wrong with the world today is ultimately Google's fault.

Now, that's a little strong. But Google taught us—they were the first ones to really teach us—that traffic was the deity that we all had to worship.

Speaker 2

Mm-hmm.

Matthew Prince

Again, I think Google's actually been a massive net positive for the world, but Google begat Facebook, which begat TikTok, which is just kind of spiraling down toward the attention economy: How do I create a cortisol response and get people to click on my stuff so that we can sell ads to them? I think that's the wrong direction for really furthering humanity.

If, on the other hand, as we create this, the giant block of cheese has holes in it, we can say that if you as a creator go out there and create something that fills in one of those holes, we'll compensate you for that. We'll pay you more for it. What I'm hopeful about is that we get a lot more really interesting, long-form, knowledge-generating content, which is what we all want.

Speaker 2

Mm.

Matthew Prince

The only reason we're not getting it is because all of the incentives are not “How do we do great things?” but “How do we actually just rage-bait people into clicking on links?” And so I think it's super important that, as we go through this, it shouldn't just be us figuring this out.

I've tried to spend as much time with publishers and AI companies. We've been working with some of the leading academic economists and others to say, “As we think about this market, how do we make sure it's as healthy as possible?” Because the dark-mirror version of this is not that journalists go away, and not that researchers go away. It's that they're all employed by the 5 big AI companies, right?

Speaker 2

Mm.

Matthew Prince

And so it's not that we're going to go back to kind of the media times of the early ’90s. It's not that we're going to go back to the media times of, like, the Medicis, where you have just these 5 powerful families that control all of academia and research and journalism.

And that's not actually that hard to envision. There will be a conservative one, sort of the Fox News version. There will be the liberal one. There will be the Chinese one. There will be the Europeans' attempt at one, and those will be the things that are out there, and knowledge gets consigned behind that.

I don't want that to happen, and I think the key to making that not happen is figuring out how we can have a healthy market with lots of sellers of content, lots of buyers of content, and make that as robust as possible.

Speaker 1

Yeah. Matthew, I know in the past you have expressed some concerns about the power that you and Cloudflare have by virtue of sitting in front of 20% of the internet, as you say. When we've talked before, some of it has been in the context of the decisions you've made around content moderation, basically deciding whether or not to protect websites with extreme or violent content on them. You famously pulled protection from 8chan, which led that site to go down after some mass shootings.

I think a lot of people understood that, but you were concerned at the time that this sort of unilateral power that you had to wake up and make a big change to the structure of the internet wasn't something that maybe anyone should have. So I'm curious: when you were deciding to implement this change, to push the button, to block all the AI crawlers by default, did you worry about that exercise of power and how unilateral it was?

Matthew Prince

Totally. And I think we take that responsibility really seriously. But what we realized was all the publishers were sitting around saying, “We're dying, we're dying, we're dying,” and no one was doing anything about it. And so, if you're in that situation and you see something that really matters—the internet really matters—and we should be fighting for it, we should be protecting it, and it's dying because the business model behind it is dying.

At no step have we said we're the only solution. In fact, we've tried to work with as many different competitors as possible to say, “This is important. Let's do it.” And we don't pretend that we have the answer or that it won't evolve over time, but we do know that the first step in any market has to be creating scarcity, and so that's what we did on July 1.

Speaker 1

Mm. Got it. Well, Matthew, thank you so much for stopping by. Really interesting experiment. I'm excited to follow it and see how it plays out.

Matthew Prince

Thank you guys for having me.

Speaker 2

Thanks, Matthew.

9. Hat GPT Takes Over

Speaker 1

All right, Casey, it's time to pass the hat. We are playing another round of our favorite game, Hat GPT. Hat GPT, of course, is our game where we pick tech headlines out of a hat, riff on them, and eventually yell at each other to stop generating.

Speaker 2

Which they don't even say in ChatGPT anymore, sadly. It's already become a throwback expression.

Speaker 1

It will not have slips of paper inside it, though.

Speaker 2

Or will it?

Speaker 1

There's only one way to find out.

Speaker 2

Only one way.

Speaker 1

God, we are merging into the same person. It's terrible. Okay, so I'm going to put the slips in.

Speaker 2

Okay.

Speaker 1

Jiggle them around a little bit there, mix them up, and then why don't you pick the first one?

Speaker 2

All right. LeBron James' lawyer sends a cease-and-desist to AI company making pregnant videos of him. This is from our friend Jason Keebler at 404 Media. He writes, “The creators of an AI tool and Discord community called InterlinkAI that allowed people to create AI videos of NBA stars says that it got a cease-and-desist letter from lawyers representing LeBron James. AI-generated videos of James,” Kevin, “included scenes like James as a homeless person, James on his knees with his tongue out, and James lying on a couch clutching a pregnant belly.”

So I guess my first question here is, who is LeBron James? Just kidding. He's a very famous basketball player. Kevin, what do you make of this?

Speaker 1

So this is obviously going to be a thing for celebrities. They do not like their names and likenesses being used without their permission, especially if what you're doing with LeBron James' name and likeness is turning him into a pregnant person. Did you watch any of these videos?

Speaker 2

I did see a couple of them.

Speaker 1

They're very disturbing.

Speaker 2

Some of them, I think, are at the end of the spectrum that is just surreal and funny. And then some of it is also just racist and horrible. Regardless of all of that, it's clear that LeBron James did not give permission for this. And it's an interesting story because, as far as we know, this is one of the first known times that a celebrity has objected to the misuse of their likeness by an AI company.

Speaker 1

Yeah, and I am worried that it seems like they have automated the jobs of the mpreg community.

Speaker 2

Mm-hmm.

Speaker 1

Mpreg is, of course, a niche fan-fiction thing—

Speaker 2

Yeah, what is that? I've been meaning to ask you about that.

Speaker 1

For years, people have been creating these animated fan-fiction cartoons of Sonic the Hedgehog becoming pregnant. Why are they doing this? I don't know. Couldn't tell you. Not an mpregger myself. But there were a lot of hardworking mpreg artists out there who now have been put out of jobs by these AI tools. So for that reason and that reason alone, I think we should take a hard stand.

Speaker 2

I now want to say retroactively to my parents: do not listen to this segment.

Speaker 1

And definitely don't Google mpreg.

Speaker 2

All right. Stop generating.

Speaker 1

Okay. Next up. Oh, this is a fun one. Sam Altman warns there's no legal confidentiality when using ChatGPT as a therapist. This one comes from TechCrunch. Sam Altman, the CEO of OpenAI, went on the popular podcast by Theo Vaughn last week, where he acknowledged that OpenAI might be legally required to produce a user's conversations with ChatGPT in the case of a lawsuit.

I saw this clip going around, and I thought, Sam Altman continues his terror campaign against the Hard Fork podcast. As you will remember, at our live show, he burst onto the stage and then peppered us with questions about this lawsuit between OpenAI and The New York Times, which has resulted in OpenAI having to retain conversations between ChatGPT and its users. He's very upset about this, doesn't want this kind of document retention to be required, and has started advocating for the privacy guarantees that AI companies should be allowed to make with their users.

Speaker 2

Look, this one is important because a lot of people are already using these chatbots as therapists. They're having therapy-style conversations, and I think most people are not thinking a lot about what is happening to that data. Some companies may want to erase that data for user protection reasons. Other companies might want to keep it forever and create a detailed profile about you, and then rent some information to advertisers.

So I would love to see some kind of legal intervention or regulatory intervention come down and say, “You're allowed to use information submitted to chatbots in these ways. Users should have a full view into what a chatbot knows about them and what kind of information is being stored about them. They should be able to delete it.” Right? We just need a lot of data and privacy stuff around this sort of thing.

So I have to say, I'm grateful to Sam Altman for at least saying, “Hey, by the way, you don't have legal protections here.”

Speaker 1

Yes.

We are allowed to use data however we want in training our models, but God forbid anyone else wants to use data about our users after the fact.

Speaker 2

All right. Stop generating.

Speaker 1

All right. Your turn.

Speaker 2

How to catch a wily poacher in a sting: a thermal robotic deer. This is from James Finelli at The Wall Street Journal. Wildlife enforcement officers turned to a Wisconsin taxidermist to make remote-controlled robots that look like wild animals to catch poachers, Kevin.

Speaker 1

Hmm.

Speaker 2

To make his decoys, a man named Brian Wolslegle applies the skin of a dead animal to a mold made out of polyurethane. He affixes glass eyes and plastic ears. The circuitry to make the decoys move comes from parts for remote-controlled cars. With some AA batteries, officers can remotely operate the Bambi-bots.

Kevin, would you be interested in one of these to ward off some of the poachers on your property?

Speaker 1

Yeah. I think there's been a lot of poaching attempts made against my livestock, and I won't be standing for it. So I'll be buying one of these.

Speaker 2

Here's where this man's real opportunity is. He needs to use the same technology to build decoy AI researchers and put them in OpenAI so that when Mark Zuckerberg comes onto the property, he mistakenly poaches one of the robots instead of one of the human researchers.

Speaker 1

Yes.

Speaker 2

So if I'm Sam Altman, this is what I'm doing.

Speaker 1

I think it's a great strategy.

Speaker 2

Yeah, yeah.

Speaker 1

In general, I think robot animals—I'm not a big fan of them. On my block, there's a guy with a robot dog. He runs a STEM program for kids, and I've seen him walking his robot dog out on the street. It's very disconcerting. Let's just say—

Speaker 2

Top 10 signs you live in the Bay Area, by the way.

Speaker 1

Yes. Let's just say these things are not entirely lifelike yet, and they do still give you the willies.

Speaker 2

All right. Stop generating.

Speaker 1

Next up: Did a guy just save a picture of a bird to a bird's brain? This one comes to us from Sean Hollister at The Verge. It's about a YouTuber named Ben Jordan, who has a popular channel about music and acoustic science and was able to get a bird—a starling—to reproduce a spectrogram image in sound.

Casey, did you hear about this story?

Speaker 2

This is actually my favorite story of the week.

Speaker 1

What happened?

Speaker 2

So, if I get this straight, the YouTuber converts an image into a sound. He then plays the sound for a bird. The bird then repeats the sound. In that way, you can say that the YouTuber was able to store an image in the bird's brain. He used the bird as a data transfer device, like it was a disk drive.

Speaker 1

Yeah, you almost got it right. What happened is he takes this drawing of a bird, converts it into a spectrogram, and then plays that sound for the bird. The bird mimics it back, and then he takes that recording and turns it back into a spectrogram. So he's essentially going image to sound, transferring the sound to the bird, having the bird transfer it back, and converting it back into an image.

Now, Casey, why would you do something like this?

Speaker 2

Because you have a YouTube channel and you're trying to get a lot of subscribers. How close was the second image to the first image?

Speaker 1

Apparently, it was quite close. The bird is a little bit of a lossy compression device, but you were actually able to recognize it as a line drawing of a bird after the fact, which is kind of amazing.

Speaker 2

Honestly, hats off to this person. What a bizarre idea, but fantastic.

Speaker 1

You know, I uploaded an episode of our podcast to a bird.

Speaker 2

Oh, yeah?

Speaker 1

Yeah. It's called “Hard Stork.”

Speaker 2

Okay, sure. Why not? Stop generating.

Speaker 1

All right, you're up.

Speaker 2

Okay. Meta is going to let job candidates use AI during coding tests. This is from Jason Keeler at 404 Media again. Meta told employees that it is going to allow some coding job candidates to use an AI assistant during the interview process, according to internal Meta communications.

On an internal message board for the company, there was a post that called for mock candidates because apparently they're going to let employees do one of these AI-assisted interviews so they can work out all of the kinks. It looks like Meta's going all in on using AI coding agents to write code and also is just not going to try to stop people from cheating on their job entrance exams anymore.

Speaker 1

This is a total Roy Lee victory.

Speaker 2

Absolutely.

Speaker 1

This is a vindication of what Roy told us when he came on the show several months ago, which is that these LeetCode interviews are totally cooked because now people can just use tools like the ones Roy Lee is developing to cheat on their interviews. I guess Meta is seeing the writing on the wall and saying, “You know what? Go ahead. Use your AI. We'll design our new test,” which I think is probably a good outcome. What do you think?

Speaker 2

Yeah, I think I'm interested to see how this affects the candidates that they attract and the quality of the engineers that they can recruit. I am ultimately persuaded that people are going to be using these tools in the workplace anyway, so why not use them when you're doing the actual test to get the job?

Speaker 1

Now, do you think they make the people who are making a billion dollars a year take the tests when they come in?

Speaker 2

Yeah, they just give them the really hard version. They say, “If you do the job, we'll give you a billion dollars.”

Speaker 1

God, what a weird thing.

Speaker 2

Yeah.

Speaker 1

Can you imagine being onboarded to a new job, and you go through the training and it's like, “You know, here's your benefits package,” and you're just sitting there thinking—

Speaker 2

Yeah.

Speaker 1

“I'm making a million— a billion dollars?”

Speaker 2

Yeah. They're like, “Do you want to put anything in your health savings account this year? What about the commuter benefit? Do you want the $40 for BART this month?”

Speaker 1

You're like, “I'm making a billion dollars, people. I'm not sitting through the IT training.” Okay.

Speaker 2

Okay.

Speaker 1

Stop generating.

Speaker 2

Stop generating. Here's one, Kevin. This is from Taylor Lorenz at UserMag. Substack sent a push alert encouraging users to subscribe to a Nazi newsletter that claimed Jewish people are a sickness and that we must eradicate minorities to build a white homeland.

Speaker 1

Oh, boy.

Speaker 2

Yeah. I have to say, Kevin, rarely have I felt so smug in my entire life as I did when I read this story. Longtime listeners of the Hard Fork show may know that I moved Platformer off of Substack last year after some other folks had found a bunch of pro-Nazi websites on the network, and Substack would not commit to doing any proactive searching for these blogs to get rid of them.

The main reason that I wanted to leave was that I thought, “These people are building amplification features, and inevitably they're going to just start promoting these things. It might be unwitting, and it might be intentional. But either way, I don't want any part of it.” So now, sure enough, a bunch of people who had the Substack app installed yesterday just got a ringing endorsement for a Nazi blog.

Speaker 1

Oh, boy.

Speaker 2

Yeah. Substack did basically say that this was a huge mistake, and they took the offending recommendation system offline. They're going to rejigger it so it doesn't happen again. But, well, it did happen.

Speaker 1

Yeah, and they did not see it coming.

Speaker 2

Yeah. That's a good way of putting it.

Speaker 1

Yeah.

Speaker 2

All right. Stop generating.

Speaker 1

All right, last one. Why Amazon wants an AI bracelet that records everything you say. This comes from Nicole Nguyen at The Wall Street Journal. Amazon is acquiring a company called Bee. Bee makes a wearable device—B-E-E—that transcribes all the conversations in your day, including when you talk to yourself.

It then uses AI to turn that giant word soup into a searchable history, offering up key events and even to-do lists based on your chatter. Friend of the pod and Wall Street Journal reporter Joanna Stern reported on her own experience testing out the Bee bracelet earlier this year. She described it as impressively useful and also, quote, “Really fucking creepy.”

Casey, what do you make of this bracelet, and will you be buying one?

Speaker 2

I'm not going to buy one for myself. Generally, I don't want a detailed record of everything that I say during the day. One of the reasons why we podcast is so that most of what I say can be edited out, you know? So the idea that I would just have this unfiltered record stored in an AWS bucket doesn't really appeal to me.

Also, as I read the reviews of these devices from Joanna and others, nobody really seemed like they were getting a lot of value out of it. It was like, “Oh, yeah, I told my husband I should get milk, and now I get an email that's like, ‘Hey, remember, you want milk.’” Is that really worth giving up all of your privacy in perpetuity?

Speaker 1

Well, and all the privacy of everyone that you talk to—

Speaker 2

Yes, exactly.

Speaker 1

Throughout the day—that is the worst part of this. I have at times been recording interviews and accidentally left the voice memo running for an hour afterward, and it's never that it's that interesting, but sometimes it does catch other people's conversations in there, and then I feel bad about it and delete it. But with the Bee bracelet, this is the whole point. It's just logging you all the time. I don't think people are going to be that excited about that.

Speaker 2

Well, here's what I'm looking for: keeping the Bee on you at all times as a condition of continuing to work at The Washington Post, owned by Amazon founder Jeff Bezos. Because they're going through a lot of turmoil right now, and there's a lot of people who are leaking stuff to the media. So I think it's going to be like, “Hey, we need to check your Bee.” “What have you been saying about us?”

Speaker 1

So it sounds like you're not going to be on the early beta tester list for the Amazon bracelet. Maybe I will.

Speaker 2

I'll be minding my own beeswax.

Speaker 1

Stop generating.

Speaker 2

Okay.

Speaker 1

And that's ChatGPT.

Speaker 2

Yay.

Speaker 1

Thanks for playing.

Speaker 2

We won again. It is a competition.

Speaker 1

What's that they used to say on Whose Line Is It Anyway? Where everything's made up and the points don't matter?

Speaker 2

That's right. ChatGPT, very similar.

Speaker 1

You know, my three-year-old is very obsessed with winning and losing.

Speaker 2

Oh, yeah?

Speaker 1

Yeah, every time we do anything, he says, “I won. You lost.” So I'm going to start doing that with you.

给互联网设年龄门槛 + Cloudflare 对阵 AI 爬虫 + HatGPT — 文字稿与摘要 | BidClub