太空数据中心 + 右翼AI政策 + 一桩Gemini历史谜案
Google正把轨道AI基础设施视为应对地球土地、审批、电网和社区约束的长期方案。 Project Suncatcher计划将TPU集群和太阳能阵列部署到晨昏低地球轨道,近乎持续的日照可能让面板效率达到地面面板的8倍;数据回传或许只会增加“几毫秒”。这是应对爆发式算力需求的真实选项,但还不是当下的生意:发射成本仍是建设同等地面数据中心的数倍,维修可能需要机器人。
第一个具体里程碑是2027年的双卫星原型,而不是轨道超大规模扩建。 Google称,新一代TPU在质子束辐射测试中承受的强度,已经超过5年任务预期;公司正与Planet合作测试。Starcloud、Axiom Space以及一个信息模糊的中国项目也在探索这一领域,Eric Schmidt和Jeff Bezos则释放了兴趣。Google把Suncatcher与Waymo、量子计算并列,称这是一个需要投入“8、10、12、15年”的项目;发射经济性和运营可靠性才是关键门槛。
共和党的AI政策仍是一组直觉的光谱,还没有形成定型的MAGA教义。 Dean Ball说,政府内部共享几条宽泛直觉:AI可能是数十年来、甚至有史以来最重要的机会;它既有熟悉的风险,也有“更陌生”的风险;并将塑造美国领导力。阵营从聚焦对华竞争的国家安全官员、关注儿童安全的保守派,到主持人所对应的David Sacks与Steve Bannon的相反立场不等。Ball预计,围绕垃圾内容、电力、水、就业、儿童安全、灭绝风险,以及“AI是假的”等说法,最终可能形成一锅“奇怪的冷汤”。
基础设施可能是各家实验室最强的竞争壁垒,而它们的监管立场则与自身激励一致。 超大规模云厂商未必反对出口管制,因为中国买家、甚至对TSMC有限晶圆制造产能的间接需求,都会与它们竞争;前沿实验室需要芯片来出售token,但Ball认为“模型参数不是你的护城河”。Anthropic宣布投入500亿美元建设数据中心,加入Google、OpenAI的Stargate、Meta和xAI共同参与的资本开支竞赛,最终均衡需要政府裁定。
Ball划出的监管边界是:联邦政府预防灾难性尾部风险,普通伤害在证据出现后再反应式监管,政府采购的模型则只受采购规则约束。 他认为“Woke AI”行政令针对的是意识形态偏见条款,以及联邦版本的系统提示词披露,同时承认右翼反对政府施压与自身施压之间的矛盾,并称强迫改变公开模型训练方式“明确违宪”。他希望对训练成本达10亿美元、面向全球提供服务的模型设立国家标准,但也认为国会必须处理儿童安全问题,并称加州SB 53透明度规则“总体上相当合理”。
一个身份未明的Gemini测试模型,似乎把手写转录从“令人印象深刻”推进到了专业可用。 在5份、合计约1,000词的基准文件上,Mark Humphries测得约1%的词错误率;相较Gemini 2.5 Pro在这些测试中的结果,错误率下降约50%,接近人工转录专家的水平。结果来自Google AI Studio的A/B测试,有时需要提交20或30次,因此Kevin Roose所说的“可能是Gemini 3”仍然只是押注,而非身份确认。
更深层的信号并非转录,而是一次看起来具备符号推理能力的账簿推断。 面对一条Albany记录,模型推断“14 5”指的是以每磅1先令4便士售出的14磅5盎司糖,并将其与19先令1便士的总价核对一致;Humphries说,“模型不应该能做到这一点”。如果结果可复现,模型就能汇总整本账簿并承担更广泛的档案级任务;但目前样本很小、模型身份未知,仍需在模型发布时复现。
1. 地面电力稀缺让轨道算力变得可以想象
Roose的前提是:地面数据中心需要土地、审批、电网容量和建设速度,而居民越来越反对其电力、水资源和环境成本。面对每一个拟议项目都可能遭遇“字面意义上,真的没有足够容量”的电网,AI需求的指数级增长让太空不再只是笑话,而成为一个默认AI使用量近乎无限的行业所需的应急方案。
Newton对投资者的表述更直接:这说明行业已经进入“这个泡沫的阶段”,企业开始相信地球无法为其雄心提供足够电力。即便轨道算力最终失败,对它的考虑本身也衡量了行业预期需求和资本开支增长的激进程度。
Google将Project Suncatcher称为“未来的、基于太空且高度可扩展的AI基础设施系统设计”,同时也称其为登月计划。今天没有任何公司在运营这类数据中心;该项目仍是面向一个可能在5年、10年或15年后到来的世界的主动研究,届时AI或许会持续服务几乎所有人。
2. Suncatcher把持续日照变成8倍能源押注
能源逻辑从规模开始:太阳的输出约为人类总产出的100万亿倍,但地面面板在夜晚到来后会失去发电能力。晨昏低地球轨道可以提供近乎持续的光照,让轨道面板的产能达到地面面板的8倍。
设想中的中心不是轨道上的仓库。Starcloud的模型看起来像“一只巨大的鸟”:薄薄的太阳能板翅膀为结构中央的计算机集群供电,整个结构则在地球上空运行。
把结果传回地面,可能反而是较不科幻的问题。业内人士告诉Roose,这可以类比Starlink:低地球轨道距离足够近,利用已经存在的卫星通信,传输延迟可能只增加“几毫秒”。
Google让一块标准TPU接受模拟轨道辐射的高强度质子束测试。新一代芯片承受了远超5年任务预期的辐射强度;硬件故障则更难处理,研究人员对Roose说:“我们得想办法”让机器人完成维修。
3. 2027年原型仍无法改变地面经济账
Google计划在2027年与测绘卫星公司Planet发射2颗原型卫星。Starcloud也在准备原型,因此近期竞争的目标是验证耐受性、连接能力和运行能力,而不是立即复制地面超大规模算力。
Roose的核心保留意见是:目前把大量芯片和卫星送入太空的成本,仍是建设同等地面算力的数倍。他预计至少数年内不会出现有意义的部署;更低的发射成本和可行的维修系统是前提,而非细节。
Google把Suncatcher放在Waymo和量子计算的同一条路线中,表明公司愿意投入“8、10、12、15年”,再等待其进入主流应用。只有当AI需求近乎无限、而地球最终同时缺少合适土地和电力时,这笔押注才成立。
这一领域还包括Axiom Space、一个细节含糊的中国项目,以及Eric Schmidt和Jeff Bezos可能的兴趣。当地居民的“不要在我后院”式反对,可能变成Newton所说的NOPs——“不要在我的星球上”;但太空碎片,以及富裕公司逃离地球问题的观感,最终也可能催生轨道反对者。
4. 共和党AI联盟拥有直觉,却没有定型教义
Ball在写完Hyperdimensional并在X上发帖后进入白宫,随后牵头起草其AI行动计划。他发现的是“连贯的直觉”,兴奋、担忧和困惑大致各占相同分量,而不是一套成熟的保守派政策目录。
第一条直觉:AI是数十年来、 “很可能也是有史以来”最重要的技术、科学和经济机遇。第二条直觉:AI既会产生可以用既有框架处理的熟悉风险,也可能带来政府缺乏清晰概念应对的“更陌生”风险。第三条直觉:AI将实质性地塑造美国的全球领导力。
Roose把David Sacks和Steve Bannon描绘成大致相反的两极:前者攻击“末日论”论点,后者谈论生存性风险。Ball接受这种光谱划分,但强调中间地带很大:包括聚焦对华竞争的国家安全官员,以及担心聊天机器人精神病、青少年自杀和社交媒体教训的保守派。
对于是否存在一种明显的MAGA式AGI理论,Ball的诚实回答是:“还没有。不,基本没有。”网上的MAGA讨论目前可能更偏向末日论,但许多参与者还没有形成对AGI的看法,更谈不上将其转化为具体的国内监管。
5. AI反弹会把矛盾熬成一锅政治浓汤
Ball说,在AI行业聚会中,反复出现的问题是“干草叉什么时候会出现”,以及什么会触发它们。他的答案不是一场单一灾难,而是一团混沌:内容垃圾化、面向儿童的不安全产品、电力和水资源消耗、失业、灭绝恐惧,以及同时存在的“AI是假的”这一信念,混合成“这锅奇怪的冷汤”。
他不认为会出现一套干净利落的党派立场。“AI政策”会像“互联网政策”一样分裂为数据中心、对华竞争、软件监管、儿童安全等问题;单个议题可能两极化,但AI内部差异太大,整体技术不可能形成一条持久的党派路线。
6. 基础设施可能成为前沿实验室的长期护城河
Ball反对把“行业”视为单一政治行动者:超大规模云厂商、前沿模型公司和其他供应商处于产业链的不同位置。“没有谁的论点是不正当的。每个人都在按自身激励行事。”政府要做的是在这些参与者之间找到均衡。
Microsoft、Google和Amazon Web Services未必憎恨芯片出口管制。当中国公司无法竞争同一批加速器、较少中国需求进入TSMC晶圆厂有限的产能时,它们都会受益,即使各家公司采购的是不同设计的芯片。
前沿实验室需要获得芯片,因为它们“想靠出售token赚钱”;但在Ball看来,比模型参数更可能形成护城河的是基础设施。他把Anthropic的500亿美元数据中心承诺,与Google、OpenAI的Stargate、Meta和xAI放在一起:每家公司都在通过拥有或锁定算力来建立防御能力。
7. “Woke AI”行政令暴露右翼反对政府施压的矛盾
Ball强调,行政令管的是联邦采购,而不是卖给消费者或私营企业的模型。政府要求供应商不要在提供给政府机构的特定版本中,自上而下植入意识形态偏见;它并未正式监管面向公众发布的模型。
他认为“客观”不是可行的技术标准——自语言诞生以来,人类就一直在争论真理——并称该行政令明智地避开了这个问题。它更狭窄的要求是,开发者不得额外强加一种世界观,并且在采购过程中披露政府版本的系统提示词等要素。
主持人的反驳是:共和党曾谴责拜登时期围绕新冠错误信息向平台施压,但特朗普政府同样在告诉科技公司产品应如何回应。Ball承认,这体现了特朗普之后的内在矛盾:一方面坚持反对政府施压,另一方面又使用政府权力“把它还给左派”。
Ball选择了原则立场:“任何人都不能施压。”但他区分了客户要求,指出政府模型本来就要承担更严格的《信息自由法》《总统档案法》和数据管理义务。他说,强迫改变公开模型的训练方式将“明确违宪”,同时侵犯公司的言论权和用户的言论权。
8. 联邦标准与正当的州级紧迫性相冲突
Ball的联邦主义论证取决于规模:模型训练成本可达10亿美元,且面向全球服务,因此其训练、评估和测量标准都涉及州际商业。各州分别建立制度并不现实;在国会缺位时,最大的州实际上可以制定全国规则。
因此,加州成了美国事实上的中央AI监管者。Ball称,这是制宪者无法预料的宪法失灵模式,因为当时还不存在现代经济的规模效应。不过,他总体支持SB 53这部仅适用于最大型开发者的透明度法律,称其“总体上相当合理”。
Newton的挑战聚焦现实伤害:聊天机器人精神病、儿童安全和青少年自杀倾向都是他描述为当下已经存在、且在一定程度上受到市场产品鼓励的伤害,而国会仍无力或无意监管技术。一名州议员完全可以合理地说:“我不想让本州的孩子自杀。”然后采取行动,而不是等待华盛顿。
Ball认同这种激励,并澄清自己的观点是主动性的:国会需要解决问题,而不是仅仅要求各州停止行动。他有时会批评立法者起草法律不佳——“让法院去解决”不能成为借口,因为议员同样宣誓效忠宪法——但他不会因此批评保护儿童的诉求。
9. 尾部风险值得预防,普通伤害可以等待证据
Ball借用了Ezra Klein对政府的描述:政府是“一项宏大的风险管理事业”。灾难性尾部风险和国家安全威胁需要成熟、最好是两党共同支持的预防措施;如果政府无法管理这些风险,就已经辜负了基本职责,不如“把钱还给股东”。
Ball说,前沿实验室中很多人确实真诚希望处理安全问题,但他无法代表这些实验室作为公司发言。他认为公司本身也有激励,因为如果它们引发一场大流行之类的事件,可能会破产。近期的生物和网络风险对资深风险官员来说“完全可以处理”——严重,但更像一场正在向佛罗里达逼近、可以追踪的飓风,而不是不可知的抽象威胁。
对于当前和近期技术,Ball不接受安全与加速之间必然存在取舍:政府可以改善生物安全,而不必实质性拖慢发展。他预计未来确实会出现真正的取舍,但届时应根据具体问题作出判断,而不是现在就预设可接受边界。
对于非尾部伤害,他主张在满足4项条件后采取反应式法律:伤害已经发生、很可能再次发生、普通法责任不足以解决、针对性法规能够产生实质帮助。儿童安全符合这一描述。他承认,一场灾难可能促使国会行动,但即使没有灾难,渐进式推进仍然可能。
10. 神秘Gemini进入人工专家级转录区间
Humphries和研究伙伴Lianne Leddy利用AI,把数万份关于毛皮贸易中普通人的手写记录连接起来。这些证据包括账目、合同、洗礼、婚姻和死亡记录,重建了大约1760年至19世纪初北美西部一批支离破碎的人生。
2023年的GPT-4只能“算是读懂”手写内容,但大多数结果都是错误。随后模型很快达到约90%的准确率,却难以突破;最后的10%恰恰包含历史学家最需要的姓名、金额和地点。之后,Gemini 2.5 Pro达到约95%。
他们的基准集包含50份被认为、但无法保证、不在训练数据中的文件。Humphries测试了其中5份,合计约1,000词;由于AI Studio会随机分配实验性A/B比较,他有时需要提交同一份文件20或30次,才能得到具有揭示意义的并排结果。
神秘模型的词错误率约为1%,其中包括大小写和标点错误;在这些测试中,错误率下降约50%,接近专业人工转录员的水平。Roose只能确认Google会在AI Studio中测试尚未发布的模型;他关于“可能是Gemini 3”的推断仍未得到验证。
11. 一本糖账暗示知识工作可能出现更大跃迁
更难的测试是一本文本来自18世纪Albany的账簿:手写快速潦草,表格结构即时形成,单位则是磅、先令和便士。这类记录类似为即时记账而非日后解读制作的收银条目,过去的模型在这类任务上表现很差。
Samuel Slitt的一条记录写着1条糖、简写数字“14 5”、每磅1先令4便士的价格,以及19先令1便士的总额。模型解释“14 5”代表14磅5盎司,并正确核对了数量、单价和总价。
Humphries惊讶的是,表面上的转录错误其实是澄清。模型必须从随机数字中推断含义,而不是预测一个可能出现的短语;识别进制不同的单位,并反向推算一套相对罕见的历史货币体系。对他而言,这看起来像符号推理:“模型不应该能做到这一点。”
如果结果能够复现,历史学家就可以要求模型找出并汇总整本账簿里的每一笔糖交易,而不只是转录文字。Humphries将这一机制推广到知识工作:转换信息、连接不同格式并推导含义;Newton则认为,这显示规模扩张仍可能产生涌现能力。但这两项结论都取决于对正式发布模型进行更大规模测试。
Casey Newton
What’s going on?
Kevin Roose
Oh my gosh. The other day, I was walking down Market Street. For context, this is one of the main thoroughfares in San Francisco, and over the past year, someone has recognized me from the podcast four or five times and stopped me to ask for a picture. It always makes my day. Hard Fork listeners are the best.
It had happened to me just the previous week. Then this weekend, I was coming home from the gym, and you know how you are when you’re coming home from the gym: Your face is flushed—
Casey Newton
Yeah, you’re sweaty.
Kevin Roose
You’re sweaty, your hair’s all over the place. This very sweet young woman comes up to me and asks for a picture. Of course, I’m thinking, “I kind of look gross right now,” but anything for a Hard Fork listener, right?
She’s there with a guy who I assume is her boyfriend or her husband, so I put on a show and I’m introducing myself: “Hey, what’s your name?” and all that. She hands me her phone, and they go and stand up against the street with their backs turned so they can get San Francisco in the background. That’s when I realize these people have no idea who I am. They’re just tourists, and they want a picture of themselves in San Francisco.
I’m Kevin Roose, a tech columnist at The New York Times.
Casey Newton
I’m Casey Newton from Platformer.
Kevin Roose
And this is Hard Fork.
Casey Newton
This week, Google’s crazy new plan to build data centers in space. Is this the final frontier of the AI bubble? Then, former Trump White House policy adviser Dean Ball tells us what Republicans really think about AI. And finally, it’s a history mystery. Professor Mark Humphries is here to talk about how an unidentified new Gemini model offered mind-blowing results on a challenging research problem. It was about Canada.
Kevin Roose
It was not about Canada.
Casey Newton
It was basically about Canada.
Kevin Roose
It was about sugar.
Casey Newton
It was about the sugar trade in Canada.
Kevin Roose
Okay, fair. Well, Casey, today we’re going to start by talking about space.
Casey Newton
Finally, the final frontier, some call it.
Kevin Roose
Yes, because I have been looking into this story that I have become obsessed with: We’re going to build fricking data centers and put them in space.
Casey Newton
I’m very excited to talk to you about this. I have been skimming the headlines, so I have a lot of questions for you about this. Whenever we can start an episode in space, that is a great place to start, because I don’t know if you’ve looked around lately, but who wants to be on planet Earth right now? I’d like an alternative, I’ll say that much.
1. The Space Data Center Bet
Kevin Roose
This has been quietly percolating in the tech industry. Obviously, we have this giant data center build-out going on here on Earth. Every company wants to build these giant data centers, fill them with GPUs, use them to train their AI models and do things like that. As you may have noticed, it is not easy to build data centers here on Earth.
Casey Newton
No, I’ve tried. I got nowhere. I felt like I was building IKEA furniture. It’s like, “You want me to do what?”
Kevin Roose
You need land, permits and energy to power the data center. You need to do all of this relatively quickly, and people sometimes get mad when you try to put up a data center where they live.
Casey Newton
Mm-hmm.
Kevin Roose
Also, we are facing an energy crunch for these data centers. There is literally not enough capacity on our terrestrial energy grid to power everything. That may get worse as people demand more and more AI and the growth continues exponentially.
Casey Newton
Yes.
Kevin Roose
So a couple of companies, including Google just recently, have announced that they are exploring a data center in space.
Casey Newton
Which sounds like a joke when you say it. Building anything in space seems so impractical, so expensive and so doomed to failure that it truly does just sound like a joke. But what you’re saying to me right now, Kevin, is that there is a legitimate, serious plan to try to do this.
Kevin Roose
Yes. I also thought this was some kind of crazy science-fiction moonshot thing. It is an experimental thing. No one is doing this today. But Google has put out a paper on what it calls Project Suncatcher.
Casey Newton
Yes, Suncatcher, which sounds like a lost Led Zeppelin single, but is somehow a project to build data centers in space.
Kevin Roose
Yes. They’re calling this a moonshot. They’re saying this might not happen for several more years, but it is an active area of research for them. There are a couple of other companies that have been doing this. Jeff Bezos, Eric Schmidt and other big tech figures are really interested in this idea, and I think we should talk about it today just to give people a sense of what the future may hold if we continue to demand all of this power and all of these data centers to run these giant AI models.
Casey Newton
I think it is so worth talking about because, among other things, it indicates that we are at the stage of this bubble where people have come to feel like we cannot provide enough electricity for the future we want to build on the planet that we live on. We actually have to get off the planet to realize our ambitions.
If nothing else, that tells you how ambitious these companies are getting and the crazy big swings that they’re about to take.
Kevin Roose
Totally. Where should we begin?
Casey Newton
Well, let’s talk about Project Suncatcher first. What exactly is Google proposing to do, and what did it say about it last week?
2. Google Builds Its Space Blueprint
Kevin Roose
This was a blog post and a paper that came out last week. They are calling this “A Future Space-Based, Highly Scalable AI Infrastructure System Design.” Basically, they have started doing some testing to figure out if a space-based data center would actually be possible.
The problem they’re trying to solve here is twofold. First, as we mentioned, it’s very hard to build stuff here on Earth. You need all of the permits, approvals and energy. Second, the sun is a really fricking good source of energy, right? It emits something like 100 trillion times as much energy as the entire output of humanity.
But building solar panels on Earth has some issues. Mainly, the sun sets for half the day, so you can only get power for half the day.
Casey Newton
Which has long been one of people’s primary criticisms of the sun.
Kevin Roose
Yes.
Casey Newton
Yeah.
Kevin Roose
But if you put the solar panels and the data centers into low-Earth orbit, and you put them on something called the dawn-dusk orbit path—which I did not just look up this week; I definitely knew what that was from my high school astronomy class—you can effectively give them nearly constant sunlight. The solar panels can be much more productive, up to 8 times as productive as solar panels here on Earth.
Casey Newton
Let me ask you this: When you say “data center,” I picture one of these giant, anonymous office complexes that’s the size of six football fields and that they’re building all over the heartland right now. I assume they are not going to build something like that in space.
Kevin Roose
No. If you look at some of the mock-ups that some of these companies have made, there’s another company called Starcloud that’s a startup with some funding from Nvidia. Its mock-up kind of looks like a giant bird, but the wings are very thin solar panels—arrays of solar panels—and the center is clusters of computers, essentially.
It’s just orbiting in space. The wings are catching all of the sun, and they’re feeding that energy into the computers at the center of the cluster.
Casey Newton
Got it. So we’re in one of these giant, terrifying, bird-like structures that are sort of swarming over the Earth in this future. They’re getting so much energy from the sun, and it’s so efficient, and that is driving all of the compute happening inside the computers.
How does whatever is happening inside the giant, terrifying bird get back to us down here on Earth in a timely fashion?
Kevin Roose
That’s a great question. I asked this to a couple of people I talked to over the past week or so who have been working on this stuff, and what they told me is that this is actually not that much different from something like Starlink, right? You’re sending data from a satellite or a series of satellites back to Earth.
It’s not that far away. It’s not like these are light-years away. It might take a couple more milliseconds than it would take to transmit something here on Earth. That is actually something that we know how to do.
Casey Newton
Got it. Okay, Kevin, so last week, Google put out a blog post about this. Give us a sense of where they are in this experiment.
3. Orbital Compute Faces Barriers
Kevin Roose
I would say they feel like they are pretty early in this process. There are still some technical barriers to overcome, and we can talk about those. But they have started actually running tests to figure out things like: If we send our TPUs, our AI training chips, out into space, will they just fall apart because of all the radiation out there?
They did an experiment that they describe in this paper, where they took a normal TPU, like the kind they would put in their data centers here on Earth, to a lab and hit it with a proton beam that was supposed to simulate a very intense kind of radiation that these chips would experience if they were floating out in space. They found that their newer TPUs actually withstood radiation much better than they thought.
Casey Newton
Hmm.
Kevin Roose
So these things can apparently handle radiation well beyond what's expected of them in a 5-year mission.
Casey Newton
Hmm. Now, if you watched The Fantastic Four: First Steps earlier this year, you know that cosmic radiation is what transformed the Richards family and Ben Grimm into the Fantastic Four. Has Google addressed that at all, about any of those concerns?
Kevin Roose
They did not address that, to my knowledge.
Casey Newton
Okay.
Kevin Roose
They did address some other potential hurdles. One of them is: If these chips glitch out or break, how do you fix them if they're in space? I asked a couple of people who have worked on similar projects, and they basically said, “Yeah, we've got to figure out how to get robots up there to fix the data centers.”
Casey Newton
Hmm. Got it. So they'll focus on using robots for that. I guess that makes sense. Now, am I right that Google is actually planning to do some kind of test launch within the next couple of years?
Kevin Roose
Yeah. They are planning to test this in 2027 by launching 2 prototype satellites in partnership with Planet, the company that sends up these little tiny satellites into space for mapping and things like that. That is their plan.
There are also other companies, including Starcloud, which is also planning to send up some prototypes pretty soon. So they are moving forward with testing on this. I will say, I think this is probably not going to happen in any real way for at least a couple of years, in part because things are still very expensive to send up into space.
Right now, it is not economically feasible to send up a whole bunch of chips and a whole bunch of satellites into space. It costs many times more than what you would need to build a comparable data center here on Earth.
Casey Newton
Yeah, and people here on Earth are saying that building the data centers that we're building here on Earth is not economically feasible, right? So I can't imagine how much more out of control the costs are going to be once you leave orbit.
One thing I thought was interesting in the Google blog post was that the company tried to place Project Suncatcher in the lineage of self-driving cars, so what is now Waymo, and quantum computing, which hasn't quite become a mainstream technology yet but has made a lot of strides. Within the past year, we did an episode on it not all that long ago.
They're sort of saying, “Project Suncatcher is one of those things where we are willing to work on this for 8, 10, 12, 15 years to make it into a mainstream technology.” I took that as Google saying, “Hey, this is not just some crazy little experiment that a couple of engineers are working on in their spare time.” It seems like they're serious about this.
Kevin Roose
I think they're serious about this, and I think they are looking out to a future 5, 10, 15 years away where the demand for AI and AI-related tasks is essentially infinite, right? This is not something that 10% of people are using every day. This is something that 100% of people are using constantly, where there are entire companies or sectors of the economy that have been fully turned over to AI.
Maybe that happens, and maybe it doesn't. But if it does happen, we're going to need a lot of energy and a lot of data centers, and we may run out of land and power here on Earth.
Casey Newton
Yeah. Now, something that I did not realize until after I had read about Project Suncatcher is just how many other companies are looking at doing the same thing. Can you give me a high-level overview of who else is playing here, and does it seem like anyone else is further along than Google is right now?
Kevin Roose
Yeah. As I mentioned, there's this company Starcloud, which is a Y Combinator startup that got some funding from Nvidia. They are sort of the main ones here doing this.
There's also a company called Axiom Space that is doing this. We think that there are some Chinese companies, or at least one Chinese effort, to do a space-based data center, although they've been a little bit vague about the details there.
The Information had an article about some comments that Eric Schmidt and Jeff Bezos have made, suggesting that maybe they are also interested in or looking at doing something like this.
Casey Newton
Well, you know, Jeff Bezos just put Lauren Sánchez into space.
Kevin Roose
Yes.
Casey Newton
So you have to wonder if that was kind of a first step toward something in this vein.
Kevin Roose
Yes.
Casey Newton
One thing I think is interesting about this approach, Kevin, is that, as you know, we've seen an increasing amount of resistance from people in local communities to having data centers put in their towns or near their towns. They're worried about how it's going to affect the cost of energy for them, right? They're worried about water usage or the environmental impact.
If this sort of thing comes to pass, we'll have gone from the NIMBYs saying “not in my backyard” to this new group of people that I'm calling the NOPs, who are saying “not on my planet.” They want all the data centers just built up in the sky. So do you think NOPs are going to become a major political force?
Kevin Roose
I do, although I also think that eventually people may start to not want them in space either.
Casey Newton
Hmm.
Kevin Roose
But it's going to be harder for them to protest. You have to get in a rocket and go up there into low Earth orbit. It's very inconvenient.
Casey Newton
Now, why wouldn't people want them in space?
Kevin Roose
There are various people who think that this is going to create a lot of space debris and things like that, which would eventually be bad. I talked to some folks who work on this stuff, and they said they don't think that's really going to be a big deal. There's all kinds of stuff up in space now. We generally don't pay much attention to it.
But I can see this sounding to people like Elon Musk proposing to build colonies on Mars or something. It's too futuristic, it's too sci-fi, and it sounds like these very rich companies and individuals trying to flee from their problems here on Earth by sending stuff into space.
Casey Newton
Here's what I would say: I would love to be living in a time when one of the top 10 concerns I had in my life was space debris. If I ever get there, Kevin, I will be in heaven. Heaven.
Kevin Roose
Well, you'll be in low Earth orbit, technically.
Casey Newton
I'll be low Earth orbit. Exactly.
Kevin Roose
Now, I have a question for you.
Casey Newton
Yeah, yeah.
Kevin Roose
Would you go to space?
Casey Newton
Yes, absolutely.
Kevin Roose
Would you go to space to fix a data center?
Casey Newton
I mean, what is the salary for that job?
Kevin Roose
Very high.
Casey Newton
I mean, there's probably a certain price for which I would do it. But here's the thing: You know I'm not handy around the house.
Kevin Roose
Yeah.
Casey Newton
If ChatGPT doesn't know what to do, I'm calling the handyman.
Kevin Roose
Yeah.
Casey Newton
Okay?
Kevin Roose
I will just say that I think we should make an offer to Google, which is: If you guys get Project Suncatcher up into low Earth orbit, we will do a podcast episode where we go up there and cut the ribbon.
Casey Newton
You're just dying to be exposed to massive levels of solar radiation.
Kevin Roose
You know, I just think it'd be fun.
When we come back, the ball is in our court. Dean Ball talks about how he crafted the AI Action Plan.
4. Washington Takes On AI Policy
Well, Casey, recently we've been talking about some state-level AI regulations that have been passed and signed into law. But today we're going to have a discussion about national AI policy.
Casey Newton
Yeah, I think that the states have been acting because the federal government has not really passed any legislation related to AI just yet, and that's left us with a lot of questions around how the administration has been thinking about AI.
Kevin Roose
It's been a little confusing. I think especially in this administration, it has not been particularly clear to me what President Trump and his allies believe about things like whether we are headed toward some kind of AGI moment or how the federal government should try to protect against some of the risks of very powerful AI systems.
So the conversation that we're going to have today, I think, will help us answer some of these questions and just get a better sense of what is happening in Washington, especially on the right, when it comes to AI and AI policy.
Casey Newton
Yeah.
Kevin Roose
So earlier this year, Dean Ball spent several months working as the White House's senior policy adviser for artificial intelligence and emerging technology. He was brought into the White House in order to lead the drafting of the White House's AI action plan.
And in that role in the White House, Dean not only got to see how the AI policy sausage was made at the highest levels of government, he actually got to make the sausage himself. He was sort of responsible for taking all these different ideas from the various parts of government and putting them together into a document that would represent the administration's official view on AI.
Casey Newton
Yeah, and while he was there, Dean also got a good sense of who the various factions on the right are when it comes to AI policy. What do they believe? What are the competing incentives? Who has whose ear?
And I think if you want to understand the likely path forward for AI regulation over the next few years, that's a really important part of the conversation.
Kevin Roose
Yeah. So Dean left the White House in August after the AI action plan was released, and since then he's become a senior fellow at the Foundation for American Innovation and the author of Hyperdimensional, a newsletter about AI and policy.
Casey Newton
And because we're going to be spending a lot of time in this segment talking about AI, let's do our disclosures.
Kevin Roose
I work for The New York Times. We're suing OpenAI and Microsoft over alleged copyright violations.
Casey Newton
And my boyfriend works at Anthropic.
Kevin Roose
Let's bring him in. Dean Ball, welcome to Hard Fork.
Thank you both for having me. It's so good to be here.
Kevin Roose
So how did you end up at the White House earlier this year working on AI policy? What was your background before that?
I was a think tanker. A lot of it was not tech policy. A lot of what I did was state and local policy, but I was always very interested in tech.
Basically, when the AI policy conversation really took off in early 2023, I made the decision to start writing about AI as a part-time gig, purely on the side. I wasn't being paid for it or anything. Eventually, I decided I really liked it and was finding my voice, and I was hired by the Mercatus Center at George Mason University to go spend some time there.
I spent about a year there, and then was recruited to the White House on the basis of primarily my writing on Substack. My Substack is called Hyperdimensional. It's where I talk about AI stuff.
Kevin Roose
The Substack-to-White House pipeline. I feel like you are not the only person who has posted their way into a job in the federal government.
You can post your way to the federal government. It's really true. Probably a big chunk of it was my posts on X, which is maybe even more scary.
Kevin Roose
Ooh.
Yeah.
Kevin Roose
So, okay, you get this call. You go to the White House. What did you find there with respect to AI policy? Was there a coherent single view of how AI should be governed and regulated?
I would say there are coherent intuitions, but the field is so nascent, and there haven't been a lot of fights where dividing lines have really firmed up yet. I think, by the way, this is true on the left as well.
Kevin Roose
Mm-hmm.
I don't think that those intuitions have formed yet into a lot of different, very specific policy positions. I don't think they've concretized yet is really what I'm saying.
I think, though, there's a combination of excitement and some worry and some confusion, probably equal parts. In a macro sense, that's probably roughly where I am too, actually, and that sounds about right to me.
Kevin Roose
You say there were some coherent intuitions about AI in the administration. What were those intuitions?
I think coherent intuition number 1 is AI is the most important technological, economic, and scientific opportunity that this country, and probably the world at large, has seen in decades and quite possibly ever. Basically everyone shares the assessment that this is going to be extremely powerful and really important.
The second intuition that directly follows is that there are going to be some risks associated with this that are familiar to us and cognizable under existing policy frameworks, and others that might be more alien and might be risks that we don't really even have concepts for clearly yet.
And then maybe the third intuition is, regardless of those risks, it feels like AI is going to play a very big role in the future of American global leadership.
Kevin Roose
Yeah, that's really helpful and kind of helps me get a sense of the lay of the land when you arrived.
5. The Right Splits On AI
I'm wondering, Dean, if you can help me understand the intra-right factions when it comes to AI, because I think I've identified at least 2 different views of AI that I've heard coming from prominent Republicans. Maybe you could call them the David Sacks view and the Steve Bannon view.
Uh-huh.
Kevin Roose
David Sacks, the president's AI czar, is constantly talking online and on his podcast about these AI doomers, who he thinks are ridiculous and are overhyping the risks of AI and trying to get their way on policy. He calls them woke, implying that they're trumping up these fears of, no pun intended, job loss and things like that to get their way when it comes to policy.
Mm-hmm.
Kevin Roose
Then there's Steve Bannon, who has been out there talking about the existential risks from AI. You and I were both at this Curve Conference—actually, all 3 of us were there—a few weeks ago, where one of Steve Bannon's guys was there and gave this very fascinating talk about how he thought he was sort of in league with the so-called doomers who believe that this could all go very badly very soon.
Are there more views on the right than those 2, or are those the primary camps?
No, I think that there's a whole spectrum. I can't speak for either David or Steve, of course, but I would put them at roughly polar opposites in terms of how conservatives talk about this issue. I think there's a whole spectrum in between.
First of all, you've got national security people. You have national security people who don't actually know a ton about AI—and this is true on both sides here. They're just thinking of this as a strategic technology that's important for U.S. competition with China and other things. Maybe they think there are some national security risks, but they're not really thinking about domestic policy. They're not really thinking about regulation. They're not thinking about EA versus doomer. So that would be one.
I think also, related to the Bannon viewpoint but maybe more toward the middle, would be people who are worried about kids' safety primarily. There are a lot of conservatives who would distance themselves from the AI doomer view, but who would also distance themselves from the pure accelerationist view. They would use the lessons we've had with social media as an example.
For these people, the issues of things like LLM psychosis and, of course, teen suicidality with chatbots are very salient. For everyone, I hope. There are others in between, and I would put myself somewhere in the middle in a weird fusion.
Kevin Roose
Where does industry fit into that spectrum? My sense from the outside is that industry groups and lobbyists have had a lot of success in this administration in getting what they want.
Kevin Roose
Where are they in those conversations?
I think it really depends on incentives. People in policy conversations very often will refer to industry as this kind of monolithic, coherent entity. It's of course not. There are different people who have different incentives. So, if you're a U.S. hyperscaler, you don't hate the export controls. You don't want more competition for the same chips that you're trying to buy.
Casey Newton
Meaning a Microsoft, Google, or Amazon.
Yes, Microsoft, Google, Amazon Web Services, et cetera. You don't hate that because, A, you don't want Chinese firms competing for your chips. But even if it's not the same chips you're competing over, you don't want to be implicitly competing over space at TSMC fabs to make the chips. So hyperscalers will definitely have nuanced positions on export controls. But by and large, their incentives are not to hate them, and they largely don't.
Kevin Roose
Yeah.
Frontier labs want to make money selling tokens to people. So they want access to chips. But I think there are some people who believe—and from a political theory perspective, it's not wrong to believe—that ultimately they want to create moats. And I think there are a lot of ways you can make moats. It seems to me like the main way they're trying to make moats right now is through infrastructure.
Anthropic today announced a $50 billion commitment to build their own data centers. Google obviously does this. OpenAI does this through Stargate. Meta does this. xAI does this. Everyone does this. Everyone's building infrastructure.
The basic view is, well, the models maybe are not your moat per se. The parameters of the model are not your moat, but perhaps the infrastructure is. So these are all competing interests, and no one's making illegitimate arguments here. Everyone's operating from incentives. And, of course, the job of government is to solve for the equilibrium.
Casey Newton
Is there a MAGA view of AGI?
Not yet. No, not really.
Casey Newton
Hmm.
I don't know that there's any political persuasion view of AGI. I think MAGA might actually be the closest to having one. And I think, at the moment, maybe the persuasion, at least from what I see online, is that maybe it's more doomer-y.
Casey Newton
I believe we saw a bipartisan bill introduced over the past week that would require reports of job losses due to automation, which suggests that there—
Yes.
Casey Newton
—is some increasing attention to that likelihood.
Yeah. There's this big question in the AI field. At places like the Curve and places like Light Haven, there are gatherings of various doyens of the AI community. They get together, and the main question that people talk about is, when are the pitchforks going to be out for this technology, and what is going to cause the pitchforks to come out?
I have come to the conclusion that, rather than it being a singular issue, it's going to be this kind of miasma of issues. It's going to be slopification. It's not safe for kids. It's driving up your electricity prices. It's using all the water. It's—
Casey Newton
Taking your job.
It's taking your job, and also it's going to kill everyone, and also, by the way, it's fake. It'll be all those things in this weird vichyssoise.
6. The Fight Over Woke AI
Casey Newton
The aspect of the AI Action Plan that I find the most annoying is the attention on the ideology of the chatbots and the suggestion that they should be able to respond in some ways but not in other ways. Can you illuminate the discussions that were being had and what the administration actually wants out of these models?
Yeah. The main point here, first of all, the most important thing—you're talking about the Woke AI executive order—
Casey Newton
Yeah.
—is what it is, how it's traditionally phrased. This is an executive order that deals with federal procurement policy. In other words, this is not a regulation on the versions of AI models that a company like Anthropic or OpenAI or any other company ships to consumers or private businesses. This is purely about the versions of their models that they ship to the government.
The government is saying, in this case, “We do not want to procure models which have top-down ideological biases engineered into them. We would like our government employees to have access to models which are objective.” I think objective is a really hard word. Obviously, we've been debating what truth is since there was language, right? So I don't think we're going to resolve that.
I have a feeling the General Services Administration guidelines will not resolve that issue. I think it's folly to even try, and I think the executive order doesn't try. The executive order steers clear of doing so. The executive order says instead, “We just don't want you as the developer imposing some sort of worldview on top of the model.”
Well, good luck with that, I guess.
Well, I want to ask one follow-up on that because my sense is that the Trump administration and Republicans in Congress have been very upset with how the Biden administration jawboned—how it applied pressure to social media companies to take down misinformation, or what they considered misinformation, about the COVID vaccines or things like that.
Mm-hmm.
That was seen as very inappropriate. In fact, there are ongoing investigations of the contacts between the Biden White House and the social media companies over this issue.
Yes.
And then we turn around, and we see this Woke AI executive order. I understand the subtle point you're making: This is not regulating the models that the companies are releasing to the public. It's just the ones that they're selling to the government. But we all know that there's one set of models, right? They get built, and they get sold to various customers.
I think it's reasonable to see that and think, “Okay, this is the Trump administration doing exactly what it got so mad at the Biden administration for doing, which is to contact the tech companies and tell them, ‘Hey, this is how your product should be working. This is the kind of thing it should be allowing and not allowing.’”
Mm-hmm.
Well, look, I think there is an inherent tension here, and this is a tension that has existed on the right, particularly post-Trump 45, post-President Trump's first term. There is this argument: Should we stick to our principles that the government shouldn't be doing this kind of jawboning, or should we accept that the government has this power, and now we need to throw it back at the left, right?
I can tell you that I personally have always definitively been on one side of that argument—
Which one?
—which is the former view. We should stick to principles.
No jawboning from anyone. Yeah, you shouldn't do that. At the same time, I think the government totally has a right to say—again, I wouldn't think of this as a model-training thing. I would think of this as the sort of thing that can be trivially easily changed by the developer, right?
Models that are sold to the government already have compliance burdens that are significantly higher than this executive order, right? They have to comply with the Freedom of Information Act. They have to comply with the Presidential Records Act if they're sold to the White House. There are all sorts of data stewardship laws that are way more difficult than anything in the Woke AI executive order.
The Woke AI executive order basically says, “You need to disclose in the procurement process to the agency from whom you're procuring what the system prompt is.” You can change a system prompt for a specific customer. It's not that hard.
I would only point out that if you did try to use federal law to compel a developer to change the way they train the models that they serve to the public, that is unambiguously unconstitutional. It is a violation of the First Amendment. You are violating that company's speech rights, and you are violating the speech rights of American citizens who might use that model.
So it would be quite dire and grave for the government to do that, and I am confident that the Woke AI executive order was not intended to do that.
Hmm. So, Dean, I really enjoy your newsletter. I've been reading it since before you joined the government. I continue to read it today. One point of view that you advocate for with great frequency is that most, if not all, AI regulation should be done at the federal level. And you spend a lot of very valuable time looking into how states are attempting to regulate AI in ways that I think you believe are mostly bad.
Could you give us a high-level overview of your interest in this subject and what you see states doing that concerns you so much?
Yeah, so I come from a state and local policy background, I should say. My view is that a lot of the real governance in this country happens at the state and local level. And now that I live in D.C., I mostly say thank God that’s the case.
That being said, there are some things that inherently implicate interstate commerce. I think that, for models that are trained to be served to the entire world and cost $1 billion to train, the standards by which those models are trained, evaluated, and measured have to be federal standards, because you can’t have competing standards. Now, maybe we don’t end up having competing standards. Maybe what happens is the biggest state regulates, and that happens all the time in America.
There are many, many technologies where the state of California or the state of New York or somewhere like that—Texas sometimes—has an implicit federal effect from one state doing lawmaking. I think that’s a failure mode. I think it’s a structural issue of our Constitution that the founders couldn’t possibly have contemplated, because the notion of economies of scale didn’t quite exist for them. And so I think it’s a really, really difficult issue of Supreme Court jurisprudence.
Right now, California, by default, is the central regulator of AI in America. Thus far, I think they’ve done a better job than I would have guessed, but still not a great job.
Hmm.
So I was broadly supportive of their flagship AI bill from this year, which was called SB 53. It is a transparency bill that applies only to the largest developers of AI models, and to me, it seems rather reasonable overall.
Let me bring it back to some more contemporary AI concerns, though.
Yes.
Earlier, when you were describing some of the landscape in Washington and who’s concerned about what, you mentioned there’s this group of Republicans who are very concerned about chatbot psychosis, child safety, and teen suicidality.
Mm-hmm.
Those are all harms that are present today and seem to be encouraged on some level by products that are out on the market.
Mm-hmm.
We have a Congress that is very loath to pass really any regulation at all when it comes to the tech industry, whether that’s for ideological reasons or just logistically—it’s very difficult to get Republicans and Democrats to agree.
Or the government’s shut down half the time.
That’s also been increasingly an issue. And so, in such a world, I can very much understand the point of view of a state lawmaker who says, “Well, I don’t want the kids in my state to kill themselves. We’re going to do something about this right now, and we’re not as dysfunctional as the federal government. So we’re going to get in there, and we’re going to try to do something.”
So how do you view that dynamic? And is your desire truly that the states would just say, “Hey, we’re not going to get involved, and that’s on Congress”?
No. I understand the incentives of the state lawmakers, for sure. I think Congress needs to act. My view is more proactive: Congress needs to deal with this. This is a problem that Congress needs to deal with.
I don’t blame the state lawmakers. Sometimes I do. Sometimes I blame them for poor statute drafting. There’s no excuse for that, right? It’s your job.
Hmm.
I say this sometimes to legislators, and they’re like, “Well, we’ll let the courts figure that out.” And I say, “No, you took an oath to the Constitution, too—not just the judges.” But in the general case of, “I want to protect kids in my state,” no, of course, I don’t blame them for that.
Yeah.
Yeah.
I want to zoom out a little bit and ask a question about AI and polarization. It feels to me right now—
Mm-hmm.
AI is kind of in this weird, confusing, pre-polarized state. There’s this sort of machine that, when an issue gets important enough or salient enough to enough people, gets run through the polarization machine, and it comes out the other side: Republicans take one position, and Democrats take another position.
Do you think something similar is going to happen with AI, where it will become very predictable which view you hold on AI based on which party you vote for?
I think what’s more likely is that over time, it splinters, and there are different things that people talk about. So there are going to be data centers, and there is going to be China competition that’ll be an issue. There will be software-side regulation. There will be the kids’ issues.
Just like today, we don’t talk about computer policy or internet policy. We used to talk about internet policy—in the ’90s, internet policy was a thing. But now it’s social media, privacy, whatever else. I think it’ll splinter in that way.
Will those issues themselves be polarized? Yeah, probably. In some ways they will.
I do hope, though, that there are certain parts of an issue—and this is a very important part of the action plan, in my view, too—
Mm-hmm.
Not every single aspect of an issue has to be polarized. There are legitimate tail-risk-type events, national security issues that I think it is the obligation of the federal government to deal with in a mature and responsible way.
I’ve heard Ezra Klein describe this before—I love this turn of phrase of his.
Who? I’ve never heard of him.
Yeah, we’re not familiar with his work.
Yeah.
I’ve heard him describe government as a grand enterprise in risk management.
Casey Newton
Mm.
I think that’s true. In a fundamental sense, I think that’s very true. And so there are certain things that we just need to deal with, and the action plan tries to make some incremental progress on some of those things.
And of course, there are a lot of things we need to do to embrace the technology and let it grow and all that, too, and I think that’s an important part as well, but that’s less controversial to say as a Republican. I think the thing that’s maybe more controversial right now to say is, yeah, there are legitimate risks, and I hope those things can be bipartisan—that dealing with those risks can be bipartisan.
Because really, if we can’t deal with catastrophic tail risk, then we do not have a legitimate government. The whole point of government is to deal with this issue, and we should, as Michael Dell said about Apple in the ’90s, throw the thing out and return the money to the shareholders if we can’t manage these things. I really do believe that.
Casey Newton
So let’s talk about that point specifically. When I look at AI policy in America today, I mostly see the big frontier labs getting just about everything they want, right? It seems like there is a high degree of alignment between the labs and the government.
And when it comes to safety restrictions, for example, I don’t see a lot that is holding them back from building their next 2 or 3 frontier models. So there are components of the AI action plan that are meant to address some of those catastrophic risks that you mentioned. Tell us how you envision that actually working. Where is the moment where the industry stops getting everything that it wants?
Well, I would say there’s so much you can say here. I think the first thing is that many of the people who work at the frontier labs—I can’t speak for the labs, of course, but having known a lot of them personally, including at very senior levels—I can say that they have an earnest desire to deal with these problems, and they invest real resources as companies. Part of the reason they do that is because they have incentives, because their companies would be bankrupt if they, e.g., caused a pandemic.
Kevin Roose
Mm-hmm.
And the other thing is that a lot of these problems are super tractable. We don’t have to act as though these things are the hardest problems we’ve ever dealt with.
To me, as someone with experience in public policy—and, by the way, this is the posture of people that I met in government who are 30-year veterans of thinking about tail risks—to them, you bring up AI bio risk or AI cyber risk, and they’re like, “Yeah, sounds like a serious risk. Okay, there’s a hurricane that’s tracking toward Florida. Let me go deal with that,” right?
These things come across your desk every day when you’re in government. These are eminently tractable problems in the near term with current technology and technology that I think we’re going to have in the near future. Without spending a ton of money, there’s a lot of traction you can get on them that doesn’t involve, really in any meaningful way, slowing down AI development.
I want to push back on the idea that there’s this trade-off between mitigating tail risk and slowing down AI development. Now, will that always be the case? No. At some point, there will be trade-offs. We’ll have to make those trade-offs, and they’ll be hard, and it’s hard for me to know where I’ll come down on that, because it’ll depend on the particulars.
But right now we have this great opportunity: we can accelerate AI development, and we can also have better biosecurity, which, by the way, was a problem before ChatGPT existed. There was a whole pandemic about it. So, yeah.
Casey Newton
Sometimes I talk to people who work on AI policy or just work on AI and think about policy, and they'll say things like, “You know, I don't think we're gonna get any meaningful AI regulation until there's a catastrophe.” Do you think that it will take something like that to really catalyze significant movement on AI policy in Congress?
Possibly. Certainly, a catastrophe is plausible and could catalyze movement in Congress, for sure. I think there are other ways to achieve this. I really do. I think you can make incremental advancements in the absence of a catastrophe.
Now, it depends on a lot of people in the AI safety community will say this. People at labs who care about AI safety will say this, too. That's a very Anthropic-type position. I don't say that as a pejorative, by the way.
Casey Newton
To be totally transparent—
I just say it descriptively.
Casey Newton
I've heard this from people at—
Yeah.
Casey Newton
—at lots of different labs where they're sort of like, “Yeah, I don't really think we're capable of...” And it's not so much a knock on this particular Congress or anything. It's just—
No, it's just Congress.
Casey Newton
—I don't think the government is capable of regulating things in advance.
I am okay with government being in a mostly reactive posture, particularly with respect to things that aren't tail risks. Tail risks are the one exception, because those things can be very, very damaging, and so you wanna do some stuff in advance to mitigate that.
But when it comes to most other harms from AI, I'm comfortable with government just really reacting to realized harms in areas where it's like, okay, well, it's a realized harm that we've seen. We think it's gonna continue happening. It doesn't appear to be resolved adequately by the existing system of common-law liability that allows people harmed to sue the people who harmed them, and it can be meaningfully addressed through a targeted law.
And if all those conditions are satisfied, then we should totally pass that law. I think kid safety is in this category.
Casey Newton
Yeah.
Kevin Roose
Yeah.
Casey Newton
Mm.
Kevin Roose
Well, Dean, thanks so much for coming. Really fascinating conversation, and people should check out your writing. Your website is Hyperdimensional.
It was a real pleasure, guys. Thank you.
Casey Newton
Thank you.
Kevin Roose
Thanks, Dean.
Casey Newton
When we come back, we'll have more to say about the Canadian fur trade than we've ever said before.
Kevin Roose
It was not the Canadian fur trade. It was the upstate New York sugar trade.
Casey Newton
They're related—in ways I don't understand.
Kevin Roose
Well, Scooby Gang, it's time to get in the old Mystery Machine, because today we've got a mystery.
Casey Newton
That's right, Gumshoes. Grab your notebook and your magnifying glass, because there are a few clues, and we're about to crack the case wide open.
7. The Gemini History Mystery
Kevin Roose
And this one is a history mystery. It involves an experiment that a historian ran using an AI model, and we're gonna talk about it all with the historian in just a second. But Casey, to set the scene here a little bit, there are a lot of rumors going around right now about this new Google Gemini 3 model.
Casey Newton
There really are. Gemini 2 came out almost exactly a year ago, last December, and while Google has updated it throughout the year, we have been hearing an increasing number of whispers this fall about Gemini 3 and rumors that it really is pretty great.
Alex Heath reported a few weeks back that he expected Gemini 3 to come out in December, and one thing that happens in the run-up to the release of new models is that companies quietly test them. That brings us to our story today.
Kevin Roose
Yes. So Mark Humphries is a history professor at Wilfrid Laurier University in Ontario, Canada. He does research involving a lot of old documents and trying to decipher the handwriting on these documents, and he is also kind of an AI early adopter. He's got a Substack called Generative History where he's been writing about his experiments using AI to solve some of his research problems.
Recently, he had a post that really caught our attention called “Has Google Quietly Solved Two of AI's Oldest Problems?” in which he explained a really fascinating experiment that he ran using one of these test models inside Google's AI Studio, which is a Google product where you can experiment with different models.
And he says that the responses that he got back from this mystery model made the hair on the back of his neck stand up. This was so astounding to him, not just because they were very good, but because they seemed like a different kind of capability than ones he had seen in any other AI model.
Casey Newton
Yeah. And so the mystery is what model Mark was using, but I think the bigger story is what it means that this historian was as impressed as he was with this very unusual thing that he found a large language model doing.
Kevin Roose
Yes. And we should say that it is very hard to determine exactly which model anyone is being shown at any given time, the way these pre-release tests go. Companies will show 1% of users one model and another 1% of users a different model and ask them to compare the two.
Casey Newton
And they give them weird code names. They don't tell you what you're using.
Kevin Roose
Exactly. So there's still some uncertainty around this. This may have just been a one-off. We will obviously need to see what Gemini 3 actually does when it comes out.
But for now, I think this is a very interesting story because it points to the way these AI models are starting to do things that surprise even experts in their fields.
Casey Newton
Yes. And so for those reasons, it's time to bring in Mark Humphries and talk about what he found. Kevin, you know the difference between an American and a Canadian historian?
Kevin Roose
What's that?
Casey Newton
Canadian historians process data, while American historians process data.
Kevin Roose
Is that true?
Casey Newton
Yeah, that's true.
Kevin Roose
Well, let's talk to Mark, and he can pronounce it however he wants.
Casey Newton
Hell yeah, brother.
Kevin Roose
Mark Humphries, welcome to Hard Fork.
Thanks for having me.
Kevin Roose
Where are we catching you today? Are you up in Canada? What's going on up there?
I am. I'm in Waterloo, Ontario, in Canada, in my office at Wilfrid Laurier University.
Casey Newton
So Waterloo, you must just be surrounded by AI computer scientists at all times.
There are a lot of startups, a lot of AI researchers, and a lot of computer companies in Waterloo, yes.
Kevin Roose
Home of the BlackBerry.
Casey Newton
That's right.
That's right.
Casey Newton
Yes.
RIM Park.
Kevin Roose
Before we get into the specifics of your most recent brush with this new mystery AI model, can you just tell us how you've been using AI in your history research over the last year or so?
Sure. My research partner and I—Lianne Leddy, whose lab this all comes out of as well—have been working on trying to develop ways of processing huge amounts of data, mostly handwritten, related to the fur trade.
That involves a couple of things. It involves trying to recognize the handwriting accurately, but it also involves trying to basically generate metadata for tens of thousands of records to try and understand what's in those records and make connections between them.
So we're operating at tasks that are just at the threshold of what AI models are capable of doing. It's been interesting to watch, over the last couple of years, the models get better and become capable of doing some of these things, and then finding out new limitations as we go along.
Kevin Roose
Yeah. Tell us a little bit about the kind of work that you do in general. I know you're really focused on using older documents in your work. What kinds of stories are you trying to put together?
I've always been really interested in stories of ordinary people. In the fur trade, when you're trying to understand what happened to ordinary people in the 18th and 19th centuries, the problem is that many of them were illiterate and didn't write.
Although they appear in a lot of documents generated in the course of living—marriage and death records, account books, and so forth—it's a lot of detective work. It's a lot of trying to piece together stories from fragmented documents: what somebody bought in one place, a contract they signed somewhere else, a baptismal record somewhere else.
A lot of this is trying to do that. That's what Dr. Leddy and I have been trying to do with our graduate students: piece together what these stories about ordinary people can tell us in the fur trade and in the western part of North America, from about 1760 through the early 19th century.
Casey Newton
You know, it's interesting, Kevin, because every time I go to a Starbucks and they try to give me a receipt, I think, “I don't need any paperwork about what just happened here. I'm just going to take my mocha and get out of here.” But what you're saying is that that document could be of huge value to a future historian trying to understand our lives.
Kevin Roose
Exactly.
Casey Newton
Yeah.
Kevin Roose
Yes.
Casey Newton
All right.
Kevin Roose
They will want to know.
Casey Newton
Let's get into it.
Kevin Roose
Tell us about this experience that you had with Gemini, the AI model that you were trying to use for this transcription—basically taking this very old document about the fur trade, plugging it in, and saying, “Transcribe this. Tell me what this says.”
8. Gemini Reveals Symbolic Reasoning
I think to understand why this is a significant, or could be a fairly significant, development, it's important to understand where we've come from in the last 2 years.
When GPT-4 first came out in 2023, it could sort of read handwritten documents. It was mostly errors, but you could see that it was beginning to be able to do this. It's been really easy for companies and systems to get up to about 90% accuracy, and then everything above 90% has been pretty difficult.
The problem is that that last 10% is the most important part. If you're interested in people's names, amounts of money, or where they were, you've got to get that stuff right in order to make it useful.
Up until about when Gemini 2.5 Pro came out last spring, we were still in that era. Gemini 2.5 Pro got up to about 95% accuracy, and that's really good. What I was interested in was, when we began to see reports on X that there were new models being tested by Google in AI Studio, which is its playground app, how much better would this get?
Kevin Roose
So, okay, you're hearing these rumors that there's a new mystery model inside AI Studio, where Google tests new models before they're released. What do you do?
Dr. Leddy and I have a corpus of 50 different documents that we've been using to benchmark how these models improve over time. They're all documents that we're pretty sure are not in the training data, because we've either taken them ourselves or they've come from sources that are not typically online. You can't be 100% sure, but that seems to be the case.
I started to put a few of those documents in. For your listeners who may not be aware, the way that testing of these types of models often works is that you have to put in the document dozens of times before you get a hit on the model you're hoping to test, because it randomly pops up. It's not an easy thing to do.
I managed to test about 5 of our 50 examples—about 1,000 words—and the results were impressive, to say the least. The error rate declined by about 50% from where it had been with Gemini 2.5 Pro, and it got to about a 1% word-error rate. That means you're getting 1 in every 100 words wrong, but that can include capitalization errors, punctuation, and things like that.
That in itself is really significant. No models come close to that. Human experts who do transcription for a living offer about a 1% error rate, so that itself is fairly important.
Casey Newton
And your sense, having used this new experimental model, did that just come from inputting dozens and dozens of queries and, every once in a while, getting a result that was radically better than the others, so you thought, “Aha, I must be getting the new one”? Or were there any other signs about what Google was showing you?
It's A/B testing. Normally in AI Studio, you put in a query and get a response. When you get the A/B test, you get 2 responses, and it asks you to rate which one is better.
The labs do this to get feedback on whether a model is actually better at specific types of tasks than other models. You might have to do that 20 or 30 times until you get one of those 2 responses, and then the differences were pretty notable.
Casey Newton
Hmm. You said the overall error rate fell by about 50%, but that was not actually what impressed you the most about this new model. What impressed you the most?
First of all, that was impressive. Then I was curious: If it's gotten to this point, how's it going to do on tabular data?
As historians, one of the things you work with, to go back to your Starbucks example, are receipts and ledgers from merchants in the past. A lot of that is fairly boring, but if you want to know where somebody is, where they bought their coffee one morning, and trace that person's movements, you can use these types of documents to do that. You can see what they bought and all of those kinds of things.
The thing is, to this point, models have been pretty bad on tabular data. It's often kept like a cash-register receipt system is kept, so it's done on the fly. Nobody's expecting people to necessarily read it down the road.
Kevin Roose
Mm-hmm.
It's difficult to interpret just by looking at it. It's also sometimes quickly written, so it's even worse handwriting than people are used to.
Because these are historical documents, in this case I'm dealing with records from 18th-century New York State, in upstate New York, in Albany. Those records are written in pounds, shillings, and pence, so that's the old system. It's a different base than we're used to using, with a different form of currency measurement.
When I dropped in a page at random from this ledger, I was curious to see what I'd get back. Suddenly, it not only came back in a near-perfect transcription, which was remarkable given how difficult it is to make sense of what's actually on the page, but as I started to go through it, I was looking for errors. I was trying to find errors, and I began to realize that some of the things I was seeing that looked like errors were actually clarifications. They required the model to do some really interesting things.
Casey Newton
Give us an example.
Sure. In the actual ledger document, what we're dealing with is a series of entries made in a daybook. People come into a store, they're buying things, and it's being recorded just like on a cash-register receipt.
In the particular entry I was looking at, it basically says, “Samuel Slitt came in on the 27th of March,” and it says, “To one loaf of sugar at 4, 14 5, at 1 4 0, 19 1.”
What that means when you break it all out is that this guy named Samuel Slitt came into the store. He bought one loaf of sugar. If you're not aware, in the 18th century, sugar came in hard, conical shapes. They broke off pieces and sold them to you.
It says, “14 5,” meaning 14 pounds and 5 ounces of sugar, sold at 1 shilling, 4 pence per pound, and then the total is 0 pounds, 19 shillings, and 1 pence. This is the old notation.
What I saw in the model's response, though, was that it had figured out that it was 1 loaf of sugar measured out at 14 pounds, 5 ounces, sold at 1 shilling, 4 pence, and then the total.
What's significant about that is that, in order to figure out that the random number written on the page—14 5—meant 14 pounds and 5 ounces, the model had to work backward from a different currency system with a different base.
The thing that makes that important is that models shouldn't be able to do that. The way these models are trained is in pattern recognition. What they're trying to do is predict the next token.
The first problem here is that predicting numbers is actually very difficult for models. The model has no idea whether Samuel Slitt is buying 14 pounds, 5 ounces or 13 pounds, 6 ounces. That's effectively a random number.
It's not probabilistic. The other problem is that although there would be a lot of material in the training data related to this kind of old currency system, the reality is there's not that much of it in terms of the actual percentage of the material that's there, because there's so little of this out there in terms of the overall sum total of all the records that exist.
And so when we're thinking about it, the model's having to do some interesting things there. What it looks like to me is a form of symbolic reasoning. I have to know in my head that I'm dealing with different units of measurement that don't have a common kind of base pair to multiply or divide by, and then I have to abstractly realize that these units of measurement are, in fact, comparable as long as we do some conversions. We have to then move them around in our heads to figure it out.
This was something that I had to think about for a second and realize: In fact, the model had done something that was mathematically correct and unexpected.
Casey Newton
So what are the implications for you in your work of a model being able to do this kind of abstract reasoning?
Yeah. As a historian, what it means is that, assuming that this replicates once we start to see the actual model come out, you're going to be able to trust the models to do a lot of stuff that historians would normally need to do.
It's one thing to transcribe a document. It would be another to say, “Here's a ledger. Go through and add up all the sugar that was bought and sold in this ledger.” Right now, you can't trust a model to do anything like that. You can't trust it to necessarily recognize sugar, come up with quantities, or do that type of math. If we're getting to a point where models can begin to do that, you can begin to get them to do tasks that would take humans a very long time.
Kevin Roose
It sounds like the equivalent of the moment when AI coding tools went from being a useful—
Yeah.
Kevin Roose
—assistant for a person who's a professional programmer to actually being able to go out and program things just on their own with very minimal instruction. It's like that for history, right?
Yeah.
Kevin Roose
Well, Casey is—
And that's—
Kevin Roose
—but he's a special case. Yeah.
Yeah, that's fair—
Casey Newton
I'm really interested in this Samuel Slitt and why he needed 14 pounds of sugar. Take it easy, Sam.
It's true. Well, he's a merchant; he also wants to go and sell it to other people, right?
Casey Newton
Oh, he's a dealer.
There we go. Now I understand.
He is. He's a sugar dealer.
Casey Newton
Yeah.
But the interesting thing about this is that the stuff we do as historians with these historical records is what all knowledge workers do, right? You take information and synthesize it, take it from one format, put it into another, realize the implications of the things that you're reading, and draw conclusions and analysis based on that.
It can be 18th-century sugar, but it can very easily be any other kind of widget that a knowledge worker uses. So what I'm seeing turning on here for historians is highly likely to start turning on in other areas as well.
Up to this point, we've been getting this sense that the models are starting to get good enough where you can feel like, “Yeah, I think I can trust the outputs on this,” but you're getting to the point where it just works. As somebody who uses coding assistance all the time now, it's a very similar situation. You used to have to cut and paste back and forth, and it would never run the first time. You'd have to run it three or four times, paste the errors back and forth, and eventually it would work. Now you can just hit the button and it almost always works, right? And that's what—
Kevin Roose
Yeah.
—we're going to see here with knowledge work.
Casey Newton
So I want to zero in on what makes this so interesting. We don't know at this moment that this is Gemini 3, but I think Kevin and I feel like it's highly likely to be Gemini 3, right? We also don't know a lot about how, if it is Gemini 3, exactly how it was trained, but I think we can assume that it was trained in a way that its predecessors were, which was in part by just feeding it lots more data and lots more compute, right? Just following the scaling laws.
There's been so much debate over the past year about whether we're seeing diminishing returns, right? Have we figured out the limits of what we can get out of these scaling laws? The story that you're telling us, Mark, is a suggestion that, no, we have not gotten everything there is to be gotten out of this increased scaling. In fact, we should expect to see continued emerging properties from this ongoing scaling, and you've just given us an example of it right there.
So that's why I think this is so fascinating.
Kevin Roose
Yeah, I was fascinated by this experiment, and I wanted to see if I could actually get to the bottom of what happened here. So I asked some folks who would be in a position to know, “Hey, there's this history professor in Canada. He thinks he stumbled onto an unreleased Gemini 3 A/B test, and it was really good.”
They said, “Lose my number.” No, they were very tight-lipped. They did not want to talk about it. They are keeping things very secretive over there.
But I was able to confirm that Google does test new models in AI Studio before they appear elsewhere. So if I were a betting man, it's a pretty good bet that what you experienced was, in fact, an unreleased model, probably Gemini 3.
Casey Newton
Hmm. So, Kevin, I have not been in AI Studio myself recently to see if I could try this model. Have you made any efforts to try to access whatever this model is?
Kevin Roose
Yes. I use AI Studio. People don't know this, but Google has 800 AI products right now.
Casey Newton
Mm-hmm.
Kevin Roose
There are, like, 800—a billion ways to use Gemini. The most effective way—the best way—to use Gemini is inside this product that basically no one except—
Casey Newton
Hmm.
Kevin Roose
—developers and nerds like us uses, which is called Google AI Studio. If you go in there, I don't know, for whatever reason, Mark, do you find this too? But the version of Gemini in AI Studio is better than the one on the web.
Casey Newton
Hmm.
Kevin Roose
I don't know why.
Yeah.
Kevin Roose
But this is something I'm consistently able to get AI Studio to do, like transcribing long interviews, that the regular old Gemini won't do.
So anyway, I was in there this morning, actually doing some research for our segment about Suncatcher, this Google project about putting AI stuff in space. I was trying to have it summarize this research paper and give me some ideas and comparisons to what other companies are doing, and I got this A/B test, this choose-between-these-two-answers thing.
I'm looking at it right now. It says, “Which response do you prefer?” And it has these 2 side-by-side things, and they basically both look pretty good. I think the problem I'm identifying is that, unlike you, Mark, I'm not smart enough to come up with problems that are challenging enough where the difference between one pretty good model and a very good model is readily apparent.
So maybe you can help me with that.
Casey Newton
Well, I mean, here's an idea. I know Mark really focuses on the 1700s and the 1800s in the fur trade. What about the 1500s?
I bet you can make a dent.
Kevin Roose
Yeah. Well, I'll look into that.
All right. Well, totally fascinating experience, and I can't wait to hear more about what you're doing with AI and history. This is a really interesting mystery that I hope we've shed some light on. Thank you, Mark.
Kevin Roose
Thank you very much for having me.