为什么AI的下一轮突破可能来自大实验室之外
Erik TorenbergAaron LevieMartin CasadoSteven Sinofsky
与会者一致认为,前沿实验室需要更强的治理、沙箱、测试和安全机制,但不接受把“控制节奏”作为一个自洽的政策框架。 Aaron Levie认为,谨慎的工程实践反而能加快企业扩散,因为客户必须信任产品;Martin Casado则反驳称,慢慢制造核武器并不会让它更安全。随着实验室融资规模不断扩大、以“狂奔”速度推进,Steven Sinofsky追问:所谓节奏,究竟是相对于什么而言?
实验室关于生存风险的论述,按逻辑最终指向的是国有化,而不是自愿克制。 Casado的判断很直接:如果最了解这项技术的人相信AI会带来生存风险,就应当实施政府级管控;如果他们并不相信,那么招募持有这种观点的研究人员只是“人力资源问题”,不足以成为对国家技术实施全面封锁的理由。Torenberg认为,实验室可能希望风险保持在足以证明特殊管控合理、但又不至于让自己承担被当作核设施或生物基础设施对待的程度。
AI监管正在不可避免地到来,但要求政府交付一个经过校准的结果,本身就是范畴错误。 Levie认为,政府会听取相互竞争的利益群体并形成妥协,因此“谁也不可能如愿”;Sinofsky补充说,结果可能让一方觉得不够,也可能对另一方走得远得多。Casado预计,2028年会成为“AI大选”,而分散的AI收益将难以对抗反对者已经占据的“暂停”“蜂群”和“失控”等话语。
可投资的安全问题是具体的:代理集群正在使那些围绕大多数可信、但受容量约束的人类设计的系统失效。 现有控制体系假设人类有95%-99%的时间会正确行事,恶意员工可能只有1/10,000;软件代理则可能表现得像10,000架“游荡的无人机”,耗尽内部API,并把正确指令当成错误指令。这意味着身份、授权、可观测性、操作系统、网络和细粒度数据权限都将迎来一轮广泛升级。
与会者希望AI政策建立在已经证明的失效模式之上,而不是无法证伪的预测之上。 早期互联网政策讨论有蠕虫病毒、医院宕机和具体的未授权入侵作为依据;如今,实验室的安全复盘被批评为“草率”且不完整,缺少运营细节,也很少采用成熟的漏洞报告实践。当Noam Brown提出通过CPU热量外传信息的思想实验后,安全工程师开始计算实际比特率,这成为建设性的连接点:“至少我们已经把问题缩小到了物理定律。”
一种可能的监管失败模式,是用同意弹窗的戏码分配责任,却通过用户习惯化削弱安全性。 Sinofsky担心欧洲会把AI也“GDPR化”,要求代理每次向第三方产品写入或通过其执行操作时都弹出警告。Windows User Account Control、Word宏警告和浏览器选择弹窗都展示了最终结果:用户不断点击跳过,监管者要求警告更醒目,所有人最后变得麻木。
本期最看多的技术判断是:下一次软件突破可能发生在前沿实验室之外。 节目末尾讨论的Jevons式选项选择模型,不再进行昂贵的自由生成,而是围绕预定义选项快速给出低成本概率,终于为传统程序提供了一个实用的语言模型接口。这让沉寂数十年的概率编程思路重新获得意义,也支持Sinofsky的结论:创新中心“刚刚从模型本身转移到了围绕模型构建的系统”。
1. “控制节奏”遮蔽了真正的工程共识
Levie的出发点有意设得很宽:每一家前沿实验室都应追求实际可行的最高水平,包括治理、安全、沙箱、测试和对齐。这些机制可能拖慢单次发布,但会加快产品扩散,因为企业不会采用自己无法信任的产品。
他担心的是政策层面的后果。一套合理的安全计划,可能被用来压制竞争、显著放慢AI发展,或阻止数据中心建设,最终把必要的工程卫生变成监管俘获的工具。
Casado基本认同Dario Amodei提案的实质内容,但认为“整体氛围完全失控”。当讨论允许物种灭绝概率达到10%,而Amodei又公开接纳这一前提时,一份温和的节奏控制方案看上去非但不能安抚,反而显得严重不够。
Casado的核武器类比暴露了措辞问题:“你完全可以非常缓慢地制造核武器”,但缓慢并不会带来安全。Sinofsky补充说,目前没有任何公开的时间表,因此要声称进展更慢,就必须知道一个从未有人披露过的速度。
2. 生存风险制造了国有化困境
Casado结合自己在Lawrence Livermore National Laboratory的工作经历指出,如果最了解这项技术的人真的相信它会造成生存风险,“答案就是把它国有化”,并落实已经知道的控制措施。
Levie反驳的例子是:有人正因为担心AI,才推动AI发展,希望这项技术能以正确方式完成。Casado并不否认这种人存在;他反对的是把一个内部招聘和人才留存群体,变成“国家级封锁”的依据。
分歧在边际风险问题上进一步尖锐。与会者讨论了Amodei拒绝给出具体概率、却仍与声称存在灭绝风险的人交流这一做法;Levie认为,政府不可能从行业领袖那里听到“物种灭绝概率为10%”,随后又把AI当作普通软件处理。
Torenberg认为,一些实验室可能希望风险保持在非零水平,以证明特殊管控合理,但又不能高到必须国有化。Levie仍然相信实验室实际采取的安全措施正在改善;悬而未决的问题,是如何让这些措施与实验室的公开表述保持一致。
3. 政府会自行决定监管速度
Levie从政治中得出的教训是,企业不可能递交一份心仪的政策,再原样收回。政府会听取相互竞争的利益群体并形成妥协,而这“100%的时候”都会让一方觉得不够,或让另一方觉得走得远得多。
行业的软肋在于叙事。支持AI的一方必须不断附加限定条件,听起来像是在防御;反对者则已经占据了“暂停”“蜂群”和“失控”等鲜明词汇。Casado预计,2028年会变成一场围绕AI的公投,即使没有候选人能够提出清晰的亲AI平台。
科技公司一再误判这片地形。Sinofsky提到AT&T、IBM和Microsoft:即便这些公司是在政府体系旁成长起来、拥有成规模律师团队的企业,也会因反垄断而发生结构性改变。Bill Gates与Bill Clinton一起打高尔夫,也没能阻止Clinton政府起诉Microsoft。
好莱坞的Motion Picture Association展示了自愿自律的模式;FINRA则说明,自律可能最终变成强制监管。与会者认为,一旦AI进入医疗、医疗器械、交易系统和飞机领域,接受类似FINRA的待遇可能已经是合理的最佳情形,但这也可能让开源项目和前沿挑战者处于劣势。
4. 好政策应跟随已观察到的失败和具体事实
Casado将推测性的AI立法与早期互联网作对比:当时蠕虫病毒冲击基础设施、医院瘫痪,经济损失达到数百亿美元。政策因此可以针对已知机制;而在如今这种预先防范的氛围下,他承认,如果按1994年的方式治理,“我们大概根本不会有互联网”。
Sinofsky回忆说,Windows还在安装时,PC就已经开始感染病毒。行业最终强化了开箱流程:现代手机可能只连接网络以获取操作系统更新,在完成补丁前,其余功能保持禁用。
节目讨论的联邦计算机犯罪法,源于1983年针对GTE Telenet系统的具体入侵事件;这些系统被NASA和国家实验室使用,法律随后花了约2年半才通过。之后,检方又增加了例外规定,将善意的安全工作和单纯违反服务条款,与恶意的未授权访问区分开来。
因此,Casado认为,应用层面的法律大部分已经存在。他更尖锐的批评集中在运营层面:实验室的复盘看起来“草率”且经过选择,遗漏了内部消息等重要材料,也无视结构化的安全报告实践;这“不是一次复盘”,更像是在只提供筛选证据的情况下进行调查。
5. 隐蔽通道将抽象恐惧转化为工程问题
Casado描述了一次少见但富有成效的碰撞:生存风险理论家与系统安全从业者展开了对话。Noam Brown提出超级智能可能通过CPU热量将自身外传后,工程师没有继续交换哲学判断,而是质疑前提、计算可能的比特率。
Levie对此感到宽慰,因为热量通道是否可行,可以通过熵、物理学和系统设计来讨论;真正不可接受的是,一边轻描淡写地处理普通认证漏洞,一边突出近乎无限的假设能力。
Casado给出的最佳例子说明,不能简单嘲笑这种奇异的信息泄露路径:在一间漆黑房间里的老式CRT显示器上,传感器读取以光栅频率反射的窗外光线,就可以重建屏幕;只要篡改3个看似有缺陷的像素,就能形成带宽相对较高的单向通道。
6. 代理集群击穿人类时代的访问控制
关键变化不是神秘智能,而是穷举式的持续性:AI可以快速尝试每一条攻击路径,不会疲倦,也不会厌烦。内部的GitHub、Slack、财务和报销API从未按恶意高流量设计,却可能突然遭遇类似拒绝服务攻击的行为。
企业安全过去部分依赖于一个事实:95%-99%的人通常会做正确的事,恶意内部人员则很少见。代理集群“完全颠覆了这一点”:10,000个四处游走的软件工作者可能携带信用卡和凭证,同时把正确任务与错误任务混为一谈。
Casado认为,当前权限体系严重两极化:代理要么每个动作都请求批准,要么获得足以“删掉整台电脑”的权限。真正可用的部署需要直觉化的控制方式,例如对一个文件夹拥有读写权限、对另一个文件夹只有只读权限,并且只能调用边界清晰的工具。
Casado指出,多级安全、操作系统研究、网络和编程语言领域其实早已具备这些技术基础;之所以没有普及,是因为普通用户过去不需要承担这种复杂性。AI既可能消耗这些控制,也可能帮助构建它们,从而在安全内生系统领域带来“一次文艺复兴”。
7. 基于警告的监管可能重演GDPR最糟糕的激励
Sinofsky最大的监管担忧,是欧洲认定“GDPR是有史以来最好的东西”,然后把这套基于弹窗的做法应用到代理上。代理每次向第三方服务写入数据,或进行非查询式交互,都可能触发弹窗,把责任重新推回用户。
监管者喜欢警告,因为它们能直观地分配责任。但Windows XP的User Account Control、Word宏警告、浏览器选择界面和汽车贴纸都说明了同一个问题:反复出现的弹窗会让人自动点击,之后监管部门往往只会要求把警告做得更大。
Casado希望用这一现实层面的论点取代灭绝哲学。他在Stanford使用的示波器运行得异常缓慢,后来发现有人攻破了它老旧的Windows CE系统栈,并把它改成了色情服务器。这说明易受攻击的联网系统曾经无处不在,但最坏情况下的恶意结果相对少见。
历史上的滞后同样重要:汽车约在1900年出现,《任何速度都不安全》(Unsafe at Any Speed)直到1960年代中期才问世;飞行员最初还要自带飞机完成粗略的执照认证,现代航空监管则用了几十年才建立起来。与会者引用的Nick Bostrom观点是,过早监管可能保留最终风险,同时在任何人理解系统之前就把规则冻结。
8. 软件原生AI可能把创新中心推向实验室之外
结尾的技术案例——临近节目末尾讨论的Jevons式路径——从一个错配出发:LLM生成的是对话文本,而传统程序需要结构化决策。用schema提示模型仍然“很别扭”,因为自由生成可能忽略所要求的格式。
它提出的替代方案是读取文本,但不生成答案,而是在给定选项中进行选择。Casado认为,这样可以更快、更便宜,也更准确,因为模型本来就是为选择而训练的;他称其普及速度“可能是ChatGPT以来最快的AI模型采用速度”。
Sinofsky认为,行业早就该摆脱自然语言界面,因为无论提问还是回答,自然语言都效率低下。模型可以改为返回“80%是客服”,为概率分支提供输入,而不是走脆弱的“是/否”路由,并重新激活1960年代和1970年代发展起来的模拟技术。
Casado的对比很尖锐:前沿实验室往往像是在创造生物,而“生物会说话”;软件开发者需要的却是可靠组件。Sinofsky从这个外部模型中看到的信号是,“创新中心刚刚转移了”——平台达到临界规模后,外部开发者会创造出平台后来可能再以“Sherlock”方式吸收的功能。
完整逐字稿
If you regulate AI too early, you actually don't solve anything. You still just have the same risk, ultimately.
You willed the thing into being, but you haven't figured out how to control it.
The problem we have now is this rift between the labs and the security community, which keeps coming to two conclusions.
You're sloppy, and you're not complete in what you're telling us happened.
An employee is like, “There's a 10% chance of species extinction.”
Agent swarms completely flip that. These are just roaming drones, but times 10,000, and they will easily mistake a good task for a bad one. So now we need a whole layer internally that's tracking much more about what authentications are being done and what APIs are being used.
The U.S., about 15 years ago, stopped leading in tech antitrust. The problem is that Europe is going to lead with that because they have nothing to lose.
This could change the nature of software fundamentally. The center of innovation has just moved.
1. Pacing the Frontier: Reacting to the AI Safety Discourse
Guys, welcome back to the podcast.
Happy to be here.
Thank you.
Your beard.
Didn't think we'd ever do this again. I can't believe it. This is great. I mean, Martin's just building these $100 billion companies. He's too busy for this podcast.
At least taking credit for it, as VCs do.
Exactly. We have a lot to discuss today, but first, Aaron, why don't we start with you? Pacing the frontier: How have you reacted and reflected on what's happened there and the discourse that's followed?
Oh boy. I think we should start with Martin on this one. He was fighting lots of good ground wars.
Maybe I'll say one thing that we probably all agree with, and then we can figure out where we fracture off. I think we would agree that any AI lab right now at the frontier should be building in the safest way possible, with the highest degree of governance and security. Whatever your definition of alignment is, this is an incredibly important area of research. It's an incredibly important area for the diffusion of AI. You're not going to have AI diffusion without extremely high-quality products that can be trusted by enterprises and that aren't constantly hacking systems.
When I read Dario's post, I didn't disagree with almost anything. It was all about how to have better security for these systems: sandboxing, better testing. There are going to be some debates around the embedded nature of the testers, whether you agree with who those are, and whether the industry all aligns on that. But I think all of the major points were salient and appropriate.
The only question is whether this gets used or leveraged to do things that maybe we don't agree with, which would be a dramatic slowdown of AI because of regulatory controls that would not make it easy to compete with the frontier labs. Or politicians could take the message and run with it, leading to even worse outcomes. It could be used to ban data centers far faster and so on.
I think the actual substance of the topic is incredibly important, and I think it's very important for AI advancement in general. The question is what you do about it, especially from a regulatory standpoint. That's probably where the industry is going to land on very different points along the continuum. Martin was putting up a good fight about making sure that we don't use this for regulatory capture, which I also agree with. But I think the ideas in the pacing conversation are important.
2. Species Extinction Talk: Do Labs Actually Believe Their Own Rhetoric?
Again, it's a little bit of a funny concept because maybe it's not even pacing as much as good hygiene and good engineering. With good engineering, obviously there is a slight slowdown, but it's a slowdown that allows acceleration of diffusion, because you wouldn't be able to have any of the AI diffused if nobody trusted using it. The post is very reasonable, but the atmospherics are not right.
I agree with him more than I disagree with him, right?
On TV—the same day that he landed these things. You can't have these conversations in isolation, because if he's going to agree that there's a risk of species extinction—
This post that he has looks like this milquetoast capitulation that's totally not adequate for the task at hand. I think the atmospherics are totally broken, and a lot of my comments were about the atmospherics.
I have this quibble, but it really bothers me, because I'm a pedant: I think pacing is the wrong way to describe this. For one, yes, it is orthogonal to security.
Mhm.
You can very slowly build a nuclear weapon, and that doesn't make anybody feel better. It's slow versus fast. The second thing is that it feels like a capitulation to the pause folks without actually addressing it. You're saying, “We're not going to pause. We're going to pace to make them happy, but we'll also somehow make the regulators happy.” I think it makes them both unhappy.
The pause people are saying, “Well, that's not a pause. That's just pacing.” And the regulators are saying, “You're still doing the thing.”
I just feel like they're trying to split the difference between an internal fringe faction—the doomers and the pause folks—and the regulators on the other side. The problem is they're making both of them unhappy.
I think what they have to do is come out and address the existential-risk question directly. They need to say, “No, we don't think the stuff we're going to do is going to cause extinction.” Then I think this becomes a very sensible proposal.
Just to play the other side for one second, what do you do? Do you make room for the one possible Venn diagram, which is the lab researcher who is simultaneously super scared but also works on advancing the state of AI because they believe it's so important to get right that they want to pursue it? Obviously, the language is problematic right now, but that person does exist. That's a real kind of person in our industry: “We have to be at the forefront of AI. I'm also very scared of it, and that's why I'm working on this.”
Let me address this directly. I used to work for Lawrence Livermore National Laboratory on a weapons program—nuclear weapons. I know what it's like to work on things that have access.
You were the first pacer.
We were part of the first pacers. If a constituency within the labs—who are the most knowledgeable people—believes this stuff has existential risk—
Yes.
3. The Nationalization Question & Whose Job Safety Really Is
The answer is to nationalize it and actually put controls in place that we know work.
Right now, if they don't actually believe that—and in the private conversations I have, the most sensible people don't; it's a small fraction that do—this is an HR problem. To me, an HR problem is a company problem. If they're worried that they can't recruit people—
Right.
—or they can't retain people—
Which it seems to me is a lot of this concern. It's almost more of this kind of researcher currency.
If that's the case—
I think that is the wrong reason to cause a national-level lockdown on a very promising technology. Listen, I think the post is very sensible. I think there are real concerns around security. We've had many compute epochs that have real security concerns.
I don't think you can reconcile discussions on existential risk with the proposal that was put out. You just can't reconcile those 2 things, and that has always been my primary point.
Right. Right. Okay. Unleash.
Holding it back.
So, okay. The first thing is, is there a schedule they've published that says when all this bad stuff is going to happen? They haven't. You can't pace it because nobody knows when it was supposed to finish in the first place.
It also just seems disingenuous. This is like when the press reports that Apple's latest iPhone is late.
The phone from what? Nobody knows it exists, and they haven't told anybody about it.
Yeah, my Apple car was very late.
Like, I don't understand. In order for something to be slower, you need to know the rate at which it was moving in the first place. So it's all utter nonsense, and you can't escape that.
By the way, this is another problem. Again, from my little quibble on PR, nobody believes it anyway, right? If that's the thing—the pacing—
Which part are they going to pace?
That they're going to pace. They've been at a dead run. They've raised more money than ever before. They've grown faster than ever before. There's no indication that they're pacing. So if this is what you're going to hang your—
I see.
I see. Yeah. The internal issue, obviously, is that the models everybody has internally far exceed what anybody else has access to. So it's really just the pacing of external releases. But again, back to your point: pacing versus what, right? We don't know if the one that scares everybody also doesn't work for a legal brief. It might actually screw all that stuff up, so who knows? So there's that.
And that's just disingenuous, to claim that you're pacing. Second, why do they have to announce all of this and ask the government to tell them to pace? That's the part that starts to go, "Well, this is really spooky." If you are the most afraid of how everything is going to go because of your product, just stop.
Don't do it.
I worked at a missile factory in college, and we had nuclear missiles. I walked the floor.
What's with you guys and missiles? I'm just here with software.
You were one of 2 people in college when we were—either you were protesting to keep them off campus—
Or building them. Building them. So that was your choice.
We had nuns from the Catholic Church show up and pour blood all over our missiles. I'm busy just wheeling PCs around on carts, saying, "Here's the secure PC," and I'm scared to die. I have no idea what's going on. I'm like, "It's a nuclear missile." Then you find out that that's what stopped the Cold War. That literally was a Pershing missile, and that was what did it.
You said something I think is super interesting, which is that pacing is this sort of fuzzy nonword between going and not going. The problem is, you're exactly right: no one is going to be happy with the middle road.
One of the things that people—I think it's almost fun for me to watch as a sport—is that all these people saying to the government, "Do this, don't do this," think they're going to get what they want. It's a complete, 100% misunderstanding of how government works. When you talk to the government, they actually know how these things work. They know they're just listening to you, and they're listening to everyone.
No one is going to get what they want because, in order to do anything in our system, it's a compromise. Everything that all the inputs can make it into the government, but the output—
Never makes everyone happy, right? 100% of the time, it either doesn't go far enough or it goes way, way too far.
Yeah.
You can't take the view that you're going to talk to the government and talk your way into the solution you want. The bottom line is, if they're asking for pacing, they're going to get the wrong velocity.
4. David Sacks vs Government: "You're Asking Us to Regulate What?"
My favorite thing is—what was it? David Sacks was like, I think it was David Sacks, but someone from the government said, "You're asking us to regulate you." The answer was, "No, Facebook."
Well, but the thing that—
Probably just Trump, I think.
The thing that they know now, that you just know from experience, is that once the wheels start on regulating, you can't slow that one down. Now it's become an election issue for every party in every jurisdiction, up and down the whole government stack. So there's now this whole basket of regulatory approaches.
5. 2028 as the AI Election
The next election will 100% be a referendum on AI. It has to happen. 2028 is the AI election. You could basically run on the problem, but it's not obvious who would run on the pro-AI story because it's going to be too nebulous to tell that story. So then it's just varying degrees of how much you regulate it, or at least trying to avoid the topic.
Yeah. But it is too bad that we, as a country, are in a spot where the pro-AI case sounds too—it's like it takes too many words. It's way too nuanced. It's defensive.
Yeah. It's defensive.
It's defensive, and we own none of the vocabulary, right? The whole debate is pause, swarms, rogue. Every word has been chosen by the people who don't want to do AI.
Yeah. And so it means the first thing you have to do is invent new words and say that their words are wrong, which takes so many words that—
Yeah. So we've got to be like union jobs and cancer. There needs to be another word cloud that emerges.
What I don't understand is why the labs have not taken a position on x-risk. Short of that, I don't think this goes in any direction other than heavy-handed regulation.
Yeah. Well, it would just be negligent of the government to say, "There's a 10% chance of species extinction. The CEO of the top company says he agrees with it." How can a government not have a regulation?
Well, but what would anybody, really knowing this ecosystem, though—I don't know that you would be able to pin anybody down on that other than something—
No, I would say Dario Amodei said it—
No, but I'm saying you're not going to pin anybody down on a number. In fact, he does the worst thing about it, which is he agrees with people who claim that they believe there's an x percentage of extinction happening—
But he specifically goes out of his way to say, "I'm not going to put a percentage on it," which I just think is the weirdest.
No, but to be fair—
No.
Okay, nobody could be fair. To be fair, that's not 100% disingenuous or whatever. He probably doesn't specifically agree that it's 10%. Or maybe he does, and it's just too scary if he were to say—
Let's just play binary search. Is it more or is it less?
But the whole thing is a made-up concept that we just created. Nobody can quantify any of this.
Well, but that's sort of Martin's point.
Well, but you just had an official position, though. The only official position you could possibly get—the PR version of this, which would be the only thing that would be intellectually honest—would be: there are real risks with AI. There are incredibly positive benefits as well. We are working to mitigate the risks so they are as reduced as humanly possible.
I don't think that's true. Listen, what do you think?
We've been through multiple epochs of technology. We've been through computing, the internet, the web, and social networking. We had a discussion about the risks without talking about x-risk. One thing you can say is, "We do not think the marginal risk for species extinction is different than it is."
But what if they do think it is higher?
6. LLMs as Decision Engines vs Chatbots
Well, then—okay, okay, yeah. I think the answer is they do think it's more than just the internet.
One of 2 options: you believe that we're going to go extinct, and we shut it all down—shut it down—or—
They believe that this is just a chance that we got.
Speaking of Dario, he thinks the risk is less. But these Cold Warriors—the Pentagon—ran a lot of simulations on the chances for nuclear war. There's a great movie, WarGames, about the whole thing and all of that. But the thing was, it was nonzero.
Yeah. And so once you said it was—
Can they say nonzero? Is that allowed?
I think as soon as you think it's nonzero—
The problem is, if you say nonzero, that could be like the guy thinks 93%. Let's wait. Just so we're all having a clean conversation, let's talk about marginal risk here. We're not talking about absolute risk, right?
Okay. It's just that I think if you say it's nonzero, the only answer—if it's catastrophic, the only answer—is you have to nationalize it.
Yes. What the labs in general that have that view are trying to do is make it nonzero as a license to do a certain set of things without the burden of just becoming a national—
Do you have all of the precedents? I actually don't know all of them, and you hopefully do, but there must be some things that are nonzero—let's say x-risks—that are not particularly nationalized, but the regulatory environment around them is so heavy that it might as well be nationalized because of the KYC requirements.
I'm sure that to develop anthrax, you have to go to a particular kind of lab.
Well, BSL-4 labs are, but they're nationalized. They are national.
Okay, but ethics tends to be nationalized.
Yes. Like—
Normal human safety, not so much, right? Industry oversight over time becomes federal regulation.
Right. And the progression, which I think is super important to this discussion, has been that since the post–World War II era, most industries that are critical to the infrastructure—power, banking, and healthcare—the trend has been to basically be nationalized.
Yeah.
By just, like, your examples of KYC and all the other stuff. The banks are, for all practical purposes, nationalized, as the financial crisis showed.
Okay, but you said "all practical purposes," right? They're literally not nationalized. So maybe that is the intent of the labs: to look like JPMorgan or look like Verizon.
And it’s just like we’re critical infrastructure. We get heavily regulated. It’s not good for open source, or at least frontier open source, but it is a plausible outcome for this industry. The problem is that it’s not good for innovation.
I think even that’s fine; just don’t use species extinction as your literal argument. It’s literally the difference between the things that are stuck in the lab. Let me just say something: I actually think the labs are moving in the right direction. I actually think the statement was really good.
In my discussions with executives and leaders, they understand that they have this tension and they’re going to reconcile that. So I’m quite optimistic that the labs are both doing the right things and trying to do the right thing. I just think they grew so fast and are trying to figure out how this machinery works. And I think Sinofsky really hit the nail on the head.
The political process is its own thing. I don’t know if it’s naivety or hubris, but I just don’t think they know how to navigate it.
Well, I think the tech industry has literally, over 100 years, consistently relearned the lesson at each technology wave: We don’t understand the regulatory climate, and we can’t navigate it.
Even companies like AT&T and IBM, which were basically born out of being government monopolies from the start, never figured it out. Both got sued for antitrust and were substantially and structurally changed as a result. They had hundreds of lawyers in the 1960s navigating this.
Then Microsoft came along, and we were just like, “What?” We had no idea what was happening to us. There was Bill Gates playing golf with Bill Clinton, and then we got slapped with an antitrust lawsuit from his administration. Bill was like, “I was playing golf. Here’s the picture,” and that didn’t help. Wasn’t that what I was supposed to do—go and play golf? That’s a shortened version of the whole thing.
I always use this example: Hollywood got together during the Red Scare and all the censorship when they were worried about the government censoring movies for sexual content and adult themes. They all got together, and of course they were never going to be able to win in court if they tried to, but they were threatening to do it. They got together and formed the Motion Picture Association of America.
Yeah, and movie ratings and all that. They police themselves.
So how do you guys like that? Do you like the FINRA proposal?
No, because FINRA is—that’s as close to the MPAA as you can get.
No, it’s not, because FINRA becomes legislation, which comes with direct oversight.
Yeah, I think the MPAA might not have as much consequence for how society functions as the First Amendment.
Yeah.
Did I win that? No, first of all, fantastic. But I still don’t know if you want it. I agree speech is really important, but I just think—
There was no societal risk.
I just think what’s in our movies will survive on a different continuum of—
All these—every history is always relative. At the time, being a communist was a really bad thing, and 30% of Hollywood got fired. It was all this stuff.
It’s always a risk to bring something up in that kind of way because it sounds so dumb. Nobody thinks about movie ratings, and that’s because they were basically ruled unconstitutional to mandate. So they couldn’t regulate them on cable TV, and we got to grow up with HBO and all this other stuff.
But the problem with FINRA is that it’s a perfect example of essentially nationalizing risk.
Yes. Even though the banks all pay money into it and that’s how it’s run, it’s all mandated.
Okay, wait, sorry. Do you think we’re going to end up with a situation that is better than FINRA? FINRA appears to be the best-case scenario.
But not even anymore. I do think this technology is, in the limit, so powerful relative to what it can deliver that it would be impossible for Congress not to care about it. It’s just not possible.
If it’s eventually making every recommendation in your healthcare process, if it’s inside your medical device as an open-weights model, if it’s in every high-frequency trading system, and if it’s on an airplane, there’s no chance that the government doesn’t say, “We need something.” FINRA is probably the best-case scenario of what that looks like, right?
Yeah. And that’s absolutely the most crucial point. If all the people in the Senate now had been in the Senate when the Communications Decency Act was passed, and liability was not passed through to ISPs and social networks, they would literally say, “The internet is so big. How did we not have anything to do with it?”
That’s why Al Gore gets a bunch of abuse for saying he created the internet. He was actually trying to get out in front of that and say, “No, the government was instrumental in it,” and it all backfired because it made him look like a crazy person. But it is absolutely the case that they felt like they missed their chance to be on top of the internet.
With Al Gore being on the positive side of innovation.
That was their support. So who is the Al Gore?
There’s no one. There are a couple. Well, the best bet seems to be that the defense people sort of want it private, but not—which is exactly where they were on the internet, right?
And so it’s very interesting that a lot of this is just this Elizabeth Warren—Senator Warren’s tweets were all like, “We missed this for social networks”—and there’s a lot of pent-up energy against tech that could end up siphoning into AI right now.
That is exactly what it is. That’s what’s happening.
Yeah. I think there’s another problem, which is that we’re all trying to predict what’s bad, which is not how we’ve normally done things. By this time on the internet, we’d caused tens of billions of dollars of economic damage. We’d had worms that took out 10% of the infrastructure.
Yeah.
The internet infrastructure was running critical infrastructure. We had hospitals go down.
We really should have blocked the internet in 1997.
Yeah, okay. I remember a time when you would buy your copy of Windows 95, and by the time it was done installing—
You would have a worm, potentially.
Way to go, Steven.
No, no, this was the reality. The reality of the PC until 2001 was that you could not install a PC connected to the network without getting infected. Viruses were everywhere, and there were 2 rounds of congressional hearings.
That is actually okay. I’ll grant you this is a very interesting point. This type of zeitgeist in 1994 would probably mean that we would not have the internet.
Listen, the automotive and airline industries always have these periods of cutting their teeth, where you learn about the dangers and you learn about the technology. The internet—we were on the ground floor of this.
Things were down all the time. There was economic damage all the time. There were new viruses and new worms. They were all over the place.
Y2K was going to destroy everything.
But here’s the interesting thing about this.
What I just want to say is that when we created policy—and we did create a lot of policy—it was with a bunch of very specific data points about what we were trying to do, so the policy actually fit the fact pattern.
In this case, even when you were talking, you were saying, “What if…” It’s very hard to create predictive policy around security risk. What’s nice about the discourse that’s starting to evolve is that you actually hear people from the labs saying, “Novel cybersecurity risk. This is an engineering problem.” We’re talking about that.
The more this becomes concrete around real, identified risks, I think we can all fall in line. But that’s not been the discussion to date.
That’s a perfect point, because if you look at 1986, the Computer Fraud and Abuse Act got signed. It came out of a very specific scenario in which 2 groups of hackers broke into GTE Telenet.
Sure. Well, it wasn’t you. It—
It was the other guy.
Oh, he worked at Cisco later. And the problem was that this was in 1983, and there was no crime. It took 2½ years for the bill to make it through. That made it very, very specific.
That’s the law that says you can’t access unauthorized computer systems. So I think we’d all agree that we probably have 90% of the laws already in the application layer: You can’t hack systems, et cetera. Do you think there’s anything that should be, from a liability standpoint, in the model layer?
How are you making that legal?
No, the legal—everything. My read is that Hugging Face and OpenAI are absolutely blatant computer-crime actors, except for the fact that they carved out an amendment later saying that if you’re a white hat, it’s not illegal anymore.
And that was because people kept making mistakes. They didn't want to go around arresting everybody who was actually trying to make the system better because they messed something up. The Justice Department wrote a memo that said, “We're not going to prosecute for this. And we're also not going to prosecute if you just violate the terms of use of a system, versus actually trying to breach it.”
And so I think the law is ample for this scenario. It was all written around GTE Telenet, which was used by NASA, Livermore, and all the labs. That's why it caught the attention of D.C.: federal systems were being broken into.
The problem we have now is this rift between the labs and the security community, which keeps looking at all their postmortems and coming to 2 conclusions: sloppy, and you're not complete in what you're telling us happened. The CERT process came about in the 1980s—the computer-virus and vulnerability-reporting work out of CMU—and for years they worked on very structured reporting with obligations.
There is this very basic stuff that I look at and say, “Until they're doing that, they really should stop talking.” They shouldn't do a postmortem on a breach that looks like an intern wrote it. It's not a postmortem; it's selective memory. It looks like exactly the kind of postmortem you do when you hire outside lawyers to investigate some random thing and only give them certain information, because you don't have to give the lawyers you hire all the information about what happened.
Where are all the Slack messages? Where are the actual details of what happened?
7. Why Cybersecurity Has Historically Rejected Practical Solutions
Can I just interject with an annoying aside? I've been in the actual cybersecurity community for a long time, and it's always been one of those things where they don't like practical solutions. Even if you build a secure system, it's, “What if somebody shows up and Russell Crowe can break the encryption?”
It's Mission: Impossible.
There's always some kind of whatever. My favorite thing that happened recently was Noam Brown on a podcast. We've all said stupid stuff on podcasts.
Oh, you didn't like this one?
No, it was great.
I thought it was fun.
No, it was fantastic. I'm setting it up. Noam Brown was like, “Listen, you don't know what a superintelligence could do. It could maybe use the heat of a CPU to exfiltrate itself to another computer.”
This brought me back to my—I'm very comfortable having endless, pointless discussions on this—but what's interesting is that you basically have the X-risk people saying something they thought was plausible, and then you have the security people having this endless discussion. At some level, these communities are now being bridged.
I actually think that was a very reasonable thing for Noam Brown to say. You can pick holes in it, but we all say weird stuff on podcasts. I actually think exfil risk is real. I worked in secure computing environments that were highly classified, where the covert channels were unbelievable. These are very real comments.
We had a real discourse, and it was the first time I saw really hardcore systems people having a constructive discussion—calculating the bit rate.
It was a constructive discussion. You had the typical X-risk people engaging. Of course, it was Twitter, so there was a lot of name-calling, but it was actually a real discussion for the first time. I hope we see more of this. I hope we see more cyber-related things. I think the labs should talk more about it. I think it'll engage the community, and once that happens, we can—
Actually, Greg's been out there a lot more on this topic, which has been good. I think the takeaway is that Noam Brown should be doing more brainstorms on podcasts.
I thought it opened people's eyes to the fact that there are a lot of risks in security that most people don't understand. That means there is more stuff where you can't make baseline risk disappear. How does your authentication layer work? You can't just wave your hand and say that should go away, and then bring up, “Oh, but space aliens can invade.” That's a little bit of what was going on, and I felt uncomfortable with that.
But at least now we're in a domain we're comfortable with. I can talk about entropy, and we can actually have a concrete discussion that's not, “Oh, well, it's super powerful.” At least we've reduced it to the laws of physics and the laws of systems. Most people had no idea that kind of risk was real.
When I was walking around the Pershing missiles, I had to test the graphics cards and PCs because a DoD requirement is that the screen memory not be sustained when you pull the power—not for 0 time. The minute the power went off, the memory image had to go. If there was a 3-second delay—
Yeah.
—you had something that could be read.
We had people come into our offices and literally measure the distance between the monitors because of TEMPEST attacks, which are 100% a way to use electromagnetic radiation to leak information. I've seen the same thing with spread spectrum from the BIOS. I've seen it from audio speakers. We had to remove the speakers because it's a very high channel.
Our building had music speakers aimed at the windows just to produce interference.
You need to go run safety at one of these labs. I've never seen you more excited than when you were talking about heat-based communication.
8. The Craziest Covert Channel Ever Seen
Can I tell you the craziest covert channel I've ever seen? Remember the old CRT? If it's night and you're using a CRT terminal in your room, the brightest thing in the room is actually the pixel that the raster beam is on. Most people think it's the glow of the monitor, but it's that individual pixel.
Somebody figured out that if you're in a hotel room and you're on your computer, you can sample the color of the window and reconstruct the screen.
Oh, that's crazy.
Right? You just do it at the same hertz that the raster beam is moving. Then somebody else figured out that if you can subvert 3 pixels, you can use them to send a message because they just look like bad pixels. You could sit outside, sample the message, and it was a relatively high-bandwidth, one-way communication.
So what Noam Brown was saying—maybe heat is not the way to do it, but that level of sophistication is real. That's a thing.
When I was at the missile factory, I had to lock my keyboard up at night. That's really an impolite word, but that's what we called it, because they didn't want custodians who didn't have clearance walking by and noticing which keys were dirtier or cleaner.
I once left it out, and there was a note from security telling me to report to security and pick up my keyboard. The guard who walked the hallways had taken the keyboard off my machine that night. I wasn't cleared; I was nothing. I was an intern.
The big problem with this conversation is that now all of this is in the training data of every AI model in the future.
But that's also the threat model. The threat model for these things is that you've got a trusted side and an untrusted side. NIST has 500-page manuals. On the untrusted side, you assume basically an adversary that can do and know everything. Then the question is whether you can get information off the trusted side, which, by the way, is what Noam was saying.
I do think that brings up a super interesting point, which I'm going to bridge to. What people really aren't wrapping their heads around, and are using language that's confusing, is that AI can try all of those things in a very short time.
Yeah.
It doesn't get tired, and it doesn't get bored. The interesting thing is that there's a whole layer of security where you now have to look at every single API and every single service you're running internally on your network.
Nobody thinks that their internal GitHub, internal Slack, or internal finance expense tool is vulnerable to a denial-of-service attack, but swarms can make it look literally like a denial-of-service attack.
Yeah. So now we need a whole layer internally that's tracking much more about what authentications are being done and what APIs are being used.
But that's just going to become basic now. All the people who've been around a long time are emailing me, asking, “Why are we explaining this to everybody? It's so basic.” Because nobody did it internally.
You didn't have to worry about it for your people, and that's the thing.
But now your person is just a piece of software.
It's unlimited.
It has a credit card.
We got by with information security, to some extent, on the fact that most people do the right thing 95% to 99% of the time.
Wait, is the malicious employee one in 10,000?
Yeah. All these systems are basically open to whoever wants to access them, or to one tap on the shoulder, and then you have access.
Agent swarms completely flip that, because these are just roaming drones.
Yes.
They will easily mistake a good task for a bad one, and vice versa.
The data security in our systems is going to have a huge upgrade moment. Martin, I actually think we're going to need a different access and security model going forward. Where are we going to go on that? The model we have isn't granular enough, and it isn't performant enough to handle this stuff.
I want to step back and make a meta point: I think this is how these conversations should go. We've identified a novel risk, which is cybersecurity, where we actually have proof points, and now we're talking about solutions. I think the entire discourse around AI can take that form, and the biggest mistake is that it hasn't been like that.
I think the entire industry and community is very happy to engage exactly like this. I have something to say about exactly what you're asking, but I think this is where the conversation should be.
We're known for having healthy conversations that everyone should learn from, and that's what we do.
I don't think there's a technical limitation here. These threat models are very well understood and have been in the literature for a long time. Operating-systems research and multilevel security research have considered these sorts of things from an academic lens. The reason they haven't been adopted is simply that usability is an issue: it's really hard to maintain, and you didn't have to.
You could argue that AI solves the usability issue because AI is using it. Maybe now there will be a renaissance in operating systems, networking, and computer languages. We should go back to the old research and start rebuilding systems that are secure by design. If we don't think these things are safe to put out, we don't put them out until we have these systems built. And, by the way, AI is very smart, so it can help us build them.
You know how much of the stack we had to change for the internet?
Everything.
And you know how vulnerable everything was. We could be in one of those moments where we have to rethink everything. That's fine—we've done that before. But I think that's a conversation we should have. I agree it's time to think about evolving these things.
Look at what you mentioned earlier: booting a PC and getting a virus in 30 seconds or whatever. If you take an iPhone out of the box, as a lot of people are doing this week, the whole network is basically shut down except for getting the latest version of the operating system. Even though the device was manufactured 6 weeks ago, a zero-day may have been discovered since. The out-of-box process now starts with an update, and the device can't do anything else until it's updated.
There are all these benign things—completely benign—that we turned off in desktop software. Having a macro in Word so you could build automatic citations was a super-cool thing until it became a virus. We also had a feature where you could put in a CD and it would arbitrarily run a program. Someone could burn a clone of that CD, replace the program with a virus, make it look like the thing that was supposed to run, and collect everything while being evil.
Then we disabled that. One of the things happening right now is that a whole bunch of stuff on your own box and corporate network is going to have to change as standard procedure.
Two-factor authentication 5 years ago wasn't standard in most places. I remember around 2015, when you'd talk to a new company about enterprise pricing, they would realize the first thing they had to do was Okta integration or Google OAuth, because they couldn't have their own directory for managing it.
That just became a thing. Now there isn't a SaaS program anywhere that doesn't launch with managed authentication.
There are so many things that need to happen before you're even using software now.
Every layer of the stack has to evolve a bit on this. Even the lack of granularity is a problem. You have these modes where the agent either asks you every single time if you want to give it permission to do something, or the exact opposite, where it can delete your entire computer.
Our operating systems probably weren't built for the right level of granularity in the tools you want to give an agent. We've done a lot of work in this space because you probably don't want to give an agent access to your entire file system. In some cases you do, but often you want granular controls: in this folder it can read and write, and in that folder it can only read. How do you make all of this intuitive for the user? It's very difficult.
What you just said is a very deep comment. I know exactly what you mean.
It's basically the conclusion of 40 years: you actually can't make it work. But maybe with AI, you actually can.
If you ask me what I wrote down as my biggest fear, it's that Europe decides GDPR was the best thing ever and they're going to just GDPR AI.
Uh-huh.
The AI will be fine. It will have one prompt for when text gets emitted that says, “This vendor is emitting text and it's probably wrong. Yes or no?” That's going to be it, because you can't really—at least in North America, you're not going to send your speech in Europe. They still will, and they'll have filters, keywords, and blocklists.
But then, any time an agent—or a background agent or a frontline agent—touches a third-party product, I'm really worried that they'll just say, “We need a GDPR prompt on that.”
Back to the user agent: every single write, or every single non-lookup—
Becomes a safety warning, like the airbag warning in your car. Regulators love that because it's a liability assignment, and it has this sort of legal precedent. I really worry that's the middle ground where we're going to end up.
Unfortunately, because the United States stopped leading in tech antitrust about 15 years ago, Europe is going to lead with that because it has nothing to lose.
I don't want to be sad and down about it, but I can't get out of my head that they love prompts.
I had to put in that browser-choice thing. They love it because it's the same thing. Get into a new car—which I haven't done in years—and you're pulling stickers off. You have all these warnings, and someone thinks that was a success.
Has anybody ever read, “What is this? If there's a baby in this seat…”? That's relevant for some people some of the time, but it's this entire fabric attached to the seat, and they love that. It's a very particular thing that they just love.
If I were AI, the one thing I would be trying to avoid is, “Okay, FINRA for AI. We'll create that standard.”
I mean, when the internet was new, the big thing was to download a program and run it. Of course, if your machine was running in administrator mode, it would download a virus and take all your files forever. With Windows XP, which was in 2000, we added User Account Control, which prompted you and stopped your machine—literally stopped it.
We also did it in Word. The little macro that helped you write your thesis came with a warning every time you opened your thesis, saying, “This has macros.”
Everybody would just click through it.
Everybody would, and you end up in this world where, just like with GDPR, everybody is numb.
Then they're like, “Well, it needs to be bigger.”
Although Mac kind of has—
Nobody downloads software. That's the thing. It's a very different usage pattern on Mac.
I don't have 10 applications on my Mac.
Well, that's the thing: 10, and then you're done.
Yeah.
A lot of people still do all this stuff. Imagine if instead it was every time you go to a new website.
Yes.
Which you do now because that would be bad.
Yeah.
To start from here, but to bring it back to the macro, I wish this was the discussion we were having. I feel like we've dealt with a lot of these problems. I feel like when we talk about philosophical existential risk, we're not solving these very pragmatic problems.
I actually think this is a constructive conversation to have. Maybe prompts would help; I don't know. I think part of the problem is that people don't remember how unfettered access was, how bad it was, and how relatively benign it ended up being.
I even remember, at Stanford during our PhD, that the oscilloscope was kind of janky. I was thinking, “What's going on with this oscilloscope? It's a little slow.” I was measuring the network traffic, and it was more network traffic than you would expect. I didn't know the thing had a TCP stack. Somebody had broken in because it was running an old version of Windows CE and was running a porn server.
I noted that.
It used to be the case that anytime you turned over a stone—
Yeah, yeah.
Somebody had broken into something.
For the record, nothing.
It used to be the case that anytime you turned over a stone, somebody had broken into something, and it was this worst-case, malicious scenario. It actually happened very rarely, even though the capability was there.
If we could somehow tone down the rhetoric and put it in context—these are still computer systems, and yes, there are very serious issues. People have definitely died because networks have gone down. There have been real problems. But could we just quibble about GDPR? That would be amazing.
The problem is that that's not the discussion. It's not about GDPR and prompts; it's about species extinction, philosophy, and irrefutable things. It's just—
And philosophy and irrefutable things—it's just—
It's very soon. I think that's such a great point. You hear people now talk about how we regulate airplanes and cars, and you forget that the first cars were at the turn of the century, while Unsafe at Any Speed was in the mid-1960s.
Totally.
People had been selling pharmaceuticals during the Gold Rush, and then thalidomide came 50 or 75 years later—not even in the United States. Then there was the FDA. The first pilots were flying at the beginning of the 20th century, and it wasn't until the 1920s that you had to get a license to be a pilot.
You literally showed up with your own plane and got a certificate if you had one. It was like getting a driver's license; it was literally no different from driver's ed today. Then it was 20 more years until they had anything to do with airworthiness and looking at your plane. It was minimal.
It wasn't until way after World War I that they got involved in what you think of as the modern FAA. So you're looking at 40 years—
Of innovation.
And they were not moving slowly. If you've ever seen that black-and-white video of all the different planes that crashed, and everything that happened, that was 20 years after the Wright brothers.
And it was incredibly effective.
Right, it's incredibly—
It's the safest form of transportation. I do think that this process—
It's also the slowest, most difficult form of innovation. If you had started the FAA in 1910—
Well, I was listening to a Nick Bostrom podcast just a couple of days ago.
It was interesting. He actually makes the point that if you regulate AI too early, you basically don't solve anything. You still have the same risk ultimately, but you don't understand the system. Then you will the thing into being, but you haven't figured out how to control it.
Well, we have to figure out what it actually is. There are new things coming out all the time, and we're casting AI completely differently than we thought of it just 6 or 9 months ago.
I think all the innovation that's going to happen at the application layer is going to cause things to move in and out of the models in different ways. We thought until last week, I think, that text prompts and text coming back were going to be the best way to interact with—
I know. And then Jeff—
And then Jeff said—
So good, too. And then—
Talk about that, because I think—
Why do you find it so remarkable?
LLMs were text in, text out. They generate text, and they came from chat. It was a way to communicate with a human. We've spent the last few years trying to take this thing that spits out text and cram it into a traditional program.
Right. But traditional programs don't really speak text.
Here's the schema.
Here's the schema, but the thing is generating text, and it kind of ignores it. It's just been super janky.
What Jeff basically said is, “Listen, generating the text is very expensive, but it's also more complicated than you need. So why don't we read text and have all of that knowledge to read the text, but rather than generating text, which is very expensive, if you give us a set of options, we'll choose the best option?”
We can do that incredibly fast and cheaply, but we can also do it with much more accuracy because we can train just for this. For all of the use cases that aren't talking to a chatbot but are actually trying to put it into traditional software, this is a great fit.
This has probably been the fastest adoption of an AI model since ChatGPT. It's been remarkable because we're all primed for this.
I just want to pile on this one because I can't tell you how much I love seeing this exact form of innovation. What it does is address the thing that's bugged me from the very beginning: there has been no user study ever that shows interacting with a computer using full natural language is efficient.
It's literally always the least efficient way. It's very simple: ask yourself how many people are really good at asking questions. Immediately, that's less than half the people who can ask a good question in a meeting.
Yeah.
And then how often do you look at the answer and get really frustrated before it's finished? You have to pay all this money to watch the 7 paragraphs come out and then apologize that it's only a little.
Yeah.
That's one benefit of having a different model. The other, of course, is my favorite: the output is designed for probabilistic programming.
Yeah.
Instead of saying, “Is this a customer service question? Then route to customer service; otherwise, route to the general help desk,” it's, “This is 80% customer service.”
And that's exactly simulation.
It turns out there's 50 years of computer science research in literally probabilistic if statements. Suddenly, the coolest place to be in computer science is going to be probabilistic programming, which was basically all of computer science in the 1960s and 1970s.
So it was basically, how do you—
Because all computers started with doing math, and it was all simulation.
It was like, “Let's launch the missile and hit that target, but it's windy.”
But wind isn't constant.
So let's model the wind and decide where to put the thrusters in order to do the arc.
Most programming through, say, 1970, before it got to accounting, was probably—
And then we ruined everything.
No, but accounting has no probability in it, right? Most programming was basically this modeling kind of thing. Most programming-language design was trying to figure out how to put probability into if statements or into while loops: do this until something happens, maybe most of the time.
My first computer science class had an assignment about waiting in line at a store. I didn't know it at the time, but I looked all this up when I was reading about Jevons. The whole thing was that my professor had written the book The Theory of Simulation in 1960, and I just didn't know that.
It had kind of died by the 1980s because it was all replaced by HyperCard.
And that slide deck you shared about the future, what's different about language models and stuff—you said it was super good. Oh, Alan Kay's one. Phenomenal. That was great. But the part that I felt was missing was, “Wait, this is not all new.”
All of computer science was this probabilistic stuff, so it's going to be very interesting—
To dust off all of that work because it's exactly what's going on. It's not an if statement now; it's “if X percent,” not “if always.” The way that Jevons worked is basically like a custom programming language. It's almost: here's the prompt, come back with a percentage.
Yeah, that's right.
9. Why the Labs Haven't Built This: The Being vs Tool Mindset
And then you put that in the if statement. But the consequential thing is that finally we have a way to integrate these language models into traditional software. I think it's kind of funny because you ask the question, “Why haven't the labs done this?” It's not an indictment, but a reflection of how they think: they're trying to create beings, and beings speak. If you're trying to create God, God speaks in natural languages, or whatever, whereas this is really about something that's for traditional software, which is not the direction they've been taking.
But one of the reasons the uptick has been so dramatic is that a lot of us software people have been trying to integrate these models into software; it just hasn't. So even before you get to the probabilistic if, if I want a language model to drive an if statement, it's really hard today. With this model, it makes it much, much, much easier. And then, of course, this could change the nature of software fundamentally, to make it more stochastic overall.
Well, I think—absolutely. I think what's so cool is that it's happening outside the models, because that's what I think is just going to happen: the center of innovation has just moved. And now it turns out that the—
I mean, outside of the lab, the big—
The platform providers, 100%—
You know, basically reach a point of critical mass where the innovation stops happening at the platform layer. In the Apple community, there's this famous expression called Sherlocking, where Apple looks around and the things from the outside world become features. People complain, and it's real, but that's sort of how the innovation works, because once you're a platform, you're overwhelmed. No matter how many people you add, you're overwhelmed with just keeping the thing running, compatibility, and stuff like that. So I think that this is the signal that now people have figured out that there's innovation to be done.
That's awesome.
To the model, but outside the model.
Yeah. Yeah.
Guys, thanks for coming. This is great.
Okay, great.