[BidClub_]
Machine Learning Street Talk · · 77 分钟

智能是集体性的,而非人工的——Michael I. Jordan 教授(UC Berkeley / Inria)

Tim ScarfeMichael I. Jordan

YouTube
TL;DR
  • Jordan 的可投资核心判断是,持久的 AI 价值将来自协调人类、模型、数据所有者与激励机制,而不是打造一个单一的“超级智能”。 数十亿人已经在生产底层数据,最终服务也面向数十亿消费者;经济学提供了在这张潜在网络中分配价值所缺失的机制。“如果对数据做梯度下降时,没有深度的微观经济学视角相伴”,规模化 AI 就会错误定价隐私、劳动、真实性与参与。

  • 最强的机会在于垂直领域系统:将强大的预测能力、新鲜的真实数据、不确定性估计与人类控制结合起来。 AlphaFold 预测的2亿个蛋白质结构支持了高统计功效的假设检验,而规模小得多的实验数据做不到;但针对所提具体问题,它给出了一个极窄、却“远远偏离真值”的置信区间。Jordan 提出的 prediction-powered inference 通过加入少量真实数据修复这一失败,提供了一种可投资的科学工具架构,而不是声称模型“理解”了什么。

  • 单纯扩大数据规模并不能替代市场设计,因为最有价值的知识往往是本地化、情境化的,而且经常会被有意隐藏。 Jordan 预计数据仍将具有竞争价值,参与者会要求获得报酬、隐私与控制权,而不是把数据直接交出去;因此平台必须诱导准确的信息,并让激励机制保持一致。“数据是自下而上来的”,未来的偏好仍会过于短暂,无法由一个中心系统编码成普适的人类价值函数。

  • 始终在线的 AI 助手,可能不如改善医疗、交通、金融、科学与市场的基础设施更适合作为商业模式。 Jordan 将“坐在你肩上的秘书”称为“愚蠢的商业模式”,认为许多用户最终会把它关掉;相比之下,现有统计系统每天为1亿人协调数十亿种商品,才是真正具有深远影响的系统。他的判断标准不是一个盒子听起来是否聪明,而是周边生态能否降低不确定性、创造可靠服务并产生就业。

  • 只要周边系统能让黑箱模型的行为变得可预测、可质疑且可执行,黑箱本身就可以接受。 被拒绝贷款的人不需要参观内部神经回路;Jordan 会展示大约50名可比申请人、其中哪些人获贷,以及申请人与他们的差异。“我不认为有必要理解所有细节”,但运营者必须理解输入、输出、约束、错误率,以及失败会带来的后果。

  • 平台经济学将决定生成式 AI 是扩大人类创意市场,还是掏空这些市场。 Jordan 认为 Spotify 可能接近垄断,艺术家获得的报酬很低,因此合成内容在经济上更具吸引力;他为 UnitedMasters 提供建议,后者让音乐人保留作品,并与品牌建立连接。他同样认为 Google 通过广告中介 YouTube 生产者与消费者关系是“一个巨大的错误”:系统确实创造了市场,却没有把足够多的价值直接导向创作者。

  • 近期风险在于劳动与资本的集中,以及设计糟糕的自主系统,而不是某个递归自我改进的实体选择征服人类。 Jordan 称 AGI 是“公关术语”,认为“超级智能—灭绝”的二元叙事让年轻建设者“极其泄气”。后续 David Krakauer 环节将自动驾驶仪视为人机协作的例子,而 Jordan 仍看好积极、以人为尺度的 AI 系统。

摘要 · 为研究而整理的核心内容

1. “AI”重新包装了更长的机器学习传统

  • Jordan 认为自己是一名统计学家和认知科学家,“从来没真正把自己当作 AI 研究者”。1950年代的 AI 计划强调逻辑推理等目标,但这些目标“并没有真正实现”。

  • 真正带来产业价值的方法——决策树、近邻法、逻辑回归、隐马尔可夫模型和基于梯度的方法——主要在1960年代至1980年代从统计学与运筹学中发展出来。供应链、商业和交通早已在使用机器学习;Amazon 的云服务最初就是为处理机器学习工作负载而发展起来的。

  • 大约在这次对话发生前5年,“AI”这个标签重新流行起来,因为模型开始输出流畅语言,而不再只是预测库存或价格。在图灵测试这类狭义定义下,Jordan 承认确实发生了变化;但他反对的是,这个复兴的标签扭曲了研究优先级和商业思维。

  • “对我来说,AGI 只是——一个公关术语。”Jordan 说,它带有警报式和狂热式的两种变体,让20岁、25岁的年轻人仿佛只剩两种姿态;但技术还有许多不那么戏剧化的路径,可以改善家庭、机构与经济。

2. 智能属于社会系统,而非孤立的盒子

  • Jordan 的集体主义前提是,人类智能来自意见汇总、继承而来的文化与具体情境:“在一个情境中聪明的行动,换到另一个情境就未必聪明。”智能无法与那些彼此合作、竞争、发出信号、甚至相互利用的人彻底分开。

  • 从物质基础看,这项技术本身已经是集体性的。它的输入来自数十亿人,预期用户也以数十亿计,尚未解决的问题是:这张潜在网络如何分配权力、金钱、隐私与机会。

  • Jordan 被经济学吸引,是因为它能把社会互动转化为“可执行的数学思想”。参与者不必暴露全部动机;机制可以让他们检验意图、传递信号,并在信息不对称的情况下安全合作。

  • 这不是对 AI 的外部批评:“我想把它做对,也想把它做得更好。”Jordan 希望将机器学习的预测能力嵌入正式系统,尊重人类既是生产者也是消费者,而不是把人类当成免费的数据来源。

3. AI 助手竞赛把输出误当成系统

  • 当被问及为什么那些令人印象深刻的文本、编程和解题系统没有像预期那样带来帮助时,Jordan 说这“完全不奇怪”。旧有前提——造出一个智能的东西,收益自然会出现——如今只是变成了“坐在你肩上的秘书,帮助你、在你耳边低语”。

  • 搜索是重大进步,但一个侵入式、永远在线的伴侣“就是一种愚蠢的商业模式”:人们希望自己思考,也可能更喜欢偶尔收到一份摘要。Jordan 预计许多用户会直接“把这破玩意儿关掉”。

  • 医疗、交通和金融本来就是数十亿参与者之间的数据流,合作与竞争始终嵌在其中。这些系统需要更好的市场与决策流程,而不只是几个扮演升级版搜索引擎的大模型。

  • “我希望这东西能创造就业。”Jordan 关注的分析单位是生态系统——谁在互动、互动频率如何、质量怎样、价值在哪里产生——而不是一个把输入映射到输出、却忽略劳动与所有权的统计盒子。

4. 黑箱需要行为契约,而不是拟人化叙事

  • 主持人的行为主义担忧以那只母鸡为例:它根据过去的安全经验不断外推,直到脖子被扭断;同时,机制可解释性研究则在寻找模型内部的推理回路。Jordan 的态度“没那么悲观”:只要围绕系统建立结构,工程师可以使用自己并未完全理解的系统。

  • 化学工程师历来会利用尚未完全理解的现象,同时摸清它们的输入、约束与可观测行为。Jordan 反对的不是不透明本身,而是依赖“AI 安全”这类口号,却不明确周边系统所需的行为、透明度和约束。

  • 经济学提供了一个实用范式:人们无法解释自己的每一个选择,但其行为仍足够可预测,足以让他人据此规划并交换价值。有效系统需要的是相关的规律性,而不是对每个参与者进行完整的机械式解释。

  • 如果模型拒绝某人的贷款,展示内部电路图并不能构成解释。Jordan 会展示模型嵌入空间中大约50名相近的人、哪些人获贷、哪些人未获贷,以及其中的关键差异——这些信息申请人能够理解并据此采取行动。

5. AlphaFold 说明,不针对具体问题校准的规模会造成误导

  • Jordan 称赞 AlphaFold 是一个有明确目标的系统,而不是类似 LLM 的通用产物。它大约2亿个预测蛋白质结构,可以支持仅凭讨论中提到的大约20万个实验结构无法完成的分析。

  • 他的团队利用一个2×2统计表,检验蛋白质中的量子涨落是否与磷酸化相关。仅凭实验结构,统计功效不足以拒绝“无关联”假设;2亿个预测结构则提供了足够的功效来拒绝该假设。

  • 但最终得到的置信区间极窄,而且相对于金标准真值“远远偏离”。可能的原因是:针对具有这些涨落的蛋白质,训练数据覆盖十分稀疏;同时,AlphaFold 并没有为一个它原本不是为之设计的问题提供量身定制的误差线。

  • Prediction-powered inference 只需向庞大的预测数据集加入少量真实数据,就能保留相对较高的统计功效,同时让置信区间转而覆盖真值。科学家工作的正是“知识的边缘”,也正是在那里,基础模型可能存在最大偏差,因此每个此类模型都需要一条本地验证路径。

6. 预测与实验比“理解”更重要

  • 主持人问,AlphaFold 的迭代回收与精炼是否可以算作一种理解过程。Jordan 的回答很直接:“AlphaFold 为什么需要理解?”它负责预测,也可能帮助人类实现控制;人类可以在模型产物上做实验,再形成自己的理解。

  • 2000年前后,Jordan 看到 Amazon 使用当时的方法——访谈中结合随机森林进行了描述——来预测供应链中船舶延误等中断事件。没有人能够理解每天将数十亿种商品运送给1亿人的完整系统,但系统的预测降低了不确定性,并支持了规划。

  • 当人们追问这张网络是否“理解运输和物流”时,Jordan 的回应是“谁在乎”。拟人化语言为媒体提供戏剧性,却没有改善工程系统的优化、规划或约束。

  • Fosbury flop 是他对宏大认知理论的反例:一名运动员尝试向后跳,结果奏效,其他人便跟进采用。工业化 A/B 测试同样把创造力与经验筛选结合起来;工程师首先应该问的是:“你想实现什么?”是取代教师、增强教师,还是让医生变得更强?

7. 药品监管是一个带有策略性数据的统计问题

  • 药物开发形成了一张纠缠的网络,参与者包括科学家、相互竞争的制药公司、蛋白质、患者和监管机构。监管机构希望整个系统保持较低的假阳性率和假阴性率,但药企提交的材料并不是来自中立来源、相互独立同分布的观测。

  • 公司拥有私人证据,也有混合动机:既想帮助患者,也想赚钱。面向10亿人市场的药物,一旦监管环节出现假阳性,可能带来极其丰厚的利润;因此,即便无差别提交会损害系统整体的错误控制,提交策略仍可能是理性的。

  • 主持人用航空公司的类比把机制讲得更具体。航空公司无法观察每名乘客的紧迫程度或支付意愿,因此会提供不同服务与价格,去包围这些隐藏偏好;乘客的选择会透露足够信息,让系统无需强迫任何人的行为也能运转。

  • 监管同样需要设计激励,让企业提交那些经过测试、且企业认为有前景的化合物。随着本地数据获得竞争价值,参与者必须被诱导去分享真实、有用的信息,而不只是分享带有对抗性的噪声。

8. 数据市场是均衡问题,而非优化问题

  • Jordan 的3层模型包含用户、服务平台和第三方数据买家。支付平台可以利用交易数据改善服务,但可能还需要额外收入;Mastercard 等公司就会向市场研究机构出售行为信息。

  • 加入买家会改变均衡,因为用户将失去隐私。平台可以提供不同级别的差分隐私与相应价格;重视隐私的用户可能选择更强的保护,从而为该平台带来更多数据,但新增噪声也会降低数据对买家的价值。

  • 当主持人把这个过程描述为一个需要模拟的动态系统时,Jordan 纠正说,在这个模型中,方程会直接给出 Stackelberg 均衡。监管机构可以比较不同均衡,包括它们的效用总和或社会福利,并检验最低隐私要求或异质化隐私要求。

  • 机器学习擅长优化,而经济学提供不动点、Stackelberg 博弈、Pareto 前沿、群体效应与策略适应。“未来必须是这些分支走到一起”:数据可以放松经济学中过于僵化的理性假设,但没有经济结构的数据仍然可能“把事情搞得一团糟”。

9. 情境知识击穿自上而下的价值函数

  • Jordan 在哥本哈根街头行走时看到,价格、商品、需求和购买意愿每时每刻都在变化。即使拥有数艾字节的数据,也不可能提前10秒确定每一个选择;社会知识“非常短暂,也非常即时”。

  • 因此,一个安全的市场应当承认无知,同时让自下而上的偏好在参与者愿意表达时显现出来。它不需要“站在顶端的上帝形象”,从假想的全知视角为每个人的生活定义一个人类价值函数。

  • 主持人指出,市场早于资本主义存在;Jordan 回应称,资本主义只是让市场运行起来的一种方法论,而不是市场的定义。Jordan 认为,人类文化与个体会创造抽象概念;有用的抽象可以被传播并提升为文化,而未来系统应允许新的抽象自发出现,而不是从上到下安装一套抽象。

10. 创作者平台暴露出 AI 尚未解决的价值分配问题

  • 主持人指出,Spotify 可能有动力用 AI 生成歌曲;Jordan 表示同意。他认为艺术家获得的报酬“非常、非常少”,而 Spotify 可能接近垄断,其价格似乎也不是通过竞争性机制形成的。

  • Jordan 为 UnitedMasters 提供建议。这是一个让音乐人保留作品,并与品牌及其他机会建立连接的替代平台。目标是让一个人能够成为“真正的艺术家”,而不只是每次歌曲被播放时收到一小笔钱。

  • YouTube 是一次错失的转折点:收购 YouTube 让 Google 真正拥有了生产者与消费者之间的市场,而不只是网页链接。Jordan 认为,观众本可以与创作者建立直接的经济联系,由可识别的受众基础提供激励。

  • 但 Google 通过广告导流价值,保留大部分收入,只返还一小部分;“Facebook 让情况变得更糟”。Jordan 将这一模式与 Amazon 具体的包裹配送业务作对比,同时承认 Amazon 仍然带来真实的劳动力市场问题。

11. AGI 宿命论掩盖了劳动、制度与有用的自主性

  • 面对讨论中与 Geoffrey Hinton、Stuart Russell 相关的递归自我改进、具备代理能力的超级智能论点,Jordan 的回答是:“非常科幻。”更深层的伤害在于告诉年轻建设者,上一代已经完成了算法,他们剩下的只有灭绝或即将到来的超级智能。

  • “这太让人泄气了。太让人泄气了。”Jordan 更担心劳动与资本的关系,而不是机器接管世界:规模化模型会强化垂直领域工具,让数学家工作更快,但他不认为模型会简单地让数学家失业。

  • 在随后 David Krakauer 的环节中,人类被描述为能够创造美、爱与艺术,但也会误解意图、发动战争、运行失灵的政治制度。Krakauer 对积极 AI 的设想,是改善信息流、信号传递和决策支持,让人们避免在不确定性下做出有害选择。

  • 主持人坚持认为自主软件确实带来安全风险;David Krakauer 回答“是,也不是”,并指出自动驾驶仪可以减少事故。即便主持人认为飞行已经被充分定义,Krakauer 仍列举了飞机、天气和人为错误:更有建设性的模式是系统层面的人机协作,而不是“把一个超级智能放到驾驶位上”。

12. 不确定性只有嵌入激励与群体后才会变得有用

  • Jordan 把博弈论比作 F=ma:先明确一个博弈,再计算 Nash、相关、序贯或 Stackelberg 均衡,并检验它是否能够预测行为。它在工程上的反向形式是机制设计——从期望实现的分配、分配公平规则或市场出发,再设计能够实现目标的博弈。

  • 在 Jordan 将保序预测归功于其教授 Vladimir Vovk 后,Peter Grünwald 解释了 e-values,并将统计契约理论与证据联系起来:只有当统计对象是 e-value 时,激励相容性才成立。

  • 与反复查看 p-value 和 p-hacking 相比,非负超鞅和 Ville 不等式支持 anytime inference:研究人员可以持续收集数据、反复检查证据,并自适应地停止,同时保留数学上的控制。

  • Jordan 的鸭子例子把统计学与群体情境连接起来:当谷物以2:1的比例出现时,贝叶斯鸭会以概率1前往更好的那一侧;而真实鸭子大致按2/3对1/3分布,这才是同时利用两种资源的 Nash 均衡。

  • 另一类不确定性来自私人信息与数据溯源:10年前收集的医学证据,其不确定性应当根据证据年龄进行调整,但当代模型很少以定量方式携带这类元数据。“可怜的 LLM 什么都没有”;当被要求给出置信度时,它主要是在模仿人类过去如何回答这一问题。

  • Jordan 最后举了披萨餐厅的例子:餐厅可以放心使用番茄,因为市场会奖励其他人去发现番茄并把它们供应过来。“市场缓解不确定性”:分布式探索能够稳定资源,让每个参与者都可以在其他人完成的工作之上构建自己的能力。

Tim Scarfe

Nature said that you are the most influential computer scientist.

Michael Jordan

It exists in the real world. This is an abstraction, but it’s a real thing. It’s like F = ma. It’s a set of equations; it’ll make predictions. So if I write down a game, just like I wrote down F = ma in some coordinate system, I can now predict what’ll happen.

I don’t think we need to say it understands. I think this anthropomorphizing of intelligence and understanding is not necessary or appropriate; it’s a distraction for many problems. Why say it understands? I think it’s science fiction, and I think science fiction is important for society, but at the level it’s being promoted and by those kinds of voices, it’s really hurting 25- and 20-year-olds.

These young folks, of whom there are huge numbers, are excited about technology, and they want to build things that help their family and help their country—actually, more their family than their country, honestly. They see real opportunities in doing that, and they’re being told by the leaders, “Well, we had our fun and we developed a bunch of algorithms. We did it, and we were just interested in the pure understanding of intelligence, even though they didn’t understand intelligence. They built gradient-descent algorithms, and now you guys can’t do this because it’s dangerous. It’s going to wipe out humanity with a high probability, or superintelligence will arrive soon, so there’s nothing left to do. That’s in your lifetime.”

That is so demoralizing. So demoralizing. That’s the thing I think bothers me the most. The second part that bothers me is that there’s no economic thinking going on there.

The current generation is just way too—there’s not much thought going on, not much intellectual stuff. It’s possible to build it. It’s possible to steal the data from wherever you want to, because that’s what the internet allowed to happen, and not return any value to the person who originated the data. It’s possible to run gradient descent on that, but you need huge amounts of money. But it’s now possible to get it from people who aren’t thinking very deeply.

I don’t think it’s bad to build systems you don’t understand. But I think this level of detachment from reality is unusual in human history.

This episode is supported by Cyber Fund. If you're building at the frontier of AI, they want to hear from you. Cyber Fund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund.

Tim Scarfe

And what do you think about the term AGI, by the way?

Michael Jordan

AGI, to me, is just a bit of a PR term. Some people think it’s fun because you have to have these great aspirations. I think it’s just distortion. I think it confuses young people.

As I’ll talk about a little bit today, one of the things I find most alarming about the so-called thought leaders one often sees on podcasts and other venues is the alarmist tone or the exuberant tone. I think 20- and 25-year-olds are watching that and saying, “Am I going to be exuberant, or am I going to be alarmist?” Those are the two choices.

I hope that this conversation we’re about to have makes it clear to young people that there are other ways to approach life and technology. I’ve never actually thought of myself as an AI researcher. I didn’t read an AI book.

The term was coined in the 1950s, and John McCarthy and others had particular goals in mind for coining it. They had particular methods in mind, like logical inference and so on, that didn’t really pan out.

In the meantime, in the 1960s, 1970s, and 1980s, something called machine learning arose. The actual methods, like decision trees, nearest neighbor, logistic regression, and hidden Markov models, were developed in other literatures, mostly statistics, operations research, and so on, and that led to industrial success stories.

Supply chains, commerce, and transportation systems all used, and still to this day use, vast amounts of machine learning. They used gradient-based methods, and the cloud was developed to handle machine-learning workloads at Amazon, in fact. That’s the tradition I came up in. I was trying to think about systems building at scale that would also serve multiple people.

The AI buzzword returned, I think, maybe 5 years ago, because the data that started to be used was language data. The box now is not just making predictions about supply chains, commerce, prices, or whatever; it spits out human-fluent language, and people said, “Oh my God, we’ve solved the old AI problem.”

In fact, in some ways, if you define the AI problem narrowly, like the Turing test, yeah. But there was this ongoing tradition of machine learning, and by that time it had incorporated people from all different kinds of fields and was really having an impact in industry. It still is.

The AI buzzword returned because of loss, and now, to my view, it’s had a distortionary effect on the path of research, on how we think about where research should go, but also on how we think about business models and where technology is going.

AI wasn’t enough; they had to create this big, hyped-up buzzword, AGI, which we’ll talk a lot about. Economics is a source of intelligence—social intelligence—and when it’s put together with machine-learning-style intelligence, you can now talk about scale, not just in terms of the numbers of computers and amount of data, but in terms of the numbers of humans.

That’s critically important to me: The role of humans as producers and consumers in these emerging systems should be respected, amplified, and thought about.

Tim Scarfe

Professor Michael Jordan, it’s such an honor to have you on MLST, especially given that Nature said you were the most influential computer scientist a little while back.

Michael Jordan

It’s funny because I was trained as a statistician and a cognitive scientist, but I’ll take it.

Tim Scarfe

Amazing stuff. Well, Michael, you’ve just published a paper called A Collectivist Economic Perspective on AI. Give us the elevator pitch.

Michael Jordan

I was never an AI person. In some ways, it’s easy for me to come in and look at people who are self-professed AI researchers and say, “What are you doing? What’s your point? What’s your goal?”

Sadly, they often don’t have a very clear goal. It’s that humans are intelligent, humans are a computer, the brain is a computer, and if we mimic that, take aspects of it, parallelize it, and make it more powerful, it’ll just do great things. It kind of stops there.

It’s not that there’s a goal in society that we’re going to try to do this or that. It’ll just solve problems for us, and then we’ll be happy. I got away from Silicon Valley partly because that’s just the way people talk, and I got tired of it. There’s not a lot of deeper, long-term thought going on. It became a rat race and a money race and all that.

My perspective comes from a long tradition of other people having social-science perspectives on intelligence. We are social animals, and a lot of our intelligence comes from the fact that we aggregate opinions and thoughts. We have cultures and so on that retain them.

Moreover, society provides a context for our intelligence. Smart action in one context is not smart action in another context, and it’s all very fleeting and contextual in the moment. Social-science ideas are needed to appreciate what that means.

When I say social science, I include economics. Game theory says that in context X, somebody else is trying to take advantage of me, or maybe to collaborate with me, and I don’t really know. I’ve got to put out feelers and send signals and create mechanisms where we can interact effectively. Economics studies that in a mathematical way, and that attracts me because I’m a mathematically inclined person.

I’m not a critic of AI. I want to make it right, and I want to make it better and understand what it means to be intelligent in this world, and safe and interesting, and think about long-term issues. To me, you have to do that formally or mathematically at some level. It’s not enough just to build things and put them out there.

When I say collectivist, I just mean that most of this technology is based on inputs from billions of people. There’s already a collective putting input in, and it’s meant to serve billions. So there’s a collective it’s serving. There’s really a big network that’s latent there.

Economics is critical. I don’t want to just say words. I want to write down actionable mathematical ideas.

Tim Scarfe

This is interesting, isn’t it? I think in the 1970s, Dreyfus came up with this idea of the first-step fallacy, and it’s related to the McCorduck effect as well. We create something amazing, and we just think we’re only one step away from being able to do anything.

These systems are incredible, right? They produce beautiful text, they can solve problems, and they can do programming. Isn’t it weird that they don’t actually help us that much? We thought it was going to revolutionize.

Michael Jordan

It’s not weird at all, because the model there is the old AI model. Let’s just build something intelligent, and it only got upgraded a little bit.

It's going to be a better search engine. That's fine. I do think the search engine was major progress for humanity. But now it became more than search. It was like a secretary sitting on your shoulder, helping you, whispering things to you. And it's just a dumb business model. I don't think many people really will want that. They'll turn the damn thing off. They want to think for themselves. They want, maybe at the end of the day, a summary or something, but they don't want this all the time. They're interacting with this entity thing. It's not a very good business model.

In the meantime, we have huge healthcare systems, transportation systems, and finance systems that are all based on data flows among many billions of agents, and are ripe for thinking about in a more economic way. They already have a lot of machine learning in them. What are the agents, what are they trying to get out of it, and what kind of cooperation and competition is latent there that you could improve? Markets arose thousands of years ago, and we learned about some of the principles, but we can improve them. Thinking arose billions of years ago or whatever, but we're not perfect—not just in terms of thinking, but also in terms of narrowly following our own agenda and hurting other people even though we don't want to.

Michael Jordan

Humans are wonderful. We want to prize human life, creativity, emotion, love, and so forth. That's fundamental, but humans also do bad things, and that's where technology should be able to aid you. You need to think about the system—the whole set of these systems. They're not really systems. They're big statistical boxes that take inputs and produce outputs. That's not a systems way of thinking.

There's a lower-level system, of course: the computer system. But I want to be above that. I want to say, what ecosystem does this belong to? Who is it interacting with, at what rate and with what kind of quality, and what kind of values are being created? When I say value, I often mean money. I want jobs out of this thing. I don't want it just to answer and do things for us. I want it to create opportunities for work and creativity, and so on.

Tim Scarfe

Rich Sutton is quoted quite a lot in respect of design versus evolution. I watched a wonderful talk by David Deutsch the other day, and he was talking about explanations. He said that physicists obviously go for these principled, low-level explanations, but sometimes you get these high-level coarse-grainings that are just really good, and maybe economics is one of those. What do you say to people from Silicon Valley like Ilya Sutskever? They're talking about human value functions and saying, "Okay, well, you've got these LLMs, and we just turn them into multi-agent systems and get all of the economic stuff that you're talking about for free." What would you say to those people?

Michael Jordan

It's just not a good way to think about engineering. If you were a chemical engineer back in the 1940s and 1950s saying, "We're just going to throw a lot of stuff together and make it work," well, you could do it, but you'd get a lot of explosions and a lot of economically nonviable things. You'd hurt a lot of people.

I think a lot of these people are not thinking about all the people who are already being hurt. Facebook and so on have damaged a lot of young people. A lot of teenagers are having mental health problems, and this is just not something that has been talked about by computer scientists at all. Now we're talking about yet another level of displacement. Jobs may go away, but, you know, that's tough. It'll create new ones, of course, like always. I just don't like to talk that way.

You've got to say, "Step back a moment. What is your point? Are you trying to create a new kind of market where people could come in and have their talents valued and appreciated, where bids could be put out for things that people might need, collaborations could emerge, and producer-consumer relationships could be explored, understood, and developed?" This could all be a mix of computation and humans. I think eventually we'll all merge, but along the way, just doing something so disruptive with all of these metaphors—that's not good social science or good mathematics. It's just metaphors.

Yes, you can build it because the previous generation of people created these amazing things that collect data, and we can do gradient descent on it and build ad hoc architectures. Yes, that works. It's amazing, but let's not give so much credit to the people who did that. It was the people 20 or 30 years ago who did that. The current generation is just way too—there's not much thought going on, not much intellectual stuff. It's just, "Yeah, it's possible to build it."

It's possible to steal the data from wherever you want because that's what the internet allowed to happen, and not return any value to the person who originated the data. It's possible to run gradient descent on that, but you need huge amounts of money. It's now possible to get it from people who aren't thinking very deeply. I may seem darker than I want to. There's a lot of good in the builders, but every previous era of engineering development—electrical engineering, chemical, mechanical—all had some builders, but they also had a lot of concepts and a lot of thinkers.

In fact, all of those engineering disciplines had something like Maxwell's equations or Newton's equations to help them. Here, it's just people who are very smart, who can code, and who have lots of intuitions, and it seems to me—I don't ever see anything that feels deeply intellectual. It feels like science fiction.

Tim Scarfe

Well, I suppose another thing that doesn't help is that these systems are like soup. There's even a field called mechanistic interpretability that tries to dig into the soup, and it's almost like they're searching for UFOs. They're trying to find these principled circuits that do reasoning or do whatever the thing is. I guess you could say, cynically, that it's not like when engineers build a bridge.

Michael Jordan

Well, I'm a little less negative than that. I don't think it's bad to build systems you don't understand. But then you've got to put things around them. The things that are around them are buzzwords like AI safety. It's a buzzword.

What you really need—I mean, you can't explain to me why you picked this Airbnb over another one or whatever. All the choices you've made today are inexplicable to me. They come out of your brain. I don't need to know all the whys and wherefores of your choices. What I need to know is that you're somewhat predictable, and that if I make certain options available to you, you're likely to take this one versus this one. Therefore, I can make my own plans, and we can start to interact, and so on.

That's part of economics. The economic style of thinking says, "I don't understand all these other entities out there, but there are certain rules of thumb that I can use, or quantitative predictions I can put in place, that allow me to interact, not get hurt, and even get value out of it." So, no, I don't think it's necessary to understand all the details.

Now, the input-output behavior you often have to understand better than we can now. For example, if I'm denied a loan at a bank and the bank used this big AI program based on past data, I want to know why. And "why" doesn't mean that you look in the internals and show me some circuit. No one's going to want that. They're going to want, "Here are 50 people who are pretty much like you, according to the embedding we're using in this big network. Of those 50 people who are like you, some got the loan and some didn't. Here, let me just show you what those people are like." You start to see, "Oh, I see that they differ from me in this way." That's actionable to me. I could now change things.

You have to build systems around this predictive system. That's a nearest-neighbor system, for example, and that system will supply what people might consider more like an explanation. It's not just about trying to go into the internals of something. Again, chemical engineering has thermodynamics, and lots of things are understood, but lots of phenomena were not understood for a long time. You mix up a bunch of stuff, certain waves are created, and certain things happen. You exploit that and move on. But you understand something about input-output behavior, constraints, and so on.

I think the current generation of neural networks will continue to be there. They have very nice scaling behavior. They'll continue to be part of the picture, but they really have to be thought of as part of a bigger ecosystem. Then you ask, "What can the neural network do in this context, and what's it missing? What if I have multiple of them? How do they engage with each other and with us? What is needed? What transparency is needed for the overall interaction to be an effective one, whether or not I understand all the details?"

Tim Scarfe

For some reason—and correct me if I'm wrong—I have an intuition that behaviorism is bad. Just by not having any mechanistic understanding and only looking at the outputs, there's the famous example, isn't there, of the hen that didn't know its neck was going to be broken.

And one example of this, actually, is AlphaFold. I interviewed John Jumper last week at Google, and you did some analysis on those 200 million predicted proteins, and you found they were very good, but there was something missing: you could robustify them.

Michael Jordan

You could robustify them. That’s correct. And I think that’s a good example. I’m a big admirer of AlphaFold. I don’t think it’s like an LLM. I think it’s targeted; it was for a particular set of problems, and it does them very well.

The issue that we found empirically was that when you ask certain kinds of questions—in particular, we did one where we were looking at whether quantum fluctuations in a protein were associated with phosphorylation, meaning the protein was active or not in the cell—you might think that these fluctuations, which lead to strands hanging off or kind of bad proteins, wouldn’t be used by evolution. But it turned out that a lot of them seem to be phosphorylated, meaning they’re reactive in the cell.

That suggests a hypothesis test: is there an association between yes-or-no phosphorylation and yes-or-no quantum fluctuation? So that’s a little 2-by-2 table, and you do a statistical test on that. The problem is that if you just use known protein data that’s been crystallized—you know, there’s a crystal structure known—you don’t have enough data to test that hypothesis with high power. You can’t reject the null hypothesis that there’s no association, even though there looks like there is.

If, on the other hand, you use 200 million proteins from AlphaFold, you can test the hypothesis with high power, and you reject the null hypothesis. But what we found is that the confidence interval on the statistic from that 2-by-2 table was extremely narrow and way far from the truth—the true, gold-standard value. And we found this in domain after domain.

So why is that? What’s happening there is that there are probably not many examples in the training set of proteins with quantum fluctuation, because it hasn’t been studied that much in the past and it’s hard to crystallize. Not many examples means it’s quite possible AlphaFold won’t give out a great answer, but it won’t tell you that. It doesn’t give you error bars, and it doesn’t specifically know the question you’re asking. That’s where I want the error bars, and it didn’t know about that question when it was built and designed.

Okay, all right. So now I have a good statistical question: what if I add a little bit of ground-truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth.

The answer is yes. There’s a methodology—we’ve developed something called prediction-powered inference—that does exactly that. It’ll cover the truth just like in a classical statistical setting, but it’s using this rather highly biased architecture. Now, it’s not biased overall. In fact, its accuracy is high overall. But for the question I’m asking, it might be very biased.

And that’s going to happen a lot in science, because scientists are rarely interested in just studying the past over again. They’re interested in brand-new things on the edge of knowledge, and that’s where, specifically, these foundation models will be most poor and most highly biased. So there needs to be, around any foundation model, the ability to maybe collect a bit of ground-truth data, merge it in with some procedure like this, and then give out a more trustable answer.

That’s all not science fiction. That’s what can be done and what really needs to be done, and I’m sure the AlphaFold people are on board with that. They would not find that weird or surprising. But a lot of other people out there talk about bias and all that, and they either don’t worry about it—I say it’ll go away if we have enough data—or they just critique the architectures and critique the outputs, but they have no scientific method in mind that’ll help us go forward. So that’s kind of the state we’re in.

Tim Scarfe

I challenged John a little bit about the extent to which AlphaFold understands, and he was basically allergic to the word “understands.”

John Jumper

We are not trying to tell you everything. We are not a model of the entire cell. These machines let us predict. They let us control. We have to derive our own understanding at this moment.

Right? We can experiment now on the artifact. We can look at the 200 million predicted structures, not just the 200,000 experimental structures, in order to help us understand. But it doesn’t do the act of understanding for us. It does the act of prediction and maybe control.

Michael I. Jordan

Why should AlphaFold understand?

Tim Scarfe

He was sketching it out to me. He kind of said that this is a weird alien artifact, and it’s not like it’s created; it’s refined. There’s this recycling pathway: you can put the thing through multiple times, and you can kind of corrupt it halfway through. The network is just iteratively refining—it solves the complex bit first, and then it’s refining, refining, refining. Could we interpret that as an understanding process?

Michael I. Jordan

I don’t think we need to. I think this anthropomorphizing of intelligence and understanding and all that is not necessary, not appropriate, and is a distraction for many, many problems. Why say it understands?

Some of my heritage comes from seeing, in real life and in industrial settings, machine-learning algorithms being rolled out 20 or 30 years ago. When I first went to the West Coast, I visited Amazon around 2000. They were using huge amounts of data to do supply-chain modeling using the neural networks of the day. It was random forests, and it was really working.

They could make fantastic predictions of whether certain ships would be delayed in the Indian Ocean or whatever, and so certain parts wouldn’t arrive in time. The overall supply chain takes billions of products and sends them to 100 million people per day. You cannot—there’s no way that any human can understand what’s happening in that big, big black box.

But it’s not necessary. In fact, you can ask, “Does that overall system understand transport and logistics?” And the answer is, who cares? It does a very important optimization and prediction process that allows an engineering system to be built around it. It brings down uncertainty. It makes it possible to do stockpiling and planning, and that’s what you ask for.

You don’t care whether it has to have a word like “understand” or “intelligence” applied to it. That’s for the media. The media—that’s kind of my problem with a lot of these people rolling out AGI and AI terminology. The media laps it up, even though we don’t have a clue what understanding or intelligence mean, and we, in our own research, realize we don’t care or need it. We want to build good systems.

Tim Scarfe

Yeah. It’s interesting, because I agree that we live in this complex, adaptive, irreducible system. We can’t essentialize it, and folks like François Chollet or even David Krakauer talk about intelligence as the adaptation and synthesis of coarse-grained representations.

But what if there is a bit of a step? Let’s not anthropomorphize it. Let’s say that understanding is not the endpoint; it’s about the path that led us there. We know that in the real world we’re a collective intelligence, and there’s the blind men and the elephant. We all take our own paths and lives, and we have different perspectives on the same whole.

So what if a better form of understanding is just being able to reconstruct the thing from your perspective using building blocks, rather than trying to essentialize it?

Michael I. Jordan

That all sounds great. It’s just not the language that those of us who do research would use. We would think in those terms a little bit, of course, but we would try to turn it into some kind of equilibrium or optimization problem. Here’s the information that’s available, and here’s the data, and here’s the power and the error rates, and we try to put a little bit of structure around it in that form.

There’s always this creative moment. I remember when I was interested in high jumping as a kid. You would go up to the bar and jump over it in various ways, and the barrel roll was the technique the Olympians were using. Then this guy, Dick Fosbury, came along and said, “No, if I go backward, I can do better.”

No one had thought about doing that. As soon as he did it, everybody did it, and the bar went up by half a meter or something—I don’t know. What process led to that? Was it an understanding process? It was just a little bit of “Let’s try something different” mixed in with the ability to try it out and do tests.

A huge amount of industrial planning is “Try it out and see what works.” Those are called A/B tests, and they’re done all the time. I’ve got nothing against that. It’s not based on understanding, but it’s led to optimized systems that can do things that people hadn’t thought about before.

So, a blend of that with understanding—but just understanding, you know, I was a cognitive scientist, and I was interested in neuroscience. One should be interested in those things. They’re fascinating, but they aren’t at the leading edge of thinking about how to build systems that work in the world, and they’re not at the leading edge of trying to build even the next-generation systems.

A lot of people keep saying, “Well, we’ve got to put logic back in, or symbols, because that came from our previous kind of view of what humans are doing.” Probably humans are capable of doing some logical reasoning, and probably have some symbols, whether they’re built into some complicated network or they’re reified somehow.

Michael I. Jordan

I don't know. My intuition is as good as yours. But really, the goal is to—I tend to be an engineer at heart, a mathematically inclined engineer. I want to say: What are you trying to achieve? Are you trying to displace teachers? Are you trying to make doctors better? What are you trying to do, what would be the abstractions and the points of entry into that problem, and how can you pull back from that and do it in some general way that's elegant and will inspire others?

Tim Scarfe

It's so interesting seeing different scientists, from a multidisciplinary perspective, attack this problem. Physicists, for example, work very, very low-level, and they talk about the dynamics of particle systems and whatnot. I'm really fascinated because you come at it from an economics perspective, which is traditionally dominated by this agential lens, and you talk about equilibria and incentives and so on. How does that come into it? How would you take a very complex system and almost decompose it into this new frame of thinking?

Michael I. Jordan

It's not been done really enough for me to have tons and tons of great examples, but we've been looking at modestly scaled examples. For example, we looked at a little bit of drug discovery and the regulation. I'm a pharmaceutical company; I test all kinds of proteins and throw them in animals and maybe a few humans to see what's working. I have some understanding—quote unquote. In other words, I know something about the evolutionary biology behind it, and so on, and that guides me. But at some point, someone has to really test this out in the real world and decide. Regulatory agencies have to come in and say, “Yeah, that goes to market,” or it doesn't. Okay.

Tim Scarfe

So now you've got a tangled web of scientists and pharmaceutical companies—not just one, but many, many of them—and proteins, and now you've got to think about how that system is behaving. Hopefully, the regulatory agency is trying to have the number of false positives and false negatives be low across the entire system. That's the goal of the problem, so it's a statistical problem.

Oh, but wait—a classical statistical problem. You would just go gather IID data, independent and identically distributed data, from some source. Here, no: The data is coming from self-interested pharmaceutical companies. What's their motivation? Money and whatever—maybe to help. They want to help people and make money. All that's hidden from you as a regulatory agency.

Now the economic mindset comes into play. An economist says, “Well, it's hidden from me, but it's not arbitrary. I can probe in various ways.” So that becomes very economic. Economists think about how you set prices. If I've got a lot of people coming onto my airline—say, 1,000 people have just arrived who want to go from here to London—every one of them has a different price point, and that price point will shift in the moment. How eager are they to get there? It's not just because they have a lot of money; it's because they have needs, and I don't know what those are.

So what I do is set various services and various prices that bracket the possibilities, so that overall it's likely I'll make enough money, everybody will be happy, and the service will go forward in life. It's a blend of knowing a few things, admitting that you don't know other things, but putting it in a system that can actually work with that kind of mix of asymmetries and incentives.

The incentives there are that there's a certain service and price. If you pick that one, you're likely to be able to get on the airplane and have the goodies you need, or whatever, and that doesn't make you do something; it incentivizes you.

In the pharmaceutical world, if I could get them to be incentivized to mostly send in drugs they've done some testing on, or that they have some belief is a pretty good one, and not just throw arbitrary ones at me, then maybe the overall system will actually have the error rate you want it to. Okay? Because if you don't do that, if it's a drug that a billion people will use, then you're going to make money whether it really works or not.

So you need it to get to market. How do you get it to market? Well, you just throw it at the regulatory agency, and maybe there's a false positive. They just got a false positive and put it on the market. You make a ton of money. If there's enough of that, the incentives are all wrong, and the overall system will not control Type 1 and Type 2 errors.

Those are examples we've actually worked on. But I just hope you can appreciate that when all this stuff starts to really roll out in society, it's not going to be that there are a few big LLMs and everyone consults them like a search engine. That's just not the model. It's going to be that there's local data, like I told you about with prediction-powered inference. Everyone's got to vet what's coming at them.

There's also going to be local data because I collected it with some expense and I don't want to just give it away. Thank God, finally, Anthropic is paying people for their data. That has got to be the future. So I'm going to have some competitive value in my data and not just give it out.

If you start interacting with lots and lots of people who want to get some value out of the interactions, you have to talk about the incentives. What's the incentive for them to send the data? But not only send the data—send correct data. Send truthful data. Don't just add noise; don't be adversarial.

I cannot imagine a fully fledged version of all this rolling out in society, and all of our decision-making throughout our lives, without a deeply microeconomic perspective accompanying the gradient descent on data.

Tim Scarfe

You spoke about this three-layer model. There was an example where you might have consumers, and they might have their data, and you've got Google. Google is using the data, consumers are getting a service, and then Google might sell the data over here. That's kind of like a traditional model. Let's start with that.

Michael Jordan

Okay. So those are really kind of Bohr-atom kinds of things. We're being scientists there; we're trying to say, what's a minimal model that exhibits some of the behavior that we want to study here? So let's think about a data market, because data is not just something you analyze to build a big LLM. It's also something you'd sell and buy, and it has value, and there are also privacy concerns about data. So let's put a little minimal model together where we could study that.

One we've done is—we call it a three-layer data market. It exists in the real world. This is an abstraction, but it's a real thing. You've got a user, or multiple users, coming into some platforms. The platforms provide a service—imagine a payment service—and as I use that, they get data from me. They learn about what kind of purchase I've made and so on, and they use that data to make their service better. That's a good little loop there, right?

The problem is that rarely do they make enough money off of that service. They take a small cut that the merchants don't like to give them, so they have to do other things to stay in business. So typically, for a long time now—probably 20 years—they've been selling their data to third-party data buyers.

These are not evil people just trying to ruin people's privacy. They're trying to do market research, learning what would work and what people are really doing. This is behavioral studies, and there is value to them. They pay for it. All right.

Google doesn't need this because they created this artificial advertising market, which has superpowered all this nonsense—we could talk more about that. But other companies, like Mastercard, do have to sell their data.

So now it's a three-layer thing, and as soon as that third layer was introduced, the equilibrium had to shift, because the user who's sending their data in just lost something. They lost a little bit of privacy. Some third party that I don't know anything about is getting data about me, and I can't just accept that. But I can't walk away, so there's stress on the system now.

In an effective economic system, what would happen is that you wouldn't just wait for the regulator or the government to come in and say, “No, this can't be done.” What you would do is the platforms would say, “We'll offer you a tunable level of differential privacy for some cost,” or, “I'll offer you level 0.3,” and some other company says, “I'll offer you level 0.7.”

The user looks at that and says, “Ah, 0.7. That's better.” I really care about my privacy, so I'll go there. That company will then start to get more data, and its service will get even better. Ooh, you've got a nice little feedback loop there.

But now the data buyers will look at the data from that person. 0.7 means more noise has been added to the data; it's less valuable to the data buyer. The data buyer will say, “Oh, I'll spend less. I'll give you less money for that. I'll give more money to Google.” So now you can see there are conflicting tendencies here. The incentives are aligned, but they're not optimal for everybody. And so now the mathematics is not just an optimization problem.

The mathematics is an equilibrium problem. But it’s an equilibrium problem that involves statistical assertions, data, and how much you can predict with this data. You quantify that with error bars and statistical predictions.

You put that all together in a big mathematical system, and you can find the equilibria as a function of various system parameters. For example, is there a minimal level of privacy the regulators could require? Or is there some heterogeneous privacy budget? You can put in various ones, and now you do a little plot of how the equilibria move.

The equilibria have overall utilities for all 3 players summed up. That’s the social welfare. You can ask how high the social welfare is at that equilibrium versus this one versus this one. Another regulator could look at that and say, “Well, I prefer this one because it’s overall higher social welfare,” and laws could be made at that level.

Even though this is a little toy model, it has the ingredients that I’m very interested in: predictive models, data markets, money incentives, and a real system that’s already kind of working, but people aren’t thinking about very well, just like in the drug discovery domain. If you take an economic point of view, you can make the system better, right?

Tim Scarfe

Cool. So, modeling it as a dynamical system, which we can simulate, and then we get these modes of—

Michael Jordan

No, we don’t have to simulate it in that case. You can actually write equations and calculate equilibria. It’s a Stackelberg game, and you can actually find the equilibria. In other cases, you would simulate, but the point is that a lot of my machine learning colleagues don’t know much about fixed-point algorithms, finding equilibria, how they shift as you shift various parameters, and all that. That’s economics stuff.

Machine learning people are really good at optimization, but this is not an optimization problem. There are all these algorithms in other branches of mathematics that find Pareto frontiers and do it statistically, as a function of the size of various markets and the size of populations and all that.

It’s kind of amazing in this era that the 2 have almost never met. Economists never had a lot of data to inform the design of their market, so they just wrote down a bunch of equations, made rational assumptions, and then found equilibria mathematically or otherwise. Machine learning people never thought about the equilibria. They just had a lot of data and used it to do the obvious thing: predict the next word in a string of words.

The future has got to be that those branches come together. The economics-equilibria perspective is critical, but the “it’s got to be adaptive” perspective is critical. You also alluded earlier to some of the machine learning Silicon Valley types just saying, “Well, we’ve got all this data, therefore all the behavioral stuff is already built in.” That’s too naive, obviously, but it is a useful point of view in a certain sense.

Economists do make rational assumptions they shouldn’t have to make, and if you put in data instead of that assumption, you’ll probably do better. You’ll have some of the behavioral economics already built in. But if you do it outside of any economic thinking whatsoever, you’ll just make a mess of things. Of course, that’s what Silicon Valley seems to be pretty good at.

Tim Scarfe

What is the difference between data and the kind of knowledge that you’re talking about?

Michael Jordan

Social knowledge is very ephemeral, and it’s very much in the moment. I walk down the streets of Copenhagen here, and there are all kinds of little markets out there, what’s available and at what price, what I might like, and all that. It’s all super-ephemeral, and that’s a better way to think about all this.

You can’t just gather enough data to know that the person walking down the street there is going to come buy this product. All of our decisions and choices cannot be reduced to having enough data to cover everything that’s going to happen, even in the next 10 seconds.

You have to be a little more humble about that. I have a lot of ignorance, but that doesn’t mean I can’t build a safe system, like a market of some kind, where people could come in, not get cheated, get value, and where it can evolve over time. It can shift in ways, and I’m not the god figure at the top designing the human value function or whatever, so that it responds particularly well to what humans really want in my God’s-eye view of the world.

No, it’s got to be a system that permits bottom-up preferences to be expressed in the way humans want to in the moment. That’s not built in. The system respects those things, wants to maybe learn more about them, and then use them in the moment. Maybe it keeps some of it; maybe it’s all ephemeral and goes away.

There seems to be a huge naivety about what data can do. Even if you have whatever exabytes of data, you’re going to miss all the details that are probably the main thing that matters for a particular class of decisions. Even AlphaFold, based on huge amounts of data, doesn’t do very well on certain queries. They’ll patch those, and it’ll do better and better, but the new questions that people ask will always be on the edge of knowledge.

People aren’t thinking about that. They’re thinking, “Well, I just have to replace the teacher, because a teacher is working not at the edge of knowledge. They’re working back in all the stuff that’s already known.” That’s fine. You can aid teachers, but good teachers also know how to migrate to the edge of knowledge.

Tim Scarfe

Yeah. When we do abstraction and idealization, it’s always a little bit lossy. I’m fascinated by this observation that markets were around before capitalism. It’s this bottom-up thing. It kind of—it’s a—

Michael Jordan

This is not capitalism. That’s one methodology for making markets work, but it’s not the only one.

Tim Scarfe

Exactly. So, it’s a natural phenomenon, the emergence of something we might call markets, and it’s constructive, divergent, and diverse. You’re saying somewhere the rubber can meet the road, so we can create abstractions, we can do some kind of modeling, and we try to do it in such a way that it’s not too lossy.

Michael Jordan

Absolutely. Human culture creates abstractions. Individual humans create abstractions, too, that work for them. When those abstractions are useful enough, and people can communicate them and get them promoted into the culture, that flows up and down all the time.

Indeed, that’s something that systems could perhaps help with. I’m not going to just trust systems to take on that burden, but they could be helpful. It’s not just the individual cognitive entity that creates abstractions and that we should reify. Cultures create abstractions, and you can study the microeconomics of that or whatever, or you can just say those abstractions that emerge are part of the culture. They’re useful, and they’ve stayed around because they’re useful.

Going forward, it’s not that we’re going to just keep the old ones. We’re going to build systems that allow new ones to emerge. That is not the god figure figuring it all out and putting them in there.

Silicon Valley says, “We’ve got so much data, and we’ll have so much that we can do all of it top-down.” They’re forgetting, first of all, that the data came bottom-up. The data was all contextual, and the data was supplied by people. They’re also forgetting that it’s all got to continue to be at a micro level that is going to be beyond their ability to sense, and we’re not going to want them so much in our lives.

After the search engine, which I thought was a fantastic piece of technology that allowed access and all that sort of thing, a lot of it became very creepy. “We’re going to put glasses on you. We’re going to put cameras around your home, and we’re going to know all the details of your life, and we’re going to make your life better somehow.” That equation did not calculate for me.

Tim Scarfe

Yeah. This idea that culture is the abstraction hard drive in the sky, but culture is very adaptable. We can delete strategies. Knowledge decays very quickly, and organizations maintain knowledge. Really good bits of knowledge stay around for a long time, and down at the bottom-up level, we’re creating new bits of knowledge.

How does that whole ecosystem work? How do we designate good things that stick around, and how do we find new bits?

Michael Jordan

We don’t. You and me—the answer is, you and me don’t. But that’s something that good intellectuals do.

There’s a whole field of organizational behavior, or how organizations effectively emerge, and that’s really interesting. It’s not everything, but there’s a lot known there. Some of it is mathematical, some of it is not, and some of it is best practices.

Those are the kinds of ways that these AI ecosystems should be talked about, not just in terms of neuroscience and metaphors of neurons, physics metaphors, and all this stuff. That was part of my heritage, too, but it felt so lacking when we actually saw these things—the rubber hitting the road, as you say.

And so, yes, behavioral organization—how people are organized into things that promote not only good revenue for companies but also promote democracy, and so on. There are people that talk about all these things, and I just don't think they seem to have much presence in Silicon Valley. Maybe that's for good: Silicon Valley, let them go. They'll just burn a lot of money and cause some headaches, and they'll also create things like search engines.

I think a lot of the companies are also really focused on creating value. I do think Amazon is different from Meta. Amazon has got a business model: bring packages to doors. Behind that, they create some technology to support that, and it's mostly doing good, in my view. You have to worry about labor markets and so on, but those are all good things to worry about.

Just creating computational artifacts that make predictions and that, if you were to wear goggles, you would be able to live in their world—it's not a business model. It's a science-fiction dream that may or may not be helpful for humanity.

Tim Scarfe

Well, can we explore that? You gave the example of Spotify. We were talking about this three-layer thing before, and now we've got this weird incentive structure where Spotify are actually incentivized to generate the songs with AI, right?

Michael Jordan

Yeah, they are. I have a project; I'm a scientific adviser to something called UnitedMasters, which is an alternative where musicians keep their work. UnitedMasters connects them to brands and other kinds of opportunities, so that they are more like real artists—not just that their song got streamed and they got a little bit of money.

Spotify, indeed, is close to perhaps a monopoly, but it's not incentivized to pay. There are monopoly prices, if you will. One would hope that somehow the market will fix that—that enough young artists will say, “I'm getting screwed here. I'm not making any money,” and another service will emerge. But we are in an era where some of these services do become monopolies pretty quickly.

I'll leave that to my economist friends to think through and to look back at historical examples and think about whether regulation is needed, or whether there's another market-making mechanism that'll make this healthier for human beings. I'm not against Spotify, but it should be part of an ecosystem that actually rewards the artist more.

Right now, an artist is getting paid very, very little, and I don't believe the prices are being set under competitive mechanisms. I think there's this broader macroeconomic view of what these systems are doing and what their role in society is going to be. With the search engine, many of us were puzzled about how they would make money. It just didn't seem like there was a money-making model, and then the whole advertising thing was a bit of a surprise, at least to me, that it would become so huge.

Of course, the underlying thing is that people expect things for free. All right? Google couldn't make people pay, but I think they made a mistake at some point. With YouTube, when they acquired YouTube, YouTube was more than just pointing people to a website. YouTube was incentivizing creators to create things that people would watch.

At that point, I think a socially responsible Google—to critique them a little bit—would have said, “Oh, we've created a market here. We've created a producer-consumer relationship. We've got to make that market a little bit more valid, and we could actually have it so that when someone's watching things, they can have some sort of economic connection to the person who made it, and there can be incentives flowing. This person is now incentivized to make more because here's my audience connected directly.”

Instead, it was all going through Google, and then Google was putting advertisers next to it to make a ton of money for themselves. There was a modest incentive to give back a little bit of money. That, to me, was a huge mistake, and then Facebook made it even worse.

Tim Scarfe

So you've butted up against folks like Geoffrey Hinton and Stuart Russell at Berkeley, and these guys are painting a picture that this technology is recursively self-improving, that it is an entity, that it's not a cultural technology—it's a thing in and of itself. This seems a little bit science fiction on the first read.

Very science fiction. What do you think?

Michael Jordan

I think it's science fiction, and I think science fiction is important for society, but at the level it's being promoted, and with those kinds of voices, it's really hurting 25- and 20-year-olds. These young folks, of whom there are huge numbers, are excited about technology, and they want to build things that help their family and help their country. Actually, more their family than their country, honestly. They see real opportunities in doing that.

They're being told by the leaders, “Well, we had our fun. We developed a bunch of algorithms. We did it, and we were just interested in the pure understanding of intelligence,” even though they didn't understand intelligence; they built gradient-descent algorithms. “And now you guys, you can't do this because it's dangerous. It's going to wipe out humanity with a high probability, or superintelligence will arrive soon, so there's nothing left to do. That's in your lifetime.”

That is so demoralizing. So demoralizing. That thing, I think, bothers me the most. The second part that bothers me is that there's no economic thinking going on there. It's zero. It's really about a cognitive-science mentality or neuroscience: “We figured out how the brain works. It's gradient descent with a lot of distributed neurons, and the fact that these LLMs are working so well shows that we figured it out. It wouldn't work so well otherwise.”

Well, I think that's dubious. We don't know; the brain is way beyond. I mean, you ask a neuroscientist if this has anything to do with the brain, and basically they'll say no. It's a nice metaphor. It's a cartoon. Does gradient descent work at massive scale? Yeah, more than we would have ever imagined.

But is it showing its weaknesses? Yeah. Can it be fixed in certain areas? Yeah. You build certain verticals, they'll do good things, and it'll make mathematicians go faster, but it won't put them out of business, and so on. It's having a big effect on society. I worry more about labor-and-capital relationships than I worry about it deciding to take over.

The rest of it, to me, is more on the ground: how does the next generation take technology and work with it? I don't think that voices like that are actually helping that generation to perceive what they should work on and why. Superintelligence versus extinction—those are your 2 options. Goddamn it, those aren't the only 2 options. There's a huge number of very positive things that can be done at human scale.

Let's hope that enough of the young mentalities get behind that. They don't have enough examples of people out there who made money by making things. Did Sam make life better, you know? Not clear. I think in previous generations there was a little bit more: “Here are people that are out there making things—vaccines or whatever. Oh, I want to be like that.” And right now, not so good.

Tim Scarfe

I don't know if I could get you to be a psychologist for a minute and try and understand why these—I don't know whether it's the search for purpose—but have you noticed as well that some folks think that there's going to be a utopian future? When you speak with them, they have quite a similar DNA. They also think that it's recursively self-improving, it's going to be superintelligent, and so on.

If I was to press you to say, “Why do they believe it?” I can understand, right, these things are so clever, but why is it? Why do they believe it?

David Krakauer

I mean, they're clever in a recognizable way, at some level. They're taking all this human cleverness and packaging it in a new way. Again, I think it'll always be missing a little bit of the point because it's not in the moment; it's not the ephemeral stuff. But that doesn't mean it can't be even more clever, and I kind of think that that's okay.

For me, the goal here is not to build a superintelligence and have it dictate or tell or anything like that. It never was. I'm kind of shocked that some people seem to think that was always the goal. To me, it just never was. Rather, again, I think I said this in the very beginning: humans are wonderful.

I'd hate to have robots taking over from us, and I don't think that's going to happen. There's just too much good about human nature and about what humans are able to produce—things that are shockingly beautiful, creative, and inspiring. We need to support all that.

Now, the issue, though, is that the flip side is that we're far from perfect. People really hurt a lot of people, and they're being empowered to do yet more of it. We are very narrow-minded. We also don't understand people; people hurt other people often because they don't understand their motivations. They've got a misunderstanding.

How many wars are created because someone didn't understand the intentions of the other side and they said, “Well, let's just proactively—let's just bomb them”? That's just all the time. That's how humans act and think.

What’s missing there is an appreciation of uncertainty and information signaling. Eventually, game theory arose to help people think it through a little bit, but it’s still extremely rough.

If you look at our political system, an aged charlatan leading a country—this is our optimized human system for making decisions at the highest level. There’s so much room for improvement of the human being, and democracy’s got to be the way. But democracies right now are a few aged people sitting in various rotundas in various capitals, not knowing what they’re talking about, mostly.

We have a very broken human system in many, many domains. We have a few that are pretty good. I think the universities are pretty good, and I think a lot of companies are pretty good, and a lot of human associations of various kinds and various skills are pretty good. But we have so many broken ones.

And so, to me, that’s what AI is about. AI is about helping with the things that were too hard for humans and aiding the information flow so humans could actually make the good decision in the moment that most of them really wanted to make, and not make the bad decision that they were afraid they had to make because they didn’t know enough.

There’s so much opportunity if you think about it at that level. That’s what AI is about to me. AI is not about replacing the human with the computer, or the recursive self-improvement stuff. It just feels like a metaphor. We work with recursive algorithms. We work with improving algorithms. I don’t see that getting out of control like a virus.

We’re going to work with these systems and, as I say, hopefully mostly focus on getting right some of the things that evolution didn’t quite get right for the human being, especially at a scale of 7 billion. Evolution perhaps didn’t prepare for that. Focusing on that, to me, is what AI can be about.

So I’m positive. I’m bullish about AI in that sense. I’m appalled by the dialogue that has developed between the people who have all the money and want to just build something for the sake of building it, and the people who are just anti-intellectually saying it’s terrible, it’s going to destroy all of humanity. That’s the dialogue in the public eye right now, and I find that so harmful.

It bothers me that people who worked on it for all these years think that we’ve reached the end. Somehow, gradient descent is like the brain, and therefore you could take multiple brains and fuse them together, and, oh my God, it’s going to do incalculable things. That’s just science fiction. Whether it’s true or not, it’s not even worth thinking about. What’s the path? How do we engage younger people to do things that are actually positive? What mechanisms are we going to talk about? What kind of education are we going to talk about? What goals are we going to set?

The thought leaders are not talking in any of that kind of language. I think it’s unusual for human history. The thought leaders are heading off in these two directions.

Tim Scarfe

It’s also complex because there are very real security and safety risks of having any autonomous software just doing things without direct human supervision. So we should—

David Krakauer

Well, yes and no. Think about airplanes. That’s the classic example, but there are very, very few airplane crashes at massive scale these days. There used to be a lot when I was a kid, and it’s because of the autopilots.

Tim Scarfe

Yeah.

David Krakauer

Mostly now, planes are flown by autopilots, and the human can come in as needed. It’s because of that. This blend of automation with humans is actually the most effective way to go. Again, it’s improving. Humans didn’t evolve to be flying this big thing up in the air, so you can improve upon human ability there. You put the two together, and you can do something that’s helpful for everybody.

Tim Scarfe

I suppose in that case it’s quite a well-specified problem. We want to go from A to B, and here are the parameters.

David Krakauer

Yes, to know—I mean, you have multiple planes in the air. You have clouds, changing weather patterns, and some person who did something stupid. It’s easier because, up in the air, there’s a lot of room. In 3D, there’s a lot more room than in 2D.

But in 2D, you’ve got all these cars flying around. You’ve got tens of thousands of people dying each year in each country. It’s a mess at some level, even though it’s very important and effective for many of us. So we do it.

A hybrid system that had a lot of autonomy with some human involvement, and so on—that’s what we should be thinking about. But you’ve got to think about it at the system level. Just putting a superintelligence behind the wheel of a car is a dumb way to think about technology.

Tim Scarfe

Is there any hope? I don’t know what you think would be the thing that would make these folks update.

David Krakauer

I think that Ilya Sutskever and others have done some great things. They built systems that all of us are not only using, but that are changing our thinking and everything.

I think that’s what I get out of what they’re saying: I’m a builder. You think I’m a guru and a thinker, and maybe I think I am too, but maybe I’m really better as a builder. I can build things with the resources that are now available to you. Again, it’s not just the money; it’s the whole internet and all the things that previous generations of people did.

One thing that bothers me a lot about these people—not the Elon Musks or the Sam Altmans; they’re just coming in and taking the cream off the top from all this effort that people put in—is that a lot of these people rightly are not just wanting the credit. They’re annoyed that this is the direction these people are now taking it, without appreciating why these people were building these things. They weren’t building them for you; they had other goals in mind.

I think these are some builders, and there are some very impressive builders. But I think there’s this system that we call Silicon Valley, or whatever, that these people live in, where the more outrageous, the more far-flung, and the more physics-, biology-, and neuroscience-inflected your language is, the more you sound like a guru.

People enjoy that posture and that activity, and it creates a great amount of money. They don’t care about the wealth, perhaps, but it allows them to be even more prominent because they can now have another company that tries some other crazy thing. If it doesn’t work, that’s a sign that you had a great idea.

It’s a world I wouldn’t want to be in. I’m trying to become a bit of a historian. I mentioned chemical engineering and electrical engineering, but you look back at the history, and there were some glimmers of these kinds of things. I think this level of detachment from reality is unusual in human history.

This level of “my crazy science-fiction, 25-year-old dreams are all I’m going to pursue for the rest of my life, come hell or high water,” and then at some point I’ll flip because I realize, “Oops, I didn’t really have a great goal in mind at all. What have I got here? I’ve just spent a lot of money, and I’ve got this thing, and I don’t really know what to do with it, and I’m worried about it”—to me, that’s a sign of a certain level of immaturity, frankly.

Tim Scarfe

Circling back, you were talking about statistical contract theory, which is when we have things with an information asymmetry and we model incentives. A lot of folks in the audience would have heard of game theory, right? What’s the difference?

Michael Jordan

Game theory is a mathematical discipline. It started with von Neumann in the 1920s, and it’s got many, many branches to it. It’s a mathematical way of thinking, really.

One way I like to think about it is that it’s like F = ma. It’s a set of equations that will make predictions. If I write down a game, just like I wrote down F = ma in some coordinate system, I can now predict what will happen.

In the case of F = ma, I integrate a differential equation. In the case of game theory, I write down the game and calculate the Nash equilibria, the correlated equilibria, or some other equilibrium concept. I say, “Here’s what will happen in nature,” because my little mathematical model captures the appropriate ingredients.

For F = ma, the thing follows a parabolic curve. It means the theory is right, and then Einstein says it’s not quite right and makes a better one. In game theory, it’s the same thing: You look at whether those equilibria actually characterize how systems, organizations, and people behave. Sometimes yes, sometimes no.

Those aren’t the end-all. There are all kinds of other equilibria: Stackelberg equilibria, sequential equilibria, and various kinds of figures of merit. There are various social-welfare constructs, various regret constructs, and all sorts of things. It’s a whole huge field of its own.

Let’s think about it eventually becoming as big as physics, because it’s all about strategic interactions and so on—not molecular interactions, but strategic interactions. You can also ask the inverse question. In physics, the inverse question would be: I want to build a bridge. My goal is not just to see whether something follows a parabolic path or something.

I want that bridge to stand up. So I invert F = ma. I go from the goal back to the design that would ensure that the thing stood up. All right? Most engineering fields are inverse problems: they go from the goal back to the design, whereas the forward direction is science. You say, “Here’s the setup. Here’s the prediction. Is the prediction realized or not?” So, yes, it is. That means the model must be good.

What’s the inverse of game theory? It’s outside of economics, perhaps not talked about that much. Game theory sounds like it’s everything, but the inverse of game theory is what’s called mechanism design. Mechanism design says, “I want a certain outcome in the world: that this person gets paid, that the wealth is divided equally, that there’s some fairness, or that some market is created. What game do I design so that outcome is realized?” So, I’m the designer of the game. I’m not just taking the game as given and looking at what it predicts.

Mechanism design has many pieces, too. I work in contract theory, which is a part of mechanism design. It says, “What if I have 2 entities interacting? They’re not symmetric. One knows more than the other, and they have to interact with each other.” That’s contract theory. Auction theory is another part of mechanism design, where I’ve got a bunch of people coming in and I think of them as symmetric. I don’t know who’s got more money or who wants to bid more than others, but I have this mechanism called an auction that reveals their value. The outcome is that the person who wanted the painting the most got it. That’s the desired outcome. That’s 1 desired outcome.

Anyway, long story short, game theory is a super-rich, not-so-old discipline—100 years now—that’s continued to evolve and supply all kinds of algorithmic ideas for those of us who are in the business. I’ve been mostly a statistician in my career, kind of worried about uncertainty, probabilities, decision-making, and uncertainty. But when I go to equilibria and games or economic ideas, game theory is part and parcel of the thinking.

Tim Scarfe

You’ve said that we need to be thinking about—I mean, we’ve spoken about incentives, we’ve spoken about collectives. The other big one is uncertainty quantification. Now, there’s this wonderful field in machine learning called conformal prediction.

Michael Jordan

It was invented by my professor at university, Vladimir Vovk. We learned about the transductive confidence machine, which—

Tim Scarfe

Oh, nice, yeah.

Peter Grünwald

These measures of strangeness would be, for example, the distance from a hyperplane on an SVM. You can basically calculate something like a p-value and have a confidence region.

Tim Scarfe

An e-value. Go on, tell me more.

Peter Grünwald

I don’t want to get into technical talk about e-values, but Vladimir is fantastic. I don’t know what he thinks of himself as, but I think of him as a statistician, with a game-theory background, too. He’s in the school of Phil Dawid and David Blackwell, who spilled out of statistics to do all these other things.

Classically, p-values were a one-shot quantity that statisticians would talk about. Fisher said, “I’ve got a model of what’s going to happen in the world. It gives a probability distribution on the outcomes. Some outcome arrives, and it looks very improbable under that model. The model must be wrong.” That’s the p-value. The p-value is the tail probability.

The problem is, if you do that repeatedly and look at perhaps the smallest p-value along the way, that’s called p-hacking, and it gives you wrong answers mathematically and in practice. E-values are different. An e-value is the expectation of some nonnegative random variable, or a nonnegative supermartingale in more generality. You’re watching this evidence accrue, and you make sure the expectation of that evidence is less than or equal to 1 at each step.

You can think about multiplicative evidence gathering: if it’s always exponentially less than or equal to 1, then it will stay below 1. If it’s nonnegative, it will just decay away. Under the null hypothesis, I’ve got this stochastic process that is decaying away. I can look at that at any time and assert that it’s decaying away. I can look at it repeatedly, keep asserting that, and have control.

There’s something called Ville’s inequality that Vladimir and others have exploited. It says that this can be controlled over the entire path of this thing. Now we can do statistics in a new way. It’s called anytime inference. We can peek, we can change, and we can gather new data. We can do this in a very liberating way.

Vladimir is one of the leaders of that, and an e-value is one of those martingales stopped at a particular time by the optional stopping theorem. You can stop it whenever you want. That has opened up a lot of connections. In fact, our statistical contract theory—what is a contract? Remember, it was like services and prices. The services are like evidence gathering, and the price is also part of it. It’s a random variable.

It turns out that we can have incentive compatibility in contract land if and only if we have e-value validity in statistics land. There’s a nice, tight connection between game-theoretic probability and the theory of incentives. To me, uncertainty quantification is rarely just, “Here’s an error bar.” That’s classical statistics. It’s more about what the context is. The context might be a contract, or it might be some other evidence-gathering mechanism. This way of thinking opens you up to a broader class of evidence gathering.

Tim Scarfe

Very cool. I should say that in your paper, you have this figure of a triangle, which we’ll put on the screen now. You’re saying there’s economics, computer science, and statistics.

Peter Grünwald

These are thinking styles. I don’t even call them by those disciplines. There was a paper by Jeannette Wing a couple of decades ago talking about computational thinking. It says, “Computer science has developed these thinking styles that are more abstract than just computers.” It’s things like modularity, abstractions, APIs, and all that. Why don’t we teach everybody in all the sciences and all the disciplines to do computational thinking? I think that’s totally right on. That’s great.

But lots of algorithms don’t come about from those kinds of computer science principles. They come about from thinking about inferential uncertainty: how do I gather data to make predictions about things that don’t yet exist? And they come about from thinking about incentives: how do I make sure that incentives are in place? I called those 2 kinds of thinking inferential thinking—not just statistics, because a lot of fields have inference in them—and economic thinking. It’s not just economics; it’s social scientists of all kinds, legal scholars, and so on.

When you put those 3 together, you get a pretty good platform for training the next generation and a pretty good platform for problem-solving of the kinds that we’ve been talking about this entire time. Just 1 of the fields—computational logic and optimization—kind of gives us LLMs. Fine, great, but it doesn’t give us any of the context around the LLM. The incentives give you the whole thing we’ve been talking about. Statistics, to me, is critical. It thinks about what kinds of errors I’m going to make and how to make sure the data is controlled so I don’t make those errors.

When we put the 3 together, they also bring some partners. The economists talk to the behavioral psychologists. The computer scientists talk to the physics people or whatever. The statisticians talk to the legal people or whatever. There are whole subcommunities that come together.

To me, if you put this triangle there and think about what’s around it, it starts to become a new way to think about academia. This is the liberal arts of the era. This is the core. My colleagues in the humanities might disagree—the core is still the humanities—but I don’t think it’s touching the core intellectual issues of the era, which are about data, compute, and all that. I want to put the ingredients in place so that those things are thought about in a socially responsible way.

Tim Scarfe

But could you bring this to life? You famously spoke about having a language model and asking it, “How confident are you about the answer?” It tends to be quite bimodal, so it’ll either be 1, 0, or not. What’s the difference? Why doesn’t the language model really have any idea about its confidence?

Michael I. Jordan

You should ask the language-model builders, because all they’re doing is predicting the next word, and there isn’t any thinking about uncertainty quantification in doing that. You can graft in ideas, but they’re dubious. They’re often putting in a dubious prior, and you can go to the statistician—that’s what people have done—and say, “Okay, I can just treat it as a black box and put conformal prediction around it.”

It’s a nice method. It doesn’t require a lot of assumptions. So, yes, that’s true, but it’s not assumption-free: there’s an exchangeability assumption.

The data, if you scramble it, you get the same. While I think all that’s crucial and important, I tend to think more about the broader context. I gave an example in the article you mentioned of a duck who goes to a lake—a statistician duck.

It has calculated that, over the last year, there tends to be twice as much grain on that side of the lake than on this side—a 2:1 ratio. So the next day, the duck needs to decide which side of the lake to go to. The Bayesian duck, who has those probabilities, would then maximize expected value and go to the left side of the lake with probability 1, because it’s always right.

The actual ducks don’t do that. They go to that side of the lake probably 2/3 of the time and to the other side 1/3 of the time. They’re hedging. But it’s not just a hedging thing. Hedging would mean occasionally going to the other side of the lake; they’re actually getting the right ratio.

The explanation is that you weren’t thinking about the context of this uncertainty. It’s not just you, the individual duck. You probably evolved in a world where there are many ducks, and if all the ducks went to the same side of the lake, obviously you’d miss out on a resource.

So, is there an algorithm that allows many ducks to cooperate here? If they all have that same uncertainty, then they can sample with probability 2/3 and go to this side, versus 1/3 to the other side. That’s actually a Nash equilibrium of the bigger system.

The right way to think about uncertainty there is: in the context of the population, how should I use my uncertainty? Another kind of uncertainty is on the economic side. One uncertainty in economics is the one I’ve alluded to: information asymmetry.

There are things I don’t know, and you have expertise that I don’t know about, but we’re going to work together. I might give you a contract—a menu of options. But even if I interact with you for a while, I still might not know. There are things you’re going to know that you’re not going to give away to me, and maybe you’ll hedge—you’ll lie a little bit.

I don’t know about that. That’s not just sampling. That’s a different kind of uncertainty.

Finally, there’s what I like to call provenance. That’s more like a database kind of uncertainty. If I want to have a medical operation, and you’re a doctor, you look at the data for people like me: here’s the probability of survival if you do the operation this way versus this other way.

I look at that and say, “Great,” but then you tell me, “That data was gathered 10 years ago.” I’m going to say, “Okay, my confidence interval should go up.” Classical statistics could talk about that. In fact, it would be more Bayesian to think about that, but it doesn’t. It just says the data is the data.

In a bigger system where data is flowing around, it should always be tagged with metadata about how old it is, and that should be quantitatively brought into the uncertainty quantification. We’re not doing anything like that right now.

The poor LLMs, which are basically doing none of the above, have to strike out a little bit in all these directions if they’re going to start to do what humans do. We’re pretty good at getting a little bit of provenance: “Oh, it’s old data. I discount that.” We get a little bit of context: “Oh, there’s a social environment here. I should just do the same thing. I should randomize.”

There’s some sampling uncertainty, and so on. We put all that together almost seamlessly. Then we do this in a social context, where if I don’t know how to get from here to the other side of town, I’ll ask someone who looks Danish. I know something about how to gather more data, and so on.

The poor LLM has none of the above. What should it say when you ask, “How sure are you?” All it’s doing, to the best of my knowledge, is this: in the past, someone asked a human on the internet, “How sure are you of that equation you just wrote down?” and someone said, “I’m very sure because of this or that.”

I think it just mimics those kinds of assertions, but that’s not reasoning under uncertainty.

Tim Scarfe

And if we did have epistemic uncertainty quantification, what would be the main uplift from that? Is it about, “I know I don’t know something, so I’m going to lean in and try to do more epistemic foraging in that area”?

Michael I. Jordan

Again, I think we’re now in statistics land. The statisticians are all about what species are present on the island. Have I sampled enough to know that there’s not a new species? These are classical areas of statistics.

Optimal experiment design—for that subpopulation, I don’t have enough data. I’m making a bad inference, and data collection in the context of inference, in the context of making assertions and doing that repeatedly—that’s what statistics has long focused on.

I think we should give them credit for handling a kind of active form of uncertainty reduction. But for me, uncertainty reduction in the large comes about from much broader sets of components, like a market.

I used an example in the paper where I want to have a restaurant like this for pizza, and I need tomatoes. If I had to forage for tomatoes every day, it would be pretty uncertain whether I would have pizza that evening. But because there exists a market where someone else did the foraging, there’s a stable amount of tomatoes every day.

I can build my restaurant assuming that’s true. My uncertainty about finding tomatoes went down. Therefore, I can build on top of that and do other things.

Markets mitigate uncertainty. They don’t do it because someone designed an optimal experiment or ran a multi-armed bandit—not directly, but because the market tried various things out. There are incentives for people to explore and exploit.

Tim Scarfe

Professor Jordan, it’s been an honor having you on the show. Thank you so much.

Michael I. Jordan

All right. It’s been my pleasure. I’ve enjoyed talking to you.