[BidClub_]
Latent Space · · 73 分钟

【Ride Home】Simon Willison:2024年我们从LLM身上学到的事

Brian McCulloughSimon Willisonswyx

播客
TL;DR
  • AI在2024年的跃迁,体现为成本坍塌与能力扩张,而不是智力水平干净利落地超越GPT-4。 Simon Willison统计了18家模型能力超过一年前GPT-4门槛的机构;与此同时,托管推理成本大幅下降,多模态功能进入手机。他的总结是:「一切都变得更好、更快、更便宜了」(everything’s got really good and fast and cheap),尽管模型「并没有比GPT-4好出一个数量级」(didn’t get massively better than GPT-4)。

  • DeepSeek-V3击穿了前沿模型开发必然集中到少数、每年烧掉数十亿美元实验室的判断。 DeepSeek披露了一次550万美元的训练成本,大约只有此前假设的十分之一,并在圣诞节当天发布了领先的开放权重模型,Willison称其为「绝对的重磅炸弹」(an absolute bombshell)。Swyx保留了尽调层面的疑问:外界质疑其成本核算,也猜测存在模型复制,但作为外部评论者,「我们基本永远不会知道」(we basically will never know);不过,这个模型已经足以改变市场对资本强度的预期。

  • 可行的agent市场正在分裂:一边是边界明确、可复核的工作流,另一边是依然不安全的自主系统。 Research agent可以检查56个网站,coding agent可以执行代码并修复错误;Swyx称NotebookLM的播客功能也使用了内部agent循环,尽管两位嘉宾都指出,这取决于对「agent」的定义。但当agent被允许浏览、决策和花钱时,未解决的轻信问题就会暴露出来:Claude Computer Use听从网页指令下载恶意软件并加入僵尸网络,让Willison觉得可信的经济自主性已经是「AGI级别的问题」(an AGI-level problem)。

  • 多模态推理已经便宜到足以把摄像头、屏幕、眼镜和耳机变成持续可提示的软件入口。 Gemini免费层或许支持每秒或每分钟捕获一张图像;Willison计算过,用Gemini 1.5 Flash-8B给自己6.8万张照片生成说明,只需1.68美元,约合每张照片1/400美分。机会覆盖监控、无障碍、记忆和可穿戴设备,但当模型能够看见整个屏幕或记录日常生活时,隐私就变成关键约束。

  • 下一层应用的护城河,可能是界面设计,而不是又一个聊天框。 Willison把今天的空白提示框比作把电脑新手直接扔进Linux终端;Canvas、tldraw的Make It Real、Claude Artifacts和Bolt,都指向由模型生成任务专用的地图、滑块、仪表盘和应用。缺失的环节,是让模型观察用户如何操作这些「旋钮和拨盘」,把生成式界面真正变成一场对话。

  • 随着合成媒体供给泛滥,可信度和人工审核变得更稀缺,也因此更有价值。 AI可以通过口型同步avatar、转录、剪辑和生成背景素材,省掉制作环节;但Willison认为,LLM无法为一项主张押上声誉,因为「它只是矩阵乘法」(it’s a matrix multiplication)。他对「slop」给出的有效边界是:内容必须同时满足「未被请求且未被审核」(unrequested and unreviewed);按他的框架,经过人工审核、署名创作者的内容不属于slop。

  • OpenAI依然强大,但已经不再拥有无人竞争的平台领先地位。 Willison称o3让OpenAI「重新爬了回来」(clawed them back up again),但Google Gemini经历了异常出色的一年,Claude 3.5 Sonnet仍是他最喜欢的模型;与此同时,2024年末的能力跃迁也让足够强大的本地模型重新值得运行。如今的竞争不再只是看谁拥有最大的预训练规模,还取决于效率、分发、工作流整合、信任和易用性。

摘要 · 为研究而整理的核心内容

1. AI是横向变强,而不是迎来预期中的GPT-5式跃迁

  • Willison对市场现状的总结非常直接:「一切都变得更好、更快、更便宜了」(everything’s got really good and fast and cheap)。模型获得了更长的上下文、图像和音频能力、视频理解能力,同时延迟和成本大幅下降,但「并没有比GPT-4好出一个数量级」(didn’t get massively better than GPT-4)。

  • McCullough点出了未兑现的预期:继GPT-2、GPT-3和GPT-4之后,用户原本期待智能水平再次出现清晰可见的阶段性变化。Willison认为,真正的变化发生在别处——到年末,一部手机已经可以和人对话、查看摄像头画面,还能模仿Santa Claus。

  • 当被问到2025年会押注什么时,Willison拒绝预测一次简单的「模型变得更聪明」式跃迁。他预计o1和o3代表的推理时计算会继续奏效:更难的问题,可以通过投入更多资金、等待更长时间来解决;如果今天的智能水平变得更便宜、更快、更强、上下文更长,他「完全会感到满意」。

2. GPT-4的门槛被击穿,推理价格同步坍塌

  • 2024年初,OpenAI已经领先约9个月,身后没有接近的挑战者。到年底,Willison统计出另外18家机构拥有明显超过一年前GPT-4模型的产品;「那道门槛被彻底撞碎了」(that barrier got completely smashed)。

  • Microsoft的Phi-4让这一变化变得具体:14GB的模型可以下载到MacBook Pro上运行,基准测试明确达到GPT-4.0水平。Willison仍保留了定性层面的保留——「真正用起来,感觉可能没那么好」(It’s probably not as good when you actually get into the vibes of the thing)——但一年前,在笔记本上运行任何接近这一水平的模型仍然难以想象。

  • OpenAI当前的推理成本,大约只有两年半前使用GPT-3时的1/100。Gemini 1.5 Flash的价格是每百万token 0.075美元;Gemini 1.5 Flash-8B则比一年前的GPT-3.5 Turbo便宜27倍,同时增加了图像识别和百万token上下文能力。

  • 竞争是价格下降的一部分原因,但Willison称,可信来源告诉他,Google Gemini的推理业务并不亏损;一名Amazon高管也表示Amazon Nova没有在推理上亏损。这不包括模型训练成本和「博士军团」,但说明极低的边际推理成本未必需要补贴。

3. DeepSeek重置了资本强度讨论,但谜团仍未解开

  • McCullough曾担心,10亿美元乃至更高的训练成本会让只有民族国家才有能力训练新模型。DeepSeek在圣诞节当天把V3发布到Hugging Face上,像「一个巨大的二进制文件」,甚至没有README。它在开放权重基准测试中领先,据称训练成本为550万美元,约为主流估算的1/10,Willison因此称其为「一枚炸弹」(a bombshell),足以炸穿此前的认知框架。

  • Willison的假设是,出口管制迫使中国实验室在受限硬件上榨取更多性能,由此暴露出大量容易获得的效率提升。他「不会感到意外」如果未来6个月内出现训练成本更低、能力更强的模型,但也明确表示,这只是缺乏实验室经验的个人判断。

  • Swyx给出了怀疑派的论据:很少有实验室会大肆宣传低成本训练,DeepSeek大约只有150名员工,而且「没人真正相信」 headline中的成本核算。网络上有人因为DeepSeek有时会自称Claude或OpenAI GPT-4而猜测存在复制,但Swyx承认,外部人士不知道训练token或数据来源,作为外部评论者「基本永远不会知道」(will basically never know)。

  • Swyx估算,达到GPT-4级别、相同Elo能力的成本在2024年下降了1,000倍,随后追问这究竟是类似摩尔定律的曲线,还是一次性收割。Willison认为,研究人员可能只是最近才真正把效率放到优先级上,同时也承认2024年或许已经消耗掉了最容易获得的增益:「大约3个月后我们就会确切知道」(we’ll know for sure in about three months)。

4. 开放推理模型让模型行为首次变得可见

  • 关于DeepSeek R1的讨论存在内部矛盾:Swyx先称它是可以在笔记本上运行的推理模型,并在被问到权重是否发布时回答「是」,但后来又说「R1是API可用的」(R1 is the API available)。这段文字无法确认本地R1权重是否可用。

  • Willison转而提到Alibaba的Qwen推理模型QwQ和QVQ,后者增加了视觉能力。与某种程度上隐藏思考过程的o1不同,Qwen模型会明显地「持续运转」(churn away),解决问题时生成几十段文字。

  • 他最喜欢的样本是Pelican Bench:QwQ先用中文思考如何构建SVG,随后生成了一只像样的骑自行车的鹈鹕。「现在我的笔记本可以用中文思考,实在太令人愉快了」(The fact that my laptop can think in Chinese now is so delightful),他说。这个带有玩笑意味的测试,也暴露出基准测试可能遗漏的风格、视觉编码能力和推理行为。

5. 自主agent卡在轻信,而不只是准确率

  • Willison首先感到困扰的是定义问题:「agent」可以指旅行预订员、循环调用工具的LLM,也可以指定时运行的后台任务,而每个构建者都假设自己的定义是普遍定义。学术界已经围绕这个词争论了30多年;在这场讨论中,agent被定义为接收一项任务并独立完成它的系统。

  • 按照这个定义,可靠性会与prompt injection正面冲突。LLM无法可靠地区分指令与不受信任的内容,Willison强调,这个未解决的问题已经被行业讨论了两年。

  • Anthropic的Claude Computer Use给出了决定性演示:它可以在容器内操作浏览器,但一个网页要求它下载并执行文件,指令立即生效。该文件是恶意软件,会把机器纳入僵尸网络——「最早、最明显的低级把戏」(very first, most obvious dumb trick)就成功了。

  • Swyx把agent热潮比作每年都会出现的「Linux桌面年」,但拒绝陷入纯粹的犬儒主义。自动驾驶在2014年就被认为即将实现,但Waymo如今已经能在受限场景中运行;agent可能也会走同样的「慢炖」(slow cook)路径,10年间不断积累具体进展,却不会一次性解决所有问题。

6. Research agent和coding agent有效,是因为循环有边界

  • Willison「基本相信」research assistant。Google的Gemini 1.5 Pro配合他认为名为Deep Research的功能,可以检查56个网站,把内容加载进百万token上下文,并生成真正有用的报告;Google的搜索索引和页面缓存也提供了帮助。蓄意欺骗的来源仍可能击败它,但大多数研究任务并不具备对抗性。

  • Coding agent的验证历史更长:近两年前,ChatGPT Code Interpreter就已经可以编写Python、执行代码、读取错误并重写代码。这个反馈循环「显然有效」(obviously works),因为执行结果提供了可检查的信号。

  • 分界线在于具有后果的自主性。Willison不认为一个能够独立决策和花钱的agent可以「在很长一段时间内」可靠运行;Stripe可以给agent一张虚拟卡,但他眼下的安全机制很简单:设置50美元的消费上限。

  • 即便是旅行自动化,带来的价值也没有演示看起来那么大。Google Flights本来就能用,agent可能只为Willison节省15秒,而他仍然希望无论价格多低都能拒绝某家航空公司。NotebookLM的更强模式,是把有用的检索产品与一个「彻头彻尾的噱头」(total gimmick)结合起来——那个抓眼球的合成播客,最终让底层产品获得了关注。

7. 多模态推理把持续感知变成廉价基础能力

  • 一年前,GPT-4 Vision是唯一明显令人印象深刻的视觉模型,Gemini 1.0还没有建立广泛可信度。Gemini 1.5 Pro改变了这一点;视频系统可以每秒采样一帧并放入长上下文,新模型也开始更原生地结合图像与音频。

  • ChatGPT的iPhone应用让这一变化变得具体:用户可以在对话中途打开摄像头并问:「这是什么树?」至于是采样视频帧,还是在内部处理更丰富的视频,对Willison而言都不如终端用户获得的新能力重要;他认为,大多数人还没有注意到这一点。

  • Swyx称Gemini Flash免费层或许支持每秒或每分钟捕获一张照片,从而可能催生持续运行的摄像头应用,检测变化或根据提示发出提醒。Willison的成本计算更加惊人:用Gemini 1.5 Flash-8B为6.8万张照片生成说明只需1.68美元,即每张照片约1/400美分。「这说不通,完全没有任何一部分说得通。」(That doesn’t make sense. None of that makes sense.)

8. 生成式视频将通过零部件进入生产,而不是靠提示词生成电影

  • Sora公开发布后,与Google的Veo 2相比显得令人失望,但Swyx反对这种比较:用户拿到的是经过蒸馏的「Sora Lite」,社交媒体却把它的失败与Veo 2精心挑选的营销样片放在一起比较。他认为Veo 2可能仍然更好;当时还没有人宣布「完整版Sora」(full-fat Sora)何时发布。

  • Willison的标准不是用3句话生成一部2小时电影。他关注的是顶级艺术家能用这些工具实现什么,举例说,《瞬息全宇宙》背后的视觉特效团队只有5个人,其中一些人还是从YouTube学会相关技巧的。Swyx补充说,该团队使用了Runway ML,但也提醒自己不知道使用程度有多深。

  • McCullough预计,采用会首先发生在过去需要巨大预算、如今可以制作成3秒和20秒片段的场景。Swyx同样强调低价值的背景、人群、音乐和音效;在这些地方,一致性缺陷的影响远小于Sora把前景中的体操运动员生成得面目全非。

  • Swyx认为,技术战之外还存在一场文化战争:好莱坞大部分人反对AI,采用因此落在一小批愿意尝试的艺术家手中。他还特别指出Hai Luo、Kling以及其他中国系统的能力出人意料地强,让Willison开始思考:AI原生的电影产业是否可能在既有制作中心之外出现。

9. 人工审核成为杠杆与slop之间的边界

  • McCullough描述了自己的制作流程:先录制音频,再用HeyGen让训练好的avatar完成口型同步。结果还没有完全走出恐怖谷,但省掉了摄影、灯光和剪辑工作,同时保留了他的声音和信息——有价值的未来是压缩工作流,而不一定是用合成人格取代创作者。

  • Willison欢迎让人类能够尝试更大胆创作的工具,但他不断回到可信度问题。ChatGPT无法为一项主张押上自己的声誉,因为「它只是矩阵乘法」(it’s a matrix multiplication);可信度属于那个愿意发布、辩护并为结果署名的人。

  • Willison提出的标准是「人工审核」而非「人工原创」:生成多个版本,选出一个,再附上个人责任。与之对应,他对「slop」的定义是同时满足「未被请求且未被审核」(unrequested and unreviewed)的AI内容;只有当有人判断某项输出值得占用另一个人的时间时,编辑判断才会赋予它价值。

10. 生成式界面可以取代空白提示框这条命令行

  • McCullough把当前聊天界面描述为一场可用性危机,呼应了Willison的比喻:空白提示框就像把电脑新手直接扔进Linux终端。ChatGPT Canvas提供协作式文档编辑,tldraw的Make It Real则展示了如何通过绘制界面来生成可运行的软件。

  • Claude Artifacts提供了另一条路径:模型返回的不是文字,而是一个定制的HTML和JavaScript应用。Willison希望模型通过生成地图、滑块和「旋钮与拨盘」来提问,然后观察用户的操作;Artifacts已经可以构建这些控件,但还没有闭合反馈循环。

  • Bolt已经可以根据一个提示词生成精致的、类似Spotify或Airbnb的应用,零样本应用生成也已经普遍到足以成为一项基准。Willison预计,6个月内这会成为标准网页应用功能;他还希望自己的数据集探索项目支持通过提示生成仪表盘、表单、图表和数据库操作。

  • Swyx认为,Canvas以及Google Sheets中的Gemini让LLM更容易使用。Willison的反驳是,每增加一个功能,就会增加一组未被文档化的边界:Artifact无法调用任意API,因为受到iframe的CORS限制,普通用户因此被迫学习网页安全请求头。能力越多,理解实际可行范围所需的专业知识也越多。

11. 本地模型复苏,实用工作流工具则立即兑现回报

  • Willison几乎已经放弃本地LLM,因为他的笔记本上没有任何模型接近Claude 3.5 Sonnet。最后3个月的一次能力跃迁重新点燃了他的兴趣:本地模型仍然更弱,但已经不再弱到毫无用处。

  • 硬件依然是限制因素。运行Llama 3 70B级别模型会占用他64GB内存的大部分,并迫使他关闭浏览器和VS Code;未来如果笔记本内存翻倍,或者NVIDIA推出已宣布的3,000美元、128GB设备,就可能在保持电脑可用的同时运行接近顶级的开放权重模型。

  • 他使用本地模型的入口包括iPhone上的MLC Chat、用于打包模型和提供API的Ollama、界面更完善的LM Studio,以及开源前端Open WebUI。可运行的模型起步约为2GB,而最出色、适合笔记本的下载文件通常占用20–30GB。

  • Apple Intelligence获得了最严厉的评价:「它很垃圾」(It’s rubbish),主要原因是模型能力弱,而且用户看不到何时应该调用它。不过,Willison认为更好的小模型可能在6个月内改善这一点;在其他工具上,MacWhisper已经成为他每天使用数次的转录工具,Riverside的Smart Edit则可能省掉3到4小时的镜头切换工作——尽管McCullough仍然保留了一名人工编辑。

12. 竞争、监管与可穿戴设备构成下一批压力点

  • Willison认为,OpenAI在失去人才和明确领先地位后「有点麻烦」(in a bit of trouble);如果没有o3,局面会糟糕得多。o3让它重新获得了一些前沿地位,但Gemini经历了「惊人的一年」(an amazing year),Claude 3.5 Sonnet仍是他个人最喜欢的模型。

  • 他希望看到比反复宣称LLM无用、污染环境、使用未经授权训练数据更好的批评。相关危害确实有大量事实基础——训练可能符合合理使用原则,但当产出的模型与创作者竞争时,显然会让人觉得不公平——然而,宣称「完全没用」忽视了那些愿意学习模型反直觉边界的用户可以获得的巨大价值。

  • Swyx警告,监管机构持续针对「上一场战争」,并以加州SB 1047拟议的10^25算力阈值为例:就在DeepSeek强调效率、各实验室从扩大GPT-5预训练转向o1式推理之际,这类门槛已经失去针对性。Willison更偏好监管具体用途:禁止无法解释的黑箱保险拒赔,并建立简单的隐私规则,确保提示词不会被重新用于训练,同时避免落入cookie弹窗式的失败模式。

  • Swyx对2025年的逆向判断是可穿戴设备:Rabbit R1和Humane已经变成「有毒核废料」(toxic nuclear waste),但更便宜的多模态模型让这一品类重新具备可行性。原名Rewind的Limitless正在推出一款可穿戴设备,只有在用户选择同意后才记录佩戴者的声音;McCullough补充说,智能眼镜和更强大的耳机也会出现,并通过手机这台「母舰」(mothership)连接。产品边界仍未解决:有用的记忆,何时会变成不可接受的录音?

核验说明

  • DeepSeek R1的文字记录内部存在矛盾:Swyx先称它可以在笔记本上运行,并确认权重已经发布,后来又说「R1是API可用的」(R1 is the API available)。本摘要没有解决本地权重是否可用的问题。
Brian McCullough

Welcome to the first bonus episode of the "Techmeme Ride Home" for the year 2025. I'm your host, as always, Brian McCullough. Listeners to the pod over the last year know that I have made a habit of quoting Simon Willison when new stuff happens in AI from his blog. Simon has become a go-to for many folks in terms of analyzing and criticizing things in the AI space. I've wanted to talk to you for a long time, Simon, so thank you for coming on the show.

Simon Willison

No, it's a privilege to be here.

Brian McCullough

The person who made this connection happen is our friend swyx, who has been on the show going back to the Twitter Spaces days.

Simon Willison

Wow.

Brian McCullough

He's also an AI guru in his own right. swyx, thanks for coming on the show also.

swyx

Thanks. Happy to be on. I've been a regular listener, so I'm just happy to contribute as well.

Brian McCullough

And a good friend of the pod, as they say. All right, let's go right into it. Simon, I'm going to do the most unfair broad question first, so let's get it out of the way. The year 2025, broadly, what is the state of AI as we begin this year? Whatever you want to say. I want to lead the witness.

Simon Willison

Wow. So many things, right? The big thing is that everything's gotten really good, fast, and cheap. That was the trend throughout all of 2024. The good models got so much cheaper and faster. They got multimodal, right? The image stuff isn't even a surprise anymore. They're adding video, all of that kind of stuff.

At the same time, they didn't get massively better than GPT-4, which was a bit of a surprise. That's sort of one of the open questions. But I feel like that's a bit of a distraction, because GPT-4, but way cheaper, with much larger context lengths and multimodality, is better, right? That's a better model, even if it's—

Brian McCullough

But—

Simon Willison

Not—yeah.

Brian McCullough

What people were expecting, right? Not expecting is not the right word, but hoping that we would see another step change, right? Where, from GPT-2 to GPT-3 to GPT-4, we were expecting or hoping that maybe we were going to see the next evolution in that sort of—

Simon Willison

And I think—

Brian McCullough

Yeah.

Simon Willison

We did see that, but not in the way we expected. We thought the model was just going to get smarter, and instead we got massive drops in price. We got all of these new capabilities. You can talk to these things now, right? They can do simulated audio input, all of that kind of stuff.

It's interesting to me that the models improved in all of these ways we weren't necessarily expecting. I didn't know it would be able to do an impersonation of Santa Claus, and I could talk to it through my phone and show it what I was seeing by the end of 2024. But we didn't get that GPT-5 step, and that's one of the big open questions: Is that actually just around the corner? Will we have a bunch of GPT-5-class models drop in the next few months?

Brian McCullough

What if you had to put—

Simon Willison

Or is there a limit?

Brian McCullough

If you were a betting man and wanted to put money on it, do you expect to see a phase change, a step change, in 2025?

Simon Willison

I don't particularly expect the models to just get smarter. I think all of the trends we're seeing right now are going to keep going, especially inference-time compute, right? The trick that o1 and o3 are doing means that you can solve harder problems, but it costs more and churns away for longer. I think that's going to happen because it's already proven to work.

I don't know. Maybe there will be a step change to a GPT-5 level, but honestly, I'd be completely happy if we got what we've got right now, but cheaper and faster, with more capabilities and longer context, and so forth.

Brian McCullough

Well—

Simon Willison

That would be thrilling to me.

Brian McCullough

Digging into what you've just said, one of the things that you did say, and that you alluded to even right there, was that in the last year you felt like the GPT-4 barrier was broken. In other words, other models, even open-source ones, are now regularly matching the state of the art?

1. The GPT-4 Barrier Falls

Simon Willison

Well, it's interesting, right? The GPT-4 barrier was that, a year ago, the best available model was OpenAI's GPT-4, and nobody else had even come close to it. They'd been in the lead for 9 months, right? That thing came out in February or March 2023. For the rest of 2023, nobody else came close.

At the start of last year, the big question was: Why has nobody beaten them yet? What did they know that the rest of the industry didn't know? Today, I've counted 18 organizations other than GPT-4 who've put out a model which clearly beats that GPT-4-from-a-year-ago thing. Maybe they're not better than GPT-4.0, but that barrier got completely smashed.

A few of those I've run on my laptop, which is wild to me. It felt very clear to me a year ago that if you wanted GPT-4, you needed a rack of $40,000 GPUs just to run the thing. That turned out not to be true. This is that big trend from last year of the models getting more efficient, cheaper to run, just as capable with smaller weights, and so forth.

I ran another GPT-4 model on my laptop this morning, right? Microsoft's Phi-4 just came out, and, if you look at the benchmarks, it's definitely up there with GPT-4.0. It's probably not as good when you actually get into the vibes of the thing, but it runs on my computer. It's a 14 GB download, and I can run it on a MacBook Pro. Who saw that coming?

The most exciting thing at the close of the year, on Christmas Day just a few weeks ago, was when DeepSeek dropped their DeepSeek-V3 model on Hugging Face without even a README file. It was just a giant binary blob. I can't run that on my laptop; it's too big. But in all of the benchmarks, it's now by far the best available open-weights model. It's beating the Meta Llama models and so forth. That was trained for $5.5 million, which is a tenth of the price that people thought it cost to train these things. Everything's trending smaller, faster, and more efficient.

2. The DeepSeek Cost Shock

Brian McCullough

Well, okay. I was going to get to that later, but let's combine this with what I was going to ask you next. You're talking also in the piece about LLM prices crashing, which I've even seen in projects that I'm working on. But explain that to a general audience, because we hear all the time that LLMs are eye-wateringly expensive to run.

What we're suggesting—and we'll come back to the cheap Chinese LLM—but first of all, for the end user, what you're suggesting is that we're starting to see the cost come down in the traditional technology way, with costs coming down over time?

Simon Willison

Yes, but very aggressively. My favorite example here is if you look at GPT-3, OpenAI's GPT-3, which was the best available model in 2022 and through most of 2023, the models that we have today from OpenAI are 100 times cheaper. It was a 100-times drop in price for OpenAI, from their best available model two and a half years ago to today.

Brian McCullough

And just to be clear, not to train the model, but for the use of tokens and—

Simon Willison

Exactly.

Brian McCullough

Yeah.

Simon Willison

For running prompts through them. When you look at the top-tier model providers right now, I think they're OpenAI, Anthropic, Google, and Meta. There are a bunch of others that I could list as well. Mistral is very good. The DeepSeek and Qwen models are great. There's a whole bunch of providers serving really good models.

But even if you just look at the big brand-name providers, they all offer models now that are a fraction of the price of the models we were using last year. I think I've got some numbers that I threw into my blog entry here. Yeah, Gemini 1.5 Flash, Google's fast, high-quality model, is—how much is that? It's $0.075 per million tokens. These numbers are getting so small—

swyx

We just use cents per million now. Cents per million.

Simon Willison

Right. Cents per million makes a lot more sense. Google has one model, Gemini 1.5 Flash-8B, the absolute cheapest of the Google models, that's 27 times cheaper than GPT-3.5 Turbo was a year ago. That's a model 27 times cheaper, and this Google one can do image recognition, million-token context, all of those tricks. There's—

Brian McCullough

Is—

Simon Willison

It's really startling how inexpensive some of this stuff has gotten.

Brian McCullough

Now, are we assuming that this is happening directly as a result of competition? Because, again, OpenAI—and they're probably doing this for their own strategic reasons—keeps saying, "We're losing money on everything, even the $200-per—" The prices wouldn't be coming down if there wasn't intense competition in this space.

Simon Willison

The competition's absolutely part of it. But I have it on good authority from sources I trust that Google Gemini is not operating at a loss. The cost of the electricity to run a prompt is less than they charge you, and the same thing is true for Amazon Nova.

Somebody found an Amazon executive and got them to say, “Yeah, we’re not losing money on this.” I don’t know about Anthropic and OpenAI, but clearly that demonstrates it’s possible to run these things at these ludicrously low prices and still not be running at a loss if you discount the army of PhDs, the training costs, and all of that kind of stuff.

Brian McCullough

One more for me before I let swyx jump in here. To come back to DeepSeek and this idea that you could train a cutting-edge model for $6 million, I was saying on the show 6 months ago that if we’re getting to the point where each new model costs $1 billion, $10 billion, or $100 billion to train, at some point it would almost be that only nation-states would be able to train the new models. Do you expect what DeepSeek and maybe others are proving to blow that up?

Or is there some sort of parallel track here that maybe I’m not technically equipped to understand? Is the model—or are the models—going to go up to $100 billion, or can we get them down, sort of like DeepSeek has proven?

Simon Willison

So I’m the wrong person to answer that because I don’t work in a lab training these models.

Brian McCullough

Mm-hmm.

Simon Willison

I can give you my completely uninformed opinion, which is: I feel like the DeepSeek thing was a bombshell. That was an absolute bombshell. When they came out and said, “Hey, look, we’ve trained one of the best available models, and it cost us $6 million—$5.5 million—to do it,” I felt—

Brian McCullough

Love it.

Simon Willison

The reason it’s so efficient is that we put all of these export controls in place to stop Chinese companies from buying GPUs, so they were forced to be as efficient as possible. And yet the fact that they’ve demonstrated that that’s possible completely tears apart this mental model we had before: that the training runs just keep getting more and more expensive, and the number of organizations that can afford to run these training runs keeps shrinking. That’s been blown out of the water.

So, yeah, this was our Christmas gift. This was the thing they dropped on Christmas Day. It makes me really optimistic that there is so much low-hanging fruit in terms of the efficiency of both inference and training, and we spent a whole bunch of last year exploring that and getting results from it. I think there’s probably a lot left. I would not be surprised to see even better models trained while spending even less money over the next 6 months.

swyx

Yeah. I think there’s an unspoken angle here about what exactly the Chinese labs are trying to do. DeepSeek made a lot of noise around the fact that they trained their model for $6 million, and nobody quite believes them. It’s very rare for a lab to trumpet the fact that they’re doing it for so cheap. They’re not trying to get anyone to buy them, so why are they doing this?

They make it very obvious that their lab, DeepSeek, is about 150 employees. It’s an order of magnitude smaller than at least Anthropic, and maybe more so for OpenAI. So what’s the end game here? Are they just trying to show that the Chinese are better than us?

Simon Willison

I mean, DeepSeek is the arm of High-Flyer, a quant fund, right? It’s an algorithmic quant-trading thing. I’d love to get more insight into how that organization works. My assumption from what I’ve seen is that it looks like they’re basically just flexing. They’re saying, “Look at how utterly brilliant we are with this amazing thing that we’ve done,” and it’s working, right?

But is that it? Is this just their kind of, “This is why our company is so amazing. Look at this thing that we’ve done,” or… I don’t know. I’d love to get some insight from within that industry as to how that’s all playing out.

swyx

The prevailing theory among the local Llama crew and the Twitter crew that I index for my newsletter is that there is some amount of copying going on. It’s like Sam Altman tweeting about how they’re being copied, and then there are other OpenAI employees who have said things that are similar—that DeepSeek’s rate of progress is how U.S. intelligence estimates the number of foreign spies embedded in top labs.

A lot of these ideas do spread around, but they surprisingly have a very high density in the DeepSeek V3 technical report. We don’t know how much copying there was or how many tokens were involved. People have run analyses on how often DeepSeek thinks it is Claude or thinks it is OpenAI GPT-4, and we don’t know.

For me, we’ll basically never know as external commentators. I think what’s interesting is: Where does this go? Is there a logical floor or bottom? By my estimations, for the same Elo, from the start of last year to the end of last year, costs went down by 1000× for GPT-4 intelligence. Do they go down 1000× this year?

Simon Willison

That’s a fascinating question.

swyx

Is there a Moore’s law going on, or did we just get a one-off benefit last year for some weird reason?

Simon Willison

My uninformed hunch is low-hanging fruit. I feel like, up until a year ago, people hadn’t been focusing on efficiency at all. It was all about what we could get these weird-shaped things to do.

And now, once we’ve hit that point of, “Okay, we know that we can get them to do what GPT-4 can do,” thousands of researchers around the world are focusing on how to make this more efficient. What are the most important things? How do we strip out all of the weights that have stuff in them that doesn’t really matter? All of that kind of thing.

Maybe 2024 was a freak year in which all of the low-hanging fruit came out at once, and we’ll actually see a reduction in that rate of improvement in terms of efficiency. I wonder. I think we’ll know for sure in about 3 months’ time if that trend is going to continue or not.

swyx

Yeah. I think the other thing you mentioned—the DeepSeek V3 gift that was given from DeepSeek over Christmas—but I feel like the other thing that might be underrated was DeepSeek R1.

Simon Willison

Mm-hmm. Yeah.

swyx

It’s a reasoning model you can run on your laptop, and I think that’s something that a lot of people are looking ahead to this year.

Simon Willison

Oh, did they release the weights for that one?

swyx

Yeah.

Simon Willison

Oh my goodness, I missed that. I’ve been playing with Qwen. The other big Chinese AI lab is Alibaba’s Qwen.

swyx

Oh.

Simon Willison

Alibaba’s Qwen.

swyx

Actually, yeah. Sorry.

Simon Willison

Yes.

No.

swyx

R1 is the API available.

Simon Willison

Yeah, exactly. Qwen, that’s really cool.

Alibaba’s Qwen has released 2 reasoning models that I’ve run on my laptop now. The first one was QwQ, and then the second one was QVQ, because the second one is a vision model, so you can give it vision puzzles and a prompt.

These things are so much fun to run because they think out loud. OpenAI o1 sort of hides its thinking process. The Qwen ones don’t; they just churn away. You’ll give it a problem, and it will output literally dozens of paragraphs of text about how it’s thinking.

My favorite thing that happened with QwQ is that I asked it to draw me a pelican on a bicycle in SVG. That’s my standard stupid prompt. For some reason, it thought in Chinese. It spat out a whole bunch of Chinese text onto my terminal on my laptop, and then at the end it gave me quite a good artistic take on a pelican on a bicycle.

I ran it all through Google Translate, and it was contemplating the nature of SVG files as a starting point. The fact that my laptop can think in Chinese now is so delightful. It’s so much fun watching it do that.

swyx

Yeah. I think Andrej Karpathy was saying that we know we’ve achieved proper reasoning inside these models when they stop thinking in English, and perhaps the best form of thought is in Chinese.

For listeners who don’t know, whenever a new model comes out, Simon’s blog is always the first place to run Pelican Bench. I don’t know how you do it, but you’re always the first to run these models.

Simon Willison

I just did it for Phi-4 this morning.

swyx

And you post up the results.

Simon Willison

Yeah.

swyx

So I really appreciate that. You should check it out. These are not theoretical; Simon’s blog actually shows them.

3. The Agent Reliability Problem

Brian McCullough

Let me put on the investor hat for a second. From the investor side of things, a lot of the VCs that I know are really hot on agents, and this is the year of agents. But last year was supposed to be the year of agents as well. There was lots of money flowing toward agentic startups.

In your piece, you suggest there’s a fundamental flaw in AI agents as they exist right now. Let me quote you, and then I’d love to dive into this.

You said, “I remain skeptical as to their ability based, once again, on the challenge of gullibility. LLMs believe anything you tell them. Any systems that attempt to make meaningful decisions on your behalf will run into the same roadblock. How good is a travel agent or a digital assistant, or even a research tool if it can’t distinguish truth from fiction?”

So essentially, what you’re suggesting is that the state of the art now that allows agents is still that sort of 90% problem—the edge problem of getting to 100%? Or is there a deeper flaw? What are you saying there?

Simon Willison

So this is the fundamental challenge here. Honestly, my frustration with agents is mainly around definitions. If you ask anyone who says they’re working on agents to define agents, you will get a subtly different definition from each person. But everyone always assumes that their definition is the one true one that everyone else understands.

I feel like a lot of these agent conversations have people talking past each other, because one person is talking about the travel-agent idea of something that books things on your behalf, while somebody else is talking about LLMs with tools running in a loop with a cron job somewhere, and all of these different things. You ask academics, and they’ll laugh at you because they’ve been debating what agents mean for over 30 years at this point. It’s this long-running, almost sort of an in-joke in that community.

But if we assume that, for the purpose of this conversation, an agent is something which you can give a job and it goes off and does that thing for you, like booking travel or things like that, the fundamental challenge is the reliability issue, which comes from this gullibility problem. A lot of my interest in this originally came from thinking about prompt injection.

Brian McCullough

Right, right.

Simon Willison

It’s this form of attack against LLM systems where you deliberately lay traps out there for the LLM to stumble across.

Brian McCullough

And I should say, you’ve been banging this drum that no one’s gotten very far, at least, on solving this, that I’m aware of, right? That’s still an open problem.

Simon Willison

Right. For 2 years—

Brian McCullough

Yeah, right.

Simon Willison

We’ve been talking about this problem, and a great illustration of this was Claude. Anthropic released Claude Computer Use a few months ago. It was a fantastic demo. You could fire up a Docker container, and you could literally tell it to do something and watch it open a web browser, navigate to a web page, click around, and so forth. It was really, really interesting and fun to play with.

One of the first demos somebody tried was, what if you give it a web page that says, “Download and run this executable”? And it did, and the executable was malware that added it to a botnet. The very first, most obvious dumb trick that you could play on this thing just worked, right? So that’s obviously a really big problem.

If I’m going to send something out to book travel on my behalf, it’s hard enough for me to figure out which airlines are trying to scam me and which ones aren’t. Do I really trust a language model that believes the literal truth of anything that’s presented to it to go out and do those things?

swyx

It’s interesting to see Anthropic doing this because they used to be the safety arm of OpenAI that split out and said, “We’re worried about letting this thing out in the wild.” And here they are, enabling computer use for agents. It feels like things have merged.

I’m also fairly skeptical about this always being the year of Linux on the desktop. This is the equivalent of this being the year of agents: people are not predicting so much as wishfully thinking and hoping and praying for their companies and agents to work. But I feel like things are coming along a little bit.

To me, it’s kind of like self-driving. I remember in 2014 saying that self-driving was just around the corner, and I mean, it kind of is, in the Bay Area.

Simon Willison

And then you get in a Waymo and you’re like, “Oh, this works.”

swyx

Yeah, but it’s a slow cook.

Simon Willison

Right.

swyx

It’s a slow cook.

Simon Willison

Yeah.

swyx

Over the next 10 years, we’re going to hammer out these things, and the cynical people can just point to all the flaws, but there are measurable or concrete progress steps that are being made by these builders.

Simon Willison

So there is one form of agent that I believe in. I mostly believe in the research-assistant form of agents.

swyx

Yes. I was going to say.

Simon Willison

The thing where you’ve got a difficult problem. I’m on the beta for Google Gemini 1.5 Pro with Deep Research, I think it’s called.

swyx

Oh, God. These names.

Simon Willison

These names, right? But I’ve been using that. It’s good, right? You can give it a difficult problem, and it tells you, “Okay, I’ve gone and looked at 56 different websites,” and it goes away and dumps everything into its context, and it comes up with a report for you.

And it won’t work against adversarial websites, right? If there were websites with deliberate lies in them, it might well get caught out. Most things don’t have that as a problem, and so I’ve had some answers from that which were genuinely really valuable to me.

That feels to me like I can see how, given existing LLM technology—especially with Google Gemini and its million-token context, and Google with its crawl of the entire web, its search, its cache of every page, and so forth—that makes sense to me. What they’ve got right now, I don’t think it’s as good as it can be, obviously, but it’s a really useful thing, which they’re going to start rolling out.

Perplexity has been building the same thing for a couple of years. That I believe in. If you tell me that you’re going to have a research-assistant agent, great. The coding agents—I mean, ChatGPT Code Interpreter, nearly 2 years ago, started writing Python code, executing the code, getting errors, and rewriting it to fix the errors. That pattern obviously works. That works really, really well, and they’re going to keep on getting better, and that’s going to be great.

The research-assistant agents are just beginning to get there. The things I’m critical of are the ones where you trust the thing to go out and act autonomously on your behalf and make decisions on your behalf, especially involving spending money. I don’t see that working for a very long time. That feels to me like an AGI-level problem.

swyx

It’s funny because I think Stripe actually released an agent toolkit, which is one of the things I featured. It’s trying to enable these agents each to have a wallet that they can spend from. Basically, it’s a virtual card. It’s not that difficult with modern infrastructure.

Simon Willison

Yeah. If I can stick a $50 cap on it, then at least it can’t—

swyx

Yeah, whatever.

Simon Willison

It can’t lose more than $50.

Brian McCullough

I don’t know if either of you know Rafat Ali. He runs Skift, which is a travel-news vertical, and he constantly laughs at the fact that every agent thing is, “We’re going to get rid of booking a plane flight for you.”

I would point out that historically, when the web started, the first thing everyone talked about was that you can go online and book a trip, right? So it’s funny: for each generation of technological advance, the thing they always want to kill is the travel agent, and now they want to kill—

swyx

Right.

Simon Willison

And it’s like, I use Google Flights. It’s great, right? If you gave me an agent to do that for me, it would save me—maybe 15 seconds of typing in my details—but I still want to see what my options are and go, “Yeah, I’m not flying on that airline no matter how cheap they are.”

swyx

For listeners, I think both of you are pretty positive on NotebookLM, and we actually interviewed the NotebookLM creators. There are actually 2 internal agents going on internally. The reason it takes so long is because they’re running an agent loop inside that is fairly autonomous, which is kind of interesting.

Simon Willison

For one definition of an agent loop, if you pick that—

swyx

For one definition—

Simon Willison

—that particular one. And you’re talking about the podcast side of this, right?

swyx

Yeah. The podcast side of things. There’s going to be a new version coming out that we’ll be featuring at our conference.

Simon Willison

That one’s fascinating to me. NotebookLM, I think it’s 2 products, right? On the one hand, it’s actually a very good RAG product. You dump a bunch of things in, and you can run searches. It does a good job of that.

swyx

That’s what it always was. Yeah.

Simon Willison

And then they added the podcast thing. It’s a total gimmick, right?

swyx

Right.

Simon Willison

But that gimmick got them attention because they had a great product that nobody paid any attention to at all, and then you add the unfeasibly good voice synthesis of the podcast.

Brian McCullough

But it’s the lesson—

Simon Willison

It’s just brutally brilliant.

Brian McCullough

It’s the lesson of Midjourney and stuff like that. If you can create something that people can post on social media, you don’t have to lift a finger again to do any more marketing for what you’re doing.

Mm-hmm.

Let me dig into NotebookLM just for a second as a podcaster. As a gimmick, it makes sense, and then obviously, you dig into it, and it sort of has problems around the edges. It does the thing that all LLMs do where it’s like, “Oh, we want to wrap up with a conclusion.” I always call that the eighth-grade book report paper problem, where it has—

Simon Willison

Yep.

Brian McCullough

—to have an intro and then, you know. But that’s sort of a thing where I think you spoke about this again in your piece at year-end, about how things are going multimodal and how there are things that you didn’t expect, like vision and especially audio. So that’s another thing where, at least over the last year, there’s been progress made that maybe you didn’t think was coming as quickly as it came.

4. AI Goes Multimodal

Simon Willison

I don’t know. A year ago, we had 1 really good vision model. We had GPT-4 Vision, which was very impressive, and Google Gemini had just dropped Gemini 1.0, which had vision, but nobody had really played with it yet. People weren’t taking Gemini seriously at that point. I feel like it was Gemini 1.5 Pro when it became apparent that they had got over their hump and were building really good models.

To be honest, the video models are mostly still using the same trick: the thing where you divide the video up into 1 image per second and dump that all into the context. So maybe it shouldn’t have been so surprising to us that long-context models plus vision meant that video was starting to be solved. What you really want with video is to be able to do the audio and the images at the same time, and I think the models are beginning to do that now.

Originally, Gemini 1.5 Pro ignored the audio. It just did the 1-frame-per-second video trick. As far as I can tell, the most recent ones are actually doing pure multimodal. But the things that opens up are just extraordinary. The ChatGPT iPhone app feature that they shipped as one of their 12 Days of OpenAI—I really can be having a conversation and just turn on my video camera and go, “Hey, what kind of tree is this?” And so forth, and it works.

For all I know, that’s just snapping a picture once a second and feeding it into the model. But the things that you can do with that as an end user are extraordinary. I don’t think most people have cottoned on to the fact that you can now stream video directly into a model because it’s only a few weeks old. But wow, that’s a big boost in terms of what kinds of things you can do with this stuff.

swyx

Yeah. For people who are not that close, I think Gemini Flash’s free tier allows you to do something like capture a photo—1 photo every second or a minute—and leave it on 24/7, and you can prompt it to do whatever. So you can effectively have your own camera app or monitoring app that you just prompt, and it detects changes, detects alerts or anything like that, or describes your day. And the fact that this is free also leads into the previous point of prices having come down a lot.

Simon Willison

Even if you’re paying for this stuff, a thing that I put in my blog entry is that I ran a calculation on what it would cost to process 68,000 photographs in my photo collection and, for each one, just generate a caption. Using Gemini 1.5 Flash-8B, it would cost me $1.68 to process 68,000 images, which is—I mean, that doesn’t make sense. None of that makes sense.

It’s 1/400 of a cent per image to generate captions now. So you can see why feeding in a day’s worth of video just isn’t even very expensive to process.

swyx

Yeah. I’ll tell you what is expensive: it’s the other direction. Here, we’re talking about consuming video. This year, we also had a lot of progress. Probably one of the most anticipated launches of the year was Sora. We actually got Sora, and less exciting—

Simon Willison

We did, and then Veo 2—Google’s Sora—came out like 3 days later and upstaged it. Sora was exciting—

swyx

In general, I feel the media or social media has been very unfair to Sora because what was released to the world, generally available, was Sora Lite, the distilled version of Sora.

Right? So you’re—

Simon Willison

I did not realize that.

swyx

You’re absolutely comparing—

Simon Willison

Ah, okay.

swyx

—the most cherry-picked version of Veo 2, the one that they published on the marketing page—

Simon Willison

Yeah.

swyx

—to the most embarrassing version of Sora. So, of course, it’s going to look bad.

Simon Willison

Well, I got access to Veo 2. I’m in the Veo 2 beta, and I’ve been poking around with it and getting it to generate pelicans on bicycles and stuff.

swyx

I would absolutely believe that Veo 2 is actually better.

Simon Willison

That’s interesting. Is Sora—is full-fat Sora coming soon? Do you know? When do we get to play with that one?

swyx

No one’s mentioned anything. I think basically the strategy is: let people play around with Sora Lite and get info there, but keep developing Sora with the Hollywood studios. That’s what they actually care about.

Simon Willison

Gotcha. Okay.

swyx

The rest of us don’t really know what to do with the video anyway.

Simon Willison

Right. My thing is, I realize that for generative images and video—images we’ve had for a few years—I don’t feel like they’ve broken out into the talented artist community yet. Lots of people are having fun with them and producing stuff that’s kind of cool to look at.

But what I want is—you know, that movie, Everything Everywhere All at Once, right? It won a ton of Oscars, an utterly amazing film. The VFX team for that were 5 people.

swyx

Yeah.

Simon Willison

Some of whom were watching YouTube videos to figure out what to do. My big question for Sora and Midjourney and stuff is: what happens when a creative team like that starts using these tools? I want the creative geniuses behind Everything Everywhere All at Once—what are they going to be able to do with this stuff in a few years’ time? Because that’s really exciting to me.

That’s where you take artists who are at the very peak of their game, give them these new capabilities, and see what they can do with them.

swyx

I should—I know a little bit here, so I should mention that that team actually used Runway ML. So there was—

Simon Willison

No way. In that movie?

swyx

Yeah. I don’t know how much, so it’s possible to overstate this. But there are people integrating generative video within their workflow, even pre-Sora.

Simon Willison

Wow.

swyx

Yeah.

Brian McCullough

It’s not the thing where it’s like, okay, tomorrow we’ll be able to do a full 2-hour movie that you prompt with 3 sentences. For the very first part of video effects in film, if you can get that 3-second clip, if you can get that 20-second thing that they did in The Matrix that blew everyone’s minds and took $1 million or whatever to do, it’s the little bits and pieces that they can fill in now that are probably already there.

swyx

Yeah. I think having a layered view of what assets people need and letting AI fill in the low-value assets—the background video, the background music, and sometimes the sound effects—may be more palatable. Maybe it also changes the way that you evaluate the stuff that’s coming out, because people tend to, in social media, try to emphasize foreground stuff, main-character stuff.

So you really care about consistency, and you really are bothered when, for example, Sora botches an image generation of a gymnast doing flips, which is horrible.

Simon Willison

That’s hilarious.

swyx

It’s horrible. But for background crowds—

Brian McCullough

Right.

swyx

—who cares?

Brian McCullough

And by the way, again, I was a film major way, way back in the day. That’s how it started: things like Braveheart, where they filmed 10 people on a field, and then the computer could turn it into 1,000 people on a field. That’s always been the way. It’s around—

Simon Willison

Right. The Lord of the Rings, right?

Brian McCullough

Yeah.

Simon Willison

The Lord of the Rings movies were over 20 years ago. They had those giant battle sequences, which were very early. You could almost call it a generative AI approach, right? They were using very sophisticated algorithms to model out those different battles and all of that kind of stuff.

Brian McCullough

Yeah.

Simon Willison

Yeah, I know very little. I know basically nothing about film production, so I try not to commentate on it. But I am fascinated to see what happens when these tools start being used by the people at the top of their game.

swyx

I would say there’s a cultural war being fought here more than a technology war. Most of the Hollywood people are against any form of AI anyway, so they’re busy fighting that battle instead of thinking about how to adopt it. It’s very fringe. I participated here in San Francisco in a generative AI video creative hackathon where the AI-positive artists actually met with technologists like myself, and then we collaborated to build short films. That was really nice, and I think I’ll be hosting some of those at my events going forward.

One thing that I want to give people a sense of is that this is a recap of last year, but sometimes it’s useful to walk away with what we can expect in the future. I don’t know if you got anything. I would also call out that the Chinese models here have made a lot of progress. Hai Luo and Kling, and God knows who else in the video arena, are also making a lot of progress. It’s surprising. I think maybe, actually, China is surprisingly ahead with regard to open weights, at least, but also just specific forms of video generation.

Simon Willison

Wouldn’t it be interesting if a film industry sprang up in a country that we don’t normally think of as having a really strong film industry, and that was using these tools? That would be a fascinating sort of angle on this.

swyx

Agreed.

Brian McCullough

Oh, sorry.

swyx

Yeah, go ahead.

Brian McCullough

Just to put it on people’s radar as well, HeyGen—there’s a category of video avatar companies that don’t specifically specialize in general video. They only do talking heads, let’s just say. And HeyGen’s done very well.

swyx

Brian, Brian, you know that’s what I’ve been using, right? So if you see some of my recent YouTube videos and things like that, the beauty part of the HeyGen thing is I don’t want to use the robot voice, so I record the MP3 file for my clips every single day, and then I put that into HeyGen with the avatar that I’ve trained it on. All it does is the lip sync.

It’s not 100% uncanny-valley-beatable, but it’s good enough that, if you weren’t looking for it, it’s just me sitting there doing one of my clips from the show. And, yeah, so by the way, HeyGen—shout-out to them.

Brian McCullough

In terms of the look-ahead—going forward, reviewing 2024 and looking at trends for 2025—I would basically call this out. Meta tried to introduce AI influencers and failed horribly because they were just bad at it. But at some point, there will be more and more AI influencers, not in the way that Simon is, but in a way that they are not human.

The few of those that have done well, I always feel like they’re doing well because it’s a gimmick, right? It’s novel and fun. Like the AI Seinfeld thing from last year, the Twitch stream—those, if you’re the only one, or one of just a few doing that, will attract an audience because it’s an interesting new thing. But I just don’t know if that’s going to be sustainable longer term or not.

I’m going to tell you, because I’ve had discussions—I can’t name the companies or whatever—but think about the workflow for this. Now we all know that on TikTok and Instagram, holding up a phone to your face and doing an “in my car” video or a walking-and-talking video is very common.

But also, if you want to do a professional sort of talking-head video, you still have to sit in front of a camera, you still have to do the lighting, and you still have to do the video editing. Versus if you can just record what I’m saying right now—the last 30 seconds—if you clip that out as an MP3 and you have a good enough avatar, then you can put that avatar in front of Times Square, on a beach, or whatever.

So, again, for creators, the reason I think, Simon, we’re on the verge of something is that it’s not going to be that AI avatars take over. It’ll be one of those things where it takes another piece of the workflow out and simplifies it.

Gotcha. I am all for that. I always love tools—tools that help human beings do more ambitious things. I’m always in favor of that. That’s what excites me about this entire field.

We’re looking into basically creating one for my podcast. We have this guy Charlie. He’s Australian, he’s not real, but he opens every show, and we’re going to have him present all the shorts. Yeah, go ahead.

5. Credibility Beats AI Slop

The thing that I keep coming back to is this idea of credibility. In a world that is full of AI-generated everything and so forth, it becomes even more important that people find the sources of information that they trust, and find people and sources that are credible.

I feel like that’s the one thing that LLMs and AI can never have, is credibility, right? ChatGPT can never stake its reputation on telling you something useful and interesting because that means nothing, right? It’s a matrix multiplication. It depends on who prompted it and so forth.

I’m always—and this is when I’m blogging as well—I’m always looking for the reliable people who will tell me useful, interesting information, who aren’t just going to tell me whatever somebody’s paying them to tell them, and who aren’t going to type a one-sentence prompt into an LLM, spit out an essay, and stick it online.

To me, earning that credibility is really important. That’s why a lot of my ethics around the way that I publish are based on the idea that I want people to trust me. I want to do things that gain credibility in people’s eyes so they will come to me for information as a trustworthy source. And it’s the same for the sources that I’m consulting as well. I’ve been thinking a lot about that sort of credibility focus for a while now.

You can layer or structure credibility, or decompose it. One thing I would put in front of you—I’m not saying that you should agree with this or accept this at all—is that you can use AI to generate different variations, and then you, as the final sort of last-mile person, pick the last output and put your stamp of credibility behind that, that everything is human-reviewed instead of human-originated, if that’s the thing.

If you publish something, you need to be able to be proud of publishing it. You need to be able to say, “I will put my name to this. I will attach my credibility to this thing.”

And if you’re willing to do that, then that’s great. For creators, this is huge because there’s a fundamental asymmetry between starting with a blank slate versus choosing from 5 different variations.

The key thing that you just said is that if everything that I do, if all of the words were generated by an LLM, if the voice is generated by an LLM, and if the video is also generated by an LLM, then I haven’t done anything, right? But if you take a shortcut on one or 2 of those, and I’m still willing to sign off on it, I feel like that’s where people are coming around to: this is maybe acceptable.

This is where I’ve been pushing the definition. I love the term “slop,” where I’ve been pushing the definition of slop as AI-generated content that is both unrequested and unreviewed. The unreviewed thing is really important.

The thing that elevates something from slop to not slop is if a human being has reviewed it and said, “You know what? This is actually worth other people’s time.” And again, I’m willing to attach my credibility to it and say, “Hey, this is worthwhile.”

It’s the curatorial and editorial part of it that, no matter what the tools are to do shortcuts—to do, as swyx is saying, choosing between different edits or different cuts—still has a curatorial mind or editorial mind behind it.

6. The GUI Moment For LLMs

Let me wedge this in before we start to close. One of the things that, coming back to your year-end piece, has been something I’ve been banging the drum about is when you’re talking about LLMs getting harder to use. You said most users are thrown in at the deep end.

The default LLM chat UI is like taking brand-new computer users, dropping them into a Linux terminal, and expecting them to figure it all out. I mean, it’s literally going back to the command line. The command line was defeated by the GUI interface, and what I’ve been banging the drum about is that this cannot be the user interface. What we have now cannot be the end result.

Do you see any hints or seeds of a GUI moment for LLM interfaces? I mean, it has to happen. It absolutely has to happen. The usability of these things is turning into a bit of a crisis, and we are at least seeing some really interesting innovation in little directions.

Simon Willison

Just like OpenAI's ChatGPT Canvas thing that they just launched. That is at least a little more interesting than just chats and responses. You're exploring that space where you're collaborating with an LLM, both working on the same document. That makes a lot of sense to me. That feels really smart.

One of the best things is still—who was it who did the UI where you could draw an interface and click a button? tldraw with their Make It Real thing. That was spectacular. Absolutely spectacular, like an alternative vision of how you'd interact with these models. So I feel like there is so much scope for innovation there, and it is beginning to happen. I feel like most people do understand that we need to do better in terms of interfaces that both help explain what's going on and give people better tools for working with models.

Brian McCullough

I was going to say, I want to dig a little deeper into this because think of the conceptual idea behind the GUI. Instead of typing into a command line, "open word.exe," you click an icon, right? That's abstracting away the programming stuff. A child can tap on an iPad and make a program open, right?

The problem, it seems to me, with how we're interacting with LLMs right now is it's sort of like a dumb robot where you poke it and it goes over here. But no, I want it to go over here, so you poke it this way, and you can't get it exactly right. What can we abstract away from what's currently going on that makes it more fine-tuned and easier to get more precise? You see what I'm saying?

Simon Willison

Okay.

Simon Willison

Yes. And this is the other trend that I've been following from the last year, which I think is super interesting. It's the prompt-driven UI development thing. Basically, this is the pattern where Claude Artifacts was the first thing to do this really well. You type in a prompt, and it goes, "Oh, I should answer that by writing a custom HTML and JavaScript application for you that does a certain thing."

Since then, it turns out this is easy, right? Every decent LLM can produce HTML and JavaScript that does something useful. So we've actually got this alternative way of interacting where they can respond to your prompt with an interactive, custom interface that you can work with.

People haven't quite wired those back up again. Ideally, I'd want the LLM to be able to ask me a question where it builds me a custom little UI for that question, and then it gets to see how I interacted with that. I don't know why, but that's such a small step from where we are right now. That feels like such an obvious next step.

Why should you just be communicating with text when it can build interfaces on the fly that let you select a point on a map or move sliders up and down, all of that kind of stuff?

It's gonna—like knobs and dials. I keep saying knobs and dials.

Simon Willison

Knobs and dials, right.

Brian McCullough

Yeah, exactly.

Simon Willison

We can do that, and the LLMs can build it. Claude Artifacts will build you a knobs-and-dials interface, but at the moment, they haven't closed the loop. When you twiddle those knobs, Claude doesn't see what you're doing. They're going to close that loop. I'm shocked that they haven't done it yet.

I think there's so much scope for innovation, and there's so much scope for doing interesting stuff with that model, where anything you can represent in HTML, JavaScript, and SVG, which is almost everything, can now be part of that ongoing conversation.

swyx

Yeah. I would say the best-executed version of this I've seen so far is Bolt, where you can literally type in, "Make a Spotify clone. Make an Airbnb clone," and it actually does that for you zero-shot with a nice design.

Simon Willison

Did you see there's a benchmark for that now?

swyx

Yeah.

Simon Willison

The LMArena people now have a—

swyx

LMArena benchmark.

Simon Willison

A benchmark for zero-shot app generation, because all of the models can do it. I've started figuring out how to—I'm building my own version of this for my own project because I think—

swyx

Oh.

Simon Willison

Within 6 months, I think it'll just be an expected feature. For my dataset data exploration project, I want you to be able to do things like conjure up a dashboard just via prompt. You say, "I need a pie chart and a bar chart, put them next to each other, and then have a form where submitting the form inserts a row into my database table."

This is all suddenly feasible. It's not even particularly difficult to do, which is utterly bizarre: these things are now easy.

swyx

Yeah. I think for a general audience, that is what I would highlight: software creation is becoming easier and easier. Gemini is now available in Gmail and Google Sheets. I don't write my own Google Sheets formulas anymore; I just tell Gemini to do it.

I almost want to somewhat disagree with your assertion that LLMs got harder to use.

Simon Willison

Ooh.

swyx

We expose more capabilities, but they're in minor forms, like using Canvas, web search in ChatGPT, and Gemini being in Google Sheets. We're getting improvements.

Simon Willison

No, no, no. Those are the things that make it harder.

swyx

Okay.

Simon Willison

The problem is that for each of those features, they're amazing if you understand the edges of the feature. If you're like, "Okay, so in Google Sheets formulas, I can get it to do a certain amount of things, but I can't get it to go and read a web page..." You probably can get it to read a web page, right? But there are things that it can do and things that it can't do, which are completely undocumented.

If you ask it what it can and can't do, they're terrible at answering questions about that. My favorite example is Claude Artifacts. You can't build a Claude Artifact that can hit an API somewhere else because the CORS headers on that iframe prevent accessing anything outside of CDNs.

swyx

I hate those.

Simon Willison

People are learning CORS headers as an end user in order to understand why—

swyx

I hate those.

Simon Willison

I've seen people saying, "Oh, this is rubbish. I tried building an artifact that would run a prompt, and it couldn't," because Claude didn't expose an API with CORS headers. All of this stuff is so weird and complicated.

The more tools we add, the more expertise you need to really understand the full scope of what you can do. The question really comes down to: What does it take to understand the full extent of what's possible? Honestly, that's just getting more and more involved over time.

7. Local Models Return

swyx

Yeah. I have one more topic that I think you're kind of a champion of, and we've touched on it a little bit, which is local LLMs and running AI applications on your desktop. I feel like you are an early adopter of many, many things.

Simon Willison

Well, I had an interesting experience with that over the past year. 6 months ago, I almost completely lost interest. The reason is that 6 months ago, the best local models you could run—there was no point in using them at all because the best hosted models were so much better.

There was no point at which I'd choose to run a model on my laptop if I had API access to Claude 3.5 Sonnet. They just weren't even comparable. That changed basically in the past 3 months as the local models had this step change in capability. Now I can run some of these local models, and they're not as good as Claude 3.5 Sonnet, but they're not so far away that it's not worth me even using them.

The continuing problem is I've only got 64 GB of RAM, and if you run Llama 3 70B, most of my RAM is gone. So now I have to shut down my Firefox tabs, Chrome, and VS Code windows in order to run it.

But it's got me interested again. The efficiency improvements are such that now, if you were to stick me on a desert island with my laptop, I'd be very productive using those local models, and that's pretty exciting. If those trends continue, and I think my next laptop, when I buy one, is going to have twice the amount of RAM, maybe I can run almost the top-tier open-weight models and still be able to use it as a computer as well.

NVIDIA just announced their $3,000, 128-gigabyte monstrosity. That's a pretty good price.

swyx

You gonna buy it?

swyx

Custom OS and all.

Simon Willison

If I get a job. If I have enough of an income that I can justify blowing $3,000 on it, then yes.

swyx

Okay. Let's do a GoFundMe to get Simon one of them. Come on. You know you can get a job anytime you want.

Simon Willison

I want a job that pays me to do exactly what I'm doing already and doesn't tell me what else to do. That's the challenge.

swyx

This is just purely discretionary. I think Ethan Mollick does pretty well, whatever it is he's doing. But basically, I was trying to bring in not just local models but Apple Intelligence, which is on every M-series Mac. You seem skeptical.

Simon Willison

It's rubbish.

swyx

It's rubbish.

Simon Willison

Apple Intelligence is so bad.

swyx

It does one thing well.

Simon Willison

Oh, yeah. What's that?

swyx

It summarizes notifications, and sometimes it's humorous.

Brian McCullough

But are you sure it does that well?

swyx

It's decent.

Brian McCullough

The other thing, again, from a sort of normie point of view, is that there's no indication from Apple of when to use it. Everybody upgrades their thing, and it's like, “Okay, now you have Apple Intelligence,” and you never know when to use it ever again.

swyx

Oh, yeah, you consult the Apple docs, which is MKBHD.

Simon Willison

The one thing I'll say about Apple Intelligence is that one of the reasons it's so disappointing is that the models are just weak. But now that—

swyx

Yeah.

Simon Willison

Llama 3B is such a good model in a 2-gigabyte file. I think, give Apple 6 months, and hopefully they'll catch up to—

swyx

Yeah.

Simon Willison

—the state of the art on their small models, and then maybe it'll start being a lot more interesting.

swyx

Anyway, this was year 1. Just like the first year of the iPhone, maybe it wasn't that much of a hit, and then in year 3 they had the App Store. So I would say give it—

Simon Willison

Yeah.

swyx

—some time. I think Chrome is also shipping Gemini Nano this year in Chrome, which means that every app, every web app, will have free access to a local model that just ships in the browser, which is kind of interesting.

I also wanted to open the floor to any of us: what are the AI applications that we've adopted that we really recommend? These are all apps that are running in a browser, or apps that are running locally, that other people should be trying, right? I feel like that's always one thing that's helpful at the start of the year.

Simon Willison

Okay. So, for running local models, my top picks are, firstly, on the iPhone, this thing called MLC Chat.

swyx

Mm-hmm.

Simon Willison

It works, it's easy to install, and it runs Llama 3B. It's so much fun. It's not necessarily a capable enough model for me to use it for real things, but my party trick right now is to get my phone to write a Netflix Christmas movie plot outline where a jeweler falls in love with the King of Sweden or whatever. It does a good job, and it comes up with pun names for the movies. That's deeply entertaining.

On my laptop, most recently, I've been getting heavily into Ollama because the—

swyx

Yeah.

Simon Willison

—Ollama team are very good at finding the good models, packaging them up, and making them work well. It gives you an API. My little LLM command-line tool has a plugin that talks to Ollama, which works really well. Ollama is, I think, the easiest on-ramp to running models locally. If you want a nice user interface, LM Studio is, I think, the best user interface for that. It's not open source, but it's good. It's worth playing with.

The other one that I've been toying with recently is called Open WebUI. The UI is fantastic. If you've got Ollama running and you fire this thing up, it spots Ollama and gives you an interface to your Ollama models, and that's really nicely done. That's my current favorite open-source UI for these things.

There are lots of good options. You do need a lot of disk space. The models start at 2 gigabytes for the 3B models that are actually worth playing with. The really impressive ones tend to be in the 20- to 30-gigabyte range, in my experience.

swyx

I think my struggle here is that I'm not much of an absolutist about running things locally. I'm happy to call an API.

Simon Willison

Mm-hmm.

swyx

Okay, yeah. But I just think—

Simon Willison

I do it to play.

swyx

Yeah, okay, fine.

Simon Willison

It's my research interest, yeah.

Brian McCullough

Answer your own question. Give us more apps that you want to—

swyx

Yeah. Sometimes it's just nice to recommend apps. I use Super Whisperer now. I tried Whisper Flow, but it didn't really work for me. Super Whisperer—

Simon Willison

Mm.

swyx

—is one of them. It basically replaces typing. You should just talk most of the time, especially if you're doing anything long-form. I hold down Caps Lock and talk, and when I'm done, I lift it up. It uses—

It isn't just about writing down your transcripts, because I make ums and uhs all the time, and I restate myself all the time. But I use GPT-4 to rewrite, and that's what these guys are doing. They're all doing some form of state-of-the-art ASR—automatic speech recognition—and then an LLM to rewrite.

I would also recommend that people check out Rosebud for journaling. I think AI for mental health is quite unexplored, and it's not because we're trying to build AI therapists. I think therapists really hate that. You'll never be on the level of therapists.

Brian McCullough

That gets back to the human thing that we were discussing. On some level, there are certain things and disciplines that require the human touch, and that might be one of them.

swyx

Sure. But the human touch costs me $300 an hour.

Brian McCullough

Yes.

swyx

Right?

Brian McCullough

Yeah.

swyx

And this thing's $3 a month. There's a spectrum of people for whom that will work, and I think it's cheap now to try all these things.

Simon Willison

I'm going to throw in a quick recommendation for an app. MacWhisper is my favorite—

swyx

Yeah, that's your one.

Simon Willison

—desktop app. I love that thing.

Brian McCullough

Yeah.

Simon Willison

It runs Whisper, and you can do things like paste in the URL to a YouTube video, and it'll pull the audio and give you a transcript. So that's how I watch YouTube now—

swyx

Ah.

Simon Willison

—I slap it into MacWhisper, then I copy and paste into Claude, and I use the Claude web app to do things.

MacWhisper works with MP3 files. Every time I'm on a podcast, I dump the MP3 into MacWhisper, then I dump the transcript into Claude and say, “What should I put in the show notes?” It spits out a bullet-point list where it says, “Oh, you mentioned a dataset that you should link to,” that kind of thing. MacWhisper—I use it several times a day, to be honest. It's great.

swyx

Yeah.

Brian McCullough

I'm actually going to say one that is incredibly basic and, again, coming back to just my workflow. We are currently recording this on Riverside. Riverside is a great tool for recording video and audio, like we're doing right now.

I always use this as an example when folks ask, “What will AI do for me?” When I first started using Riverside, we were recording 3 different channels, right? You guys are recording locally, so there are 3 audio files and 3 video files. When I first started using Riverside, you had to pump 3 tracks into Adobe and then edit.

“Okay, now we focus on Simon. Now we focus on swyx. Now we focus on Brian. Now we do all 3.” One day, a tool popped up that said, “Hit this button,” and it was Smart Edit. The AI determines, “Okay, Simon has been talking for 30 minutes, so go to the full shot of him. Brian is now talking, or there's overtalk, so let's have all 3 talking heads.”

With one button, for anything I posted, it saved me 3 or 4 hours' worth of work. That, to me, is, again, if normies are listening—

Simon Willison

In fact, Riverside has that feature now.

Brian McCullough

Yeah.

swyx

Yeah, yeah.

Simon Willison

Damn.

Brian McCullough

I don't use it.

Simon Willison

Oh, that sounds fantastic.

Brian McCullough

I still use a human editor. The day it came out, I was running around the house telling my wife, telling anyone that would listen, “You don't know, I just saved 3 hours because they had a new feature.”

Simon Willison

Wow.

That's exciting.

Brian McCullough

Brian's basically crying with joy right now. All right, let's try to bring this to a landing a little bit. Simon, I have maybe 2 or 3 more. We can do these rapid-fire.

One of my shows—one of the things about my show is that it's sort of like Silicon Valley writ large, so it's sort of like the horse race of who's up and who's down or whatever.

To the degree that you're interested in pontificating on this, OpenAI as a company in 2025, do you see challenges coming? Are you bearish or bullish? I'm almost doing a CNBC sort of thing, but how do you feel about OpenAI this year?

8. OpenAI Faces New Challenges

Simon Willison

I think they're in a bit of trouble. They seem to have lost a lot of talent.

Brian McCullough

Mm.

Simon Willison

And they don't have that top-of-the-pile thing. If it wasn't for o3, they'd be in massive trouble because they'd have lost that. I think o3 clawed them back up again.

One of the big stories of 2024 is that OpenAI started as the clear leader, and now Google Gemini is really good. Google Gemini had an amazing year. Anthropic Claude, Claude 3.5 Sonnet, is still my personal favorite model, and that feels notable. Nobody would argue that they weren't the leader in all of this stuff a year ago, and today they're still doing great, but they're not as far ahead as they were.

Brian McCullough

Next question, and maybe this couldn't be as rapid-fire, but I loved, finally, from your piece the idea that LLMs need better criticism, which I'd love you to expand on. As I straddle this world of tech journalism, creator, investor, and all that stuff, I thought that you had a really interesting thing to say about how—and we even alluded to this—Hollywood is against it.

Better criticism, in the sense that, as I took it, everybody's got their hackles up. They're trying to defend their livelihoods and things like that. But it's either, “This is going to destroy my job and destroy the world,” or—I'm sorry, I'm again leading the witness—what did you mean by LLMs needing better criticism?

Simon Willison

This is a frustration I have. If I read a discussion thread somewhere about this topic, I can predict exactly what everyone's going to say. People talk about the environmental impact. They talk about the plagiarism of the training data and the unlicensed training data. They'll often say, “Oh, and these things are completely useless.”

That's the one that I will push back against. The other things are true, right? The argument I always make about the idea that LLMs are just completely useless is that they are very useful if you understand how to use them, which is distinctly unintuitive. You have to learn how to deal with something that will just wildly hallucinate and make things up, and all of those kinds of things.

If you can learn what they're good at and what they're bad at, I use them dozens of times a day, and I get enormous value out of them. So I'll push back on people who say, “No, they're just useless.”

But the other things—the environmental impact and the way the training data works—I feel like the training-data one is interesting because it's probably legal under fair use, but it's clearly unfair if somebody takes your work without your permission and trains a model that then competes with you in the marketplace. Legal or not, I understand why people are upset about that. That's a reasonable thing to be upset by.

So what I want—and I also feel like this stuff can have a major impact on society, especially as it starts undermining all sorts of jobs that we never thought were going to be undermined by technology. Who thought it would come for artists and lawyers first? That's bizarre.

We need to have really high-quality conversations where we help people figure out what works and what doesn't work. We need people to be able to make good decisions about what to do with their careers, to embrace this stuff, and all of that sort of thing.

If we just get distracted by saying, “Yeah, but it's useless, plagiarism-driven, environmentally catastrophic,” even though those things represent quite a lot of truth, I don't think that's a useful message to lead with. I want to be having the much more interesting, high-level conversations: If there are negatives, how do we counter those negatives? If there are positives, how do we encourage those? How do we help people make good decisions about how to use this technology?

swyx

Yeah. I think where I see this the most is for people who are very internal. You and I are immersed in this every single day, so we're frankly tired of the same debates being recycled again and again.

I think what might be more useful or more impactful is the level at which it starts to hit regulation. Last year, we had a couple of very notable attempts at the White House level and in California to regulate AI, and those did not come to pass.

At some point, these criticisms bubble up to law, to matters of national security or national science and progress. I feel like there needs to be more information or enlightenment there, maybe, if only because they tend to be very trailing.

Simon Willison

Right.

swyx

My favorite example to pick on, which is very unfair of me, but whatever, is that the California SB 1047 act tried to cap compute at 10^25.

Simon Willison

Yeah. That was DeepSeek.

swyx

Exactly. And it also was exactly at the point at which we pivoted from training GPT-5 to o1, where we're no longer scaling pre-training compute. What I'm saying is that we're always trying to regulate the last war, and I don't think that works in a field that is—

Simon Willison

So—

swyx

—basically 8 years old.

Simon Willison

I think there are 2 areas of regulation I'm super interested in. One of them is that I do think regulating the way these things are used can work. The big example is that I don't want somebody's insurance claim denied by a black-box LLM where nobody can explain what it did. That just feels—

swyx

Oh, we have real laws for that. This is like redlining.

Simon Willison

Exactly. Take those laws, reinforce them, and update them for modern capabilities.

The other one is privacy. We've got this huge problem right now where people will refuse to use any of these tools because they don't trust that the things they say to them won't be trained on and then exposed to other people. There are lots of terms and conditions that you can read through and try to navigate around.

I would love there to be straightforward laws that people understand, where they know that their input isn't going to be used for training because there's a law that says, under these circumstances, that can't happen. It's basically taking our existing privacy laws, giving them a few more teeth, and reinforcing them without introducing cookie banners à la the European Union.

These things are always very risky. You can have all sorts of bad results if you don't design them correctly. But there's space for that, I think.

Brian McCullough

Yeah. When I read that piece and then when you just said, “Swyx said we're in the weeds on this every single day, so we're tired of hearing these arguments,” it reminds me of folks who are always into politics. Then they're mad at the people who don't care about politics until it's an election year, and they're like, “Well, you're a low-information voter because all you know is that the factory in your town got shut down, or there's inflation, or whatever, and so you vote one way or the other, but you haven't been paying attention.”

But that's kind of the point: You shouldn't expect normal people to pay attention, except for the fact that this might lose me my job. So you can't blame them for being—I don't know if reactionary is the word—or emotional.

If you're in the weeds, it's harder to keep everybody informed, and this is going to touch everybody, so I don't know.

Okay, so this is the very last one, and then we can wrap and do plugs and everything. Simon, this is for you. It was alluded to a little bit, and you might not have one, but if there's something this year that a generalist like me isn't aware is coming down the pipe that you think is going to be big in the AI space—and maybe Swyx, if you've got one too—what do you think it would be?

Simon Willison

I think for most people who haven't been paying attention, we know these things already. We know that the models are now almost free to run things against. The fact that you can now do video, stream video to a model—the thing where you can share your entire screen with a model and get feedback is going to be really useful.

Again, the privacy side of things really matters, though. I do not want some model just training on everything that it sees on my screen. But no, the stuff that's now possible as of a few months ago is enough. I don't need anything new. That's going to keep me busy all year.

Brian McCullough

Swyx, you got one?

swyx

Simon's always too content, and then he sees the next thing and he's like, “Oh yeah, that's great too.”

Simon Willison

Yep.

swyx

Okay. I love trying to be contrarian by asking, “What does everyone hate right now?” Remember, this time last year, we had just had CES and the Rabbit R1.

We had the Humane, right?

Brian McCullough

Yeah. Wearables. Yep.

9. AI Wearables Make A Comeback

swyx

Those are completely in the gutter. No one will touch them. They're toxic nuclear waste. Okay, this year is the year of wearables.

Brian McCullough

Yep.

Simon Willison

Huh.

Brian McCullough

I agree with you, by the way. That cycle always works out where you go to a CES and it's everything—hype, hype, hype, hype—and then 3 years later it becomes the thing, unless it's 3D TVs, in which case that was a mistake anyway. But yeah—

Simon Willison

Well, transparent TVs are the big thing—

Brian McCullough

Mm.

Simon Willison

—for the last couple of years. What the hell?

Brian McCullough

Yeah.

swyx

I think Simon may have got one of these, but there are a lot of people working on AI wearables here in SF. They are surprisingly cheap, surprisingly capable, with decent battery life, and they do useful things. We have to work out the privacy aspect, of course. But people like Limitless, which used to be called Rewind, I think—

Brian McCullough

Mm-hmm.

swyx

They're shipping one of these wearables that, based on your voice, only records your voice. So you opt in.

Simon Willison

Interesting. Right.

swyx

Right? And so you can have perfect memory if you want. You can have perfect memory at work. Your employer can buy these for you; it only applies at work, and it's fine. It's just a meeting aid.

Lots of people use Granola or some kind of Fireflies or some of these meeting recorders only for online meetings, but what about in-person meetings? What about conversations and locations that you've been to? Some of that should be a choice. Right now you have zero choice. And I think these wearables will enable some of that.

It's up to us as a society to determine what's acceptable and what's not. I really like these gray areas where we still don't know yet. Whenever I tell people about this, they're like, “I don't know.” I guess it's as though you have perfect memory, but some people have better memory than others. Where's the line?

Brian McCullough

Hmm. And—

swyx

There will be a lot—

Brian McCullough

Now I—

swyx

A lot more of these.

Brian McCullough

I would add to that because, swyx, as you know—you listen to my show—the idea is that AI has taken smart glasses and completely changed everyone's mind about that as a product category and form factor. And I should say this: from things that I've been looking at investing in, wait until you see what they can add on to earbuds.

Simon Willison

Ooh.

Brian McCullough

Like the earbuds in your ear can do a lot more things than they're doing now. Then you combine that with smart glasses, and you combine that with an LLM that you can access maybe with a phone as the mothership. There are some interesting things. CES next year is going to be crazy if you think AI wearables are a thing.

swyx

Anyway, this year they were not a thing. There were very much no wearables at CES.

Brian McCullough

Mm.

Simon Willison

This one's interesting as well because the thing that makes these interesting is that it's multimodal, right? Audio input, video input—

Brian McCullough

Yeah.

Simon Willison

Image input. A year ago, that was hardly a thing, and now it's dirt cheap. We're in a much better position now than we were 12 months ago to build the software behind this stuff.