面向金融的 Claude Code + 全球内存短缺:Doug O’Laughlin,SemiAnalysis
Doug O’Laughlin 的核心判断是,Claude Code 在 Opus 4.5 左右从新奇玩具跨越成了具备经济价值的信息工作者。它能把原本看似需要“博士项目”的工作压缩到1—2天,但“这东西一直在犯错”;如今它更像初级分析师,而不是自主专家。Shawn 把这款付费工具比作一个完全听话的初级分析师。眼下最直接的收益属于有经验的审阅者,他们能识别出“手工打磨的最后5%”;初级数据分析岗位则越来越暴露在风险之下。
Claude Code 的采用曲线表明,编码代理正成为所有信息工作的接口,而不只是编程工具。Doug 描述称,Claude 署名的提交在“2周左右”已占公开 GitHub 活动约4%;Shawn 说更新后的图表约为5%。Doug 刻意保守地预测年底达到25%,而 Shawn 认为当前速度更接近50%,25%只能算95%置信区间内的下限。更大的判断是:“Excel 是分析师的 IDE”;一旦代理能够检索可信数据并直接呈现答案,Excel、Bloomberg、PowerPoint 以及其他面向人的界面都会受到冲击。
最明确的硬件判断是,Doug 认为内存短缺大约还要2年才能缓解。考虑到工艺复杂度和良率,生产1 bit 的 HBM 实际上会消耗3—4 bit 的传统 DRAM 产能;与此同时,供应商刚刚走出一次严重衰退,期间无尘室投资被冻结。SemiAnalysis 认为 DRAM 价格还有再涨100%的空间,这将迫使需求遭到破坏,推迟数据中心和设备建设,通过 CXL 重新启用老 DDR4,并让“上下文配给”成为一种可能真实存在的产品逻辑。
Google 的 TPU V7 拥有暂时性的成本窗口,但 Nvidia 对内存和更广泛供应链的控制,可能会随着 Rubin 到来而关闭这一窗口。Doug 认为,如果供应不受约束,几乎所有主要实验室都会消耗更多 TPU V7;双方讨论给出的潜在 TPU 业务规模接近1万亿美元,对应约30%份额。但 TPU V8 相对 Rubin 的 HBM3 对 HBM4、Nvidia 激进的供应商管理,以及 Google 有限的可用产能,意味着这是一场存量基础竞争,而不是一次干净利落的架构胜负。
Microsoft “最可能失去最多”,因为 AI 正在攻击人类进行信息工作的横向软件,而 Azure 却在为攻击者提供融资。Doug 把 OpenAI 比作“城门外的蛮族”:Microsoft 可以继续做一家高利润的算力供应商,同时任由 Office 的抽象层被颠覆;也可以把投资转向保卫 Office 和打造内部模型。Doug 认为,Oracle 试图承接被挤出的建设需求,其底层投资未必是坏生意,但规模巨大且节奏失控的债务发行冲击了市场流动性,把融资节奏变成了瓶颈。
AI 基建热潮已经更像铁路,而不是互联网,这意味着未来可能经历多轮繁荣与破灭,而非一个干净利落的周期。Shawn 提到,铁路投资一度约占 GDP 的4.8%、固定资本形成总额的25%;他同时把 Stargate 与美国 GDP 约2%的规模相比较。Doug 不会断言需求会永远上升,但 Claude Code 改变了他对需求弹性的判断:他愿意每年支付2万—3万美元甚至更多,因为它表现得像许多名并行工作的初级分析师。
Doug 的半导体框架仍然是核心锚点:摩尔定律放缓,而 scaling law 驱动的需求加速,把价值转移给仍能交付系统级性能的公司。他在2020年得出的 Nvidia 是主要受益者的结论,方向上被证明是正确的,尽管连他自己也没预料到它会成为“全球最有价值的公司”。真正持久的投资经验,不是精确维护 EPS 预测,而是找出能够推动数十亿美元收入的1—3个技术变量,再用代理去研究它们,同时不把判断力外包出去。
1. ASML 让一名注重质量的投资者变成半导体迷
Doug 纠正了那个过于简化的起源故事:不是某位导师把他“nerd-sniped”进半导体行业的。2018年,他在寻找高质量复利股时发现 ASML,“爱上了它”,随后沿着产业链向下研究教材、制造设备,以及那些能够生产每一层结构的公司。
真正留住他的是,半导体制造看起来“全是科幻”。即使在智能对话机器人出现之前,每一代新芯片的制造都需要看似不可能的技术,但投资者却把这些进步视为一个所谓成熟行业的日常特征。
这种着迷最终变成了职业选择。Doug 曾多次提前捕捉趋势,包括2019年开始沉迷 TikTok,但 ASML 让他看到了“一轮非常大的浪潮”,并形成了足够强的信念,足以重组生活、全情投入。
2. 摩尔定律放缓颠覆了成熟硬件的投资范式
Doug 真正被“激进化”的信念是,摩尔定律正在终结。行业入门材料仍把半导体描述为集中度高、成熟且增速缓慢的行业,而主流投资范式则假定 CPU 会一代代自然改善,硬件的价值仍会低于软件。
他的因果链条很简单:scaling law 会把算力需求大幅推高,而晶体管缩放放缓,将移除过去自动带来、并不断降低性能成本的供给侧增益。当“所有这些免费的增益”消失后,芯片、网络、封装和完整系统设计方面的专业能力就获得了定价权。
Nvidia 是最清晰的例子,因为它掌握了“从芯片到网络、设计再到扩展规模的每一个环节”。并行计算很重要,但当性能不再能自动来自下一代 CPU 制程节点时,系统级能力更为关键。
3. Nvidia 投资论断的兑现超出了 Doug 自己的信念
2020年,在一篇关于 GPT-3 和“曙光已现”的文章中,Doug 认为 scaling law 驱动的需求加上摩尔定律失效,将利好半导体,而 Nvidia “基本上是唯一一家”有能力捕获上行空间的公司。这仍然是他最喜欢的长期判断。
重要的是他的不确定性:他深信这一论点,却没有预见到其规模。即使有人当时告诉他 Nvidia 会成为“全球最有价值的公司”,他仍会觉得难以置信,尽管前提都是他自己写下的。
Doug 认识 Dylan Patel,是因为 Dylan 看起来像“全世界唯一一个和我一样痴迷半导体的人”。两人的视角互补——Dylan 以技术为先,Doug 受过更多金融训练——支撑了他们早期的判断,也促成了他们在 SemiAnalysis 的合作。
4. SemiAnalysis 专注于足以撬动数十亿美元的拐点
Doug 对卖方研究的批评是结构性的。这个行业源自银行需要名义上独立的分析师帮助分销证券,后来演变成围绕买入、持有、卖出评级以及分毫级 EPS 预测的“机械维护”,而不是差异化的技术研究。
“没有人会说,‘我的预测总是比所有人精准1美分,这就是我擅长选股的原因。’”准确率高5%很少能决定一笔投资;判断一个产品拐点是否真的发生,却可能让收入上下波动数十亿美元。
他举的例子是 AMD 的 Helios 机架:它是否真的按时交付、能否在发布时准备好生产 token,远比季度表格精度的又一次迭代重要。同样,某个网络瓶颈的出现时间,也可能决定整个部署节奏。
因此,分析师与投资组合经理的对话最终会归结为找出“真正重要的1件、或者说3件事”。难点在于,既要跨越多个供应链环节的博士级深度,又要保留足够的宽度,理解某个不起眼的组件或光学技术如何改变其上层经济性。
5. Opus 4.5 跨过了 Doug 的一次完成门槛
Doug 长期把 SemiAnalysis 的招聘案例当作非正式基准:一个代理何时能完成一项人类大约需要24小时的多步骤金融分析任务,何时能击败最弱的应聘者。早期系统展现出潜力,但还没有稳定达到要求。
Opus 4 能搭建副项目,但需要大量反馈,而且经常出错。Doug 早先也试过 Codex,却没得到他想要的无缝代理体验;真正的质变出现在12月下旬——大约12月20日—27日之间,Doug 说是在圣诞节返程后——Claude Code 搭配 Opus 4.5 开始能够一次完成完整任务和小型项目。
关键变化并不是输出完美。它可以处理仪表盘、电子表格或类似项目,接受改进要求,并持续迭代,而不需要 Doug 每一轮都重新拼装整个项目。
这段经历催生了他现在的口头禅:“你就是可以直接做事。”当通用项目可以从自然语言意图出发完成时,每个尚未实现的功能至少部分都成了“能力问题”——取决于如何设定目标、提供上下文、进行审阅和使用工具。
6. 投资组合分析变成了可扩展的判断系统
Doug 从一个很普通的请求开始:导入自己的持仓和笔记,整理内容,计算基本的投资组合风险,并持续维护结果。成功之后,他让 Claude Code 编码自己的投资风格、搭建框架、给持仓打分,并攻击底层假设。
价值来自迭代式扩展。最初的复制粘贴笔记,变成了对其思考方式的可复用表达,能够接收新数据、应用评分标准、比较持仓,并在无需传统开发项目的情况下生成仪表盘。
评分标准有助于管理随机性:明确重要维度,每项按10分制打分,并让代理暴露工作薄弱之处。Doug 描述称,他既会让代理按评分标准执行任务,也会在单独一轮中进行评估;讨论表明,新的评估上下文可能减少偏见和迎合。
Opus 4.6 倾向于赞同,使这种隔离更有价值。当创作和评分共享同一上下文时,先前的推理可能污染批评;干净的上下文可以减少迎合,也能防止模型为已经做出的决定寻找合理化解释。
7. Claude 署名的提交揭示了指数式采用曲线
为了寻找超越线上“集体癔症”的证据,Doug 注意到 Claude Code 可以为公开提交签名。他让代理系统性地抓取这些签名、查询可用的 GitHub 数据、计算每日总量,并将其表示为整体活动的百分比。
Doug 描述称,结果在“2周左右”达到约4%。Shawn 说更新后的图表约为5%。Doug 表示,自己从未见过如此明显的指数趋势。
图表本身也是代理生成的——Doug 提到 Opus 4.5 或4.6——而不是手工制作。他还让代理阅读约70本可视化书籍,将其中有用的经验压缩成一个很小的风格技能,并套用 SemiAnalysis 的颜色、格式和水印。
他的重点是经济性,而不是美观:“做这件事的成本为零。”如果综合70本书与综合3本书的成本基本相同,代理就可以吸收更大的参考集,只保留重要规则,并基于新数据反复生成图表。
8. 代码正在成为所有信息工作的底层语言
Shawn 观察到,Anthropic 的生产流量中软件工程仍接近50%,但他追问,数据分析等表面上独立的类别,是否也正在变成软件工程。Doug 的回答基本是肯定的:目前,代码是机器与外部世界进行交互的接口。
金融行业早已在使用抽象层。Excel 模型编码了各种关系,并表达某项资产可能值多少钱;Doug 说,编程可能更难,所以“你是说难的已经自动化了,那为什么简单的不能自动化?”
这也是 Doug 觉得 Claude for Excel 不如 Claude Code 使用 Python、再在必要时把结果写回 Excel 的原因。把代理塞进传统工作簿,就像把“汽车发动机”硬装进“马车”;面向机器的表达应当成为主表达层,人类格式只在边缘生成。
9. 上下文卫生胜过复杂的代理编排
Doug 当前的设置强调少量强技能、API 访问,以及通过受控服务暴露的 SemiAnalysis 内部数据。每次会话开始时,他会指定一个应在单个上下文窗口内完成的具体目标,然后让专注的子代理收集有边界的信息。
新的100万 token 上下文是一次巨大改善,因为项目指令占窗口的比例更小,子代理也可以在自己的上下文中工作。相比之下,反复压缩会开启“对噪声的压缩”,并持续损害保真度。
他承认自己对 hooks 的使用不足。早先对 Ralph 循环和 Gas Town 式编排的热情,后来让位于“少即是多”:当前模型的保真度还不足以支撑奢侈的多阶段自动化,而紧凑的技能和明确的上下文更可靠地把任务推到完成。
上下文腐化仍然明显。代理可能变得语无伦次、偷懒,或者忘记
CLAUDE.md中已经记录的 API;Doug 借用了《Of Mice and Men》的梗:一个精疲力竭的会话最终必须被“放倒”。
10. Kimi K2.5 的 swarm 在 Claude 的代理团队失效之处奏效
Shawn 的争议性判断是,Claude 的实验性代理团队功能尚未接受协调所需的强化学习。把一个大型、跨公司的 KPI 仪表盘拆给团队后,表现“明显变差”,因为基于提示词的分工缺乏情境意识。
子代理表现更好,因为每个代理都会收到更干净、边界明确的任务,并返回结果。Doug 说 Kimi K2.5 的 swarm 确实很好;Shawn 表示,在他的实验中,swarm 显著改善了模型表现,并将其与 Claude 较弱的代理团队作对比。Doug 提到 Anthropic 曾在一篇文章中谈到强化学习和游戏,因此具体训练解释仍有争议。
这让 Doug 可以将一组问题运行20次,衡量不同模型的表现,再比较其中的定性差异——这是3个月前“普通人”还无法完成的内部基准测试。Shawn 说,运行 swarm 需要约16个 H100 节点,凸显出 swarm 同样是一个横向扩展算力的故事。
11. OpenClaw 的能力先于其安全模型到来
Shawn 第一次使用 OpenClaw 时感到“非常兴奋”:它可以读取邮件、查看日历,并在个人信息之间执行操作。随后他意识到这套系统极易受到 prompt injection,于是撤回了暴露敏感账户的计划,认为现阶段专注使用 Claude Code 已经足够。
Shawn 通过使用独立邮箱来降低风险,只有在代理证明有用之后才会进一步授权。他说 Clawdbot 没有给他留下深刻印象;那些对 Moltbook 做出反应的人忽略了一个事实:终端代理仍经常无视明确工具,或丢失任务焦点。
Doug 仍然区分完成任务与制造新奇感。Zapier 也能更安全地实现部分相同工作流,但通常需要数小时僵化的点击;一个代理只用4分半钟就达到同样结果,代表的是另一种机制。Shawn 的总结很有记忆点:“你的先验会变成你的牢笼。”
12. 一个失败的内存价格模型,仍把博士项目压缩到了几天
Doug 让代理收集历史 NAND 和 DRAM 价格,选择并微调 Chronos-2,加入协变量,识别租用的 GPU,评估结果,最终还通过内部 Vercel 仪表盘暴露这项工作。
预测论点失败的原因具有典型金融特征:内存市场会经历规则可能反转的不同制度,而把每个周期视为独立样本又会摧毁样本量。Shawn 正是基于这一点反对 LLM 股票游戏——过去的关系只能成立到“某个基本面发生变化”为止。
Shawn 把这套失败的预测方法概括为最终只剩启发式规则——“祝你好运,玩得开心”——而 Doug 的结论是,时间序列模型可能根本不会奏效。但这个项目仍然留下了他能找到的所有历史序列、付费 API 数据、宏观协变量,以及描述每个周期开始、中段和结束的框架。
此前,他曾通过旧年报、GDP 数据和叙事重构,手工整理1980年代、1990年代、2000年代和2010年代的历史。过去看似需要“一辈子”或一个博士项目的工作,如今用1—2天就能完成,尽管最后仍然需要专家判断。
13. 专家能看到垃圾,因此能捕获上行空间
Shawn 的反驳关乎声誉:SemiAnalysis 不能把流畅但错误的工作发布给付费客户,而一个没有亲自收集证据的分析师,可能缺乏挑战这些内容所需的知识。Doug 没有淡化这个问题,而是表示:“这东西一直在犯错。一直都在。”
他使用的比喻是,一个初级分析师为高级决策者收集痛苦的信息。缺失的,是初级分析师在反复处理案例后逐渐内化经验、发现自己在哪些地方命中率可靠,并将其提升为真正专业能力的历史过程。
Doug 希望未来系统能够获得这种学习能力——“全世界花费1千万亿的人都认为它会”——但他不认为今天已经具备。在 SemiAnalysis,代理能够大幅放大现有专家的能力,因为专家脑中已经积累了模式、例外和隐含的修正机制。
人类的贡献是“手工打磨的最后5%”:发现虚构的假设,知道某个前提何时与分析师自己的估值框架冲突,并修正前提,而不是把输出润色得更漂亮。这使代理使用成为“一场卫生游戏”,而不是接受任何看起来完整的东西。
14. 自动化威胁的不只是初级岗位,还有学徒制
Doug 担心,跳过证据收集的初级员工,可能永远无法建立审阅所需的内部模型。检查答案有没有明显错误,不等于亲自与原始问题搏斗,直到其中的例外和层级关系变得直觉化。
因此,相比已经成型的研究员,SemiAnalysis 可能会对新员工使用 AI 更加谨慎。“你仍然必须做其中一部分”;否则得到的只是没有认知过程的输出,而其粗糙之处对于已经付过学费的人显而易见。
Shawn 把自动化重新表述为增加掌舵次数:人类可能只尝试一次分析,而代理可以并行生成多个版本,把人的注意力转向审阅。Doug 接受这种杠杆效应,但又回到同一个风险——如果审阅者从未亲自做过这项工作,就可能不知道该关注什么。
他们讨论了始终在线的 heartbeat 设计:它可以审阅会话、提取经验,并把经验带入下一项任务。一个能够检索每次历史互动的专业客服代理,或许拥有比任何人类都更多的案例上下文,前提是验证机制能阻止累积错误永久写入记忆。
15. 经济意义上的 AGI 先于“机器之神”到来
Doug 在一个实用定义下变得“AGI 上头”:系统能否自动化、改造或消除一批有意义的信息工作。Opus 4.5 开始完成更长项目后,他的答案变成了“是,100%”,尽管他拒绝把当前代理描述成完美或超级智能。
初级数据分析是他最清晰的案例。给定一个设计良好的代理系统,持续扫描季度数据并寻找有趣变化,他无法想象平均22岁的人能稳定地“把它打得落花流水”。
Shawn 提到 GDPval,其中50%代表模型与行业专家打平。他说,包括 GPT-5.2 和 Opus 4.5 在内的新系统已经进入70多分区间;他的解读是,在抽样的专业任务中,模型超过专家的次数已经更多。
双方的分歧主要在词汇。Shawn 称这符合 AGI 定义;Doug 同意它适用于白领工作,但将其与 ASI 和“机器之神”区分开来。把目标推向超级智能,会掩盖大量普通经济工作已经跨过有意义门槛这一事实。
16. AI 可能提高产出,却让 GDP 更难解读
Doug 描述了一条演进路径:经济体从农业进入制造业,再进入白领工作和成熟金融部门。人类会创造新工作并适应变化,但即使最终出现新的经济层,5—10年的转型仍可能在社会层面突然发生。
他的“异想天开理论”明确带有保留:AI 可能具有极强的通缩效应。信息工作的数量可以扩张,但在供给充裕的情况下,其市场价值会崩塌,使 GDP 中隐含的服务小时逻辑不再能代表实际生产性产出。
这可能在调整阶段形成一场“AI 大萧条”——不是因为完成的工作更少,而是因为社会尚未为大量廉价认知定价,也尚未吸收它。Doug 强调不确定性;他观察到的近期反应只是,人们会用生产率增益来更加努力地工作。
17. 铁路意味着多轮 AI 资本开支周期,而不是一轮
Doug 更喜欢用铁路建设类比互联网。他说,互联网以实际美元计算的建设规模约为1万亿美元,而 AI 的绝对规模已经超过互联网;铁路则是一个长达约45年的基础设施周期,其中包含3次繁荣与破灭。
铁路做的不只是运输农业产出:其融资需求帮助现代银行业成形,铁路债务一度占据票据市场主导地位。这类基础设施规模巨大、部署缓慢,周围不得不形成新的资本机构。
Shawn 提到,铁路资本开支接近 GDP 的4.8%和固定资本形成总额的25%;他把 Stargate 与美国 GDP 约2%的规模相比。Doug 预计 AI 最终会超过铁路,但更快的信息流动会压缩周期。
他的基准情景不是一条不间断的上升线。供需曲线最终会交叉,每一轮资本繁荣都会经历“必须建设,价格无所谓”的阶段,直到投资者发现,“那是一个高得离谱的价格”。
18. Claude Code 终于让 Doug 看清 token 需求
Claude Code 改变了 Doug 对需求曲线的信心,因为他自己变成了重度用户。标准 Max 套餐“远远不够”;他说自己使用 Fast,并在 API 上花了“100万美元”,但逐字稿没有说明单位。
当被问到这款工具每年值多少钱时,他估计“轻松2万—3万美元,如果不是更多”。相关比较对象是一名约9万美元、完全听话、并且可以被并行实例化多次的初级分析师。
这种支付意愿并不能证明需求无限,也不能证明“数字会永远上涨”这种不诚实的论断。但它确实表明,一项新能力一旦能够稳定完成有价值的专业工作,其需求弹性可能远高于消费者订阅价格所显示的水平。
19. 分析师的 IDE 正走向程序员 IDE 的命运
听到传统编码 IDE 已死的说法后,Doug 认为同样的逻辑适用于金融。“Excel 是分析师的 IDE。Bloomberg 是分析师的 IDE。”二者都保留着围绕人类操作员能够手动导航而建立的界面。
Shawn 斩钉截铁地说:“我再也不会在 Excel 里做图了。”Doug 认同更广泛的结论:代理可以查询可信来源、分析关系,并比人类操作单元格更快地返回 Matplotlib 图片或定制仪表盘,即便输出与旧模板略有不一致。
Shawn 早期创业时曾试图挑战 Bloomberg,并由此认识到 Bloomberg 的护城河更多来自消息、新闻和数据源,而不是界面。Shawn 说 SemiAnalysis 正在转向 FactSet API 加 Claude Code,同时承认交易员和受监管的交易工作仍依赖基本分析师工作流之外的信息网络。
Shawn 估计,对部分分析师而言,降低终端规格每年可节省1万—2万美元。更重要的是,这一工作流用直接表达意图,取代了对文件夹、密钥和函数的隐性记忆。
20. Claude 主导通用工作,而 GPT-5.3-Codex 已经“编程上头”
Doug 对 Claude 署名 GitHub 份额的年底预测是25%,而且他说自己刻意压低了预测。Shawn 认为25%落在95%置信区间内,但当前轨迹更接近50%;双方都承认,公开署名追踪会漏掉一部分活动。
GPT-5.3-Codex 出现后,Doug 修正了此前关于 Anthropic 优势主要来自 token 效率的看法:“他们回来了。”它的强化学习栈看起来非常适合编程,而 Doug 预计,未来更强的预训练运行搭配这套栈,可能会扭转竞争格局。
局限在于专业化。Doug 只是希望 GPT-5.3-Codex 导入并理解网页,它却反复尝试搭建抓取软件;Opus 4.6 更自然地在研究、评分标准、分析和编码之间切换,因为它没有那么“编程上头”。
Shawn 提到 Conductor 这类多模型界面,让 Claude 和 Codex 互相审阅。Doug 没有亲自使用过 Conductor,但同意跨模型审阅可能有用,同时不信任夹在快速迭代、资金充足的第一方平台之间的中间件:看似干净的超集,往往会被其中一方“吃掉”。
21. Microsoft 正在把算力租给城门外的蛮族
Doug 否认 Microsoft “退出 AI”这种吸睛说法,但坚持认为它“最可能失去最多”。Office、PowerPoint、Excel 和邮件都是人类进行信息工作的横向接口,而这正是代理威胁的抽象层。
Azure 让防御变得复杂,因为 Microsoft 正在通过向 OpenAI 出租基础设施赚钱,而 OpenAI 可能成为颠覆者。Doug 的类比是罗马雇佣蛮族:每一年,雇佣兵都变得更强,而它们最终可能攀上的城墙却更加破败。
战略分叉非常痛苦。Microsoft 可以突出 Azure,承担沦为高溢价“哑管道”的风险;也可以把资本转向自研模型和 Office 防御,同时牺牲部分云业务增长。
Claude for Excel 和 Claude for PowerPoint 让指控更加尖锐,因为“Microsoft 本来就应该造出来”。Doug 把管理层关于内部能力投资的表态解读为正在把资源拉回城墙,但在执行表现参差不齐之后仍持怀疑态度,并表示 Microsoft 必须选择方向。
22. Oracle 的融资节奏把资本供给变成瓶颈
Shawn 问 Oracle 承接 Microsoft 放弃的产能是否不负责任。Doug 的回答是,这套安排主要是一次自摆乌龙:Oracle 承诺了巨大规模——他记得 RPO 约为4000亿美元——却在建立更平滑、部分由已部署 GPU 收入自我融资的爬坡路径之前就激进融资。
Shawn 做了一个粗略且明确不精确的比较:投资级 TMT 债务规模接近5000亿美元,而 Oracle 约为1350亿美元。如此大规模的发行必须提供更好的条款来吸引买家,重新给竞争性债务定价,并通过单纯的供给压力导致整个指数下跌。
因此,意外的瓶颈是“进入市场的债务供给”。超大规模云服务商过去一直能够自我融资;突然要求信贷市场提供数量级更多的资本,即使底层项目最终可能赚钱,也会制造流动性问题。
Microsoft 本可以内部为同样的产能融资,或者以接近美国政府的利率借款,从而比 Oracle 获得约2个百分点的资本优势。Doug 称放弃这一优势是失误,而 Oracle 的突然发行又助推了 CDS 焦虑,迫使资本主义“踩刹车”。
23. TPU V7 只有一个狭窄窗口来把 TCO 转化为存量基础
Doug 把 Google 愿意向外部销售当前一代 TPU 解读为一次争夺市场份额的决定。在 Gemini 3 之前,如果 Google 自有产品表现不佳,囤积硬件的意义并不大;如今 Ironwood TPU V7 在 Nvidia Rubin 到来前,提供了预期中最宽的 TCO 优势。
双方讨论给出的 TPU 业务规模接近1万亿美元,对应约30%份额。关键在于建立存量基础:拥有当前 TPU 的客户有升级动力,而 AMD 必须反复争取那些此前没有 AMD 加速器、并非在替换旧设备的用户。
TPU 的硬件、网络和软件已经足够成熟,Anthropic 又是能力非凡的外部客户。Doug 认为,如果供应不受约束,Anthropic、OpenAI 和其他实验室都会消耗更多 TPU V7,因为它当前的性价比是“街区里最热门的孩子”。
供给限制让理论需求无法转化为份额。TSMC 可能是最大的瓶颈,而最初 TPU 项目的一些工程师已经分散到行业各处;这个窗口可能只有1—2年。
24. Nvidia 的供应链控制可能关闭 Google 的硬件差距
SemiAnalysis 预计,TPU V8 相对 Rubin 的表现会不那么有利,尤其是 HBM3 对 HBM4,以及 Nvidia 更强的内存扩展能力。更大、更快的内存直接支持更长的上下文和更强的系统。
Doug 认为,如果 GB200 能按时、稳定交付,它会“彻底碾压” TPU V7。它的延期和可靠性问题制造了当前窗口,但 Nvidia 有多种方式修正路线。
Nvidia 像一支 F1 车队,把每个组件都推到极限;Google 则按照自身基础设施文化,设计稳定、可复制的 pod。Nvidia 出现失误时,这种取舍有利于 Google;当最高性能系统干净利落地到位时,则有利于 Nvidia。
Jensen Huang 与供应商的关系本身就是产品的一部分。Doug 提到他与 Samsung、SK hynix 及其他亚洲合作伙伴会面并互相“示爱”:Nvidia 以 Google 管理层尚未明显做到的方式,为 HBM、封装、连接器、密度和未来路线图争取优先权。
25. HBM 让每一颗 AI 加速器都成为 DRAM 短缺的放大器
内存机制始于一个约3:1到4:1的交换比例:在增加工艺步骤且良率不完美的情况下,生产1 bit 的 HBM,实际上会消耗数 bit 的普通 DRAM 产能。Doug 的比喻是,提炼一种新近变得不可或缺的航空燃料,需要消耗更多传统燃料。
这轮需求出现在 Doug 所称的史上最严重 NAND 和 DRAM 短缺之后,或许只有1996年的情况可以类比。供应商的自由现金流大幅转负,停止支出,并推迟建设无尘室和设备,而这些设备的交付周期可能达到2—3年。
HBM 吸收高端产能,KV cache offload 和更广泛的 AI 系统则消耗中间层。衰退期没有增加备用产能,如今“比上帝还多的需求”正沿着每一个内存等级级联。
SemiAnalysis 的结论是,DRAM 价格还可能上涨100%,供应大约2年内难以追上需求。届时,Doug 预计会出现真正的需求破坏,而不是组件成本上升后无痛地转嫁出去。
26. 短缺将配给设备、上下文,甚至老内存
近期需求破坏可能意味着超大规模云服务商推迟边际采购,尤其是在电力导致数据中心延期、建设周期从12个月拖到18个月的情况下。客户会先提前采购并对所有东西重复下单,随后在库存和电力约束追上来后暂停——这正是内存价格崩溃的典型触发器。
消费者也会感受到分配效应。Shawn 建议尽早购买 iPhone,因为手机厂商最终会进入现货市场并传导更高成本;低端手机、游戏 GPU 和其他产品可能直接因价格过高而被挤出市场,而 AI 基建赢得稀缺供给。
CXL 可能在此前败给 HBM 后获得“一次射门机会”。运营商可以收集老 DDR4,把它放入扩展机架,再通过 CXL 连接——因为如今每一个可用内存 bit 都有价值,一种接近消亡的架构因此被重新启用。
100万 token 的上下文窗口可能因此更像“一栋豪宅”,而不是普遍默认配置。Doug 认为,物理内存可能让完整上下文在未来5年甚至10年都维持在今天的规模附近,从而推动“上下文配给”;Shawn 则提出,稀缺上下文可能出现差异化定价,以及递归语言模型的可能性。
27. 固定权重芯片和 CPU 暴露了次级短缺路径
Shawn 问到 Taalas 将模型权重直接烧录进硅片、从而消除反复内存传输的做法。Doug 认为逻辑很有说服力:“加速的方法就是永远不要传输任何东西”,尤其是生产模型通常比前沿训练规模更小,或经过蒸馏。
不确定性在于市场广度。一种高效固定模型可以大规模扩展推理,但权重变化和性能取舍留下了困难的设计空间;Shawn 对多数加速器初创公司仍持怀疑态度,因为此前太多公司没能走到量产。
CPU 面临更安静的挤压。云厂商在2020—2021年购买了约1000亿美元的 CPU 及相关设备,如今在将2年资本开支转向 GPU 后,正接近5—6年的刷新周期;与此同时,代理生成的软件、生产环境代理和强化学习模拟正在提高 CPU 利用率。
Doug 提出一个明确称为“精神分裂式锡箔帽脑”的猜测:近期网络服务不稳定,可能是氛围编程产生的生产错误与老化云基础设施叠加的结果。更确定的论点则简单得多:少量新增需求撞上一个严重投资不足的品类,就可能制造另一轮短缺。
28. Doug 先私下积累,再完成一次干净利落的写作
在 LLM 之前,Doug 认为高吞吐量阅读和综合能力是自己最核心的信息技能。他可以按不同深度读完教材或朋友的博士论文,而写作则完成闭环,迫使他表达大量阅读最终形成的理解。
从2021年10月开始每周发布文章,建立了这一习惯。他不喜欢 LLM 生成的散文,但会用模型进行构思、列提纲,并根据《On Writing Well》的经验进行编辑;文字本身仍然必须来自他自己的思考。
他最好的方法是收集证据、列出大纲、深入思考,然后睡觉。第二天早上会提供一个“新鲜的上下文窗口”:他打开新标签页,一次性完成写作,通常先写到60%—75%,再补齐剩余空白。
29. 6个月的 Continental Divide Trail 让他看清了工具的使用者
2021年,Doug 选择了 Continental Divide Trail,因为在美国3条主要长途徒步线路中,它最让他害怕。他在6个月内走了约2800—2850英里,大部分时间独自完成,并接受了自己可能再也不会遇到同等机会的事实。
机会成本是真实存在的——他认为自己错过了 Substack 增长的重要一年——但这段旅程带来了职业加速无法提供的东西。那是一场真正的“冒险”,包含无聊、恐惧、饥饿、孤独,以及被压缩在一起的人生“最低谷和最高峰”。
回到马斯洛需求层次的底部后,相比保暖、进食和活下去,抽象的信息工作显得“完全虚假”。Doug 更清楚地理解了自己的边界和动机,也让 Shawn 最后的总结显得恰如其分:“自我掌控是你最重要的工具。”
This crap makes mistakes all the time. It is still just like a junior analyst. The analyst goes and does all this really painful information gathering, and you bring it all together to make a good decision at the top.
Historically, what happens is that the junior analyst who I once was gathers all that information, and after doing this enough times, there’s a meta-level of thinking happening: “Here is what I really understand and what type of analysis I’m an expert in. I’m actually very good at this. I consistently have a hit rate. Now I’m the expert.”
I don’t think that meta-level learning is there yet. We’ll see if L1s do it. Everyone who’s spending $1 quadrillion in the world thinks it will. It better happen, right? If you’re spending $1 trillion and there’s no meta-level learning, that’s a problem. But for me and our firm, that massively amplifies everyone who is an expert, because you still have to do something. You can’t just slop it up. It’s very obvious to me when it’s slop.
swyx
Doug, welcome to Latent Space.
Yeah, thank you for having me. After all this time, is it okay if I just call you swyx? I feel like that’s where my brain is. I’ve known you for so long.
swyx
You can call me Mule if you want. Yeah, I mean, it’s been a long time.
It’s been a long time coming. I think I first met you at one of the NeurIPSes. I think it was in Vancouver.
swyx
Right. Yeah, I think it was at some party.
And you were like, “Hey, who’s this cool dude?” I was like, “Oh, okay.”
swyx
[Laughter] Well, I knew about you, and we’ve been internet pen pals for a long time, so it was cool meeting in person. I didn’t go to the New Orleans one. I really wish I had. I love New Orleans, obviously.
Yeah.
swyx
So, we had two New Orleanses in a row. Honestly, we should go back there. Are you guys going to Melbourne, the Australia one, this year?
I haven’t even thought that far out, but that sounds pretty interesting to me. I can’t remember which one there is, but there’s something in Korea this year, right?
swyx
Yeah, I think ICML.
ICML. I think I’m going to try to go to ICML in Korea. I know ICLR is—
swyx
I don’t know, man. There are so many conferences. Honestly, I hate to say it, but I’m not much of a travel guy. I’m glad to catch you. I am traveling to you.
Yeah, thank you. It was fun.
swyx
I really appreciate it.
I did not know I’d be caught in a snowstorm.
swyx
Yeah, it’s funny. I feel like people have been coming here recently, and they keep getting stuck in these snowstorms. This is the first blizzard in 4 years or something like that. Thank you for coming.
Yeah, it’s a pleasure.
swyx
You used to be anonymous. You used to be Value Mule, which is how I know you.
You know what’s funny? Value Mule was the very first one. Do you know how I noticed you? I was just like, “Oh, this guy seems smart.”
swyx
Yeah, I don’t know, dude. I remember noticing you, too. This was in the early, primordial days of Twitter, around 2017 or 2018. I miss those days the most.
Yeah, I remember Value Mule. If you’re even aware of what that is, that’s the deepest cut you’ll possibly have. I have another account, and I actually have a third account, which is my main account these days.
swyx
Wait, which one is that? I don’t have one.
I don’t want to dox your other account.
swyx
Just semi-dox it.
Yeah, it is there. That’s my oldest finance account. I think of it as my legacy account. I want to have some privacy.
swyx
Yeah. So now you’ve gone all in on the brand and everything.
Yeah, I’ve got the brand and everything.
swyx
Same profile picture, you know.
Yeah.
swyx
Let’s do a little bit of the Doug story, because a lot of people hear about Dylan, and I wanted to make this the Doug story—the Fabricated Knowledge story. You used to be a value investor. That’s kind of how you became Value Mule, and you had a mentor or something who nerd-sniped you into semis. Is that the story?
No, actually, I solo nerd-sniped myself. I wouldn’t say I was a value investor—although, for everyone listening to this podcast, I might as well have been value, right? Maybe quality-focused back in the day, but we had this whole thing where we wanted to buy quality compounder companies.
The one that nerd-sniped me, that single-shotted me, was ASML. I fell in love with it in 2018. After ASML, I read about how complicated it was to make these machines, who the people were who were able to make them, and the semiconductor ecosystem downstream from there. It all started with ASML in 2018. I really fell in love with it, read textbooks, and just kept going deeper.
swyx
I was going to pull up Asianometry.
Yeah, that’s perfect. Jon’s a monster, honestly.
swyx
The thing that’s crazy is that he has a whole playlist about it—every single aspect of what goes into it. What’s truly great about it is—
It’s all science fiction. That’s my favorite thing: science fiction exists, other than the talking, perfectly intelligent robot that is the information LLM. ASML is all science fiction. Semiconductor stuff has always been science fiction. I’ve always loved it and thought it was cool. I thought it was the most important thing we ever made, and everything followed from there.
swyx
I don’t know if you know this, but obviously you know I used to be an analyst myself. I covered TMT.
Which is a freaking huge sector to cover. It’s absurdly huge.
swyx
Yes, very large. I was covering Sprint.
Yeah.
swyx
And I covered Viacom.
Yep.
swyx
And then there’s ASML.
Yeah. I feel like the T, the M, and the T are actually 3 completely separate industries. Once upon a time, I think in 2000, they were really close together, but ever since then they’ve split off.
swyx
Yeah.
swyx
My reflection is that I used to be, I guess, our tech-sector guy. I did the flights to Taiwan, and I took those meetings with Credit Suisse and all those guys who would tour you around. I never really felt like I got it, because I was always being filtered through investor relations and all that. I think you have to do what you did, where you go into textbooks and actually learn about the technology.
But it’s really hard as an investor to make the connection to, “Okay, what does that mean for this quarter, or at least this year?” There’s so much foundational knowledge. Then you’re like, “Okay, everything here is taken for granted. It’s already priced in.”
You assume that all the Taiwanese people buying and selling based on the rumors of capacity are pretty well informed. You assume all the people who are TMT investors in the United States are pretty well informed.
I think the foundational difference for me was being young and brash and believing in yourself enough to say, “No, this is something that really matters, and everyone else doesn’t see it.” But what really radicalized me was that I believed Moore’s Law was dead.
I thought, “Oh my God, not only is this cool new technology super hard to make and very interesting and technologically fun to understand, but I also get it intuitively. Everything about the old playbook is about to be thrown out because of this.”
This was a super-mature industry. You’d read the primers about it—that’s how you learned about things before ChatGPT knew everything—and they’d say it was a very mature industry. It used to be really immature in the 1980s, 1990s, and 2000s, but now it was consolidated, and growth didn’t go up much.
A lot of people had this old playbook from the early 2000s. They hated hardware. There was a perception that semiconductors weren’t valuable—or at least weren’t as valuable as software. Software was considered the most valuable thing. Now software is getting shit on, but that’s outside the scope of this.
People thought semiconductors were an old, mature business with nothing new under the sun. Meanwhile, every single day, just making a new chip was science fiction. People took that for granted.
When the science fiction ends because you can’t make the chips as small as you could before, all of a sudden all those free gains you got go away, and you have to think about it.
And what happened specifically for semiconductors is that it created a lot of pricing power, or value, for everyone who knew how to make a good chip. Nvidia is probably the best case. You could talk about parallel compute and all that stuff, but it's not just that they know every aspect of it—from the chip to the networking to the design to the scale-up—the whole thing. In the past, it was just, “The CPU gets better,” right?
I think I had a really deep belief that Moore's Law was ending and everything would change. Coming in with that thesis at the top level made me want to attack every little assumption. Something that really changed as well—and, dude, this is honestly my favorite post I've ever written—was “ChatGPT-3 and the Writing on the Wall in 2020.”
I got an early pitch for Fabricated Knowledge, and I'm like, “Hey, you know, Moore's Law is over. Scaling models seem like a big deal.” If you simplify it all the way through, it's like, okay: supply versus demand. Demand is growing a lot because of scaling laws. Supply is actually slowing down because Moore's Law is completely screwed. That's probably really good for semiconductors, and parallel compute is going to be a big deal.
My conclusion then was that Nvidia was pretty much the only one who was going to benefit. That's my good long-range prediction. I just don't think I would have expected the magnitude. I think that's been the craziest part about this whole story: I had all these beliefs and theses, and I really, really, really believed them.
The reason why I met Dylan Patel is that I felt like he was the only person in the entire world who was as semiconductor-pilled as I was. I remember yelling at him and arguing about all these kinds of things in our DMs and stuff like that.
Was it just online, or in person?
We met online, and we met in person. No, no, no. I've actually only been to Taiwan with him one time, I think. So, look, we just met in person. We yapped. We went to conferences. I think we were both really early to the thesis, with different backgrounds and perspectives.
Dylan Patel is technology-first and, obviously, the technology matters. I have a little bit more of a financial background, but I was always around him. He was the only guy who cared to the same level. So, yeah, the thing that's crazy is that we called it, we were right, and so on. But what still shocks me all the time is the magnitude of how right we were.
You can say, “NVIDIA was good,” right? NVIDIA was pretty good. And then it's like, no, NVIDIA is now the most valuable company in the world.
If you had me read that and say, “You wrote that. You believed it,” I still wouldn't have put that together. I wouldn't have believed it.
This was one of many theses at the time.
Exactly. Yeah, yeah. There were so many things.
What else were you writing at the time that didn't work out?
We can look back, but I'm pretty happy with my long-term track record. I really am. I'm just really surprised by the magnitude of how everything happened. It's crazy to me that CoWoS is not a household term, but is relatively well-known. It was an exotic technology.
All this stuff has been a learning journey: really believing where technology is going, why chips are so important, and then obviously understanding the big scheme of all the things and putting it together. That was the early days, and I think it's all been downstream of that one goated insight, pretty much.
Yeah. Probably a career-maker right there. I love those kinds of decisions—sort of quarterbacking career decisions for other people who are also weighing a bunch of things. I have ADD, and I just chase whatever is interesting, but at some point you really have to choose.
I think one of my skills has always been trend-following and trend-watching. If we're talking about my account, ValueMule, I was always pretty good at trends—being relatively early. I remember loving and being obsessed with TikTok in 2019, and everyone's like, “Why are you so obsessed with the dancing music app?” I feel like I've always been decent with trends.
But when you see a really big wave that you have a lot of conviction in, it's worth going all in. That's kind of what it came down to: I saw this really big wave, and it was worth going all in. So I reoriented my life around it.
Cool. We're going to talk about other trends that you've spotted, primarily the memory cycle, but also optics, which is an amazing story. But we wanted to focus this on the Claude Code launch and the Claude Code anniversary, and you've been a big Claude Code shill.
Yeah, I am. Where's the chart with the 4% of code?
It's actually—go to the top left, this one. Yeah, yeah, Android. You know what's really crazy? We've updated that chart. I think it's 5% now. As you know, it's really easy to generate code now, so that number will continue to climb. It's staggering, the rate at which this is happening.
Let's recap for people. I think this is one of the most important pieces I've read in a long time. You laid it out, and it's weird because I think of you as an analyst. One of SemiAnalysis's alphas is that you're kind of the fun millennial semiconductor firm while everyone else is super boring and old.
What are you doing getting so into Claude Code? Shouldn't you be reading reports and stuff? Tell the story of your Claude Code psychosis.
I think here's the thing: if you want to be good at any game, we're tool users at the end of the day, right? Obviously, this is outside of my job at SemiAnalysis. I have all these other things I need to do to grow and make SemiAnalysis the best research firm ever. But let's say you're a fund manager or an analyst. Your job is to find information edges and new ways to put information together that no one else has done.
I've always thought it's really important to know the most important, weapons-grade tools that you can use all the time—essentially ChatGPT, Anthropic, and all this kind of stuff. I'm an early adopter of tools as much as I can be.
For example, I've been running our case study through Claude Code since it first came out—over a year, I want to say since March or April.
Which case study?
The case study we use when we're hiring a financial analyst for our core research seat. It's basically, “Can you take this company and do some analysis, and give us this format back?”
I've been running it through agentic tools. When agents really come around, they should be able to one-shot difficult, multistep tasks—things that would take a human 24 hours to do, right? I always wondered because there are some good submissions and some bad submissions. We pride ourselves on the case study being good, and I always joked, “Well, they're going to start to beat the worst submissions.”
That was always my baseline. My baseline was: is it better than ChatGPT's agent mode, Anthropic's Claude Code, or Gemini CLI, whatever? So I started running these benchmarks a little bit, and I was very familiar with how good it could be. But then I was like, “It isn't quite there.”
I vibe-coded some stuff on Opus 4 for sure, but they were kind of interesting projects on the side. It was really hard, it took a lot of feedback, and it would mess up. It just didn't feel seamless. I also tried Codex 2 before this, like Windows 5.2, but I never really got it to work seamlessly and agentically.
Of course. So this was recent?
Oh, no, no, no. My most recent awakening—when I was like, “Oh, man”—was probably December 27.
Oh, December 20th. You know, it’s a day—
Something like that. I’m thinking because it’s between the days. I got home from Christmas, and my fiancée wasn’t feeling so well, so I had some time to mess around by myself. There were also 2× usage limits. Oh my God, I miss those days, but now I’m addicted to fast.
I was playing around with these coding agents, just like everyone else in the space should. I was doing simple tests to see if they could make a thing, and it never really one-shotted a total idiot’s thing. Then Claude Opus 4.5 just started one-shotting stuff, and that, to me, was a huge difference. I was like, “Wow, it can just one-shot stuff. I have all these interesting ideas.”
Is it primarily Excel sheets?
No, not primarily Excel sheets. I would say it’s usually a mix of a dashboard, Excel, or something like that. A good example where I think Excel is moderately okay is one-shotting a basic financial model, or just taking information from one place and putting it into another.
It’s not at a human level, but honestly, if you know much about investing and being in the business, is your model being 5% more accurate really ever going to make a good investment decision or not? No. Never. Not once. No one is saying, “My estimate is always one cent tighter than everyone else’s, and that’s why I’m good at stocks.” It doesn’t matter.
Sell-side is ridiculous because everyone is like, “I’m bullish because my EPS estimate is 10% higher than the Street’s.”
And I’m like, “Who cares?”
Well, as you know, sell-side—if we’re going to take shots across the bow at sell-side—one of the reasons why SemiAnalysis has such a successful business is that I think sell-side, as a concept, is very broken. If you’re talking about waves and things that are changing, sell-side is, in a lot of ways, the hereditary child of 30 or 40 years of banking, where a company would go public, so you needed someone to talk about it, issue securities, and sell the stock.
You’re literally selling the stock. You have to be independent-ish, so your ratings are buy, sell, or hold. One of the biggest sales you could do is when a company IPOs: you talk about it so people know who they are. That’s the core original part of the sell-side.
The problem is that all the research has really fallen apart. It’s just not differentiated. A lot of banking regulations have changed, and the primary information process is a 40-year-old business model on its last legs.
That’s one of the reasons why SemiAnalysis is so good. We’re not focused on being a one-cent EPS shop, which I would argue isn’t exactly a skill. It’s just mechanical maintenance. We’re really good at understanding when technology changes and how that impacts everything.
It doesn’t really matter if one EPS estimate is slightly higher or lower. It does matter if AMD’s Helios rack is on time and out of the gate, ready to make tokens on a certain day, because that’s going to be billions of dollars of difference in revenue for AMD. The same is true for some networking technology or some other bottleneck. Being right on the timing and magnitude of those inflection points will make a huge difference in the stocks.
That’s our business. We’re a research firm, we’re independent, we’ve had a really good hit rate, and we care deeply about the technology.
Exactly. I didn’t mean to characterize you as just young and fun. You are young and fun, but you’re also extremely damn good. It’s almost like a triple threat.
I always wonder if it’s, one, that you have a deep understanding of the technology; two, maybe you’re financially literate; and three, there’s this X factor of focusing on the things that matter and ignoring everything else. I don’t know what that is, but that obviously is the alpha.
Yeah, 100%. That’s always been the analyst-PM conversation: there are really only one to three things that actually matter.
Right. Find me those 3 things.
Find me those 3 things. Then there’s all this information. What’s actually what? That’s the hard part.
The thing is, we’re really focused on finding the things that actually matter. This thesis is better than this thesis; this case doesn’t matter, but this one does because now you have a giant opportunity. That’s what the game is all about in terms of research and finance.
When you do so much research, all these different industry parts are so hard to understand. You go to some networking conference and talk to a guy who works at a company, and they’re talking about their new emitter versus whatever laser is replacing it. They have PhDs, and you don’t. Everyone has a PhD at the deepest level, and they’re all doing something different.
You have to understand all these deep parts of these supply chains, but you also have to have a big-picture understanding, because this little part at the bottom of the supply chain is actually going to impact this giant business at the top. It’s all interconnected, but it’s so complicated. Just paying the tuition to show up is very expensive.
I think one way to bridge this for listeners is that this is the complexity of the problem domain: there’s extreme depth and extreme width, and you have to throw human attention at all of it to find what matters. You’re saying you noticed some kind of breakthrough in December, where it was suddenly clicking for you. I really wanted to figure out the tasks that I was nailing and the tasks that I was still not great at.
Yeah, so let me specifically talk about my use case. I’m still a stock guy. I can’t trade or do anything in the semiconductor or AI world, but I do still really enjoy stocks. It’s one of the reasons I’m passionate about it, and it’s probably my defining skill. The people who are really into stocks are lifers. They just love this stuff. It’s an addiction.
Here are all my positions, and here are some thoughts on them. Can you start copy-pasting notes and putting them all together? Claude Code does that. Then I said, “Okay, add it, make the portfolio run some basic risk analysis.” It did that, too. Then I said, “Can we make an investment framework for my investment style, start to grade all this stuff, attack it, and do things like that?” You can do iterative work, and I was like, “Whoa, this is a crazy useful tool. It systemized how I think really quickly.”
Then I thought, “What else can I do with it?” The answer is anything. My joke on the podcast is that it’s all a skill issue now.
I’ve been doing this systematically for every aspect I can think of. A perfect example is this chart. Claude Code is a really big deal, everything is one-shotting, and everyone on the internet is going into psychosis at the same time. How do I actually know what’s real and what’s—
I wonder, right? I wonder. So I’m like, “Okay, I heard that Claude Code puts commits onto your public repository. It says, ‘Signed off with Claude Code.’” I thought, “Why wouldn’t Claude Code scrape all the commits?”
Lo and behold, it pretty much did. I was looking for that signature. I copy-pasted it and asked, “How would you systematically go about doing this?” It did a big BigQuery pull for everything and pulled all the data every single day. The APIs are relatively open.
Then I thought, “Oh my God, let’s see how much this is growing.” The chart goes up. You ask how big it is as a percentage of GitHub, and the chart goes up. It’s a huge deal. I have a cron job updating it every day, and I’m watching it. This is the biggest deal.
I love watching trends. I love watching exponential trends, and I’ve never seen one even remotely at this rate—4% in 2 weeks or something.
Do you have a PR Arena?
It’s a previous attempt before yours, but somehow they didn’t talk about it. They just talked about merge rates, but they didn’t plot it as nicely as you do.
Yeah, and you also want to—
You asked the question: What is this as a percentage of GitHub? This guy didn’t.
Yeah, that's it.
Yeah. And also, the other thing, too, is that I have a lot of those as well. But I thought the Claude Code, because I'm trying to really focus on that specifically.
Yeah. Well, and also, you want to give an example.
Bro, I didn't make that chart.
Opus 4.5 did.
Yeah. Or I think 4.6. I'm like, “Hey, I want you to do this in this style. This is the SemiAnalysis color scheme. I like it to summarize books about visualization and put little style tips in here.” It had me go read like 70 books or something. I'm like, “Give me the—”
It's probably a waste. Part of it is a waste. [laughter] Look, tokens are free. The cost of doing this is nothing. That's the part that's so amazing.
The cost of doing this is nothing. The information gathering and synthesis is like—if it costs effectively the same, doing 70 instead of 3, who cares, right? And so, whatever.
The answer is, I'm like, “Oh, this is too many tokens. You better really summarize this into like 90 tokens or something really basic.” And then you have all the skill. Like, okay, now you can put all that into a skill for how to make charts in the SemiAnalysis format using any kind of data, and then you can systematically just push this out again. I'm like, “Hey, data analyst, please consider all the relationships you can,” and you generate information. I think that one was not generated, which I hate. Honestly, I don't like that one as much because it doesn't have the guidelines.
That was generated. And so, what you can do is just ask it, “Hey, here's all the dates that we have. Can you visually brainstorm with me a way to better represent this information?” It's like, “Yeah, actually, I'm going to generate you a timeline.” You can just do things.
Yeah. I mean, it's your catchphrase, right?
Yeah, that is my catchphrase right now: you can just do things.
People were looking at this from the perspective of people who are coding, and they're like, “Hey, just programming is automated, right?” But all information work is. I would argue that coding is a big subset of all information work.
I think there's a Byrne Hobart tweet or something from forever ago. He's like, “Coding and finance are actually very different types of abstraction, but you are doing abstraction.” Excel is a ginormous abstraction. You're building these relationships and describing what you think a financial thing is worth, right?
I think coding is a little harder, if I'm being honest with you. And you're telling me the hard one got automated. Why can't the easy one get automated? So I started to ask myself, how much can we do?
The answer is, it feels like a skill issue. It makes errors on the margin, but you can force it into—for me, I love using rubrics, right? “Hey, I care about X, Y, Z. Out of 10, score this.” And then you can really do multiple things. It helps with the stochastic dance.
Do you put it all in one prompt, like the task and the rubric for the task, or do you put the rubric after the task is done?
I actually have 2 versions of this. I'm like, “Hey, you can pull all this stuff together. Just run the rubric,” or you can do the task and the rubric. It just depends on how you want to do it.
Yeah. Because obviously, if you put the task and a rubric together, then it can iterate itself. But if you put it after, then it's probably more like you pay attention to the rubric.
Yeah, exactly. And the other part of it, too, is that it will iterate, but context rot doesn't matter. I kind of like it to be separate because the thing is, okay, it needs to be this fresh look at it. You have to think of it kind of like it would perceive anything anywhere, right? Each context window is just opening it up.
I think sometimes, if you do it together, it commingles the information to the point where it becomes biased or susceptible. Opus 4.6, as you know, is super sycophantic. It loves to say, “Yes, okay, yeah, I'll do this for you.” I think having it separate keeps some of that drift away, and that's one of the things that, personally, I like—the results are better. But it's just complicated.
Part of this is really weird because I'm weirdly now opinionated on taste in terms of how you should design things. For example, the context rot thing—until someone explained it, I was like, “Oh my God, thank God someone said it. This is a huge deal.”
There's this meme where it's like—do you see the meme? It's like Of Mice and Men, and at the end of the book, I can't remember which character shoots the other. I never read it.
Yeah. So one character shoots the other guy, and some guy made a meme about it: “Oh, this is after your Claude Code has garbled 5 million tokens. You're like, okay, it's time to put you down.” Because context rot is huge.
So, yeah, this is an example where—
What are your compact practices? Do you aggressively compact manually?
Personally, with the new 1M, I feel like I try to do it all in one context window. I'm not doing ginormous projects.
The 1M is very new, right? Yeah, 1M, very new. Okay. Very—but it's a big deal, too, because your skills and whatever your CLAUDE.md is, as a percentage of 1M, is so much smaller.
Yeah. Yeah. And also, with how the agents are working, the subagents will have their own context window, and then the pasting kind of really saves that big, you know, the 1M. You just want a really high-quality project within that. That’s the best, in my opinion. Compacts just kind of start the compression of the noise.
Mentioning subagents and multi—first of all, I wanted to give a shout-out to this thing from Anthropic Research where they were like, “Here's our production traffic.” They did a report that was kind of like their equivalent of the METR chart.
There's a lot of people saying that software engineering has PMF, but here's the next list of everything else. But what if they're all also just software engineering, right? Software engineering is like 50% right now, but what's stopping it from continuing to go to 80%?
I think maybe what's going to happen—this is maybe a giant take.
It has data analysis in here, which is what you were doing.
Yeah, in my opinion, that is downstream of—so I think how we should think about it is, software engineering might all be downstream of chips, which is downstream—chips is upstream, and then it's AI, and then it's software engineering. It is all an extension of that same compute hierarchy.
I think the teaching of where machine and code intermingle right now is code, and so that's just going to be the leading language that's used to figure out everything else. That's my belief. It doesn't make sense to build—for example, this is a perfect example: Claude for Excel is much worse than Claude Code using Python to use the Excel skills and then deposit into—it's all much worse.
It's much worse.
Even all the work they're doing.
Yes, 100%. Because if you think about it, it's legacy. Why make a car engine fit into a horse carriage? It should just be in a car. It's a backwards-compatibility thing where it does work because LLMs are relatively generalizable like this, but why bother?
That same abstraction of information in Excel is just in there because it's human-formatted for us to understand. And I think that's the important distinction: all of this information stuff, all this software stuff, is just to be consumed by humans. It doesn't matter. If they're just as good at putting the data together, we should be much more concerned about machine-focused software consumption, so the LLM and the agents can put together and synthesize all the information and deposit it, God knows however you want it to be.
I don't need to make a chart in PowerPoint or Excel. It will just deposit the Matplotlib in a chart to me, in an image. Fine. Yeah, Matplotlib.
Matplotlib. Wow. You're trying to use Matplotlib.
Yeah.
Why? Why? You know, it's better—it's better at understanding that code.
Yeah.
So why ever make a chart again?
Yeah.
If it's better—
It's just that it could be inconsistent with the other charts that you do.
Yeah.
I don't care that much about it.
I don't think we would care that much, but I think, one, our new charts are better than our old charts.
Yeah.
And number 2, I think if it increases the speed of information, that matters a lot. And so I think pretty much the new charts will outweigh the old charts because they'll just grow. [laughter]
So, yeah, I think it is a little inconsistent. We have the same watermarking. Honestly, I think it's better than our old formatting anyway.
Well, the first thing this reminds me of is Bloomberg. I was like, “You guys are just becoming Bloomberg.”
Which is a nice view. A couple of things I wanted to double-click on, because this is just a Claude Code brain dump for one of the biggest Claude Code shows in the world, which is subagents and agent swarms. I don’t know if you’ve tried—
I have tried them.
Pick either one you want.
I have a controversial opinion that Claude does not do RL on agent swarms or agent teams or something.
Yeah, it’s just an experiment.
It’s just an experiment. Thank you. Exactly, because the problem is it’s just via prompt, and it’s actually very bad. I think subagents are okay because they usually have a CLAUDE.md to go do whatever, but the agent team is actually really—
Okay, well, we can’t knock it; it’s experimental. So—
Yeah. No, no—what did you try it on?
Well, it was some big data analysis of many different companies with different KPIs into a dashboard, all in one. I was like, “Hey, can you just make this all—whatever—split up the teams?”
Speaking of that, though, you say that, but Kimi K2.5’s agent swarm is actually good. I have also tried that. It’s actually really good. So I did some examples of things that were never available to me, like internal benchmarking of these models: “Here’s a set of problems I’d like you to do 20 times. Can you do them?” Then I can measure the performance between them and do qualitative analysis, like, “What’s the difference between X and Y?” Yeah.
That was completely out of the hands of me, a normal guy, 3 months ago. Now it’s completely available to me. That’s awesome. I care about this stuff, and now I have tools that are able to automate and do a lot of it, because all of software engineering is partially automated.
My experience is that the Kimi K2.5 swarm actually improves the model’s performance meaningfully. The agent team makes it meaningfully worse because there’s clearly no RL done, so it isn’t context-aware of what’s the best thing to do. I think it’s interesting. I like subagents because it’s usually a little cleaner to give it a task and have it come back, but the agent team is just very—
They had some post about how they did it, where there was—yeah, there’s a bunch of RL for this, and I tried it myself. I thought it was pretty cute how they do all these little games and stuff.
Yeah. Also, it’s crazy how much compute you need just to run the swarm. I think it’s like 16 nodes of H100s. Okay.
And you’re just like, “Dang.” So you and I are not going to be running this, and this is just to run it. I’m sure there’s concurrency available. But yeah, I think it’s really cool, and I think that’s the sign of what’s next, because these agents are going to get better to a certain extent. It’s another benchmark to hill-climb, right? But then it’s going to be: how many of these together in a bigger chain can you get to work?
You could argue it’s kind of a scale-out of the reasoning problem, too. How do you get this one agent to essentially get a verified whatever, put it into a bigger process, and do more information work? That’s the next thing, and it’s important to have context windows that don’t garble up into random stuff and are able to do things well enough with token efficiency. I think that’s a huge part of it. Yeah.
So yeah, that’s what our experiments have shown, at least in terms of the agent swarm versus not. I think it’s very clear the agent team in Claude is an experiment. Kimi will definitely do better, but Kimi K2.5 tells you that this is already a perfectly great new place to do more work, completely available to us right now. I think it’s huge, because if these agents get any better, I don’t know—I’m never going to sleep again.
Honestly, Moonshot AI is very interesting, and this is a tangent. We’re not really going to focus on this very much, but you know how the AI tigers out of China were DeepSeek and Qwen, and—
Then you were like, “Well, who are these Kimi guys?” And these newer names—MiniMax as well would matter.
And Z.ai has been around longer, but only recently became much more active. I noticed that Kimi is much more in the productization phase, as opposed to the Qwens of the world and the DeepSeeks of the world, who don’t really care that much.
I mean, Qwen, because of how it’s attached to Alibaba, right? They have a way to productize it, but it’s kind of like the Gemini version: they have so much stuff to do elsewhere, right? But yeah, Kimi’s pretty interesting.
They’re pushing so hard. They got everything.
I know.
They got Kimi Manus, Kimi Claw. Kimi Claw. [laughter]
Yeah, I know—Kimi Claw. Have you messed around with OpenClaw? Because I did. Oh, I remember—what was it first called?
Clawdbot.
Clawdbot. Yeah, dude. I was going to say, it was really euphoric. I was having it read all my emails and my calendar and do all this stuff, and I was like, “Wait, this is really, really prompt-injectable.” This is pretty secure and important stuff.
Claude Code psychosis is good enough for me at this point in time. What I do is just have multiple emails. There’s a safer email to give to bots, and I can let it use that; if it impresses me, then I can upgrade it.
But Clawdbot didn’t impress me. I’m being honest with you: I wasn’t impressed either. That was the reason why people were freaking out about Moltbook. I was like, “Bro, have you actually used this shit? Because even right now, in Claude Code in a relatively focused terminal, it will be like, ‘Oh, blah blah blah.’ Dude, in the .env there is an API I told you to use for this subcase of problems, and it’s in your CLAUDE.md. Please focus up.” It still makes mistakes.
This is not truly AGI, and there is a harness—you still have to wrangle this thing—but it’s not a perfect skill follower. The context in each attention window is going to change; sometimes it’ll be lazy, sometimes it won’t, but it’s definitely good enough to do a lot of information work.
I use our Discord as basically a way to bring information in and out. I saw this too: a lot of people are setting up things they could have done in Zapier with Clawdbot because they’re like, “Well, you know, now I’m using AI,” but they could have done it more securely with Zapier. [laughter]
Okay. I think it’s kind of interesting. But the difference, though, is Zapier. I remember I’ve tried to use Zapier before.
Yeah, and it’s also not very—
It’s also not very good. The difference, though, is—and that’s okay. It’s okay to be early to something and just wrong because you weren’t the one that made it happen, right?
Claude Code, Clawdbot, whatever—all this stuff—the reason why it’s so powerful is that it gets to completion. Zapier, maybe you can get to completion all the time, but it probably took you 8 hours of clicking through things and copy-pasting crap to make sure it all works and is secure. Clawdbot or Claude Code did it in 4.5 minutes, and that’s good enough for me. That’s a faster achievement.
It’s totally okay that they were right, but they were just not the right mechanism, right? You see this happen in information—in the history of computers. I think there’s also an Innovator’s Dilemma thing, where Zapier, as a preexisting business, had this view of the world of automations as very strict, on-rails workflow-type things that their giant user base already uses. They couldn’t really pivot that much. Yeah.
So that’s why I think one of the co-founders left: they were like, “Well, I can’t exist within this—”
You end up becoming—you know, the box will control you. You are—
It’s your golden handcuff.
Yeah, it’s just like your cage. You’re going to act like how you are in the cage. And so, yeah, that sucks. Honestly, I feel like that sucks for—
The framing I have is, “Your priors become your prison.”
That’s pretty good. Your priors—yeah.
I haven’t blogged that yet, but I should.
You should. Your priors become your prison. I like that a lot.
Coming back to Clawdbot, I also want to make this the sort of Clawdbot—
No, no, no. I want to indulge, because that’s how natural conversation goes, and I think people enjoy that, right? And probably that’s the only time we’ll talk about Kimi.
Yeah.
So, do you use hooks? Give me the Doug O’Laughlin Claude Code setup.
I had just essentially a few base skills.
Then I have a lot of APIs, and we've also made sure—and this is all a work in progress as well—to have APIs for some of the SemiAnalysis information. That way, we have an internal server that is accessed by people with an API, so all the SemiAnalysis researchers are able to hit some basic level of context, because I think the context is really what matters. I'm too dumb to be really smart in order to have—well, I guess I do have some hooks, if it makes sense.
I think hooks are very underrated, right?
Yeah, I do think so.
Because you can do a Ralph loop just with a hook.
Yeah, yeah. I feel like I underutilize hooks. I think that is true. But I do run some version of them on skill calls, effectively: on this, you have to start pulling all this stuff. In the beginning, I tried to do all this hook and compound stuff, and I found that the Gas Town Ralph loop era is a sign of what will come, but I just don't think there's enough fidelity to make crazy, multi-turn things happen. So, actually, less is more.
Try to have a strong set of smaller skills with a good amount of context information to be pulled in. Then, at the beginning of every session, ask and focus on what you want to do, so that it prompts the—not, like, a Claude within a Claude, whatever—the goal to finish within this single context window and then get it done.
This is my generalized research thing: I want to look at the price of NAND since 1984 or something like that. This is what I want to do. Actually, no, let me just give you the best example, which is probably not going to work. I would like to fine-tune a time-series foundation model to predict NAND and DRAM prices.
Okay, I'm going to first start by gathering as much information as possible from all this stuff, and then we're going to fine-tune it and evaluate which ones we're going to use. I chose Chronos-2 because of covariates. Try to set this whole project up. We'll make it a Vercel dashboard internally for SemiAnalysis. Maybe we'll externalize it if we want, if it's a good enough product.
Then it does all this stuff, and I just start plowing away. Hey, can you go research? Here's a search API—Serper or Exa, or whatever you want to use—to go look for all these different information sources and then bring it together, right? So this agent goes and gathers all this information. This agent goes and works on considering the fact that the price isn't perfect to do all this fine-tuning on, and then we throw it in.
I also had it—well, what do I use? It showed me which GPU we're renting on an hourly basis. So, yeah, we just pull all this stuff together and fine-tune it. I'm like, okay, cool, how did this work? Then we just have this constant iterative loop until I try to finish something. I got to the point where I was like, okay, this time-series LLM is probably not going to work, unfortunately.
You said it was because of regimes or something else?
I think so. Regimes, yeah. There's no way; it's so messed up.
For a lot of people who are new to finance, this is why I have an issue with all these kids doing stock-trading games with LLMs. They have no idea. They've never studied finance, and they don't know that the past predicts the future a lot until something fundamental changes and the macro shifts—risk-on versus risk-off. They've never heard those terms. I had to explain it to people at Cognition. The rules invert completely: what used to work is exactly the opposite of what you need to do when you have a regime change.
Exactly. It's very, very hard because the other thing, too, is you'd be like, okay, each of these is almost like a one-off on its own.
Right, which reduces your sample size.
Yeah, which reduces your sample size. So then, at the end of the day, you end up being like, well, it kind of just—I guess it's heuristics. Good luck, have fun, right? Here's your checklist to see if it might be over, but you really don't know anything until then.
An example of where this project was helpful is that I'm not going to have the magic LLM tell me what the price of memory is going to be. It was a good weekend project, and I did burn quite a few tokens, but I do happen to have, after all this information synthesis and analysis, all of the memory prices of everything I could possibly find, plus the things behind APIs that we paid for, plus enhanced data sources. I have all the covariates, so, hey, what was consumer sentiment? Every macro thing of all time.
What's really interesting is that I'm going to be like, okay, well, now can you go make a summary of each and every memory regime and what it looked like, what created the beginning, middle, and end, and put that in a dashboard so it's relatable, easy, shareable, and consumable within my firm and company? Yes. I'll probably be done with that today.
Then you're like, well, that's just gathering information—you don't understand. No one's ever done that in the history of time. I know for a fact, as the semiconductor-cycle guy, I've written and done more work on the cycles than I think anyone else has at this point, especially for the older ones, like the '80s, '90s, 2000s, and 2010s.
When I did it the first time, the human grind was that I went and read these old annual reports, put it together, and tried to string an era through it. I went through all of it, like, okay, what was GDP growth? What was it this year? What was all this stuff? You have to make this giant sheet and then make the narratives. No, none of that shit, dude. I mean, this is too much information to gather. It's like a lifetime of work. It's like a PhD project. I did it in a day—2 days.
Yeah. I mean, I think the pushback would be that then you don't have enough expert information to criticize the reasoning that went into the report that you're slopping out, you know?
There is some slop. I definitely agree with the sloppiness. So I think of it this way. Right now—
By the way, that's also essential for you guys. If you get caught putting some slop in front of your clients, right, you have to—
At one point, be extremely AI-pilled and number 1 in the world at applying AI to your productivity. Great. But also, you have to. I think the thing that's really interesting is this whole thing is a game of hygiene now, because I think it's—
This is really hard, and I think about it all the time. I feel very comfortable with doing all this work because, at the end of the day, I've done the work. I have a lot of embeddings in my brain and a lot of information. The vibes that got me in here are tons and tons and tons of information, setup scenarios, and pattern recognition.
But, yeah, you're right. This crap makes mistakes all the time. All the time. I think of it, once again, as a junior analyst, right? The analyst goes and does all this really painful information gathering, and you bring it all together to make a good decision at the top.
Historically, what happens is that the junior analyst, who I once was, went and gathered all that information. After doing this enough times, there's a meta-level of thinking that's happening where it's like, okay, here's what I really understand, and this type of analysis I'm an expert in—actually, I'm very good at it. I consistently have a hit rate. Now I'm the expert, right? I don't think that meta-level learning is there yet.
We'll see if LLMs do it right. Everyone who's spending 1 trillion dollars in the world thinks it will. It better happen, right? If you're spending 1 trillion dollars and there's no meta-level learning—
For me, in our firm, that massively amplifies everyone who is an expert, right? We are a firm filled with experts. It's this hard part where I wonder if, for new people, we will be less lenient in terms of how much AI tools—
Are you, like, junior or new to the firm?
Junior, because you still have to do some of that. You can't just slop it up. It's very obvious to me when it's slop, right? When it's slop and there's no cognition, then it's like—whatever the artisanal last 5% is, that really matters.
But for me, I know inherently what the 5% is. I can fix it right away with some really easy heuristics and time and be like, “Okay, well, this is the last 5%. You fix this. This is what I believe. Just make up these assumptions instead. Press Enter. Okay, cool, we're good to go.” That's the hard part. That's a real hard part.
There is still a human in the loop right now. One day, someday, it'll be superhuman, but I definitely believe where we're at today—it's not there. You'll just compound all this noise, and it becomes garbled, just context rot.
But in terms of the capability that is over here, the human CPU in this agentic swarm is very, very powerful now—a huge, huge, huge multiplier of what you're able to do. For me, that was enough to make me feel AGI-pilled, honestly, because if I define AGI as many common jobs—not ASI; that's religion—can it automate, change, take over, or completely shift a lot of the information work?
Yes, 100%. Data analysis is a perfect example. Every quarter, I want you to find me some examples of information that might be interesting. I just can't imagine that if I were an entry-level worker doing data analysis, an average 22-year-old would murder the hell out of a relatively well-thought-out agentic system. So you're like, yeah, that job actually does seem at risk.
That 4.5 capability is enough that we hit some level where it seems to work and do bigger information work. That's when I'm like, okay, yeah, this does change everything. There are all kinds of mistakes, and it's a new level of hygiene that we have to do. You're going to have to understand what the output of agentic work is.
I catch it making errors all the time. It doesn't always pull skills. You can definitely tell that context windows get dumber over time. It's not AGI today, but it can do these crazy long tasks, and as long as you finish it at the end and deposit it as information work, that's very valuable.
Yeah, amazing. So you do a lot of client visits, obviously. By the way, Transistor Radio is amazing for understanding what your world is like. Are you also Claude-coding your analysts and, you know, on the other side?
I've definitely Claude-coded the analysts. Everyone in the New York office must try it. I really tried Claude-pilling.
I mean, not the SemiAnalysis customers and all that. My perception is they don't adopt any of this stuff.
Okay, so yes and no. Some people are interested, but you have to remember it's relatively more conservative. If you ask any analyst if they're using AI, every single one of them will tell you, “Yes, I use it every single day. Of course. How could I not? This is a vital skill.”
The basic inference that I'm doing is: I am a bleeding-edge adopter. I'm a relatively smart dude who knows what he's doing and knows if a tool is useful or not. I've evaluated the tool, and I'm like, “Wow, this is an amazing tool that I literally pried out of my dead, cold hands.” Even if it makes mistakes, I will be using this for all kinds of work forever.
Then I look around at everyone else and think, most of these guys are enough like me that if they have an opportunity and an edge, they will obviously apply it. They look at this tool and start to use it. If they start to use it and they're thinking like me, they're obviously going to adopt it. I'm like, well, I don't understand why everyone doesn't adopt it.
I would argue—we'll see—in the 24-month view, it will be a base level. I think Claude Code, Cowork, whatever, is going to be a base level of all information work very soon. My friend was telling me how his portfolio manager found Cowork, and he's getting it to read his emails. He's like, “Oh my God, I love this.”
Everyone's moment is going to be a little different, but my moment feels like GPT-3.5 or GPT-4 for me. There's that first time where you're like, okay, I know it made some shit up, but this is better than if I went for hours searching and putting information together.
It can also do analogy. You can say, “Hey, this is the setup.” It has these really strong pattern-matching skills that are really powerful. I just think it hits some level of capability. I can't tell you what it is. It's my personal taste where I'm like, “Oh, wow. This is completely over the chasm of what needs to happen for it to be a very, very powerful tool.”
That's my Claude Code moment.
There's some kind of automation chart—xkcd has this automation chart—and I think we need a version of this for Claude Code.
What's crazy is that this Claude Code thing murders the axes.
Exactly. It just shifts everything to the right or something. What I was trying to figure out is, well, okay, it is maybe dumber, with less human attention, but because you can spin it up so quickly and it can spin in parallel so quickly, and it gets done, you get more turns at the wheel.
Yes.
Whereas as a human, you get 1 turn. You get 1 turn—
But with Claude, maybe you get 3 turns, and the sort of review process is the thinking.
Yeah.
And you just need to get very good at review, or hygiene.
Yeah. I think of it as hygiene. The thing that's really going to be painful, though, is a lot of my expert opinion has been built by pre-phones and now, right? Your attention span—the children are cooked, okay? The attention spans are really bad. I read this really sad thing that we're getting dumber or something.
You see the Coinbase earnings—all the coins.
Yeah, it's so funny.
I think you should just do that. Well, we do deal with some of the SemiAnalysis memes, you know. The thing is, you say some of this brain rot is so bad—which it is, it's terrible—but some of it is also hitting some attention mechanism in my deep, primordial monkey brain.
Stimming you.
Yeah, it's stimming me, and you're like, I can't look away from the Subway Surfers.
You couldn't look away. I had to pause it.
Yeah, yeah. I was literally—well, hey, have you ever been at a bar where they play these weird—there'll be TikTok videos, for lack of a better term, and you just watch and find yourself engaged with it? TikTok bars in New York?
No, not TikTok bars. Not TikTok.
Okay, it's essentially a B-roll channel that they'll sometimes play in public spaces, and you will just find yourself engaged with it. There are certain things that just work. Sorry, that's completely off-topic, but I wonder—this Claude Code pill is very powerful for me. I believe it will shift all of that over massively—the chart. But it's just really weird because if you didn't pay any human cognition to get there, I don't think you're going to be a great reviewer.
One of the reasons what makes that human feedback loop work well is because, once upon a time, you did that. You could say to me, “Yeah, idiot, you're not thinking about this problem in this way. You're missing this. You're not considering this 90%—the 10% tail,” something like that.
I know you said this, but I know that I told you the valuation is the only thing that matters, but it's also a fraud. You can't do both, right? If you think about the analysis stuff, you have to know when your own personal embedded model is like, “Yeah, actually, this one overwrites this one.” That's through learned experience. I wonder if, when we're just reviewing, we won't be building and embedding those assumptions to understand judgment.
Right, right, because you're just checking for mistakes rather than trying to do original thought by just doing the work.
Yeah.
Yeah, I think that is a danger.
Yeah. And that's what hygiene sounds like to me. It's really addicting to press the button over and over and over, but sometimes you do actually have to think.
Have you tried—so, the way to model the meta-learning system is, once a night, you do a batch job: look over everything I've done and extract some learnings. OpenClaw, I think, has this heartbeat, and people aren't excited enough about this because this is the first instance where the agents are just always on, always living, always reflecting.
Yes. SOUL.md, I think, is much more for character and whatever, but HEARTBEAT.md is the crown.
Yeah, I think that's a good way to put it. The powerful thing about all this stuff is that, yes, we know that the context gets garbled. We know that OpenClaw doesn't always do everything you asked it to do initially, but you can see the design patterns. HEARTBEAT.md is a perfect example.
Are all of our tasks every single day actually us having this genius thing, or do we sit down in a single session, finish a single project, get up and get some coffee, then come back? If it's that, you could make HEARTBEAT.md consider the session to session and say, “Hey, meta-learnings,” all this stuff, and have it specialized and focused on one form of doing something.
Doug O’Laughlin
So it actually does have the context of all of it. I’m thinking of a customer service agent or something like that. It does have the context; in fact, it can look at every single time it’s ever happened. That’s information and context no human could ever hold.
You’re like, “Wait, that feels effectively good enough to do a huge information test and have enough context to be able to fetch it. Maybe there would be some verification to make sure it doesn’t totally mess it up.” But that, to me, feels like a design pattern that you can build something on.
And so that’s the vibe: we’ve hit some capability where you can build these much bigger blocks now. Those bigger blocks are not just this single line of code. It might actually be a business. It’s kind of crazy. I wouldn’t have put myself as AGI-pilled. I think 4.5 is actually—
I think my own timelines have moved up a lot. Yeah.
Doug O’Laughlin
Are you guys watching GDPval?
I do, to the best I can, but I’m feeling like I’m mostly just trying to—
Doug O’Laughlin
No, no, no. So, to me, when GDPval came out—I mean, GDPval is basically a broader SWE-bench, let’s call it, applied to every discipline, every white-collar profession that you can model, and that’s above something like 2% to 5% of GDP. That’s why it’s called GDPval.
They had human experts do the tasks, as well as GPTs, and here are the results. Fifty percent is parity with an industry expert. You can see the nice increase from 40 to 41, and since then, obviously, GPT-5.2 and Opus 4.5 have already exceeded that. We’re at 70-something now.
swyx
Yeah.
Doug O’Laughlin
Coin flip. Exactly. You can see that nice increase from 40 to 41, and since then, obviously, GPT-5.2 and Opus 4.5 have already exceeded that. We’re at 70-something now.
swyx
Which means models are consistently better than industry experts at these tasks.
Doug O’Laughlin
Yeah.
swyx
So, to me, this is the AGI definition, isn’t it?
Doug O’Laughlin
Yeah. Yeah, this is—and so I think the problem, though—yeah, I would say that that is the definition. The thing that’s crazy is that there’s this ASI element that people are really, really focused on.
swyx
Yeah, we’re moving the goalpost.
Doug O’Laughlin
Yeah, we’re moving the goalpost.
swyx
But I’m like, bro, the goalpost—I mean, we’ll see if this is actually the machine god and the shoggoth will come and talk to us and vibrate on our same—
Doug O’Laughlin
I don’t think so. Yeah.
swyx
Okay. I’m going to be honest with you: I’m very open. I will change my mind often. This is not something I feel intuitively in my gut today. Maybe it’s the next, next, next thing, but when it comes to the GDPval version of this, yes—
Doug O’Laughlin
Yeah, this is doing white-collar work—the white-collar work, which is most of the—
swyx
Knowledge work. Actually, it’s almost all—not almost all, but it’s a huge portion of all work in the world. It’s like now we just made this massive shift where technology is going to massively change the relationship with all of that, and it’s going to be this 99-to-1 thing.
I don’t know if it’ll be quite that drastic or whatever. Maybe everyone’s just doing leisure. So far, my experience is everyone just works harder. That’s been my experience, but it just feels like a massive moment has happened. The steam engine’s invented, the trains are here, and everything’s going to change in knowledge work.
Doug O’Laughlin
And it’s kind of crazy. There’s a sort of economic cycle from my macro days that I can’t remember the name of—I can’t look it up—but it’s basically that there are these stages of economic development where your economy starts out majority agriculture, then it discovers manufacturing, then it discovers white-collar work, and then it builds a very mature financial sector. These are like a layer cake, all declining over time, with the new things increasing.
So my theory is that there’s this fifth layer that has to open up and start to happen, because I do fundamentally believe we just invent new work.
swyx
I do believe that. Yeah, 100%. Humans are very adaptable. That’s my favorite thing I’ve learned.
Doug O’Laughlin
You’re able to adapt to the coldest place in the entire world and the warmest place. Humans are in every latitude. That’s in a physical sense, but I think we’re going to find a way to make utilization go up. We’ll invent more work for sure.
swyx
It’s happening in our lifetimes. It’s happening right now. It’s really crazy.
Doug O’Laughlin
I think the thing that’s crazy is just how quickly things change, and that 5-to-10-year period—that 10-year gap—can be drastic and crazy. That’s just society. It’s wild, but yeah, it’s happening in our lifetimes.
swyx
I’m really curious about when we start to see it in a much bigger way in the real economy. That’s my pet question.
Doug O’Laughlin
Yeah. Why is it not showing up in GDP yet, right?
swyx
So there are going to be some people who are like, “Oh, the effects of the internet—same thing, information transfer, whatever.” I think I’m actually scared of a third, worse thing, which is—now, this is a complete crackpot theory, so please don’t hold me to this—but what if AI is massively deflationary?
I think one of the more interesting conversations I’ve had in a bit is that GDP was invented once upon a time as a way to figure out how much we could divert from the normal economy to war during World War something like that.
Doug O’Laughlin
My spiciest take is that GDP itself is going to be very, very challenged by AI, because information work—how we capture it effectively—is all an economic good, and then the service is hours divided by hours. So there isn’t a widget-to-widget difference.
But in theory, if we could break all information work down into units, we’re going to have a lot more information work for sure. More work will be done. I don’t know what the value of that’s going to be. Is it going to be such an increase in supply that it’s deflationary? That seems to be a real concern. It’s possible.
And then we’ll figure out how to use it. But there may be a Great Depression of AI where we figure it out.
Doug O’Laughlin
Yeah. Well, I wrote this whole thing about railroad stuff because it’s my favorite—my favorite capital. It’s on Fab. I can’t remember; it’s like “Railroad Fab.” It’s about all the railroad stuff over time.
Pretty much because everyone was first looking at the internet. We’ve massively passed the internet in terms of the absolute size of the buildout. It’s not even close. Like, we—
swyx
What numbers are you thinking about?
Doug O’Laughlin
I think $1 trillion all-in was essentially the real-dollar version, and I think we are well past that. Whatever this year is—and it’s cumulative, right?—we were well past that.
I think railroads are so interesting because, honestly, it’s way crazier. But part of the problem and craziness, too, is that railroads were literally one of the first added layers of the layer cake. Before that, it was agriculture, and railroads were like, “Okay, well, how do we move this agriculture around faster?”
Then banking effectively got invented by railroads.
swyx
Because there was a need to finance it. So much money was needed that effectively 85% of all paper, whatever, was essentially just railroad debt.
Doug O’Laughlin
Yeah. One of my favorite anecdotes was that before there was a Federal Reserve, Andrew Carnegie was the Federal Reserve.
swyx
Yes. Yeah. There were individuals.
Doug O’Laughlin
Yeah. And so, all this stuff—I kind of did some work on the Gilded Age and all this stuff. My takeaway is that it was a really interesting cycle because it was so big and took so long to deploy. It was actually 45 years. There were 3 cycles, actually—3 boom-busts.
swyx
Okay.
Doug O’Laughlin
I don’t know if it’ll be quite that long. All the cycles kind of collapse, and information moves around faster.
swyx
Exactly. Yeah. So you have all this stuff where I think it’s going to happen faster, but I’d be really shocked if it was all in 1 go.
Doug O’Laughlin
That’s my vibe. It’s not all in 1 instantaneous up-down. I think it’s going to look like multiple cycles. I kind of just wrote about railroads. There was a baby railroad cycle, then there was a huge railroad cycle. The modern railroads were invented out of it.
That’s my favorite analogy for this, because I think capex as a percentage of GDP each year was in the high single digits, sustained for about 10 years. What’s crazy is that amount of spending—we’re well on track for that. Did you do the percentage of GDP? Because I think that’s the way you make it comparable.
swyx
Stargate itself is 2% of U.S. GDP. I mean, it’s going to go up.
Doug O’Laughlin
Yeah. Yeah. And it’s not all going to be in 1 year, right?
swyx
So total capex for it was 4.8% of GDP and 25% of total gross fixed capital investment.
Doug O’Laughlin
Okay.
swyx
So 25% of investment every year and 4% to 5% of GDP.
Doug O’Laughlin
I think we’re there. Stargate plus Anthropic plus whatever—we’re right there. xAI.
Yeah. So, we’re at the railroad buildout. The thing is crazy.
We should exceed it.
Probably. Yeah.
Yeah. No, no, not probably. We should. But, okay, this is bigger.
Yeah.
Okay. I would like to say, “Yeah, sure. We will do it.” But I’m worrying. Where are we going to get all the money? That’s such a pedestrian concern.
Yeah, it’s not a pedestrian concern. I mean, that’s what happens every capital cycle. I’m worried, like, we must have hands in the Middle East. You’ll flip the thing. We must—we must—what? This happens every single time. That’s the reason why the bubbles happen, right? We essentially get so big that it’s like, “This must be built. It doesn’t matter the price,” and then all of a sudden we look at it and we’re like, “Ooh, that was a steep-ass price.”
But I think the way I think about the big picture is that there’s a demand curve and a supply curve, and we have no idea when they cross. They will cross one day. Every single year, we’re refining that demand curve, and for the supply curve, we’re just doing our best to deploy it.
For me, I don’t know when that number is. I don’t want to say the number goes up forever, because I feel like that’s intellectually dishonest. But Claude Code is the first time where I’m like—and we’re bringing you all back together—where demand goes up so much that I’m now guzzling it as an individual.
For example, I’m off Max. It’s not enough. It’s not even anywhere near enough. Some people buy, like, 5 Maxes.
Yeah. So, I’m on Fast with $1 million on the API, which is addiction-level, if that makes any sense. I really think it’s the first time we’re like, okay, how much is this worth to me on a yearly basis? I think it’s $20,000 to $30,000 easily, if not more. I can’t price it; I have no idea about the elasticity.
Yeah. You pay for a perfectly compliant junior analyst.
Right. And so, what’s the equivalent at that cost? Like, get $90K—
That’s able to work in parallel. You can have 100 of them. It’s kind of crazy.
It’s a skill issue if you cannot manage a junior analyst that is $20K a year.
Yeah, 100%.
Which—I mean, okay, “skill issue” is like, it’s your fault. But no, we have to learn how to do this.
Yeah. It’s 2 or 3 months old.
Exactly. It is 2 months. That’s the correct way to put it. It was definitely a skill issue that you didn’t know how to get your settings on your iPhone to work. We know how to do that now, but in the very first month of us having it, no one’s going to be like, “Yeah, you idiot, you rube. You don’t know how to use your completely new technology that got birthed last month.”
I think it’s just a bit of time. What’s cool is that if you’re on this absolute bleeding edge, you get to see the design patterns blossom in real time. We have this older guy who’s been through the history of technology forever. He’s one of the most interesting, intelligent people at SemiAnalysis. He talks about how—who is he? Tanch.
We had this conversation one time. He was talking about the early internet, when it wasn’t actually clear that the browser was going to win. Some people thought it was going to be a remote web file service: “I’m just going to reach in and play with someone else’s web files remotely.” Who knows? It kind of is a remote web file service. They were searching for design patterns back then.
I think we’re at that again, where all the design patterns are open. It’s really interesting because there are many different ways this could go. We’re going to have to collectively agree on the best set of hygiene and design patterns, what the level of abstraction should be, and how much SaaS will disrupt everything else. Who the hell knows? But you get to watch it from a front-row seat right now.
Yeah. Yeah. My biggest one—and I do want to bring it to SemiAnalysis in a little bit—is the IDE. 2 months ago, we had Steve from Gas Town talk about how IDEs would be dead, and 2 or 3 weeks ago I was like, “Fuck, he’s absolutely fucking right.”
It’s over. [laughter]
I’m really wondering, too. My daily driver was never an IDE; it was Bloomberg or Excel or something like that.
Excel is the IDE for analysts.
Excel is the IDE for analysts. Bloomberg is the IDE for analysts. I believe every one of these IDEs is done. It’s dead—over and dead.
I just think, why wouldn’t it? Imagine the concept of having an agent with information that can perfectly retrieve and analyze things, with the ability to pull it all together in a better UI than it was before, with no legacy whatever. I think all of that is dead.
My spiciest take of all is that Microsoft has a lot to lose.
I think they have the most to lose of everyone.
Because Excel is a human IDE for information work that’s generalizable. So is PowerPoint, and so is email. Those are the base, core-level abstractions that are broadly generalizable. But I just don’t think that matters anymore.
I think Claude Code, Cowork, or whatever is going to be the year that destroys all of that information work—the work where you sat every single day. It’s over. I think that’s the one that’s more shocking and scary, and people don’t believe it. I believe it in my stomach, with conviction, because I’ve already had that moment for myself.
I will never make a chart in Excel again. I actually believe—
It’s hard to let go because I have so much ingrained knowledge of manipulating things directly in Excel. I have so much with Bloomberg.
There’s no way that you know this, but my very first startup was an attempted Bloomberg killer: Sentieo.
I remember. I remember you.
Yeah. No, no, I—
Oh, yeah. You were one of the few. You had—
Sentieo. I was a Sentieo customer. How does a Sentieo customer roll in, dude?
I remember. How dare they acquire Sentieo?
I had a patent. We filed for a patent for similar tables. Anyway, one of my conclusions was that Bloomberg is just 3 things: it’s Slack, the journalism—which is amazing—and the data feeds. It’s actually not really the UI. But I think, for the first time in my life, I just wonder if that—
Okay, so you’re telling me that the undisputable future is just IB and nothing else, and then a terminal that types in some stuff. I think that if you are marginally curious and not hyperconnected—which I would argue that I am at SemiAnalysis—for example, I’m trying my absolute best to just rip Bloomberg out.
We’re going to FactSet API, like all-in API, with Claude Code. That’s my belief of the future: a verifiable data source that you trust, at scale.
For you guys, you can do it for traders. We’re—
No way. I understand that there’s an information network that’s outside of this.
And you do deals in IB, right? They’re tracked by the regulators.
But as an analyst, yes.
But as an analyst, yeah. And so, I just think that, okay, that doesn’t really— You’re right: the core cash-flow calculation will continue onward. But each iteration of this AI thing, I was like, “Yeah, I’m still going to be using Bloomberg,” right? This is the first time it’s actually, “No, I don’t care anymore.”
The utility—the marginal value—from IB is now outweighed by how clunky this is, and I want to just make some charts, right?
Immediately, you save $10K to $20K.
Yeah.
By switching down.
Yeah. There you go. [laughter]
It’s amazing.
Yeah.
By the way, what was your Claude Code end-of-year prediction? 25?
Yeah. I sandbagged the ever-living fuck out of that.
Oh, okay. I just believe 25 is very— The rate it’s on is whatever, 50 or something like that, but I wanted to give a 95% confidence interval.
Mhm.
I think 25 is within the 95% confidence interval.
Sure. So, it’s between 25 and 50—
Something like that. Yeah.
Yeah. It’s just absurd. But, you know—
It could also be Codex. Are you also watching? To be clear, I’m actually even willing to comment on that because I know we’ve done a lot of shit-talking and been Codex haters.
Yeah. I think, by the way, when I put Claude Code, Codex, and whatever agent all-in percentage we can publicly see, I would argue the ratio outside of that is probably higher too. But whatever. Yeah, I think, together, we're watching Codex. I actually think GPT-5.3-Codex is pretty good, I think. So we had the whole thing—because I wrote most of the articles—I was like, “Oh, token efficiency, the context throttle...” Is this the same one?
Yeah, yeah. It's in the bottom. It's in the paid section. Okay, but TL;DR, I was like, well, you know, the reason why Claude Code is so good, and Anthropic is so good, is because of all this token efficiency. The token efficiency is better than ChatGPT, all this stuff.
And then GPT-5.3-Codex came out, and it was like, “Yeah, that completely doesn't matter anymore.” They're so back. I really think GPT-5.3-Codex is awesome in coding, though.
But you can watch it. The reason why I like Opus 4.6 so much is because, when I'm using it, I'm using it for coding, and the way I interact with it is that I'm using it for broad, generalized information work, right?
But I think the difference is that Codex wants to code because it's RL'd to be so good at coding, to win on SWE-bench. You're trying to use it for general information: “Hey, can you go research and search all these websites?” I don't even think they have web search in it, or whatever. Maybe you can give it an API or whatever, but it's like, “Great, I'm scraping—I'm creating a piece of scraping software to go look at these websites.” I was like, “No, no, no, no. Just ingest tokens of what's on the website.”
It's like, “Okay, great.” I'm still like, it's so coding-pilled from the RL that I think it isn't generalizable in the way that 4.6 is, where it's like, “Oh, I could have it make some rubric or do some research or do something like that,” versus Codex. It's very—
It's very coding-pilled. Codex is coding-pilled, and so that's what... But I am very optimistic, actually, on Codex, and we do track them quite a bit. You can see that they have a meaningful amount of market share on Bloomberry.com. Claude Code is definitely in the lead, but I think part of it, too, is that the like-to-like comparison is off. There's a ratio for Codex that's not available because it doesn't sign off every commit; it does sign off on pull requests. That ratio is much closer. So all the OpenAI people, like Roon, will tell you, “We're not accounting for it.” Yes, we didn't account for it, but—
I think Codex is better. I think there are some real problems and issues, but I bet the second that they have a new pre-training with the RL—because the RL stack on GPT-5.3-Codex is amazing, it's very coding-pilled—that's when it flips over.
And, yeah, look at the other players here. My favorite thing is that GitHub Copilot is number 1. I've never heard of anyone who uses GitHub Copilot. Do you know anyone who uses GitHub Copilot?
Yeah. Look, okay. That's a bubble talking, right? That's a bubble. Yeah, yeah. That's the San Francisco bubble talking. There are all these Windows users, and you don't talk to them, right? You do, but we don't in San Francisco, and that's fine. But GitHub Copilot has $1 billion in ARR, I think—at least.
What's crazy is Claude Code has a ratio—their attribution of Claude Code and ARR is 2.5—
Yes.
So on the daily install counts, right, it's an order of—
Which is, by the way, just the VS Code extension, right? Yeah, I know. That's not even the default way to use Codex.
Yeah, you're right. You're right. CLI, npm, and downloads are other ways to track it, but I think they have their own CLI, their own installer now. Anyways—
All in all, definitely. I understand it's very hard for us to actually track it, but—
I'm not criticizing. I'm just saying, I think the big thing I'm watching is: Is Codex back? They reported that from January to February, they doubled users.
Yeah.
Okay. So I have some—not skepticism, just because they have such a big ChatGPT portal. That could be, like, the modal that pops up can really move big users. They're not quite Google.com in terms of having so much ability to siphon off users, but I wonder. That's my skepticism, but—
I have an answer for that. Alexander Embiricos was just on the Lenny podcast saying that they actually haven't invested enough in the web experience. So I think the attribution for that is zero.
Okay. Yeah. I guess I just saw a modal be like, “Oh, try Codex. Try Codex.” I mean, but a modal isn't—and to be clear, Codex on the Mac is great. I'm actually—I mean, like—
Yeah, yeah. The app launch is actually pretty good. So, yeah, and I think I'm pretty bullish on Codex, honestly, especially for coding, because it's very coding-pilled. I just can't get it to work as well for non-coding stuff. Then, you know, you use Conductor.
No, I've not used Conductor.
Oh, okay. I thought I heard you say on a podcast that you had.
No, no, I've not used Conductor.
So basically, the argument for any AI lab—any first-party app—is that they're only going to prefer their own first-party—
Yes, 100%.
Which they're already doing.
They're already doing it. I feel like this is how they're going to differentiate, right? They're going to—
Well, then you have Conductor, where you can use Codex and Claude Code for different tasks as you see fit. And so this is the clean superset. No—
In theory.
Yeah. But I mean, this is like—okay. So then you can argue this is the clean superset. It feels kind of like... I guess my design pattern on that is really skeptical of building on top of something that is growing very quickly and has all the money and whatever. I just think my favorite one is platform as a service. If you remember when it was infrastructure as a service, platform as a service, SaaS—software as a service—and, like, “Oh, this platform as a service”—and it always just ends up being in the middle. So it just gets eaten by one or the other.
I think of that middleware layer as something that often dies, unless it's a really, really, really compelling case. But that being said, in this moment, I agree. I actually like to have them review each other. Having them yell at each other is really great. I might actually try this soon. I haven't used Conductor personally. I've mostly just been going deeper into the psychosis.
Yeah. And this, as a former cloud analyst, is very typical: do you want multicloud, or do you want to go all-in on one cloud? The classic argument for multicloud is, well, then you can use the best of everything. Exactly.
But if you go all-in on one cloud, you can exploit—
—the sort of minor features of everything, and, you know, the—
It makes a market, and there's no right answer for everyone.
Exactly. [laughter] Yeah, yeah. I mean, even the really small percentages in AI still really matter because they're huge, and people are very happy, very productive, and make money.
Okay. It's good to be an analyst in the space because it's fun to keep up with it, right? Like, I agree. I think everything—
We like the horse race. I like the number 1, number 2. Ooh.
Yeah, yeah. But then your brain also has to be like, number 2 is really big too.
Yeah. No, I know. But then I just think, for me, someone who likes the history of all this—who likes the history of innovation, competition, disruption, and new technology—it's a very fun time to be following this stuff all together. Tech during, like, the 2017 to 2020 years was so boring.
Yeah, at least for me anyway.
I thought it was pretty boring too.
Yeah, yeah.
Sorry, I interrupted you in mid—
No. I was talking about it being just a fun time. It's a fun time to be here. Things are happening.
Okay. I wanted to transition to a little bit of a spicy thing. You were on TBPN, and the title they chose for you was, “Doug O'Laughlin thinks Microsoft is out of AI.” Oh, did you not see this?
Okay, so I wouldn't say out of AI. No, I did. Okay, so I didn't watch it. I never rewatch these things. How I think about it is—
But you said things like Microsoft is scaling back investment, and—
So, so, so, it was the previous conversation I was talking about, yes.
How Microsoft has the most to lose—
They had the most to lose of everyone in the entire world if—
They're the horizontal software company—
Yeah, exactly. They're the horizontal software company that humans use their software to do information work. Okay, no, I cannot paint a bigger target. I cannot paint a bigger—
And Salesforce. Yeah. Well, okay. That's another $200 billion one.
Microsoft is automatically 2× bigger than Salesforce.
Yeah.
But the other thing, too, is they have this Azure business. I don't think you're completely out of the race. It's a really great clickbait title, but the problem is the Azure business with OpenAI, right? You're essentially renting barbarians at the gate. You're like, you know, this is ancient Rome, and you're like, “Hey, we need some extra guys, so we're going to pay money for these barbarians to burn.”
The golden army. Exactly. From Game of Thrones.
Yeah, yeah, the Golden Company. The problem is that each year they become more powerful, and then, at some point, they're just like, “You know, we could just scale these shitty walls.” So that's the problem: the wall and the moats every year are getting more dilapidated as they continue to rent GPUs to the barbarians.
So it's just like Google and Yahoo again. Yeah, it is exactly like that. And so it's just this weird process where that's a terrible setup, too, because what happens in the history of that is you have to choose one or another. If you do either poorly, you're somehow in a third, worse place.
You either go all-in and become Azure, like maybe in the telecom era, right? Because you're a telco guy, you become dumb pipes. Okay, that's the—Azure becomes—what is it?—Charter, right?
Yeah, Azure and Office 365. There we go.
Yeah.
Okay, I don't think you have the answer, but this is the most bizarre—I want to call it a blunder, but I don't know if it's a blunder or not, even because it's a clear decision where they were the lead investors in OpenAI. They had the deal, and they consciously, obviously, stepped back. They're still good partners, but what happened?
I think the biggest blunder of all time—the part that's kind of crazy to me about that one is, yeah, I definitely think there was a financial decision, because when you look at it, it looks like a conversation about shareholders, ROIC, and how much cash you're willing to burn. You look at all the other peers, and Google, I would argue, is going to free cash flow zero. I think Meta will go to free cash flow zero. Microsoft is still—Satya did not make the company. He is a professional manager, and there is a board, and there's a conversation.
Being responsible. Yeah, yeah, he's being responsible, right? But the problem is that responsibility is like an innovator's dilemma, right? Do I maintain and maximize shareholder value and cash flow today, or do I have a deep belief that AI will kill the hell out of my core business and I need to go all-in and invest? Am I ready to bet the entire company on a trend? It seems like Satya is not a believer. We've been talking about AGI; he is not on the ASI pill, okay? He doesn't have any fear of the shoggoth.
He thinks it's just like a new Lotus. It's a new tool. Lotus and Excel came around, right? It's just a new tool.
But I think, at the same time, this conflict between renting GPUs to the barbarians who will disrupt your actual core business—it's clear how they're feeling. On the earnings call, they talked about how they could grow a lot faster if they wanted to, but they're trying to reinvest in internal capabilities. That, to me, sounds like: “We are not going to hire as many barbarians. We're going to pull in. We're going to reinvest in these walls, pull in together, and try to defend the core moat,” right?
Because the dream of this, in theory, is you're like, “Oh, remember in ’23 when they did the first big deal?” You're like, “Wow, Microsoft's going to win it all because they already have all the distribution, and they're going to have the perfect product, and boom, they're going to have this giant business that makes them, you know, whatever, $100 billion, $100 trillion”—okay, whatever number you want to say. But the reality is Claude for Excel and Claude for PowerPoint are literally exactly what they're supposed to be.
Microsoft should have built it.
Microsoft should have built it. Yeah. And so now you see the barbarians, and this isn't even your primary barbarian issue. You know, this is like the tribe over the hill barbarian.
Yeah, this is the tribe over the hill. You know, on a nightly raid, they can easily sack the hell out of your castle, and you're like, “Dang, this is an issue.” So Microsoft is now super stuck in the middle, and how they're going to have to do this is totally different.
I think they're going to keep pulling back in. We're starting to see that: they're going to do internal training, and they're going to try to do more foundational models. They're going to try to use the weights that they have access to.
This MAI. Okay.
Yeah, but I'm very skeptical because their execution has been kind of dismal.
Well, you know, it remains to be seen. They do have—they are one of the biggest companies in the world with all these resources.
Yeah. I always want to push back on the responsibility part. Oracle picked up the slack. Is Oracle being irresponsible?
Actually, if we're going to talk about Oracle, let's talk specifically about Oracle, because this is where we're going to go. I think Oracle was irresponsible because the magnitude of what they did—the thing is, I think they should have done it, but the whole setup, in my opinion, on Oracle is an own goal.
They messed up the messaging. They messed up the fundraising, and, in my opinion, one of the things that happened is they went so aggressive out of the gate and did the quarter where they said, like, $400 billion, right? They said RPO raised the roof. They promised the world.
Then they proceeded to raise as much money as possible, and this is the first time they've ever done these giant build-outs. And so now there are delays. Everyone's like, “Whoa, whoa, you did this much, right?” Capitalism is kind of like, “Hey, hey, pump the brakes.”
Seriously, I think that if they just tiered it out better—meaning that they didn't do it all in one period and played a little bit of expectations management—this year's revenue from the deployed GPUs should partially help start to keep self-funding it. And that's how you make this work on a glide path without going up, down, up, down—a big bang.
I think what really happened is the big bang that really screwed them up was the debt side. They offered so much debt. It's kind of funny because, in high-yield TMT, it's such a big part of the entire index. The issuance is so big.
Debt indexes? I have zero familiarity with them. I'm pulling some numbers up. I did the numbers forever ago; I hallucinated whatever. Forget all the precision. Let's just say all of investment-grade TMT is like $500 billion, okay? I think Oracle is like $135 billion of it. That's so big.
Each time you put up a huge new issuance, you have to give someone an incentive to go buy your debt instead of someone else's. So you're screwing up the liquidity because these issuances are so big, diluting the whole pie. It makes all the terms a little better, or more favorable, for investors. So literally, the entire index is selling off because it's a supply thing.
Yeah, it's a supply thing, right? And that's the thing that's crazy to me: we're at a weird bottleneck I never, ever, ever thought we'd hit. I think you could appreciate this uniquely. One of the bottlenecks is the supply of debt into the market.
Capital markets cannot absorb that much capital demand because the order of magnitude is totally different. These hyperscaler businesses have been completely self-funded since the history of time. They had never gone out and issued anything. The first time they want to, they turn around and say, “Hey, can you give me a $10 billion loan limit? We've never done that before,” right?
So the absolute size is kind of screwing it up. I think Oracle specifically was way, way, way too aggressive in a relatively illiquid market. You have to leg yourself into it if it's going to be like that. But they did these big, huge incremental adds in a super-jolty way and kind of flipped the whole thing.
Oracle CDS people are all freaking out. I think a lot of it's mechanical, specifically because of how badly it was done from a supply-and-demand perspective, and I think they can pay for it. Microsoft could have just internally funded this and, like, fine.
Microsoft could have internally funded this. It would have been totally fine, 100%. And, in this example, yeah, I think that's a blunder. That's a perfect example of a blunder, because Microsoft's cost of debt is the same as the United States government's. It's the cheapest you'll get anywhere else.
Correct.
And just from the math perspective, no one else is better than Microsoft. They, just by their credit rating, have 2% more profitability on a capital basis.
That's like—you can't beat that. I don't know why they decided not to, but now they're in this weird thing where they're wavering. To win, you have to be really bold, right? They're doing one thing over here, being really defensive with Copilot. Satya is now the product manager of Copilot, and then they're also pulling back from Azure.
Meanwhile, the competitors are pushing in for the supply. It's a really weird game. I think Microsoft has to choose a direction.
We'll see. We'll see.
We'll see. And that's what's going to make it fun. I'm more than happy to change all of my opinions when new information comes around.
Yeah. And I'm sure we'll have more information that emerges. I wanted to touch on TPUs and then go into memory.
TPUs will hopefully be a short one, but for a long time you could not buy TPUs, at least current-generation TPUs, externally, and now you can. Google is open as a supplier, I guess.
I think Sergey doesn't want to lose, and I think the thing that happened was—he wasn't there; no one was there. Part of the whole DeepMind story was, “We will hoard all the TPUs because we were first,” and so why give anything to Anthropic?
I think it's because, at least last year, pre-Gemini 3, it was like, “Dude, we have all these TPUs. We're going to hoard them all.” But people aren't using our products anyway, and what's the good of all these TPUs if we're getting our asses kicked in consumer? I think it's an interesting thing, too, because there are a lot of different ways to break this down.
One, we wrote about it in TPU v8. Whatever we think, Rubin will be much more competitive. I think Ironwood v7 is the peak gap in TCO between NVIDIA and TPU, right? If you're at your absolute strongest point, what do you do? There are 2 ways you could do it. You could try to maximize, squeeze the juice, and make margins, or you can gain market share.
You must have done the math.
I've done the math. It could be like—it's like—
A trillion? Yeah, it's like a trillion or something like that, assuming it gets 30% market share or something like that. Everyone has been trying to crack the merchant silicon model, right? And now they have the biggest absolute outperformance. A lot of the people who did the original TPU program are now at OpenAI.
Some of them are—yeah, you're exactly right. Some of them are all over, right? The core team that did most of the engineering has really dispersed. And so I think the gap might close over time.
At this absolute period of time, they're going to win market share. And then what happens is, if you have an install base, you have an incentive to upgrade your install base. That's the hugest problem with AMD, for example. No one wants to buy new AMD chips because it's not like they have old AMD chips. No, they're not upgrading from anything.
And so when you're in that number-two place, you have to win definitively, and then you have an opportunity to win again next year. I think the install-base issue has been a huge one. And so TPU is at the point where the software ecosystem is mature enough, the hardware is definitely mature, and the networking is really mature. You have a really good external customer who actually knows how to use your product. If you want market share, now's the time.
Yeah, that would be insane if they actually pumped the gas on that stuff. Are you also hearing—I don't know if this affects your analysis at all, because I don't have any appreciation for the sizes we're talking about here—but JAX is helping TPUs win, or JAX is winning relative to PyTorch, at least in the academic arena, which is a leading indicator of what it's going to be used in.
I don't have a special purview on that. The thing I'm most excited about—and very much TBD, we'll see—is that InferenceX will have TPUs eventually. That's something we want to do longer term. I think that will really show in the numbers what's—
As a benchmark.
Yeah, as a benchmark.
How do you expect them to come in?
Pretty good on a price basis. Our expectation is that they're the best TCO by a meaningful amount right now. Anthropic is very clear about how they feel. Everyone is very clear. I think even OpenAI would take—I think everyone would eat as much TPU v7 as possible if you had it in a perfectly unconstrained world.
It would probably be, at this exact moment, the hottest kid on the block until Rubin comes out. But the reality is that supply chain really matters, and that just isn't available. So that TCO advantage is at its absolute biggest aperture. Then NVIDIA essentially gets its stuff together, is competitive, and boom, it closes.
This door is only open right now. Probably TSMC is the biggest blocker.
Yeah. Yeah. What can you do? It's this cascade, right, which I think you've talked about, where it goes all the way back to the fabs.
Yeah. Yeah. Well, it's interesting because it's even more than the fabs—it's about the optical side.
Is there a link I should be pulling up?
Yeah, that's it. That's it. TPU v7.
Yeah. So, yeah, it all goes back to the fabs. It all goes to who's making the chips, and I think one of the big differences, too, is the performance. It's just a really cleverly designed system architecture, and it's relatively stable. It's clear that you can pretrain big models on it, which is a huge swipe at OpenAI right now.
That being said, I think OpenAI will get its act together very quickly, and so that's kind of the narrative. I think it's going to be a good story for probably a year or 2, but then the real question is TPU v8. We just don't think it will be as competitive with Rubin, and that's when your special window starts to close.
What's the technical reason why?
HBM4 versus HBM3.
And that's a strategic decision by NVIDIA?
Yeah, I think so. NVIDIA is always—if you think about NVIDIA—they're always trying to gas it as hard as they can. It is a high-performance chip. It is an F1. It is as maxed out as possible.
TPU is kind of like this replicable pod in a very large system, with very high stability. If you know the history of Google, that's what they do. That's what they do with infrastructure.
I think GB200 would have completely mogged v7 if it came out on time and stable. It came out a little delayed, and it wasn't stable. So I think there are a lot of different ways to course-correct that.
The one thing that's important is that, on the supply-chain side, bar none, NVIDIA is the best. They own the entire supply chain. They really do. You think all those HBM price increases are going to come for TPU just like NVIDIA, but NVIDIA was literally in Asia. You saw him drinking with everyone—with SK, with all the Korean guys, with all the TSMC people. He's doing shots with everyone. Why do you think he's doing love shots with everyone? It's because he needs to get the chips.
So, yeah, this is Samsung's chairman.
Yeah, this is Samsung's chairman.
Yeah. And who's the other guy? I—
But let's put it this way. That's a huge deal. That's a huge, huge, huge deal. Do you think Google was out—do you think Sergey was out in Taiwan drinking to get supply? No. 100%.
There's an opportunity here, but there's only so many TPUs that can be made because of all the bottlenecks, right? NVIDIA has all the supply chain locked up. And so they're going to have so much of that constraint there.
They're going to get the best, most performant HBM. They're going to be first on the road maps for even more rack density. They're going to have the best connectors. The whole system will once again be turbo-jammed as hard as it can be.
The people who made v7—the chip was done 3 or 4 years ago. The talent-dispersion aspect, where people who worked really hard on this team to make this great chip have gone all over, starts to get worse. If that gets better, which takes some time, I think our current read is that HBM specifically and memory scale-up are going to really go in Rubin's favor.
That's the big difference. And I think, as you know, that's what makes the context windows able to do bigger—bigger everything. They're really going to jam it, and that's going to be a huge advantage in performance.
One thing I love about your analysis is that it's not actually just the context windows.
It's not just the KV cache. We also have to offload it to non-HBM.
Yeah. Every other part of the memory—it's such an interesting cascade, a waterfall of a short squeeze and everything. It's not a short squeeze. It's a surprise squeeze. I just want to know the ratio.
Yeah.
Yeah. Okay.
It's a 3:1 to 4:1 ratio. I think it's in the “Memory Mania” post that we just put out—the 4:1, or the trade-off ratio. Scroll down somewhere and you'll see.
Yeah. So, basically, for listeners, it's the idea that when you convert to HBM, because there's a huge amount of HBM, it takes 3 times—1 HBM unit is like 3 times the other sort of DDR or whatever, right?
Yeah. Some amount will always be lost in production because yield isn't perfect. Effectively, you're trading some—I actually wrote a funny piece. I called it “Super Oil,” but this is a better one. Pretty much, a higher grade of jet fuel has been invented, and the only way to make it is to get rid of all your other fuel and massively condense and refine it.
Now, if there's any demand here, it's an instant shortage. Hilariously enough, we came out of the biggest shortage ever in NAND and DRAM—terrible, catastrophic, the worst one ever. The last analog I could point to is like 1996 or something like that. Seriously, it's a historical one.
Meanwhile, we have all this new demand—HBM specifically, the highest end, where you need the most memory. The trade ratio is crazy. Each bit of HBM is essentially a 4× multiplier onto DRAM. And now we completely constrained all the DRAM capacity. We just came out of this shortage, so no one invested in clean rooms, capital equipment, or anything like that. People got massively free-cash-flow negative. No one's spending a cent. People could go bankrupt, you know? They haven't invested in these 3-year-long lead-time items, and now there's more demand than God. It also evaporates the middle layer because of the KV-cache offload, and boom, you're looking at the supply and demand, and you're like, “Yeah, this is not going to catch up for 2 years.”
I think the thing that's so interesting is the supply-chain squeeze, because these clean rooms take 2 years to make, man. Effectively, everyone paused, and how bad the last cycle was really forced everyone to completely pause altogether in terms of adding any new capacity. Now we're a few years later and all the supply is gone. It's crazy. Our post's conclusion is that we could see DRAM prices go up 100% again. I think you will start to have demand destruction.
What does that look like?
Where hyperscalers might purchase less, or something like that, on the margin. They're like, “Okay, what if I just really focus on this energy aspect instead?”
Ironically, many—most—of the data centers in America are delayed. You had this thing that's supposed to come online in 12 months; it's coming online in 18. Maybe what you can do is play chicken with memory prices and push it out. Of course, everything you have in the pipeline, you pull forward as hard as you can. You double- or triple-order. Then the DRAM and HBM guys are like, “Oh my God, look at all this demand.”
At some point, you say, “Well, we pulled this all forward. The power is going to constrain us anyway, so we're going to chill out the orders.” Historically, that's when the memory market crisis happens—that's what causes prices to drop. Realistically, just looking at the aggregate demand of how much we've purchased in terms of power, it seems like the gap is huge. It's completely off, to the point where the most obvious, logical leg of the AI trade is effectively investing in memory capacity. Not just SK hynix, Samsung, and Micron—all the semicap companies have been ripping.
Which, by the way, when I was in Basne[?], a majority of the money we made was just being long Micron.
Yeah, it's a good example. You have all the semicap stuff—everything even remotely related to investing in capacity for memory, which is the ultimate bottleneck right now.
And also, for listeners, it's going to affect your phones.
Yeah. Apple—I think Apple's moving.
I had to buy an SD card for this thing. It was 8 bucks.
Yeah, that's nothing, too. That's just the NAND side. Dude, have you looked up 64 GB of DRAM?
I'm moving up. I need to refresh my iPhone.
I'm moving it up because I'm doing this research.
Oh, yeah. You need to buy your iPhone now.
Yeah. You buy your iPhone now, because what's going to happen is when iPhones go into the spot market—
Prices are going to go up 100%.
That's insane.
And so they have to pass it on.
We're going to be buying old iPhones and taking them apart for the memory.
There's actually a whole super-deep, in-the-weeds technology that was very focused on the cloud era called CXL. It's a memory expander for CPUs, in order to have elastic pools of CPU and DRAM compute, or whatever memory is attached. It never really took off because, essentially, HBM was the way that really crushed it all—high performance wins.
But this CXL technology that never really took off is going to take off because they're going to take DDR4, the oldest—every bit of spare memory they can find—and put it into racks, then attach it via CXL.
Oh, exactly that.
Yeah, it's exactly that. But the thing that's so crazy is that this dead technology is having a shot on goal because of how bad the shortage, or how bad the memory constraint, is. I was a CXL bull once upon a time, then it became very clear it was going to die, and I was like, “It's back,” but only because the entire express intent is to take these old DDR chips and attach them to something new. That's what it's going to be like. The memory shortage is just crazy.
Yeah, it's incredible.
So obviously, this is lower level than I usually go to, which is why I'm having so much fun. One thing I do tell people about is that everyone, including Sam Altman, is predicting longer context windows. We've been effectively stuck at 1 million for 2 years now.
I've actually been thinking about that a lot. This is not going to go to 100 million context windows. It's not going to go to 1 trillion. This is it, for 5 years, 10 years, pretty much.
Okay, so the question is, will—yeah, I mean, probably, actually. Will capitalism work? Will there be a way for supply to show up? Probably.
But on top of that, I wonder if there's going to be—like, in his history of compute, what happens is you have to make a curve of the context windows. Does free context go to 1,000? You can use ChatGPT for free now, but your context window is 1,000 tokens or something like that. Then you somehow do a tiny parcel for that so that you can charge 100 times more for 1 million. The 1-million-token context window is like a mansion, you know? That's the real—
You live in a mansion, right?
I live in a mansion right now. Yeah.
Oh my God. The word “context rationing” just came to me. I'm like, “[expletive].” We're going to have vouchers: “Okay, you can have this amount of context today.” It's like, yeah, you have to learn how to use it well because of the DRAM.
So, I actually have a question. Long context to me makes a lot of sense, right? That's the memory-scale-up version—if you think about chips, but in the AI world. I've always been curious because, at least in my stated experience, really long contexts, like you see in the papers, kind of drop off. They actually don't use all the context.
That's what I've been most interested in: does 100 million-context actually matter if it's not possible to use it all? Versions of 100 million do exist today; they just suck in various ways. They're not actually applying full attention, right? You can use state-space models or even an LSTM to process 100 million tokens, but you're not paying full attention to those 100 million tokens.
I think the way we have context today—and those curves—will improve over time, and they have been improving a lot.
But we're never going to use all of them. We'll improve on the algorithm side. I think, for me, what matters is that you represent the physical constraints that those of us on the software side can never surmount, because it's a physical constraint.
And, well, physically, we can’t even double it—say, 10×.
Yeah. What’s the point of talking about that?
Yeah. What’s the point? I was going to say, we could invent a lot of things. Context rationing is pretty good. I really like that one. Context fatality, or a context budget, or something.
I feel like everyone’s going to be like, “Whoa, you’re running out of context window today.” Maybe that’s what happens next year, where we’re charged for context window. One of the more recent things is Recursive Language Models, which again is just reusing the same context window on—
Yeah, over 100 artifacts. I’ve been pretty interested in that. But, to be clear, I’m a total idiot. I have no idea. Claude tells me what’s going on.
You’re the semis guy, man. You’re really good at this. One thing I wanted to spot-check was Taalas.
I haven’t messed around with it.
You don’t have to mess around with it. Just this general theory of custom ASICs burning the weights into the chip, so you don’t need memory.
That’s pretty good, actually. I think that makes sense to me.
This comes at the perfect time.
It does. But I guess, historically, the question is: how big does it scale? A lot of the models are actually smaller than you think, right? So that’s—sorry, what do you mean?
A lot of the production models—
Yeah. They get distilled to [__].
Yeah, they get distilled to [__]. So the push and pull there is going to be: can you just burn in an efficient Pareto frontier, in terms of performance, straight onto the silicon that doesn’t need memory? Then, boom, you can scale this forever, versus the performance edge of the long-context thing.
Just so the compute is even remotely okay.
It makes sense to me, and TBD on the practical implementation. But otherwise, burning their weights into the chip—why didn’t Etched or some of the other guys get there first? Etched is pretty interesting. I don’t know.
Yeah, I’m going to speculate.
I’m not going to super-speculate. The thing is, their thing is: how do we have a big systolic array?
I mean, look, I just think the way to speed things up is to never transfer anything.
Yeah, that’s the fastest way possible. But the bet on this really large systolic array is effectively that everything is compute-bound, right? I don’t think that’s really the case in terms of where we’re actually seeing issues in production markets today. You’re actually seeing all the issues in memory.
I just don’t know if that’s going to be the perfect solution. There is definitely a world and a space—a design space—where they’re going to be very valuable and cool. But the reason my hit rate for every AI accelerator chip is so low—I just don’t believe in them—is because where are they?
Until Cerebras and Groq, honestly, they were all considered failures. And even then, we’re like, “What are they going to do with Groq? What are they going to do with Cerebras?”
Is it SambaNova?
No, I think SambaNova is a much more interesting one, but I think there are all kinds of deal issues with that. I haven’t been keeping up with that one as much.
Yeah. I always try to mention them as part of that cohort. I kind of forget about them, too, but honestly, I was going to say they were—
Once a year, they show up.
Yeah, they do, and they’re not so bad.
You mentioned some CPU shortage stuff. What’s been going on there?
I think—okay, I have one. We’ll start with the conspiracy theory that I think is really funny. Have you been noticing that web services have become really unstable?
This is pure schizophrenic tinfoil-hat-brain stuff, because I have a schizophrenic tinfoil-hat brain. I’m wondering if it’s 2 things. Shipping vibe-code slop to production—that’s number 1. That’s definitely possible. But it’s happening to all the clouds at once. I feel like it’s not just an AWS thing. It’s not just a GitHub or Azure thing.
We’re right at the exact 5- to 6-year period of the refresh cycle of CPUs. During COVID, in 2020 and 2021, you bought something like $100 billion of CPUs and stuff like that. We’re right at the natural end of life for these chips, and usually what you do is have this big refresh of all these chips.
But what’s been happening instead is everyone has essentially stretched all of their budgets as hard as they can. They’ve invested as much as possible in AI and done maintenance capex on CPUs. Ironically, at the same time, with all this Claude Code stuff, if you have a coding agent generate God knows how much compute—how much software—where is that software going to run? On CPUs.
I think we’re going to see increasing utilization, as well as the fact that RL is heavily used for RL gyms. You have to simulate software, and it uses a lot of CPUs. Not quite orders of magnitude like GPUs, but it’s such a big trend that even when it steps slightly in one place, it creates massive amounts of demand.
We might actually be seeing a CPU shortage partially because of this refresh cycle, but partially also because I legitimately believe Claude Code is increasing software creation. And on top of that, there is real demand from—
RL. Yeah. And general production agents as well. Every RLM takes compute, and OpenClaw takes more compute. It’s just a different slope but the same direction.
It’s still an upward slope, and to be clear, it’s had massive underinvestment for the last 2 years because everyone—
How did the same problem happen? Massive underinvestment because they’re like, “Screw it, we’re doing maintenance only. All we’re going to do is maintain the past; we’re not going to add anything else.” And then all of a sudden, just a tiny slope on top of it—boom, shortage.
Yeah. Amazing. Semiconductor guys say, “Semiconductor numbers go up.” [laughter]
That’s one way to put it. The thing that’s crazy is, we talked about the demand—
But you’re right. For sure. Show me where I’m wrong. Definitely not. The thing that’s crazy is memory prices are going to go up so much that we’re going to have to choose what we want. That’s the crazy part to me.
Historically, memory has never been a constraint like this, where actually you’re not going to get your low-end phone, you’re not going to get a GPU this year for gaming—none of that stuff. You can’t do these things because you’re priced out of the market. That’s what’s crazy. That’s the first thing that’s happened in a long time. It’s going to be really interesting to see where that shortage goes and how it’s digested and felt.
It’s amazing. Thank you for that breakdown. I feel like I really understood it talking to you. Let’s transition to a couple of personal things, and then, yeah, as we end. How do you write? Because you write a ton.
Yeah, I do. I’ve been writing a little bit less these days now that I’m in the SemiAnalysis megamind. I definitely write a lot.
And you kept going with Fab?
Oh, wow. Okay, dude. To be clear, that was really—so, look, I’m still trying to do Fab because I do feel deeply connected to writing. Let’s specifically talk about this a little bit.
Yeah. Just explain yourself, you know.
Okay. Before LLMs came around, the thing I felt strongest about—my No. 1 information skill—was that I was able to read, synthesize, and process at really high speed, really high throughput, with decently high comprehension.
The advantage is speed in terms of comprehension—almost anything. When my friend gets a PhD, I go read their paper, and I’m like, “Oh, I have a pretty good idea of what you’re doing.” When I was interested in semiconductor books, I literally raw-dogged some textbooks. Whatever—the comprehension was not very high, but whose comprehension is? I was able to push through these books and learn.
I’ve always loved reading. That’s my No. 1 original competitive skill-set differentiator, and also something I loved as a kid. I was a crazy reader when I was a kid, and always have been.
Starting the Substack, which has been really fun because I just really wanted to get my story out—the things I cared about—closed the loop for writing for me, because I love reading so much. It makes a lot of sense that I love writing. What really helped is that I wrote every single week since October 2021—a consecutive streak for a long time.
The streak has been a little broken as of late. SemiAnalysis plus Fabricated Knowledge is pretty hard to do, but all of 2024, I think, we were talking every single day, every single week. I would put something out.
Was it a hard rule, like one a week?
It was a hard rule: one a week, at least an attempt to do 2. All the people who write about writing say the same thing: you need to just be writing. That's how I started writing every week.
It really helps. What's crazy is that it's kind of hard these days, and LLMs have really changed things. I don't like LLM writing. I do like it for ideation, like making an outline.
Yeah. Here's my unorganized thoughts: make it into an outline. I'll even say, "Put bullet points in the outline," and I'll literally read the outline and ideate and write in parallel. That's how I feel about writing, I guess: write more.
For nonfiction writing, I really like this book called On Writing Well. It's just a really good classic book. It's actually summarized and synthesized into a skill for me.
Oh, yeah, yeah, yeah. "Please edit this. Use this style guide. Use the learnings from this book." Stuff like that.
Okay. Do you have a topic-idea list that you groom? I've put mine in Apple Notes now.
Bro, it's nowhere. I'm just a one-shot-on-your-head kind of person. Usually I one-shot the idea all the way through. I think about it for quite a bit, so it's been bouncing around in my brain. At some point, I've condensed enough information to make a really crappy outline, and that's usually when I just one-shot it and go.
It's hard to one-shot and bounce because you will forget. Sometimes you have really good stuff that you forget. I call this "mess plus writing," where you basically just have a store where you're writing and working your ideas in parallel, and every now and then you cook.
Just like all the little things? Then you search it up when you need it?
Yeah, search it up when I need it or something like that. But I do most of the prewriting in my brain, and I have places where I put things that I reference later.
My favorite tip when it comes to writing is to do the prewriting, think about it, and go to sleep. Wake up to a fresh context window in the morning. That's my number-one piece of advice on writing. It helps so much. If I'm like, "Hey, I need to write something right now," I'll write it all down, make outlines, and do all kinds of things except write it. Then I'll go to sleep, wake up, open a new tab, and write it.
Got it.
Usually that will get me to 60% or 75% of something, even if it's an outline where I've gotten all the ideas down enough to know how to fill it out the rest of the way. That's how I take it from there.
Cool. Amazing. Last thing: hiking.
One bit of context for me is that I've never taken a break. Never. I feel like if you take a break in this time, you're going to be so behind. You're going to miss out on so much. I just found out that my friend from OMI took a break—a year off—to bike through Japan.
How could you? You're going to miss it. You're going to miss everything. But he's like, "I'm good. I'm having kids," whatever. You did a sabbatical as well, and it was pre-AI, but it was interesting. You did the Appalachian Trail—which one?
There are 3 big ones in the United States: the Appalachian Trail, the Pacific Crest Trail, and the Continental Divide Trail. I did the Continental Divide Trail, which is the longest and most remote of the 3. Sometimes it's considered the older, bad one, whatever, but honestly, the PCT and the AT are all different trails.
I'm pretty steeped in hiking culture. I think, mile for mile, the AT is actually the hardest, but I did the CDT as my first trail, as my first thru-hike. You learn a little bit about the 3 when you're choosing which one you want to do, and the CDT was the one that scared me the most.
I thought, "Hey, this would be the hardest, biggest accomplishment I could possibly imagine." If I never have an opportunity to do this again—which so far seems to be pretty correct—which one am I going to do to feel the most like, "Hey, I did the thing that I really wanted to do"? I've always wanted to do a long-distance hike, so I chose the Continental Divide Trail.
I did that in 2021, pre-AI.
But after the GPT-3 essay?
After the GPT-3 essay. Yeah, I felt like I was missing out a lot. It was a huge year for Substack, and I feel like I missed out on a very big year of growth.
You're doing okay.
I'm doing fine. I just think that, for me, this was something I always deeply wanted to do from an intrinsic perspective. I think it's about fulfillment and life fulfillment. I would definitely do it again, but probably—
And, to be fair, for people, it's 4 months, 5 months?
6 months.
6 months?
6 months, 2,800 miles. We'll call it 2,850 miles on the route, or whatever the mileage was. You meet people along the way, but you're mostly alone.
Mostly alone? Did you do it alone? You get a trail name; it's a whole thing.
I did it alone. You get a trail name; it's a whole thing. I listened to audiobooks until I hated them and listened to music until I hated it. I was bored as hell. You just go through all of it.
Yeah, it was awesome. Six months.
I think about how, in most of my life up to that point, you get kicked from situation to situation. You create a view or a form of yourself, think you know yourself, and have ideas about what motivates you and how you react in situations.
With the CDT, I was like, "I like the outdoors. I like hiking. I'm good at it," whatever. It was something that really appealed to me from an adventure perspective. When in modern life do you get to say, "Hey, I'm going on an adventure"?
Never.
That's what it was. It was an adventure for me, and one that I got to really experience. It's like, "The journey is the destination," or whatever. You learn a lot about yourself.
It didn't grow me up, per se, but I feel like I am more well-defined in my view of myself. I understand how I react. I know exactly where my line is. You're like, "Oh, I'll go do this," and then it's actually, "No, I know my exact line. I would not do that. I know exactly where I'm not going—that's too scary, too hard, too whatever." I know my limits a little better. I feel like I know just more about myself.
It is a very condensed version of a very intense life. I wouldn't give up that experience for anything in the entire world. It was extremely personally meaningful to me.
I think it's very fun to go back to the lower part of Maslow's hierarchy of needs. All this stuff we're talking about today is so abstract. It's totally fake, and we were not born and built for it. We were born to scrape a living in the mud.
Hunt and gather and just not die.
It's kind of interesting to go backwards and see what feels real. I was so hungry, so scared, so alone, so low. The phrase is "lowest lows and highest highs." These crazy lows are when you're like, "What am I doing? What does all this mean?" The highest highs mean, "Holy crap, it's so good just to be alive."
All these things—the raw experience of life—is so meaningful, and you don't get to experience it without doing it that way. I highly recommend it. I would do it when you're younger. I wish I had done it right after college instead of getting kicked out for a year or whatever. I think it's good to learn about yourself.
It's really important. Self-mastery is your most important tool of all.
Yeah, self-mastery is your most important tool of all.
Amazing. Thank you for jumping on and covering everything. I feel like I got to go through the Claude Code psychosis, all the way to the semis, all the way to hiking.
Yeah, thank you. Thanks for having us. It was great to catch up.