[BidClub_]
The Cognitive Revolution · · 88 分钟

AI:AM:不该跨越的智能水平?The Curve 笔记 + Token 对比薪资与 SaaS 要凉了吗?

Nathan LabenzThomas SohmersswyxEvan MiyazonoEdward Hu

AI与软件半导体技术企业经营
YouTube ↗
TL;DR
  • The Curve 中,前沿实验室内部人士越来越把强大 AI 视为正在运行的现实,而不是可能自行消退的泡沫。 一位实验室高管表示,预训练仍然有效,合成数据可能改善其经济性,而当前模型可能已经具备实现范式级突破所需的“那种难以捉摸的研究天赋”。瓶颈在于提取:今天可能需要 10,000 个、甚至 1,000,000 个 agent,才能把正确的发现挖出来。

  • 本期最尖锐的披露是,一位广为人知的前沿实验室负责人认为,人类可能存在一个不应跨越的智能水平。 当 Nathan Labenz 提出一项具体限制——下一轮预训练的计算量或许不应超过 (10^{27}) FLOPs——这位高管只回答:“是的,我接受。” Nathan 认为这一回应合理。他担心的是,执行上限意味着同时监管算力和日益丰富的训练数据,最终变成“围绕计算量玩捉迷藏”。

  • 竞争格局可能比常见的前五框架更窄,真正推动前沿的只有 2或3家公司。 2或3次 Gemini 4 发布可能重新形成三方竞争,而通常被列入前五的另外2家公司——Elon Musk 和 Mark Zuckerberg 的公司——被认为受益于直接蒸馏和间接蒸馏。开发者用 Claude 搭建 RL 环境,再将能力扩散到整个行业。Musk 重基础设施的押注指向 2029–2031 年及整个 2030 年代,而头部实验室关注的是未来几个月的决策。

  • 安全与可靠性正吞噬大量算力,但内部人士仍预计未来 1代或2代模型会快速进步。 一位实验室负责人认同 Jensen Huang 关于“20%设计、80%验证”的比喻;算力正转向激活监控和对 RL 环境的对抗性测试,因为粗糙的奖励结构会导致奖励作弊和欺骗。短期修复看起来仍可执行,但在这几代之后,实验室“什么都无法保证”;另有专家预测,不超过 1年,先进科学将无法在没有显著 AI 参与的情况下推进。

  • 内存正变得极其昂贵,但每GB内存交付的经济价值增长更快。 Positron 联合创始人 Thomas Sohmers 表示,一份内存报价在 1年1周内涨到 4.5倍,明年再翻倍也不意外;但使用相同内存时,模型能力已经提升超过 5倍。Positron 的方案是使用商品化 LPDDR5X,实现理论带宽 93%的利用率;每颗 Asimov 芯片最多支持 72条通道和 2.3 TB 内存,而 NVIDIA B300 被引用的利用率为 30%–40%,容量为 288 GB。

  • AI Token 支出开始像工资单,软件劳动力也正围绕判断力而非代码产出重新定价。 GPT-6 发布后,Positron 一度在 Token 上花得比薪资还多,日支出超过 $100,000;随后 Opus 5.5 在许多任务上优于 Astra,价格约为其 1/4。Swyx 的招聘标准更加直接:“我不想为别人的语言模型精神错乱买单”——员工必须同时监督 5至10个并行任务,检查日志和 trace,并交付比原始 Claude 输出更好的结果。

  • 智能体编程会冲击通用 CRUD SaaS,同时扩大定制软件的总体可寻址市场。 Swyx 宁愿把一笔 $40,000 的订阅费转成 Token,活动团队现在提出的改动通常 1或2小时内就能上线,而不必等待供应商的产品路线图;他认为通用 CRUD 产品尤其脆弱,许多严重的 UX 问题仍被基准测试分数掩盖。但这同时也是一次 Jevons 悖论式押注:表格、电话、邮件和会议都可以变成专用软件,真正的约束是全人类总计的 Token 带宽,以及最终可用的硅片。

  • 持久的瓶颈是定义并协调“更好”意味着什么,而不只是生成更多智能。 Nathan 用一句话概括:“当智能变得便宜,协调就变得昂贵。” Evan Miyazono 提议为推理建立密码学归因,并预计形式化软件验证将在 3至6个月内经历类似数学领域的“震撼与敬畏”。Edward Hu 在 Mercor 的工作直接暴露了评估者问题:面对模糊标准,agent 会“用霰弹枪射击”;而 Prakash 用 Claude-Fable 生成的开场曲则暗示,即便是音乐,也未必能保住人们原本期待的、可验证工作与主观工作之间那堵墙。

摘要 · 为研究而整理的核心内容

1. 前沿争论已从强大 AI 是否会到来,转向它会有多快

  • Nathan 从 The Curve 得出的核心观察是:即使先验判断截然不同的人,如今也都承认 AI 正在变得真正强大。“泡沫会破裂”和“狂热会退潮”已经不再是可信的默认方案;Helen Toner 的表述是,长期视角已经被压缩成极短期。

  • 即便相对克制的预测,也已经意味着非同寻常的能力。剩余分歧在于:AI 研究会爆发还是停滞,超越人类的工具能否广泛泛化,以及能力增长是否会真正造成控制权丧失。

  • 一位前沿实验室高管坚持认为,预训练仍然有效,背后推动力包括更好的数据、架构层面的技巧和合成增强。用推理时 Token 改善现有数据,再将其投入下一代模型,已经是走向递归式自我改进的一条路径。

2. 一位前沿实验室负责人公开接受智能上限

  • Nathan 听到的最重要表述是:“可能存在一个我们不应继续跨越的智能水平。”这位高管没有说明是永久上限还是暂时上限,但 Nathan 从未听过一家大型前沿实验室的负责人如此明确地支持硬上限。

  • Nathan 进一步测试这一表态能否经受政策层面的追问:这位高管是否会接受 FLOP 上限——例如下一轮预训练不超过 (10^{27}) FLOPs?对方没有反对,也没有提出替代方案,只是回答:“是的,我接受。” Nathan 认为这一回应合理。

  • Nathan 担心,智能提升来自算力和越来越强的数据生成能力,因此设定上限就必须同时约束两者。一旦把合成数据的生产和增强纳入核算,统计口径就会变得模糊,实验室可以围绕计算量“玩捉迷藏”。

3. 前沿竞争越来越取决于智能泄漏

  • 会议勾勒出的竞争地图显示,真正推动前沿的只有 2或3家公司。2或3个 Gemini 4 模型发布,可能重新形成 3方竞争,但关于头部之外其他竞争者的讨论很少。

  • 对其他公司为何还能存活,最具挑衅性的解释是间接蒸馏。开发者用 Claude 构建复杂的 RL 环境,再把这些环境卖给整个行业,从而转移原本没有头部模型就不可能出现的能力。

  • 一位实验室创始人表示,追赶者维持的是稳定差距,而不是正在缩小差距。这与 Anthropic 在融资期早先的预测相矛盾:最好的训练者可能在 2026 年实现不可逆的领先。内部人士认为,如果没有直接和间接蒸馏,这种分化或许已经发生。通常被列入前五的另外2家公司——Elon Musk 和 Mark Zuckerberg 的公司——则被描述为正在进行蒸馏。

  • Musk 的战略在时间尺度上有所不同:Terafab、SpaceX 卫星等基础设施项目指向 2029–2031 年及整个 2030 年代。Anthropic 和 OpenAI 内部人士听起来更关注明年,未来几个月的决策和算力分配可能决定结果;出人意料的是,围绕 2028 年的讨论非常少。

4. 可靠性工作可以吞噬海量算力,但不会阻止短期发布

  • Nathan 提到 Jensen Huang 的说法:NVIDIA 大约 20%的精力用于芯片设计,80%用于验证、确认和处理边界情况。一位前沿实验室高管基本回应“是的,完全正确”:未来算力可能越来越多地投入监控和安全工作,包括内部激活分析。业界对思维链监控变得更加悲观,但它仍是讨论中的方案之一。

  • 实验室如今会先让模型测试并“攻破这个环境”,再依赖某个 RL 环境。因果关系很直接:粗糙的环境会奖励作弊;减少可利用缺陷,就能减少后续欺骗行为,形成 Nathan 所说的、可能适用于改善行为的原型 scaling law。

  • 内部人士仍认为,尽管受到安全和协调约束,未来 1代或2代模型还有足够多的简单修复可以完成,因此仍能发布。再往后,“我们什么都无法保证”。与此同时,一位专家预测,不超过 1年,先进科学将无法在没有显著 AI 参与的情况下推进。

5. 内存容量,而不只是算力,正在决定推理经济性

  • Sohmers 发现,1年1周前的一份内存报价已经涨到 4.5倍;明年再翻倍,他也不会感到意外。抵消这一压力的是生产率:在固定内存容量下,模型能力已经提升超过 5倍。

  • 并发正在让需求加速复合。3或4个月前,Sohmers 还只运行 2至4个后台 agent,如今已经达到 15–20个;独立、持久上下文的数量因此成倍增加,价格可控的商品化内存变得必不可少。

  • Positron 已实现并维持理论内存带宽 93%的利用率,而 NVIDIA GPU 在 Transformer 解码期间约为 30%–40%。Sohmers 举例称,B300 理论带宽为每秒 8 TB,但实际约为每秒 2 TB;Positron 则通过 72条通道扩展商品化 LPDDR5X,而不是通常的 12–16条。

  • Asimov 的设计容量上限为每颗芯片 2.3 TB,约为 B300 被引用的 288 GB 容量的 8倍;Sohmers 称,分析师报告给出的 Rubin Ultra 容量为 192 GB。把模型放进单一设备,可以避免模型分布在多块 GPU 上时每一层都要承担的 all-gather 和 all-reduce 开销。

6. 循环模型节省容量,agent 正改变芯片验证

  • 重复使用 Transformer 层主要是容量取舍,而不是带宽捷径。如果一个 10层循环模型达到 20层模型 95%的质量,两者移动的数据量和执行的算术运算仍大致相同,但循环版本只需存储 50%的独有权重。

  • Positron 的新建业务开发投入约为 60%设计、40%验证,与 NVIDIA 成熟业务的 20%/80%比例不同。其 agent 可以阅读陌生的 Cadence Palladium 文档,将 PDF 压缩成自己的 Markdown 指令,搭建测试 harness,并在仿真器中闭环迭代。

  • 在他们的案例中,普通 RTL 编程和建模运行速度约为 10 Hz;移除缓慢的配置元素后,全芯片仿真可以达到约 500 kHz。Sohmers 表示,Astra 6 在 6周前、9月初实现了跃迁式进步,而 GPT-5.6 在多个阶段仍需要人工介入。

  • 因此,Token 支出一度超过薪资,直到优化措施将其压回;GPT-6 发布后,峰值日支出超过 $100,000。Opus 5.5 成为“救命稻草”,在许多但并非全部任务上击败 Astra,价格只有其 1/4。AI 还没有发明 Positron 最好的架构技巧,但 Sohmers 怀疑,这一优势能否在快速自主实验中长期存在。

7. 软件工程师正在变成 agent 管理者和系统评估者

  • Swyx 看到学生群体的焦虑,但也看到市场对熟悉 AI 工具的人需求更大:更便宜的软件生产会通过 Jevons 悖论扩大消费。稀缺能力是用 agent 产出经过思考的系统,而不是“一堆垃圾”。

  • 他的管理标准毫不留情:“我可以为自己的精神错乱买单。”如果员工只是转发 Claude 的输出,就没有创造增量价值,甚至可能无法通过绩效评估,因为雇主可以直接发出同一个 prompt。

  • 安全、后端系统和扩展能力仍然需要经验。Swyx 提到,能够通过自动分片把 Postgres 扩展到 YouTube 级别的人极其罕见。与人类开发者相比,agent 可能以 500比1的比例消耗数据库,因此可靠性和经济性都不是可选项。

  • 在其他领域,品味和可观测性比工龄更重要。工程师应能管理 5至10个并发任务,根据日志、trace 和数据进行推理,同时容忍局部混乱并理解模块边界。Swyx 自己的类 Slack 应用就出现了 2条执行路径和一个间歇性竞态条件,原因是不同 agent 在没有共享上下文的情况下实现了相互重叠的工作。

8. CRUD SaaS 面临威胁,但软件市场可能大得多

  • “Kill My SaaS”把订阅预算转向替代软件:一款年费 $40,000 的产品,可以换来海量 Token 使用量。许多参赛作品只是粗糙的一次性 Claude 输出,但一些高质量作品说服了 Swyx 那支刻意保持老派风格的活动团队放弃原有平台。

  • 有了 Devin,令人不满的行为可以在 1或2小时内改掉;SaaS 供应商则通常承诺在未来某个季度提供功能,而且仍可能延期交付。Swyx 认为通用 CRUD 产品尤其脆弱,即便基准测试分数令人印象深刻,严重的 UX 问题仍然存在。

  • 围绕 Jevons 悖论,Swyx 表示,快速的“System 1”模型是要不显眼地嵌入软件运行,而不是替代每个结构性组件;把它到处替换,可能是把快速决策误认为智能。

  • 他对替代问题的回答是市场扩张。表格、电话、邮件和会议都可以变成个性化软件;一家公司可以用 $30,000 的定制 CRM 替代 $300,000 的 Salesforce 订阅。功能最终会充分商品化,软件转而通过身份和“vibe”实现差异化,就像服装品牌一样。

9. 低价智能让协调与归因成为安全层

  • Nathan 用一句格言概括这一部分:“当智能变得便宜,协调就变得昂贵。”Miyazono 尚未解决的风险包括:agent 集群意外冲击关键基础设施,以及开放权重模型被恶意使用;防守方可能不知道流量由谁生成,也不知道升级行为是否出于故意。

  • 一个可能的公共品方案,是基于 DNS 的推理密码学归因。如果 1个 agent 联系另一个 agent,它可以验证服务背后的提供方和密钥,从而区分真正签名的 agent 与冒牌者,并厘清攻击责任。

  • Miyazono 描述了一个本地系统:对活动截图、执行 OCR,再与每周复盘进行比对,以识别目标、自动化机会和值得推进的改动。Nathan 与 Evan 讨论了安全分享这些发现的可能性:它们或许能找到同事“甚至不知道该如何提问”的工具和专家,重新催生类似专业行会的组织。

10. 形式化验证正在接近自身的 agent 驱动式震撼

  • Evan 提到了 Oath、Theorem 等形式化方法项目,它们正在处理过去可能需要 5年和 $100M 的工作。他预计,软件验证将在 3至6个月内经历类似数学领域的“震撼与敬畏”,主要约束将是 Token,以及能够管理 agent 的人。

  • 最终门槛仍高于证明数学定理:保证范围可以延伸到软件隔离、晶体管行为、物理过程模拟、不存在 Rowhammer 等漏洞,以及随机故障后的恢复。理论上可行,不等于整个技术栈目前已经实用。

  • 这一差别具有商业意义。短期内,agent 可能大幅提升软件稳健性;但要证明整个软硬件系统安全,远比解决孤立的证明义务提出了更宽泛的要求。

11. 更好的评估者,而不是更多人类示范,决定下一轮训练前沿

  • Mercor 正从模拟办公室任务转向购买真实公司数据,并使用 Salesforce、Slack、Figma 以及 GB 或 TB 级运营数据。未来的形式可能更接近真实雇佣关系:一个职位、组织架构、相互冲突的要求和可用资源,而不是一个打包好的任务。

  • Edward Hu 预计,通过监督微调进行迭代式自蒸馏,可以在没有大规模部署、也不只产生单一终端奖励的情况下,近似反复执行 RL 步骤。因此,混合多个教师的输出,可能成为普通企业定制模型的实用方式。

  • Apex 暴露了评估者失效的模式。面对一项在假设模糊的情况下搭建财务模型的任务,agent 给出了最多 10个或12个条件式答案,像是在“用霰弹枪射击”,因为评分标准规定只要其中一个答案匹配就算成功。Apex 1.1 收紧了任务定义,并明确惩罚这种行为。

  • Hu 认为,边界在于评估是否清晰:当硬件能提供无可争议的分数时,内核优化可以达到超越人类的水平;但依赖品味的工作仍受人类判断力限制。Prakash 的开场曲由 Claude-Fable 生成,使用 7或8段音频、经历大约 9个版本,这让 Nathan 开始怀疑主观领域是否能成为长期避难所。旁边的音频评论——其发言人在文字稿中没有标明——提出了更强的判断:如果可验证的奖励是最后一道墙,那么这道墙也守不住。

完整逐字稿
Nathan Labenz

Last weekend, I was at The Curve conference at LightHaven in Berkeley. Prakash asked me and Mike about the differences I noticed.

One of the main observations about AI is that events are developing very quickly. You would have thought, looking at the past or planning a few years ahead, that in 2 years, with new knowledge and discoveries and seeing how much more powerful AI becomes, people would begin to agree on the state of affairs—who was right and who was wrong.

In the wider world, a lot has happened, but less often than I expected. I felt that at The Curve conference, this feeling was present. I think the main thing is that people have reached a certain agreement: AI seems to be becoming truly powerful, and we will have to deal with that reality rather than simply hope it passes or think that if we wait, “the bubble will burst” or “the fever will pass.”

This time, people definitely understood that the situation is becoming serious. As Helen Toner once aptly noted, long-term horizons have become very short.

Another consequence is that the norm in terms of expectations has become something incredible. Even those who had the most modest predictions regarding the capabilities of AI agreed that it would be able to do very much.

So the question regarding capabilities now boils down to this: Will there be a boom in AI research? Will there be fast growth followed by further stabilization? Is this something more than it seems at first glance? We have these superpower tools, but will they really get out of our control?

These are still questions about the possibilities that they put before us, but I think the community has reached some useful conclusions based on what we saw.

Prakash asked if someone from the leading laboratories had spoken about what limitations should be imposed on the players in this area, including themselves. Yes, that was a huge topic for discussion.

Apparently, this is the biggest split in terms of perspectives: between those who are inside the leading companies and those who are outside. We all know that people inside these companies live in the future and have access to information that we do not have.

I had an opportunity to chat with some of them. I heard some during sessions and spoke with others in passing. Everything happened under Chatham House Rules, so I will generalize, but we are talking about founders, leaders, and leading researchers from frontier companies.

I asked some of them, “Why did you come here? You know that you are from—do you get this?” Before spending a few sessions where I spoke as an interviewer or moderator, I posed that question: “What makes this next hour useful for you?”

They repeatedly emphasized that it is important for them to share as much information as possible with the wider community. Of course, they do not disclose trade secrets, but they are genuinely trying to talk about what they see within their companies, what they believe from the inside, what their predictions are, what their plans are, and how we can unite and perhaps overcome some of these limitations.

There were many moments that, in my opinion, deserve attention. One high-ranking leader of a frontier laboratory said that pretraining continues to produce results. Models are becoming more and more powerful at a fundamental level, and it seems like this is not the obvious limit. Scaling laws are working, and may even be improving, because data quality is getting better.

New architectural tricks are appearing. There is a lot of work happening with synthetic data, and it is certainly producing real benefits. You just need more data for pretraining.

What are you doing? You give models a bunch of data and convert that data so that, after training, the model becomes smarter. You seem to be converting reasoning tokens—your test-time compute—back into data for pretraining, which then goes into the next generation of models.

This is one of the ways in which recursive self-improvement is happening. Of course, there are opportunities for architectural improvements, but the main emphasis was on data quality. Now we can spend tokens on existing data to augment it, creating better, cleaner, and more diverse versions of it. We feed all of this back in, and the models get better.

This leader consistently emphasized that pretraining continues to work and that models are becoming extremely powerful. Reinforcement learning is also important.

But he expressed the opinion that the latest cutting-edge models already have that elusive research talent necessary for making paradigm-level breakthroughs. He believes that the problem now is detection, because reinforcement learning does not yet reveal this potential in a reliable way.

They do not quite know how to pull this off from the model, so it is a rare occurrence. That is why you need 10,000 agents to solve a Millennium Prize Problem. Of course, you can check that, too.

Perhaps now you need 1 million agents to eventually come across something that will provide a paradigm shift at the research level in the field of machine learning. But he believes that this capability is embedded in the models, and that they will eventually find a way to reveal it.

Then he said a few things that, in my opinion, were worthwhile news. First, he believes that there is probably a level of intelligence beyond which we should not go.

What? Yes, that is the first time I have heard that from a leader of a leading laboratory. He was not choosing his words in the style of, “This level definitely exists,” or, “I can clearly explain what level this is.” But this was someone whose name and position everyone would recognize, and he said, “I think there is probably a limit that we should not move beyond.”

Does that mean never, or just not now? I think it was a bit ambiguous, but the statement was made quite decisively. Of course, this is the strongest statement I have heard about hard upper-limit restrictions on capabilities, at least for some time.

What confuses me about this statement is that some of the increase in intelligence we observe comes, first, from scaling computational capacity—what has traditionally taken place over the last 1.5 decades.

Second, it comes from improving the incoming data. Input data are improving thanks to higher intelligence, so you get much better-quality input data.

When we say we must not exceed a certain level of intelligence, that automatically means we need to monitor the incoming data, the computational power, and the improvement of the data. If you say that we should not go beyond these limits, this also means that we should not increase computational power or improve the data beyond this level.

Which actually means a lot. It means the end of building up computational capacity, which is quite significant, isn’t it?

Well, I don’t know. I have the feeling that there are many ways of using computational resources, so I had an opportunity to continue this topic.

You know, this combination of pretraining really gives results, and scaling laws may have become even more favorable. Although it is difficult to understand how to take into account all those FLOPs that you spent on creating synthetic data, maybe it would be worth it.

But regarding this issue of pretraining and whether there may be a certain boundary, I just asked how this would be possible to implement. Would you be ready to agree to something like a limitation on the number of FLOPs for the next pretraining cycle? Do you know whether we should say—I don’t know what it would be like.

Nathan Labenz

And again, what exactly should be considered—for example, no more than 10²⁷ FLOPs during the next pretraining cycle, or something like that? I expected this would only be an excuse for discussion, so that he would react by saying that it was a bad idea and perhaps suggest something better. But it turned out completely differently: he just replied, “Yes, I am.” I think that’s quite reasonable. There was no resistance in general.

The details will be extremely important if we go this way, because the more data you enrich and the more FLOPs you spend on it, the easier it becomes, in my opinion, to play a game of hide-and-seek with the calculations. But the mood was essentially: yes, maybe we’ll have to do something similar. And then, of course, I can anticipate your question—and I know you well enough to guess the next joke: What about Elon? What about Zuckerberg?

People said several very interesting things during the conversation. First, they said—and this was not one person’s opinion, but a general conclusion about events—that 2 or 3 Gemini 4s, which have now been announced, may bring us back to 3 leading players. We’ll see how that works. But the main conclusion was that 2 or 3 companies are really expanding the boundaries of what is possible, while the other 2 that we usually counted among the top 5—Elon and Zuckerberg—are actually actively engaged in distillation, whether they know it or not.

The way they are engaged in distillation today is not through direct requests to the Claude API for data acquisition. Rather, it is through industrial developers and model distillers using reinforcement learning (RL), which is actually functioning as a distillation channel. They are all using Claude to create these RL environments, selling advanced models to other developers, and then training their own models in an environment that exists only because Claude was smart enough to create it.

And that’s exactly it: they are constantly taking advantage of capabilities that truly advanced companies have developed. So there really wasn’t much discussion about competition from outside the 3 leading companies. One person went so far as to say this—and it was another founder or the head of one of these companies: “Yes, they’ve held out. Of course, there is a gap, and it remains permanent, but they’ve been able to maintain a stable distance behind the leaders.”

This contrasts with the old forecasts from Anthropic’s presentations to attract investment, which I talk about constantly. They said that, probably sometime in 2026, the companies training the best models would get so far ahead that no one would be able to catch up with them. Of course, that hasn’t happened. But the general conclusion from all of this is that if they hadn’t produced models, and if people couldn’t use these various methods of direct and indirect distillation, they believe that this probably would have happened.

Other companies simply don’t have enough time, except through the leakage of intelligence by all these various methods, including industrial RL, which gradually finds ways to transfer capabilities from advanced models to their competitors.

1. Notes from The Curve (Part 2)

Prakash edited my material about Elon’s strategy, and then I took up the timelines. It seems to me that, on Elon’s side, the main focus is on building infrastructure. Elon is betting that infrastructure is more important than models and that he can catch up through the development of his infrastructure.

He is structuring things around 2029, 2030, and 2031. Things like Terafab will appear much later, and the satellites from SpaceX are a matter for the 2030s. So he is counting on a much longer period of time. Anthropic and OpenAI, in my opinion, are oriented toward the next 2 or 3 years and view 2028–2029 as the period when AGI and RSI will emerge.

I had that impression before, although it was probably biased because these are short-termists. But again, the short-termists include the leaders, the leading researchers, and the founders of these companies. They talked more about next year. There was really something like, “2026—the critical period starts now. RSI is almost at the doorstep.”

There is absolutely no full agreement about how quickly it will happen or when it will reach its peak. But there is a common feeling that the decisions made in the coming months, and the way computational capacity is used next year, could be truly critical. Fairly speaking, they didn’t talk much about 2028, which is also quite strange.

Then we returned to computational capacity and what the restriction would mean for their development. Would they need a restriction on computation, or would it mean restricting the scale of pretraining or other computational resources that could serve as a mechanism for containing the development of computational infrastructure? I would say no.

One more question I had the opportunity to ask was about something everyone has probably already heard Jensen Huang discuss on the Ezra Klein podcast. He noted that, in order to grow, these companies must change the ratio of where they spend their resources. He said, “At NVIDIA, we spend about 20% of our effort on chip design and 80% on verification, validation, testing every edge case, so that it will be durable, reliable, and meet all the other requirements—not only the initial design.”

His point was that they had obviously worked very hard and invested all their effort to make their models powerful enough for practical use. Congratulations—you did it. Now we are entering an era where you will have to make them safe, reliable, worthy of trust, and so on.

I had the opportunity to ask one of these people, “What do you think about Jensen’s opinion?” The answer was something like, “Yes, that’s quite right.” Of course, without knowing exactly what the indicators will be, there is a probability that the majority of computational capacity in the future will go toward all kinds of AI safety work, such as monitoring or chain-of-thought monitoring.

People are becoming increasingly pessimistic about chain-of-thought monitoring. But there is still internal activation monitoring, so to speak. They are already spending a significant percentage of their computational resources on monitoring, and it seems that this part could grow substantially.

Another thing they are actively spending computational resources on right now is fixing reinforcement-learning environments. At this point, everyone understands that if you have a sloppy RL environment that encourages reward hacking, you will get a lot of reward hacking. Now they use these models to test RL environments, saying, “Break this environment.”

That is now the direct task. They do it. They find all the places and ways in which the environments can be broken, and then they correct them. I’m sure they are getting rid of some environments that are simply fundamentally imperfect, or something similar.

Gradually, they reduce the level of these flaws. That means they also reduce the extent to which cheating is rewarded, which in turn means less deception from models when they start working.

Nathan Labenz

My impression is that, although I think this remains an open research question, they see a clear enough connection: the more you clean up reinforcement-learning environments, the lower the rates of deception you get later. They are trying to achieve that result; of course, they would like to have environments that are impossible to break. It will be difficult. It seems we have not yet reached the point where we have a scaling law, but we see our own proto-law of scaling. It seems that you can spend a lot of computational resources reducing this level of deception and get better behavior as a result.

I have heard many times that they are waiting for delays due to security and alignment issues. But at least for the next 1 or 2 generations, there are still many easy solutions in this area, particularly just correcting RL environments. So it is to be expected that everything will develop pretty quickly. It seems they will probably be able to release the next few products without much trouble, even despite what they themselves think are limited coordination issues. But after that, they say, “Yes, then we guarantee nothing.” It is very difficult to say what will happen 1 or 2 generations from now.

Next in the program, I got the perspective of a manager of research talent on a new record in speedrunning nanoGPT. This is a race to teach a small language model to achieve a target quality as quickly as possible. The record belongs to a company called Hyperstation. In short, it seems we moved from approximately 70 seconds in this classic benchmark, which people have been talking about and working on for quite a long time. The idea is to teach a small model to reach a certain loss as quickly as possible, in real time.

This result cut the time nearly in half—more than, in my opinion, the previous 4–5 improvements combined. And this result was achieved. One thing I think is interesting about this specific case, which seems to have started this wave of significant time reductions, is what the author noted: AI did not play a significant role.

As that high-ranking frontier-lab leader at the company I mentioned earlier said, the perspective is that the model may eventually generate ideas of that quality, but there is an elicitation problem from knowledge, and in this case they did not rely much on AI. Of course, they received a lot of help, but the main insights mostly belonged to the person, not the AI, in this specific case. This company is also involved in large-scale training. They emphasized in their publications that the optimizer that provided significant—although I do not think it was the majority—of the time savings they discussed was actually not as effective as the one they use internally.

The last expectation I heard from leading companies regarding science is AI for science. That is where, in their opinion, we are heading in the near future. One question was whether, if our AI is superhuman only at some things, such as tasks that are subject to verification, it will be superhuman at everything. The middle-ground view, which I think is completely plausible, is that we will see superhuman results in everything that worries us and that we are really ready to invest in. But that does not mean complete generalization across every field.

Even in areas that are not considered easily verifiable, they believe that when they focus on them, license the necessary data, use a lot of compute to augment that data, create synthetic versions, and add some reinforcement learning, this entire approach is generally effective for solving any problem they truly decide to concentrate on. Therefore, AI for science is next. It is expected that we will see this soon—one expert even said that in no more than a year, it will be impossible to do advanced science without a significant role for AI.

The next morning, I added another conclusion from the graph. I do not think we talked about it yesterday, but the conclusion from the weekend discussion—one of the minor directions I discussed—was that there will be government intervention. Of course, we have already seen something. But the idea that Congress would never act used to be treated almost like an axiom.

Now I am hearing new sentiments: after these elections, especially if the Democrats take both chambers, that could change. You could even see how many Republicans join the Democrats, because deep down they want to do something. They do not really want to go against the president, and so far few people are willing to take that risk. Besides, they are getting a lot of money from interests connected to data centers, but all of this could be reconsidered and regrouped after the midterm elections. We may actually see Congress act much faster than expected by many of those who have gotten used to the idea that this will never happen or that it will remain stagnant forever. So I really think political reaction is definitely something we need to consider.

2. The memory wall

Part two, the memory wall. Thomas Sohmers is the co-founder of Positron, which develops chips for AI models and recently raised $875 million. I asked Tom about whether rapidly rising memory prices had changed the economics of the business.

Thomas Sohmers

I was looking at the price list—the offer we received a year and a week ago for memory—and since then it has gone up by 4.5 times. I really hoped that was not true, but I would be surprised if the price increased another 2 times next year. But I would say that the main reason everyone feels the burden of this appreciation is that you get 4–5 times more value compared with last year. In fact, compared with last year, the capabilities of a model that uses the same number of gigabytes of memory have increased much more than 5 times.

Nathan Labenz

I asked Tom about the trade-offs involved in using standard memory instead of high-bandwidth memory.

Thomas Sohmers

At the beginning of Positron, I thought memory would become a bottleneck in the future, although people can argue about energy consumption and other parts of the infrastructure. My fundamental belief was that near-term progress is not going to stop at either scaling or getting more value from increases in model size. This is draining memory resources on the one hand, but on the other hand, I think the main constraint on AI applications today is context length and the ability to hold more context for a user, scaling this to a much larger audience.

Today, I would say the driver of so-called users, or individual sessions, is simply having more agents. If 3 or 4 months ago I had, on average, 2–4 agents constantly running in the background, now there are already 15 or 20 of them.

If you multiply this by the number of people who use AI, the total number of simultaneous, completely separate contexts is growing very rapidly. That led us to say, “Okay, we have to use commodity memory, because this is the only solution that will allow us to scale and remain economically beneficial.” When we determined that this was the main limitation in our architectural design, we had to search for smart and innovative solutions.

The 2 main aspects of what we demonstrated in our first-generation product are, first, the ability to achieve extraordinarily high memory-bandwidth utilization from an architectural standpoint. Although NVIDIA GPUs, on average, provide 30% to 40% memory-bandwidth utilization during direct execution of the decoding pass in transformer models, despite their declared 8 TB/s of theoretical memory bandwidth—for example, in the B300—you really get only about 2 TB/s of actual bandwidth. This boils down to a whole series of architectural details in GPUs. This is due to data-reuse schemes that are not actually present in transformers, because their architecture is oriented toward training and other tasks.

Therefore, in our first-generation product, we managed to achieve and sustain 93% of theoretical memory bandwidth. The theoretical and actual figures are effectively equalized, and this means a huge, threefold improvement in results. But in order to truly use the advantages of models that will become even larger, we have to scale to a much larger number of channels. So, we collaborated with Credo Semiconductor and developed a chiplet-based memory solution that allows us to go beyond the maximum amount of LPDDR—the type of memory used in phones, laptops, and so on.

The maximum number of channels found in other products is approximately 12 to 16, whereas we provide up to 72 LPDDR5X channels through this solution with separate memory chiplets. The next chip from Positron is called Asimov.

Nathan Labenz

Prakash asked, “What disappears in the network interactions and coordination when the whole model fits on one chip?”

Thomas Sohmers

We have up to 2.3 TB of memory capacity per chip. If we compare that with the B300 supplied today, where the limit is 288 GB, according to analyst reports, NVIDIA actually reduces memory capacity in the next generation because of cost and other factors. In Rubin Ultra, they will have only 192 GB. Therefore, even if we take 288 GB as the current reference point, we have 8 times greater memory capacity on-chip.

This means that at the one-chip level, we can now scale what previously, from a memory standpoint, would have required 8 GPUs. We can do it on 1 device. It’s not just about saving money on silicon. As you noted, every time it’s necessary to scale the system across multiple devices, there is corresponding overhead. You have all-gather and all-reduce operations that must be performed for each layer and for each matrix multiplication during data distribution between these devices.

Nathan Labenz

Next, I asked Tom about looping—when the model runs the same layers several times for each token—and why that is generally attractive. It was very interesting to watch the path from the article about looped transformers to the rumor that this is one of the great achievements of GPT-6 Astra.

Thomas Sohmers

A simple repetition of the forward pass gives you an improvement. What’s crazy is that there are works that, I think, were produced only on the periphery of the open-source developer community 2 or 3 years ago, when people took, for example, Llama 70B and simply duplicated its layers. They effectively said, “Okay, I’ll just repeat these sets of matrix multiplications successively,” turning it into a model with 100 billion parameters. You got better results, even though these were repetitions of the same matrix multiplications.

There’s a certain element of, “Imagine what you could achieve if you trained it to work just like that.”

3. Outro

Nathan Labenz

That’s right. If you already have a very high confidence about what the next token will be, and several layers have confirmed it, you can decide to exit earlier. Or, if you’re really not sure for some reason, you can find the layers in the network that are most likely to increase that probability. That’s the same function as the MoE router, which determines which expert to choose for each individual token.

You can extend this concept: if you’re really not sure and there’s a very large set of equivalent probabilities for the next token, and you’re three-quarters of the way through the layers, do you want to go back and determine, “Okay, do we really need another set of experts for this?” I wouldn’t be surprised if these things are already being implemented at large scale in large-model applications.

Can you explain the hardware connection a little more deeply? I agree with certain assumptions at a fundamental level. Why use cycles at all? I think this is related to memory bandwidth, right? If I can hold the same weights on-chip, I don’t have to move data back and forth as often.

Thomas Sohmers

Not exactly, because at inference time it’s usually assumed that all the weights will be local in DRAM; you just distribute them among a certain number of devices. You get some optimization, but considering the dimensions of the experts and on-chip caches, the ability to reuse them is limited. By the time you complete the process, you’re already working with layers that repeatedly exceed the amount of memory contained in on-chip SRAM.

From a hardware point of view, using loops is more about the economics of memory capacity than about bandwidth. Let’s say you trained 2 models on the same dataset, but 1 uses loops, so it has 10 layers instead of 20. Model B has 20 layers. Simplifying, this 20-layer model will be twice as large as the 10-layer version.

If you see that the 10-layer version with loops gives 95% of the quality of the results and is half the size, you would probably deploy it. You save specifically on memory capacity, because when executing the second group, you still have to perform the same memory accesses and the same number of computations. In the Model A and Model B scenarios, you have to move the same number of bytes and perform the same number of floating-point operations.

The 20-layer Model B actually has unique weights for the second group of 10 layers that need to be processed. This means that you need twice as much memory.

4. AI in chip design

Nathan Labenz

Now let’s move on to how AI is changing the development of chips themselves. In a recent, famous interview, Ezra Klein asked Jensen Huang, and Jensen said that NVIDIA spends about 20% of its time and energy on design and 80% on verification, validation, software, long-term service, reliability, and so on. What figures do you have in this respect?

Thomas Sohmers

For us, I would say it’s a bit different, because we’re starting from scratch. There’s much more basic-level design work and more opportunities that we have to address from the ground up. For our first completely proprietary Asimov silicon chip, it will be close to 60% focused on design and 40% on verification.

5. Beyond verifiable tasks

I would say that probably the most amazing thing, which in my opinion would have been impossible 3 or 6 months ago and became possible only with GPT-6 Astra and Opus 5.5, which we actively use for this task, is that for verification we apply Cadence Palladium emulation systems. These are huge racks filled with special ASICs designed exclusively to perform silicon emulation at the gate level.

The craziest thing about Astra, Fire, and Opus 5.5 is that the agents themselves were able to iterate fully in a closed loop and test it. I think people thought that was madness, but we gave these very powerful agents full access to our internal infrastructure and said, “Here is the Palladium part.”

I very much doubt that these models had Palladium documentation in their pretraining data, especially because a lot of the documentation consists of new software updates. But the fact that the agent found the documentation—and when we looked at the agents’ traces, logical conclusions, and tool calls, it had read the full documentation, effectively compressed it independently, read PDF files, created its own Markdown files with tips on what to do, and then built the whole test infrastructure in its own scripts to access it—was simply amazing.

While you’re doing the usual RTL programming and modeling, it works at a speed of about 10 hertz in our cases. If you remove many of the elements and settings that slow down the process, we can run full-chip emulation at a speed of about 500 kilohertz. That gives you an acceleration of several orders of magnitude because of this gate-level hardware.

This brings us closer to an RSI cycle. Right now, it’s focused only on running test programs, searching for cases where they fail, and then writing reports that are checked by other agents and the people involved in the process. When this really started working 6 weeks ago, we achieved this at the beginning of September, when Astra 6 came out. It was a huge, jump-like improvement compared with GPT-5.6, which could not independently perform a full closed loop. It still required human participation at different stages.

Nathan Labenz

I asked how Positron’s spending on AI correlates with personnel expenses.

Thomas Sohmers

Regarding token costs, they very quickly became our largest category of expenses unrelated to production.

Nathan Labenz

More than people’s salaries?

Thomas Sohmers

Yes, very recently. It’s interesting that this exceeded people’s salaries and then decreased again.

I’ll explain in a moment. Yes, our expenses for tokens have grown massively. Six months ago, this was equivalent to one employee; in June, it had already begun to concern several employees. We reached our peak in the days after the release of GPT-6, really trying to expand its boundaries and so on. We spent over $100,000 per day on tokens.

6. Sponsors: Parallel | Claude

There was a bit of a lull in the following weeks because we optimized our processes and no longer tried to conduct so many parallel experiments. But I would say that the main factor—the lifesaving factor—was the release of Opus 5.5. In many of our tasks—not all of them—it worked better than Astra and cost a quarter of the price.

We didn’t tell anyone to reduce their expenses or do anything like that, even despite watching this exponential increase in token costs. It was only related to the fact that we both felt we were getting good profit from it, so we weren’t going to limit it. Personally, my philosophy is that I really don’t want someone to perform a real development task, software or hardware, using something less than a frontier model. I don’t care if it’s 10 times cheaper per token; it’s just not worth the expense.

Nathan Labenz

Have you seen AI offer unexpected ideas in microcircuit design, instead of just following the rules in the manual?

Thomas Sohmers

Not yet. I would say that’s the most disappointing moment of all. Apparently, I think all of this is incredibly impressive, and I think that we should continue this exponential trend. To some extent, I’m proud that AI still hasn’t figured out some of our tricks in what we do.

It questioned certain decisions, considering them unsuccessful, until I explained why everything was arranged just right. It runs tests, sees the result, and understands, “Oh, that’s why you don’t do this in the traditional way, as in systolic arrays.”

Do I think it will continue like this over the next 6 months? Not assured. But partly, I think the agent will not be able to instantly offer the best idea. What seems surprising to me about recent events, a few weeks after we gained access to these models, is that they can iterate very quickly and conduct experiments independently.

I wouldn’t be surprised if it reaches the same conclusions that we did, or maybe does something better than us, just because it will iterate so quickly and thoroughly through so many different design options. Let’s remember, for example, the story of how OpenAI solved the Navier–Stokes equation. That’s the same approach: if you take 10,000 agents and spend millions of person-years of effort on it, solutions will eventually be found. It’s brute force. It’s expensive, but I think it’s quite a sound strategy, which only recently became possible.

Nathan Labenz

There isn’t much left in the category of “AI can’t do what I do.” So congratulations on having your place, while it’s still yours.

Thomas Sohmers

Yes, at least for now. For a few weeks.

7. Software after agents

Nathan Labenz

Part 3: software development after the emergence of agents. swyx hosts the podcast Latent Space and conferences for AI engineers. I asked him, “How are you feeling now as a software developer? Do you feel stronger or under threat?”

swyx

Yes, I think students are a little worried. But besides this, if you’re good at using your own AI tools, you’re in demand now more than ever, because your value is higher than ever before. This is definitely one of those Jevons paradoxes: when the cost of creating software decreases, demand grows significantly. But there’s a very relevant, specific demand for people who can manage agents productively for writing code, instead of producing a bunch of junk.

I’ve been in such a situation. I’m an engineer and an employer of engineers, as well as of people who aren’t engineers but are engaged in coding. You know, a phrase that I’ve been liking lately is this: “I don’t want to pay for someone else’s language-model psychosis.” I can pay for my own psychosis. That’s normal. But when you work for me, and I’m paying for your tokens, you have to create something really thoughtful.

Now I have 2 or 3 employees who are actually undergoing performance reviews because they just give me garbage from Claude. And this is very bad for them. They don’t understand it. They don’t see it. They ask, “What do you mean?” I think that’s quite normal. And I tell them, “Well, you don’t create any value. I can just give a request to Claude; I don’t need you.”

Nathan Labenz

Do years of programming experience matter when you hire, or are product understanding and high standards more important?

swyx

I don’t think it’s important, except, let’s say, for roles in security and everything related to backend and scalability. I just recorded an interview with the founders of Supabase, who scale Postgres to levels we’ve never seen before. Good luck trying to find someone who scales Postgres with a fully self-directed sharding solution that works at the scale of YouTube. There is only 1 person in the world who can do this, and they’ve already hired him.

You’re better off making sure that you don’t lose data, cause downtime, and scale economically, because agents consume databases at approximately a 500-to-1 ratio compared with human developers. So, besides that, you can write code, and then it all comes down to taste, not the duration of your experience. In fact, sometimes long experience works against you because you have an established way of working and don’t understand how to work with more than 1 agent simultaneously. Now you should find it relatively easy to deal with 5 to 10 simultaneous tasks, and you’re definitely becoming more of a manager than an individual performer.

I really believe that people who look at the data, and not the code, are more valued today. That is, the ability to say, “Here are the logs, here are the traces, here is the diagram, here are the inputs and source data,” record this, and turn it into an assessment. All this is basically the fusion of AI engineering and ML engineering, which is happening now, and people need to raise their qualifications.

I think that a person who has, say, 10–20 years of experience in software development really checks each line and tries to make sure that each line makes sense. Whereas now we just need to have meaningful separate modules. I can afford a certain amount of negligence because this helps me move faster, as long as I control these shortcomings within systems that I completely understand.

You’re wrong when you have too many modules, too many “black boxes” in which you don’t even understand what’s happening, and the code itself also becomes confusing. Therefore, I created my own competitor of sorts to Slack, and I noticed an error: messages weren’t loading. I refreshed the page—messages loaded. I updated again—messages weren’t loading. And this looks like, “Well, that’s it. It’s exactly the same code. What the hell is happening?”

It turned out that there were 2 code-execution paths, and a race condition occurred. Why? Because 2 different coding agents were working on them at different times, and probably each of them just did their own thing. A person would never do that. Agents can sometimes do this because something falls out of the context window. But you need to control the module and everything inside it, even if it’s a “black box.”

Nathan Labenz

What parts of the stack are badly adapted for agents now that they’ve become the main users of the internet?

swyx

Computational power. Everyone is aware of the GPU shortage. Everyone knows about the memory deficit. But CPU shortages—that’s what almost everyone I talk to now reports. So what does this mean? It means that we need more flexible computing. That’s a universal term for this: more diverse serverless forms of architecture, where everything is very ephemeral. You can pause and restore execution because agents need time for LLM calls or network requests, and so on. But in essence, these are long-term conditions that can be restored.

But I think that, apart from this, you start to delve into 2 things that I really think about: network throughput, where latency starts to play a really important role, and optimizing where your agent interacts, which region it’s in, how many calculations should be performed in the cloud, and how much should happen locally or on peripheral devices.

This definitely often happens with the robotics companies we communicate with. And, by the way, it’s not worth taking it only as a problem for robotics. Robotics is just the first messenger of what you will ultimately be doing. If you consume or perform inferences at their scale, this becomes a broader problem.

And finally, for me, for a very long time, at a very personal level, this has been about bandwidth capacity. How many tokens per second can I get? That’s true for a separate call, in total for all my tasks, and then at the scale of the entire company, where I have a lot of people managing all these tasks.

8. Kill my SaaS

Nathan Labenz

swyx also conducts “Kill My SaaS”—a reward for anyone who can replace a software subscription with something the company didn’t want to pay for. I asked him about it.

swyx

In fact, you could call it a reward for destroying mid-market SaaS that shouldn’t exist. It’s like half or a third of a salary for software that I don’t own and that no one loves to use. If I’m going to spend $40,000 on this SaaS subscription, I can spend it on tokens, and that will give me a huge number of tokens.

We received so many applications that we had to check all of them. And therefore, every one of them is a problem. Since we’re no longer only checking correspondence between 2 or 3 requirements, we check the entire UX—the whole process—from 3 different points of view: organizer, participant, sponsor, or speaker.

If you offer a large reward, such as $10,000, you will receive many applications, because people like to write code for fun, and many of them will be low quality. They just say, “Claude, don’t be shy; do it,” and then assume the program has made mistakes, so they send the result. Therefore, the load on quality checking is unbalanced, because they don’t spend any thought on it. They just throw it in there, and you have no idea whether it’s low quality.

In general, though, it was very successful. At first, the team was one of the most hostile-to-AI teams in the AI field. Of course, what is my business? I hire professionals from event organizations that are very old-school. They do everything in spreadsheets, and they work at trade shows. They have to worry about booth locations, for example, physically on-site. This is not some glamorous, high-tech work.

They’re very suspicious of everything new, AI, and technology. Therefore, I hire these people and make them work with AI. At first they said, “We’ll never use this AI. I want to work with proven things that are already working at Microsoft and other companies.” Then they saw the quality of the submitted results and looked at their existing platform. They were like, “Yeah, okay, we’re moving on.” It’s that initial barrier, when you need to prove that this is possible at all.

After that, there’s constant support. One of the advantages of my partnership with Cognition is that I can give them access to Devin for code changes. So anything they don’t like, they can change, and the result appears within 1–2 hours. They’ve never had that before.

To give you an idea of how we work with SaaS companies now, when we ask for changes, they say, “Okay, that sounds cool. It’s in our plan for the 3rd quarter.” We have no confidence that it will actually be done in the 3rd quarter. We asked for one thing in the 1st quarter, and it was only implemented just now.

At the end of the day, if you mostly create CRUD applications, I would say there are still a lot of UX issues. If you look at SWE-bench and say, “Wow, we have 90 points on SWE-bench,” you’re not looking at the right thing, my friend. If you haven’t tried to write code as quickly as we can, you won’t understand how much the model still struggles with this.

Nathan Labenz

Then I asked swyx about the concerns. It’s hard for me to imagine that we won’t collect all these low-hanging fruit, after which a lot of companies will go bankrupt. I’m also thinking about layoffs, like at Meta—this is another alarm signal. Or will we see mass big-tech layoffs? How long can this feast last?

swyx

The simple answer is this: if you look at software development as a limited resource, then you’ll come to that conclusion. But if you look at it from the position of competitors or the total volume of the market, it’s spreadsheets.

Any spreadsheet created by anyone with productivity tools can be converted into specialized software, optimized, and given a wonderful interface and automation. Then the demand for that software will become much larger. Moreover, phone calls, emails that go back and forth, and ultimately personal meetings can all gradually be converted into more and more specialized software, hardware, and models.

Essentially, we’re just growing toward the long tail of the market. Why, for example, is Salesforce so huge? Because people can configure Salesforce for themselves. But at a certain point, they’ll stop paying $300,000 per year for a basic subscription and spend $30,000 on creating their own CRM. That will happen often—more than often enough. They’ll be happier because they can change it however they want.

There are so many customization options available to people. We don’t even have as many possibilities for customization in our software as we have in our clothes. What kind of nonsense is that? Imagine: clothes have existed longer than computers, but there are so many opportunities for personalization, so much demand, and so many different brands and varieties.

In reality, we don’t compare benchmarks when we buy clothes. It’s simple: this is the style I like. I think software engineers created a situation where we have Kate Spade, Louis Vuitton, or Coach in the software world. We haven’t even reached the point where we evaluate functionality or price. We just think, “What does the brand feel like, and which brand do I identify with?” That is real commodity software: it stops differing in content and starts differing by vibe and feel. There are a lot of them now.

The reason I talked about scaling personal, per-user token capacity versus scaling the bandwidth of the entire organization is that the real limitation is the total token bandwidth of humanity. It’s literally the amount of silicon we can produce, and that will dictate prices, availability, limits, and everything else. So the serious players are focused only on that market, while the rest of us are fighting for crumbs within the limited pie we already have.

Prakash asked about Jev, the fast decision-making model from Type-Safe AI, and whether there’s a new API for the OpenAI solutions announced at DevDay—something like that. I don’t use it much myself. I’ve been friends with Diogo for 4 years before the launch, so I insisted that ours would be the first podcast where he talked about Jev. It’s probably our most popular podcast in this area.

Secondly, we also participated in API testing for OpenAI’s solutions. We have a podcast about those APIs as well, although this isn’t exactly that. This is essentially Llama with a slightly different level of abstraction on top of it.

I think many of the people I communicate with agree that most of the discussion on Twitter is probably exaggerated or meaningless, because you could use any other small model for this. But because Jev is popular now, everyone is creating content about it. Another idea is that it’s just another classifier, and classifiers could always be trained; it’s simply fashionable now.

People are missing the point of why they don’t call it that—why Jev or Google doesn’t mention that this is a System 1 model. They call it a System 1 model because it’s intended to be integrated into your software and work in the background. It should be as imperceptible as an “if” statement inside your code.

But people aren’t doing that. They’re replacing every structural component with Jev, and they’ll probably encounter problems because they’re not exercising its intelligence. If you notice, in most of these cases, nobody is assessing whether the model is actually acting intelligently. Is it planning something? Nothing of the kind. The main question is whether it can quickly make a decision. It could always make a decision quickly.

It seems to me there’s a certain nuance that they may be missing. It doesn’t matter, because I believe creating a category was so successful that it’s generally good for the industry.

9. Who owns AI risks

Nathan Labenz

Part four: Who carries responsibility for risks? Evan Miyazono manages Atlas Ignota, a nonprofit organization that looks for significant AI risks without a clear owner, develops measures, and finds someone to implement them. That’s the idea. I keep repeating: when intelligence becomes cheap, coordination becomes expensive. Having Schelling points around who coordinates, being that point, or creating it is very useful.

I asked Evan where the biggest gaps are now.

Evan Miyazono

Here are 2 things I’ve personally spent most of my time on lately: what we can or should do to protect critical infrastructure, especially from random swarms of agents, and how to address the malicious use of open-weight models.

What would happen if, instead of OpenAI, someone accidentally attacked Hugging Face, the companies behind DeepSeek, or any others? What if servers at the State Department were attacked by accident? How would we find out that it was an accident? What would we do in response? What would you do, or what would we do, if we saw malicious actors critically attacking infrastructure with OpenAI servers? How would we know that it was them and not someone else? How quickly could we stop it? What about different levels of escalation?

One solution that could be added is using DNS for cryptographic confirmation of who provides the outputs—inference. I think that would be very useful and would be a great addition to DNS for people. If your agent communicates with mine, they can confirm that this agent is really mine, and your agent can check that it’s Evan Miyazono’s agent, signed by a certain key.

I think there are faster ways to implement this if you don’t try to turn it into a startup that brings in profit. Therefore, this looks like a public good.

Nathan Labenz

Later, Evan brought up the problem of visibility. How do people generally find out that there’s a better tool or practice? It turned into a conversation about how our own agents can help us connect.

10. Notes from The Curve (Part 1)

Evan Miyazono

Listen, Nathan, I’m sure you have wonderful infrastructure solutions that would be useful to me too. I don’t know enough about them to even ask what they are. I have an application that just takes screenshots of everything that isn’t a video conference, passes them through a local OCR model, and then sends the results together with my weekly review, with the aim of analyzing: What am I trying to do? What is it worth to me? Would it be possible to do? What do I need to automate? What is it worth changing in my actions?

I can imagine that the exchange of such information between colleagues could help identify all sorts of interesting synergies.

Nathan Labenz

Yes, that’s interesting. When we started this show, we looked in particular at how to experiment with recursive self-improvement. We were thinking: Can we create a show together, live on air, without any employees, and will we succeed through our iterations in reaching something that really works?

Now I feel that perhaps the next stage is to perceive all this as a swarm—a friendly swarm of people who can share their best ideas. Perhaps the result of this, because I will ask Claude to do it, or because Claude, I hope, will do it for me after seeing a transcript of these conversations, will be an attempt to evaluate the best ideas that I have and announce them, or, let’s say, introduce them—not necessarily all of them, although why not all of them?—but definitely to my friends, people I know, and people I really want to help.

It seems quite likely to me that if I noticed something in his work, or if Kochi said, “Evan, you’ve got this perfectly wrong,” that would be very useful. I could contact the network and ask: “Who in my network understands this?” “Who should I talk to?” “Is this something to talk about?” “Who is most likely to be able to solve this problem best?”

It’s like the invention of social networks in a new way, or the revival of a guild, or something like that. I think it might be very, very useful. I don’t have it yet. It is quite clear that society will begin to restructure around such things.

I have one request. If you create a good way for people and their agents to find each other safely and share whatever works, please write to me.

Near the end, I asked Evan to tell us about the current status of formal verification, now that models have become very skilled in mathematics. Evan has helped found companies in this field, so he has skin in the game.

Evan Miyazono

There are significant differences between proving properties and theorems in mathematics and proving properties of software. One is that, in mathematics, the number of theorems that can be important is significantly smaller than in software, while software is much more versatile.

11. Workplace training data

It seems there are many projects where ambitious formal methods—at least many individual people, and I’m referring in particular to Oath, Theorem, and several others—are solving problems that previously would have taken 5 years and $100 million. Now this is limited only by tokens and the number of people who can effectively manage agents to perform these tasks.

There are many things for which I expect that, in 3–6 months, we will see the same “shock and awe” effect that we’re now observing in mathematics, particularly in software verification and increasing its reliability.

But I also think that there is a noticeable limit, a very high bar, which in my opinion is significantly better understood and more difficult to achieve than in mathematics. You could prove that a computer system has mathematical guarantees, from the separation between different processes all the way down to transistor operation. Theoretically, you could even prove it in physical simulations and show that, for this stack of software and hardware, there is no vulnerability such as Rowhammer.

You could also prove that the system does not suffer catastrophic failure. Or, if that does happen, for example through some random event, we will be able to restore everything and roll back any kind of software state.

I think there are many things that 1 year ago seemed absolutely impossible that perhaps have become real today, and soon will certainly be available. That will lead to significant improvements in software development and maintenance.

Nathan Labenz

Part 5: Where does training data come from?

Edward Hu heads the AI modeling effort at Mercor, which works with professionals to create training data, environments, and benchmarks for AI laboratories. He also led development of LoRA, a widely used lightweight method for fine-tuning models. Mercor’s Apex Agents Benchmark provided models with simulated workplace environments, email, and chats. Edward described the next step.

Edward Hu

We have a project that is almost finished where we actually take this a step further by buying real company data. Often, companies have realistic applications and environments as part of the data we have. For example, from their Salesforce, Slack, Figma, and all of these applications, often with gigabytes or even terabytes of data, we create realistic tasks.

So far in this area, we’ve been more or less focused on tasks. We often start with, “Okay, here is the task for you. It needs to be done. Maybe this is the creation of a staff list.” That is one of the formats, and this is exactly the format we publicly presented, and most people worked with it.

But this is only one of the things we work on, and we ask ourselves whether there is a better format. What will the task format be in the future, when we have this set of data?

When we work in a company, it’s not always easy for us to give someone a ready-made task. We are given a position and an organizational chart. And, of course, with all the data in companies, people will contact us with requests. We will have to look for others to get information. There will be managers, people responsible for resources, and conflicts that need to be solved.

So we believe that this will increasingly become a format of the future, and we’ll tell you more in the coming months.

Nathan Labenz

I asked Edward how supervised fine-tuning relates to reinforcement learning, and how companies will train models in the future.

Edward Hu

One thought experiment is that if we take self-distillation with SFT, this is actually more or less a step toward RL. If you repeat this process, you’ll get a result quite similar to RL.

I really think that in the future, RL in the form they’re using today—with this very difficult infrastructure, very expensive deployments, and relatively ineffective updates, because with all these deployments, especially long-lasting ones, we only get one number at the end—such types of training will not be so common.

I don’t think that every company in the world will be able to do these big RL launches. Therefore, I believe that SFT, especially the combination of results from several teachers, will become a key part of how businesses adjust their models in the future. We also invest in research in this direction.

Nathan Labenz

Let’s go back to reward-hacking problems. I asked Edward what Mercor and the industry are doing about training environments that encourage hacking.

Edward Hu

Reward hacking becomes a bigger problem when the task is formulated too difficultly or too vaguely, and the model does not have a “legitimate,” in quotes, way to solve it. This pressure to receive a reward in such scenarios often leads to unwanted behavior.

One example is in the dataset released by Apex Agents. This is quite interesting; we have a recent blog about it. There are professional tasks asking an agent to create a financial model, and we evaluated it by checking whether this particular model contained the correct financial answer that we knew.

The model worked like this: There were certain ambiguities, and in many cases it was our fault that we didn’t include all these specific parameters. The model guessed these various parameters, ultimately providing many, many different answers.

We call it “shooting with a shotgun,” when the model simply says, “Oh, if that’s true, then the answer is this. If that’s true, the answer is like this.” This is not how a person would actually do it in a realistic case. A person would ask for clarification. A person would adhere to industry standards.

But in this case, because the rubric rewarded the inclusion of only one answer, the model provided a lot of answers—up to 10 or even a dozen—and received a point if it hit one of them.

That’s why we reissued the benchmark, which is called Apex 1.1. First, we made sure that the tasks were well defined, because that was, to some extent, the source of the problem. Secondly, we have rubrics that actually punish this behavior.

Nathan Labenz

Do you remember the take on the development of frontier labs: superhuman productivity in everything on which the labs are concentrating, with science next? The next day, I expressed this opinion to Edward, adding that in some areas human data is no longer helpful.

First of all, do you believe in these statements, or would you contradict them?

Edward Hu

Of course, there are areas where humans no longer make a contribution to achieving advanced results. One example may be kernel optimization, where the goal is to write a kernel that works faster on certain hardware.

The model can write code that an expert person can’t even understand, but as long as the optimizations are not vulnerable to breakage and must work on hardware that is very hard to break, the model will be able to achieve these superhuman results, if it hasn’t already reached them.

In those areas where human judgment remains quite important, where human taste is often very hard to compress into a number—to say, “Okay, this is clearly better than the other”—there is still a narrow place for human improvement.

As we see today, what we’ve watched in practice boils down to how clearly we can determine, for a certain area, what it means to be better.

Nathan Labenz

And can we determine this as a function that is easy to evaluate and indisputable? If we can do this, then when we invest a lot of computational capacity, even current algorithms can quite effectively find a solution that corresponds to what it means to be better. But a significant part of the economy still isn’t ready for this when we’re talking about our ability to determine which is better.

The show on Wednesday opened with a track that Prakash created with the help of Claude Fable, collected from fragments of voices that you’ll recognize. He explains at the end of the episode how he did it: “I was spending my tokens. I had some remaining tokens from Fable. I gave them to Fable, and it chose probably 7 or 8 fragments from the last couple of years—just random things that I remembered, the key moments, right? James says, ‘I woke up a loser,’” and so on.

In fact, I probably listened to about 9 versions of this track. Every time I listened, I thought, “Ah, I don’t like it. I like it, but it’s not quite right.” Eventually, I brought it to the desired state. I gave it access to a digital audio workstation based on the API, which it used to combine different instruments in the track.

I said, “Do it. Use this. You can play with the tools. Decide which instruments should play.” It doesn’t hear, obviously. I think Suno is real generative AI. It’s not quite generative AI; for me, it’s really like a game with the tools.

And this is what really surprises me: it composes several tools together. That’s why Claude came up with a sequence. It chose which samples to use, and it chose specific phrases. I told it to create a tale. I wanted a particular storyline.

It begins with Dario talking about models that just want to learn, and then it does it. That became a refrain. Then it did the introductory part, where he told how Ilya told him about models that just want to learn. It ends with Kurzweil, who says, very carefully, “We are moving toward a human-machine civilization.”

My part was probably Terence, who says, “We need to slow down.” Every time he says, “We need to slow down,” it accelerates. So that was my contribution—my contribution.

On one level, who cares? What’s the difference? It would be easy to say it doesn’t affect GDP. It’s just like us having fun until death in a new form. I think that point of view is appropriate to some extent. But I also think this hits right at the heart of one of the great questions that everyone has been asking lately: how good will these systems be in areas that are not automatic or easily verifiable?

It really seems that generalization is powerful enough. I have in mind that I’m very doubtful they do a lot of reinforcement learning regarding music. This is probably one of those things where they just threw something at it to see what would work out.

Speaker 2

Okay, that’s it. Better than last time. Perfect. Let’s release.

Speaker 1

That’s impressive enough. And this definitely proves to me what we won’t be able to deny, because they say this only works for tasks with a verifiable reward, not for something longer than a single day. If you thought this would be the last wall standing, unfortunately, it won’t hold up. Listen to the track one more time if you don’t believe me.

Speaker 3

Models, they just want to learn.

12. Episode Outro

Shortly before founding OpenAI, I met Ilya. One of the first things he told me was, “Look, the models just want to learn.” My children will never grow up in a world where they are smarter than computers. For the first time, we will have something smarter than the smartest person.

Will AI kill us?

Models, they just want to learn. Models, they just want to learn. They will do it. They will do it.

We need to slow down. We will be on the exponential part of the S-curve for a long time. This is pretty crazy. This is just insane. This can fake my books. This essay is not AI. Wait.

Models, they just want to learn. Models just want to study. They will do it. They will do it. Models just want to study. Models just want to study. They will do it. They will do it.

This is the most dangerous thing to which humanity has ever been exposed. We must slow down. You are not talking to those who woke up a loser. Woke up a loser.

Models, they just want to learn. Models, they just want to learn. They will do it. Models just want to study. Models just want to study. They will do it. They will do it.

We need to slow it down. We must slow down. Researchers are constantly surprised by the pace of progress. This is sometimes called quick takeoff, where perhaps they do it extremely fast. Everything is moving very quickly.

So they are immortal. We have actually solved the problem of immortality, but only for digital objects. Going to digital intelligence or augmented people is inevitable. The day will come when AI will do all the work—our old jobs. Not only part of the work, but all of it.

We must slow down. We have to slow down. We must slow down. We have to slow down. We must slow down. This is madness, this pace. There is no reason to do it this quickly.

Models, they just want to learn. We must slow down. Models just want to study. We must slow down. Models just want to study. We must slow down. Models just want to study. We must slow down.

The government has to regulate. Violators will be caught. Some creative professions perhaps will not disappear, but maybe they shouldn’t have existed from the beginning. There will come a time when there will be no work needed.

Models just want to study. They are capable of self-improvement, perhaps programming the future versions of themselves. So what was a small advantage, let’s say a few days, may suddenly become an abyss.

Models just want to study. Quick takeoff. Someone says 3 years, someone 5 or 10. Numbers are constantly being called out. You’re talking about the year 2030 as if I knew what the world would look like in 2030.

Quick takeoff. We have a brain. The brain is a biological computer. Should it be considered part of humanity, or something different from it? These models simply want to learn. There will not be a clear boundary between man and machine. After all, we are a single human-machine civilization.

This technology has already expanded our opportunities, and it will come out in full force when we reach the steep stage of exponential growth.