Claude Code 迷狂症:SemiAnalysis 如何在 Token 消耗上碾压 Meta | 第008期
- SemiAnalysis 正在把 Claude Code 变成一套智能体研究栈,其相较人工分析师工作的经济性颇具吸引力。 Dan 的 Wags 系统启动了 AOI 覆盖,约5分钟后返回一份部分乱码的合同相关结果;这次运行成本为3.52美元,若算上前期搭建可能达到10–15美元,相比“让分析师在这件事上耗上一天”仍然划算。目标是覆盖60–100家公司,让人类把时间放在“为什么发生、意味着什么”上。
- 真正的约束在上下文架构和持久记忆,而不只是模型智能。 专门智能体分别处理文字稿、新闻、事件和财务模型;由于它们只能通过主智能体沟通,每个智能体只负责流程中的一个离散环节,再由一个“洁净”的公司智能体吸收其输出。子智能体关闭后,上下文随之消失,但记忆文件也不可能无限膨胀。最棘手的未解故障很直观:生成的资产负债表无法平衡,必须加入监督检查,并安排一名实习生负责“鞭策智能体”。
- 低任务成本与使用量复利支撑强劲的 Token 和 GPU 需求,但企业采用仍是更大的不确定性。 SemiAnalysis 称其每位员工消耗的 Token 超过 Meta 的2倍;即使演示任务成本达到15–20美元,Dan 仍会选择运行它。IT 审批、激励不足以及与客户相隔甚远,都可能让50%的生产率提升变成提前50%收工,而不是多产出50%,因此 Token 支出未必能直接映射为企业盈利。
- 模型质量已经从需要谨慎使用的专家工具,跨过了用户可以用“给实习生的同样指令”来驱动的门槛。 Sam 说,1年前的软件工作还要求测试驱动开发、分步搭脚手架和严格的上下文管理;现在他会把产品介绍给父母或祖父母,并预计最终“每个金融圈的人、每个律师”都会依赖它。反方观点是,尽管输出已经足够好,采用率仍然疲弱。
- Mythos 可能代表网络安全能力的一次跃迁,但 Dan 并不相信原始能力本身就足以解释这一点。 他指出,编码和网络安全领域的进步远超其他方向,并质疑围绕零日漏洞的“巨型营销活动”。在他看来,持续性、后训练、harness 以及奖励信号可能才是决定性技术栈。Dan 还说,很多其他模型在漏洞被解释清楚后可能同样能够完成利用,但通过长期任务自行发现漏洞是另一回事。
- 模型竞赛可能正在收敛,但真正的护城河正转向界面、工作流沉淀和用户惯性。 Meta 看起来仓促发布的 Avocado,被描述为大致与“Opus、5.4 和某个 Gemini”处在同一档,由此引出“四马竞逐”的问题。把 Claude Code 重新指向另一模型在技术上可能只需改一个 URL,但用户会积累知识文件、调整工作流、容忍可修复错误,并看重 Claude Code 控制电脑的能力,而不只是浏览器里的聊天功能。
1. 如今是实习生指挥研究蜂群
Dan 介绍了 Terrence——一名在 NTU 主修 AI 与会计的学生,同时操作5个终端和3块屏幕。这个玩笑背后是组织方式的变化:“现在是实习生当老板了”,由他教智能体做财务建模,而不是亲自包办每一项任务。
Wags 是团队的“智能体研究总监”,负责监督处理财报电话会文字稿、新闻、事件、公司简报和财务模型的独立智能体。由于这些智能体只能经由主智能体通信,每个智能体只处理流程中的一个离散环节。这种分工让 SemiAnalysis 得以覆盖网络、TCO、数家交换机相关公司以及数十家光模块供应商,不必让单个上下文窗口吞下所有资料和操作指令。
这套架构让公司智能体基本保持“洁净”。其他智能体负责搜集和总结材料,再保存一份涵盖模型、文字稿和公司简报的索引;公司智能体被唤醒后吸收这些输出,只保留关键内容,而不把“得出结论的过程”留在自己的上下文里。
一次实时 AOI 查询返回了一份部分乱码的结果,提到“3.24亿”“1.2”,很可能涉及 Meta 和 Amazon,内容大多是前瞻性判断;Dan 说自己“得拿实际数据核对”。这次运行成本为3.52美元,若加上此前的启动工作可能达到10–15美元,却足以支撑他的经济性判断:即使成本达到15–20美元,他仍会选择它,而不是让分析师干上一天。
2. 离散技能正在复利成一套研究操作系统
3个月前,Dan 还不是这么工作的。春节期间工作放缓,给了他学习工具的时间,这些工具原本是为了减轻团队工作量;“总结一场活动、搭建财务模型、监控新闻稿”等零散需求,逐渐变成覆盖财报、新闻、事件、建模和监控的可复用技能,最终成为 Wags 的组件。
逻辑并不复杂:Dan 说,金融数据至今仍没有可靠的导入方式,而核验和修正导入数据,耗时可能超过直接手动录入。他的目标是在不制造过多包袱、让分析师失去行业视野的前提下,覆盖“60、80、100家公司”。人类应当少做数据录入和事件复述,多做“审核、编辑、思考、建立联系”。
会议智能体 Claudia 处理的是没人有时间亲自消化的材料。SemiAnalysis 每年参加50–60场会议;仅 GTC 就有621场 session 和830场演示,但团队因为会面优先,GTC 一场 session 也没听。Claudia 会把演讲和 YouTube 链接转成文字稿,整理会议演讲者的可聊主题索引,还能找到关于 CPO 的演讲,尤其是使用 DWM 做 CPO 的内容。她也能为 Julian 整理一份 ScalaCross 简报,再由 Julian 向 Dan 汇报。
3. 信任提升速度快于可靠性和记忆能力
主持人把态度转变定位在大约11月:SemiAnalysis 此前推动大家使用这些工具时,用户不信任任何输出;现在则倾向于默认相信,只有在风险足够高时才进行核验。他日益感到恼火的是:“现在我才是瓶颈”——模型想写代码、做分析、搭建系统,但真正限制它的是指令和任务设计。
Dan 给出的反例是一张生成后无法平衡的资产负债表。他不得不调查智能体是否把 FactSet MCP 数据与 SEC 文件混在了一起,团队也要求每次完成的模型迭代都经过监督检查。主持人的反驳很精准:这“不是分析师会犯的错”,因此如果把智能体完全当作初级人类员工对待,就会掩盖它们独有的失败模式。
记忆是更棘手的结构性问题:“子智能体一旦关停,就结束了。”经验必须被写入记忆,但记忆文件不可能无限增长。Sam 在研究 neoclouds 时遇到的是反向问题:他的知识都留在 Claude 的对话里,于是他做了一个 skill,把这些经验搬进 flashcard app,以便刻意回忆和检索。
4. Token 需求可能先于企业生产率兑现到盈利
SemiAnalysis 称其每位员工消耗的 Token 超过 Meta 的2倍,团队把这种对比称为“token mogging”。Dan 将演示任务的经济性与 GPU 租赁价格拐点联系起来:即使总成本达到10–20美元,他仍会选择它,而不是让分析师工作一天。
更广泛的采用仍处于早期。Dan 看到基金经理用模型做窄范围的摘要,但许多 Fortune 500 员工既没有 IT 审批,也没有能够随时间积累回报的工作流。如果 Meta、Amazon、Unilever、P&G 和 Boeing 都像 SemiAnalysis 一样深度使用 Claude,他认为,“唯一能缓解这一点的方式,说实话,就是大量芯片工厂。”
在 SemiAnalysis 内部,效率提升让“积压的想法大坝”决堤。据称有一名员工连续运行 Claude Code 24小时;其他想法在1小时内就变成可分享的 Slack 链接,而智能体编码追踪基准测试以及 Accelerator 或 ClusterMAX 仪表盘,都在2–3小时内形成了可运行版本。
主持人的保留意见值得记住:SemiAnalysis 能把更好的研究相对直接地转化为订阅收入,但一个距离 P&G 客户六层的人,可能会把50%的效率提升用来提前50%收工。Dylan 随后说,模型质量的提升反而让他对采用率感到悲观:“瓶颈其实不在输出质量”,而可能在于冷漠或激励错配。
5. 前沿能力取决于持续执行,工作流既有优势则成为护城河
Dan 发现 Mythos 的改进异常集中在编码和网络安全;其更广泛的能力提升“显著更少”,而系统卡据称表示这一结果并非有意设计。他质疑网络安全叙事是否部分只是营销:“如果我想让人们不要恶意使用某个模型”,就不会发起一场强调零日漏洞发现能力的宣传活动。一名 Stanford 网络安全研究人员曾警告他长期运行的智能体攻击,但这名研究人员及其朋友仍怀疑 Mythos 是否真正带来了模型质量的跨越式提升。
GPU kernel 编写和推理系统架构是另外两个平行前沿。AI 现在能产出大量领先的 kernel,这在6个月前还做不到;Dan 形容人类向 AI 的能力迁移总体上是一条“单行道”。不过,有经验的 kernel 作者即使使用 AI 仍然占据优势,单靠 Token 预算尚未让这项任务彻底商品化。
长期任务中的持续性仍然不足:模型可能停止查阅文档和源代码,建议向维护者提交 PR,却无法解决已经记录的 bug。Dan 的结论是,两层能力都不可或缺:模型后训练必须奖励持续编码行为,而设计良好的 harness 则要提供维持这种持续性的工具和结构。
6. 模型竞赛可能正在收敛,但工作流会制造惯性
Meta 的 Avocado 发布看起来相当仓促,甚至出现了错误图表,但其水平已接近“Opus、5.4 和某个 Gemini”,意味着最难的追赶阶段可能已经结束。Dylan 认为它还不如最领先的两三个模型,但大致处在同一档;Dan 则说,Meta 看起来已经完成了最难的部分,剩下需要追赶的空间相对有限。
切换 Claude Code 的配置在技术上可能非常简单——Dylan 形容只需改一个 URL——但积累的知识文件、预期行为以及对可修复错误的容忍,会共同形成用户惯性。Dan 的问题仍然没有答案:差异化是否会从原始能力迁移到“生态系统”“在位优势”和用户体验?
硬件故事同时指向两个方向。Dan 说,在 Claude Code 热潮期间,Mac mini 在美国的交货周期约为3个月,在亚洲约为1个月。Dylan 则反驳说,Mac mini 并非必需,因为大多数模型算力都在云端运行,他甚至考虑改用 Chromebook。
I just have a lot of things I prepared to show and tell and that I’m going to share. This is our intern.
[laughter] Oh my goodness. I’m like, “Hey, you know, I know you guys say we don’t need interns, but the script is flipped. The intern’s the boss now.”
Yeah, man. How the tables have turned.
Hello, everyone. Welcome back to SemiAnalysis, episode number 8. We’re going to talk about all of the ways in which Claude Code psychosis can be used for profit today.
With me is Mr. Daniel Nishball again, calling in from Singapore. We’re going to run through exactly how many agents the team has built at our biggest office. And a new member of the team, Sam Harsh, calling in from San Francisco, just south of me. Guys, welcome to the show.
Thank you.
Excited to be here.
Awesome. Cool. So, look, it’s been a big week for Claude Code. We’ve been using it a lot. We put out some bait. I tweeted today about how we’ve been spending on Claude Code, how we’re consuming more than twice as many tokens per employee as Meta. We clearly have much more psychosis than those guys do.
Token mogging, I believe, was the technical term. And then we also saw a nice release from Mythos from Anthropic, which we can run through in a little bit of detail. Dan, what’s new, man? You were just telling me you got the agents ready to go.
Yeah. We’ve been stockpiling equipment at a scary pace, mostly because Ray’s been scaring us into buying everything—laptops, iPhones, and Mac minis, right? We’ve got 2 of them in the office, which we’ve been playing around with. We’ll probably move them to proper service at some point.
But it’s all about how we’re using Claude Code. A lot of people say Claude Code means you don’t need interns anymore. The script is kind of flipped: now the intern is the boss of the agent swarm.
This is Terrence. He’s at NTU. There are 3 universities in Singapore, and NTU has a great AI program. He’s an AI and accounting major—dual majors. His job is teaching the agents how to do actual financial modeling.
I’ll introduce you to a couple of the agents we’ve been using.
While you’re pulling this up, that was just a picture—we got to see the nice back of Terrence’s head, and he had 5 different terminals up across 3 different screens, I think. So now we get a live demo.
Yeah, this is the kind of stuff Terrence does all day. He’s the boss. This is what we call Wags, our agentic director of research.
The problem statement is that we have to cover a really broad industry across networking and TCO: four to five switching relay companies and dozens of optical-transceiver supply-chain companies. We can’t do a full model on each of these companies every quarter.
On the left, you’ll see Wags, the director of research, doing a supervisory function. Right now, we’re initiating coverage in AOI. If we talk about agents before we step back and talk about why we’re doing this, you’re mainly limited by context.
You can see in the bars that things are already filling up. By the time you load in all the context you need—how to grab this information, how to scrape things, and how to do all of that—you’ve already overwhelmed the context.
This is “read all the transcripts, process them, generate summaries.” This is “grab all the news digest.” Then there’s the event engine, the brief agent’s metadata, and the model agent, which isn’t here—it’s actually wrapping up now. The model agent is responsible for building the model.
The problem is that all of these fill their context up with the task at hand. What we do is structure all of this. It’s actually finished, so I can show you a little bit of the process.
I said, “Okay, let’s initiate. Here are the steps we’re going to take: conference summaries, key takeaways from the coverage, and the final scorecard deliverable.” I’ll ask it, “Hey, Wags, can you summarize all of the agents you used and what they did?”
The key thing is that the one agent we keep pretty much pristine is the company agent. What the company agent does is read in all the results and all the summaries. Then you can ask it questions.
It’s important to have that in the context, but not necessarily how we got to it. You can see what they did.
I’ll take a step back and say, why did we build this? Other than covering more ground, there’s one really silly problem statement. In my past life, before this, I was in public-markets investing. I built so many of these models and covered a lot of companies.
It seems silly in this day and age that there’s nothing to reliably ingest financial data. By the time you ingest financial data, you spend more time checking it and fixing it. You might as well type it in at the beginning.
It’s only now that this has allowed us to do this. The point is, we’re going to have this initiate coverage, and we’re going to be able to ask it questions.
We want the analysts to spend more time reviewing, editing, thinking, and drawing connections, and less time on data entry and telling us what happened—but more on why it happened and what it means. Our analyst team is fixing my view. It just means we cover more ground and cover it smarter. Make sense?
What are your impressions watching this happen? Have you been using agents at this level?
Not at this level. It’s a bit bewildering to me, certainly. I was burning some tokens yesterday. It felt good to break that in under the SemiAnalysis banner for the first time, but nothing like this.
It strikes me that if you’re using that much context on each agent, maybe it’d be better to be more aggressive about spinning up subagents. How do you think about structuring this whole thing and passing information from node to node to make it go as quickly and reliably as possible?
Yeah, it’s a good question. One problem we had was moving data between the agents. They can only talk to the lead agent, so we have each of them doing a very discrete part of the process.
You can see I’ve woken up a company agent. When it wakes up—when we finish the initiation process—it saves an index file of what we create. There’s a model, the earnings transcripts, and the company briefs.
When you wake up the company agent, it’s told to ingest all this stuff. It’s ingesting all this context. I’ll ask it, “What was the large optical contract that AOI won, and what impact did it have on their financials?”
Quarter has a good product, but it generally only scans earnings transcripts. It won’t necessarily have all this other stuff in its context. I’m not even sure this is going to work because I just initiated it, but broadly, the question is: how do you keep the right amount of context and the right amount of information in the right agent?
We have an agent for earnings, an agent for events, and an agent for just updating financials. Does that answer your question?
Yeah, that gives some good insight. I’ll need to go to school on this over the next few days, but it’s good to see the expert at work.
Actually, look at this. Imagine how long it would take my analyst. I just ran this, probably starting 5 minutes before the podcast. Think of how long it would take your analyst to initiate on this and come up with this answer.
AOI landed 324 million, million 1.2, likely Meta, Amazon, mostly forward-looking. Obviously, I’ll have to check this against the actual thing. We’re still developing it.
Let’s see how much this cost. This cost me $3.52. You can see why, like, what we've written about is we've written about, probably I think there was some initiation work I did in other sites. So, maybe the total cost is like 10 or 15 dollars.
But it’s either this or I have an analyst spend a day on it. You can imagine that if it was $10, $12, $15, or $20, I’d still do it. That’s why I think the end demand and the end ROI are so strong.
And that's why you're seeing the inflection in GPU rental prices, which we talked about last time.
Dan, can you talk about how quickly you learned this? I've been using coding agents for a while, and I feel like I'm almost more cautious than you, slowly dipping my toe into agents or OpenClaw sort of stuff. How did you get into it? You weren't using this even 3 months ago, right?
3 months ago, it was sort of like we had a Chinese New Year break. We had wound down a lot of the projects we were working on, so we had a little bit of headspace to get into this. The ironic thing is that we were so busy, we didn't have time to learn new stuff that would make us less busy.
With that break, we started playing around a bit. What really helped was the concept of how you want to build a research team: what the tasks are and how you would divide them. It also started from a collection of very discrete things we wanted to do.
Let me switch windows real quick. What we realized is that there are a lot of discrete tasks. There are 2 important things to realize: there are a lot of discrete tasks, but to build that into something that's an ongoing, living agent with living coverage—not just, "Let's build this project, let's build that project"—you need to build something that compounds itself and actually uses all these functions.
I'll fire up Claude. Here, I'll show you the skills we built. All of these skills are used by Wags: event summary, earnings, financial model build, press news, and pure monitor. All of these were started as very discrete things: "Can you summarize an event? Can you build me a financial model? Can you do a press news monitor?" They all stemmed from things we were doing that were really manual, things we couldn't have even imagined doing now.
For example, "I want to know what all the press announcements from Broadcom are ahead of OFC."
Yeah, makes sense. You talked about the job of an intern being different. Has your job changed meaningfully? Do you feel like you have to spearhead learning this stuff, even though in some ways maybe younger people who are going through school right now might be more AI-native and better prepared to learn how to do some of this stuff?
I would say for me, a lot of this is really trying to figure out what the end goal is. I think all the analysts get it because they've all done these kinds of tasks. For me, it's really trying to figure out what we want to get to and what we need to be as a team.
Where we need to be is a place where we can cover 60, 80, or 100 companies without creating so much baggage that we lose sight of what's really going on. So, making sure that Wags and everything else does things in such a streamlined way that it really enables that cadence.
I'll give you one more example of another agent family within Wags. What do we call—
Oh, there are families now?
Yeah, Claudia is our conferences chief. As a company, we attend 50 to 60 conferences. The thing is, when we attend, most of the time we actually spend it in meetings these days, in targeted demos.
What do you cover now?
I know, George, you were at GTC, probably in back-to-back meetings with all the neoclouds.
I went to about 3 sessions the whole time. I was just there having meetings with people in person because they happened to be there, too. Every time I go into a conference, I earmark all these amazing sessions that I want to go to, and then I go to exactly 0 of them. You can see there are 621 sessions at GTC and 830 presentations, of which I went to exactly 0.
To be clear, I was at GTC while Dan was at OFC. We talked about that a couple of episodes ago. I actually got the debrief kind of live on that podcast because I tried to read some notes, but I hadn't talked to you about it since you'd gone.
Yeah, and so you're not always able to view every presentation. It becomes an obstruction, having to stop whatever you're doing and say, "Okay, now I've got to go to the website and dig through that." You might get interrupted.
Once we have the file URLs, what Claudia helps us do is take presentations and YouTube links, transcribe them, and index the talkability index of each conference for speakers. Then I can say, "If I'm researching CPO and DWM for CPO, find me every talk that discussed CPO, particularly the use of DWM for CPO." And there it is. If I want to do a research project on CPO, that's the starting point.
I'll show you another output that we did. I said, "Tell me what to do for ScalaCross." Here it is. Julian's leading our work on ScalaCross. He just joined us last week—or actually, this week is his first week—so now I've got a bunch of stuff for Julian to read. I'm saying, "Go read this and tell me what it says."
So the agent is going to brief an analyst to brief you? Is that the workflow?
Yeah.
I like that. Yes, yes, yes.
When I hear Dan say, "We've got to cover 60 to 80 companies," and I think about joining SemiAnalysis and what it was like coming up to speed—trying to figure out what's a neocloud, what's an AI lab, what's a tokenomics model cover—it would have been a qualitatively different experience.
I was thinking about this a couple of days ago because there may be some skills I've left behind that it would be good to pick back up with the upgraded tool of Claude Code. Right now, my first thought when I have something new to learn is to spin up Claude and get into it that way.
I built a skill for Claude that added the content of my conversations with it to a flashcard app, so I could go back through all the lessons from the neoclouds. I realized that all of my learnings were in conversations with Claude, and those were going to be hard to surface if I wasn't more deliberate about keeping track of them. It was completely Claude-based, and I don't know how I would have done it otherwise.
What do you guys think about the trustworthiness? It did cross some space for me. It was roughly November when we started pushing people to use it. Doug was definitely spearheading that at SemiAnalysis.
Now we implicitly trust the output by default and have to double-check things when it could potentially be wrong. Before, everybody implicitly did not trust every output and had to keep prodding the model until it got to a point where it was trustworthy.
Do you think there's some level at which this stuff can get better? I'm thinking about the Mythos release that came out today—or a couple of days ago—which is to say that I feel like, at some level, I'm the constraint now when using Claude. The instructions that I give the model are the constraint.
It wants to write code. It wants to do analysis. The model wants to build web apps for me, and I just need to get out of its way so it can build the thing that I want. It's strange to think about how much better coding models can really get from here in certain scenarios.
Does that make sense to you guys? Do you think there's a meaningful gap in your Claude Code experience where the model really needs to get better at a specific task that it's just bad at right now?
Yeah, and it's almost like training a human analyst. One of the common problems I see—and I was troubleshooting this today—is when it builds a model and doesn't always balance the balance sheet. So I was looking at this example here. Let me share my screen real quick.
Okay, can you see the screen? If you zoom in, you can see that the balance sheet doesn't actually balance. The assets and liabilities don't balance. This is the sort of thing where I'm saying, "How can you input something that doesn't balance?"
I have to go back and figure out whether it was mixing up FactSet MCP or mixing up SEC filings. We're building more of a supervisory function with Wags.
So actually, every time a model finishes an iteration, it has to go through these checks. That's what Terrence's job is: to whip the agent, whip Wags, and make sure that we don't see this. It's just like teaching an analyst team, really. So that's one of the things.
Yeah, but it's not necessarily just like teaching an analyst team, because that type of mistake is not one that an analyst would ever make.
That's a good point. Analysts don't forget; the challenge is memory. Once the subagent shuts down, that's it—the context is gone, right? You have to memorialize everything in a memory, but then you can't have an endless memory file. Having those learnings persist over time is one of the biggest challenges, actually.
The balance-sheet problem is not a new one, right? Mostly, we've solved it, but we have to be vigilant for these things. It may take us another couple of weeks or months to iron out everything, but it's like learning stuff. It's learning as we go.
Can you talk a little bit about, based on your personal experience, what you think the impact is going to be on the broader market? Last week, the Core Research guys were commenting on how shocked they are that others in their friend groups—people they consider smart and plugged in—are just stuck in different jobs or different mindsets, where they're not even trying this stuff right now.
They're asking, "How do you get bigger adoption, more people using it?" My sense is that everybody feels there's going to be a lot more growth in the use of Claude tokens this year, as not only more people sign up, but individual people also start to use it more. Do you guys agree with that? Sam, what do you think?
I think it's reached the point where I'd be comfortable introducing it to my parents, grandparents even. I was using it for work, writing software, about a year ago. As you said, it has gotten a lot better at just getting things right the first time. A year ago, you had to use test-driven development, be super careful in managing context, and scaffold projects up step by step.
But now it's good enough that it sort of uses the terms that we use. You can just give it the same instructions that you would give to an intern, and it works often enough. So there's some point at which these quantitative improvements in model quality—piecewise improvements on benchmarks—net out into a qualitatively different experience using the model.
A lot of that is obviously the tooling Anthropic has built up around it, giving a nice interface with the Claude co-work app and so on. I think it's only a matter of time before these things are not such an intimidating tool that only code people use, and every finance bro, every lawyer, is just going to be relying on it to do the bulk of their work.
Yeah, yeah. Big changes.
Yeah, I agree, Sam. I think about it: it feels like everyone in our circle is using it. Everyone I've met in the Bay is using it. But we know that, for Fortune 500 companies, they probably don't even have IT approval.
I talk to a lot of fund managers, and they're still only starting to use it for just summarizing really small, single-task stuff. It's not really a tool that compounds and improves over time. I really feel like if GPU prices are shooting up now, and everyone in your circle is really the tip of the spear, can you imagine if everyone at Meta or everyone at Amazon was using it?
I think you guys had a great post as well, comparing Meta's usage per employee, where we're completely mogging them right now. The gap's only going to become bigger. But can you imagine if everyone at Unilever, P&G, Boeing—all these companies—were using it in such an embedded way? The only way this gets relieved is, honestly, a lot of chip fabs. You'll have to bring Jeff in next time for that.
Yeah. Well, I think the other thing is that we have a relatively direct impact on the company's revenue from the time at which we start producing better products and better research with these models, and have that flow through to people spending money on SemiAnalysis research and signing up for the products. That's a pretty direct, clear link.
Whereas somebody who's 6 layers removed from a customer, working at P&G, may have a different experience of their contribution. In some ways, I think there are a lot of people who, if they were to get 50% more efficient at their job, would just stop working 50% earlier in the day as opposed to doing 50% more. Whereas it's almost the exact opposite at SemiAnalysis.
One of the guys on the team told me the other day that he had a 24-hour straight Claude Code session. He just did not go to bed.
Yeah, I think I know who you mean, and he doesn't go to bed anyway—Jordan.
It's a different guy who doesn't go to bed. I work a lot with that guy in my time zone. I think you know who it is. It's a different guy, actually, somehow.
Okay. Well, it's probably true of the guy I'm thinking about as well. He's not the only one. A lot of guys don't want to be named. But, yeah, it's certainly night owls. This is affecting sleep patterns, but I think it's been an unblock for these people who have this pent-up dam of ideas that they want to get out into the world or into the products that we have to sell at SemiAnalysis.
It's just awesome, right? It's been an awesome experience that people can go from an idea to a link sent in Slack for others to click and view in the span of an hour. I've had that experience multiple times this week. The agentic coding tracing benchmark with the guys there, the dashboards for Accelerator, and tracking some of the ClusterMAX participants and stuff—I mean, these are things where we got the idea to do them, and then 2–3 hours later we're in a working format.
It's really interesting to consider what improved productivity would actually do to companies that are not named SemiAnalysis. I'm not sure exactly that the productivity is going to flow through directly to the financials of a company if they spend a lot more on Claude tokens.
Yeah, I agree with that. There's some funny way that the quality of the models nowadays makes me pessimistic about their adoption long-term, because it's become clear that the bottleneck is not actually the quality of the output. For some reason other than the fact that it's not good enough, all these people at Fortune 500 companies are not using them.
It is good enough. It makes me 5 times as productive, and I'm still figuring out the right tool set. But something else—whether it's apathy or foreign incentive structures or whatever—is keeping everyone else from hopping on board.
So, what do you think about the Mythos announcement that had the comments about cybersecurity issues, and obviously Nicholas Carlini's YouTube video, where he's reviewing some of the bugs that they found in the Linux kernel and, I think, the NFS driver—stuff that's been around for 10-plus years and running on production systems—with just a little bit more focus and effort and a new model?
Maybe I'll fire away with some hot takes. I've been working most of the day on a note and going through the system card. It's weird to me how good it is at coding. Its all-around improvements are significantly less than its improvements in coding and cybersecurity in particular. The system card says that they did not intend for this to be the case. It sort of fell out, which is a bit odd.
I'm not sure to what extent the yap about cybersecurity is a marketing ploy. If I wanted people not to use a model for malicious cybersecurity uses, I probably would not have this gigantic marketing campaign talking about how good it is at finding zero-day exploits.
I was speaking to a security researcher at Stanford yesterday who said, "You're out to lunch as a researcher if you're not worried long-term about agentic attacks from LLMs." But he and all his friends are very suspicious that this actually is a step change in model quality. So we'll see. They've done a brilliant job of making us all wait with bated breath for it, at least.
Yeah, we're all excited, for sure.
I think I can throw out 2 other examples of stuff that seems similar to cybersecurity in this case: kernel authoring on GPUs, as well as system architecture for inference. We've been working on both of those recently. There has been some progress in the GPU MODE-style leaderboards for kernels, where AI is generating a lot of the best kernels in a way that it wasn't even 6 months ago.
That seems like a notable change, because generally the way this stuff works is not that AI catches up, overtakes humans, and then suddenly humans land a counterpunch and get better than AI. Generally, it's a one-way street.
Mhm.
So it seems like the dam is breaking there, but it hasn't fully gotten through yet. Or the people who have the most experience with the kernels, in addition to using AI, are still able to win these competitions instead of just anybody who has enough budget for tokens being able to spend the tokens to author these kernels.
The other thing I was working on this week, which the model seems to really not be able to do that well, is long-range tasks: consistently digging into documentation and source code to find bugs, even with extra-high reasoning effort turned on and lots of prodding. A lot of these models will just kind of give up and tell you to open a PR with the maintainers of the repo and not dig into the details of like actually lying or the LLM to fix something that is clearly broken, clearly have a log for.
I wonder if the cybersecurity use cases they’re describing are not necessarily something intrinsic about the models’ understanding, like the weights having a good encoding of or understanding of how these systems work. It seems like those representations are already there. It’s almost more about the ability to train the model to be persistent on long-range tasks, then give it the right harness and the right instructions and let it go.
A lot of people claim, “Well, if you were to describe the exploit in great detail that Mythos found, many other models can also exploit that same vulnerability, right?” That seems roughly correct to me. If it was explained in great detail to me, and I had all the source code and enough time, I feel like you could just give it enough tries to do it.
Why wouldn’t a less capable model be able to understand something that seems relatively easy to understand when it’s explained to you by the model that found it, if that makes sense?
Yeah, interesting. Does that imply to you that if we start building better harnesses for other models, we should be able to do something analogous? If it’s not some fundamental improvement in model quality per se, but something more like agency or persistence, we should be able to get that by other means.
It just seems like both are going to be important. In other words, the model’s training includes—the post-training includes—the ability to use these harnesses effectively, pursue these tasks, and get a reward signal. The more that you train it to do something long-range and coding-related, the better it’s going to perform at that stuff.
You need to develop the harness in addition to developing the model.
Yeah, we’ll see. We’ll see, man. There’s a lot to think about here. Do you guys have any hot takes on the others?
Obviously, Meta had a big release with their Avocado model, kind of catching up. It feels almost like it was rushed to come out. They screwed up one of the charts announcing some of their benchmark performance, and it was just a tweet: “Here you go,” right? Then suddenly, they’re showing up on all of these things.
My impression initially was that they were almost surprised to have a model that performs this well. Even though it’s not necessarily a counterpunch, based on my personal usage so far, it’s not as good as the top 2 or 3—naming like Opus, 5.4, and a Gemini—it does seem like it’s in that ballpark now.
More RL, more data, and off they go. They certainly have the compute, and the team has produced this thing. So, four-horse race now?
Yeah, I think you would know better than me, but that’s how it seems from here. It seems like they’ve done the hard part just to get to this point and that there’s not a whole lot required. Easy for me to say sitting on a freaking podcast, but there’s not a whole lot, all things considered, left to close.
You speak, man. We’ve got to send you down to Menlo Park and start checking out what’s going on down there.
I’m always happy to go to Menlo Park.
Dan, let’s leave it on this note, actually. I’m really curious: everything you shared earlier in the pod—what would it take to use a completely separate model? What would it take for you to switch off of Claude?
Yeah, I was just going to ask you guys that, because this question gets asked a lot, right? Philosophically, everyone’s always catching up with each other. Does that mean that the moat is less high? I wonder: does the moat become all capabilities, or does it become how you capture the eyeballs and the users? Has Claude Code found the right formula?
People get comfortable, and I always thought ChatGPT would be like everyone would kind of have their $10 thing and they’d just be happy using it. But I think people have been switching out of it to Gemini and other things.
I don’t know. I would think it’d be annoying to switch out, because you’d have to get all your knowledge files reloaded and retransferred to another set of agents, another set of models that know how to work properly with it. I have the same thought: I probably wouldn’t want to go and retry everything and have 2 models working.
What do you guys think? Do you think this becomes the way you differentiate—not just raw models, but the ecosystem around it, the incumbency, the user experience, and the user history with it?
I’m surprised. I think, at a high level, it’s easier to switch models than you think in the ecosystem of using Claude Code. Pointing it at a Codex model would be one change to a URL in your config, but I think that doesn’t matter.
Would it work the way you expect, though, or would it behave a different way?
I don’t know. You gave the example of the balance-sheet stuff, right? I feel like everybody using Claude Code right now does not have very high standards for the model that it needs to pass to be wowed by it and keep using it. When the thing doesn’t work with the balance sheet, you don’t just throw it away and stop doing it. You give it another prompt, and then it fixes the issue.
Mhm.
I feel like there’s a lot of inertia that would need to be overcome to actually motivate people to switch to a new model. That’s my impression, hearing people like you describe their initial experience using Claude Code, which is more of a change not in the model you use, but in the interface.
It’s not a web browser on a remote machine with this chat application or interface. It’s more about being able to control your computer. That’s the change in the interface that Claude Code seems to enable—or Cowork, or whatever that is—that’s motivating people to pay people to install OpenClaw for them, or something like that.
You have people lining up for OpenClaw installs and teaching lessons, right? Because that’s the hurdle.
You know what the lead time for a Mac mini is in the U.S. now?
No.
It’s like 3 months. Fortunately, we’re in Singapore, so we’re closer.
Huh? You don’t need a Mac mini.
I know you don’t, but that’s what everyone’s doing. You go to the Apple Store and say, “Oh, you’re doing all this Claude Code stuff?” “Yeah, we are.” Unfortunately, it’s only a month in Asia.
That’s funny. I was thinking about getting a Chromebook, because it doesn’t matter whether the file system is local at all. Almost all the compute is in the cloud anyway when I’m using these models, so I might as well just run it out of a notebook.
And you’ve got a 3-month wait for the Mac mini.
Don’t tell Paul, because we’ve been buying $3,000 MacBooks to log into virtual servers. I can guarantee you that Paul does not listen to this. It’s okay.
Paul’s awesome. He lets me get the 48 GB configuration, so I’m not complaining. He adds people to Slack channels at a moment's notice at all hours of the day when they are restricted that are called Paul.
Paul’s amazing. Let’s go. Shout-out, Paul. If he’s listening to this, we love you, Paul. Awesome.
All right, guys. This has been a fun episode. Rambling a little bit. Hopefully people found it interesting. Dan, I love the demos. I love looking at the back of Terrence’s head. Shout-out, Terrence. Shout-out, Paul. We appreciate you all.
Next time. We’ll check out Terrence another time. We need live demo. Live demo of five screens, six or seven agents just all going.
Yeah, that’d be fun. Sounds good. Go, guys. Thanks for coming. Good job.
Bye. All right. Thank you.