[BidClub_]
Moonshots · · 89 分钟

AI Now:Elon 的1万亿美元激励、Apple 为 Trump 承诺的6000亿美元,以及小型创业公司如何胜出——与 Dave、AWG 和 Blitzy 对谈

Peter DiamandisDave BlundinAlexander Wissner-GrossBrian ElliottSid Pardeshi

YouTube
TL;DR
  • 只有在创造远超自身价值的前提下,Elon Musk 潜在的1万亿美元 Tesla 奖励才具备合理性。 Musk 必须帮助 Tesla 达到8万亿美元估值,而他的角色同时涵盖 CEO、首席营销官和注意力引擎;Diamandis 指出,传统车企通常将约7%的营收用于营销,Tesla 的投入却是“零”。更广泛的创始人启示是,沟通已经成为产品的一部分,但“它还会再次改变,而且还会再次改变”。

  • AI 这轮万亿美元资本周期,必须靠劳动力自动化和变革性科学来创造与投入规模匹配的回报。 讨论中提到的一场晚宴上,美国敲定了6000亿美元投资承诺;Dave Blundin 表示,由于视频片段经过剪辑和混杂,无法确认 Tim Cook 还是 Mark Zuckerberg 先提出承诺,但 Zuckerberg 随后进行了匹配;他还提到,到2029年 OpenAI 预计将获得约1190亿美元追加投资,以及此前对 AI 芯片及相关基础设施5万亿—7万亿美元投入的预测。Alexander Wissner-Gross 的逻辑是:先把一类类劳动力和服务的成本推向可忽略水平,再创造足以支撑资本开支的发现。

  • 丰富供给会让货币贬值,却不会消灭稀缺,只会把瓶颈转移到别处。 Peter Diamandis 想象分子组装器以接近零的边际成本制造一辆电动 Ferrari;Wissner-Gross 则反驳说,Star Trek 有复制器,但星际旅行仍然稀缺。他提出的核心问题是:当能源和智能趋近于零成本时,“什么仍然稀缺”——这提醒人们,不要假设所有受限资产会同时消失。

  • Mercor 证明,极端 AI 估值仍然可以由经营表现驱动,而不只是纯粹投机。 Blundin 表示,公司在2年内从约3000万美元初始估值成长到100亿美元,同时达到5亿美元营收年化运行率;创始人起步时年仅18岁。他的创投原则是寻找“被低估、未获充分认可的人才”,更大的变化则是,20—23岁的创始人如今可以进攻过去通常只向经验丰富得多的团队开放的市场。

  • Blitzy 对企业市场的判断是,代码生成正在变成商品,但理解并安全改造规模庞大的代码库仍然不是。 其平台声称支持超过1亿行代码,已成功接入约6000万行代码的代码库,单次任务运行时间从12小时到数周不等。输出目标是交付已经完成验证、编译和测试的代码,因为“拉取请求的另一端”是昂贵的人力劳动。

  • Blitzy 报告其 SWE-bench Verified 得分为86.8%,高于视频录制时排行榜的75.2%,但更大的信号是,这项基准可能已经接近耗尽。 公司表示,结果可以通过 SWE-bench CLI 复现,并未使用针对基准测试的专用脚手架;此前它处理了500个分支,合计摄入约4亿行代码。Wissner-Gross 认为,86.8% 实际上已接近100%,因为剩余部分很大程度上可能是有缺陷的测试,由此需要针对 Linux、VS Code 等代码库开展企业级测试。

  • 创业公司的打法,是成为前沿实验室的大客户,深入解决它们无法产品化的问题,并在模型升级时同步升级。 Blitzy 负责编排 Gemini、Anthropic 和 OpenAI 的模型,而不是训练前沿模型;Diamandis 将实验室的投入概括为“为 Blitzy 进行的1万亿美元研发”。公司当前的企业级主张比单纯的代码生成叙事更扎实:在适合的项目中,约80%的工作可以自动化,端到端开发速度约提升5倍,同时明确标出剩余任务交给人类完成。

摘要 · 为研究而整理的核心内容

1. 1万亿美元奖励,给 Musk 的定价同时包含操盘手与分发渠道

  • Diamandis 先抛出标题,Blundin 随后解释条件:只有 Tesla 达到包括8万亿美元估值在内的一系列指标,Musk 才可能获得约1万亿美元 Tesla 股票奖励;Blundin 称,这相当于把 Microsoft 和 Nvidia 的规模合计再翻一倍。他的辩护很简单:“这只是你创造价值的一小部分。”

  • Diamandis 认为,Musk 改变了 CEO 的经济职能。他“不只是公司的领导者,还是公司的营销声音”;传统车企可能将约7%的营收用于营销,而 Tesla 的投入为零,因为 Musk 个人创造了注意力、势能和客户需求。

  • Diamandis 将这一观察延伸到风险投资选人:如果创始 CEO“非常、非常害羞”,可能会陷入困境,因为投资人如今需要能够公开传递信念的沟通者。但他也提醒,复制 Musk 并不是永久有效的公式——媒体正在变化,CEO 的范式“还会再次改变,而且还会再次改变”。

2. 万亿美元资本开支,必须对应万亿美元级的经济后果

  • 讨论中的 Trump 晚宴据称敲定了一项6000亿美元美国投资承诺;Blundin 认为 Tim Cook 先提出了承诺,但由于视频经过剪辑和混杂,他无法确认,随后 Zuckerberg 进行了匹配。Blundin 戏剧化地还原现场:那场面像一场筹款会,而某位硅谷 CFO 在一旁嘀咕:“他到底承诺了什么鬼东西?”

  • 讨论还提到,到2029年 OpenAI 预计将获得接近1190亿美元的追加投资。Wissner-Gross 回忆说,大约18个月前,市场曾认为 AI 芯片,以及晶圆厂、数据中心和能源等相关投入合计5万亿—7万亿美元的预测荒谬至极;但如今,这种规模的支出已经“完全变得可信”。

  • Wissner-Gross 进一步追问资本需要获得怎样的回报:既然市场投入了数万亿美元,就会期待规模同样巨大的营收,合理来源可能是自动化整类劳动力或服务,将其成本推向零。再往后,AI 可能还必须带来变革性发明和科学发现;否则,他问道:“为什么要在这上面投入数万亿美元?”

  • 这引出了关于丰裕的争论。Diamandis 想象分子组装器把能源、土壤、钛和锂转化为一辆边际成本接近零的电动 Ferrari;Wissner-Gross 则以 Star Trek 回应:每个人都有复制器,但并非每个人都能进行星际旅行。丰裕不会终结稀缺,只会改变“什么仍然稀缺”。

3. Mercor 证明,应更早押注极致人才

  • Blundin 将 Mercor 描述为一段由经营加速支撑的估值故事:公司在2年内实现约5亿美元营收年化运行率,获得100亿美元估值报价,最初融资估值接近3000万美元。“我不希望大家觉得这是泡沫,”他说,因为营收增长本身也打破了先例。

  • 他的选人原则是寻找“被低估、未获充分认可的人才,而不是太关注概念”,尽管 Mercor 当时已经拥有正确的概念。真正不寻常之处在于,投资对象是18岁的创始人:他们大学只读了约1年,就开始对教育进度感到不耐烦。

  • Diamandis 将今天20—23岁的创始人与他记忆中10年前风险投资支持的独角兽创始人平均30多岁的年龄进行对比。Wissner-Gross 预计,教育激励还会进一步扭曲:如果学生相信先进 AI 即将到来,那么积累学分可能不如立即创业划算。

  • Diamandis 援引研究称,诺贝尔奖级别的工作往往发生在研究者20多岁的前半段。Wissner-Gross 则不愿把这一点变成可预测规律:纯 AI 或人机混合体可能很快承担大部分创新工作,年龄与生产率统计充其量只能说明“过去是怎样的”。

4. 一款面包店应用,暴露出后来成为 Blitzy 的瓶颈

  • Elliott 和 Pardeshi 在 Harvard Business School 相识,但两人的运营经历截然不同:Elliott 经历了 West Point 和军旅生涯,Pardeshi 则曾在 NVIDIA 工作,并从 Attention Is All You Need 时代起就参与生成模型研究。两人的友谊后来异常深厚;Pardeshi 还是 Elliott 小儿子的教父。

  • GPT-3 到 GPT-3.5 期间,一项为 Boston 面包店无偿完成的项目成为火花。面包店原本预计要为移动应用支付30万—40万美元;Elliott 和 Pardeshi 用一夜或一个周末,通过如今被称为 vibe coding 的工作流把应用搭了出来。Elliott 开玩笑说:“我们装得好像花了更久。”

  • 他们观察到,瓶颈已经变成人类:两人反复把错误反馈给一个模型,再把输出交给另一个模型,持续迭代,直到得到可以编译、通过测试的代码。公司的创始假设,就是把这种多模型精炼过程自动化,将商品化的开发劳动力从链路中移除。

  • 如今 Elliott 将 Blitzy 定义为“企业级自主软件开发平台”。平台吸收整个代码库,把 COBOL 转 Java 或持续开发等需求转化为长时间运行的推理任务,并力图交付已经完成验证、编译和测试的高质量代码。

5. Blitzy 进攻规模过大、历史过久、已超出人类理解范围的代码库

  • 金融机构可能仍然依赖 COBOL、PL/I 及其他运行了数十年的系统:它们确实能工作,却令人不敢轻易修改。了解这些系统的人正在成批消失,而历史上整体替换往往不具经济性;管理层也常常无法看清数千万行代码究竟在做什么。

  • Blitzy 首先为源代码建立索引,为企业绘制功能地图,再利用这套上下文进行迁移或开发新功能。Blundin 称,让1000万行没有文档的代码用通俗英语解释自身、并找出其中的 bug,这种体验“令人震撼”。

  • Elliott 表示,2000万行代码的代码库很常见,成功接入的最大代码库约有6000万行;平台声称的上下文容量超过1亿行。任务运行时间可以达到12小时,在超大代码库上甚至持续数周——这是一套有意与交互式 copilot 分离的工程机制。

  • 这正是 Blundin 的投资逻辑。Cursor、Windsurf、Replit 和 Lovable 擅长快速交互,但一个运行5—7分钟的 agent 处在“无人区”:对话太慢,对实质性工作又太短。Blitzy 则可以运行一夜甚至整周,最终返回一个大型拉取请求。

6. 86.8%的 SWE-bench 得分,几乎让当前标尺饱和

  • Wissner-Gross 解释,SWE-bench 衡量的是 AI 系统能否解决典型 GitHub issue:理解代码库、修复 bug 或性能问题、提交拉取请求,并通过测试。SWE-bench Verified 则进一步缩小范围,只保留500个由 OpenAI 研究人员审核、确认可解且值得解决的问题。

  • Pardeshi 表示,视频录制时的排行榜最高分为75.2%;Blitzy 通过 SWE-bench CLI 验证后取得86.8%。尽管基准测试使用12个代码库,但这500个分支被 Blitzy 摄入时合计约4亿行代码;Blitzy 估算,其累计代码语料库约为10亿—20亿行,取决于是否将更新计算在内。

  • 可复现性是这次发布的核心。Pardeshi 表示,Blitzy 没有增加专用的 SWE-bench 脚手架或仅针对基准测试的功能;他还对比了一些案例:某前沿模型宣称接近80%,但复现结果只有约60%。“我们高度重视可复现性,以及实际应用价值。”

  • Diamandis 质疑,80%以上的基准得分是否已经意味着测试饱和。Wissner-Gross 认为,在86.8%的水平上,剩余部分很大程度上可能是有缺陷的测试,而不是更难的问题。他希望下一代留出测试集包含约2000万行的 Linux 和约400万行的 VS Code;结尾讨论提到,Blitzy 正与 MIT 合作开发后继基准。

7. 编排机制让每次前沿模型发布都变成产品升级

  • Elliott 对 Mag 7 威胁的回答是:它们已经位于 Blitzy 内部。公司让 Gemini、Anthropic 和 OpenAI 的模型相互协作、相互校验;数百种模型、提示词和工具组合,将质量推高到单一模型无法达到的水平。

  • Blitzy 将这一机制称为“扩展推理时验证”,并辅以领域特定的上下文工程。在每一个阶段,agents 都必须在企业级代码库中保留正确的功能上下文,验证候选输出,并持续迭代,而不是把第一次生成的代码当成最终成果。

  • Pardeshi 表示,他们在模型上下文窗口只有5000个 token 时就押注了这一方向:上下文会扩展,编程能力会提升,因此没有必要通过训练另一个模型来与 Mag 7 竞争。他们可以“站在巨人的肩膀上”;Diamandis 将前沿实验室的投入概括为 Blitzy 可以借力的1万亿美元研发,Elliott 则表示,底层模型每次升级,Blitzy 都会随之改进。

8. “大重构”在技术上可行,但商业资金尚未到位

  • Wissner-Gross 建议,让 Blitzy 处理支撑现代文明的 Linux、Python、GNU 及其他依赖组成的“羊皮卷”。他偏好的宏大项目,是用 Rust 或其他内存安全语言重写存在漏洞的遗留库:通过重建软件供应链,消除整类弱点。

  • Pardeshi 表示,早期实验已经包括把一款20年前、仅适用于 Windows 的 MATLAB 库转换成 Python,并使其实现操作系统无关。在另一次测试中,Blitzy 从 NVIDIA 代码库中选取一个公开 issue,并返回了能够解决该问题的拉取请求。

  • 一次更具商业说服力的演示,使用了一个 AWS 代码库:它被刻意设计成杂乱的主机代码,混合了多种历史时期的编程风格。Blitzy 自主将其从 COBOL 迁移到 Java,并实现开箱即编译;运行时工作仍然存在,但针对 Elliott 认为通常需要数年的项目,推理过程只花了约1周。

  • Pardeshi 精确地描述了护城河:“从 AI 获取代码是商品。”只有当重写必须保持既有行为、完成编译、通过每一项单元测试、不引入新的安全漏洞,并满足企业特定约束时,价值才会出现。没有这些验证,聊天机器人生成的重构代码就不是可部署资产。

9. 企业经济学偏好经过验证的输出,而不是最便宜的代码行

  • Elliott 表示,Blitzy 生成一行代码的成本大约是其他供应商的100倍,但可能只有人工编写成本的百分之一。企业有意为质量付费,因为便宜代码一旦触发昂贵的审查、诊断和返工,只是转移成本,并没有消除成本。

  • Wissner-Gross 用 AI“超级通缩”重新解释这项溢价:如果推理成本每年下降约10倍,那么100倍溢价相当于2年的成本下降。Elliott 认可文明级软件重写今天就有价值,但眼下的商业问题是“谁来付钱?”因此银行和保险公司会成为优先客户。

  • Elliott 反驳 Wissner-Gross 将劳动力成本推向零的判断。软件需求可能接近无限,而 Blitzy 在一个平均规模的大型问题上可以自动化约80%的工作,明确知道哪些部分无法完成,并将一份干净的拉取请求及清晰的剩余任务清单交给人类开发者。

  • 因此,公司的生产率主张是速度提升5倍,而不是假设代码生成速度提升1000倍。企业可以提前启动下一轮开发冲刺;开发者拿到的主要是已经写好的代码,以及剩余工作的指南。Elliott 强调,协调、需求、文档和发布流程的开销,对实际生产率的影响远大于 token 输出速度本身。

10. 真正的事实来源,从代码转向功能意图

  • Diamandis 观察到,传统软件把可运行代码视为事实来源,文档只是外围辅助。如果代码可以一夜之间重新生成,那么规格说明就会成为核心资产。Pardeshi 回应说,规格说明仍然只是抽象层;Wissner-Gross 则补充,人类可读的规格说明是系统更深层功能理解之上的中间表示。

  • Blitzy 更深层的表示形式,是一套客户专属的混合图—向量数据库,用来捕捉软件必须实现的功能,而不依赖具体实现语言。这套由企业拥有的功能模型,既可以生成可读规格说明,也可以支持从一种语言迁移到另一种语言,同时不丢失必要行为。

  • 递归式自我改进已经存在边界。Blitzy 的大量冲刺工作由 Blitzy 自身生成,但 Elliott 区分了可替换的软件与核心发明——例如 Pardeshi 设计的、用于确保编译并防止循环依赖的算法。重写整个代码语料库,可能复现其形状,却未必能复现底层发明。

  • 关于 METR 式自主性时间跨度,Pardeshi 表示,部署、CI/CD、调试、安全分析、监控和追踪已经分别存在;MCP 和 A2A 可以把这些组件连接起来。他预计,在满足明确条件的项目中,端到端自主交付“还有几个月”,目前产品经理、架构师、QA 和提示词编写 agents 已经在生产环境运行。

11. 创始人可以通过成为巨头不可或缺的大客户来竞争

  • Elliott 预计,政府现代化将成为一个重要市场:政府机构拥有影响身份识别、出行和关键服务的大量遗留系统。Blitzy 目前尚未服务美国政府,但他认为“12个月”内实现是可信的;Diamandis 则认为,只要成功部署一个政府机构,就可能向其他机构扩散。

  • 他的战略规则是建立相互依赖:成为前沿实验室的大客户,从而“你希望他们成功,他们也希望你成功”。Blundin 补充说,大型模型公司也会从繁荣的合作伙伴生态、营收和竞争中受益,并可避免遭遇反垄断拆分;创业公司应该与它们保持足够近的距离,以理解其发展方向。

  • Elliott 应对复制的保护,是领域深度。亲自经历过企业安全壁垒、流程缺口和产品失败的创始人,能够跑在被官僚体系拖累、资金充裕的 incumbent 前面。最基本的配方是“正确的人才、适量的资本”,以及由真实企业验证过的问题。

  • Elliott 将 Blitzy 的高强度文化与他2017年的 Ranger 任务联系起来:约100人被秘密派入 Syria,在 Raqqa 对抗约2000名 ISIS 武装人员,并必须在当地招募力量夺回这座城市。Blitzy 首先筛选雄心和创造力,自称是“997”文化,并将软件自动化定义为 GDP 扩张;Wissner-Gross 最后的号召只有一句:“解决一切。”

Peter Diamandis

A piece of news I saw yesterday had me scratching my head: the trillion-dollar pay package for Elon.

Speaker 1

Elon Musk could officially become the first trillionaire, with just about $1 trillion in stock.

Speaker 2

It's a striking number, but the benchmarks that he has to meet are also equally striking.

Dave Blundin

He's not just the leader of the company; he's the marketing voice.

If we really do expect to find ourselves in an abundant society soon, we should expect to have a lot of trillionaires in our society.

Peter Diamandis

Money will start to have far less value than ever before.

Alexander Wissner-Gross

If we're really on the verge of abundance, then what comes after that? What's going to remain scarce even as energy and intelligence—the cost of both of those—goes to zero?

Peter Diamandis

Stuff is changing so quickly. Media is changing quickly. Elon is paving a new path for what it means to be a great corporate CEO. But it's going to change again, and it's going to change again, and it's going to change again.

Speaker 3

A trillion dollars here, a trillion dollars there. How do you compete? What's your mode?

Peter Diamandis

So, Brian and Sid, welcome. It's a pleasure to have you both.

Now, that's a Moonshot, ladies and gentlemen.

Everybody, welcome to Moonshots, the news that really matters in your life.

I'm here with my Moonshot mates, Dave Blundin and Alexander Wissner-Gross. We're going to have a conversation today about David versus Goliath. We're talking about trillion-dollar commitments and trillion-dollar pay packages. It's insane.

But first of all, one of the most important pieces of news is Alex. I'm reading the comments, and what I keep hearing is that people want me to ask you a question: Have you talked for the entire episode? They love what you have to say. I'm not going to do that today, but everybody, if you're new to Moonshots, Alex has been receiving huge fan mail because of his brilliance.

Alexander Wissner-Gross

It's an honor to be here, Peter. I think I got maybe 2 adoption offers.

Peter Diamandis

I think you did. Dave, are you jealous?

Dave Blundin

Am I jealous of Alex? I'm surrounded by so many people who are so brilliant. I'm used to it by now.

Peter Diamandis

That's one of the most important things: finding incredibly brilliant people to have around you in life, right? The old saying is, “You're the average of the 5 people you spend the most time with.” If you've got a community of individuals who are really uplifting you and challenging you, who have you doing the best you can, that's critically important.

One of the things I love about Moonshots and our WTF episodes is measuring toe to toe and trying to have this conversation in a meaningful fashion. We're missing Celestial again. Selee, we miss you, buddy. I think he's back in India.

Dave Blundin

He's probably got a crate full of iPhone 17s coming, but we'll be talking to him soon about that.

Peter Diamandis

I want to open up with a piece of news I saw yesterday that had me scratching my head, which was the trillion-dollar pay package for Elon at Tesla. That's extraordinary. Have you guys gotten an offer from your boards for $1 trillion?

Dave Blundin

A trillion. Well, hey, if you hit the metrics, the metrics have to be a multitrillion-dollar market cap, but sure, why not? It's just a fraction of what you create. If Elon grows Tesla to an $8 trillion company, so it doubles the size of Microsoft and Nvidia, he'll earn $1 trillion. It's not like he needs help becoming the first trillionaire on planet Earth, right?

Peter Diamandis

Yeah, but he's unique. He's really changed the definition of what it means to be a CEO. He's not just the leader of the company; he's the marketing voice. Most car companies will spend 7% of revenue, or thereabouts, on marketing. Tesla spends 0% because Elon is a one-man force of nature driving consumers to the product.

We're going to meet a couple of entrepreneurs in a minute here who have that same attitude: “We embrace social media.” We're creating morale and momentum like you wouldn't believe. That's a 20% pay package as opposed to the normal 5%.

Alexander Wissner-Gross

It's worth it.

Peter Diamandis

I agree with you. If you're a founding CEO and you're very, very shy, that's not going to bode well for the company. I know when I'm investing in companies, I'm looking for a CEO who's a great communicator, who's able to go out there in front of the crowd and convey their passion—what he or she is loving in life.

For all of the insanity Elon does with chainsaws or whatever the case might be, it grabs attention, and people either love it or hate it. One of the many reasons we love Alex so much is because he's not just brilliant, but he has very, very high situational awareness. That's really rare in the brilliant community.

Stuff is changing so quickly. Media is changing quickly. Elon is paving a new path for what it means to be a great corporate CEO. But it's going to change again, and it's going to change again, and it's going to change again.

If you map your behavior to the change—but most people study backward in time and say, “Well, what did Jack Welch do?” or “What did Genghis Khan do?”—that's okay, but that's not going to—

Anyway, we'll go on. This branch will go too long if I go too hard on this.

Dave Blundin

I also think, Peter, it's worth noting that if we really do expect to find ourselves in an abundant society soon, we should expect to have a lot of trillionaires in our society.

Peter Diamandis

We will. Then we'll have an expectation that money will start to have far less value than ever before, right? You and I have had this conversation, Alex, about a post-capitalist society. Do you still believe that's going to be the case? What are your thoughts on that?

Alexander Wissner-Gross

I think that's always the question: What does so-called late-stage capitalism even look like, to the extent that the concept makes sense?

If we're really on the verge of abundance, then what comes after that? I think what comes after abundance is closely tied to what remains scarce in an abundant society. In Star Trek, a common foil is that energy is relatively abundant, while intelligence is relatively scarce. The ability to travel between stars is relatively scarce.

The question I would ask is: What's going to remain scarce even as energy and intelligence—the cost of both of those—goes to zero?

Peter Diamandis

That is a critical question. My endpoint here, my mental experiment, is this: If I build a number of assemblers—in Eric Drexler's parlance—that are able to rearrange atoms, and I drop an assembler into my hand and say, “Make me 5 copies of yourself,” then I give each of you an assembler.

The assembler is able to use energy and matter and build anything. I drop an assembler into the soil here, and it starts pulling the atoms together to make me an electric Ferrari. It says, “I need a little bit of titanium and a little bit of lithium.” You add it, and all of a sudden you've got an electric Ferrari.

Everything starts to become effectively zero-marginal-cost, and that becomes a pretty cool society where anyone can do anything.

Alexander Wissner-Gross

Does it? Again, not to overindex on Star Trek, but in Star Trek everyone has replicators, and not everyone gets to travel between the stars. Maybe the new post-scarcity ability is the ability to travel outside the solar system.

Peter Diamandis

We're going to find out, because we're getting there really fast.

On the news item of “a trillion dollars here, a trillion dollars there,” there was a dinner with Tim Cook, Sam Altman, Mark Zuckerberg, and Trump. I guess during this dinner, an offer was made—was it Zuckerberg first or Tim Cook first?—to invest $600 billion into the U.S. economy.

Dave Blundin

I think it was Tim first. I can't tell from the clips, actually, because they get cut and mingled.

Peter Diamandis

Then the other one matches it, and all of a sudden, over dinner, Trump is getting $1.2 trillion of commitments into the U.S. economy.

Dave Blundin

You missed the really fun punchline there. Tim Cook had it all scheduled, planned, and budgeted. Then Mark said, “I'll match that.” It was like a YMCA fundraiser: “If you can do that, I can do that.” I'm sure a CFO is back in Silicon Valley going, “What the hell did he just commit to?”

Peter Diamandis

Oh, my God. There's a third piece that comes out in this related story, which is Altman announcing to his employees that he expects OpenAI to be the most capital-intensive company in history.

What was the number, Alex? Was it $119 billion of additional investment between now and 2029?

Alexander Wissner-Gross

It was something like that. I mean, do you remember? It was only a year and a half ago or so that the numbers of $5 trillion, $6 trillion, or $7 trillion of capex into AI chips were being floated, and a lot of people laughed at that.

Yet we're finding ourselves, a year and a half later, in a world where it's entirely plausible that the true amount of capital expenditure in fabs, AI chips, data centers, and new energy sources completely exceeds that.

Peter Diamandis

File that away, because you're dead right. This is the effect we see all the time: Something insanely mind-blowing is predicted 6 months in the future, and everybody's like, “Impossible.” Then it actually happens, and they're like, “Oh, yeah, well, it's just part of life.”

This trend is happening over and over and over again. The numbers you just quoted—everybody was like, “Oh, Sam's just blowing smoke. There's no way that's—” You know, “Trillion” being thrown around isn't a real word; it's just sort of a euphemism. But here we are, just a few months later. You're going to hear some benchmarks, actually, like SWE-bench later in this podcast, where things have just been crushed. The timelines will blow your mind.

Well, we’ll get to it when we get to it.

Dave Blundin

Well, then I guess the point I’m making here is there’s a huge amount of capital flowing here. I mean, go back to when all of us were starting our careers in the ’90s and in the dot-com era. The idea that there’d be trillion-dollar movements of capital in any particular company or industry was just mind-boggling. And here it is, routine.

Alexander Wissner-Gross

But this is, in some sense, a wonderful opportunity, with trillions potentially in capex being invested. There’s going to be an expectation, I would assume, by capital markets that there’s going to be enormous revenue generation that pops out of those trillions in capex. The question you have to ask yourself is: What form does that take?

At some point, with trillions invested, I think there’s probably a reasonable expectation that entire classes of labor or services are going to be automated, and the cost of what we currently construe as labor is going to be driven down to zero. Then perhaps, immediately after that, you start to need transformative science, inventions, and discoveries that will really justify the trillions in capex.

So it’s a blessing in disguise. I would argue that trillions in capex are going to motivate the demand and supply of utterly transformative discoveries and inventions soon. Otherwise, why invest trillions in this?

Peter Diamandis

Yeah. The concern, of course, that a lot of folks have is around inflation. Are these dollars really inflated dollars? We’re going to find out. But it’s interesting, Dave, in particular, as a venture capitalist, seeing the valuations of companies going at this level.

For the average person, how do you get into any of these companies when they’re coming out at multihundred-billion-dollar and trillion-dollar valuations? Being able to get in early is one of the areas that I think you’ve been focusing on. One of the other companies that came out of Link Ventures, where you were an early check, was Mercor. I just saw that Mercor has gotten an offer at a $10 billion valuation. You must be pretty happy about that.

Dave Blundin

It’s a $10 billion valuation, but it’s also a $500 million revenue run rate after 2 years, which is completely unprecedented.

Peter Diamandis

So go back—when did you invest in Mercor?

Dave Blundin

2 years ago, in the first funding. What you really want to look for is undervalued, underappreciated talent, and not so much concepts. But they had the concept right already. It’s rare, but at 18 years old, not a lot of people invest in the 18-year-old gang.

Peter Diamandis

So a couple of 18-year-olds come forward with this idea. Do you remember what the opening valuation was when you invested?

Dave Blundin

$30 million, plus or minus $5 million.

Peter Diamandis

Okay, so $30 million to $10 billion in 2 years’ time. It’s got to shatter all kinds of records. But again, I don’t want people to feel like that’s a bubble, because the revenue growth also shattered all kinds of records.

And so, from a cold start, I don’t think anything like that’s ever been done before. But you’re going to see a lot more of them now, too. They’re setting the trend for many other companies. I think what’s different about them is they’re inspiring an age group that normally would have been uninvestable 5 or 10 years ago, and now it’s kind of, “Wow, mainstream.”

We’ve made the point that the average age of a VC-backed unicorn a decade ago was sort of the mid-30s, in terms of the average age of the founders. Today, Dave, what you found out of the investments we’re doing, especially out of MIT and Harvard, is that it’s ages 20 to 23—and these guys were 18 when they started.

Dave Blundin

Yeah, they were 18. They got through 1 year of college and then got frustrated with the pace, like everybody.

Peter Diamandis

Which goes—

Dave Blundin

They met in high school.

Peter Diamandis

Which goes to the point, Alex, you made last time on the last WTF episode: If you really believe we’re post-AGI, on the verge of ASI—artificial superintelligence—going to college during those years and trying to get credits versus building something is not the right trade.

Alexander Wissner-Gross

It’s going to distort all sorts of societal cues and societal expectations. The best fiction treatment that I’ve seen of this is a novella by Vernor Vinge, “Fast Times at Fairmont High,” where you see this start to completely distort the way secondary education is run in this country. You start to see high school students and middle school students suddenly spending all of their time doing startups. I think it’s entirely plausible that we find ourselves in a near future that looks a lot like that.

Peter Diamandis

I agree. One of the points here that we need to realize, or that people need to realize, is that the max—sort of peak creativity—if you measure it by when a Nobel laureate does their Nobel Prize-winning work, not when they get the prize but when they actually did their work, is typically in the first half of their 20s. Alex, do you have the data there off the top of your tongue?

Alexander Wissner-Gross

Not at my fingertips. I’ve seen those statistics, too, and I’ve seen how they vary from field to field, purportedly between math, physics, and other fields. I also tend to discount this notion because I expect that in the very near future, most of the innovation is actually going to come either from pure AIs or some sort of human-AI hybrid.

I view those statistics, maybe self-servingly, as more of a retrospective—this is how things used to be, at best—versus how they’re going to be in the future.

Peter Diamandis

Every week, my team and I study the top 10 technology meta trends that will transform industries over the decade ahead. I cover trends ranging from humanoid robotics, AGI, and quantum computing to transport, energy, longevity, and more. There's no fluff, only the most important stuff that matters, that impacts our lives, our companies, and our careers. If you want me to share these meta trends with you, I write a newsletter twice a week, sending it out as a short two-minute read via email. And if you want to discover the most important meta trends 10 years before anyone else, this report's for you. Readers include founders and CEOs from the world's most disruptive companies and entrepreneurs building the world's most disruptive tech. It's not for you if you don't want to be informed about what's coming, why it matters, and how you can benefit from it. To subscribe for free, go to demand.com/tatrens to gain access to the trends 10 years before anyone else. All right, now back to this episode.

Well, there’s another company I want to talk about here, and we have a couple of guests to join us. If you’re an entrepreneur, listen to how they built this company. This is a company on the doorstep of being a unicorn itself, a company that you’re going to hear a lot about in the coming years. Dave, you want to introduce our guests and Blitzy?

Dave Blundin

I cannot wait. We have Brian Elliott and Sid Pardeshi, the founders of Blitzy. My son Jack interned with them this summer, and I tell you, it drove my wife a little nuts. She started thinking, “Wow, we’re going to play a lot of tennis this summer and have a great time.” Jack got so wrapped into the culture of Blitzy so quickly. It’s the most high-energy place I think I’ve ever seen. Morale is off the charts.

He pulled in his best friend from high school, Yash Blesetti. They pulled in a couple of other young computer science majors at Northeastern or Northwestern and a couple of other places. The whole gang worked all summer on SWE-bench, crushing numbers. I tell you, the morale of this group is like nothing I’ve ever seen. The mission is incredibly cool and fun.

Brian is a West Point alum. My one experience in life with West Point was Rick Dalzell, who was my biggest and most important customer ever. He actually ran all things complicated at Walmart—massive logistics, half a million people moving around—and then he got poached by Jeff Bezos to work at Amazon. He was the number 2 guy at Amazon right when they were in total hypergrowth, and he had a West Point background. He really understood morale, people, and logistics.

Brian went to Harvard Business School after that. Sid went to BITS, which is the MIT of India. It’s actually statistically much harder to get into than MIT, if you can believe that, Peter.

Peter Diamandis

I can. MIT let me in, so there have got to be some flaws there.

Dave Blundin

Oh, so humble. Sid, after that, spent a long time at NVIDIA and saw it go from tiny to monstrous, so that’s got to be inspiring. I actually don’t know their story before that, but they met at Harvard Business School.

They’re inspiring to a different class of people who were already in a career path, and then AI hits the world, but they’re nimble. They’re not going to watch it happen. This is way too rare—people remapping their entire lives to take advantage of what’s happening right now.

I hope a lot of the listeners today get a lot out of their backstory and their transition to building this incredible company, Blitzy.

Peter Diamandis

Yeah. I really want to frame the story here as David versus Goliath. We’ve heard about a trillion dollars here, a trillion dollars there, and how do you compete in that world?

If you’re a young entrepreneur, you’re building a company, and you’re wondering whether you’re going to get literally decimated in the wake of Google, OpenAI, or xAI just happening to release a particular feature, how do you compete? What’s your mode?

Brian and Sid, welcome.

Pleasure to have you both. Where are you guys this morning?

Brian Elliott

I am in One Kendall Square, here at The Link studio offices, with MIT right behind me. We just walk over the talent from the MIT AI Lab right over here to work, which works well for us.

Peter Diamandis

Beautiful, nearby in Cambridge. Fantastic.

Brian Elliott

The first thing we had to overcome was convincing Dave that we could be successful when we're not 19 years old. So he totally flipped his paradigm on these young people. I said, “Dave, when I was these young kids' ages, I was out across the ocean fighting at war.” I think I have enough experience to hopefully have a second career here in technology.

But, Peter, I love the question you're asking: How on earth do you compete with these frontier AI labs, with Google, with OpenAI? There are really 2 reactions you can have when there are these trillion-dollar investments. There's the reaction where you've built a company and say, “Oh, no. They're going to steamroll over me.”

Peter Diamandis

Yeah.

Brian Elliott

And then there's the reaction where, when every single model gets better and the combination of those models makes your product much better, you're jumping for joy. So we just got a trillion dollars of R&D for Blitzy, and we're—

Peter Diamandis

Riding a rising tide, and you can float on top of all of that. That's perfect. But it's critical to find that product-market fit that is able to benefit from the rise of these technologies.

Sid Pardeshi

I think we're seeing this happen for the second time. I was at NVIDIA back in 2016, and I heard all about the story of CUDA and saw Jensen believe in that. That was pre-generative AI; the term “GenAI” didn't exist, right? I've been working on generative AI models since the Attention Is All You Need paper came out.

Jensen was asked to stop investing in CUDA. He started this back in 2006, right, and it was a negative to the company. But he still invested in it. He still believed in AI. He worked with researchers and built the technology to solve problems that he foresaw, and that is very relatable to what we're doing.

So, Blitzy—we'll go into this in detail for sure—but it's very unique in terms of how the product is built. It is specific to the enterprise and based on the opportunity that we've seen over many years working at the largest companies, like NVIDIA.

Peter Diamandis

I think we should begin, for those who don't know Blitzy, by explaining what it is. I want to get into where you guys were when you said, “Aha, we're going to build this,” right? I love that story. Then I'll unleash Alexander Wissner-Gross on you to ask the most intelligent, important questions.

Brian Elliott

Well, let me tell you what it does today, and then I'll tell you the humble beginnings of all of this, Peter. Blitzy is an enterprise-grade autonomous software-development platform. We ingest and understand up to 100 million-plus lines of code, where most single-LLM tools are stuck with this finite ability to understand context. We've developed some really unique context-engineering systems to understand enterprise-scale codebases.

From there, an enterprise will express its work from a development perspective, whether that's a COBOL-to-Java upgrade—which is very common in these old financial-service institutions that use us—or steady-state development work. Blitzy will send off the most compute-intensive workload in the entire AI code-generation space. We've done a 12-hour run and multi-week runs for massive-scale codebases, ultimately delivering high-quality, prevalidated, precompiled, pretested code.

Our view, our thesis, is that we want to increase the quality of code at any cost, because the other side of a pull request that comes from AI code generation is human labor, which is exponentially more expensive. That's really the view that we have: enterprise-scale, high-quality code.

Peter Diamandis

All right, I want you to slow that down for a second. A lot of companies out there, a lot of traditional companies in particular in the finance world and insurance world, have large codebases. How old are some of the software systems that you're playing in?

Brian Elliott

I mean, we're talking about PL/I.

Peter Diamandis

Oh, my God.

Brian Elliott

COBOL. These are old-school financial-service institutions that, quite frankly, for a long time have been afraid to touch the code because the cost to get something modern just wasn't worth it.

Peter Diamandis

Okay, so you've got a company out there running COBOL from 20 or 30 years ago.

Brian Elliott

Yeah, that's right.

Peter Diamandis

And their system is operating. It's working. It's not doing anything significantly useful, given that it's 20 or 30 years old. Do they have people who can still patch that code? Are there engineers still around?

Brian Elliott

Oh, they make you and Dave look like spring chickens.

Peter Diamandis

We are spring chickens.

Brian Elliott

Oh, yeah. I know you're a longevity guy, Peter. You look great.

Peter Diamandis

So you've got this problem where you're too scared to touch it because you'd have to do a wholesale replacement. And so they call in the Blitzy guys, and you come in. What are you able to do for this 30-year-old chunk of code?

Brian Elliott

For starters, we give them visibility. We'll ingest, index, and understand the state of their underlying source code. They rarely have somebody who understands the entirety of tens of millions of lines of code. It's actually an impractical problem to know that.

Then we allow them to execute large-scale transformations, whether that's getting onto a modern technology stack or adding required functionality. These businesses, these enterprises, are stuck with the inability to layer artificial intelligence on top of their existing codebase because it is so old, so antiquated, and has so little visibility into what they're doing.

Peter Diamandis

Massive value, massive value creation.

Brian Elliott

It's a mind-blowing experience, too. If you take 10 million lines of unintelligible, undocumented code and run it through Blitzy, then say, “Tell me what it does in plain English. Explain to me where there are bugs,” you're talking to the code. It's mind-blowing.

Peter Diamandis

One of my favorite uses of GenAI is to give it a patent that is unintelligible and say, “What does this do, and how could I use it in my company?” The ability to take something that's complex and make it understandable—to grok it fully, to use that term.

Thank you, Robert Heinlein. Well, Sid's got 27 patents, so I had to use that to understand what the heck you did at NVIDIA.

So, Sid, you were at NVIDIA for 5 years?

Sid Pardeshi

8 years total.

Peter Diamandis

8 years.

Sid Pardeshi

Yeah, 6 years.

Peter Diamandis

I heard Jensen recently say that most of the executives there are now billionaires. Did you make it to that status?

Sid Pardeshi

Well, I held on to my stock. I rode the wave from double-digit billions to a trillion dollars. But then I got 2 degrees, Peter.

Peter Diamandis

And then what happened?

Sid Pardeshi

I got 2 degrees from Harvard, so they took all the money.

Peter Diamandis

No way. Wait a minute. Oh, my God. Let's not get started on the expense and value of degrees from Harvard or MIT, or any of these Ivy League schools.

This is a David-versus-Goliath story, and I want to understand it. You've got incredible success. But before we get to that story, you guys are both at Harvard Business School. When did this idea germinate? What was the positive agent that said, “Okay, let's build that”? What was that founding story like?

Brian Elliott

Yeah, if you can rewind the clock back to the GPT-3 to GPT-3.5 era, where these things could code, but it wasn't what we're experiencing today. At that time, Sid and I were doing a pro bono project for our favorite local bakery here in Boston as part of our time at Harvard. They mentioned they were about to spend $300,000 to $400,000 on a new mobile app, which got our attention as young enterprise entrepreneurs.

Sid and I went home and did what is now, 2.5 years later, called vibe coding, and we built them the application overnight—literally over the weekend.

Sid Pardeshi

Overnight? Yeah, during the week.

Brian Elliott

Which now is like no big deal, but if we—

Peter Diamandis

Did they pay you $200,000 for that?

Brian Elliott

We acted like it took longer. I think that was our first mistake. What was so clear during that time was that Sid and I were actually the bottleneck for development. You get an error and feed it back to the system, and then give it to a different model, right? Through that practice, you're able to get much higher-quality code.

And so we said, if we could just invent a system where all the commoditized development work could be removed, and we could have multiple models going back and forth, iteratively refining and getting to code that compiles and tests, that is going to be what the future looks like. We learned that by being hands-on, then having an idea of what the enterprise needs from Sid's experience and building toward that.

Peter Diamandis

Were you guys already friends, or did the all-nighter make you friends?

Brian Elliott

Yeah. We've only gotten—I mean, Sid's actually the godfather to my youngest son, believe it or not.

Dave Blundin

Okay. Investing in best friends is pretty good.

Peter Diamandis

That’s another part of the story, Dave, that we’ve talked about: some of the most successful companies are when best friends get together and build 24/7. Being an entrepreneur and having co-founders, you’re going to spend more time with your co-founders in the trenches than you do with your husband or wife or kids. It’s an intense period of time, and you better pick somebody or some buddies that you love spending time with.

Dave Blundin

Yeah. I always tell people, imagine you’re on a long international flight, one of those 14-hour flights, and you’re sitting right next to somebody. Do you walk off the plane feeling good and having fun, or do you walk off the plane not waiting to get away from this person?

Peter Diamandis

Your startup’s going to feel just like that every day. It can be work or it can be fun. It just depends on the personalities and the match. So, Dave, what do you find most exciting about Blitzy, just as an investor? They came through Link Studios and Link Investment, so tell us a little bit of that story, if you would.

Dave Blundin

Well, they’re definitely much more experienced than a lot of the entrepreneurs around the studio. They had the plan and the idea fully baked on arrival. We were still first money in, and we still gave them space and support and made a bunch of introductions, but they already had it more than figured out.

That’s not all the companies fit that profile. They’re also very different from a lot of the companies coming right out of dorm rooms. Those companies will do image generation, and they don’t understand the word COBOL or PL/I to save their lives. So when I look across the range of business plans that are right in front of us with AI, more of them fit into the “you need to understand the domain space” category than the “you could just think of it in a dorm room” category—maybe 2/3 to 1/3, rough numbers.

What I’m hoping with Brian and Sid is that they inspire a ton more people to go after these. These are still multi-trillion-dollar markets. But they’re not AI girlfriend apps. They’re not apartment search or another photo-sharing app. They get really deep. You’ve got all these areas like manufacturing and semiconductor manufacturing automation. That’s really deep. You get insurance actuarial risk adjustment. That’s very deep.

This one is actually nice in that it’s very broad. Code generation is very broad, so it’s a huge market, but it’s also deep. Refactoring 10-million-line codebases is a pretty deep knowledge set.

The other thing that’s really cool about Blitzy to me is that we have all these code-generation products. I use Cursor, and we’ve got Windsurf, Replit, and Lovable. Those companies are all worth billions now. I write a line of code, or I tell it to write a line of code, and it creates a button for me. I say, “I don’t like that button. Make this other widget,” and you’re doing it in real time.

Peter Diamandis

But you can’t build something really big. When you put Cursor in full agent mode, it’s right in no man’s land. It sits there and grinds for 5 or 7 minutes, which is too long to wait but too short to build something substantial.

Dave Blundin

So they’re getting stuck in no man’s land. Blitzy just said, “No, we’re going all the way to the other end, where it’s going to run all night long or all week long, like Brian was just saying, and come back with something really big.” That’s just fundamentally a much different engineering problem than what Lovable, Replit, or Windsurf are doing. It’s a different kind of company. I don’t know of any other company that’s there.

Peter Diamandis

Amazing. We’re here today to announce a particular piece of news as well—some groundbreaking news. Is it Brian or Sid? Which one of you wants to talk about what you guys have accomplished?

Brian Elliott

Sid, please. You are the inventor of the technology here.

Peter Diamandis

Sid, tell us.

Sid Pardeshi

Every time a new model comes out, they benchmark on this leaderboard, which is called SWE-bench Verified. The leaderboard itself was built by OpenAI. It’s a subset of SWE-bench. It contains 500 problems that were vetted by the researchers at OpenAI, and they confirmed that these are solvable problems and worth testing models on. It’s been ubiquitous. Every time a new model comes out, you always see results. The current top of the leaderboard on the SWE-bench website, as of filming, is 75.2%.

We hired a bunch of extremely talented interns. So, Dave, going back to Jack, if we could hire him today, we would. That’s how good some of these interns are. Every single intern who worked for us was amazing. We really credit this to their effort and to Nir, who led the efforts on our end.

We ran Blitzy on SWE-bench. As it turns out, these are 12 repositories, but they have 500 branches, and that equates to 400 million lines of code if you ingest them on Blitzy. We’ve ingested anywhere from 1 to 2 billion lines of code overall on Blitzy, depending on how you count it, because you count updates and the whole repo thing as well. So it’s 2 billion lines of code if you count all of the updates, including SWE-bench Verified.

We ingested all of that. We ran Blitzy to solve the problems, and our final result, accounting for everything that we’ve tested and verified using the SWE-bench CLI, was 86.8%. That is a significant jump over the current leaderboard on the website. The last time this was done was when Devin had a 13-percentage-point jump, from 1% to roughly 14%. We’ve come a really long way with the system.

The primary reason we were able to achieve this, echoing some of the points that we made earlier, is that we’re very different from the way existing tools are structured. One thing is that you can reproduce these results in production using Blitzy. We haven’t added any custom scaffolding just for SWE-bench, and we haven’t tampered with any of the features to achieve this.

We’ve seen reports from some of the other labs that claim that even though, for example, the latest frontier model claims 80% on SWE-bench, if you actually run it and reproduce it, you get 60%. We wanted to avoid that problem. We care deeply about reproducibility and practical, real-world applicability, which SWE-bench Verified has been vetted to be good at. You can reproduce these results, and it’s live as of today.

Alexander Wissner-Gross

First, to Brian and Sid, congratulations on your announcement. I would expect there’s going to be an enormous amount of interest from the community once they hear these results in trying Blitzy and reproducing those results. Congratulations in advance on the onslaught of interest that I expect you will receive.

I think software engineering is arguably the first major vertical of human labor that is very high economic value, very high productivity, and perhaps succumbing to automation. Any sort of step-function improvement in software engineering is arguably super transformative to the global economy.

Maybe just pivoting on that thought, one of the first things I was wondering when I heard that you would be announcing these results—and maybe jumping back 8 months—we all remember when DeepSeek, aka High-Flyer, launched R1 in January of this year, and there was sort of an aha moment all around the world. They didn’t just announce a new reasoning model. They announced—and maybe this got a lot less attention—a bunch of new open-source libraries at the systems level, like a new file system.

One of the first things I was wondering when I learned that you’d be making this announcement is this: the whole world is sitting on this sort of palimpsest of legacy libraries and operating-system code, billions and billions of lines of code—Linux, Python, GNU, all of these libraries. Is there something that you and Blitzy, with this new remarkable capability that you’re announcing, can do to speak to what we can do to improve the performance of this entire tech stack that the whole world runs on at this point?

Sid Pardeshi

That’s a fantastic question. We’ve been running some of these experiments. We’ve been taking some of these open-source libraries that, for example, were written in MATLAB and that one of our customers was using, and we converted them to Python. MATLAB was written 20 years ago. It was specific to Windows, and we made it OS-agnostic.

We’ve run these POCs all the time where we go from OS-agnostic to OS-specific to agnostic, from traditional to modern. But we’ve also been running other kinds of POCs where, for example, we picked an NVIDIA repo. We identified an issue that was marked as open, and we just put Blitzy at it and solved it. We created a pull request that solved the issue.

If you think about that problem, you can identify bugs, issues, and feature requests in any of the modern frameworks and systems. The system is not limited by how much code there is or how big the repo is, right? You can send Blitzy at it, and it’ll come back with a solution. If you don’t like it, you can iterate over it, create 5 projects, get 5 different pull requests, see how that works, and deploy it—all within a matter of days.

Peter Diamandis

I think that's a fundamental shift that really is going to change the way people work with open-source and closed-source technology. Amazing. What is the largest repository of code that you've actually tackled? Brian, do you want to say?

Brian Elliott

Yeah, go ahead. We frequently see 20 million lines, but I think the absolute largest we've seen is about 60 million lines that we onboarded successfully.

Peter Diamandis

Crazy. And how long, just for comparison, for fun, if you had to guess, how long would it take in terms of human labor hours to do that?

Speaker 2

It's insane.

Peter Diamandis

It's insane. Brian, I think we frequently scope these as part of the POC process, right? And I think you probably speak to that.

Brian Elliott

So, there's a question of how long it would take to grok 60 million lines of code. The reality is it's just too big for a human to understand. You might pay Accenture $100 million over 3 years, and they might come back with some diagrams covering the 60 million lines of code.

Peter Diamandis

By the time they did that, what they came back with would be out of date.

Brian Elliott

Exactly. Which is why you can see that, essentially, the industry has been stuck. This is why your airlines are always misrouted and they can't get their software updated; it's this kind of fundamental problem.

Every time we run Blitzy, we estimate for clients in production—all our enterprise clients—how many hours were automated. The CIOs love this because it's the KPI that they give the board on how many hours they've automated away through their intelligent vendor selection.

Grokking is sort of an impossible problem, but the real value is in code generation, right? Being able to accurately affect and develop code and accelerate that lifecycle for the development team on that large underlying corpus.

But importantly—and, Alex, I know you want all developers to go away and we're going to live in a society of abundance—I think it's going to go in the opposite direction, where there's almost an infinite demand for code and software development.

Blitzy doesn't do everything. Blitzy does about 80% of the quantum of work, on average, for these large-scale problems. But it knows exactly what it doesn't do, which is really the power. It hands off that batch of work to human developers to finish things out. So we get a really clean, full pull request plus human labor at the end to accelerate development, but not remove the need for developers altogether.

Alexander Wissner-Gross

So, just if I may, to pull in the thread, Brian, I would argue we're about to enter an age not necessarily of just abundance, but of great projects, when it's possible to send lots of automation loose on the world and fix all the problems, solve everything, as it were. In the case of Blitzy, this is letting AI agents loose on an enormous, sprawling legacy codebase and just fix everything. I think it's a good name for the episode: “Solve Everything.”

Peter Diamandis

Or a book, or, you know, fill in the blank. But are you familiar—maybe you're tracking this project? I love this idea: The Great Refactor. This is a project. Yeah, yeah, I love this. What is that, Alex? What's The Great Refactor?

Alexander Wissner-Gross

So, The Great Refactor—I love this. This is a classic solve-everything concept. We've built our whole civilization on a bunch of software libraries that could be better maintained and that are filled, in many cases, with memory vulnerabilities.

There are statistics out there that most of the insecurity of present-day software is due to the way the software is written, which exposes it to certain types of cybersecurity vulnerabilities—memory vulnerabilities. If we could only rewrite all of these libraries, dependencies, and software supply-chain upstreams that our whole civilization depends on in Rust or some other memory-safe language, suddenly that would fix almost all of it—that would solve everything—in terms of so many vulnerabilities.

Brian Elliott

I think I had 200 customers send me that project. They immediately saw that it became hot news that day, and they all sent it to me: “Oh, are you guys going to do this?” And I said, “Well, are you going to pay for it? Do you think it's worth it?”

Peter Diamandis

Here's the most critical thing, right? If you think about this idea where you can give these projects, refactors, or whatnot to AI and have it come back with the code, you can do that with any chatbot. You can give it to any AI and have it write code. Getting code from AI is a commodity, right?

But if you add constraints to that problem—where the code needs to replicate the existing functionality, it needs to compile, all the unit tests need to pass, it should not have newly added security vulnerabilities, and all the other items that are crucial to the enterprise or the problem itself—that's when the challenges begin, right? That is not something you can achieve.

So, Sid, my question is: I've got 3.2 billion lines of code, which is my genome. Can you recompile that for me? Can you go and identify and fix the broken parts?

Sid Pardeshi

LLMs can write as long as it's in a language that LLMs are trained to understand; we can do it. The scale is the problem that we've solved for.

The other problem we've solved for is making sure the requirements match, so that when we put you back together or edit you, you actually look like you, right? It's validated that it is you. We didn't change or break something that we shouldn't have.

Peter Diamandis

Alex, what do you think about this?

Alexander Wissner-Gross

Yeah. Maybe narrowly on the bio side, there are many other projects that speak the language of the genome and the proteome. I think, Peter, for rewriting your genome, you'll have the opportunity over the next few years to use one of these biological sequence-based foundation models to do some variant of that.

I do want, for Brian and Sid, to pull in the economics of this. I really want to press you guys: when we talk about The Great Refactor or some of these great projects to basically rewrite the source-code bases for much of our civilization today, and you think about the economics of that, there's a school of thought that says we're seeing generative AI hyperdeflate by 10× per year or so—an order-of-magnitude cost reduction every year.

At what point, in your minds, using Blitzy or maybe competitive tools, does it become reasonably economical to basically rewrite all of the legacy code out there that civilization depends on?

Brian Elliott

I would argue from a value perspective it's there today, because the value this would provide to society is just dramatic. This is a question of who's the payer. A line of code from Blitzy has a 100× higher cost basis than a line of code from any other provider, which is maybe 100× less than it would be from a human developer, right? And so we're talking about huge orders-of-magnitude differences.

Would it be worth it from a societal-value perspective to rewrite all the software today? Absolutely. But am I going to continue to serve these financial-services institutions and insurance companies first, which are readily paying me today? Yes.

If you grab the capital funding for us to break even on this, we'll start rewriting all the—we'll rewrite Linux if you need to.

Peter Diamandis

I want to insert another topic here. You guys shot a really cool podcast—it's on your LinkedIn—where you're bantering between the two of you about the fact that the definition of truth within large-scale software has always been the functional code.

Here it is: This is the final thing. It runs the PL/I code that does all the NAV accounting for the mutual funds over at State Street. It's millions of lines of legacy PL/I. But that debugged code is the core asset, and the documentation is just something around the edges.

Post-Blitzy, the truth moves to the documentation because you can regenerate the code overnight anyway. Your actual core asset has moved from code to a document, but it's going to move again. This is where it was really cool to hear you guys bantering around: What is then the foundational truth of this piece?

Because, as Alex is saying, the entire infrastructure of society is about to move and also expand 100×, 1,000×, or 1,000,000×, because code is so cheap to create. All of a sudden, we have much, much more of it, so you've got a much bigger world, but the ground truth is some other format than just PL/I, COBOL, or Python code. This human-readable spec becomes the central asset. That's a big shift.

Sid Pardeshi

The spec is still an abstraction layer, right? That's easy for a human to look at. The real source of truth or understanding is actually a customer-specific hybrid graph-vector database that understands exactly what is going on from a functionality perspective.

You could change that functionality from one language to another, but we are capturing the core essence of what is required there.

Alexander Wissner-Gross

And we can display that as a spec, which is 200 pages. But out of 20 million lines of code, that's an intermediate representation. Really, we want to get back to the core database-level understanding. That's the core asset for these folks, and Blitzy makes it the property of the enterprise.

Peter Diamandis

Everybody there's not a week that goes by when I don't get the strangest of compliments. Someone will stop me and say, "Peter, you've got such nice skin." Honestly, I never thought, especially at age 64, I'd be hearing anyone say that I have great skin. And honestly, I can't take any credit. I use an amazing product called One Skin OS01 twice a day, every day. The company was built by four brilliant PhD women who have identified a 10 amino acid peptide that effectively reverses the age of your skin. I love it and like I say, I use it every day, twice a day. There you have it. That's my secret. You go to onskin.co co and write Peter at checkout for a discount on the same product I use. Okay, now back to the episode.

Brian, given your background in the military, probably the one institution that has the largest repository of ancient code has got to be the U.S. government, right? Can you attack all of that and unlock massive productivity? There's a lot of fear that the U.S. is a falling empire—its inability to understand and legislate efficiently. Couldn't you have a single massive impact on the U.S. government?

Brian Elliott

Peter, I live in Back Bay here in Boston, and so does this lead investor at In-Q-Tel—or so he tells me—because he keeps running into me when I'm walking to work, bumping into me and seeing how things are going, which has me skeptical. But the short answer is yes.

The U.S. government is a fantastic end customer to ultimately modernize: to get your flights there on time, to make getting IDs easier, right? This is absolutely part of critical infrastructure and a target customer that we have, certainly. Do we have them today? No. In 12 months, will we be serving them? I think yes.

Peter Diamandis

That could catapult you into a deca-billion-dollar company easily, just landing that kind of a customer, because once you've modernized one of the agencies, they're all going to want it.

Brian Elliott

I think we have the most top-secret-security-clearance-to-patent ratios of any company out there.

Peter Diamandis

Fascinating. I'd like to pull in the theme that we've talked about on the pod in the past: the elephant in the room, which is recursive self-improvement. So how much of Blitzy is written by Blitzy?

Brian Elliott

A lot. This is actually a question of where the core value is. I would say every single sprint that we do from a software-development perspective is driven by Blitzy. A significant amount of the corpus of the code is driven by us.

Now, there are algorithms that are not about writing code. They're about core invention, and I think most companies' core IP is not going to be the software that it creates; it's going to be some core invention. Think about Google with PageRank. PageRank, not actual pages, is really the core source of its original IP.

Sid invented a number of algorithms at the core of Blitzy that allow us, for instance, always to compile code and never have circular dependencies. So you could actually rewrite a corpus of Blitzy with Blitzy pretty readily and pretty quickly, and it would look like it. But would it do what we do? The answer is sort of no.

It gets to the question: What is the source of IP for companies? It's got to be a breakthrough or an invention.

Sid Pardeshi

What we're doing for these large companies is telling them how to use AI. We're coaching and consulting with them on how to use Blitzy. You cannot do that unless you've actually used it yourself and perfected the process, because they're not just making the engineering-velocity changes or the tool-usage changes; they're also making the process changes. That's why this is critical.

Peter Diamandis

Alex, I'd love you to take a second and dive into the SWE-bench metric here. Again, for those who aren't familiar, a little bit of the origin—you said from Princeton and ratified by OpenAI—but what is it measuring? Who were at the top of the leaderboards before? Then I want to get into the conversation about how you compete against the Mag 7 or against the frontier models in this regard. Let's contextualize it first by understanding what the SWE-bench metric really is.

Alexander Wissner-Gross

Sure. So, for context—

Peter Diamandis

Why don't we start with the title of the white paper? That's going to drop the same day that this podcast does, so people need to find the paper. We should put a link to the white paper in the show notes here as well. Alex, please.

Alexander Wissner-Gross

Sure. Maybe let me just speak to SWE-bench. SWE-bench is a benchmark that measures the ability of AI systems to solve typical software-engineering tasks, in the specific form of responding to and solving issues on GitHub, a very popular source-code-management system.

SWE-bench as a whole—not SWE-bench Verified—consists of a couple thousand instances of tasks in which the central challenge posed to an AI is to respond to an issue in a codebase. What would happen in a normal software-engineering context is—

Peter Diamandis

What kind of issue, Alex?

Alexander Wissner-Gross

A wide range of issues. It could be bugs that need to be fixed, other performance issues, and the usual workflow in a software-engineering context is that an issue will be identified and a pull request will be submitted.

Identifying an issue, responding to an issue, submitting a pull request that responds to an issue, and satisfying unit tests—these are all standard parts of what would be archetypically considered software engineering. SWE-bench is an excellent industry standard at the moment that attempts to capture the life cycle of high-value-added labor that a typical software engineer would perform. I'll let the guys respond to their white paper.

Brian Elliott

Sure. I'll also respond by answering your core question, Peter, which is: How does one compete with the Mag 7 in this utterly important labor task?

The reality is that a significant amount of the Mag 7 is at the core of what Blitzy does. We use Google's Gemini models, we use Anthropic's models, and we use OpenAI's models. What's unique about this technology moment in time is that when you use these models against one another, the quality moves up quite dramatically.

When you use them against one another hundreds of times, in hundreds of different combinations, with hundreds of different tool sets and prompts, the quality goes up even more exponentially. Really, it's the art of orchestration through what we call extended-inference-time validation to improve the quality of code, ensuring at every moment in time that the system has the right context to operate despite the large-scale underlying codebase.

We say we're excited when Gemini or OpenAI release a new model, because our product gets dramatically better every single time. That's a really important insight: building so that the better your components are, the stronger you are as a whole.

Peter Diamandis

That's a unique niche. When did you realize that? I mean, that's sort of a fundamental for you.

Brian Elliott

I think we realized it back when we were initially figuring out this first project together, but I'll let Sid expand on this.

Sid Pardeshi

I think we made a bet. We said there's not a doubt—we were building this when models had 5,000 tokens of context—and we made a bet that there was no doubt that context windows were going to expand and the models were going to get better at writing code.

Do we want to compete with the Mag 7 and build our own model, or do we want to stand on the shoulders of giants and use the technology to solve the problems that they're actually meant to solve? That's exactly what we did.

Peter Diamandis

Amazing. Just for context again, who had the record before you? Who had the record last week? I think there were some open-source LLMs, and there have also been some other unpublished reports that have claimed around 80%, but the 86.8% that we're claiming is unprecedented and the highest number we have seen.

Brian Elliott

I think ByteDance's Trae model was the most recent to be at the top there. We can rebrand this U.S. versus China if you want to.

Peter Diamandis

Every time I hear 80-something percent or 90-something percent, I'm thinking that these benchmarks are getting super-saturated. Where does this go next? What are you going to measure when you're at 100%?

Alexander Wissner-Gross

We talk about this in the white paper. We certainly need new benchmarks, but the reason we didn't do SWE-bench Verified for a long time is that it's just not representative of the scale of problems that most people use Blitzy for.

The typical pull-request size is around 100 lines of code, and the largest repository in this benchmark is 1 million lines. We really need a set of benchmarks based on Linux and VS Code, which have 20 million and 4 million lines, respectively, with holdout sets against those, trying to do larger-scale work to ultimately show how far we can push the bounds on autonomy.

Yeah. And in terms of Peter’s saturation question, the paper does a really good job of describing the landscape of benchmarks and SWE-bench Verified in particular, as well as the need for a new benchmark. But one of the points it makes is that when you score 86% on this benchmark, that’s effectively very close to 100%, because the remaining subset of questions are flawed. They’re not harder; they’re just not structured well.

Dave Blundin

And so you’ve basically capped out this benchmark now.

Peter Diamandis

So, if folks want to look up the paper, what’s the name of the paper?

Dave Blundin

Do you have it?

Brian Elliott

Yeah, Alex is urging us to retitle it, so I’ll give you the subtitle: “Domain-Specific Context Engineering Paired with Extended Inference-Time Validation Breaks the Barriers of LLM-Driven Software Development.” That’s really what we’re talking about.

Peter Diamandis

That rolls off the tongue and onto the floor. It is a technical title.

Brian Elliott

Yes.

Peter Diamandis

You know, Sid, you can search “Blitzy” in Sid’s name. He has a searchable name. Brian Elliott—there are thousands of Brian Elliotts, it turns out.

Sid Pardeshi

But “Sid Pardeshi” plus “Blitzy” will get you to the paper.

Peter Diamandis

Okay, perfect. Alex, where do you want to take this next? What have you found important and fascinating about what Blitzy is doing? What are the implications in the long term, buddy?

Alexander Wissner-Gross

It’s so interesting. There are so many different directions to go in. Maybe just to go back to this idea of great projects, because I think Blitzy has the potential to be an embodiment of an era when we just turn all the AI agents loose on all of the problems in a discipline.

What I heard—I think, Brian, you said a few moments ago—was something like a 100× price difference per line of code or per file between Blitzy, with its state-of-the-art performance announcement, and other competing tools. When I hear you say a 100× price difference, I immediately say internally, “Well, that’s just 2 years’ worth of cost hyperdeflation.” So, you’re 2 years more expensive than the competition, is the way that I heard that.

If you project forward 2 years, 3 years, 4 years, do you think we will find ourselves in a world where AI really has completed the great refactor and we’ve rewritten all of our foundational systems with Blitzy, or maybe copycats of Blitzy? Do you think we’ll find ourselves in a near future like that?

Brian Elliott

Let’s pontificate first here. I think we will see the models get significantly better at doing this and the cost go down. The key thing that I would like to underscore is that a lot of the approach at some of the labs, and what’s happening right now, is that the labs aren’t really making money on the inference that they’re running for all the models. But what we’re doing is, because we’re using the labs and we’re able to charge a premium for the work that it does, while also providing the validations with it, we’re not losing money on the code that we’re writing.

As this equation improves over a period of time, the difference that Blitzy is able to create is going to also grow. Somewhere in that double negative, I heard the answer is yes: as hyperdeflation kicks in—call it an order-of-magnitude cost reduction per year, maybe more—not only does Blitzy become very profitable, but it also becomes very feasible to start tackling these solve-everything-level grand challenges in software engineering.

Peter Diamandis

I’ve been attempting to kick the tires on Blitzy myself. My first project with Blitzy was that I wanted to rewrite Python, the very popular programming language. I gather, Brian—sorry, Alex—I gather you, Brian, and Sid have had access to your own product longer than I have, which has only been 2 days. Have you tried to take on some large-scale project? I think Brian mentioned Linux a few minutes ago. Have you tried to take some large project and either say, “I want to ask Blitzy to improve performance by 10% on some relatively mature codebase,” or add some crazy, transformative new feature? Have you tried that?

Brian Elliott

Yeah, we did this project where we’ve done a number of these, Alex, but one of the most fun things we did was onboard VS Code and say, “Hey, add a chat experience to VS Code.” At the time, Cursor was—and still is—one of the biggest tools out there. We built a subset of Cursor’s features and tried using that internally.

Any time we consider SaaS products at this point, we’re trying to first see if we can replicate them internally using Blitzy. If we’re a few months out or a few years out, we’ll say, “Let’s just use the SaaS products as a starting point and consider rebuilding them later on.” I think a lot of enterprises will do this.

Alexander Wissner-Gross

Have you thought—I mean, sort of free marketing advice before the public—about taking all of these open-source projects that are, in many cases, starved of core development team members and hungry for human capital or human-capital equivalents? Could you take these projects and aggressively set loose the AI agents to submit very friendly, very polished pull requests to these projects to launch improvements?

We’ve actually done that for MLflow, I believe, right, Brian?

Brian Elliott

Yeah. We’ve done this, especially for one that enterprises specifically rely on. It’s quite possibly the best BDR, which is just sending pull requests to open-source libraries.

There’s an open-source one right on the homepage of our website that I think you’ll find fascinating, Alex. It goes all the way back to what we were talking about at the beginning of the episode. AWS created this repository specifically to be incredibly messy mainframe code, with all different styles to represent different decades of people working on it, and ultimately to ask how one would use code generation to move this from COBOL to Java.

Alexander Wissner-Gross

Oh, cool.

Brian Elliott

We ran that through Blitzy because mainframe is such a big problem in all these large organizations, even government organizations, and we moved it from COBOL to Java completely autonomously, with the ability to compile right out of the box.

There’s still some remaining development to work on with the runtime, but this is a multi-year-long project to move a messy mainframe codebase from COBOL into Java. The results of that have probably been one of the best business-development tools that we’ve ever created.

Alexander Wissner-Gross

How fast did you do it?

Brian Elliott

How long did that take? A couple of days. If you count everything from start to end, it was a week’s worth of inference.

Dave Blundin

Essentially, yeah.

Alexander Wissner-Gross

Amazing. So, if we project that forward, maybe I may go back to Peter’s question about the human equivalent of this. Do you have any metrics that you can point to, either for cost savings relative to humans for a given unit of code—a line of code or per file—or how much faster than humans this is, in general?

Brian Elliott

We typically see a 5× velocity difference when this is brought into the enterprise. The biggest challenge is really operational deployment.

You’re used to starting your work the same day that you pick up your IDE. What we’re having these organizations do is start the sprint for the next development work the week prior. Instead of developers starting with tickets and tasks, they’re starting with code that’s mostly written and a project guide with all of the human tasks to begin.

We really focus on this 5×. Any time we engage an enterprise to work with us, we say, “Pick a real-world project that you have coming up next quarter. Let us prove a 5× difference, and we’ll do the remaining development work so we have an end-to-end solution. If we do that, then you adopt across the organization.”

It’s incredibly successful because people don’t realize the cost of coordination and development, and all of the requirements to actually get a piece of full software out. When you can offload a significant chunk of that to agents, the velocity gains are honestly unbelievable.

Alexander Wissner-Gross

That’s a really cool insight, because a lot of the younger teams are going into enterprises and they’ve never been in enterprises before. They’re saying, “Look, the raw code generation is 1,000× or 10,000×.”

You’re like, “Yeah, but what’s it going to do for me in my enterprise?” There are very few that are credible in saying, “We’ve actually done it, and we know the final answer.” Right now, it’s apparently 5×, but they don’t know the overhead of the organization, the documentation, and just getting all of the things that happen before you can even start code generation.

It’s really nice to have at least one vendor that understands how to get the real thing. We actually need this to work in the end. It can’t just be a hypothetical 1,000×.

It’s 3 years from now. Digital superintelligence has landed. It’s come out of—you know, I’m going to put my bets on Google, but we’ll see. What does Blitzy look like?

Brian Elliott

Yeah, I think Blitzy is the core system of record and system of action for software development in the enterprise environment. The source of truth moves from documentation and code, which is this hybrid today, to organizations relying on Blitzy’s hybrid graph-vector database that understands core functionality.

Organizations are going to be able to move incredibly quickly from a software perspective, and the source of value is going to be core IP that is not easy to replicate.

Peter Diamandis

Mhm.

Brian Elliott

To add more onto that, Peter, there’s always going to be some tasks where it’s better to have a human in the loop and do them sequentially, right? You’re solving a problem that has never been solved before, and you need quick feedback from AI, right? That’s always going to be there, where you use the copilots, but there’s always going to be this other category of tasks where you can automate them away, right?

Build the code, run it, deploy to production, and execute maintenance. Blitzy now gives you the code. The code is the final output, but we’re going to go into autonomously maintaining, deploying, and keeping the applications running. So you’re not going to need humans for specific sections of the entire enterprise. It’s all going to be driven by AI.

Peter Diamandis

Interesting. If you could, I think Alex was about to paint a kind of a 2-year view. He asked a question about what’s the force multiplier today, but then I think we were going to next segue into, okay, but there’s 100× and another 100× coming. So I would love to finish that thought, Alex.

Alexander Wissner-Gross

Totally. There’s a lot of interest in the benchmark out of METR that’s measuring the effective time of autonomy—the characteristic time scale over which AI systems, including AI code-generation systems, can basically operate without human intervention, sort of like disengagement with a driverless car. How far can it drive without a human needing to take the wheel, as it were?

So I’m curious: Have you thought about the characteristic time scale over which Blitzy is able to do autonomous code generation, or the human equivalent, really, of autonomous coding, before which the human needs to step back into the loop and be involved? Right now, if I remember correctly, the current state of the art is something like 1 to 3 hours.

There’s a nice, very clean, at least on a semilog plot, expectation—and I think we’ve discussed this previously—that if you project it out a decade or 2, we get to many, many years and perhaps hundreds of millions of years in a few decades. Where does Blitzy fall in this apparent exponential trend toward exponentially increasing times without humans needing to be in the loop?

Speaker 1

I think this is a project.

Sid Pardeshi

If you think of all the pieces needed to achieve this, let’s take a very small example. Let’s take the AWS example, right? There’s the part where you identify the requirements, decide what you need to do, get the code, and there’s the part where you get it all the way to production.

If you look at those parts, each of them has already been automated in isolation. For example, CI/CD: How do you deploy the code to production? You have automations for that. Debugging, security analysis, monitoring in production, tracing and viewing the logs, ensuring that the system is not doing anything malicious—all of these items exist today.

There are these glue layers that we’re now seeing, like, for example, MCP and A2A, that allow agents to form this mesh and automate work. The only thing that’s left to be done, really, in my opinion, for these projects is to just connect the dots, and that’s exactly what we’re working on.

So, to answer your question, how far out? I would say we’re months out, actually, from delivering projects completely autonomously, as long as they meet a certain set of criteria and conditions.

Peter Diamandis

So what I just heard you say, Sid—correct me if I’m wrong—is that the documentation writers, the spec writers, are the new limiting factor for the speed of software development. Is that correct?

Sid Pardeshi

That is correct.

Peter Diamandis

Great insight. And how do you think about automating that process, if at all?

Sid Pardeshi

That is also automated. So if you go to ChatGPT and ask it to write documentation, it will do so following your criteria. The reason we have documentation writers in the first place is to have quality, right?

We have concerns that models lose context over a period of time, and they skip and omit things, or they can be gamed into adding things that you don’t want. We’re really concerned about quality control, which is why we have humans. But as you can see, we’ve solved the context problem for large codebases, and there’s nothing really stopping anyone, for that matter, from effectively adding the right safeguards and layers of protection to ensure that we minimize the need for humans.

I think it’s a matter of us becoming comfortable with AI doing that, and I definitely see that happening over the coming months.

Peter Diamandis

I’m begging for a follow-up white paper, because Alex’s question is infinitely recursive, right? If you said, “Okay, well, then that’s not the constraint,” then what? Because there’s always going to be a constraint. It’s turtles all the way down. You’ve got to just ask, okay—

Dave Blundin

Speed of light, buddy.

Alexander Wissner-Gross

Yeah, Douglas Adams, right? He famously pointed to that. The problem is far harder to pose than the solution.

To the extent that the new limiting factor is the spec writer or the prompt engineer—whatever we end up calling it in the future—I really would like to press you, Sid, on this: When do we get our automated program manager, product manager, spec designer, and documentation writer, if that really is the limiting factor for the speed of software engineering in the near future, all the way down?

Sid Pardeshi

I’ll tell you something, Alex. You know how the Blitzy platform works. We have these thousands of agents, and each of these agents has a persona. There is a product manager agent, a software architect agent, a QA agent, and an agent that writes the prompts for the other agents.

Brian Elliott

So all of the challenges that you’re describing are live today in production with the platform.

Peter Diamandis

This is such an important question, though, because I know it sounds very hypothetical, but look at this timeline on SWE-bench here. This was only 18 months ago that you got, like, 12%.

Alexander Wissner-Gross

And now it’s saturated. It’s only been 18 months. So what we think of as the distant scientific future—it’s all science fiction. It’s only a year, a year and a half in the future.

Dave Blundin

Star Trek’s coming, buddy.

Alexander Wissner-Gross

It’s really, really hard to anticipate. Yeah.

Peter Diamandis

Now, I love this question. This is another white paper waiting to happen. I want to wrap this episode with a conversation amongst all of us on a particular topic.

We opened up talking about trillion-dollar pay packages, trillion-dollar investments—numbers that are extraordinary—and the sovereign funds, the venture funds, and family offices are just supporting this with massive capital inflows. So the question that I put to all 4 of you: Think about competing in the long term with the Magnificent 7, who’ve got this incredible access to capital. How should founders consider going about that?

What’s your advice to others who are getting in here during this period of exponential growth in the AI economy? How do you compete? How do you think about that?

Brian Elliott

I think you want to be a large customer of those folks as well. I mean, we are major customers of all the AI frontier labs, and so they’re quite excited that we’re going to continue to push the bounds of autonomy. Their market caps are going to continue to grow, probably dramatically, in line with that return on investment, and Blitzy’s going to ride those waves as well.

If you’re happy when they’re successful and they’re happy when you’re successful, then I think you’re in a pretty good strategic position.

But there’s one more thing to add onto that. I’d like to go back to what Dave said, right? Mercor, for example, was able to do that because it went deep. I think that’s also the case for us, right? We’ve seen the enterprise—Ben and I—from the enterprise side of the challenge and the enterprise perspective, and the security roadblocks, the product roadblocks, and the process gaps that stop them from taking full advantage of the product.

So if you’re an entrepreneur or a founder and you’ve seen this personally, and you’ve struggled with this problem, and you think you have a solution that addresses the core of that, and you’ve been able to test that with the actual enterprise and demonstrate effectiveness, I think you’re holding on to something that is core, right?

Peter Diamandis

So you’re saying—

Brian Elliott

Understand the problem deeply.

Peter Diamandis

Yes.

Brian Elliott

Understand the problem, because look, you have the Magnificent 7. They’re giants. They have all the money. That’s fine, right? You’re an entrepreneur. You’re nimble. You can find the right investors.

We were grateful to find Dave, who believed in us the moment we pitched it, and we were able to get just the right amount of capital to get started. That’s really all you need.

If you have the right talent, the right amount of capital, and you have the right problem that you’re going after—one that you’re convinced about because you’ve experienced and solved it—then you’re going to be so nimble and make these moves and get a product out that is significantly better than anything that the Magnificent 7 can put together, because they’re struggling with their own challenges, like bureaucracy and all of the hurdles that they have to go through to actually put out—

Dave Blundin

Politics, struggling with what to say over dinner with Donald Trump.

Brian Elliott

Exactly. So while they’re distracted with all of that, you can build a kick-ass product, get it to market, solve real-world problems, and you’ve changed the world effectively.

Peter Diamandis

Alex, what’s your thought? How do founding entrepreneurs compete with companies? I remember famously, Amazon was out there as a platform for people to sell their products, but then when Amazon saw a product that had incredibly high margin and growth, they would clone the product and compete directly. How do you keep from that happening?

Alexander Wissner-Gross

Two words: solve everything.

Peter Diamandis

The name of this episode: “Solve Everything.”

Alexander Wissner-Gross

The world is filled with so many problems that a startup standing on the shoulders of the trillions of dollars of CAPEX being invested in cloud, AI, chips, fabs, and energy is now poised to solve so many problems—thousands of problems. I think Brian and Sid—and again, congratulations on the benchmark announcement—are well poised, potentially, to solve the problem we face of decades of civilizational software cruft and legacy code that’s just piled up without enough human capital to invest in reinventing it.

Now I think we’re arguably on the verge of doing that. That’s 1 of thousands of problems, entire domains that can be solved. Protein folding was solved by AlphaFold essentially overnight, transforming a subset of structural biology. So many more opportunities.

Peter Diamandis

Before I go to you, Dave, I just want to remind people: I define an entrepreneur as someone who finds a juicy problem and solves a juicy problem. The more entrepreneurs in the world, the more problems that get solved, the better the world is. It’s why we’re going to hit on this over and over again.

I think the career of the future is being an entrepreneur: finding a problem, falling in love with the problem—not the solution, not the tech. If you understand the problem deeply, as the tech evolves and continues, you’re going to use the newest version to go and solve that problem.

Again, some of my favorite lines: the best way to become a billionaire is to help a billion people, and the world’s biggest problems are the world’s biggest business opportunities. So that’s what entrepreneurship means.

Dave, you see hundreds and thousands of companies. How many companies do you have right now in the Link Studios?

Dave Blundin

Yeah, 28 in the building and about 50 total.

Peter Diamandis

Amazing. What do you look for when you’re looking to invest in a young entrepreneurial team like Brian and Sid, or like the founders of Mercor, or some of the incredible unicorns that we’ve backed out of Link Experimental Ventures? What are you looking for to make sure that company isn’t going to get disrupted in the wake of an OpenAI or Google slight jog to the right?

Dave Blundin

You know, it’s funny. Kevin Weil—we asked that exact question on the podcast we did 2 weeks ago. He answered exactly the way I had hoped he would answer, which is: in a world where the foundation-model companies get to AGI and can do virtually anything, are they just going to take over the world?

Kevin was really clear that maybe they can do that, maybe they can’t. We probably can’t anyway. But even if we could, we don’t want antitrust to come in here and break us up. That’s the last thing we want. We want a huge, thriving ecosystem of partners that give us money.

Is Blitzy 1 of those companies that gives us money? Yes. Therefore, they’re our best friend. Go conquer the world. Take over. Change the entire foundation of all legacy codebases. Make $1 trillion and give us half of it. We’ll all be happy. That’s what they want.

Peter Diamandis

I took a 2-hour walk yesterday with a dear friend of mine here who runs a large venture fund, and we were talking about the notion that his bet was that Google had so much more capability than they unleashed. They said, “Look, it’s OpenAI. Go and do as much of this as you can, because we need someone out there competing with us. Otherwise, we’ll get broken up for antitrust reasons.”

It’s a fascinating idea. You need viable competition to help you price, to help you remain on the edge, and to help you not be, you know, sort of broken down by the government.

Dave Blundin

And be a good partner. That was when Google was growing like crazy and we had all these portfolio companies. We made a ton of gains. But be a good partner to Google while they’re growing like crazy, and now it’s the foundation model. Just be a good partner. Talk to them all the time. Make sure you know where they’re going, and they’ll love you.

Peter Diamandis

Amazing. I have a selfish question. I don’t know if we’re running out of time here.

Dave Blundin

No, that’s fine. Let’s close with your selfish question.

Peter Diamandis

Okay. Well, this—I’m always looking for traits like this. This has obviously been 1 of our best investments ever, and the sky’s the limit from here. I’m always looking for traits of success, and the morale at Blitzy is like nothing I’ve ever seen, which is not a no-brainer when you’re doing video generation for a movie studio or whatever. It’s easy to keep high morale, but when you’re doing 5 million lines of core COBOL conversion, yet you guys have this crazy, thriving culture.

Sid mentioned that we were first money in. I don’t remember why we loved the deal so much. I do remember it absolutely was a no-brainer to invest in you guys.

Two things jump out at me. 1 of them is BITS, which is just the hardest place in the world to get into, and then there’s Blitzy, where you’ve seen growth. The other 1 is Brian. I think you had Army Ranger training and Bangalore Institute of Technology.

Brian Elliott

That’s BITS, Birla Institute of Technology and Science.

Peter Diamandis

Okay. Yeah, he has a cool name, though. BITS is like MIT, but it’s BITS, you know.

Brian Elliott

And, by the way, MIT designed the curriculum for BITS, so that statement was actually true.

Peter Diamandis

Oh, that’s cool. The other 1, though, was Brian. I think you had not just West Point, but Ranger training, which is freakishly hard, and then first boots on the ground in Syria—literally the first people touching a war zone.

I’ve got to feel like there’s something in those experiences that puts you a cut above in terms of building a team, managing logistics, and building morale. Any clues there that other founders can pick up on?

Brian Elliott

Yeah. I think some evidence of being incredibly mission-driven and ambitious is what you would see if you were an anthropologist looking at both of our backgrounds.

But if you take me, for instance, and if you fast-forward—or, I guess, rewind—to 2017, when I was serving in the 75th Ranger Regiment, the mandate was, “Hey, go into Syria. There are about 2,000 ISIS fighters in the whole of Raqqa. We’re going to send you with 100 guys. Recruit everybody else and take back the city, right? And, oh, by the way, we can’t let anybody in the United States know we’re here because we’re there covertly, right?”

To be able to go in and solve that problem, that’s a very ambitious undertaking. Conquering cities isn’t something that most people have spent their time doing. So when you look at the level of ambition of the company, everybody here at the business has that ethos. The very first thing we do when we interview is screen for ambition and the ability to invent and create.

Dave Blundin

We have those core values. If you talk to any single person who sits in this building right here, they will tell you—and they’re right—that what we’re doing is 1 of the most important things they will do in their lifetime.

The economic expansion that the globe gets, the GDP expansion that you get from automating software development, or at least huge chunks of software development, means there’s almost no better incremental use of energy than driving toward that goal.

Peter Diamandis

Wow, that’s a beautiful thought. Are you guys a 996 or 997 shop?

Brian Elliott

It’s Saturday today. We’ll be here tomorrow. We’re a 997 kind of crew.

Peter Diamandis

Oh, my God.

Sid Pardeshi

Just to back Dave up on his 997 joke.

Brian Elliott

That was a mistake. Well, I hope ByteDance was the last leader on SWE-bench Verified, so we can work longer than them and beat them on the leaderboard.

Peter Diamandis

Oh, my God, guys. Listen, congratulations on hitting that new benchmark. More importantly, thank you for the work that you’re doing—for the companies that will benefit, our government that will benefit, and the world that will benefit. You’re upgrading the DNA of industries and of our planet. So grateful for you, Alex, and Dave. Any closing thoughts here?

Dave Blundin

I’m super excited to see what you guys can bring to the future. Very few things would excite me more on the software-engineering front than, a few years from now, learning that the entire software stack that I run, and that companies I work with run, has been 99% rewritten by Blitzy’s agents to remove all the vulnerabilities and improve all the performance.

I think it’s the sort of challenge before you guys that sets us on the road to recursive self-improvement and abundance, and also solving everything in software.

Peter Diamandis

Solving everything. That’s my phrase for the day. Let’s solve everything.

That’s a good catchphrase. It’s better than Dave’s 997 or tang ping. Tang ping is the opposite of 996, lying flat in response to overwork.

Let me give my closing thought. Definitely, everybody read the white paper. The title may sound very complex, but the paper itself is very readable, so please read it. Your takeaway will be, “Wow, okay, now we need a new benchmark.” Inside baseball: Blitzy is already working with MIT to create the next-generation benchmark.

So, catch up to what they did right here by reading the paper.

To all our subscribers thank you for following Moonshots and WTF episodes. We're grateful for your time. We hope that, in spending the time with us, you're able to understand how incredibly powerful this technology is for transforming our world, our lives, creating a future of abundance. I hope this counters all the dystopian news you get on the 6 and 7:00 news. That stuff I don't watch. This is the stuff I focus on. I hope you do, too. I'm grateful to my Moonshot partners, Alexander Wissner-Gross, Dave Blundin, and Sid, wherever you are transiting the Atlantic to come back here to the US. And again, Brian and Sid, congratulations on your epic wins. Excited for your future success.

Every week, my team and I study the top 10 technology meta trends that will transform industries over the decade ahead. I cover trends ranging from humanoid robotics, AGI, and quantum computing to transport, energy, longevity, and more. There's no fluff, only the most important stuff that matters, that impacts our lives, our companies, and our careers. If you want me to share these meta trends with you, I write a newsletter twice a week, sending it out as a short two-minute read via email. And if you want to discover the most important meta trends 10 years before anyone else, this report's for you. Readers include founders and CEOs from the world's most disruptive companies and entrepreneurs building the world's most disruptive tech. It's not for you. If you don't want to be informed about what's coming, why it matters, and how you can benefit from it. To subscribe for free, go to dmmandis.com/metatrends to gain access to the trends 10 years before anyone else. All right, now back to this episode. [Music]

AI Now:Elon 的1万亿美元激励、Apple 为 Trump 承诺的6000亿美元,以及小型创业公司如何胜出——与 Dave、AWG 和 Blitzy 对谈 — 文字稿与摘要 | BidClub