两名哈佛辍学生融资8亿美元,向 NVIDIA 发起挑战
- 开场几秒就把核心论点摊牌:「推理将成为全球最大的市场。谁产出最多 token,谁就会成为全球最具生存力的公司。」 Etched 创始人认为,整个半导体栈都「建立在缓冲之上」——从 EDA 默认芯片运行在冰点温度,到其他面向通用场景的假设,纯推理公司可以逐层剥离,再把「这里20%、那里50%、这里2倍」的优化叠加起来,最终打造出他们声称不是强10%,而是「强10倍」的系统,超越现有 AI 芯片。
- 这套设计完全押注于2个技术方向。 Prefill 阶段采用低电压推理——利用 Dennard scaling(功耗与电压的平方成正比),以「低于任何其他 AI 芯片一半的电压」运行,在增加 FLOPS 之前先解决热降频;decode 阶段采用集群级内存——在二层以太网之上完全定制互联,将芯片间延迟相对 Blackwell 约4,000纳秒的跳转降低「超过5倍」,让整个 scale-up 集群的 HBM 和 SRAM 像一个内存池一样工作。
- 护城河不仅来自架构,也来自极限垂直整合:「最好的供应商就是没有供应商」。 Etched 自建机架、在台湾设厂,并通过「预取」推进交付:700块 FPGA 跑完整芯片,机架在没有芯片的情况下提前运往客户数据中心,用热仿真芯片验证冷板。结果是从芯片回片到机架内完成推理只需40天;一家「非常知名的 AI 芯片公司」公开披露的周期则是10个月。
- 真正险些送命的是资金:2024年初,公司账上只有1500万美元,但未来12个月需要1亿美元,「硅谷每一家主要投资机构都立刻拒绝了」。 他们最终拼出一轮1.03亿美元、以软承诺为主的 A 轮融资,债务提供方、以多年期条款提供 Synopsys 仿真器的供应商,以及 TSMC 都伸出援手——当时22岁的 Gavin 在一场会议晚宴上与一名 TSMC 高级副总裁讨论张量数学,赢得了后者的支持。早期投资人 Patrick 打破第四面墙:「你多少得诅咒一下基准概率。」
- Etched 相对超大规模厂商自研芯片的结构性优势,在于生死攸关的专注度:「Google 不会因为 TPU 失败而倒闭……OpenAI 不会因为 Jalapeño 失败而倒闭」。 Rob 还称,实验室和超大规模厂商的芯片,其「FB8 × FB8」FLOP 密度低于「Blackwood B300」(听辨如此),因为「他们没必要承担这个风险」。供应方面,Etched 认为双方是增量关系:采用4纳米工艺和不同于 Rubens 3纳米方案的 HBM,因此客户谈的是「2吉瓦」,而不是二选一。
- 这些未来判断既激进,也有明确时间表。 Rob 预测,2027年将是知识工作中 agent 数量超过人类的一年,生产率将以「每兆瓦多少个 agent」衡量,单个价值1万亿美元的数据中心「只是时间问题」。Gavin 的模型判断是:数学成本下降速度快于内存成本,因此未来模型应大手笔消耗算力——拥有10亿 token 上下文、把「人类写过的每一本书都放进短期记忆」、由巨型分布式 MoE 组成的「大脑」。
1. 共识认为21岁的人造不出芯片——答案是整个行业「建立在缓冲之上」
- Patrick 先给出尽调结论:半导体公司通常由40-50岁、已经交付过多代芯片的人创立;「2个21岁的人不可能做到这件事」。Gavin 的回应是,这个反对意见编码了过时约束:「要相信自己能造出比历史上所有其他 AI 芯片都更好的芯片,需要一定程度的天真……而我们确实足够天真。」
- 承担核心重量的洞见是:整个半导体和数据中心栈都「建立在缓冲之上」——从 EDA 工具到电源模块,每一层都面向通用场景。具体来说,时序签核的 corner 默认芯片运行在冰点温度;「我从没见过结冰的 AI 数据中心」,但推理芯片从来不会低于80°C。砍掉这些无效约束,收益会层层叠加:「这里20%、那里50%、这里2倍。」
- Rob 判断早期信徒的标准,是看一个人依赖经验法则,还是追寻事实。典型代表是 Mark Ross,前 Cypress Semi CTO,该公司后来以90亿美元出售。他先说「不行,做不到」,随后要求一份白皮书和功能仿真;两次看到结果都感到意外:「嗯,这东西能跑」。他建议公司至少融资300万美元,最终公司融了500万美元,并从顾问变成半职顾问,最终成为全职 CTO。
2. 推理两个阶段押注两件事:低电压 Prefill 与热设计优先的 FLOPS
- 产品并非一颗芯片,而是一套完整的机架级推理方案:芯片、电力输送、板卡、互联,以及「真正意义上,生产就是产品」。推理分为 Prefill(读取文本、建立模型 KV cache)和 Decode(生成 token),Etched 将两者拆分到不同集群;Patrick 将其概括为「装子弹,再开枪」,Gavin 表示认同。
- Prefill 阶段真正重要的指标是实际 FLOPS,而不是宣传口径的 FLOPS:GPU 的 MFU 只有20%-50%,而且不可能达到100%,因为芯片会因温度过高而自我降频。所以「如果我现在只是给 GPU 增加更多 FLOPS……它只会热降频」——必须先解决散热,再增加 FLOPS。
- 机制来自 Dennard scaling——功耗与电压的平方成正比(「电压翻2倍,功耗就增加4倍」)。他们向数十名半导体老兵询问如何在低于 GPU 的电压下运行,得到的答案都是「做不到」;这很难令人满意,因为「Bitcoin 矿机运行电压不到 GPU 的四分之一」。Etched 的新型供电机制低电压推理,运行电压「低于任何其他 AI 芯片的一半」。Gavin 表示:「我们认为,未来所有 AI 芯片都会是低电压芯片。」
3. Decode 是一场内存竞赛——真正该问的是集群带宽,而非芯片带宽
- Gavin 重新定义了问题:人们问「你的芯片有多少内存带宽」,但真正应该问的是整个 scale-up 集群有多少带宽。Blackwell 的芯片间点对点跳转约为4,000纳秒,因此8路 tensor parallel 每位用户每秒的 token 产出,提升幅度「远远低于8倍」。
- Etched 打造了一套完全定制的互联栈——「二层以太网以上的所有东西,全部完全定制」——将延迟降低「超过5倍」,使整个集群的 SRAM 和 HBM 像单一内存池一样工作,也就是「集群级内存」。随着 world size 扩大,每个 token 所需时间按比例下降。
- 对现有厂商的诊断毫不留情:「这些架构都是在 ChatGPT 出现之前设计的」——如果针对现代工作负载重新设计,FLOP 组织方式、电压域、电源平面、封装、板卡设计和互联都会不同。物理上限仍然极高:数学意义上的延迟极限是光速,只有「2、3纳秒」,而今天是4,000纳秒。「底部还有很大的空间。」
4. 为什么这是本十年的瓶颈:Token 还没有经历规模经济
- Gavin 的宏观判断是,真正的 AI 已经存在,当前约束变成了并发量和速度:「不可能让10亿人同时使用这些模型」。如今付费 AI 计划的用户只有几百万人,「仅占全球人口的1/1000」。更快的 Decode 还能压缩墙钟时间:一个需要消耗一年推理时间的 agent 任务,可以压缩到1个月。
- Patrick 用 iPhone 做了标志性类比:对于 iPhone,「更多钱并不能真正买到更好的 iPhone」——亿万富翁和普通美国人买的是同一部。Token 还没有走到这一步:「一个相对小的通用系统,仍在手工制造这些 token……就像文艺复兴时期制造螺丝一样。」他希望 token 制造也获得 iPhone 级别的规模经济。
- 现实影响在于,一些产品(例如 coding 模型)低于特定 tokens/秒阈值就无法使用。因此今天的选择是:「要么让世界上很多人无法使用这些东西,要么所有人的体验都会变差」——这正推动硬件从晶圆到瓦特、从晶体管到 token 全面重做。
5. 起源故事:GPT-4V 一眼识别肿瘤,以及17岁的 kernel 工程师
- Rob 的动机来自切身经历:高中二年级末被诊断为四期骨癌,生存概率低于30%,经历2年化疗并重新学习走路。GPT-4V 发布后,他上传了一张背部肿块的诊断前照片,模型立刻说「这可能是肿瘤,去做 MRI」——真实医生花了他6个月才得到这个答案。随后系统弹出「你今天的图像额度已用完」的通知。「天啊……我们显然没有服务它所需的基础设施。」
- 他的另一个切入点来自运营 Prod 孵化器:Cursor/AnySphere 等早期公司都在其中。他看到每家初创公司都把融资烧在算力上,于是判断软件 COGS「以后不会再是零……它会成为推理的函数」——推理成为全球最大市场,将经历一场「长达10年的行军」。
- Gavin 17岁时在 Xnor 做 kernel 开发,年纪太小甚至无法签合同;之后加入 ExaNous(被 Altera 以2亿美元收购)和 OctoML(被 Nvidia 以数亿美元收购——节目中听到的名称是「Octo」)。Kernel 工作带给他的经验是:「数学相对容易……真正重要的是数据搬运」——这正是集群级内存要解决的问题。
6. 机器人竞赛模板:2个人、不做外联,只管赢
- 在 FTC 机器人竞赛中,Gavin 和搭档 Sanford 放弃了标准的20人团队模式,以及围绕文档和外联建立的文化:「除了造一台得分最高的机器人,我们什么都不做」,每3个月重做一次。他们一度保持世界最高分,并按 OPR 排名世界第三。
- Patrick 追问时指出,「你们选择不做任何沟通」;Gavin 承认,这「完全就像机器人队」。迁移到 Etched 的方法论就是:速度、速度、速度——「赢的方法就是不断交付」,并相信相比拥有20,000人的现有厂商,少得多的人也能「做出世界上最好的产品」。
7. 「最好的供应商就是没有供应商」:同时造机架、建工厂、做芯片
- Rob 将「最好的零件就是没有零件」延伸为从芯片、板卡、冷板、互联到生产的垂直整合:「我们是现在唯一一家同时自建机架和自研芯片的初创公司……而且是同时进行。」他们招募了曾经打造 Nvidia 全部 HGX 和 DGX 系统的机架负责人——这些系统贡献了 Nvidia「大约80%的收入」——Bryan Loyler(听辨如此)。
- 在芯片回片前,他们先用带有预期热点、规格完全一致的热仿真芯片进行验证,并不断给冷板加压,直到其爆裂:「从那以后我们一次漏液都没有」。他们还运营着一家台湾工厂,在办公室复制测试站,地面上部署一个2 MW 数据中心,并通过昼夜两班、7×24小时推进开发。
- 整合边界由规模经济决定:「自然边界在芯片侧的底部,以及模型层的顶部;中间的整个空白我们都会填满。」他们目前刻意不建数据中心:「那实际上并不能帮助我们更快上线容量。」
8. 人才:「传奇人物」加上「肩上有刺」——加入公司得「脑子有病」
- 对于尚未解决的问题,他们采用双轨方法:通过「项目制招聘」找到全世界真正最强的人——梳理历史上最难解决的技术问题,找到完成从0到1的人,然后持续追踪。「第一次谈话后愿意答应的人很少,但聊到第20次,答应的人数会出奇地高。」Brian 的价值最终兑现:「这是我学到的价值10亿美元的经验」——指的是在错误重演前,就把经验沉淀下来。
- 另一种模式是「肩上有刺的人,才能把芯片装进数据中心」("chips on shoulders put chips in data centers")。Sanford 就是代表:他被要求在一周内造出一块冷板——「任何热工程师都会觉得你完全天真,这种东西需要几个月」——但他仍然先把关键电源问题去风险。真正的化学反应来自组合:不清楚「尸体埋在哪里」的天真、第一性原理型冒险者,与熟悉规模化的老兵共同工作。
- Rob 的招募话术原封不动:「你得脑子有点问题才会加入我们……搬到 San Jose,加入一家由2个、现在应该是24岁的人运营的半导体公司……对抗全球最大的公司,而我们的设计不是比他们好10%左右,而是好10倍。你得有点问题,才会做这种事。」他的担忧是,随着规格公开、公司逐渐获得共识,这种逆向筛选机制可能会减弱。
9. 花钱买速度:Bangalore、预取,以及40天对10个月
- 临近流片时,一家物理设计供应商落后了1年;2个显而易见的选项都意味着再损失1年,于是他们选择第三条路:把数十名顶尖工程师派到 Bangalore 工作6个月,Gavin 在那里住了4.5个月,每天早上到岗,凌晨1点离开,并在每天上午8点和晚上8点进行12小时的美方交接。与同一家供应商、处在同一阶段的其他芯片,直到今天仍未完成流片。
- 烧钱加速的逻辑是:「最大的风险是不承担风险」。在一个每天创造超过10亿美元收入的品类里,「我们每晚一天交付,就会把大量机会留在桌面上」。因此他们进行「预取」:700块 FPGA 跑完整芯片和十几个模型,机架在没有芯片的情况下提前运到客户数据中心,用于先行调通软件,生产线也在芯片回片前就准备就绪。
- 最终结果是,一家「非常知名的 AI 芯片公司」从芯片回片到在机架内完成推理用了10个月——这一数字已公开披露给投资者。Etched 做到了40天,「因为芯片回来时,其他一切都已经变得无聊」。公司超过一半员工住在办公室旁边。至于轮班工资问题,Rob 的回答是:「看不见的手很有用。」
10. 最黑暗的2周:50皮秒、离职者,以及「谜题开始了」
- FPGA 能验证数字逻辑,却无法验证模拟逻辑。因此芯片回片后,跨时钟域的背压故障导致 attention 结果错误。唯一解决办法,是让芯片上的2个时钟信号彼此对齐,误差控制在50皮秒以内——即「50万亿分之一秒」——每颗芯片每秒重复20亿次。「有人辞职了……有人真的说,这个问题无解,祝你们好运。」
- 解决方法是:「第一步,先假设问题可以解决」。他们设计出一种能够以皮秒级改变时钟相位、再将其锁定的机制。大约2周后问题解决了——「非常恐怖的2周」。Gavin 将其总结为普遍规律:当一切看起来毫无希望的时刻,「恰恰是最应该投入精力的时刻」。
- 伴随而来的另一段战争故事发生在凌晨2-3点的首次晶圆测试,他们与 TSMC 通电话,屏幕上每颗 die 都从绿色变成红色。芯片验证负责人向后靠去:「谜题开始了。」实验哲学也贯穿始终:30次板卡实验中只有3次成功,但每次都「价值连城」。「人们来找我说,Gavin,你的实验几乎没有一次成功。我会说,我只需要幸运一次。」
11. 险些死亡的融资: 「硅谷每一家主要投资机构都立刻拒绝了」
- 2024年初、A轮之前,架构已经验证,但仅物理设计阶段就至少需要4000万-5000万美元;巨型 MoE 模型又意味着必须把整个集群搭出来,而公司账上只有1500万美元,未来12个月需要1亿美元。Gavin 诚实地回忆:「你坐在那里,心想,天啊,我们付不起这笔钱……回哈佛有多难?」一份30页的技术备忘录耗时100小时;每一家主要投资机构都拒绝了他们,理由是「2个刚从哈佛毕业的孩子,还没流片……未来都会是训练……这一切可能都是泡沫」。当时半导体行业最大的 A 轮融资规模约为4000万-5000万美元。
- 生存模式启动后,他们先设计了一套由债务提供方支持的3000万美元「从吃拉面活到流片」方案,随后开始逐一打电话:「我们需要1亿美元,你认识愿意激进下注的人吗?」直到一次董事会会议上,表格显示已有1.03亿美元软承诺。「我们互相看了看,说,那就接下这笔钱。」此后公司又完成了近6轮融资,其中很多投资人进一步加码,金额翻倍甚至翻3倍。
- 供应商比资金更早相信他们:Synopsys 以多年期条款延长仿真器租用,本质上是一笔大额贷款;TSMC 则是在一次 SEMI 会议晚宴后决定合作。当时22岁的 Gavin 是现场唯一一名30岁以下的演讲者,偶然坐在一位 TSMC 高级副总裁旁边;2人都是数学专业,餐桌上直接在纸上讨论每个 tensor 的模型机制。第二天收到邮件:「Gavin,想和 Etched 合作吗?想办法让它发生。」Rob 对 TSMC 的总体评价是:技术最好,但「真正的价值全部在服务里」。Etched 提议后,TSMC 自掏腰包做了良率实验,随后将方案推广到整条产线。「如果我去钢铁厂说改变钢材成分,他们会让我滚蛋。但 TSMC 不会。」
- Patrick 打破第四面墙:「我是 Etched 的大投资人……我有严重偏见。」他反思逆向投资:专家们「用非常合乎逻辑的方式解释为什么这件事不会成功」,而经验教训是,「你多少得诅咒一下基准概率」——因为总有指数基金可买。Gavin 观察到,真正相信他们的人要么相信市场和团队,要么技术极端扎实——高频交易公司会一路审计到 RTL;如果处在两者中间,「你就是无法理解它」。
12. 软件押注:Kernel 优先、不到100个模型,以及相对超大规模厂商芯片的生死优势
- 3年前,他们在图编译器(开箱即用、性能较差)和 Kernel 优先编程之间做选择;Etched 选择后者,押注「真正重要的模型不会超过100个」,不提供任意 PyTorch、CUDA 或 ONNX 支持。跳过编译器「为我们节省了大量时间」;最早相信他们的只有高频交易公司——「他们也都讨厌编译器」——其中数十名工程师后来加入了 Etched。
- 随着 AI 开始吞噬 Kernel 编写,这一押注正在获得回报:工具围绕模型未来的使用方式设计。在一次内部实验中,「Codex 只根据我们的文档,就完全自主地让 GPT-OSS 从零运行起来……一夜之间完成」。
- 市场结构层面的论点来自一名被 Etched 从招聘中途挖走的人:他原本是某前沿实验室芯片项目的架构师,一周内被以「Uno reverse card」的方式招入 Etched。他说:「对我的公司来说,这款产品赢不赢从根本上都不是生死攸关的事。Google 不会因为 TPU 失败而倒闭。Meta 不会因为 MTIA 失败而倒闭。Microsoft 不会因为 Maya 失败而倒闭。OpenAI 不会因为 Jalapeño 失败而倒闭。」Rob 提供了佐证数据:实验室和超大规模厂商芯片的「FB8 × FB8」FLOP 密度低于「Blackwood B300」(听辨如此)——「他们不需要承担这个风险,只要造出足够相似的产品,不必支付 Nvidia 税。」
- 供应方面,Etched 将其设计成正和博弈:第一代产品采用4纳米工艺和不同于 Rubens 3纳米方案的 HBM,因此大规模部署客户看到的是「2吉瓦」,而不是替换关系。最终原则是,设计阶段就要把供应链纳入考虑,因为「如果你的产品性能最好,却造不出来,那你就只是一档播客」。
13. 未来:巨型分布式大脑、每兆瓦 agent 数量,以及万亿美元级 Token 工厂
- Gavin 的模型判断始于「机器的思考方式和人不同」——飞机也不是像鸟一样飞行。对芯片而言,加载数据昂贵,而数学计算便宜;「数学成本下降的速度,快于内存成本下降的速度」。因此未来模型应该大举消耗算力:运行大量并行副本,让巨型专家横跨多个机架,并使用10亿 token 上下文。「我很想和一台能把人类写过的每一本书都放进短期记忆的机器交谈。」Rob 补充说,近期架构主题将是动态性——按 token、按用户控制算力和内存;ChatGPT 之前的硬件只能用「蛮力」处理这些问题。
- Noam Brown(如今是天使投资人)从墙钟时间角度解释速度的重要性:6个月的 agent 任务还没等到评估,下一代模型可能已经发布;而正如单个人无法造出火箭,agent 工作也需要团队协作——可能是10个,也可能是100万个——这要求巨大的共享内存和 FLOPS。Cursor 的 agent 已经能在一周内从零造出浏览器,「很快这会压缩到不到1小时」。
- Rob 最激进、也最有时间表的判断是:推理将沿着「全球性行军」发展,最终占据全球 GDP 的大多数,可能需要超过10年;生产率将重新定义为「每兆瓦多少个 agent」;「今年是人类占劳动力多数的倒数第2年——2027年,做知识工作的 agent 数量将超过人类」。对于单个价值1万亿美元的数据中心,他的回答是:「当然。只是时间问题。」规模经济不会在价值400亿美元的晶圆厂停下,同样也适用于「生产 token 的工厂」。
- Patrick 要求「聪明的外星人」做最后总结。Gavin 的答案是:思考有价值,每家公司都依赖思考,因此必须有人为「面向10亿人、同时运行的未来千万亿参数模型」铺设路线图。Rob 的答案是:生产智能的成本远低于智能本身的价值,「我们正处在这些 token 供给短缺的多年、很可能数十年时期」。节目最后回到「别人为你做过的最善良的事」:Rob 16岁时被要求在手术(活下来,但可能永远无法行走)和放疗之间做选择,后来仍不得不接受放疗;当时世界上只有少数几台设备能完成治疗,其中一台在 Boston,于是他的父母放下一切,陪他搬到了那里。
We know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. We had people quit.
Yeah.
People literally were like, “This problem is unsolvable. Best of luck, guys.”
You kind of have to be sick in the head to join our company. You’re going to convince your family to move to San Jose for the semiconductor company run by two, what, 24-year-olds now, going against the biggest companies in the world with a design that they’re saying isn’t going to be 10% better, but 10x better.
We need $100 million to do this. If we do this, we think this could be one of the most important companies of all time. You’re sitting in that moment thinking, “Holy crap. We can’t afford this.” And I’m looking at it like, “Huh. How hard is it to go back to Harvard?”
1. Why Nobody Believed Etched Would Work
All right, gentlemen. It’s been 3 years or so, Gavin, since you and I last did this, which is nuts. At the time, I was just wildly intrigued by your story and what you were going to build. I didn’t know a lot about chips, and I was considering investing in the company, so I was calling everyone I could conceive of who could give me an opinion or something.
At the time, basically, the consensus was that these kinds of companies are not built by young people. In the semiconductor world, the best companies are founded by 40- or 50-year-old people who have had a whole career’s worth of experience, have learned all the problems, and have shipped multiple chips. Two 21-year-olds are not going to do this. It’s just not going to work.
It was indicative of a theme, which was that nobody believes in us. That’s obviously changed a lot now. You just walk the halls and talk to the people who have chosen to come work here, but in the early days, it felt like this was something that you had to face down. What was that like, facing that down, where a set of incumbents and an industry’s worth of people and investors and everyone else sort of didn’t believe in you? What did that anneal in you to build the company the way that you are? What was the impact of that?
I think there’s a certain level of naivety required to think that you could build a chip better than every other AI chip ever built and build a company to do it way faster than has ever been done. We had the naivety. There were many times when we would say, “Why isn’t this possible?” and really push on it.
It turns out that everybody’s answers are extremely siloed to a set of constraints that aren’t true anymore. The reality is that the entire semiconductor and data center industry is built on buffer. What I mean by that is that every part of the stack—from the EDA tools to the power modules to the circuit boards to the chip design and standard cells—is built to be general-purpose for everything, not just in the data center, but for IoT on the edge and so forth.
When you have a specific use case that you’re really trying to design for, you can change the constraints a lot. I’ll give you a very simple example, which is one of the things you care a lot about: the clock speed of your chip. It’s proportional to the throughput of your system.
When you’re doing sign-off for different timing and figuring out what clock speed you’re actually going to be able to run at when you tape out your chip, there’s this concept called corners, which is what temperatures you’re going to be able to run at that clock speed. The default configurations for a lot of these EDA tools assume that you’re going to be running your chips in freezing temperatures. I don’t know about you, but I’ve never seen an AI data center with ice in it.
We can feel pretty confident that our chips don’t need to run at full speed at 0°C. In fact, they’re never really going to be running below 80°C anyway. Just by knowing that’s a constraint that doesn’t matter, we can make a ton of changes throughout the entire system. That’s a very simple one, but there are many more that get you 20% here, 50% there, 2x here, and these compound to a system that can be radically better for inference.
Well, I think you found two kinds of people. There are some folks who went purely on heuristics: “Okay, young founders, they claim they can go beat the biggest company in the world on performance. It cannot happen. There is nothing you could say to me that would make me change my mind.”
But there are also people out there who were, of course, skeptical but willing to say, “I’ll spend the time, I’ll do the work, and figure out whether this is actually possible.” For example, one of our earliest supporters was this guy, Mark Ross.
Mark was a very prestigious semiconductor expert. He used to be CTO at Cypress Semiconductor, which sold for $9 billion. When we met him, we were just a couple of guys in a dorm room. We came to him and said, “Hey, you want to go build hardware for inference? We think we can be much faster than NVIDIA.” Mark was like, “No, you can’t. It will not work. But if you want to convince me, you should write a white paper, build a functional simulation, and show me.”
After a lot of very long nights, we went back to Mark and said, “Here’s a simulation. What do you think?” He was like, “Huh, this works. But to go do a company like this, you’ll need a large amount of capital. You’ll need at least $3 million even to get you started.”
We went ahead and raised $5 million and raised a lot more after that. He was again surprised but got more involved. Then he became an advisor, then a half-time advisor, and eventually our full-time CTO, as he saw more and more of the development progress.
In general, this has been a filter that really heavily filters out folks who want to be right regardless. It attracts folks who want to be very truth-seeking and say, “Sure, I’m skeptical, but I will go ahead and work through the numbers myself. And if I can figure out why this is possible, well, let’s go build it.”
The specifics that you’ve made bets on and the way that you’ve built the system are immensely interesting to me. So many people are trying to do this now—build new chips that will do a better job of serving inference at massive scale—and the world is interested in the research approaches and the different architecture approaches that people are taking to building a new AI chip.
I’d love you to start by describing what this thing is and what it does, but maybe more interestingly and more importantly, the process that you went through to decide what bets to take and what technologies to invent. Compare and contrast those with what you’ve seen the rest of the marketplace try to do.
Yeah, I think to start with the product, we’re not just building a chip. We’re building a full inference solution, and that means a rack. That means the chip, the power delivery into the chip, the board in which it sits, and the interconnect that allows the chips to talk to each other. That means the production for this massive volume of racks. Really, the production is the product.
When we think about how we get our advantage, there are two key parts of running inference: prefill and decode. We have two key techniques that match both of these things. Prefill is reading in a huge volume of text, and decode is then using that data to generate output tokens.
When you run prefill, your key job is not to predict tokens. You already know the text. Your job is to get the model’s memory, what we call the KV cache, into the right state. Then you can run decode with that same KV cache.
What we will often do is what we call PD disaggregation, or prefill-decode disaggregation. You will have one cluster of servers running these prefills. You’ll then transfer those model memories, those KV caches, over to the decode cluster and use that cluster to generate the next tokens.
So it’s sort of like loading the gun and then firing it, if I think about it in super-simple terms.
Yeah, you got it. It’s getting the models to remember the right things and then using those things to do tasks.
Generally, people think about this market a bit lazily, or they say, “Are you a prefill chip or are you a decode chip? If you’re a decode chip, are you an HBM chip, an SRAM chip, or a 3D DRAM chip? Are you using optics or using copper?”
When we started this, we just wanted to understand why extremely smart people were working on these different directions. We seriously looked at architectures like having a bunch of DDR memory in a shared memory pool and looking at advanced packaging to basically break out of the shoreline. We looked at things like, “Are there ways to put memory dies on top of compute dies?”
In doing so, we realized that there’s no free lunch. Everything has a trade-off. With 3D DRAM, you have a thermal issue, you have a supply-chain issue, you have to figure out hybrid bonding, you have to figure out the FLOPs, and now you’re a decode chip.
So we went through everything, both on the prefill and the decode side. In doing so, we realized there are a few design spaces that nobody had seriously tried to explore because they had never been done in AI chips. We asked ourselves, what are the actual metrics that are going to matter the most?
On the prefill side, the thing that matters is FLOPs and FLOPs density. People talk about FLOPs often as a headline number, but in reality, you should care about the FLOPs you’re getting when you’re running real workloads. There’s this concept called MFU, or model FLOPs utilization, which is, for every peak FLOP advertised, how many cents on the dollar are you actually getting? On GPUs, you often get somewhere between 20% and 50%, depending on the workload.
Actually, you can probably not run at 100% because you have a thermal issue. Whereas, as you increase the FLOP utilization, you have more transistors going on and off, you draw more power, and the chip will self-regulate and actually lower its clock speed to make sure it doesn't overheat.
As we looked at inference, we said, if we want way more FLOPs because we want to run at way higher throughputs, we fundamentally need to solve the thermal problem before we even think about adding FLOPs to the chip. If I just add more FLOPs to a GPU today or another AI chip, I'm not actually going to get more performance because it's just going to thermal throttle.
Fundamentally, the essence of that is this concept of Dennard scaling, which is that power is quadratically proportional to voltage. So, if I 2x my voltage, my power goes up by 4x. If I cut my voltage in half, I cut my power down by a quarter.
So, we asked ourselves, how could we run voltages lower than GPUs? We talked to a lot of people about this. We flew out to Silicon Valley after dropping out and basically asked the dozens of people in semiconductors at all these different chip companies how they did it. The answer we got was, “You can't. You can't run at voltages lower than GPUs.”
This was very dissatisfying because there were many different industries of chips that run at voltages lower than GPUs. Bitcoin miners run at under a quarter of the voltage of GPUs, so this is obviously physically possible. The question is, are there issues with GPU architectures that make them unable to run at these voltages?
When we looked at the problem for a long time, we were able to create a new mechanism for running at much lower voltages, a new type of power delivery that we call low-voltage inference. We think all AI chips in the future are going to be low-voltage chips. They're going to have to cram way more FLOPs into the same silicon area and, without thermal throttling, run at way lower voltages. So, that's prefill.
For decode, it is all a memory game. More memory bandwidth means you can load the model faster, load the KV cache faster, and serve more tokens per second per user. We think people ask the wrong question here. People often ask, “How much memory bandwidth is on your chip?” You should be asking how much memory bandwidth is on your full-scale-up cluster.
What we're able to do is add way, way more bandwidth and a much lower latency from chip to chip to our interconnects. It allows us to serve models at this much higher speed because you can use the SRAM and the HBM from the full-scale-up cluster as a single pool. That's our second key technical bit, what we call cluster-scale memory.
On GPUs today, the cluster memory bandwidth is often very badly utilized because the time to hop from one GPU to another is extremely long. For example, on Blackwell chips, it can be about 4,000 nanoseconds to go point to point. That means that if you go ahead and go to an 8x tensor-parallel setup, you will get way, way less than an 8x improvement in your tokens per second per user.
What we did was build our own totally custom interconnect stack. We took everything above Layer 2 Ethernet and built it fully custom. We can go far better in terms of latencies and bandwidths this way, too. We can go ahead and cut this by more than a factor of 5x, and that allows us to use the memory of other chips much more effectively. As you scale the world size, your time per token goes down proportionally.
That's not that surprising, given all these architectures were built before ChatGPT. If we're trying to build a chip for modern workloads, it's going to look very different. The way we organize our FLOPs, the way we do our voltage domains, and the way we do our power planes are going to look super different. The way we do the packaging is going to look super different. The way we do the board design is going to look different. On the decode side, the way we connect everything is going to look very different.
We're now bringing forward our first generation of this low-voltage inference technology, which is running at under half the voltage of any other AI chip.
2. Why Inference Is the Bottleneck
If you zoom all the way out, why is this so important? Why is the delivery of much higher throughput, much lower cost per token, better tokens per watt—all of these metrics—the universe is going to start talking about more and more? Everyone knows the supply side of the equation is a big problem right now. Why is this, in the bigger picture, looking at a decade, the bottleneck in the technology world?
Well, I think it comes down to productivity. We are at this extremely interesting moment in the history of civilization where there is real artificial intelligence—not sci-fi stuff, but models that can solve problems that most humans can't. It's going to create new scientific discoveries, instant access to medical care, and instant access to education.
Now it's just about how many people can use this at the same time, how many products can service this at the same time, and also the speed of doing different tasks. When you think about wall-clock time, if we can take an agent that can run at a certain model quality and could take a year to solve a certain task using inference-time compute, if you have way faster decode speed, you can compress that into a month. The amount of scientific innovation and the amount of actual proliferation of technology will happen much faster.
The second part is concurrency. Today, it's just not possible for a billion people to use these models concurrently. Ultimately, some people are going to get downgraded, some people's models are going to be slower, and some people just won't be able to access the hardware.
A few years from now, there are going to be giant models serving billions of users. We're very much in the early innings of AI today. With paid plans, there are only a few million users in the world using paid plans for AI models, so we're at 1/1,000 of the global population actually using this stuff.
3. The Future of Models, Agents, and Intelligence
If you want to serve at giant scale, a lot of things change. One of them is the number of chips that communicate together. People usually think about this in the context of training. You have these giant training clusters, and you have Colossus with over 100,000 GPUs that are all networked together.
On the inference side, today people usually think about it as an 8-chip cluster, or maybe just NVL72 as the scale-up domain. But very quickly, this is going to become thousands of chips and tens of thousands of chips. The way to get the most performance there—the time between sending data from one chip to another—that primitive matters way more than is getting credit for right now.
When we think about optimizing memory bandwidth for the system, you have to think about how fast these chips can communicate together. If they can only communicate really quickly with themselves and very slowly with other chips, you're not going to actually be able to serve giant models at 10,000 or 20,000 tokens per second.
We need multiple orders of magnitude of infrastructure built out throughout the entire stack, from the power—from the wafer to the watt, from the transistor to the token—to actually bring this stuff to the world.
I think that when you look at most other goods, like the iPhone, for example, you've gotten to these economies of scale. As a result, more money does not really buy you a better iPhone. If you're a billionaire or you're just the average American, you buy the same phone.
Tokens aren't like that yet. We're still in the very early days, where a general-purpose system, a relatively small one, is kind of handcrafting these tokens, like they made screws back in the Renaissance. I want to live in the world where you have the same economies of scale for token-making that you do for making, say, iPhones or cars or anything else.
I think that is one of the huge unlocks that allows a huge group of people to use the best-quality models. I think that economies of scale have made capitalism very—I don't know—fair.
I think that allows you to go ahead and have the same product in many, many different hands, and you're able to serve way more users on a single scale-up cluster. It allows you to get closer to that point for token serving, too.
Yeah, and also, certain products aren't usable if they're slow.
Yeah.
So, if you want to serve coding models and you want people to actually use them, there's a certain number of tokens per second you need to hit. The question is: while maintaining that per-token speed, how many users can I serve at the same time?
You can basically decide, "I'm going to shut off a bunch of the world from using this stuff," or everyone's going to get a worse experience. Fundamentally, you need to find ways to push out the curve, and that's why there's such pressure for new hardware.
I'd like to take some time to step back and hear both of your stories for how you came to this idea and this company, and then walk through what it's been like to build it. I think in so doing, we'll understand the system that you've built for the company itself, which will then be able to power subsequent generations of products like this one for this crazy inference future that we're staring down.
Rob, maybe starting with you, just take it however far back you want. What I'm curious about in your personal story is that the very first thing I ever heard from either one of you was your personal story many years ago, which really blew me away. I'm most interested in your motivation, ultimately, for being here doing this thing.
It starts back in high school for me. I've been very unlucky and lucky at different points in life. This was one of the tougher times.
At the end of my sophomore year of high school, I got injured at a martial arts tournament. The next day, I couldn't walk for some reason. They thought there was something wrong with my SI joint or something. I went through physical therapy and did different types of scans, but they couldn't figure it out. Eventually, they found this big bump on my back in an MRI and told me it was a tumor.
It was stage 4 bone cancer, and I was told I had under a 30% chance of survival. It was a 2-year, crazy chemotherapy, surgery, and learning-to-walk-again experience. When you go through something like that, it really changes the Overton window of human experience and makes you appreciate what actually matters. You also ask yourself, "What are you going to do if you have the chance to live?"
If you actually want to get through something like that, you need to be hoping for something. I always knew I wanted to do something very impactful if I had the chance to get through it. It took me a couple of years to figure out what that was going to be.
At the same time, as I got to college and met a bunch of other people building cool tech, I got extremely excited by AI models, especially once GPT-3 came out. I was like, "Wow, this is the first model that can kind of speak English." These things are going to get really smart.
What happened was, when GPT-4 came out, there was GPT-4V, which was the first model with image uploading. I went through my camera roll and found a picture of my back with this bump on it before I was diagnosed. I said, "Hey, ChatGPT, pretend you're an expert doctor. A patient comes in and says they have this bump on their back. What could it be?"
It immediately said, "This could be a tumor. You should get an MRI immediately. Go to the doctor." I just kind of sat there still. It was—
This took 6 months.
Yeah, that took me 6 months. Yesterday, this feature wasn't there; today, it's here. I went to show my parents, and I got this notification saying, "You're all out of image credits today. You need to get a Pro plan."
I was like, "Holy crap. This is going to change everything." We clearly don't have the infrastructure to serve it, and there are very few things you can work on that can actually bring this technology to the world at scale faster.
There are plenty of people who are super smart working on models. The fabs seem maybe unreachable to work on, but it seemed like the hardware was all designed before ChatGPT. Every GPU, every TPU, and every AI chip that was serving these models was fundamentally built before this and retrofitted to serve these modern models.
There's going to be an entire new wave of architectures that come out. What are more exciting things to work on than bringing this to everybody?
A very different angle at the same time: I was running a startup incubator called Prod, which has incubated a bunch of different companies. Some of the earliest ones were Cursor and AnySphere, which merged, and Recur and Hatchpoint through it, along with a handful of others.
At the time, as these models were getting smarter—it was 2022—I was realizing all of these companies were spending all the money they raised on compute. I had this realization as I was working on some of my own stuff: "Oh my God, all the products I want to build are going to cost tens of millions of dollars a year in inference."
This is not going to be tenable. The cost structure of every software company—the COGS—is not going to be zero anymore for an incremental user. It's going to be quite high, and it's going to be a function of inference. The OPEX of an already-built business is also going to be inference as people use more and more coding agents.
4. Gavin and Rob’s Origin Stories
Fundamentally, it seems like inference is going to be really important, and it feels like we're on a decade-long march for inference to become the biggest market in the world. When you think about that, 10 years from now, there are going to be these giant projects where fundamentally nothing in that data center has been designed today. We should go pick something to work on, and that's kind of how it got started.
Yeah, but I'm really excited for you to go back about as far—probably early in high school, maybe even earlier—and tell me your favorite milestones on the timeline that ultimately led to your ambition to drop out of Harvard and start this company.
My first job ever was at a company called Xnor, where I did kernel development. I was 17.
Yep.
A 17-year-old can't sign legally binding contracts, so I had to do a traditional NDA. They sat me down and said, "Gavin, don't share this information."
Xnor was one of the only companies that saw, “Hey, maybe this is a good trade.” ExaNous got bought by Altera for $200 million. I did the same thing at OctoML, where they got bought by Nvidia for hundreds of millions of dollars.
When you do this sort of kernel work, what you realize is that the math is relatively easy.
Okay.
But to get high-speed decode, I think what matters is data movement. Almost all the work that you do is optimizing how you move data around a single chip or across multiple chips.
That's why we went ahead and built this cluster-scale memory technology. We bring that interconnect time way, way lower. You can do way more movement and, as a result, get a much faster time to generate each subsequent token.
You can build these crazy things Rob was talking about, doing a year's worth of work in a month, or more than that in the future.
Can you talk about the competitive drive that's evident in some of the high school competitions that you participated in and won?
We did a couple. For example, I was very active in FIRST Tech Challenge robotics, and I was lucky to have a very talented partner, Sanford.
For a long time, we were part of a traditional school team, where it was about 20 guys all working together, as is often typical of FIRST Tech Challenge. The goal is to build a robot that scores the most points, along with a bunch of other things.
In FIRST, they put a lot of emphasis on collaborating with other teams, trying to do really good documentation, and trying to inspire others to do the same thing. Sanford and I decided that, rather than do it this way, we were going to win. We did nothing else besides build a robot that scored the most points.
As a 2-person team.
Rather than a 20-person team.
That you were much, much smaller than almost every other team in the competition.
We figured that if we were going to specialize, if we were going to go out and win the damn games really well, we wouldn't need to advance based on the quality of our documentation or our outreach. We were just going to go win.
And so we did. We brainstormed as a 2-person team, built a robot, and decided we were going to redesign it every 3 months. And we did. We actually had the world record for the highest score during this competition at one point.
We were rated by OPR third in the world for software development, and it was a damn good machine.
What from that episode can I translate as an analogy into how you built Etched the company?
When you're thinking about how you want to do a full rack-scale product like this, there are a couple of key ideas. One of them is velocity, velocity, velocity: you win by shipping. You're not going to go out and win by having the best outreach or the best communications.
You chose to have no communications.
Exactly. I’m just realizing how exactly like the robotics team this is. There are many ways you can win in business, but we’d rather focus on just building the best product.
Mhm.
Similarly, we think we can do it with a lot fewer folks. If you’re willing to just focus on product, product, product, and parallelize relentlessly, you don’t need 20,000 people like the big companies have. You can build the best product in the world with far fewer people.
Yeah, there's a saying that the best part is no part. I think for us it's also the best vendor is no vendor. As much as possible, we want to vertically integrate the entire product, both because we get more performance, but we can move way faster. So, everything from the chips to the boards, to the cold plates, to the interconnects, to even the production, we want to do all of it as in-house as possible. We're actually, I think, the only startup right now that's building its own rack as well as its own chips. And we did it all at the same time. A couple years ago was the last time we were public. At that point we just started building our rack team, and we brought over Bryan Loyler, who built all of Nvidia's HGX and DGX systems, which is like 80% of the revenue. And we said, "We're going to build the rack at the same time." We actually went through multiple iterations of the rack before the chips even came back. Before the chips came back, we made thermal chips that had the exact same hot spots as we expected our chips to have so we could build the cold plates, we could over-pressurize them and blow them up. We haven't had a single leak since our chips came back with the cold plates because we already validated them. Yeah, we have a factory in Taiwan. We have a few dozen people out there. We built a clone of a bunch of the test stations in our office. We have a 2 MW data center on this floor and we did 24/7 development cycles. People are doing day shifts and night shifts to actually get the hardware up and running as quickly as possible. So, it's that extreme vertical integration and extreme parallelization of the schedule that lets you get products to market way faster.
Hmm. If you think about building the early team and what it required as 2 young guys building this company, there are lots of very talented young entrepreneurs out there, perhaps building something of this scope or magnitude for the first time in a long time. All of whom could probably benefit from the lessons that you’ve learned getting very sophisticated, talented people to come join you, even after careers at the other great companies. If you were teaching this as a class—here’s how to get elite talent when you’re young and inexperienced and naïve—what would the syllabus be?
We have a pretty bimodal talent philosophy. It starts with what we call the legends. When we’re trying to solve an incredibly hard technical problem and generally do something that hasn’t been done before, we need to find the very best person in the world. Often, the number 1 guy in the world versus the number 10 guy versus the number 100 guy is a huge difference in whether it’s actually possible to solve the problem.
We created a system we call project-based recruiting, where we map out all of the hardest technical problems across all industries that anyone has ever had to solve. We look at temporality: Who are the people who did the 0 to 1? Who was in charge, quote unquote? Who actually did the work? We talk to as many people as possible, and then we just track it. You’d be surprised by the number of people who say yes after the first conversation being pretty low, but the number of people who say yes after the 20th conversation being surprisingly high.
You really have to keep at them. When you hear no from somebody who really is the best in the world, that really means, “Hey, should you come back when you have a few more milestones proven out?”
Yeah. I think one of the most convincing things to see is, “Hey, we make bold claims.” When you hit those again and again and again, that is really belief-inspiring.
When we decided we wanted to build a rack and not just a chip, we were looking at this and saying, “How many products have actually shipped at scale for a rack-scale system that actually has the power density we’re trying to solve?” We said, “If we were going to wave a magic wand, what would the best possible person in the world look like?”
We’d say, “If we could find somebody who started at NVIDIA and built the entire rack team through all their different generations, learned all this different stuff, but is still scrappy, still understands the startup culture, but has seen scale, that would be the best possible person.” So, we mapped all of the different teams that related to all of the different rack-scale products at NVIDIA, and we found 3 people that we thought could fit the bill.
We talked to all of them. Two of them had just retired, and one of them was planning to do one more generation for NVIDIA and then retire. His name is Brian. Over time, we convinced him to join.
Brian started the HGX and DGX team at NVIDIA, which was a majority of NVIDIA’s revenue—tens of billions of dollars a quarter. The other 2 guys ended up investing, by the way. When you have somebody like that, they just know what good looks like because they’ve seen it.
There were so many times where we’d talk to Brian and he’d just point to us and be like, “That’s a billion-dollar lesson I learned. Billion-dollar lesson I learned.” That just saves us cycles. You pair someone like Brian with somebody like Sanford.
Do you have a name for them? Brian’s a legend. What’s Sanford?
You know, we say, “Chips on shoulders put chips in data centers.” Sanford and Gavin, in high school, were world robotics champions. Sanford was finishing his senior year of college, and we called him up a couple of years ago and said, “Hey, can you come check out what we’re doing? We need some help on the platform side.”
He came for a week, and we said, “Can you build the cold plate this week?” If you asked any thermal engineer, they would think you were totally naïve, right? These things take months to do. To be clear, they do. But you can make real progress in a week if you put your mind to it and think it’s possible.
He built a contraption in a week that actually de-risked a pretty key power question we had. You put those 2 together, and they’ve done incredible things. One is not possible without the other because you need the extremely driven people who just keep asking why and don’t know where the bodies are buried to take tons of aggressive risks. Then you need the people who’ve seen scale and still have the startup scrappy mentality to help them along the way.
It’s really the legends plus the raw, somewhat naïve, first-principles-type talent. But it’s the combination. It’s not just that you have both in the company; it’s that they’re working together.
That’s right.
If I think about that funnel, is there anything else more interesting to say about how much better you’ve gotten at recruiting and why those metrics keep getting better?
One of the shocking things is, I wish we weren’t being so contrarian, but it kind of self-selects.
Right.
If you’re the kind of person who is somewhat opportunistic, you’re going to go join whatever the hot company is, or go ahead and do due diligence. You will not come work here.
Right.
It’s one of the things I worry about as we announce more and more of the product and its specs: We may lose some of this if we’re not very careful. You kind of have to be sick in the head to join our company.
If you think about it on paper, it’s like you—a person who is probably a very accomplished engineer, making a good amount of money; it’s liquid, it’s predictable somewhere else—you’re going to convince your family to move to San Jose and live in this apartment on this housing program for the semiconductor company run by 2 24-year-olds, pre-product, going against the biggest companies in the world in the most supply-constrained environment ever created, with a design that they’re saying is not going to be 10% better, but 10x better. Something must be wrong with you to do that.
People are just wired differently here. They really want not to prove people wrong who don’t believe, but to prove people right who do believe. They just take it personally. It’s really fun to find those people, and frankly, the nature of the company makes it very easy to weed out the people who aren’t like that.
5. Taking Huge Risks to Move Faster
One of the very first things you and I talked about, Rob, was that I started asking about Sohu, which is the name of the first product here. You said we could talk about that in great detail, but the thing you should know is that what we’re really focused on is building a machine that can, at scale, produce these things and generations of them as efficiently and at the highest possible quality levels.
So, we want to build the company—or the machine that is the company—that will produce this thing and subsequent things. I’d like to talk about a few principles or cornerstones of the company.
One of them we've talked about—we've alluded to some of them. You've said velocity, you've said vertical integration. These have become more popular topics. Parallelization is something maybe that we should talk about. But I'm especially interested in your guys' willingness to take huge risks to go faster. Maybe tell your favorite story about why this is the philosophy and what it's allowed you to do that maybe other companies haven't done.
There are a number of stories here, but one of my favorites is there was a time when we were getting close to taping out the chip. We realized, "Wait a minute, one of our vendors is way, way behind schedule." We had 2 very bad options. One option was to keep the current vendor and push our timelines out by on the order of a year. Another option was to switch vendors, start over, and also push our timelines out by a year. Neither of these was a good option. So we had to go look for option number 3.
What that was, was we figured out they're all in Bangalore—the team that's actually going and doing the work. We went out and shipped dozens of our top engineers across the world to Bangalore for 6 months. I was there as well. I lived in Bangalore for 4 1/2 months personally. Every morning, we'd walk across the crazy-busy Bangalore streets into the office. We'd be the first ones in. We'd build a wide variety of tools, both things like auditing a huge amount of the code that was going in, building a bunch of tools as well to make this go even faster, and making sure we're making the right design decisions on the spot, right there. No 12-hour back-and-forth. We could decide immediately. Then, at 1:00 a.m., we'd walk back through the now-empty Bangalore streets and do it all again the next day.
We still had a bunch of the team in the US. We ran these 12-hour-on-each-side handoffs, where we had a 24-hour cycle. At 8:00 a.m. and 8:00 p.m. every day, we'd all get on Zoom, share all the data, and say, "When I wake up, this must be done. We must get this chip out." It was extremely intense. At the same time, we saw other chips at the same stage as us with that same vendor that ended up taking years and still aren't out today. They still haven't taped out today. It's that level of extreme urgency that's required to bring products to market.
What is the key to doing this well? This has become a trope of Elon, mostly—that his special skill, and others who seek to emulate him would try to do this too, is figuring out what the binding constraint is and just flooding the zone personally on that thing, which is kind of like going to Bangalore or something. It seems like this is a central tenet of the business and of any business that's going to do this kind of vertical integration. What's the key to doing that well? Again, what have you learned about that specific act?
For me, I think there were 2 key tricks to this. The first one is that you can't build a chip alone. It's got to be a team problem. Your most important job is to go get great people to go with you and great people to be inspired and excited to do crazy things like this. It is a huge ask to say, "Hey, guys, uproot your lives for 6 months, or in one case, 12 months." We had sent one guy out well ahead. It sucks. But we're lucky to have team members who are in it for the right reasons.
I think the second big thing, too, is being able to make decisions very fast. One of the worst things is when there's a factory or a vendor who's waiting for you to make some call and is then just stalled. This happens all the time, even for very small things. So send folks, delegate a big amount of responsibility to them, and say, "Make a reasonable call. It's okay if you're wrong every now and then. But I would much, much rather be right most of the time and give an answer immediately than wait every time for the perfect response. Speed wins."
What about spending money to go faster? There's this learn-by-doing thing, which has become so interesting, and as the world has gone away from software and towards more hardware again in the world of technology, we've outsourced so much of the learn-by-doing by shipping stuff overseas and effectively just being the idea guys here in the US. It seems like that's obviously reversing, and you've adopted this way of learning by doing—you want to be in that iteration learning loop. Part of that is willingness to spend and take risks with dollars. Can you talk about that a little bit?
I think there's a great quote: "The biggest risk is not taking risk." It's very similar here. Every day there's over a billion dollars of revenue in this category, and a lot of it's inference. So every day we don't ship, we're just leaving tons of opportunity on the table. Your willingness to spend money should be extremely high if you can get a very clear ROI out of it.
We have this concept that we call pre-fetching, which is when you're waiting for one thing to get done, and you know you're going to do other things once you have it. Are there ways that you can parallelize the entire schedule? For example, we know our chip is going to come back on a certain date. We want everything possible that could be done without the chip to be done before the chip lands.
This costs a lot of money. It means that we want to build our entire software stack beforehand. We actually shipped racks to customer data centers without our chips in them, with all of the networking, all the CPUs, and all the storage set up, so we could bring all that data center software up before the chips came back. It meant that we took over 700 FPGAs and put the entire full-radix chip on an FPGA cluster and ran a dozen different models with our full inference stack on them before the chips came back.
It means that we built a thermal chip to mock the thermal profile of our chip and built cold plates based on that before the chip came back. It means we had the entire production line ready. It means we did many revisions of the circuit board. The entire product was ready to go before the chips came back.
And this is what it gives you. There is another very famous AI chip company that took 10 months to go from getting their silicon back to having it running inference in a rack. This was publicly announced to their investors, and it was a really big deal. We were able to do it in 40 days.
It's because by the time the chip came back, everything was boring. The software was already written. The rack was already there. The production line was already set up. We were just getting everything together. You don't always catch everything. You make some tweaks on the fly, and then off you go.
Although in that particular case, too, that was a big part of it. Also, I think the shift made a big difference, too.
We went out and literally had a day shift and a night shift. There were team members who would come in at 10:00 a.m. and leave at around midnight.
Yeah.
You would come in at midnight and leave at 10:00 a.m. You're running around the clock to get to those 40 days. Yeah, I mean, over half the company lives next to the office, so it makes it easier to do that type of thing.
You pay them to do that, right? Pay them extra? Do you still do that?
Yeah, the invisible hand does wonders. It works for me, too. We're both there.
I'd love to take one big step back and talk a bit about just the broader ecosystem here. The amount of shortages on the supply side, the exposure of risks in the global system, and the supply chain around this stuff has become everyday Wall Street Journal front-page news. The stocks that people are watching and investing in and excited about—if you think about the memory stocks, these were boring, commodity-like nothing burgers 5 years ago. Now they're at the center of global attention.
If you just assess the global, connected supply chain that's required to make stuff like this possible—just riff on it. What scares you? What's working well? What needs to change? What do you hope you change by virtue of how you build this thing? What's your assessment of this story right now?
I think that one of the most undervalued pieces of the supply chain story is that, in almost none of these cases, do you buy something and never talk to the vendor again. You have to collaborate. That is the most important part to being successful, I think, in chips with TSMC or with memory vendors. You need that partnership.
I think that for TSMC in particular, people don't understand why it is so valuable. People look at the technology, and the technology is the best in the world. But for me, the real value is all in the service. TSMC customer service is way better than I have seen at any other company in any other industry.
It's the kind of thing where if you say, "Hey, we're trying to improve your yields by making this change," you can make them a recommendation, and they'll run the experiment on their own dime, in our case, to see if they could actually get the higher yield. When we found that we were right and the experiment worked, they moved it over to the rest of the line. That kind of thing just doesn't happen in most industries.
If I go to the steelworks plant and say, "Hey, I want you to change the composition of the steel," they'll say, "Screw you." Not TSMC. That is why they are number 1 and why they're going to win.
One of the things that matters a ton is power availability and time to power.
And the problem is, the more power you want, the more shortage there is. It's actually very similar to chip clusters. Why is Colossus charging $12 an hour for Blackwells? It's because they're the only place you can buy 20,000 of them at once, right? Why is the 500 MW data center so hard to find? It's the exact same reason.
One of the things you need to think about is: How do we get way more juice out of each megawatt? People are looking throughout the entire stack, whether it's just improving the PUE, but also entirely new hardware to get the most tokens per megawatt to solve this problem. But fundamentally, building new buildings is hard. It's much easier to go from 100 MW to a gigawatt, and then from a gigawatt to 10, and 10 to 100. We are pushing the limits of what's possible on these timelines. So, there are a lot of people trying to scale in their data centers as much as trying to scale them out.
Yeah, I mean, one of the interesting things about a system like this is what it replaces. If I think about a rack like this versus, I don't know, a set of Blackwells or Rubins or whatever's coming next, how should I conceptualize that? It's not just watts; it's also physical space, to your point. Cerebras talked about this in their recent earnings call: This is literally the problem. There's literally no space to put the systems. How should I conceptualize what this represents or replaces in terms of other units of compute?
Here's how customers think about deploying models generally. When I'm building a data center or I'm building a cluster, it's not in the abstract of, "Oh, I like these chips, and this is the power and footprint and so forth." It's, "I have a real production workload I'm trying to serve. For my product to be useful, there's a certain speed I need to serve at. For certain products, it's really fast, and for certain products, it's really slow. Whatever the speed is, this is my speed."
The question is: In a given amount of power, how many users can I serve while guaranteeing that speed? So, another way to put it is, this is what's called interactivity: What is my throughput? We are just finishing the early innings of the AI infrastructure boom, where people really just cared about speed. GPUs were not able to reach a lot of the speeds of other types of chips, like all these SRAM chips—thousands of tokens per second—and that enabled tons of new use cases that got people very excited.
There's an entirely new wave of AI chips, us being one of them, that are all going to be able to hit these speeds. The question then is: If you're hitting these speeds, what is the number of users you can serve at the same time? By proxy, if I have a 100 MW data center, how many software agents can I run at the same time? When people are doing that evaluation, our hardware is generally going to be able to get you an order of magnitude more concurrency at a given level of interactivity. That directly translates into tokens per watt, tokens per dollar, and all the things people care about when they're actually serving these giant mixture-of-experts models at scale.
There are these now-famous interactivity curves, right? Not many people publish them, but you can see a Blackwell curve, and you can see an AMD curve, which is a little bit worse than Blackwell's, and it's still an $800 billion company. So, if you think about what the impacts are of shifting that curve—not just a little bit further out, but much further out—
Yeah.
What are the things that most excite you about what this will enable?
I want to go out and solve some of the hardest problems, and I want to go solve these in much less time. There were things growing up that I was not sure I'd be able to live to see. For example, the unit distance conjecture was one of the things I thought about in college.
Yeah.
And I was not sure I'd see that proven in my life. This was done by an AI model, and it was done over a long period of time. But if you're able to run the same model 10 times faster, you can shrink the time to go have these breakthroughs.
There are a huge number of other problems in math like this as well that I worry will take 100,000 years to prove. You can either have a much smarter model or a model of the same intelligence running much faster. You can then shrink that time, and I can see it. It's so cool seeing these breakthroughs get made. I am so, so excited to see much more of this happen.
I think too often people think about tasks and applications and stuff in these very short time horizons. Doing a chat and having it be 50% faster is nice, but it's not game-changing. As these agents go to longer and longer time horizons and the models get more and more capable, you're going to see gigantic bodies of work that would take months of compute.
When we think about this in wall-clock time, if you talk to a pre-training researcher at a lab, they'll tell you that wall-clock time is often one of the most important things that matters. What wall-clock time means is the time from starting your run to finishing it, to actually getting data back. If you can shrink this time from a 6-month run to a 2-month experiment, you're going to be able to do many more iterations, and people will make changes to the model architectures to actually improve the wall-clock time.
It's very similar here in terms of how we think about the use cases. The exciting part about super-low-latency decode is that wall-clock time on long-horizon tasks becomes much shorter. A year-long compute build would now take a month, and that month-long compute build will now take 3 days, and that 3-day compute build will now take 7 hours, and so forth and so forth.
That's the thing that I think is really hard to internalize, because the models are just getting capable enough to do this stuff. I thought it was really cool months ago when Cursor announced that they had a bunch of coding agents build an entire browser from scratch in a week. Totally nuts. That will soon happen in under an hour. There are going to be many of those types of things that happen with these massively parallel agents all working on a given task.
What are the ultimate limitations of these systems? Is it just a physics question? How many times faster and cheaper can we get theoretically? How do you think—
There's a lot. There's a lot of room at the bottom, as they say. If you think about chip-to-chip latencies on an NVIDIA product, you're looking at 4,000 nanoseconds to go from one chip to another. We'll be able to do much better than that.
What's the mathematical limit?
It's the speed of light. You can do it in just a handful, like 2 or 3 nanoseconds.
And they're at 4,000?
4,000 today. There is a lot of room at the bottom. The same sort of thing applies to power efficiency. We're able to shave a huge amount by bringing the voltage down by so much, but you could go lower. You could go much, much lower. It's very challenging, but when I think about 20 or 30 years in the future, I think it's inevitable.
6. Kernels, Compilers, and the AI Stack
The same is true for economies of scale and cluster scale-up. For a long time, 8 chips was the biggest scale-up domain. Then they had the NVIDIA NVL72, bringing it to, well, 72. But you can be way, way bigger. You look at a fab, for example: You have a $40 billion single monolithic building with only a handful of lines running through it. You could have the same kind of thing for some futuristic mega-cluster—a $40 billion or $100 billion giant mega-token factory serving one or a handful of models for a massive number of users to get that same economies-of-scale effect. Same model, massive number of people.
You mentioned kernel engineering and that being your first job. That has emerged from being something nobody had ever heard of in their lives to now something that you hear about all the time—the importance of it to eke more performance out of the bare, raw metal. When will that just be something that AI does entirely as well? Are humans still the best kernel engineers? Are they doing it with the assistance of AI systems? How far down will humans still be in the loop of designing these things? When will that go away?
Today, it's all very hybrid. The best kernels are still written by human-AI collaborations. Any AI model is built with these fundamental primitives, like matmuls, convolutions, chip-to-chip operations, and collectives. Making these overlap and making them really fast matters enormously.
It's a kernel designer's job to figure out: Where can I overlap? How do I allocate memory? How do I verify that, if there's some issue like a retransmit, it doesn't stall the whole pipeline? These things are very challenging, but they can make your overall performance 3% or 4% better per optimization, and you can do so many.
When we thought about our software stack, we wanted to see where the puck was going to be. 3 years ago, there were kind of two ways you could build software. One of them was to invest heavily in graph compilers. These things are not very performant, but they work out of the box. They don't require a human to come in and tweak all the kernels.
But we went the opposite direction. We are kernel-first programming, and that means that, for a long time, it did not work out of the box. But if you were a kernel expert, you could get incredibly, incredibly high performance.
And the thing about this is that now, as the coding models get better and better, they're doing more and more of the kernel-generation task. When the models keep getting smarter, they'll eventually do all of it. They will become superhuman. So, we're going to build for where the world is going.
Even today, we think about our profiling tools or debugging stack from the perspective of how the model will use these tools, more than how humans will use these tools. We sometimes run experiments internally, and we had Codex actually get GPT-OSS running from scratch, just based off our docs, completely by itself.
Wow.
And they did it, I think, overnight. We think about game selection a lot. What we mean by that is making sure we're investing our energy in the right bets, because regardless of what you choose to work on, it will take tremendous effort.
One of the things that we started with was the explicit decision not to build an arbitrary graph compiler, not to support arbitrary PyTorch, not to support arbitrary CUDA, and not to support arbitrary ONNX graphs. Instead, we envisioned a world where there were going to be under 100 models that actually mattered, and they were all going to look very similar from the underlying mathematical perspective. We were going to build primitives using physics that would accelerate these as much as humanly possible, and we were going to allow the most sophisticated customers to have direct access to the hardware and do whatever they wanted.
That has saved us a tremendous amount of time by not having to build a compiler, and it has allowed us to get much more performance. Funnily enough, when we started, a lot of people dismissed this idea, and the only people who took us seriously were in high-frequency trading because they all hate compilers, too. They all write their own kernels, and we've had dozens of people from high-frequency trading join the team because they saw this philosophy, too.
Mhm. What are the limits to vertical integration? How do you know where to draw the line? I'm starting with this question to talk a bit about the broader market. The circumstances of the broader market are really interesting to me, where the vast majority of AI chips get bought by a very small set of customers. Many of those customers are themselves trying to design their own AI chips.
Yeah.
OpenAI announced Jalapeño. It seems like this very funny circumstance where the most valuable thing in the world all kind of flows through a couple of chip makers and a couple of chip buyers. They all seem to be thinking about doing each other's job. Then you've got the circumstance where these things go in a data center, and then you've got neoclouds and inference providers in this other part of the stack. You've got model builders and providers.
I can imagine a world where, because you have the best hardware, you design models and build data centers—you leak outside of your current vertical. So how do you think about where to draw the lines for the business? You have a saying that production is the product.
Uh-huh. Ultimately, what matters here is that we know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. So all the decisions we make are about how we get the most token capacity online as possible.
Part of that is building a really good product that has way more throughput, that can run at way better latencies and so forth, so we can, per chip we make, get way more tokens online. Another part of it is not doing parts of the stack unless we absolutely have to in order to get to giant scale.
There are parts that we decided to do because it was absolutely required to get the scale, like building the rack instead of just building the chips, and doing a CM model instead of a JDM model. But there are parts of it that are kind of noise to us right now. We're not going and building our own data centers today. That doesn't actually help us get more capacity online.
In general, our customers are actually making power and moving their clusters around to get our chips online because they're such high-throughput. If there was a world where other things were constrained, we would totally go and integrate with them, but the reality is we're just purely focused on getting as many tokens online as possible.
I think it just comes down to economies of scale again: at certain parts of the stack, there are huge economies of scale, and at others there aren't. For example, in designing models, there are huge economies of scale there. For chip fabrication, same story. But if you think about building some small metal part inside of that rack, there's not that same effect. We think the natural boundaries are on the chip side, on the bottom, and at the model layer at the top. We'll fill the whole gap between.
A few weeks ago, there was a guy who was running a next-generation AI chip for one of the frontier companies, and he was trying to recruit one of our architects. This person actually kind of did an Uno reverse card and started recruiting the person trying to recruit our guy, and within a week we hired him.
I was going on a walk as we were finalizing the offer, and I was like, "Why are you leaving this super-important project? Why are you deciding to join?" His answer was super interesting, which was, "It fundamentally is not existential for my company for this product to win. For Google, with TPUs, their revenue comes from search."
Google won't fail if TPUs fail.
That's right. Meta won't fail if MTIA fails. Microsoft won't fail if Maia fails. OpenAI won't fail if Jalapeño fails. Ultimately, this is our product. It is completely unsurprising that the best chip in the world is built by a company that only builds that chip. It's Nvidia.
Right.
And for us, it is completely existential for us to get as much token capacity online as possible. And it recruits a set of talent and recruits support from suppliers and from customers that view it with the level of intensity that we do.
Look at the raw FLOP stats. If you compare any of these chips built by labs or by the hyperscalers, the FLOP density for, say, FP8 × FP8 is lower than the Blackwell B300.
Yep.
And that makes sense because they don't have to take the risk. They just have to build a similar-enough product and not pay the Nvidia tax.
As I think about you guys building the solution, the process of doing so is solving a sequence of really hard challenges. What has been the single episode that was the hardest to overcome?
When we were designing the chip, we built this massive FPGA cluster to verify the full chip workloads. FPGAs are digital entities. You can test digital logic, but not analog logic. It turns out that when the chip came back, we began to see issues in our attention datapath producing incorrect results.
We realized, wait a minute, there's a problem with the backpressure logic across a clock-domain crossing that is failing. This is going to cause the chip to produce wrong results. It is very, very hard to solve.
We realized there was one and only one way to solve it: we had to line up 2 clock signals on our chip to within 50 picoseconds. That's literally 50 trillionths of a second. We had to get the signals aligned to this super-small granularity and do it on every chip 2 billion times a second.
A lot of people said this was impossible. We had people quit.
Yeah.
People literally were like, "This problem is unsolvable, and best of luck, guys."
When you have a problem like that, step 1 is, okay, let's assume the problem is solvable. How would it be solved? First, we realized what we had to be able to do was find a way to move our clock phase by a picosecond, 10 picoseconds. And we had an idea: what if we had these 2 clocks and had them just a little bit apart from each other?
We figured out that if we could figure out the phase, and then use a drifting mechanism to wait for just the right amount of time to get those 50 picoseconds always lined up, we could do this extremely reliably and then lock the phases exactly where they had to be. We could guarantee this would never happen.
People were, I think, somewhat blown away that this worked, and that it worked as well as it actually did. But we made it work.
How long did that take?
This was actually about 2 weeks.
It was a dark 2 weeks.
It was a very scary 2 weeks. But it was the kind of thing where, when that kind of thing happens, that is the most important time to invest effort. That is the hardest time to do it, when you feel like things are hopeless. But the sooner you solve that problem, the sooner you can get back to building and scaling our production to mass volumes.
I think a lot of our story is, as Gavin says, assume it is possible. Assume it is possible to have a chip with way more FLOPs on it. Assume it is possible to have a system with way lower latency between chips. Assume it is possible to create a shared memory pool that can run at way higher bandwidth. How would one do it?
A lot of the time when we do experiments, we'll do dozens of experiments and all of them will fail. But we only need 1 to work. There were multiple times—I mean, Gavin, I think you were leading the charge—during our chip bring-up with, I think, 30 different board experiments, and 3 of them worked.
And all three of them were worth their weight in gold.
This is one of the things that people come to me and say, “Gavin, almost none of your experiments work.” And I’ll say, “I only have to get lucky once.”
7. Raising $100M to Survive
So one idea for one of these stories that I’m asking about—difficult moments in the company’s history—is around the ability to raise capital to fund the thing. I think when you started it, you knew you’d need capital, but you did not know you’d need the quantum of capital that you’ve ultimately raised and are spending to build the solution, and you hadn’t raised money before.
These are all new things, right? And there were moments where it was really, really difficult, because I was there; I saw it. There were moments where it was extremely difficult to raise the money that you did, without which the company would not exist. It would have died.
And like many great stories, there were many near-death moments, but money specifically in this new world—this isn’t software. You don’t just need a little bit of money. Maybe you could tell the story about the true hardest part about raising money early on, before you had something that you could show people and be so proud of, and performance that you could show them and blow their socks off.
It was just you guys talking about an idea. Talk about the early difficulties raising money, because it was pretty hardcore.
We’ve had some intense moments. It reminds me of probably early 2024, before we raised our Series A. We were at this point where we had done enough of the architecture and enough of the design that we knew the chip architecture was sound. We had to go build it.
There was a lot more to do. We were ready to go into what’s called the physical design stage. We needed to sign an agreement with a physical design vendor, which will cost you at least $40–50 million.
And then we had this realization, as the models were getting bigger and bigger and these giant MoE models were coming out, that we were going to need to build the entire cluster—not just the chip. We were going to need to build boards, build interconnects, build cold plates, and figure out all of the networking and everything. This was going to cost a lot more than the $15 million we had in the bank.
And you’re like, man, that was scary. You’re sitting in that moment and you think, “Holy crap, we can’t afford this.” And I began looking at, “How hard is it to go back to Harvard?”
At the end of 2023, we put together this memo. We spent around 100 hours on it because we had no idea how people were going to believe us when we asked for the amount of money we were about to ask for.
It was 30 pages, extremely technical and in-depth, covering all the different things we needed to build, all the milestones we needed to hit, how the market was going to evolve, all the new use cases, the cost per token, and all this modeling.
And then we went and talked to investors, and every major investor in the Valley passed immediately. They were just like, “Okay, two kids that just finished Harvard, haven’t taped out a chip, no test chip. Inference—who knows if this is going to be a big market? Everything’s going to be training. The models still hallucinate. This could all be a bubble.”
At the time, the biggest semiconductor fundraises for a Series A were around $40–50 million. We were looking at this and tallying the bill. We were like, “We think we’re going to spend $100 million in the next 12 months. If we really want to do this, if we want to actually get to scale and actually get the performance we’re talking about, this is going to be extremely capital-intensive. How the hell are we going to pull this off?”
I think one of the key ways we got started in this process was that we thought to ourselves, “What is the cheapest possible way we could do this?” And we decided, “Well, if I made almost nothing—”
Yeah.
And if I ate nothing but ramen, then we would go ahead and spend basically just the money for the mask, let it rip, and make the tape-out, and that would be that.
Right.
And if so, we could do it on $30 million—an obscenely low number. Then we actually went out and got a debt provider. They were willing to lend us the money we needed to get us across this barely ramen-to-a-chip threshold.
From there, I think it was just a matter of catalyzing a series of other steps: “Hey, maybe we can go do one more thing, one more thing, one more thing.”
Yeah, so we’re at this moment where we’re like, if we really want to build this company—because we’re not going to half-ass it. We’re not going to go do a test chip and spend years on it and let the entire AI market boom while we could be building the product. If we’re going to do it, we’re going to go all the way.
We’re going to need to find a way to get $100 million. I remember Gavin and I were sitting down in the office in Cupertino late at night, just looking at each other, and we’re like, “Could we cut $500,000 here? Could we cut $100,000 here? How long could we convince everyone not to take a salary?” The math was not going to close. We really needed to solve this.
There was a period of a few weeks where you just go into survival mode, and you call every person that could possibly know an investor. You’re like, “We need $100 million to do this. If we do this, we think this could be one of the most important companies of all time. Do you know somebody who wants to take an aggressive bet, somebody who wants to believe in us? Here’s all the information. We’re an open book. Here’s the team. They’re great people. We’ve been working super hard. We’ve done these things in record time, but we have these 100 things to go. Do you want to do this?”
The snowball starts, and you get $1 million here and $2 million here, and you’re like, “Okay, we’re not going to run out of money this month.” You get a $5 million check, a $10 million check, and you’re like, “Okay, maybe I can buy those FPGAs.”
And the snowball happened. We were very lucky that we ended up putting it all together. We had a board meeting. I showed you the spreadsheet, and we looked at it, and it was $103 million. These were all soft commits. We all looked at each other and said, “We’re going to take it.” And that was the Series A.
Luckily, I think it’s been much easier since then, and we’ve raised almost half a dozen rounds since then, many of them from those investors just doubling and tripling down. That’s allowed us to get to market so quickly. This rack would not be possible had we not been so aggressive.
I also think suppliers deserve a little bit of a commendation here.
Yeah.
TSMC was willing to work with us back before we’d raised any of the $100 million.
Yeah.
This was back when it was still really, really scary. Synopsys actually went ahead and let us get some of their emulators on extremely favorable terms, where we pay over many years.
Yeah.
Yeah, basically a big loan. And it takes a lot of belief from your partners to go do this. But I do know that you came out with this very strong team, and all the folks who back you are not in it just out of pure financial incentive. They believe.
Why did TSMC believe, do you think?
This is a great story. Even before you joined in, there was a conference, a SEMI event. I was one of the only young CEOs in semiconductors, and I think that’s kind of a novelty. They asked me to come in there and speak.
I get to the SEMI event, and I am the only speaker there—and the only person there—under 30. I was 22 at the time.
22.
So, I go up and speak. There was a speaker's dinner afterward, and by pure luck, I happened to sit next to a very senior TSMC VP. It was a very nice dinner. The former CEO of Arm was there. It was very bougie; everyone was in a suit.
I'm there with this VP, and it turns out we both studied math in college. We both got a little piece of paper and began talking in great detail about how modern AI models work at the actual, tensor-by-tensor level. The guy just gets it. We were talking about, "How do you run this very efficiently? Why is memory such a critical technology to make this work?"
The following day, I get an email from TSMC saying, "Gavin, want to work with Etched? Find a way to make it happen." They've been a great partner ever since.
Crazy.
It's amazing to think about some of the tropes. I obviously should break the fourth wall here: I'm a big Etched investor. I've been involved for a long time, and I think the absolute world of you guys, so I'm incredibly biased in this conversation. I'm trying to ask questions that are broader and interesting and could be objections to what you're doing, and we'll keep doing that.
But it's so interesting to me that when you read about investing, everyone cites this idea of being contrarian and right as the quadrant that makes all the money. It sounds really nice, but contrarian means everyone else thinks you're stupid. When you get immediate no's from literally everybody, it's a fascinating quadrant to exist in before you become consensus.
What was it like for you? I'm super curious.
Well, it's interesting. At the time, it was, by a lot, the largest first check that I'd written. Suffice to say, I'm not a math expert, a semiconductor expert, or really an AI expert at the time.
It was much more about believing in the concept of this market potentially being huge. You had made very, very clear bets on how the future was going to look and positioned the company to attack those things in a hardcore way. The two of you, and what I felt about you, were the majority of the reason why we made the bet when we did, in 2023 or whatever it was. At the time, it was the biggest.
I think the same thing you said about naivete applies to investing as it does to maybe building a semiconductor startup. I didn't know what I didn't know. When I called experts, they were basically like, "This is stupid." They laid out in very logical terms why this wasn't going to work and why it was such a low-probability bet.
I think one of the things I've learned from it is that you have to damn the base rate. If you invested on base rates, you should do something other than—
There's always the index fund.
Yeah, there's always an index fund, exactly. So it's actually never been scary for me. Most of that is probably because I don't know there's a lot I don't know. If I knew more about what you guys have done and the difficulty, I probably wouldn't have done it.
I don't know what that says about maturing as an investor. Maybe I don't want to know a lot more and have some of that healthy naivete. I don't know.
Funny. I think a lot of the traditional semiconductor funds missed the entire AI chip space—all the AI chip companies. All the coding experts missed all the coding companies. I think it's very hard to realize that the constraints have changed.
When you've looked at tape-outs for 20 years and seen so many of them not work on the first try, or the second try, or the third try, you couldn't even run a workload. You totally forget that EDA tools are way better, that FPGAs exist today in a way that they didn't before, and that all the types of validation you can do today just weren't possible before.
For us, a lot of our believers were on 2 sides of it. They were either believers in the market and the team, or they were building chips today and were extremely technical, like the high-frequency trading firms. They would literally audit everything, from the microarchitecture and the RTL to the board designs, the schedule, and the software stack.
We would sit down with 10 of their people who built their own chips, and they would ask us such detailed questions that we were wondering, "Are they going to build the chip?" It was really on either of those sides. If you were anywhere in the middle, you just wouldn't understand it.
In the investing world, they often talk about variant perception: something that you see or believe that others don't.
Yeah.
I think I've invested, I don't know, 5 or so times in that, and every time you do it, the stakes get bigger and bigger. It does get a little scarier and scarier. Because you guys have been so quiet in the marketplace, I think it's very easy to dismiss you. As the stakes get bigger and bigger, those dismissals are harder to hear.
Yeah.
I do think betting on something that you see, when what you hear from the outside world is very different—that perception gap equals opportunity.
Exactly.
The last thing I would say is that the accumulated evidence of you guys and your team's ability to solve seemingly impossible problems is one of the most interesting things a company can have. It's binary: companies do this or they don't.
That's the thing. It's a big advantage of people who have been here for a long time. You get some new joiners who are scared shitless. You see a thing like this, and there are old-timers who have been here for all of 2 years, smoking cigars in the trenches.
Another one.
Yeah, there's definitely a find-a-way mentality. If you're here, you're here because you assume it's possible, so we can't be saying it's impossible. Everything is solvable, and we're just going to work at it until we figure it out.
There's a favorite story I have about this guy who's kind of a legend in silicon validation who joined our team. We were doing the early stages of what's called wafer sort. When your chips are coming out of the fab, they go out on these wafers, and you have this thing called a probe card that attaches to the wafer before you dice it, with these probe pads. You send electrical signals to basically test which chips are good and bad.
When I dice the wafer into a bunch of chips, I can package them and only package the good ones. We go through our first wafer, and it's around 2:00 or 3:00 a.m. because we're doing it with TSMC over the phone in Taiwan. We have the screen with the wafer that's all gray, and each chip is gray. As you start running the patterns, the squares are supposed to turn green or red.
They all turn red. We're like, "Fuck. This is really bad." Everybody's like, "Guys, take a breath." He leans back and says, "The puzzle begins."
I'm like, you have to have the attitude of, "Yes, you will go out and stare into the abyss. You will go see scary things, and we'll solve them."
When did you see the first green square?
It wasn't that day. But in the moment, you're like, "I have worked for years for this. I put my life on the line. I've asked my family to stake everything on it, and then it's red." It's extremely scary.
There's a certain type of person who's just addicted to that feeling of fear and solving it, and we are lucky to have a lot of those people here.
If you think about applying all of this earned know-how from these last several years and now thinking ahead to Gen 2, Gen 3, and beyond—
Sure.
What will you be doing most differently as a result of everything that you've learned? Just from a conceptual standpoint, how will you attack designing and producing this next one based on what you learned doing it the first time?
It took us a while to get to the primitives that we think are really what matters for scaling inference. We tried a bunch of things early on, from compilers that would turn different models into FPGAs, to burning weights in silicon, to splitting your HBM into KV cache and weights, and all of these different things.
There were a lot of cycles of learning until we got to the point where we realized that fundamentally, if you want to run the majority of tokens in the world, you need to do 3 things. You need to build a chip with the most FLOPs in a given power budget. You need to build a chip that has the lowest latency between chips, so the biggest scale-up domain possible. And you need to produce as much of it as possible.
I think probably in the first half of our journey so far, we learned the first 2. That informed the design a lot, and it informs a lot about the bets we're making in the future with low-voltage inference and cluster-scale memory.
But the production part, I think in the past year, has made it extremely obvious how much people want to deploy this stuff if you can have it available today. The best ability is availability. If I have 1,000 chips today, someone's going to use them. We need to build a chip that's not just way better than what's been built before; it needs to be available at many-gigawatt scale.
We need to be able to build a product that is producible at gigawatts per month in the limit. As we think about that, a lot of the design decisions we're making with the next generation, which you've seen already, are just about simplicity. That means removing tons of parts, trying to assemble and disassemble the thing again and again, and learning how to make the cycle times as quick as possible in production. We need to make sure it's going to be reliable, serviceable, and producible at gigantic scales.
What about other problems in the ecosystem that are outside of your control, such as capacity at the leading nanometer nodes at TSMC or availability of HBM4 memory, or some of these other things where everyone is fighting for a scarce unit of capacity or whatever? How do you face up against those realities when you're trying to produce as much as humanly possible?
The people deploying the most compute in the world do think about supply as a bit zero-sum. There are only so many wafers being produced on a given nanometer node, at a given fab, right? And there's only so much memory being produced.
And that's why, actually, for our first-gen product, we built it on a different supply chain than Rubens. We're on 4-nanometer; Rubens are on 3-nanometer. We're on different HBM than Rubens, and so forth.
So it actually is not a zero-sum thing. It's a positive-sum thing where more is more. So often, when we're talking to people deploying at scale, it's not a decision between a gigawatt of a GPU and a gigawatt of us. It's 2 gigawatts. And I think, as much as possible, you need to think about supply chain early in the design decisions, because if you have the most performant product and you can't produce it, then you're just a podcast.
That's the other big thing about vertical integration, too. With certain things, like the chips and the memory, you have to go ahead and partner. For most of the other stuff, those are also very highly in-demand components. And the more that you build yourself, the more stuff you can go do on top of what the world can currently build. It is not, “Oh, you're taking availability from somebody else.” You're adding way, way more. And I think that's how you win.
One of the things we really haven't talked at all about is the models themselves, which is kind of crazy—the things behind all of this demand. Anything interesting that you would say about the way that you see models progressing based on what we've seen so far? I guess I'm more interested in how you, as thinkers about hardware, think hardware might impact where the models themselves go in the future?
One of the most important ideas that we believe in is that machines don't think like people think. You look at airplanes, for example. Airplanes don't fly like birds fly. When you think about how mechanical devices have to work, it's often very different.
And in much the same way, for people, storing data and loading memory is very cheap for neurons, and doing math is relatively expensive. It is the exact opposite for chips. Generally, loading data is very expensive, and doing math is very cheap.
As time goes on, I think you'll end up finding that math gets cheaper at a rate that is faster than memory gets cheaper. This is due to this fundamental limit on any kind of DRAM device. You should go ahead and think about: How can I make my model use a huge, huge amount of compute? What if I had, for example, many copies running at the same time? What if I activated a huge number of experts? What if I had gigantic experts that go ahead and run on multiple server racks at the same time? That is how I think you'll build models that are the next generation of intelligence.
And context, too. There's been a lot of work on very efficient inference. What if I don't load the full context in the memory? Most of the time, I think that makes a lot of sense. You want to go build a superintelligence. Why can't it go look at a billion tokens of context? Why can't it spend a huge amount of compute to go ahead and read all of it super fast? I would love to be able to talk to a machine that was able to attend to every book ever written in its short-term memory. And I think we're going to get to a point where you can.
A theme in models right now is this focus on something called dynamism, which is the ability to control the level of computation and memory spent at a per-token or per-user level when doing attention, as well as the ability to dynamically, in your chip, on the fly, send data to other chips for different MoE models doing certain types of operations.
The reason is fundamentally that, as we're scaling context length, as we're scaling model size, and as we're scaling the amount of computation per user, we're looking for ways to be more efficient. The first thing is, as Gavin says, mixture-of-experts architecture is where maybe we don't need every parameter being used for every token. But maybe there are things where, even at a token level, we can say, “Well, this token needs this context from this other token.” They can share that memory, so we don't have to have the overhead of using the memory as much.
Maybe this token is really important, so we should spend more compute; we should have longer context on that token. So hardware that really accelerates these types of very dynamic computations is extremely important. And, as you can imagine, current hardware that was designed before those types of architectures has lots of overhead in doing them.
You basically end up in these really bad worlds where you have inefficient hardware at doing this dynamism. Therefore, you can't run it very well, or you have these very blocky architectures that are kind of applying blunt force to many different tokens that all need more or less computation.
I have 2 questions about the future. We've talked a lot about what you've built so far and how you've built it. The first is about the new ways that people might start using these systems—the raw technology—over longer runtimes, things of this nature.
When inference gets much cheaper, faster, and more accessible, there's more total supply, and it's better, what are the things that you think people will use that capacity to do that are the most interesting and exciting to both of you?
Yeah, there was a viral tweet by Noam Brown where he said that as these models are having longer and longer time horizons, they can do tasks that take, say, 6 months, and there's often not enough time to go and evaluate them for such a long period of time because by that point you'll have a new model out that you'll want to go evaluate instead. And with technology like our cluster-scale memory, you can go ahead and run that 6-month job much faster.
But there's a second piece of this, too. Talking to Noam about it—he's now an angel as well—it's not just the time; it's also the number of people or agents who are working on this. If you're trying to evaluate, can a human build a rocket? You will find that the answer is no. No one person can go and build a rocket. Instead, you have to go and put a team together. And I believe the same thing will be true of agents, too.
If you want to ask, can an agent go out and build some crazy futuristic piece of software, you'll probably need a very large team. Maybe that's 10; maybe that's a million. You have to go and have this enormous amount of that colossal-scale memory to go ahead and have that very short time per token and a huge amount of FLOPs to be able to run that whole fleet.
I'm going to be a little futuristic. I firmly believe we are on a global march of inference becoming a majority of global GDP. It may take more than 10 years, but it's going to happen.
Right now, we measure productivity as a society as GDP per capita. But really, it's going to look much more like agents per megawatt, or it may be agents per gigawatt by then. And while we're being futuristic, I think this is the second-to-last year where a majority of the workforce is going to be human.
I think in 2027 you're going to see that there are going to be more agents doing knowledge work than humans. And it's going to be extremely interesting to see what happens. You could imagine a world where, for countries, a majority of their energy ends up going into data centers doing inference, and the energy efficiency of those data centers basically governs how many agents and, therefore, how big their workforce is.
So you're going to see, like Gavin is saying, right now we have 1 agent, or a team of 5 to 10 agents, working on group projects for a couple of days. So you can do pretty cool stuff because they're smart, but it's not going to be civilization-scale.
What happens when you have countries that can have literally 1 billion concurrent agents—like 1 billion people in the workforce—working 24/7 concurrently on the same stuff? I mean, it's just kind of unfathomable what's going to happen. And it's going to be the biggest proliferation of technology humanity's ever seen.
I think, well, when you have these huge amounts of demand, you get this idea of economies of scale again.
Yep.
When you think about people, I have a brain. I'm not using the whole thing all at the same time. Only part of it is going to be active, and this is the way healthy brains work. And for MoE models, it works much the same way. In an MoE model, only a small fraction of the parameters is being used for any given token at any given moment.
But if you have a large number of users on a piece of hardware, you can kind of take that brain, cut it up into many different experts on many different servers, and run a huge amount of volume through it.
You'll have a bunch of different pieces of traffic. You'll have many of them using each part of the brain at any given point in time, and you'll have to make the cost per thought, the cost per token, way, way lower. So, I think you're going to end up with these giant-scale distributed brains, where the real form factor is a big data center with a bunch of chips, a huge amount of FLOPs, and a huge amount of scale-up interconnect.
Do you think we'll see a $1 trillion individual data center?
Absolutely. It is a matter of time. It's like asking, will you see a $1 billion fab, or a $10 billion fab, or a $100 billion fab? It is inevitable that the economies of scale don't stop at, “Oh, $40 billion is the magic number for fabs.” No. The cost per wafer keeps going down as you keep spending more money. And the same thing will be true of, say, plants that go out and make steel, or plants that go out and make tokens.
A very smart alien lands on Earth and wants to know from each of you how you would frame up this opportunity that you guys are tackling. What do you say to them?
Frame it as: thinking is really valuable. Every company in the world runs on thinking. And we're entering this really unique moment in time where you have machines that can think almost as well as, and soon as well as, and soon better than, the best humans can.
Building these machines is going to be a huge opportunity. But more important than that, the way in which you run this kind of thinking is going to be very, very different as demand goes higher and higher and higher and higher. There's a unique moment right now to build a new set of solutions, a new roadmap for how you run the future 10^15-parameter models for 1 billion people all at the same time on a gigantic scale-up cluster.
We are in a new era of intelligence where the cost of producing intelligence is dramatically cheaper than the value of the intelligence, so we are in a many-year, probably many-decade supply shortage of these tokens. Basically, any chip or any system that can produce tokens is likely to be extremely valuable, and you should find some part of the supply chain of the token—it can be everything from model training down to what we're doing in the silicon and otherwise—to spend time on and push the frontier. And the companies that are the largest are, frankly, going to be the companies that produce most of the global supply of tokens and own the majority of the supply chain of that token.
And importantly, the people who build systems that, as they get more and more chips put together, get cheaper. The way you want this to scale is not that, oh, if I want to go serve 10 times more tokens, I buy 10 times more servers. It must be some solution where, if I serve 10 times more tokens, then I get some economies-of-scale benefit with my chip's scale-up memory tech that allows me to not charge as much as 10 times more for those next tokens.
What a ridiculously exciting future that you guys are building to enable. When I did this with Gavin last time, I asked him my traditional closing question, so this time I'll ask you. What is the kindest thing that anyone's ever done for you?
During my cancer treatment, there was a big decision I had to make. The doctors came to me and said, “It's time for you to decide. Do you want to get surgery, or do you want to get radiation? Here's the trade-off. If you get surgery, you're more likely to live, but you have to assume you'll never be able to walk again. If you get radiation, you'll be able to walk again, but there's not the same probability that you'll live. You may die. What do you want to do?”
I was 16, and my parents said, “You have to make this decision for yourself.” I thought about it for a long time and decided, “I'm going to do the surgery.” I got the surgery. One of the things they do when you get a tumor resection is something called a necrosis analysis, where they look at all the different cells and say, “Is the cell dead or alive?” Because if you have a bunch of cancer cells that are alive, you have a problem.
They looked at it and said, “You know, you usually want 98%–99% necrosis for us to say you're in the clear. You're below that. You should go get radiation.” There were only a few machines in the world that could actually do the type of radiation I needed. One of them was in Boston. I was in a wheelchair, and I needed to move to Boston for multiple months. Both of my parents decided to move out, drop everything they were doing, and live with me. I'm eternally grateful.
Beautiful. Thanks, guys. Amazing conversation.