[BidClub_]
Latent Space · · 72 分钟

SF Compute:将算力商品化,永久解决 GPU 泡沫

Alessio FanelliswyxEvan Conrad

播客
TL;DR
  • GPU 云无法继承 CPU 云的软件利润率,因为每增加一块加速卡都可能改进模型或带来收入,使买方对价格极度敏感。 假设机群规模为10亿美元,10%的溢价就是1亿美元;Conrad认为,客户即使花5000万美元重做这套软件,仍然划算。“他们根本不在乎你的软件。”

  • CoreWeave 的胜出模式更接近银行或房地产,而不是传统云计算:先与信用可靠的客户签下长期合约,再廉价融资购买硬件。 swyx表示,Microsoft和OpenAI合计贡献了CoreWeave收入的77%,让贷款方更有信心,也降低了CoreWeave的资本成本。把这些金融属性视为糟糕云业务的证据,恰恰误解了Conrad的观点:“GPU完全不是那样。”

  • 经济上更稳定的行业结构,是将 GPU 所有权与其上层软件分开。 CoreWeave可以作为“一家非常、非常优秀的房地产公司” prosper,Modal则无需拥有底层算力也能发展;Conrad对同时经营两类业务的公司的警告非常直接:“把这两种业务混在一起,就会被一枪爆头。”他怀疑,超大规模云厂商转售NVIDIA GPU也会亏钱,因为它们的资本有更高毛利的替代用途。

  • SF Compute 把僵化的 GPU 承诺变成可以转售闲置时段的现货市场,将健康机群的利用率推向100%。 买方可以预订数千张H100一个小时,持续按小时续订,也可以支付合约价与市场价之间的差额,取消更长期的承诺。公司将自己称为“GPU融资的电动工具”,让买方能够组合出自己需要的成本、期限和中断特征。

  • H100 过剩反映的是供给同步涌入,而不是 GPU 需求崩塌;Conrad 初步预计,冬季前后市场可能重新进入短缺状态。 基础设施瓶颈曾让机群延期,随后又集中解决;与此同时,测试时推理可能让算力消耗远超此前由少数消费级模型供应商主导的市场。他反复强调,预测取决于下一代芯片的铺货速度:“我认为自己接近市场,但也只是一个模糊的投机者。”

  • SF Compute 的终局,是建立现货价格指数,支持现金结算的 GPU 期货,让数据中心对冲价格,而不是把风险转嫁给初创公司和VC基金。 Conrad强调,SF Compute目前只是现货市场;未来的期货市场有望降低资本成本,减少对庞大且僵化客户合约的依赖。没有这种对冲,他认为算力义务推高了尚未产生收入公司的估值,并助长了AI融资泡沫:“期货是让整个行业冷静下来的方式。”

  • 让算力具备交易属性,需要的是运营标准化,而不只是一个金融界面。 SF Compute会用Linpack对机群进行约48小时至7天的压力测试,运行主动和被动检查,取得BMC访问权限,执行SLA,并推进自动退款和替代容量机制。这些物理层面的复杂性,也是Conrad不看好点对点GPU网络的原因:如果架构不变,“光速限制真的很难突破”。

摘要 · 为研究而整理的核心内容

1. GPU scaling laws destroy the CPU-cloud pricing playbook

  • Conrad首先区分了Web应用服务与模型训练的不同。Gusto或Rippling这样的公司只需购买足以满足需求的CPU,之后再增加一台服务器也“真的一分钱都赚不到”。它们的容量曲线先上升、随后趋于平坦,因此软件服务能够支撑云业务的利润率。

  • 模型开发者面对的激励恰恰相反。训练存在边际收益递减,但Conrad表示,“总会有收益”:再增加一块GPU,可能改进模型;测试时推理则可以运行更久、提升性能、服务更多客户或降低延迟。因此,每增加一块GPU,都有一条合理路径通向新增收入。

  • 这让采购逻辑从“我怎样才能尽可能便宜地买到所需CPU”,转变为“在固定预算内,我能塞进多少GPU”。结果是,客户相对于自身规模会采购异常庞大的数量,却会对单价更加敏感,而不是更不敏感,因为每节省1美元,就能买到更多有用算力。

  • Conrad给出的数字样本是10亿美元的GPU硬件。10%的软件溢价就是1亿美元,足以花5000万美元内部重做一套替代方案,并把剩余资金留下。GPU云因此不能假设服务会带来CPU基础设施熟悉的50%以上利润率:“只要做得到,他们会立刻关掉它。”

2. CoreWeave monetized credit quality rather than software lock-in

  • Conrad对CoreWeave成功的压缩式解释是:“卖出有锁定效应的长期合约,短期业务基本什么都别做。”理想客户要么预付款,要么足够可信、能够按期付款,从而在硬件融资前消除利用率和折旧的不确定性。

  • swyx表示,Microsoft和OpenAI合计贡献了CoreWeave收入的77%。这种客户集中度带来战略层面的疑问,但也形成了可融资的应收账款:Microsoft不太可能违约,而Microsoft的背书以及OpenAI展现出的增长势头,都让OpenAI的信用状况明显好于一家种子轮前、尚未产生收入的初创公司。

  • 主持人用“它是一家银行”和“它是一家房地产公司”来概括,这在Conrad的框架里并不是贬义。CoreWeave不应被当作转售NVIDIA GPU的高毛利SaaS业务来评价。把它视为拥有大量资本、并与租户签订合约的资产方,其经济模型就更容易理解,也仍然可能是一门“很棒的生意”。

  • Conrad的风险图表以GPU能否卖出和折旧敞口为两条轴,融资成本则横跨两者。来自低风险客户的长期预付承诺处于最佳区域;客户可能违约的长期合约次之;短期限、低价格、且延后收款的销售,则落在利息最高、利润空间最受挤压的位置。

3. Short-term GPU sellers are squeezed from both sides

  • 失败的云业务策略是:前期把价格定得足够高,以覆盖后期折旧,再用软件留住买方。Conrad用每小时5美元的售价对比约1.50美元的GPU底层成本来说明:前期利润必须抵消竞争最终把收入压到成本以下的阶段。

  • 实际上,价格敏感的客户会“一次次把你往下压”,而超大规模云厂商仍有足够的资产负债表空间进一步压低利润率。高价短期业务因此被摧毁,供应商被迫转向低价短合约——恰恰是现金回款下降、融资压力上升的结构。

  • Conrad怀疑Microsoft、AWS和Google转售NVIDIA GPU可能会亏掉一大笔钱,不过他明确将其作为直觉判断。它们可以从小规模实验客户身上赚钱,也可以效仿CoreWeave签长期承诺,但把同样的资本投入自有模型、产品或竞争性芯片,可能获得更高利润率。

4. NVIDIA benefits from a financing intermediary it does not control

  • 当被问及NVIDIA为何不直接创办CoreWeave时,Conrad给出了带有保留的战略解释:NVIDIA会因此与购买其芯片的云厂商竞争。他也提出,控制过多层级可能引发反垄断问题,但明确表示这只是猜测,并认为客户冲突才是更有力的理由。

  • Microsoft理论上可以把这一角色内化,但Conrad推断,NVIDIA不会希望某一家超大规模云厂商吞下过多供给。如果Microsoft、Google、Amazon和Oracle基本成为唯一买家,它们的集中度就会形成买方垄断,让少数客户对NVIDIA定价拥有越来越大的影响力。

  • 更广泛的GPU云群体能让买方彼此竞争资源分配。因此,CoreWeave对NVIDIA的价值不只是转售商,也是另一股需求来源,避免硬件市场收缩成4到5个占主导地位的采购关系。

  • 同一套逻辑支撑着Conrad最尖锐的行业主张:将硬件所有权与软件服务分开。CoreWeave可以靠GPU房地产赚钱,Modal则可以在不拥有机群的情况下构建软件。对于两者合一的运营商,他援引“死神敲响隔壁房门”的说法:能抵消十亿美元硬件敞口的十亿美元软件增值业务,实在太少。

5. SF Compute emerged from an audio lab’s recurring solvency crisis

  • SF Compute最初的目标是训练通用音频和音乐模型,而不是搭建交易所。团队没有筹集通常用于首次训练模型的约5000万美元pre-seed轮融资,因为估值、稀释以及下一轮融资的负担看起来可能致命——尽管Conrad现在认为,当时或许应该融这笔钱。

  • 他们原本预计可以按需或按月租用数千张A100,却被告知供应商要求至少承诺一年。Conrad起初非常愤怒,后来理解了原因:如果客户可以轻易取消,供应商就会独自承担庞大库存和折旧风险。

  • 他们的解决方案是消耗一年合约中的1个月,再转租剩余11个月。问题在于,这几乎是生死线:公司大约每30天要支付50万美元,而银行账户里约有50万美元。“如果我们不能把机群卖满,就会直接破产。”

  • 熬过第一年后,公司意识到自己掌握了一项能力。它开始充当GPU房地产经纪商,把需要6个月的买方撮合进供应商的1年承诺,运营裸金属机群并抽成。这些定制交易积累了买卖双方,最终演变成真正的市场,拥有买价、卖价和交易引擎。

6. Resale liquidity makes otherwise irrational GPU durations possible

  • SF Compute可以把数千张H100报价到低至1小时,尽管没有理性的资产所有者会围绕1小时客户来融资建设机群。这类容量通常来自某个客户先买下更长期限,随后因为代码失败或需求变化,转售未使用的时段。

  • 同一机制也服务于SF Compute曾经面对的那类客户:需要大机群运行1个月,却不愿为接下来的11个月承保。市场汇聚了“零散的流动性小块”,把即将到期或暂时闲置的承诺转化为可用的突发容量,而不要求每个供应商自行承担取消风险。

  • Conrad表示,健康机群的利用率会接近100%,因为价格会持续下降,直到库存出清。InfiniBand网络、交换机、NIC或其他物理组件故障时,实际利用率可能更低;但只要硬件正常运行,闲置时段就有出清价格,而不会继续被困在资产负债表上。

  • 对GPU云供应商来说,SF Compute可以在继续寻找理想长期租户的同时填满闲置机群。长期租户出现后,短期预订就转移到别处。长期买方也可以支付合约价与转售价之间的差额取消合约,在获得灵活性的同时,不把底层风险退回给供应商或贷款方。

7. Today’s H100 glut may precede another conditional shortage

  • Conrad拒绝接受对H100价格下跌的单一、确定性解释。在快速扩张期间,瓶颈并不只在芯片本身——InfiniBand线缆、NIC、发电机或数据中心容量都可能拖延交付,导致承诺的机群错过上线日期。限制因素一旦同时解除,大量机群集中上线,市场便从短缺转为过剩。

  • swyx提到过度下单和初创公司倒闭,但Conrad保留了一个重要区分:供给过剩不等于需求下降。“GPU需求显然比以往任何时候都大”;只是市场增加供给的速度更快。

  • 他的预测刻意保守:市场可能“到冬季前后”重新接近短缺状态,但高度取决于未来芯片的铺货速度。接触到更多市场数据后,他的判断不是更确定,而是更模糊:“我认为自己掌握的信息比很多人更多,而这让我成为一个更模糊的投机者。”

  • 测试时推理可能成为需求加速器。swyx估计,开源AI的规模约为闭源AI的5%;Conrad则回忆,OpenRouter中除Anthropic、Google Gemini、OpenAI等供应商之外的部分,当时只有约10个H100节点。测试时推理可能扩大算力消耗,但最终价格影响仍取决于新芯片供给。

8. Physical locality limits decentralization, while markets widen access

  • Conrad“极度怀疑”家庭或广泛分布式GPU能击败使用InfiniBand或其后继技术、且完全互联的机群。加密货币或许能降低支付摩擦,或补贴冷启动,但代币激励最终会耗尽;它们无法消除延迟,也无法把分散的显卡变成一台紧密耦合的超级计算机。

  • swyx给出了最有力的反例:架构可以通过细粒度混合专家路由、块注意力,或许还有分布在不同硬件上的200个专家来适应。Conrad承认,重新设计的模型可能更适合跨空间并行,但如果算法不变,“光速限制”仍然难以绕开。

  • 灵活的集中式供给,受益者包括传统GPU云最不愿服务的客户。研究生会毕业,资助会到期,项目也会不断更替,但一个拿到10万美元资助的研究者,可能负担得起一次大型突发算力使用,即使无法承受一整年。Conrad提到Standard Intelligence、Phind和Schmidt Futures受资助者,都是这类用户的代表。

  • VC提供的机群解决了类似的信用问题。初创公司很难直接借到5000万美元,而拥有约10亿美元资产的基金或资本合作方可以做到;因此,以股权换算力本质上是信用风险套利。Conrad认为Andromeda据称价值1亿美元的机群时机不错,但称这是一项一次性的优势,如今已经被竞争压平,并不建议盲目复制。

9. Hourly reservations let buyers engineer their own spot exposure

  • 一周价格高于一天或一个月的反常曲线,反映的是当前正在到期的容量。Conrad把即时算力比作过期牛奶:某个时段一旦过去,容量就无法再出售,因此闲置区块会不断降价,直到有人出清。起始日期在一周之后的曲线则更符合常规。

  • SF Compute不把这称为可抢占算力。每个小时都被牢牢预订,但自动化程序可以重新购买下一个小时,并持续重复这一过程。只有当下一小时的价格超过买方设定的上限时,工作负载才会失去容量;中断因此成为明确的定价规则,而不是云供应商不透明的政策。

  • Conrad的例子把最高价格设为4美元,同时以市场远低于这一水平的出清价购买。采用这种策略的历史均价接近每小时1美元,有时约为0.80美元;耐心的任务可以只在低于0.90美元时出价,并优先安排在夜间运行。买方接受价格不确定性,以换取显著更低的平均成本。

  • swyx将这些工具与批量推理联系起来。运营商不必接受一个“半价、24小时内返回”的单一产品,而是可以围绕预期现货供给,金融化设计4小时或8小时的服务水平。这种中间方案可能适合后台代理:任务具有弹性,但仍需在一个工作日的部分时间内完成。

10. GPU futures are intended to hedge risk, not amplify speculation

  • Conrad反复纠正术语:SF Compute目前是线上现货市场,“非常明确地不是期货”或衍生品交易所。最终计划是形成可信的指数价格,再支持现金结算期货,让数据中心锁定收入,而不要求实际交付GPU。

  • 他的经济学判断是,风险构成了算力边际成本的很大一部分。对冲可以稳定数据中心收入、降低融资利率,并减少对刚性客户合约的依赖。“期货是让整个行业冷静下来的方式”,只有在把衍生品单纯理解为投机杠杆时,才显得自相矛盾。

  • 没有期货,数据中心就通过长期承诺把价格和利用率风险转嫁给初创公司;初创公司再通过巨额融资需求把风险转嫁给VC。Conrad认为,这有助于制造巨额的收入前估值,而这些公司不可能全部为LP带回资本:“这就是一个迟早会彻底破裂的泡沫。”

  • 现金结算层仍需要可信的物理交付作为底座。“这些是超级计算机,不是大豆”:现货市场必须运营机群、审计性能,并证明标准化算力确实可用,之后其价格才能成为风险管理工具的锚。

11. Operational standardization is SF Compute’s real exchange infrastructure

  • 机群验收从压力测试开始,通常使用Linpack,可能持续48小时,也可能长达7天。这是对整个组装系统的集成测试,而不只是检查一张NVIDIA显卡。Conrad回忆,曾有一座机群在Linpack测试下发生电压骤降,这立即说明它无法安全承载生产工作负载。

  • 压力测试后,SF Compute会运行贴近真实场景的性能测试,并进行主动和被动监控。被动检查观察线上工作负载,并标记硬件供之后下线;主动检查则利用空闲时段进行侵入式测试。自动退款只能部分奏效,因为“总会出现新的狗屁问题”,未被捕捉的故障模式仍需人工调查。

  • 公司通常会取得BMC访问权限,从而可以远程重启和重装系统,而不是充当被动转售商。其工程师会加入客户的Slack频道,凌晨2点也可能协助排查故障;严格的供应商SLA会传递给买方,随着公司扩大,替代容量和自动按比例退款也会逐步实现。

  • 标准合约对存储及其他可变组件采用“这一规格或更好”的要求,并以硬件白名单提供支持。公司自定义的UEFI shim,可以理解为现代BIOS层,负责下载并启动客户镜像,随后支持Kubernetes、虚拟机或其他软件栈。向上控制裸金属,才能让异构机群足够相似,从而实现审计、定价和交易。

12. An anti-hype culture grew from years of founder volatility

  • SF Compute刻意采用稀疏的自然主题品牌,设定克制的预期:用户看到的是一张朴素页面,随后获得一台“便宜了几百万美元”的超级计算机。Conrad想要的,是与华丽的AI网站相反的体验——后者最终往往只是一个普通SaaS应用,尤其是在SF Compute早期产品本身还很粗糙的情况下。

  • 唯一的例外是旧金山本身。Conrad热情推介雾、桥梁、周边乡野以及“巨量的乐观主义”,用这些意象反击那些只围绕Tenderloin或苦干文化的城市叙事。因此,平静的画面既是产品定位,也是对公司名称所指向城市的真诚推介。

  • 这种克制来自一段充满挫败的创业经历。Quirq运转起来时,心理健康留存率接近于零;Room Service则面对着自行搭建分布式系统的客户;“别死掉”的建议让Conrad在4年里尝试了约40个产品。他说,放弃电子邮件并不是因为问题无解,而是因为自己已经筋疲力尽,而Intercom已经拥有理想的客服客户群。

  • SF Compute最后的招聘介绍映射出其混合型业务。公司正在招聘底层Linux工程师,代码库以Rust为主;同时,一名偏金融科技的工程师负责账本、报告,并承担“别把钱全赔光”的任务。更好的金融基础设施最终意味着更好的供应商和买方价格,也让一个只有10万美元经费的研究者能够短暂参与实验室级别的算力竞争。

Alessio Fanelli

Hey, everyone. Welcome to the Living Space Podcast. This is Alessio, partner and CTO at Decibel, and I’m joined by my co-host, swyx, founder of Small AI.

swyx

Hey. Today, we’re so excited to finally be in the studio with Evan Conrad from SF Compute. Welcome.

Evan Conrad

Hello. How goes it? How are we doing?

swyx

I’ve been fortunate enough to be your friend before you’re famous, and we’ve also hung out at various social things. It’s really cool to see that SF Compute is coming into its own. It’s a significant presence, at least in the San Francisco community, which, of course, is in the name, so you couldn’t help but be.

Evan Conrad

Indeed. Indeed. I think we have a long way to go, but thanks.

swyx

Of course. Yeah.

Evan Conrad

Yeah.

swyx

One way I was thinking about kicking off this conversation is that we’ll likely release this right after the CoreWeave IPO.

Evan Conrad

Uh-huh.

swyx

I was doing some research on you, and you did a talk at The Curve.

Evan Conrad

Yeah.

swyx

I think I may have been viewer number 70. It was a great talk.

Evan Conrad

Thank you.

swyx

More people should go see it: Evan Conrad at The Curve. But we have roughly 3 orders of magnitude more people, and I just wanted to highlight: What is your analysis of what CoreWeave did that went so right for them?

1. CoreWeave Wins With Long Contracts

Evan Conrad

Sell locked-in, long-term contracts and don’t really do much short-term business at all. I think a lot of people had this assumption that GPUs would work a lot like CPUs, and the standard business model of any sort of CPU cloud is that you buy commodity hardware, then you layer on services that are mostly software, and that gives you high margins. Pretty much all your value comes from those services, not really the underlying compute in any capacity.

Because it’s commodity hardware and it’s not actually that expensive, most of that can be on-demand compute. While you do want locked-in contracts for folks, it’s mostly just to de-risk your situation. It helps you plan revenue because you don’t know if people are going to scale up or down. But fundamentally, people are buying hourly, and that’s how your business is structured. You’re going to make 50% margins or higher.

2. Why GPUs Break Cloud Economics

This doesn’t really work in GPUs, and the reason why it doesn’t work is because you end up with super price-sensitive customers. That isn’t necessarily because it’s just way more expensive, though that’s totally the case. In a CPU cloud, you might have, let’s say, $1 million of hardware. In GPUs, you have $1 billion of hardware. Your customers are buying at much higher volumes than you would otherwise expect, and it’s also smaller customers who are buying higher amounts of volume relative to what they’re spending in general.

But in GPUs in particular, your customer cares about the scaling law behind it. If you take Gusto, Rippling, or an HR service like this, when they’re buying from AWS or GCP, they’re buying CPUs and running web servers. Those web servers buy up to the capacity that they need. They buy enough CPUs, and then they don’t buy any more. They don’t buy any more at all.

swyx

Yeah.

Evan Conrad

And—

swyx

You have a chart that goes like this and then flatlines.

Evan Conrad

Correct. It’s a complete flatline. It’s not even an incremental, tiny amount. It’s not like you could just turn on some more nodes and suddenly they would make an incremental amount more money. Gusto isn’t going to make 5% more money. They’re going to make zero—literally zero—money from every incremental GPU or CPU after a certain point.

This is not the case for anyone who is training models, and it’s not the case for anyone who’s doing test-time inference, or inference that scales at test time. Your scaling laws may have some diminishing returns, but there are always returns. Adding GPUs always means your model does actually get better, and that translates into revenue for you.

For test-time inference, you can run the inference longer and get better performance. Or maybe you can run more customers faster and then charge for that. It actually does translate into revenue. Every incremental GPU translates to revenue.

What that means from the customer’s perspective is that you’ve got a flat budget and you’re trying to maximize the number of GPUs you have for that budget. That’s very distinctly different from where Gusto or Rippling might be thinking. They think, “We need this amount of CPUs. How do we reduce the amount of money we’re spending on this to get the same number of CPUs?”

What that translates to is customers who are spending in really high volume, but also customers who are super price-sensitive and don’t give a shit—can I swear on this? Can I swear?

swyx

Yeah, yeah.

Evan Conrad

They don’t give a shit at all about your software, because a 10% difference on $1 billion of hardware is $100 million of value for you. If you have a 10% margin increase because you have great software on your $1 billion of hardware, the customers are that price-sensitive. They will immediately switch off if they can. Why wouldn’t you? You would just take that $100 million and spend $50 million on hiring a software engineering team to replicate anything that you possibly could.

That means the best way to make money in GPUs was to do basically exactly what CoreWeave did: go out and sign only long-term contracts, pretty much ignore the bottom end of the market completely, and maximize your long-term contracts with customers who don’t have credit risk and who won’t sue you—or are unlikely to sue you—for frivolous reasons.

Then, because they don’t have credit risk and they won’t sue you for frivolous reasons, you can go back to your lender and say, “Look, this is a really low-risk situation for us. You should give me prime interest rates. You should give me the lowest cost of capital you possibly can.”

When you do that, you just make tons of money. The problem that I think a lot of people are going to talk about with CoreWeave is that it doesn’t really look like a cloud provider financially. It also doesn’t really look like a software company financially.

swyx

It’s a bank.

Evan Conrad

It’s a bank. It’s a real estate company, and it’s very hard not to be that.

The problem that people have tricked themselves into is thinking that CoreWeave is a bad business. I don’t think CoreWeave is explicitly a bad business. There are kind of 2 versions of the CoreWeave take at the moment. There’s, “Oh my God, CoreWeave is amazing. CoreWeave is this great new cloud provider, competitive with the hyperscalers.”

To some extent, this is true from a structural perspective. They are indeed a real competitor to the cloud providers in this particular category. The other take is, “Oh my gosh, CoreWeave is this horrible business,” and so on and blah, blah, blah. I think it’s just a set of perceptions or perspectives.

If you think CoreWeave’s business is supposed to look like the traditional cloud providers, you’re going to be really upset to learn that GPUs don’t look like that at all. In fact, for the hyperscalers, it doesn’t look like this either. My intuition is that the hyperscalers are probably going to lose a lot of money—and they know they’re going to lose a lot of money—on reselling NVIDIA GPUs, at least.

swyx

Hyperscalers—I want to—

Evan Conrad

Yeah.

swyx

Microsoft, AWS, and Google.

Evan Conrad

Correct, yeah.

swyx

Okay.

Evan Conrad

Microsoft, AWS, and Google—

swyx

Does Google resell? I mean, Google has TPUs, but—

Evan Conrad

Google has TPUs, but I think you can also get H100s from them and so on.

There are 2 ways they can make money. One is by selling to small customers who aren’t actually buying in any serious volume. They’re testing around and playing around, and if they get big, they’re immediately going to do 1 of 2 things. They’re going to ask you for a discount because they’re not going to pay your crazy margin that you have locked into your business. They’re going to pay your massive per-hour price, and so they want you to sign a long-term contract.

Your other way to make money is to basically do exactly what CoreWeave does: have them pay as much as possible up front and lock in the contract for a long time. Or you can have small customers.

But the problem is that for a hyperscaler, selling GPUs on the low margins relative to what your other business—your CPU business—generates is a worse business than what you’re currently doing. You could have spent the same money on those GPUs, trained a model, and then turned that into a product with high margins. Or you could have taken that same money and competed with NVIDIA, cutting into its margin instead.

Simply reselling NVIDIA GPUs doesn’t work like your CPU business, where you’re able to capture high margins from big customers and so on, and then they never leave you because your customers aren’t actually price-sensitive. They won’t switch off if your prices are a little higher.

swyx

You actually had a really nice chart, again, in that talk, this 2-by-2—

Evan Conrad

Sure.

swyx

…of where you want to be, and you also had some hot takes on who's making money and who isn't.

Evan Conrad

Sure.

swyx

So CoreWeave locked up long-term contracts. Got that.

Evan Conrad

Yes.

swyx

Maybe share your mental framework. Just verbally describe it because we're trying to help the audio listeners as well.

Evan Conrad

Sure.

swyx

People can look up the chart if they want to.

Evan Conrad

Sure.

swyx

Okay, so this is a graph of interest rates, and on the y-axis is the probability you're able to sell your GPUs, from 0 to 1.

Evan Conrad

Mm.

swyx

And on the x-axis, it's how much they'll depreciate in cost, from 0 to 1.

Evan Conrad

Ah, yeah.

swyx

And then you had isocost curves, or iso-interest-rate curves.

Evan Conrad

Yeah.

swyx

So they shape in a concave fashion.

Evan Conrad

Yeah.

swyx

The lowest interest rates enable the most aggressive form of this cost curve—

Evan Conrad

Yeah.

swyx

…and the higher interest rates go, the more you have to push out to the top right.

Evan Conrad

Yeah.

swyx

And then you had some analysis of where every player sits in this, including CoreWeave, but also Together and Modal—

Evan Conrad

Yeah.

swyx

…and all these other guys. I thought it was super insightful, so I just wanted you to—

Evan Conrad

Thank you.

swyx

…elaborate.

3. The GPU Risk Framework

Evan Conrad

Basically, it's a graph of risk, and the kinds of places where you can be and what the risk is associated with that. The optimal thing for you to do, if you can, is to lock in long-term contracts that are paid all up front or in a situation in which you trust the other party to pay you over time. So if you're selling to Microsoft or something, or OpenAI—

swyx

Which are together 77% of the revenue of CoreWeave.

Evan Conrad

Yeah. So if you're doing that, that's a great business to be in because the interest rate that you can pitch for is really low because no one thinks Microsoft is going to default. Maybe OpenAI will default, but the backing by Microsoft kind of helps you. Generally, it looks like OpenAI is winning, so you can make a case. It's just a much better case than if you're selling to the pre-seed startup that just raised $30 million or something, pre-revenue.

It's way easier to make the case that OpenAI's not going to default than the pre-seed startup. The optimal place to be is selling to the maximally low-risk customer for as long as possible, and then you never have to worry about depreciation and you make lots of money. The less good place to be is selling long-term contracts to people who might default on you.

If you're not bringing it to the present—in other words, you're not saying, “Hey, you have to pay us all up front”—then you're in this more risky territory.

swyx

So this is the top left of the chart?

Evan Conrad

If I have the chart right, maybe.

swyx

Large contracts paid over time.

Evan Conrad

Yeah, large contracts paid over time is top left, so it's more risky, but you could still probably get away with it. The other opportunity is that you could sell short-term contracts for really high prices. Lots of people tried that too, because this is actually closer to the original business model that people thought would work for cloud providers.

For CPUs, it works, but it doesn't really work for GPUs. I don't think people were trying this because they were thinking about the risk associated with it. I think a lot of people who just come from a software background have not really thought about COGS, margins, inventory risk, or things that you have to worry about in the physical world.

I think they were just copy-pasting the same business model onto GPUs. I also remember fundraising a few years ago, and I know, based on what we knew other people were saying who were in a very similar business to us versus what we were saying, that our pitch was way worse at the time.

In the beginning of SF Compute, we looked very similar to pretty much every other GPU cloud—not on purpose, but accidentally. I know that the correct pitch to give to an investor was, “We will look like a traditional CPU cloud with high margins, and we'll sell to everyone.” That is a bad business model because your customers are price-sensitive.

What happens is, if you sell at high prices—which is the price that you would need to sell at in order to de-risk your loss on the depreciation curve—let's say you're selling at $5 an hour and you're paying $1.50 an hour for the GPU under the hood. It's a little bit different than that, but those are nice numbers: $5 an hour, $1.50 an hour. Great. Excellent.

You're charging a really high price per GPU hour because over time the price will go down and you'll get competed out. What you need is to make sure that you never go under your underlying costs, or, if you do, you've made so much money in the first part of it that the later end of it doesn't matter, because from the whole structure of the deal, you've made money.

The problem is that you think you're going to be able to retain your customers with software, and actually what happens is your customers are super price-sensitive and push you down, and push you down, and push you down, and push you down until they don't care about your software at all.

The other problem that you have is you have really big players, like the hyperscalers, who are looking to win the market. They have way more money than you, and they can push down on margin much better than you can.

If they have to—and they do, though not necessarily all the time; I think they actually probably keep a higher margin—but if they needed to, they could totally wreck your margin at any point and push you down. That meant that that quadrant over there, where you're charging a high price just to make up for the risk, completely got destroyed. It did not work at all for many places because of the price sensitivity and because people could just shove you down.

Instead, that pushed everybody up to the top-right-hand corner of that, which is selling short-term contracts for low prices—

swyx

Paid over time. Yeah.

Evan Conrad

…paid over time, which is the worst financial place to be in because it has the highest interest rate. That means your costs go up at the same time your incoming cash goes down, which squeezes your margins and squeezes your margins.

The nice thing for CoreWeave is that most of their business is over on the other sides of those quadrants—the ones that survived.

swyx

The only remaining question I have with CoreWeave—and I promise I'll get to SF Compute, and I promise this is relevant to SF Compute in general because the framework is important, right?—is to understand the company. So why didn't NVIDIA or Microsoft, both of which have more money than CoreWeave—

Evan Conrad

Yeah.

swyx

…do CoreWeave? Right?

Evan Conrad

Why didn't they do CoreWeave?

swyx

Why have this middleman when either NVIDIA or Microsoft have more money than God, and they could have done an internal CoreWeave, which is effectively a self-funding vehicle, like a financial instrument? Why does there have to be a third party?

Evan Conrad

Your question is, why didn't—

swyx

NVIDIA.

Evan Conrad

…Microsoft—

swyx

Or why didn't either one of those—

Evan Conrad

…NVIDIA just do CoreWeave? Why didn't they just set up their own cloud provider?

swyx

Yeah.

4. Why NVIDIA Avoids Cloud

Evan Conrad

I think—and I don't know, so correct me if I'm wrong, and lots of people will have different opinions here, or, I mean, not opinions. They'll have actual facts that differ from my facts. Those aren't opinions. Those are actually different versions of reality—is that NVIDIA doesn't want to compete with its customers.

They make a large amount of money by selling to existing clouds. If they launched their own CoreWeave, then it would make it much harder for them to sell to the hyperscalers, so they have a complex relationship with them. That's not great for them.

Second is that, at least for a while, I think they were dealing with antitrust concerns or fears. If they own too many layers of the stack, I could imagine that could be a problem for them. I don't know if that's actually true, but that's where my mind would go, I guess. Mostly, I think it's the first one: they would be competing directly with their primary customers.

swyx

Or—

Then Microsoft could have done it, right? That's the other question.

Evan Conrad

Yeah. So Microsoft didn't do it, and my guess is that NVIDIA doesn't want Microsoft to do it. They would limit the capacity because, from NVIDIA's perspective, they don't want to necessarily launch their own cloud provider because that means competing with their customers, but they also don't want only 1 customer or only a few customers.

It's really bad for NVIDIA if you have customer concentration, and Microsoft, Google, Amazon, and Oracle too buy up your entire supply. Then you have 4 or 5 customers who pretty much get to set prices.

swyx

Monopsony.

Evan Conrad

Yeah, a monopsony. The optimal thing for you is a diverse set of customers who are all willing to pay at whatever price, because if they don't, somebody else will. It's really optimal for NVIDIA to have lots of other customers who are all competing against each other.

swyx

Great.

Evan Conrad

Yeah.

swyx

I just wanted to establish that. It's unintuitive for people who have never thought about it, and you think about it all day long.

Evan Conrad

Yeah.

swyx

The last thing I'll call out from the talk, which is kind of cool, and then I promise we'll get to SF Compute, is: Why will DigitalOcean and Together lose money on their clusters?

5. Why Hardware Software Coupling Fails

Evan Conrad

Why will DigitalOcean and Together lose money on their clusters? I'm going to start by clarifying that all of these businesses are excellent and fantastic. Together, DigitalOcean, and Lambda are wonderful businesses that build excellent products.

But my general intuition is that if you try to couple the software and the hardware together, you're going to lose money. If you go out and buy a long-term contract from someone and then layer on services, or buy the hardware yourself, spin it up, and get a bunch of debt, you're going to run into the same problem that everybody else did—the same problem we did and the same problem the hyperscalers are having—which is that you cannot add software and make high margins like a cloud provider can.

You can pitch that to investors, and it'll totally make sense. It's the correct play in CPUs, but there isn't software you could make to make this occur. If you are spending $1 billion on hardware, you need to make $1 billion of software. There isn't $1 billion of software that you can realistically make, and if you do, you're going to look like SAP.

And that's not a knock on SAP. SAP makes a fuck ton of money, right? There just aren't that many pieces of software that you could make where you can realistically sell $1 billion of software. You're probably not going to do it to price-sensitive customers who are spending their entire budget already on compute. They don't have any more money to give you. It's a very hard proposition to do.

Many parties have been trying to do this—buy their own compute—because that's what a traditional cloud does. It doesn't really work for them. You know that meme where there's the Grim Reaper, and he's knocking on the door?

swyx

Mm-hmm.

Evan Conrad

And then he keeps knocking on the next door. We have just seen door after door after door where the Grim Reaper comes by and the economic realities of the compute market come knocking.

The thing we encourage folks to do is, if you're thinking about buying a big GPU cluster and layering software on top, don't. There are so many dead bodies in the wake there. We would recommend not doing that.

Our entire business at SF Compute is structured to help you not do that. It's helped disaggregate these. GPU clouds are fantastic real estate businesses. If you treat them like real estate businesses, you will make a lot of money.

The cloud services you can make on that, all the software you want to make on that—you can do that fantastically if you don't own the underlying hardware. If you mix these businesses together, you get shot in the head. But if you split them, and that's what the market helps you do, you can layer on services but just buy from the market. You can make lots of money.

Companies like Modal that don't own the underlying compute— they don't own it—make lots of money and have a fantastic product. And companies like CoreWeave are functionally really, really good real estate businesses. Lots of money, fantastic product. But if you combine them, you die. That's the economic reality of compute.

swyx

I think it also splits into training versus inference, which are different kinds of workloads.

Evan Conrad

Sure.

swyx

And then, one comment about the price-sensitivity thing before we leave this topic. I want to credit Martin Casado for coining or naming this thing. You said that you don't have room for a 10% margin on GPUs for software.

Evan Conrad

Yep.

swyx

Martin actually played it out further. He's the first one I ever saw doing this at large enough scale. Let's say GPT-4 and o1 both had total training costs of roughly $500 million.

Evan Conrad

Yeah.

swyx

When you get the $5 billion runs—

Evan Conrad

Yes.

swyx

When you get the $50 billion runs, it actually makes sense to build your own chips. For OpenAI to get into chip design—which is so funny, to say, “I would make an ASIC for this run.”

Evan Conrad

Yeah. Maybe. I think a caveat that isn't super well thought about is that only works if you're really confident.

swyx

Yeah.

Evan Conrad

It only works if you really know which chip you're going to use. If you don't, then it's a little harder. In my head, it makes more sense for inference, where you've already established it. But for training, there's so much experimentation.

swyx

You need generality, yeah.

Evan Conrad

Yeah.

swyx

Yeah.

Evan Conrad

The generality is much more useful.

swyx

In some sense, Google is six generations into the CPUs.

Evan Conrad

Yeah.

swyx

Yeah. Okay, cool. Maybe we should go into SF Compute now.

Evan Conrad

Sure, yeah.

Alessio Fanelli

You kind of talked about the different providers. Why did you decide to go with this approach? And maybe talk a bit about how the market dynamics have evolved since you started the company.

6. SF Compute Becomes A Market

Evan Conrad

Originally, we were not doing this at all. We were definitely forced into this to some extent. SF Compute started because we wanted to train models for music and audio in general. We were going to do a sort of generic audio model at some point, and then we were going to do a music model at some point. It was an early company, and we didn't really scope down to a particular thing.

The first thing that you do when you start any AI lab is go out and buy a big cluster. What we had seen everybody else do was go out, raise a really big round, and then get stuck. If you raise the amount of money that you need to train a model initially—like, you know, the $50 million pre-seed, pre-revenue—your valuation is so high, or you get diluted so much, that you can't raise the next round. That's a very big ask to make.

I also feel like we just felt we couldn't do it. We probably could have, in retrospect, but I think, first, we didn't really feel like we could do it. Second, it felt like if we did, we would be stuck later on. We didn't want to raise a big round.

Instead, we thought surely by now we would be able to go out to any provider and buy what a traditional CPU cloud would offer you—buy on demand, or buy for a month or so. This worked for small, incremental things, and I think this is what we were basing it off. We just assumed we could go to Lambda or something and buy thousands of A100s at the time.

This was not at all the case. We started doing all the sales calls with people, and we said, “Okay, can we just get month-to-month? Can we get 1 month of compute or so?” Everyone told us at the time, “No, you need to have a 1-year contract or longer, or you're out of luck. Sorry.”

At the time, we were pissed off. We were like, “Why won't anybody sell us a month at a time?” Nowadays, we totally understand why. If they had sold us month-to-month and we canceled, they would have massive risk. The optimal thing to do was to completely abandon this section of the market.

We didn't like that. Our plan was that we were going to buy a 1-year contract anyway. We would use 1 month, and then sublease the other 11 months. We were locked in for a year, but we only had to pay for each individual month.

We did this, but then immediately we said, “Oh, shit, now we have a cloud provider, not a model-training company.”

Alessio Fanelli

Right.

Evan Conrad

Every 30 days, we owed about $500,000, and we had about $500,000 in the bank. That meant that every single month, if we did not sell out our cluster, we would just go bankrupt.

That's what we did for the first year of the company. When you're in that position, you try to think, “How in the world do you get out of that position?” What that transitioned to was: We tend to be pretty good at selling this cluster every month, because we haven't died yet. What we should do is basically be a broker for other people, and be more like a GPU real estate agent, or a GPU realtor.

We started doing that for a while. We would go to other people who had a 1-year contract they were trying to sell to somebody, and we'd go to another person who maybe wanted 6 months, while somebody else wanted 6 months, and we'd combine all these people together to make the deal happen.

And we'd organize these one-off bespoke deals that looked like... Basically, it ended up with us taking a bunch of customers, assigning them to a vendor, taking some cut, and then operating the cluster for people, typically with bare metal. We were doing this, but it was definitely an “Oh, shit. Oh, shit. Oh, shit. How do we get out of our current situation?” response, and less of a strategic plan of any sort.

While we were doing this, since the beginning of the company, we had been thinking about how to buy GPU clusters and how to sell them effectively, because we'd seen every part of it. What we ended up with was a book of everybody who was trying to buy and everyone who was trying to sell, because we were these GPU brokers. That turned into what is today SF Compute, which is a compute market that we think is functionally the most liquid GPU market of any capacity.

Honestly, I think we're the only thing that is actually a real market, with bids and asks and a trading engine that combines everything. I think we're the only place where you can do things that a market should be able to do. You can go on SF Compute today and get thousands of H100s for an hour if you want, and that's because there is a price for thousands of GPUs for an hour. That's not something you can reasonably do on any other cloud provider, because nobody should realistically sell you thousands of GPUs for an hour. They should sell them to you for a year or so on.

But one of the nice things about a market is that you can buy the year on SF Compute, and then, if you need to sell back, you can sell back as well. That opens up all these little pockets of liquidity where somebody who's just trying to buy some burst capacity for a little bit of time can do so. People don't normally buy for an hour—that's not actually a realistic thing—but that's the range.

Somebody who is like us, who needed to buy for a month, can actually buy for a month. They can place the order, and there is actually a price for that. It typically comes from somebody else who's selling back—somebody who bought a longer-term contract for some period of time, whose code doesn't work, and who now needs to sell off a little bit.

swyx

What are the utilization rates at which a market like this works? What do you see as the usual GPU utilization rate, and at what point does the market get saturated?

7. GPU Utilization And Demand

Evan Conrad

Assuming there aren't hardware or software problems, the utilization rate is near 100%, because the price dips until utilization is 100%. The price actually has to dip quite a lot for utilization not to be 100%. That's not always the case because you have logistical problems. You get a cluster and parts of the InfiniBand fabric are broken, and there's some issue with a switch somewhere, so you have to take some portion of the cluster offline.

There are just underlying physical realities of the clusters. Nominally, we have better utilization than basically anybody, but that's utilization of the cluster. That doesn't necessarily translate into... Well, I actually do think we have much better overall money made for our underlying vendors than basically anybody else.

We work with the other GPU clouds. The basic pitch to the other GPU clouds is, one, we're still your broker, so we can find you the long-term contracts at the prices that you want. But meanwhile, your cluster is idle, and for that, we can increase your utilization and get you more money because we can sell that idle cluster for you.

Then, the moment we find the longer-term, bigger customer and they come on, you can kick off those people and then go to the other ones. You get the mix of selling your cluster at whatever price you can get on the market, and then selling your cluster at the big price that you want for a long-term contract, which is your ideal business model.

The benefit of the whole thing being on the market is that you can pitch your customer that they can cancel their long-term contract, which is not something you can reasonably do if you are just the GPU cloud. If you're just the GPU cloud, you can never cancel your contract because that introduces so much risk that you would otherwise not get your cheap cost of capital or whatever.

But if you're selling it through the market or you're selling it with us, then you can say, “Hey, look, you can cancel for a fee.” That fee is the difference between the market price and the price that they paid, which means that if they cancel, you have the ability to offer that flexibility, but you don't have to take the risk of it. The money's already there and you got paid; it's just being sold to somebody else.

swyx

One of our top pieces from last year was talking about the H100 glut from all the long-term contracts that were not being fully utilized and were being put onto the market.

You have $1-per-hour contracts on here, and it goes up to $2. Actually, I think you were involved. You were obliquely quoted in that article. I think you remember.

Evan Conrad

Yes, I remember.

swyx

Because this was hidden. Well, we hid your name, but then you were like, “Yeah, it's us.”

Evan Conrad

Yeah.

swyx

Could you talk about the supply and demand of H100s? Was that just a normal cycle? Was that a supercycle because of all the VC funding that went in in 2023? What was that? GPU prices have come down.

Evan Conrad

Yeah. GPU prices have come down.

swyx

Some part of that is a normal depreciation cycle. Some part of that is just that there were a lot of startups that bought GPUs and never used them, and now they're lending them out, and therefore you exist.

Evan Conrad

There are a lot of theories as to why this happened. I dislike all of them because they're often said with really high confidence, and I think the market is much more complicated than that. Everything I'm going to say is very hedged.

There were a series of places where a bunch of the orders were placed, and people were pitching to their customers, their investors, and the broader market that they would arrive on time. That is not how the world works. Because there was such a quick build-out of things, you would end up with bottlenecks somewhere in the supply chain that had nothing necessarily to do with the chip.

It's the InfiniBand cables or the NICs or whatever, or you need a bunch of generators, or you don't have data center space. There's always some bottleneck somewhere else. A lot of the clusters didn't come online within the period of time that people expected. But then all the bottlenecks got sorted out, and they all came online at the same time.

I think you saw a shortage because the supply chain got hard, and then you saw an increase, or a glut, because the supply chain eventually figured itself out.

swyx

Specifically, people overordered in order to get the allocations that they wanted. Then they got the allocations, and then they went under. Yeah, whatever, right? There was just a lot of shenanigans.

Evan Conrad

A caveat to this is that every time you say somebody overordered, there's an assumption that the problem was that demand went down.

swyx

Uh-huh.

Evan Conrad

I don't think that's the case at all, and I want to clarify that. It definitely seems like there's more demand for GPUs than there ever was. It's just that there is also more supply.

At the moment, I think there is still functionally a glut. But the difference that I think is happening is mostly the test-time inference stuff: you just need way more chips for that than you did before.

Whenever you make a statement about the current market, people take your words and assume that you're making a statement about the future market. If you say there's a glut now, people will continue to think there's a glut. But I think what is happening at the moment—my general prediction is that, by the winter, we will be back toward a shortage.

This also very much depends on the rollout of future chips, and that comes with its own... I think I'm trying to give you a good—here's Evan's forecast—but I don't know if my forecast is very—

swyx

Okay.

You don't have to. Nobody's going to hold you to it, but I think people want to know what's true and what's not. There's a lot of vague speculation from people who are not that close to the market, actually, and you are.

Evan Conrad

I think I'm close to the market, but I'm also a vague speculator. I think there are a lot of highly confident speculators, and I am indeed a vague speculator. I think I have more information than a lot of other people, and this makes me a vaguer speculator, because I feel less certain or less confident than I think a lot of other people do.

The thing I do feel reasonably confident about saying is that test-time inference is probably going to significantly expand the amount of compute that is used for inference. A caveat to this is that pretty much all the inference demand is in a few companies.

A good example is that a lot of biotech and pharma companies were using H100s to train biological models. They would buy thousands of H100s for training and then not a lot of hardware for inference—not relative to OpenAI or Anthropic, because they don't have a consumer product. Their inference event, if they can do it right, is really just 1 event that matters. Obviously, I think they're going to run in batch, and they're not literally going to run just 1 inference event, but the one that produces the drug is the important one.

I'm dumb and I don't know anything about biology, so I could be completely wrong here. But my understanding is that's kind of the gist.

swyx

I can check that for you.

Evan Conrad

You can check that for me.

My understanding is that the one that produces the sequence that is the drug that cures cancer or whatever—that's the important deal. A lot of models look like this, where they're more enterprise use cases. Prior to something that looks like test-time inference, you got lots and lots of demand for training, and then it pretty much entirely fell off for inference.

I think we looked at OpenRadar, for example. The entirety of OpenRouter that wasn't Anthropic, Google Gemini, OpenAI, or something was about 10 H100 nodes. That's not that much. It's not that many GPUs to service that entire demand, but that's a really sizable portion of the open-source market. The actual amount of compute needed for it wasn't that much.

If you imagine what OpenAI needs for GPT-4, it's tremendously big, but that's because it's a consumer product that has almost all the inference demand.

swyx

Yeah, that's a message we've had: roughly, open-source AI compared to closed AI is 5%.

Evan Conrad

Yeah, it's super small.

swyx

Super small.

Evan Conrad

It's super small. Test-time inference changes that quite significantly, so I expect that to increase our overall demand. But my question about whether or not that actually affects your compute price depends entirely on how quickly we roll out the next chips.

swyx

The way that you burst is different for test time.

Evan Conrad

Yeah.

Alessio Fanelli

Any thoughts on the third part of the market, which is the more peer-to-peer, distributed market? Some of them are crypto-enabled, like Hyperbolic, Prime Intellect, and all of that. Where do those fit? Do you see a lot of people wanting to participate in a peer-to-peer market, or, because of the capital requirements, does it not really matter at the end of the day?

8. Distributed GPU Markets Fall Short

Evan Conrad

I'm wildly skeptical of these, to be frank.

swyx

The dream is sitting at home, right? I have this 4090. Nobody has 4090s. I can rent it out.

Evan Conrad

Yeah, I just don't think this is ever going to be more efficient than a fully interconnected cluster with InfiniBand or whatever the next spec might be. I could be completely wrong, but the speed of light is really hard to beat. Regardless of whatever you're using, you just can't get around that physical limitation.

You could imagine a decentralized market that still has a lot of places where there's co-location, but then you'd get something that looks like SF Compute. That's what we do. Our general take on SF Compute is that you're not buying from random people. You're buying from the other GPU clouds, functionally. You're buying from data centers that are the same genre of people that you would work with already, and you can specify, “I want all these nodes to be co-located.” I don't think you're really going to get around that.

I think I buy crypto for the purposes of transferring money. The financial system is quite painful, and so on. I can understand its uses to incentivize an initial market or try to get around the cold-start problem. We've been able to get around the cold-start problem just fine, so we didn't actually need that at all.

What I do think is totally possible is that you could launch a token and subsidize the compute prices for a bit. Maybe that will help you.

swyx

I think that's what Nus is doing.

Evan Conrad

Yeah, I think there are lots of people trying to do things like this, but at some point that runs out.

swyx

So I would generally agree. I think the only thread in that model is a very fine-grained mixture of experts that can be—

Evan Conrad

Yeah.

swyx

—where the algorithms can shift to adapt to hardware realities. The hardware reality is, okay, it's annoying to do large co-located clusters, so we'll just redesign attention or whatever in our architecture to distribute it more.

Evan Conrad

Yeah.

swyx

There was a little bit of buzz about block attention last year that Strong Compute made a big push on. But I think in a world where we have 200 mixture-of-experts in an MOE model, it starts to be a little bit better.

Evan Conrad

I don't disagree with this. I can imagine a world in which you've redesigned it to be more parallelizable across space. But without that, your hardware limitation is your speed-of-light limitation, and that's a very hard one to get around.

Alessio Fanelli

Any customers or stories that you want to shout out about—maybe things that wouldn't have been economically viable otherwise? I know there's some sensitivity on that, but—

9. Compute For Underfunded Researchers

Evan Conrad

My favorites are grad students—folks who are trying to do things that would normally otherwise require the scale of a big lab. Grad students are the worst possible customers for traditional GPU clouds because they will immediately churn if you sell them something: they're going to graduate and not go anywhere. Or they're not going to—

Speaker 0

Mm-hmm.

Evan Conrad

That project isn't continuing to spend lots of money. Sometimes it does, but not if you're working with the university or a lab of some sort.

The ability for us to offer big burst capacity is lovely and wonderful. It's one of my favorite things to do because all those folks look like we did. I have a special place in my heart for young hackers, young grad students, and researchers who are trying to do the same kind of thing that we are doing.

swyx

Mm-hmm.

Evan Conrad

I also have a special place in my heart for startups—the people who are actively trying to compete on the same scale but can't afford it time-wise, but can afford it—

swyx

Money, yeah.

Evan Conrad

Spike-wise.

swyx

Yeah. I liked your example of, “I have a grant of $100,000, and it's expiring. I've got to—”

Evan Conrad

Cut. Cut.

Speaker 1

Spend it on that.

Evan Conrad

Yeah.

Speaker 1

That's really beautiful, and I hope interesting. Has there been interesting work coming out of that? Anything you want to mention?

Evan Conrad

From a startup perspective, Standard Intelligence and Phind—P-H-I-N-D—and then, from a grad students' perspective, we worked a lot with Schmidt Futures grantees of various sorts. My fear is that if I talk about their research, I'll be completely wrong to an almost insulting degree because I'm very dumb. But yeah.

swyx

I think one thing that's maybe also relevant, from a startups-and-GPUs perspective, is that there was a brief moment when it kind of made sense for VCs to provide GPU clusters. Obviously, you worked at AI Grant—

Evan Conrad

Yeah.

Speaker 1

—which set up Andromeda, which is supposedly a $100 million cluster.

Evan Conrad

Yeah, I can explain why that's the case, or why anybody would think that would be smart, because I remember that before any of that happened, we were asking for it to happen.

Speaker 1

Yeah.

Evan Conrad

The general reason is credit risk.

Speaker 1

Again, it's a bank. I have higher—

Evan Conrad

Yeah.

Speaker 1

I have lower risk than you. I do the credit transformation. I take your risk onto my balance sheet.

Evan Conrad

Correct. Exactly. If you wanted to set up a GPU cluster, you had to be the one who actually bought the hardware, racked and stacked it, and co-located it somewhere with someone. Functionally, it was on your balance sheet, which meant you had to get a loan, and you cannot get a loan for $50 million as a startup—not really.

You can get venture debt and stuff, but it's very, very difficult to get a loan of any serious size for that. It's not that difficult to get a loan for $50 million if you already have a fund or already have, like, $1 billion in assets somewhere. Or you can personally guarantee it or something. If you have a lot of money, it's way easier for you to get a loan than if you don't have a lot of money.

And so the hack of a VC or some capital partner offering equity for compute is always some arbitrage on the credit risk.

Speaker 1

That's amazing.

Evan Conrad

Yeah.

Speaker 1

That's a hack. You should do that.

Evan Conrad

I don't think people should do it right now. I think it made sense at the time, and it was helpful and useful for the people who did it at the time, but I think it was a one-time arbitrage because now there are lots of other sources that can do it. And also, I think it made sense when no one else was doing it, and you were the only person who was doing it.

Speaker 1

Yeah.

Evan Conrad

But now it's an arbitrage that gets competed down.

Speaker 1

Sure.

Evan Conrad

So I don't know if it's super effective. I wouldn't totally recommend it. It's great that Andromeda did it. But the marginal increase of somebody else doing it is not super helpful.

Speaker 1

I don't think that many people have followed in their footsteps. I think maybe Andreessen did it.

Evan Conrad

Yeah.

swyx

That's it.

Evan Conrad

I think just because pretty much all the value flows to Andromeda. I think the—

swyx

What? That cannot be true.

Evan Conrad

I think you had to do it—

swyx

How many companies are in AI Grant?

swyx

Like 50.

Evan Conrad

My understanding of Andromeda is it works with all the NFDG companies, or several of the NFDG companies. But I might be wrong about that again. Nat, don't kill me. I could be completely wrong. But—

swyx

It's perfect. His timing is impeccable.

Evan Conrad

Timing, yeah. Nat and Daniel are— There are lots of people who are—

swyx

Seer?

Evan Conrad

Yeah, Seer, like S-E-E-R.

swyx

Oh, Seer.

Evan Conrad

Like seers of the—

swyx

Oh, Sears, the mall.

Evan Conrad

—of the Valley. For years and years before the ChatGPT moment or anything, they had fully understood what was going to happen. Way, way before. AI Grant is like 5 years old, 6 years old, or something like that—7 years old when it first launched or something.

swyx

It depends where you start: the nonprofit version.

Evan Conrad

Yeah.

swyx

Yeah.

Evan Conrad

The nonprofit version was—

Speaker 1

Yeah.

Evan Conrad

—happening for a while, I think. It's been going on for quite a bit of time. And then Nat and Daniel are the early investors in a lot of the early AI labs of various sorts. They've been doing this for a bit.

swyx

I was looking at your pricing yesterday. We were kind of talking about it before, and there's this weird thing where 1 week is more expensive than both 1 day and 1 month. What are some of the market pricing dynamics? To somebody that is not in the business, this looks really weird. But I'm curious if you have an explanation for it that looks normal to you.

Evan Conrad

Yeah. The simple answer is that preemptible pricing is cheaper than non-preemptible pricing, and the same economic principle is the reason why that's the case right now. That's not entirely true on SF Compute. SF Compute doesn't really have the concept of preemptible; instead, what it has is very short reservations.

You go to a traditional cloud provider and say, "Hey, I want a reserved contract for a year." We will let you do a reserved contract for 1 hour, which is the part of SF Compute. But what you can do is just buy every single hour continuously. You're reserving just for that hour, and then the next hour you reserve just for that next hour. This is built in; this is an automation that you can use.

What you're seeing when you see the cheap price is somebody who's buying the next hour, but maybe not necessarily buying the hour after that. So if the price goes up too much, they might not get that next hour. The underlying part of this, where that's coming from in the market, is you can imagine day-old milk, or milk that's about to be old. It might drop its price until it's expired, because nobody wants to buy milk that's in the past, or maybe you can't legally sell it.

Compute is the same way. No, you can't sell a block of compute that is in the past. And so what you should do in the market, and what people do, is take a block of compute and then drop it and drop it and drop it and drop it to a floor price right before it's about to expire. They keep dropping it until it clears. And so anything that is idle drops until some point.

If you go on the website and set that chart to 1 week from now, what you'll see is much more normal-looking curves. But if you say, "Oh, I want to start right now," that immediate instant—here's the compute that I want right now—is functionally the preemptible price. It's where most people are getting the best compute prices.

The caveat of that is you can do really fun stuff on SF Compute if you want, because it's not actually preemptible. It's reserved, but only reserved for an hour. The optimal way to use SF Compute is to buy on the market price but set a limit price that is much higher.

You can set a limit price for $4 and say, "Oh, if the market ever happens to spike up to $4, then don't buy. I don't want to buy at that price for that hour. But otherwise, just buy at the cheapest price." If you're comfortable with the volatility of it, you're actually going to get really good prices, close to $1 an hour or so, sometimes down to $0.80 or whatever.

swyx

You said $4, though.

Evan Conrad

Yeah. So that's the thing.

swyx

You want to lower the limit?

Evan Conrad

So $4 is your max price. $4 is where you basically want to pull the plug and say, "Don't do it," because the actual average price—or the preemptible price—doesn't actually look like that. What you're doing when you're saying $4 is, "Always, always, always give me this compute. Continue to buy every hour. Don't preempt me. Don't kick me off, and I want this compute."

Speaker 1

Okay, okay.

Evan Conrad

Just buy at the preemptible price, but never kick me off. The only times you get kicked off are if there is a big price spike. Let's say 1 day out of the year there's a $4-an-hour price because of some weird fluke or something. If during other periods of time you're actually getting a much lower price, then you—

Alessio Fanelli

It makes sense.

Evan Conrad

—it makes sense. Your average cost that you're actually paying is way better, and your trade-off here is you don't literally know what price you're going to get, so it's volatile. But your actual average historically has been—everyone who's done this has gotten wildly better prices. And this is one of the clever things you can do with the market. If you're willing to make those trade-offs, you can get a lot of really good prices.

You can also do a bunch of other things, like only buy at night, for example. The price goes down at night, and so you can say, "Oh, I want to only buy if the price is lower than $0.90." If you have some long-running job, you can make it only run at $0.90, then you recover back and so on. But—

Alessio Fanelli

Yeah. So what you can kind of create as a spot instance is what the CPU world has.

Evan Conrad

Yes.

Alessio Fanelli

But you've created a system where you can kind of manufacture the exact profile that you want.

Evan Conrad

Exactly.

Alessio Fanelli

That is not just whatever the hyperscale is offering you, which is usually just one thing.

Evan Conrad

Correct. SF Compute is like the power tool of GPU financing.

Alessio Fanelli

The underlying primitives of hourly compute are there.

Evan Conrad

Correct.

Alessio Fanelli

Yeah, it's pretty interesting. I've often asked OpenAI: all these guys, Claude as well, they do batch APIs.

Evan Conrad

Yep.

Alessio Fanelli

So it's half off of whatever your thing is.

Evan Conrad

Yeah.

Alessio Fanelli

And the only contract is, "We'll return it in 24 hours."

Evan Conrad

Sure.

Alessio Fanelli

Right? And I was like, 24 hours is good, but sometimes I want 1 hour. I want 4 hours. I want something. And so based off of SF Compute's system, you can actually kind of create that kind of guarantee—

Evan Conrad

Totally.

Alessio Fanelli

—that would be, you know, not 24, but—

Evan Conrad

Mm-hmm.

Alessio Fanelli

—within 8 hours, within 4 hours, like half a workday—

Evan Conrad

Yes.

Alessio Fanelli

—I can return your result to you. And if your latency requirements are that low, actually, it's fine.

Evan Conrad

Yes.

Alessio Fanelli

And—

Evan Conrad

Correct.

Alessio Fanelli

Yeah.

Evan Conrad

You can carve that out. You can financially engineer that on SF Compute.

Alessio Fanelli

Yeah.

Evan Conrad

Yeah.

Alessio Fanelli

I mean, I think that unlocks a lot of agent use cases that I want—

Evan Conrad

Mm-hmm.

Alessio Fanelli

—which is, yeah, it works in the background, but I don't want you to take a day.

Evan Conrad

Yeah.

Alessio Fanelli

Take—

swyx

Take a couple of hours or something.

Evan Conrad

Yeah.

swyx

This touches a lot of my background because I used to be a derivatives trader.

Evan Conrad

Yeah.

swyx

This is a forward market.

Evan Conrad

Yeah.

swyx

A future is a forward market, whatever you call it.

Evan Conrad

Not a future. Very explicitly not a future.

swyx

Not yet a futures market. Yes.

Evan Conrad

Yeah.

swyx

We can talk about that one.

Evan Conrad

Yeah.

swyx

But I don't know if you have any other points to talk about. You recognize that you are a marketplace, and you've hired—I met Alex Epstein at your launch event.

Evan Conrad

Mm-hmm.

swyx

You're building out the financialization of GPUs.

Evan Conrad

Mm-hmm.

swyx

So part of that's legal.

Evan Conrad

Mm-hmm.

swyx

Part of that is listing on an exchange.

Evan Conrad

Yep.

swyx

Maybe you're the exchange. I don't know how that works. But just talk to me about that. From the legal and standardization perspective, where is this all headed? Is this fully listed on the Chicago Mercantile Exchange or whatever?

10. Building A GPU Futures Market

Evan Conrad

What we're trying to do is create an underlying spot market that gives you an index price that you can use. Then, with that index price, you can create a cash-settled future. With a cash-settled future, you can go back to the data centers and say, “Lock in your price now and de-risk your entire position,” which lets you get cheaper cost of capital and so on.

We think that will improve the entire industry because the marginal cost of compute is the risk, as shown by that graph in basically every part of this conversation. It's risk that causes the price to be all sorts of funky, and we think a future is the correct solution to this. So that's the eventual goal.

Right now, you have to make the underlying spot market in order to make this occur. To make the spot market work, you actually have to solve a lot of technology problems. You really cannot make a spot market work if you don't run the clusters, if you don't have control over them, and if you don't know how to audit them, because these are supercomputers, not soybeans. They have to work in a way that it's just a lot simpler to deliver a soybean than it is to deliver the—

swyx

I know. Talk to the soybean guys.

Evan Conrad

Sure.

swyx

You know.

Evan Conrad

Yeah, but you have to have a delivery mechanism. Somebody somewhere has to actually get the compute at some point, and it actually has to work. It is really complicated.

That is the other part of our business: we go and build a bare-metal infrastructure stack, and then we also do auditing of all the clusters. You sort of de-risk the technical perspective, and that allows you to eventually de-risk the financial perspective. That is the pitch of SF Compute.

swyx

Yeah. I'll double-click on the auditing of the clusters.

Evan Conrad

Yep.

swyx

This is something I've had conversations with Yitay. He started Reka, and he had a blog post that kind of shone a light on how unreliable some clusters are versus others.

Evan Conrad

Correct. Yep.

swyx

Sometimes you have to season them and age them a little bit to find the bad cards.

Evan Conrad

Correct. You have to burn them in. Yep.

swyx

So what do you do to audit them?

Evan Conrad

There's a burn-in process, a suite of tests, and then active checking and passive checking. The burn-in process is where you typically run Linpack. Linpack is a bunch of linear algebra equations, and you're stress-testing the GPUs.

swyx

This is a proprietary thing that you wrote?

Evan Conrad

No, no, no. Linpack—

swyx

Oh, is it? Okay.

Evan Conrad

Linpack is the most common form of burn-in. If you just type in “burn-in,” typically when people say burn-in, they literally just mean Linpack. It's an NVIDIA reference version of this.

swyx

And again, NVIDIA could run this before they ship, but now the customers have to do it. It's annoying.

Evan Conrad

You're not just checking for the GPU itself. You're checking the whole component, all the hardware, and its location. It's an integration test.

swyx

Yeah.

Evan Conrad

Yeah. What you're doing when you're running Linpack, or burn-in in general, is stress-testing the GPUs for some period of time—48 hours, for example, maybe 7 days or so—and you're just trying to kill all the dead GPUs or any components in the system that are broken.

We've had experiences where we ran Linpack on a cluster and it browns out. It sort of comes offline when you run Linpack. This is a pretty good sign that maybe there is a problem with this cluster. Linpack is the most common standard test.

Beyond that, we have a series of performance tests that replicate a much more realistic environment as well. Assuming Linpack works at all, you run the next set of tests. While the GPUs are in operation, you're also doing active tests and passive tests.

Passive tests are things that are running in the background while somebody else is running, while some other workload is running. Active tests are during idle periods, when you're running some sort of check that would otherwise interrupt something. The active tests will take something offline, basically, or a passive check might mark it to get taken offline later, and so on.

The thing that we are working on, which we have working partially but not entirely, is automated refunds, which is basically for the case where the hardware breaks so much—

swyx

Yep.

Evan Conrad

There's only so much that we can do, and it is the effect of pretty much the entire industry. A pretty common thing that I think happens to everybody in the space is that a customer comes online, they experience your cluster, and your cluster has the same problem that any cluster has—or it's a different problem every time, but they experience one of the problems of HPC. Then their experience is bad, and you have to negotiate a refund or some other thing like this.

swyx

It's always case by case, and a lot of people just eat the cost.

Evan Conrad

Correct. One of the nice things we can do as a market, and have been doing as we get bigger, is immediately give you something else, and then also automatically refund you. You're still going to experience it. The hardware problems aren't going away until the underlying vendors fix things, but honestly, I don't think that's likely because you're always pushing the limits of HPC. This is the case when trying to build a supercomputer.

One of the nice things we can do is switch you out for somebody else somewhere and then automatically refund you or prorate, or whatever the correct move is.

swyx

Yeah, yeah. One of the things that you said in this conversation with me was that a provider is good when they guarantee automatic refunds.

Evan Conrad

Yep.

swyx

Which doesn't happen, but—

Evan Conrad

Yeah, that's in our contract with all the underlying cloud providers.

swyx

You built it in already.

Evan Conrad

Yeah. So we have a quite strict SLA that we pass on to you.

swyx

Yeah.

Evan Conrad

The reason I'm hedging on this is because we have some amount of active checks and some amount of passive checks. There are always new genres of bullshit.

swyx

Mm-hmm.

Evan Conrad

The new genres of bullshit might cause a customer to have a bad experience, the active or passive checks didn't catch it, and so then it's a manual process after that. We have a literal thing on our website where you can just say, “Hey, some hardware problem. Please tell us,” and then we will go and resolve it for you.

swyx

Well, cards don't change from generation to generation. What is a new genre of bullshit?

Evan Conrad

If every component piece in the cluster has maybe a 1-in-100 chance of failing, or maybe a 1-in-1,000 chance of failing, or maybe a 1-in-10,000 chance of failing—

swyx

You discover them.

Evan Conrad

You discover them. There are ones that maybe nobody saw, maybe you didn't see, or maybe it only matters for this 1 cluster with this motherboard in this particular data center, or something. There are new interactions that otherwise don't happen. Most problems are really common, and you can adapt to them. A GPU falling off a bus is one of the most common things that can happen.

swyx

So it's not SF Compute's job to go fix those things.

Evan Conrad

No, it totally is to some extent.

swyx

You just—

Evan Conrad

Totally is to some extent. We operate the cluster. Unlike a reseller, which is what we were doing before, in almost all cases we have BMC access. So if on your laptop there's the button in the top-right-hand corner that you can hold down to re-image the machine—

swyx

Mm-hmm.

Evan Conrad

There's a similar thing in a server. It's this other box that kind of plugs in, and it basically lets you reset the machine from outside. It's a remote-hands sort of thing. We ask for this from a lot of our vendors, which means we have quite a lot of ability to solve problems for customers in a way that you might not get from a reseller.

Oftentimes, we're the person who's debugging your cluster. For most customers that we work with, we have a Slack channel. Our entire engineering team gets put in the Slack channel. If there was a problem at 2 AM, we're the ones debugging your problem at 2 AM. Not always the case, because we don't physically run the hardware cluster or the data center itself, but most problems are solvable through this.

swyx

So that's the auditing side.

Evan Conrad

Yeah.

swyx

The other side is, I think of it as standardization, or whatever you call it. Beyond auditing, the other part of the work is kind of standardizing the commodity contracts.

Evan Conrad

Yeah.

swyx

Yeah.

Evan Conrad

There are 2 ways that we do that. One is that you set a “this-or-better” list. You set a spec list, and you say, “Oh, you're going to get...” A common variable is the amount of storage on the cluster. You'll say, “Oh, you're going to get X or better,” and there's some guaranteed minimum, and sometimes you might get more.

We're working on a persistent storage layer that might sort of abstract a lot of this away, but mostly it's that, and then there's a whitelist of motherboards and various generations of things. The other part is that we run the clusters from bare metal up, and so we make a UEFI shim.

If you're not familiar with what UEFI is, UEFI is the modern version of BIOS.

swyx

Firmware.

Evan Conrad

Modern meaning it's been around forever, but BIOS is really old. It's this old IBM thing. You can write code that exists at the UEFI layer, and again, when you hear UEFI, you should think BIOS. It does the same sort of thing as a PXE boot, but in environments in which PXE boot doesn't necessarily always work for us.

It basically sits at your BIOS, downloads an image, boots into an image that's custom for the user, and then on top of that image, we can throw Kubernetes on it, we can throw VMs on it, or whatever you want. At some point, we'll probably do more stuff with that, but that's functionally what we can do.

The nice thing, though, is that because you control it from that layer, you can easily image an entire cluster, make it all the same, and run your performance tests. It's all automated, so much nicer than what we used to do.

swyx

Yeah. Yeah, I mean, that is very important work. I think, for me, as a trader, I need standard contracts.

Evan Conrad

Correct. Yeah.

swyx

And so there basically needs to be the spec of a GPU.

Evan Conrad

Right.

swyx

Yes.

Evan Conrad

What we functionally do is have a market under the hood that's focused on the buyer and the seller, and it's optimized for them. Beyond that, for a trader, you can standardize around a certain segment of it, and you can trade on that contract. That's the goal that we're trying to get to, but you start by making something that works really well for buyers and really well for sellers.

swyx

For those who are not familiar with derivatives markets, I can go ahead and say this: the point of being cash-settled, which is something that you mentioned, is that you don't have to take physical delivery of the GPUs.

Evan Conrad

Right.

swyx

And that actually does mean that, almost for certain, there will be more volume on SF Compute's marketplace than actually changes hands in GPU terms.

Evan Conrad

To be super clear, we are not a derivatives market.

swyx

This doesn't happen yet. Yeah.

Evan Conrad

We are not a derivatives market. We may, in the future, work to create a cash-settled future. We are not currently a derivatives market; we are an online spot market.

swyx

Yeah.

Evan Conrad

Um—

swyx

I just think people, normies, get really upset when they learn things like, “Oh, derivatives on mortgages are 12 times larger than the mortgages themselves.”

Evan Conrad

Yes. A common thing that people have talked to us about, or a fear or concern people have, I think, is, “Oh, you're financializing compute, and this will cause various problems.”

swyx

Subprime crisis.

Evan Conrad

Yeah. I think, first, part of this is just because crypto caused a lot of people to think about finance in a very degen way, if that's the right word. Before that, the 2008–2009 crisis caused people to think about it also in sort of a degen-y way, and this is very much not our mindset.

The reason to create a derivative at all, or the reason to create a future at all, is risk reduction. That's what futures do. The reason why a farmer wants a future is because they have no idea what the weather is going to do, and they don't want to be on the hook. They have small margins, and if things go wrong, they really, really want to have a locked-in price so that they can continue to exist for the next year.

Data centers are the same way. The way that they solve it today is you go out and sign long-term contracts with your customers. What that does for you is it means your business is de-risked. You don't have to worry about the revenue for the next year.

But that means that the customer now has to worry about what they're going to do with all this compute if they don't optimally use it, and so on. That just pushes everything onto the startups, who then in turn push it onto VCs. What the VCs are forced to do in order to invest in AI is go and write big, giant valuations for pre-revenue companies at ridiculous multiples.

So what you've done by not having a future is you've inflated the venture capital market, and that's a bubble that's totally going to pop at some point. A lot of the companies are not going to work, and the valuations are not going to work. What's going to happen is a lot of these funds aren't going to return to their LPs, and that affects the broader market.

The way that you solve that, the way that you add stability to the entire economic system in this chain, is you add a future. That's how we did it in lots of other markets. It doesn't have to be this like, “Oh my gosh, we're going to speculate on GPU prices,” and whatever. No. The whole point of SF Compute is to reduce the risk, reduce the technical risk, and reduce the financial risk.

Let's just chill out a little bit. There's so much other random stuff. It's supercomputers, there's AGI, whatever. No, let's just chill the fuck out.

swyx

I mean, also, Dan is going, raising at a $30 billion valuation for Ilia. You know, like—

Evan Conrad

Yeah. If everybody else in all of AI is pushing the hype and the extreme, everything we've been trying to do is go the other way.

The whole website is just a fucking single page. The entire brand is just, “What if we were calm in nature?” Everything that we do as the product is just calm. What if we were the opposite force of the big, hype-y, extreme thing? What if we just chilled things out?

Part of that was because, in the beginning, we were at the whim of the hype-y nature. Our entire origin is that every 30 days, if we don't sell out, we're going to go crazy and just completely bankrupt the company. Everybody in the company is just like, “What if we just chilled out? What if we stopped for a bit?”

swyx

Mm.

Evan Conrad

What if we stopped for a bit?

swyx

This is the first time I've ever heard derivatives are the way to chill out.

Evan Conrad

Yes. No, futures are the way to chill out.

swyx

Futures.

Evan Conrad

Futures are the way to chill out the entire industry. We wouldn't be doing this if that weren't the case.

swyx

I like that.

Alessio Fanelli

And you have a very nice brand with a clear sky.

swyx

Sure. We have to ask about the website.

Evan Conrad

Yeah.

Alessio Fanelli

What was the inspiration behind it? Why did you not go with the black-neon, more-cool thing and go with something more nature-oriented?

Evan Conrad

I don't think I really am a black-neon sort of person. I say this wearing black pants, and I thought I was wearing a black shirt, but apparently I'm not.

The actual thing was that a lot of companies do this thing where their website—you go there, and it's like a magical experience, and everything is extreme and amazing and incredible. Then you go to the product, and it's some SaaS app or something.

It's not actually that exciting, and that expectation of being really, really good, followed by the falloff—the drop from not being really, really good—was something that, from a product perspective, I never wanted to happen, especially because in the beginning, our product was really bad.

And so I don't want to set the expectation that it's going to be an amazing experience. I want to set the expectation that it's going to be a good price for short-term bursts. What we did instead is set the bar really low. You set your expectations really low, and then you get a supercomputer for millions of dollars cheaper than you would've otherwise gotten a supercomputer. And so you have the opposite expectation: really low expectations that are met or exceeded.

I think that's the correct way to do things. But also, we were just so sick of hype and excitement, and I really want to not do that.

swyx

It's weird: by being anti-hype, you have created hype. I would say the vibe's very immaculate—

Evan Conrad

Yeah, I hate that.

swyx

You know?

Evan Conrad

Yeah.

You just go to the Bay, The Cow Trade, and put up a banner that just says “SF Compute.”

Evan Conrad

True. That banner was created about 5 minutes before we had to actually put something up, before the deadline was there.

swyx

Yeah. You opened up Microsoft Word and did some serif.

Evan Conrad

Yep.

swyx

What is the font?

Evan Conrad

Exactly.

swyx

It's—I don't know.

Evan Conrad

Yeah, that was—indeed.

The only caveat to this, the only caveat where we ever violate this rule, is when we're pitching San Francisco. I think San Francisco is amazing, so sometimes you will see these advertisements.

swyx

You mean the city?

Evan Conrad

Yeah, the city. There's a part of SF Compute's brand that's these beautiful images of San Francisco or various San Francisco things, and I am the complete opposite about this. I am such a San Francisco promoter that any time we talk about the city, I want to show the city through the eyes that we have, which is mostly just a gorgeous, beautiful area with nature.

A lot of people think about San Francisco and they think about—

swyx

The Tenderloin.

Evan Conrad

Yeah, or the tech industry, or grind culture or something. But no, I think about the fog, the gorgeous view over the bridge, and the fact that there is this massive amount of optimism in the city. The backdrop of that optimism is the most beautiful countryside in all the world.

Any time we talk about San Francisco, you'll see that we have a billboard somewhere that's just like, “Local-friendly supercomputer,” whatever, and then the backdrop is beautiful and amazing. That's because, to some extent, we're pitching the city and the people here. I think the people in this city are actually really amazing, and so you get to earn the brand.

swyx

Mm-hmm.

Evan Conrad

Because the expectations are met. Whereas on our own product, I typically want it to be better, and so I set the brand a lot lower. Then the expectations are higher. You still meet the expectations, but you set them a little lower.

swyx

I know—are you the designer? I know you have an artistic side.

Evan Conrad

I was in the beginning. I'm a figurative artist, so I draw people. But we've worked with a design firm. Aerofoil was really excellent with us, and nowadays, John Pham—

swyx

Oh, yeah.

Evan Conrad

Head of design—

swyx

From Vercel.

Evan Conrad

Yeah. John is unbelievably amazing. I think the amount of care, craft, and attention to detail that he puts into just everything is so cool.

swyx

Yeah.

Evan Conrad

The other person is Ethan Anderson, our COO, who has this RISD design background. He's sort of an industrial designer—I'm probably going to say that wrong. He's probably not an actual industrial designer, but he has a design background. So I think between me, John, and Ethan, we're—

swyx

The source of the vibes. I had to ask.

Evan Conrad

The source of the vibes.

swyx

Okay, so we're going to zoom out a little bit. One of the last things I wanted to ask you was—I remember, I think the first time I met you was when you were kind of solo and working on your email startup.

Evan Conrad

Oh, yeah, yeah.

swyx

I have a favorite pet topic of mine. We were here with Dharmesh yesterday talking about someone building an agent that reads my emails.

Evan Conrad

Yeah.

swyx

And you did, and I think I actually paid for the first one. You were so excited in the early GPT-3 days. I was like—you were like, “I'm building the most expensive startup ever.”

Evan Conrad

It's so expensive.

swyx

Anyway, the point being: you're a very smart guy. You built email, you didn't like it, and you pivoted away. I've seen others—every year there's someone who says, “I will crack email,” and then they give up. What is so hard about email?

11. From Email To Supercomputers

Evan Conrad

I didn't pivot away because the product or the idea was bad. I pivoted away because I was super burnt out. I did a startup for about 4 years, and the first thing didn't work out.

swyx

Is this Room Service?

Evan Conrad

Yeah, this is Room Service. My startup before this originally started as Quirq, which was a mental health app, but Quirq had the same problems that basically every mental health app has: your retention goes to 0 if you work it in any capacity.

So I switched and said, “Okay, well, I will do something that's closer to my actual background.” It was a distributed systems company called Room Service. Room Service went for about 9 months and then had the same problem that I think every other competitor in that space has, which is that mostly people build it in-house.

I went back to our investors at the time, Nat and Daniel, and specifically Daniel told me that I should go start the Ocean, and that I would find something else to do and just throw shit at the wall. I think it was Gustaf at YC. Maybe it was probably actually Dalton Caldwell. Dalton Caldwell, like, just said, “Don't die.” You can just keep doing things and don't die.

I think I just got it in my head that you should keep trying things and not die. I really did not want to die and didn't really know what to do, so I threw out 40 products with the assumption that if you just keep trying things, you won't die.

This is actually not the most ideal thing to do. You should totally just pick a thing and go with it. But my brain wasn't set on, “Oh, I should do this particular thing.” It was set on not dying. So I just kept going for a very long time, for 4 years, and by the end of it, I think I was super burnt out.

I was going to do the email thing with 1 co-founder, and then they quit. Then I was going to do an email thing with another co-founder, and they fell in love and decided to get married, and you know, all that.

swyx

Okay, so it wasn't that email is intractable.

Evan Conrad

Yeah.

swyx

Is this a graveyard of ideas? Everyone wants to do email, and then nobody does because of something?

Evan Conrad

I think it's just hard to make an email client. I think it's hard to make an email client in a competitive space in which there are lots of things. I do think the better version of that is something that looks closer to what Intercom is doing, and Intercom obviously existed beforehand.

You can think about any product: should you be doing it, or should somebody else in the industry who already has the existing customer set do it? I think Intercom has pretty successfully done this. They already had the position to do it.

What do you actually need the AI to write your emails for? Most people don't need this, but support use cases are pretty much there, and the people best able to execute on this are totally Intercom. Props to Owen. I think that was completely the correct move.

swyx

Yeah.

Evan Conrad

So should be our closing thoughts.

swyx

Closing thought, call to actions.

Evan Conrad

Yes.

swyx

Like, are you—

Evan Conrad

You're hiring.

swyx

Yeah.

Evan Conrad

Oh yeah, we are. We are hiring for two roles as of this recording. I don't know, maybe this will change and we'll be hiring for different roles, so go to the website or whatever. But the first role is for traditional systems engineering. This is for low-level systems or low-level Linux-y people.

swyx

Rust.

Evan Conrad

Yeah, almost all of our codebase is in Rust, but we're not necessarily just looking for Rust engineers. We're specifically looking for Linux-y people. The pitch is that you get to work on supercomputers, and you get to work at one of the few places in supercomputing that, I think, has a pretty good business model and is a working thing.

People generally seem to think that our vibe at SF Compute is very nice. We have an unbelievably excellent team nowadays. Our CTO is Eric Park. He's the co-founder of Voltage Park, which is one of the other GPU clouds.

He's quite possibly the sweetest man I've ever met. He's extremely chill and also extremely earnest and kind, and the rest of the team feels that energy very strongly.

The other role we're hiring for is financial systems engineering, which I really should learn what to call. It's not systems engineering, but we should really find a better name for this role. It's basically a fintech engineer. We have the same problems as traditional fintech does: we have a ledger, reporting requirements, and all that stuff.

This role is responsible for the “not lose all the money” goal. We've got a whole bunch of money flowing through us. There is a bunch of stuff you need to do in order not to lose all that money. The actual outcome of that work, besides not losing all the money—which is very important—is that you end up with better prices for the vendors and better prices for the buyers.

This means that your grad student who is making the cancer cure or whatever, and needs to be able to buy 100K of compute to scale up really big, actually can do so. This is part of the reason to work at SF Compute: the things you do actually matter in a way that you don't necessarily get at all companies. Functionally, we run supercomputers, not soybeans or, I don't know.

It's a very cool place to work because the outcomes of what you do have real-deal impact—

swyx

Yeah.

Evan Conrad

—in a way that you don't always get when you're doing SaaS.

swyx

Excellent pitch. I bet you've done that a lot, but it's nice to hear it for the first time. I was going to say, have you looked into TigerBeetle, the double-entry accounting database?

Evan Conrad

We have, though—

swyx

That seems to be the thing if you want to make systems that don't lose money.

Evan Conrad

Yes. For systems that don't lose money, there are lots of other things you have to do. You have to make things in a format that your accountants can read, and then get audited and so on. It's not purely the tech.

swyx

Cool.

Evan Conrad

Yeah.

swyx

Awesome. Thank you so much.

Evan Conrad

Yeah, of course.

swyx

It's been great.

Evan Conrad

Thank you so much for having me.

SF Compute:将算力商品化,永久解决 GPU 泡沫 — 文字稿与摘要 | BidClub