开源模型 vs 前沿模型:究竟谁赢?| 每位工程师都将需要的10万美元Token预算
Clay Bavor 的核心判断是:更便宜的开放权重模型会接管昨天的前沿模型工作负载,但不会消灭市场对前沿模型本身的需求。 2023年春季GPT-4级别的智能,如今每个等价Token的成本约降至原来的1/300,前沿模型到微调开放模型之间正在形成一条“流水线”。但编程、科学、法律、发明和发现等工作,可能支撑对“姑且称之为前沿级智能”的无上限需求。
Token价格会因技术进步而下降,但未必在经济层面崩塌,因为推理需要消耗更多算力,而GPU和电力仍然稀缺。 OpenAI 的o1表明,随着测试时算力增加,模型表现仍在提升;Nebius 创始人告诉Harry,即使供给增加10倍,也可能在1天内售罄。Bavor听说并观察到,顶尖工程师按运行速率计算每年Token支出超过10万美元;他押注稳态支出会“更接近开发者薪资的20%”,而不是3.8%。
AI应该带来更小、杠杆率更高的团队,但Sierra的企业级打法说明,不能把1家149人的软件公司模式外推到所有领域。 其工程师估计,Claude Code、Codex和Sierra内部工具让他们的生产力提升3至20倍;但代表Fortune 50中40%的客户仍需要集成、信任、监管理解和前置部署支持。Bavor的限定条件很关键:FDE并非必需,但它们是“重要催化剂”。
Sierra正通过共享数据网关、全公司Agent、可复用技能,以及建立在运营语境之上的战略模型,在内部成为一家AI原生公司。 Pinecone可以检索获准访问的Slack消息、文档、演示文稿和复盘材料;Bavor则用自定义技能扫描每1份招聘材料。Sierra Brain加入20至30页的公司介绍、董事会信、运营复盘和战略信念,打造1个深入理解业务的“战略思考伙伴”。
产品版图已经从客户支持扩展到由Agent驱动的前台全链路,覆盖发现、销售、转化、服务和营销。 Rocket的部署涵盖住房发现、再融资触达、贷款准备和贷后服务;Next则用Sierra提供个性化产品推荐并搭建购物篮。终局是“对话就是界面”的世界,背后依靠可复用的行业能力,而不是无休止地堆积定制项目。
Sierra围绕里程碑而非日历惯例或最高 headline 估值来安排治理和融资。 董事会每6周交替召开3小时和90分钟会议,使用6至10页的备忘录,因为“写作就是把思考写在纸上”;当Claude 4.5和Codex 5.2改变软件开发方式时,这一节奏让管理层能够及时反应。每1轮融资都是投资人主动找上门,Sierra都“引导并接受了更低的价格”,以便为下1个明确更高的里程碑融资。
AI熟练度正在重塑谁能创造杠杆、Sierra如何面试,以及初级人才的优势来自哪里。 Sierra最有效率的员工中,有些只有22或23岁,而且“彻底AI化”;如今工程候选人会获得150美元,用于购买自己选择的编码Agent,并在自己的笔记本电脑上完成构建,再解释结果。更广泛的招聘标准仍是“聪明、友善、强悍”,将工具熟练度与系统判断、产品思维和文化契合结合起来。
1. Sierra在平台重置之际成立,没有追逐预训练
Bavor在Google度过了18段愉快的岁月后离开,原因是2022年末3个条件同时成熟:长期以来的创业意愿、对Bret能力与品格的信心,以及语言模型正在把“那副牌”重新洗牌,给小公司带来机会。两人相识20年,此前也差点合作过几次。
Sierra从Google继承的关键特质,是愿意沿技术栈一路下探到产品真正需要的深度。Sierra在2023年4月就预判语言模型Agent将出现,招募了ReAct论文背后的Princeton教授担任创始研究负责人,并在相关能力“应该可以实现”但尚未具备工程可用性时,打造新的Agent框架和架构。
Sierra曾考虑预训练,但很快放弃。Bavor称,前沿模型是“一袋高度易腐的浮点数”,初始资本开支和持续资本开支都只有极少数公司承担得起;Sierra选择紧跟实验室和超大规模云厂商,然后在开放权重模型之上构建专有微调模型。他的边界是:掌控自己的命运,但不要虚构出“必须拥有更多技术栈”的故事。
2. 开放权重继承昨天的前沿能力,智能需求却持续攀升
Harry的悲观情景挑战很直接:如果开放模型能够处理不断扩大的多数企业任务,前沿模型最终是否只剩越来越难的问题?Bavor的回答从当前分母开始:完全自动化的企业工作仍只是“舍入误差”,而差距除了模型能力,也可能来自模型、应用层和组织扩散等问题。
Bavor将常规能力过剩与高价值智能分开看。任何软件公司都会把高级工程师升级为首席或杰出工程师;退货不需要最强模型,但编程、法律、科学、材料、发明和发现表明,对有用智能的需求没有明显上限。
经济运行模式是一条“流水线”:2023年3月、4月或5月的GPT-4已经能够支撑有价值的工作负载,而如今同等智能的每个Token成本约为当时的1/300。这些工作负载可以迁移到微调开放权重模型;前沿模型则继续服务于只有更高智能才能覆盖、且收益足以支付成本的任务,企业最终会混用两者。
对于中国更强的开放模型生态,Bavor给出的解释带有保留:这是对美国前沿训练运行结果进行规模化蒸馏。如果美国实验室发布可比的开放权重模型,就会压低自家托管模型的价格;如果另一家公司无法亲自打造前沿模型,“也许下1个最优解就是把它们蒸馏出来并提供给市场”。
3. 推理和稀缺算力阻止Token经济学崩塌
Bavor认为1个被忽视的里程碑是2024年末的OpenAI o1。它的测试时算力曲线持续“向右上方移动”——以对数方式增长,因此最终会趋平——额外的推理和思考带来了更好的表现。其含义是,更便宜的Token会让模型和Agent为追求更高智能而消耗更多Token。
硬件会让每1美元对应的等价Token数量增加,部分工作负载也会迁移到更便宜的开放权重模型。但如果对智能的需求事实上没有上限,而Blackwell、H100、电力和GPU容量仍是瓶颈投入品,那么即便底层技术持续进步,基本的供需关系仍会为Token价格托底。
Harry转述Nebius的判断:即使容量增加10倍,也可能在1天内售罄;Bavor相信这一说法。自建模型会削减前沿模型供应商部分利润,但无法消除能源和算力等受约束的物理投入。
本地模型可以改善消费者体验,但单靠本地模型无法解决服务端挑战。手机会遇到散热限制;Bavor设想过针对语言模型优化的设备,或接入市电、放在家中的计算设备,可能缓解部分需求,但前沿工作仍需要“某个数据中心里1个巨大的TPU或GPU机架”。
4. AI原生工程把瓶颈推到了代码之上
深度使用Claude Code、Codex和Pinecone的Sierra工程师估计,他们交付功能的速度提升了3至20倍。Bavor预计各职能都会出现更小、杠杆率更高的团队,但他认为软件工程和数据分析效率的提升已经明确可见,而不是仍待验证的假设。
Sierra的内部基础设施是1个MCP网关,将主要系统和服务聚合到1个具备权限感知能力的接口中。接入Claude、Codex或Pinecone后,员工可以在自己有权查看的全部信息上进行推理,包括Slack、文档、演示文稿和运营复盘,同时不会获得他人私密资料的访问权限。
Pinecone在这一网关之上叠加了公司专用工具链和共享技能库。Pinecone知道如何构建Pinecone;它围绕自身工程,以及Sierra的核心Agent架构和Agent Studio——Agent在这里被构建和部署——都搭建了工具链。Bavor的私人“Clay扫描器”编码了他审阅招聘材料时关注的内容,加快了他对每1位候选人的审核和批准。这套系统“正接近不可或缺”。
Sierra Brain向任何Agent提供1份20至30页的公司说明,涵盖公司、组织、竞争、优势和弱点,以及近期董事会信、运营复盘和其他关于世界的判断。随着代码生成能力提升,Bavor预计约束会从写代码转移到审代码,再转移到决定什么可以存在,并把它编辑成什么应该存在。
5. 企业级AI仍靠嵌入客户现场的人来规模化
Harry将Sierra与Lovable作了对比。据报道,Lovable已经——用Harry的话说,“我认为”——以149人实现5亿美元ARR。Bavor认同团队趋于精简的方向,但拒绝把它当成普遍模板:Sierra有50%的客户营收超过10亿美元,30%的客户营收超过100亿美元,公司服务于Fortune 50中的40%。
这些受监管、技术复杂的组织都是“雪花”。要取得成功,必须理解业务结果,接入异构技术栈,建立关系并赢得足够信任,成为合作伙伴,而不是“把一些软件扔过墙”的供应商。
Sierra借鉴了Palantir的前置部署模式,早期设计合作伙伴包括Olakai、SiriusXM、Sonos和Weight Watchers。工程师深度嵌入客户现场,创始工程师Mihai甚至实际上成了Weight Watchers的员工,还会收到绩效评估邮件。这种贴近客户的方式让Sierra理解了部署面向客户的AI到底需要什么。
Bavor拒绝接受“没有FDE就卖不动企业级AI”的绝对判断:Sierra的平台透明、可导出,也可以独立使用。但客户与Sierra共同部署后,Next在6周内就从项目启动进入电话和聊天生产环境,Cigna则约58天上线;这使FDE成为缩短价值实现周期、提升结果质量的重要催化剂。
6. Sierra正从支持业务扩展到完整客户生命周期
Rocket展现了这一方向:Sierra与Redfin重新思考住房搜索,联系潜在再融资客户,构建Rocket Assist以设计贷款方案并收集信息,随后继续支持贷后服务。Next则通过个性化推荐搭配服装并扩大购物篮。这些是入站和出站销售动作,而不只是支持自动化。
Bavor的平台策略是让应用反哺可复用平台,使第3、4、5个应用更容易构建。如今编码Agent让构建真正独特的Fortune 50需求变得经济可行,但他的直觉是,看似只出现1次的需求通常会在其他地方重复,最终变成共享的行业能力。
巨大市场的代价就是竞争。Bavor称客户正在“用脚投票”:Sierra的规模是同一时期成立、最接近的类似初创竞争者的数倍,增长速度也更快;由于平台广度和行业经验会复利增长,他的判断是,市场更可能走向Uber/Lyft式格局,而非AWS、Google Cloud、Azure三足鼎立,Sierra希望占据更大的头部位置。
7. 每6周治理一次,用快速重定价取代仪式化董事会
Sierra每6周交替召开1次3小时董事会和1次90分钟董事会,因为“AI的时间钟”已经跑在季度治理之前。目的在于吸收新证据、更新先验判断,并在证据仍然有效时调整方向。
冬歇结束后,Claude 4.5和Codex 5.2恰逢Bavor认为编码Agent能力出现根本性跃迁。这改变了Sierra的软件开发方法和核心产品路径,正是每6周节奏希望及时捕捉的不连续变化。
Bavor和Bret不用演示文稿,而是撰写6至10页的备忘录:“写作就是把思考写在纸上”,写下来会让薄弱推理更难隐藏。在最初8个季度里,即便实际表现超过预测,每封典型信件仍会列出约7个不满意之处,让董事有时间挑战真正的问题,而不是被管理层单向汇报和引导。
Sierra早期承认过1个问题:2024年初看到了需求,却没能足够快地招人,导致本可以服务的客户无人服务。融资也遵循同样的里程碑纪律:融资金额足以到达下1个明确更高的里程碑,对稀释保持敏感但不追求极致,并接受低于Sierra本可获得的价格;每1轮融资最终都低于可获得的价格成交。
8. 工艺和强度是操作系统,不是口号
工艺意味着,1家伟大公司的总和,是“成千上万”件分别做到优秀的事情,包括人、流程、文化和产品。它也体现了Sierra会如何对待客户最宝贵的资产:在公司经历第1个黑色星期五/网络星期一期间,1名首席工程师、1名运营负责人和2位创始人中的1位,亲自实时监控每1段Agent对话。
强度反映的是竞争现实:通过复杂Agent与客户互动看起来不可避免,但没有任何公司天然有资格赢。Sierra用“聪明、友善、强悍”这一困难的维恩图招聘;Bavor和Bret则努力保持领先节奏,不断追问为什么某件事不能明天发生,而要等到下周。
创始人模式是有选择的“直接施加力量”,而不是不加区分地深入17层组织。日本就是例子:当他们追问今年怎样才能做成1项重大业务时,答案意味着需要在当地配置10人,进而促成收购Opera Technologies,并围绕日本式服务期待——例如omotenashi——展开建设。
家庭是这套体系的制衡力量,也是Sierra明确写入价值观的内容之一。Bavor反对创业文化中“有时带有表演成分的苦干”:人们可以打开加力燃烧器,同时保护孩子、父母、朋友,或任何工作之外真正重要的事。他对时间表的警告同样适用——“工作就像气体”,会膨胀到它所获得的全部空间。
9. AI熟练度已成为初级人才的不公平优势
Bavor认为,年轻毕业生获得的是1个不寻常的机会,而不只是被替代的风险。4年几乎不受限制的时间和可支配精力,可以让他们掌握1,000家公司都想要的AI工具;Sierra最有效率的员工中,有些只有22或23岁,“彻底AI化”,对这些工具的熟悉程度超过资深同事。
Sierra围绕这一现实重构了工程面试。候选人选择1个应用,获得150美元购买自己偏好的编码Agent,带上自己的笔记本电脑和工具完成构建,并解释自己的过程。架构、系统设计、产品思维、价值观,以及“聪明、友善、强悍”仍然重要;Bavor希望在2个月内让每1场面试都具备强烈的AI原生要素。
Bavor认为,线下工作能够支持学徒制和导师制,而这对学习很重要。他引用Richard Hamming的建议:“找到优秀的人,与他们一起工作,并向他们学习。”知识和努力会像利息一样复利增长,因此,早期接触优秀实践者可能改变1个人的职业轨迹。
网络安全的重要性正不断上升,因为攻击能力已经“连续提升了5个档位”。但Bavor仍未对产品方向下定论:防御端的赢家可能是专业供应商,也可能是模型和编码Agent本身。
10. 联合创始人的互补性把分歧转化为判断力
Bavor和Bret最近曾就1个进展缓慢的领域发生分歧:问题需要更强的流程,还是需要不同的领导方式。他们分别审视两种观点后得出结论:两者都需要一些。他们共同遵循的标准,不是谁拥有某个论点,而是那句追求真相的话:“这是正确的。”
他们没有把公司一分为二,而是划分主责和辅责。Bret主责销售和软件工程,在系统设计和架构上拥有Bavor高度信任的直觉;Bavor主责运营、财务、法律和公司经营。涉及重大影响的合同需要两人共同持有“核钥匙”,在错误代价最高的地方保留重叠决策。
Bavor从Sundar身上学到的是“动态范围”:既能从5年战略切换到像素、投影、声音和纹理,又不失去人性。他从Google得到的更广泛经验是:持久的使命、聪明的人、追求真相的文化,以及让合适的人照料众多实验,可以让1个组织觉得自己几乎有能力解决任何问题。
We have not yet appreciated the unbounded demand for, call it, frontier levels of intelligence. Part of the driver of the difference is probably the willingness of Chinese companies to do scaled distillation of the frontier models. If you can't build frontier models yourself, okay, maybe the next-best approach is to distill them and offer them up. Every one of our rounds, we actually guided to and took a lower price than we could have. Some of our most effective employees in the entire company are 22 or 23 years old and have been completely AI-pilled. We completely changed our engineering interview process.
When Pat Grady at Sequoia and Neil Maitra at Greenoaks tell you someone is special, well, it kind of means something. Clay Bavor joining me in the hot seat, co-founder of Sierra, one of the fastest growing AI companies in the world. Sierra has raised more than one and a half billion. They work with some of the biggest companies in the world, and they're valued at almost $16 billion, and they work with 40% of the Fortune 50. But before Sierra, Clay spent an incredible 18 years at Google, where he worked on some pretty cool projects. Google Labs, naming one. Google Workspace, Gmail, Google Drive, Google Photos. Jesus, is there anything Clay didn't work on at Google? On top of that, he's just an awesome dude. Luckily, we had a chance to do it in person in London. This was so much fun, and I can't wait to hear your thoughts and feedback on this episode.
Clay, I am so excited for this, dude. I said to you downstairs, we do a lot of shows. I often speak to people before a show. When I speak to Neil Mehta, Ravi Gupta, Sangina, and GV, and I hear what I hear, honestly, they were some of the most astounding references I've had. So thank you for joining me.
Oh, it's nice to hear. No, it's a pleasure to be here. Thanks for having me.
Listen, they paid a lot to be featured, so you've got to drop a sponsor.
They're great. I'm so grateful to be working with each one of those guys.
It only cost 500 million bucks. I want to start with—I heard that Bret tried to hire you or start a company with you several times before Sierra.
Yeah.
Why third time lucky? Why, after 18 years—
Oh, third time's a charm?
Yeah. Why, after 18 years at Google, were you like, “Ah, now?”
1. Sierra Begins At Google
Yeah. Bret and I met 20 years ago. We both started our careers in the associate product management program at Google. He was class 1; I was class 3. We met in the context of some kind of shared project that we were assigned to, hit it off, and ended up staying in touch socially through mostly a monthly poker group that, in a good year, might play 2 or 3 times, so not quite monthly.
We had always wanted to work together and almost did a couple times. I think when Bret left—I can't remember if it was for FriendFeed or Quip—he tried to get me to join that. The short answer is twofold. One, I just loved my time at Google. Culturally, it was me. I learned more than I could ever imagine having learned in those years.
The people were so extraordinary to work with. I had a series of managers and leaders I got to work with who took bets on me and gave me, on paper at least, more responsibility than I deserved. I also got to work on truly fascinating things. I was just incredibly happy and engaged and growing as a person and professional.
And then, in late '22, the planets aligned in a way that I didn't think they would probably align again. I'd always wanted to start a company. I started a very modest company when I was 13 years old and always thought I would start another. If you're going to start a company with someone, you want to make sure that they're excellent in competence and in character, and then that the timing is right.
We could see that language models were going to be a thing. If ever there was a time when the proverbial deck of cards was shuffled in favor of smaller companies, it was at the advent of a new technology. So, happy at Google, the planets finally aligned, and I took the leap. We're 3 years and change in now.
18 years at Google—
Yeah. Well, I started counting in colleges. Gosh, I've been there 1 college, 2 colleges, 3 colleges, 4 colleges. Yeah, it's a long run.
That's even more terrifying.
Yeah. It's a long run.
My question to you on the back of that is—and it's a terrible question, so you can chastise me for it—what are your single biggest takeaways from that experience that you took with you to Sierra?
Mm-hmm.
And what did you leave behind?
It's such an interesting question. Of course, the scale of a 2-, then 10-, then 100-person enterprise software company is very different from—I think when I left Google, it was roughly 150,000 people. Things that I've definitely brought with me, number 1, are a willingness to invest as far down the technology stack as you need in order to build the service and product that you want.
Google, I think, from the early days, famously built its own, if not data centers, cluster architectures, and was really the first to use commodity hardware. That required building novel distributed systems for serving and data storage and so on.
And so we could see as early as April of '23, when we started the company, that agents were going to be a thing. This was before all anyone wanted to talk about was agents. We realized, okay, this should be possible. It's not yet possible, but we're going to have to invent frameworks for building these things, our own architectures really from scratch.
So, actually, our founding head of research was the Princeton professor who literally wrote the paper on language model-based agents, the ReAct paper. So we invented and went further down the stack than I think some companies at that point would have been willing to. We're not doing our own pre-training. We'll leave the capital expense there to the labs and the larger companies.
Before we move to 2, can I ask: did you consider training your own models? Because I completely understand the desire to own as much as possible.
Yeah.
Did you consider training your own models? And what was the thought process around not doing so?
It's a great question.
2. Why Sierra Skipped Training
We did briefly and discarded it. If you recall, in late 2022 and early 2023, as a startup in AI, you were kind of nobody if you weren't doing your own pre-training and building your own foundation models. Character, Inflection, and Adept had great people at these companies, but the capital expense—the ongoing capital expense—to create what is effectively a highly perishable bag of floating-point numbers just doesn't work for any but a small number of companies.
And so our calculus was, for areas that are deeply capital-intensive, how do we slipstream behind the investments that the labs and the hyperscalers are making, take as much as we can off the shelf, while still being willing to engineer more deeply? So today we have a set of our own proprietary fine-tuned models, but these are fine-tunes on top of open-weights models. We're not going all the way down to the mega-cluster training runs.
I think it's important that you are in control of your own destiny enough, and that you don't tell yourself a story that you need to go further than you actually need to go.
Is the future open models fine-tuned to specific company needs? And if that is the future, with the realization that frontier models are too expensive, is that a bear case for frontier models?
I think it's a lot more complicated than that. If you asked any software company, "Would you like to upgrade your staff-level software engineers to principal- or distinguished-level software engineers, yes or no?" 100 out of 100 would say, "Yeah, that sounds pretty great." So I think we have not yet appreciated the unbounded demand for, call it, frontier levels of intelligence.
Now, you don't need that in every domain. For instance, in our own business, we build AIs for companies to interact with their customers. You don't need Mythos to return a pair of shoes, right? You're good. You want to do that well, but we've got some capability overhang, so to speak, for doing something like that.
But in a range of domains—coding, certainly; science; materials science; legal—where the stakes are very high and there's a high degree of complexity, I think we're going to see effectively unbounded demand for greater levels of intelligence, and therefore for frontier models. That said, there will be an assembly line of, "Cool, GPT-4, which in March, April, and May of 2023 was good enough to do some set of things, is now 1/300th the cost for an intelligence-equivalent token." And so you'll have some assembly line of taking models that were once at the frontier to perform certain workloads, and then building open-weights fine-tuned models for those.
I think you'll end up with companies using both, mixing and matching them depending on the task at hand.
As we see open models become more and more advanced, does that not mean the problem set for frontier models becomes more and more challenging? As you said, we've seen the progression of open models so much that they can actually do the majority. I get it for solving climate change, cancer treatments, and materials science, but for the majority, what percentage of enterprise tasks can be done with open models today?
Well, I think if you look at what percentage of enterprise tasks are completely automated today, it's a rounding error, right? It's very low. So is that a model gap? Is that a gap in diffusing the technology into the company? Is that an application-layer gap? I think it's probably all of these.
You're obviously correct that as the open-weights models become more capable, the set of things they can do grows larger. The set of things where, all else being equal, if they're much less expensive, you would want to point a frontier model becomes smaller. But again, I think we're not imagining just how high the ceiling is in terms of demand for frontier intelligence.
Invention, discovery, building new products, and building new services—I think it's hard to get your mind around, when you have intelligence that can work around the clock to invent, build, and discover, how you would use that and how much of it you could use.
Can you help me understand? When we look at token economics, we thought that with chat, tokens over time would go down in cost. And with the movement from pure chat to chat and agents, and an agent-economy-based model, we're seeing token costs increase, not decrease. How do we see the evolution of token costs with the evolving formats, do you think?
You missed one thing in there, which is that a large amount of token use is driven by reasoning models now—thinking out loud to themselves. I think one of the most underrated developments of the past few years was the o1 model from OpenAI in late 2024. If you recall, there was a chart that showed test-time compute, or the amount of inference done—the amount of thinking out loud—and performance, and it just keeps going up and to the right.
It's logarithmic, so it starts to level out. But what it effectively demonstrated is that if you have enough time and compute, the model will be that much smarter.
As for what happens with token economics, I think there are many drivers underneath it. One is that you're going to end up with hardware that is able to produce more tokens at equivalent cost. The cost of the inputs, so to speak, will go down. We talked about how you'll have this migration of certain workloads to open-weights models.
I think one of the drivers that's hard to predict—how it will play out across both the open-weights models and the frontier models—is just the availability of compute. It's classic economics, microeconomics 101: supply and demand. If you have unbounded demand for frontier-level intelligence or GPUs to run open-weights models, and the rate limiter is the number of Blackwells and H100s you have, you end up with a floor on the cost of tokens because you've got to pay for the energy and you've got to pay for the compute.
We had the founder of Nebius on the show the other day, and he said that if they 10x supply, they could still sell out in a day.
I believe that. And I think that makes the point, which is, okay, open-weights models will be cheaper because you're avoiding some of the margin stack in the hosted frontier models. But what is the fundamental input? It's GPU capacity and power. That's still constrained.
One thing that could slightly alleviate that is actually running models locally. People say that it could be—
Yeah, on your cluster of Mac Minis or whatever?
Yeah, or even on-device, on phones. I don't quite understand that when we think about always-on AI, 24 hours a day. That's an awful lot to run locally. Is it a pipe dream, or do we think that's actually a reality that would alleviate the server-side challenge?
Oh, it certainly wouldn't alleviate it. I think it will make some consumer applications much better. But the reality is you need petaflops and exaflops of compute, certainly for training, and you want a whole bunch of compute quickly at inference time. You just run into thermal limits on your phone.
I do think it's shocking that we're all carrying around hypercomputers in our pockets these days. Will they get better? Yes. Will you have language-model-optimized hardware rolling out in our phones and in our computers? Yes. I can see a sort of home appliance where you plug into the mains and get on-demand access to a whole bunch of compute for things in your home, and maybe that helps alleviate some of it.
Certainly for frontier workloads, though, there's one place you can go for that, and it is a giant rack of TPUs or GPUs in a data center somewhere.
We spoke about frontier versus open. Now, frontier models, obviously, you have OpenAI and Anthropic in the US, who are the dominant leaders—everyone knows.
And my alma mater, Google.
And Google, of course. We had Demis on the show. Incredible, incredible. I love Demis.
I loved Demis, too.
Also one of the most humble leaders I've ever met, so I absolutely agree there. Open models in the US have lagged behind. We see Chinese models being unbelievably advanced and impressive. Do you agree that we have a challenging open ecosystem in the US, and does that worry you?
3. The Chinese Distillation Advantage
Part of the driver of the difference is probably the willingness of Chinese companies to do scaled distillation of the frontier models from the labs. My impression is that many of the open-weights models coming from China are derived from training runs done in the US.
I think if you have the US-based labs and hyperscalers developing the frontier models, there's an obvious question: Are they going to compete with themselves and drive price pressure on the frontier models by developing and releasing open-weights models that are of similar capability? If I was running that business, that's not something I would do.
So if you can't build frontier models yourself, maybe the next-best approach is to distill them and offer them up. I think that's probably the main driver of the difference.
I have to ask: You mentioned earlier that enterprise is a team sport. I love that.
Mm.
And you mentioned earlier who wouldn't want more advanced software engineers internally. Lovable announced yesterday that it had hit, I think, $500 million in ARR with 149 people. And in a show that comes out tomorrow, Rory, who's one of my co-hosts on this weekly show that we do, says, "Well, if you're Sierra, you can't do that." I mean, as you mentioned, Sierra is an enterprise business, and you have to have a different structure for the team.
When you look at the future of teams, are we seeing a world of dramatically leaner, fewer people in teams, or is it still very much dependent on customers? Will we still have very large teams for companies like Sierra with enterprise customers?
4. The AI Native Company
I think the general direction of travel clearly is toward smaller, higher-leverage teams. We have software engineers who are completely AI-pilled and using Claude Code, Codex, and our own internal agent, which we call Pinecone and use to run much of the company. They estimate they are between 3 and 20 times more productive in terms of features shipped.
The productivity gains, certainly in software engineering, data science, data analysis, and other areas, are coming in spades. I think, in time, it will touch all parts of really every company. So that's the general trend.
I think within a company like Sierra, where we serve, in particular, the large enterprise, we work with 40% of the Fortune 50. We have 50% of our customers doing over $1 billion in revenue, and 30% doing over $10 billion in revenue. These are some of the most complex and, in some cases, regulated organizations in the world.
To be able to sell and implement our product and solution successfully for organizations that are snowflakes, the process of selling and, more importantly, successfully implementing and deploying a solution like ours into the large enterprise is still a lot about deeply understanding our customers' business outcomes and objectives. It's about understanding their technology stack, integrating with it successfully, building relationships, and earning trust to show up not just as a vendor that throws some software over the wall, but as a true partner in diffusing this technology into, in our case, all of the front office—sales, support, marketing, and so on.
I mean, there's so much for me to unpack there. I was scribbling furiously. You mentioned the internal agent, Pinecone.
Yeah.
Can you talk to me about what that is, how it was built, and what it does? I'm intrigued to see how companies change and how they operate.
Yeah. It's one of the more significant developments in how we run the company over the last 6 or 9 months. We began by building what we call our MCP gateway. This is a single MCP server that aggregates all of the main systems and services that we use to run the company.
You can add this single gateway to your Claude instance, your Codex instance, and indeed to Pinecone. Basically, via any one of those agents, you have full access, with the permissions, of course, that you have as an individual at the company. You can't read someone else's documents, but you can read your own Slack messages.
It's like having superpowers, right? You can interrogate, in essence, the entirety of the company—all information that is published, whether it's Slack messages, presentations, operating reviews, and so on—and use that access to better reason, make decisions, and get things done.
Pinecone, of course, incorporates that MCP gateway, but then is a purpose-built harness for all of Sierra. Pinecone knows how to build Pinecone. There's a whole harness around the engineering of Pinecone, and our engineers there are phenomenally productive.
We have a whole harness around the core of our platform—our agent architecture, Agent Studio, where you build and deploy agents—which speeds up software development there. Then we have a shared library of skills that anyone at the company can build. You can build one that's private to you.
I have a whole bunch of skills, including one that is basically the Clay scanner for interview packets. To date, I review and approve every single hire we make, and I get some help from Pinecone. I've basically taught it what I look for and what I scan for: flag these if there are any instances of them. It's a shortcut to a faster, deeper read of every packet.
Pinecone has just become this approaching-indispensable tool for running the company. I think we're not quite there yet, but we're approaching indispensable. I could go on about some of the other interesting things we've built. I've been working on what I call Sierra Brain and some other things in the same vein.
What's Sierra Brain?
Sierra Brain starts with a 20- or 30-page document that grounds any agent in what we are as a company, what we do, how we're organized, our team structure, the competitive landscape, our strengths and weaknesses, and all of these things.
On top of that, I've given it access to every one of our recent board letters, every one of our recent operating reviews, and other insights and observations we have about what we believe to be true about the world. I can then use it to reason about what we should be doing as a company.
So it's a bit like a strategy thought partner, if you will, that knows the company, if not inside and out, very deeply.
We're going to get to your board letters, because I heard about these.
Oh.
And how you have boards every 6 weeks, not every quarter—
Yeah.
—because the world moves too fast, apparently. Trust me, I stalk the shit out of you. But I just wanted to stay on internal builds. You mentioned the internal agent that you have. We're having a lot of CEOs who I speak to say, “I have no idea. Do I just let my dev teams run wild on token spend? Do I give them some form of budget?” Or—
Token maxing.
Yeah. What's your personal take, and what do you and Bret say around the fire? Do we put a cap on this? Do we just encourage them to go wild? How do you approach it?
Yeah. I think over the past 6 months, using a bunch of tokens was a proxy for using AI. You were leaning into it and trying to be more productive with it, so I think it's generally been a positive signal.
I have heard and observed that top engineers who are really leaning into Claude Code, Codex, and so on are spending more than $100,000 on a run-rate basis on tokens per year. That's a meaningful fraction of an engineering salary.
I think the direction we're headed is toward some amount of token budgeting on a per-employee basis. For CFOs in the future, capital allocation will look more like how we allocate OpEx and then headcount. Headcount will be both headcount for salaries and SBC, and also tokens associated with headcount.
So
here's your salary, here's your token budget, have at it. We are not yet at that point. Our usage, compared to some of those larger numbers, is modest, and I think the benefit of learning at the fastest rate possible outweighs the capital discipline at this point. We prefer to learn quickly and see what works.
It'll be interesting to see how the rate limiter in software development moves around—what the Andy Grove “breakfast factory” constraint will be, what the constraining factor is. It used to be writing code. Now it's probably reviewing code. Pretty soon it will be deciding what is worth building and editing what could exist into what should exist. The dynamics there will be interesting.
I think the core question for us to understand if everything is slightly overhyped is: what percentage of developer salary will be spent on tokens in the future? Mark Benioff said that he spends $300 million a year on Anthropic for his dev teams. That works out to about 3.8% of developer salaries—not actually as much as the headline $300 million makes you feel.
If it stays at 3.8%, a lot of the companies that we're investing in and seeing around us are actually grossly overvalued. If it goes to 20%, they're undervalued. I had Brandon at Macaw on the show, who says he spends more on tokens than he does on headcount.
Yeah. I think 3.8% is wildly off from where the steady state will converge.
Where do you think it will be? I'm not going to hold you to it in 5 years' time, but do you see it being at 20%?
Oh, I do.
So the $100,000 a year actually will be normalized? If you think about a great dev in the Valley, I presume $500,000 is where they're at?
Sure, that would be the upper end.
So it feels normal.
Yeah. I would not bet on 3.8%. I would bet on much closer to 20%.
In software engineering, the gains to me seem unequivocally there. You can debate whether it's 2X, 10X, or 20X. Even if it's 2X, you've just effectively doubled the size of your engineering team. That's remarkable.
You mentioned 40% of the Fortune 50—
Mm.
—being customers.
Mm-hmm.
You use that quite a lot in your marketing materials, and it strikes me as a very enterprise company. Is it difficult—or how do you retain a real product focus and a real closeness to customers when you're so enterprise? Is that difficult?
5. Staying Close To Customers
I think it's a little bit of a false choice you're implying there. I think being a large enterprise doesn't necessarily mean you need to be distant from your customers—in our case, our customers' customers.
Bret and I are constantly building agents ourselves. One of the more interesting things of the last 6 months is that we released Ghostwriter. This is an agent for building agents. It's agents all the way down. It's pretty cool.
We are constantly in the products ourselves. Bret is actually still an extraordinarily capable software engineer. It's remarkable. Some of the code that's in production, he has written.
I’ve probably got a couple of lines here or there, but it pales in comparison. Of course, we can’t on our own simulate the complex, multisystem environments that characterize many of our largest enterprise customers, so we have to simulate them in our heads. But we’re in the product.
One of the things I think a lot about is that we will, in short order, be one of the larger B2C companies. We’re doing that via our customers. We’re serving hundreds of millions of interactions—soon, billions of interactions. So staying close to the end experience as well—the voice fluency, latency, quality of the experience, all of that stuff—is very energizing and something that we’re close to. I don’t feel distant from the product, either from our customers’ perspective or from their customers’ perspective.
I always looked at the space itself, and I was like, “Amazing space. What a huge TAM. What a problem, and AI is perfectly suited for it.” Then I peek under the covers, and I’m like, “Oh my God.” There are 15 companies funded with $100 million, Salesforce, Atlassian, Zendesk, and all the other incumbents. What is the market maturation of this space? Help me understand how this evolves over a 5- to 10-year period.
I think, first of all, to state the obvious, the great thing about being in a giant market is it’s a giant market. The challenging thing about being in a giant market is it’s a giant market, and other folks know it too. It’s startups, long-standing companies, and the incumbents.
Your point about it being competitive is certainly right. Five or 10 years, especially in the age that we’re in, is a long time. What I would point to is that, amongst the startups, customers are voting with their feet. We are multiples the size of our next-nearest similar-vintage startup competitor, are growing faster, as I said, and are working with many of the great companies in the world.
Do you think it’s like an Uber–Lyft market, or do you think it’s an AWS, Google Cloud, Azure market?
It’s hard to know. I think because the economies of scale in terms of depth and breadth of platform, experience in specific industry verticals, and so on really compound, my hunch is that it will be more like an Uber–Lyft market. We obviously think we’re in the pole position to be the bigger of those two.
Again, you sell to some of the biggest enterprises in the world. I had a guest on the show the other day say you can’t sell to enterprise without an FDE motion.
Hmm.
Would you agree with that, knowing all that you know now, selling to 40 of the 50?
6. Forward Deployment Wins
I would like to think that, at least in the AI space, I rediscovered and borrowed this model from Palantir, and we came to it almost accidentally. We started the company, and the first thing we did was reach out to people we trusted to understand what the biggest unsolved problems were that they were looking at. We saw, “Oh, interesting—service and support as a foothold into something much broader: helping support customers across the entire lifecycle.”
We then enlisted half a dozen design partners that we built the first version of our product and platform with and for. These are, in the history of the company, legendary companies: Olakai, SiriusXM, Sonos, and Weight Watchers. We built the first version of our platform with our engineers deeply embedded inside those companies—so much so that our founding engineer, Mihai, was actually an employee of Weight Watchers, including getting performance-review-time emails and so on.
What we realized was that no one had ever deployed an AI agent. No one had ever put AI in this way in front of their customers. In order for us to build the best thing as quickly as we and our customers would like, being so close to the business—the mechanics of it, the people, their business model—so that we understood it, I won’t say as well as our customers, but approached that level, we saw so much power in it.
Starting in early 2024, we really started building out this forward-deployed team. Customers use it in widely ranging ways. Our platform is highly extensible and very transparent. You can see exactly how an agent is built. You can export agent definitions and completely build your own, so there’s no need for forward deployment if you don’t want it.
What we generally find, though, is that in getting started, having Sierra and help from our teams drive while our customer is in the passenger seat—but navigating for the first version—is what has enabled us to take companies like Next live in 6 weeks, from kickoff to live behind their phone number and chat. Or Cigna, one of the largest healthcare companies in the world, live in, I think, 58 days.
Time to market, time to impact, time to value, and then the quality of the result—we think it makes a big difference. I wouldn’t say it’s binary, though, as you framed it. I do think you can sell without a forward-deployed team. But for getting to the impact of this technology as quickly as possible and at the magnitude that we know is possible, it’s an important catalyst.
Are we at a unique time in history where, for this specific moment in time, every buyer is in the market for the product? Normally, not everyone is in the market for a product at the same time. Every CEO is being told by their board, “How are we using AI?”
Yeah.
Is it a unique time because there is a buyer pool like never before for this specific moment?
There is effectively unbounded demand, I think, in 2 areas. One, we’ve talked about coding agents. The other is the space where we’re the category leader. One of the reasons we’ve grown as quickly as we have is to meet that moment and meet that demand.
We’re now 100 people here in Europe. We recently acquired a company in Japan, Opera Technologies. You and I were talking about this. To hit the ground running there and to have a team that can be attuned to the cultural nuances of Japan and the concept of omotenashi, which is extreme hospitality—that is what is expected in Japanese service, and that’s what we intend to build there.
Starting with the beachhead in customer support and customer service, to scale into the company you want to be, you have to move out of customer support into complete lifecycle management, I guess. Is Sierra a sales platform in the future? Is it a conversion platform? Is it a marketing platform? What is it?
I think Rocket is actually a pretty good indicator of the direction that we’re headed. You think about the life of a Rocket customer: it begins with search and discovery of a home they might want to buy. We worked with Redfin to rethink their search experience. We help Rocket reach out to folks who’ve expressed interest in a refinance and make contact that way. We worked with them to build Rocket Assist, to help bring people in and help them shape and size their loan, gather all the information needed, and so on.
None of that is service and support. We do do that—loan servicing and so on. So I think that’s a good example of where things are headed.
That’s an inbound sales machine.
Inbound and outbound, you’re right. It’s not just Rocket alone. With Next, we worked with them on personalized product recommendations. How do you help someone build an outfit and a bigger basket of things they will love? Again, that’s much more sales than support.
This sounds and feels more like a Fortune 50, Fortune 500 Palantir, but more consumerized, where you’re building these amazing solutions for these products to fit their needs. Is that unfair of me?
First of all, Palantir is an amazing company. We have taken a lot of inspiration from them and copied elements of their forward-deployed approach. My understanding is that Palantir has a low hundreds of customers. We are, and intend to be, at a much larger scale than that, and I think where that will come from in particular is building real domain expertise in specific industry verticals.
Of course, our first, second, and third customer deployments were, by definition, unique and one of a kind. I think we’ve learned some things about how to help build a basket in the retail setting and, in some of these other industries, how best to handle a question about the status of a healthcare claim, a healthcare insurance claim, or questions about a fee around a checking account.
I think we’re going to have these deeper and deeper lessons in specific industries and be able to apply those in a much more scaled way.
Will you build products that aren’t uniformly applicable across customer bases? If one customer needs specific cart-abandonment product features, is that something you build? Or is it, “No, that’s not applicable to the platform”?
One of our approaches in building the company—and this goes back to where we started—is that you can build a platform and hope that people come, with the applications getting developed on it, or you can build applications to inform a platform that makes building the third, fourth, and fifth that much easier.
Wherever we can, we’re scanning for opportunities to strengthen our platform. Commonality is much better than something that is truly one-off. That said, if we’re working with a Fortune 50, Fortune 20, Fortune 10, or Fortune 5 company and there is some element of that company that is literally unique, of course we’ll build that.
One of the neat things is that it’s actually become feasible to build that because of coding agents and the pace at which you can move.
There’s a real unlock in being able to build a solution on an already deep platform, but extend it in ways that may apply to a single customer. My hunch, though, is that if you build it for one, someone else is going to have that same problem, right? And so it’s less common than you would think—true one-of-ones.
I do want to go to the way that you run the company. It was so important in so many of my conversations before this. If we start with the board meetings, I spoke to, as I said, many of the investors. Every 6 weeks, not every quarter. Can you talk to me about your biggest lessons on how to really get the most out of your board and run the best board meetings?
7. Running Sierra With Intensity
We do a couple of things. You mentioned the 6-week cadence. We have kind of a tick-tock: a 3-hour meeting and a 1.5-hour meeting. We’ve done this since the beginning of the company because we could just see that if you’re on the AI time clock, it moves a lot faster. Things are changing.
Most recently, we came back from winter break and suddenly coding agents were amazing. You had Claude 4.5 and Codex 5.2. There was a fundamental step change in the capabilities of these models. It changed our approach to software development. It changed our approach to the core product. And so having a cadence where you can take in information, even from the last 6 weeks, update your priors, and then change course, I think, is quite important.
As for running the board meetings themselves, we don’t have board decks; we have board memos. Bret and I write a usually 6- to 10-page memo. There’s a saying, “Writing is just thinking on paper,” and I think it’s very hard to hide from writing. Getting our thoughts clearly out onto paper, sending that in advance, giving each of our board members some kind of soak time to think through the issues and come prepared rather than be presented to and managed, I think, is a big part of it.
The contents of the board letters themselves, I think, are notable. We’ve done quite well in our first 8 quarters in market, and generally, the format of a board letter is like, “We exceeded forecast by a wide margin yet again. Things are going well. We landed these customers, and here are the 7 things we think we could be doing better, where we’re unhappy, where we could be going faster, where we need to hire in this area, and so on.” The board meetings then kind of take form on their own based on that.
You get the scaffolding right. You get the people right. You set the table with the big questions we’re asking, and then genuinely invite our board members in to challenge us and improve and sharpen our thinking. Those are some of the ingredients.
I heard that you also write about everything that you suck at.
Yeah.
What was one of the most memorable writings on what you suck at?
One early on was that we had such good indicators of the demand we were going to see, and we just didn’t hire fast enough to meet that demand. It was like, “We could have taken on this additional set of customers.” We had the data in front of us. We could see it, and we didn’t act decisively enough to build out a recruiting team and scale faster.
This was early 2024, so early days in the company. We’ve since corrected, but that was one that stands out.
We are going to go to hiring. You mentioned some of the people around the table.
Mm.
You and Bret—I mean, it’s the dream team of the best of the best operators. You can choose any investors at almost any price, which is kind of hard. How do you and Bret sit down and discuss price on a new round? Because investors will pay anything to get in.
Mm.
You want it to be high, obviously—
Yeah.
—but also not too high. How do you actually think about that? Is it like—
Yeah.
—okay, three is on next year's target? What does it look like?
It’s generally been inbound, is the answer. We think about it, honestly, not in terms of valuation. We think, what is the amount of capital that we need to raise to get to the next unequivocally higher watermark in terms of revenue, company scale, and so on?
We think of it as milestone-to-milestone funding. Then we’re sensitive, but not maximally so, to dilution. How do you balance those things? Every one of our rounds, we actually guided to and took a lower price than we could have.
I again spoke to them, and they said the values within the company are craftsmanship, intensity, and family.
Mm.
Craftsmanship, intensity, and family are 3 that I wouldn’t normally see.
Mm.
Can you talk to me a little bit about why those are so important?
Mm. I’ll start with craftsmanship. Both Bret and I, just because of the way we are, care about doing things well. If you’re going to do something, do it with excellence. There are 2 ways in which doing things with excellence means much more than just sweating the details.
One is, what is a great company? A great company is an aggregation of thousands and thousands of things that are themselves great. It’s great people. It’s processes that are well-designed. It’s a great product. It’s a great culture. So how do you build an excellent company? You build everything with excellence.
I think holding ourselves to the standard—if it is worth doing, it is worth doing well—is one part of that, because that adds up to a great company.
How else do you get there? The other is, you think about what our customers are trusting us with. Back to the trust value, but I’ll make the connection with craftsmanship. It is their most precious asset. It is their customers. How will a company know—how will a set of people who are considering working with us know—how we will show up with their customers? A lot of it is how we show up with them.
Sweating the details in how we show up and interact with our customers, the level of professionalism and care, dropping everything when something matters. I’ll give you an example there. In our first Black Friday/Cyber Monday with a set of retailers, one of our lead engineers, our head of operations, and either Bret or me were in real time personally reading every single conversation that our agents were having.
We wanted to make sure we were doing right by our customers. So that’s craftsmanship, and again, it adds up to a great company. It’s very meaningful in helping our customers understand the care we will have for their customers.
Intensity and family.
Yeah.
Can you expand on those? Again, 2 that I don’t often get.
It’s back to this great thing about giant market—giant market; hard thing about giant market—giant market, and others are in it, too. I think there is an inevitability to companies interacting with their customers via really sophisticated agents that capture all that they know and all they can do on behalf of their customers, get the job done on their behalf, and handle the complexity, as opposed to pointing you to websites. The conversation is the interface, right? I think there’s an inevitability to that.
Therefore, in order to win, in order to build the best company in the space, it is about pace. It is about winning. It is about building the best product. It is about being competitive and being intense about it. We don’t have the luxury of patience. There’s nothing written in the wind, right, that any particular company will be the company.
Showing up in our 5th engagement and 500th engagement as intensely as we did our 1st—you have to do that. I talk about the Venn diagram of who we hire for: smart, nice, intense. It’s hard, actually, to get all of those 3 in a single person. When you do, it’s fantastic. You can feel it in the office. Another way of translating intensity is doing things with excellence and doing things with pace. It relates to craftsmanship as well.
Is there anything that you can do or add to an organization to increase or maintain intensity, be it timelines, rewards, or incentives? How do you keep intensity with scale?
I think it starts with the founders. Bret and I are quite intense. We have to be the pacesetters, right? We have to be the examples of intensity. It shows up in how we manage the company. He and I are deep in the details and are constantly asking, “Is this good enough? How could this be better? How could we go faster on this? Why can’t it happen tomorrow instead of next week?”
I think it has to start with the founders, as one.
How do you determine what you should be in versus what you shouldn’t? We’ve seen the resurgence of founder mode, of founders being in the weeds.
Yeah.
It’s also not possible in everything, and it’s not right in everything. How do you determine that?
You have to edit it. You have to have judgment for it. I think you have to look at what is the thing that is not going to happen, or won’t happen as quickly, without direct applied force from one of us or both of us.
It’s pointless to be in, quote, “founder mode,” 17 layers into the details of something that doesn’t matter. It matters a lot if it’s our next-generation agent architecture and there’s something that we can add. We try to be selective about where we engage at that level, but it’s anything but hands-off management.
So I think it starts with this: a founder's ambitious goals have a way of becoming self-fulfilling. You set out a goal, whether it's the quality of a product or a revenue number, and ask, “What would have to be true in order to get there?” Let's suspend disbelief and just imagine: What would have to be true to cover this much ground this quickly? Why can't we do that? Okay, why shouldn't we?
Japan is an interesting example of that. Why can't we have a giant business in Japan this year and not next year? What would have to be true? We'd have to have 10 people on the ground. I was like, “Why don't we buy a company there?” You see how this stuff hangs together. Ambitious goals can take the form of a date, sure. Date-driven development can sometimes work. I think work is like a gas and tends to expand to fill all available space that you give it. And so there's a danger in setting dates as well, where it's like, “Well, we've got this long.” It may not need to take that long.
It ties to the third one: work expands to the room that you give it. I give everything to my work, and I love that. Third, family as a value.
Yeah.
I'm just interested by that one.
Yeah. Each of our values comes directly from Bret and me, and one of the best decisions we made was—I think it was when we were 5 or 6 employees—we spent half a day. Bret and I have a technique we call “think apart, think together,” where we'll initialize on a prompt, and the idea is not to groupthink one another. We want to get the best of our independent thinking.
So we did a think-apart, think-together on values, went off and spent an hour writing up what our view was, came back and compared notes, and there was, first of all, a shocking amount of overlap, which I guess shouldn't have been surprising in retrospect. We'd wanted to work together. We'd been friends. I think we were deeply similar in many of our values—really all of our core values, I would say.
Family comes from the fact that I've got 4 young kids, Bret's got 3 kids, and I married my high school sweetheart. I think for both of us, the only thing that's more important than Sierra is our families. Our belief is that you can be part of something that's growing fast. You can be intense about your work. You can turn on the afterburners when you need to.
It doesn't just mean kids. It's picking up your parents at the airport when they get in from out of town. It's going to a friend's extended birthday weekend. It's being at the parent-teacher conference or whatever it is. And I think there's too often an image of a sometimes performative grind in certainly Silicon Valley startups.
It's not that we don't believe in hard work. Boy, do we. Again, intensity. But it's in working smart and finding some balance that gives you space for—again, translate family to things that matter to you in the sense of your whole being beyond just work.
Are you literally able to work as hard, though, when you have a family and you have 4—I mean, Clay, mate, 4 kids? That's a lot of kids.
I find I work a lot. Of course, if you just magically handed me 15 hours in a week that I wasn't with kids, I could probably do something with those. I find I am intensely focused and efficient. Boy, do I get a lot out of every hour I have.
One of the things I've done is spend a lot of time on 101 going to and from the office. We are all about being in person. Coming out of the pandemic, it was something of a novelty being opinionated about being in person. So I spend an hour and a half, sometimes 2, on the road every day.
I now have a very complex networking setup that combines 2 cellular networks and a Starlink Mini so that I have uninterrupted, beautiful connectivity to and from the city every day. You're efficient. You get everything that you can out of every hour.
Why are you so opinionated about in-person?
In particular for a young company, I think it is—I won't say impossible, but very challenging—to build a culture, a set of shared norms, and camaraderie. So many of the things we talked about, like enterprise software as a team sport, feel great when you're part of an amazing team, and it's different when your connection to that amazing team is via a Brady Bunch of Zoom squares.
We have rituals that we've developed that we only would have developed if we were all working in person. I think, for younger employees, apprenticeship and mentorship happen. So much of what I learned, and I think so much of the initial conditions of my career, came from experienced people taking me under their wing or letting me, in some cases, literally look over their shoulder at how they were doing something.
I think there's an element of paying it forward that's important, and in-person has a role to play in that as well. There's a talk that I love by the renowned computer scientist Richard Hamming, “You and Your Research.” For any new graduate, it's probably the single best thing, on a per-word basis, that you can read. One of the central theses is: Find great people, work with them, and learn from them. It sounds obvious, but there's something deeply correct about it. How do we learn as human beings? We observe someone doing something, and we effectively copy it.
So find great people and copy them. That's how you accumulate skills and capabilities. One of the other points Hamming makes in this talk is that knowledge and hard work are like compound interest, and we all know that the earlier you start saving, because of the miracle of compound interest, it can massively change the trajectory of your life. And so any young person should be intensely focused on learning as much as they can as early as they can, locking in those lessons and capabilities, if you will. Because it is literally trajectory-changing.
Do you want to hear 2 funny things? One, we used to do 5 shows a week. I just worked harder than anyone else when I was starting out.
That's a lot of shows, Harry.
Yeah, it was 11 years ago, but it was 5 shows a week. And two, I didn't have 1,000 listeners per show for 3 years, and I never made a dollar on the show for 3 years. It was never about money or recognition. I only cared about using this as a method to learn from you, and especially when I was 18, it was harder to meet amazing people. So I completely agree with those 2.
There are a lot of young people today. You have kids; I can picture them leaving university, uncertain about where the world is—
Yeah.
—what to do. What would you advise them, knowing all that you know?
8. The AI Native Advantage
The obvious tsunami that's coming is AI. What are the implications for jobs? I think there's been a lot of concern, understandably, about what happens to entry-level jobs. How do you apprentice, and so on?
I think the unfair advantage that young people have coming out of university is that they've just had 4 years to spend effectively unlimited time. You've got to go to class and pass some exams and stuff, but you have huge control over your time and disposable hours. Coming out of university as a master of these AI tools, let me point you to 1,000 companies that would love to have you infuse what you know into how they're doing things.
I can't remember a time when a young person with no work experience but with the right mindset and experience using some of these tools has ever been so valued. Some of our most effective employees at the entire company are 22 or 23 years old and have been completely AI-pilled, with a comfort and facility with these tools that many of our more experienced folks don't.
Has the way that you hire changed for the profile that wins in this AI-pilled world?
Yes. We completely changed our engineering interview process. It now looks much more like this: Here's a kind of prompt: Think through an application you would like to build. Cool. Okay, here's $150 to spend on—choose your coding agent. You can use whatever setup you want. Use whatever tools you want. Bring your own laptop, bring your own tools. We're going to pay for your tokens. And then build it.
Tell us how you went through building it, and so on. So at least in engineering, it is an AI-native interview. Of course, we test for architecture, systems design, product thinking, culture, smart, nice, intense—the extent to which we think people manifest our values. But that's changed very significantly. I will be disappointed if, in no more than the next 2 months, not every one of our interviews has some strong AI-native component to it.
Do you think we are entering a golden age for cyber and for cybersecurity, given the proliferation of code generated by AI that may not be as secure as it needs to be?
In terms of importance, it's obvious to me that it has never been more important, given that the kind of offensive capabilities have just ratcheted up 5 notches. I think cybersecurity seems like a pretty good bet to me. The question is whether the offensive tools turn out to be defensive tools if actually Mythos and Codex 55 cyber, if they themselves are the solution, not a more narrowly focused cybersecurity product.
What was the most recent disagreement you and Bret had?
A couple of weeks ago, we were trying to figure out how to get something to move much faster in one space. Interestingly, we basically always converge. We're highly truth-seeking. We have a funny expression.
It's like, “This is correct.” Okay, what does that mean? From some objective, truth-seeking perspective, this is the right way to do it. So we try to get to: What is the correct solution? I was on one side: “I think we need better kind of process and structure around this thing.” Bret was on the side of people: Maybe we need different leaders, or a different leader in this space.
The answer, as with most things, turned out to be some of both, right? Turned out to be some of both. We started from the idea that it couldn't just be solved with people. No, it's not just people, and we pulled on those threads. This wasn't “think apart, think together” so much as just interrogating each other, again, with the goal of getting to the right and best approach to something.
When Bret says something, what are you like, “Yep, I'm sure. He's a G at that”? And when you say something, is Bret like, “Yep, Clay is the expert”?
Rather than dividing up the company, we think about majors and minors for every part of the company. Bret's majors are definitely sales and then engineering. He is really good at selling software. He is really good as a software engineer still. I major in what I've called the running of the company: operations, finance, legal, and so on.
I do a lot of first calls. He understands our most important contracts for things that are highly consequential in how we run the company. We have 2 nuclear keys that we turn on for those. Bret, having spent time at Salesforce, really learned from the best. Mark is extraordinary.
Is that a sell?
He's the best seller I've ever met.
Unbelievable.
Yeah.
The GOAT. The GOAT. Unbelievable.
How's the weather, Mark?
Have I told you about Agentforce?
But, honestly, respect when it comes to instincts on how to sell. It's, “Yep, okay, makes sense.” Bret's instincts on system design and architecture are second to none, and so I trust his judgment more than I trust my own. On people stuff, on building and running the company, whatever Clay says, I would go with that. So I think that's probably the rough yin-yang major-minor split.
Dude, I would love to do a quick-fire round, if it's okay.
Yeah, let's do it.
What was your biggest lesson from working with Sundar?
It's such a gift. Sundar has a remarkable ability to look at a problem from wildly different zoom levels. His dynamic range in thinking is second to none: zoomed all the way out, at the highest level of strategy—how is this gonna unfold over the next 5 years?—all the way into the details, the pixels, right, the drop shadows, the sound, the texture of something. And I have tried to emulate that.
Talk about surrounding yourself with great people or having the privilege of working for someone. I observed a leader who is extraordinarily focused on the product, the work, and building something great, and is also just a wonderful human being, deeply focused on the humanity and folks around him.
What does no one know about Google that you think everyone should know?
What people underestimate about Google is that when you have the alignment of an ambitious, enduring mission, incredibly smart people, and a culture that values truth and building in service of that mission, that company can kind of solve anything. People sometimes criticize Google for “1,000 flowers bloom.” If you have smart, well-meaning people caring for every one of those flower beds, and they're directed in the right way, it is quite a force for invention and discovery and building new things.
I got asked to ask you about your book list.
Oh.
I hear you read a lot. That's a shit question. Forgive me for it. What's the must-read for me leaving this conversation?
Oh.
Give me one.
David McCullough's The Wright Brothers. It's so good. It's so good. It's a tight history of, obviously, the invention of the first heavier-than-air aircraft. And to me, it is as accurate a portrait of entrepreneurship and invention as has been written anywhere.
The aircraft could not have existed without this kind of network of pre-existing inventions, most importantly, a lightweight internal combustion engine. And then it was: try, it didn't work; try, it didn't work. There are scenes of them stuck out in North Carolina being eaten alive by mosquitoes. It's the hardship and then the triumph of having built something that flies.
I've never said this before on a show, but have you seen a wonderful film called Those Magnificent Men in Their Flying Machines?
No.
I'm gonna send this to you.
Okay.
It is about mankind's pursuit of flight.
Sounds good.
It's amazing.
Fantastic.
Well, thanks for the tip.
It's like the 1940s, '50s.
That's great. I'm looking forward to that one.
Parenting.
Mm.
4 kids and an unbelievable operator-founder. What's your biggest advice?
First of all, having kids is the greatest gift. It is such a privilege. There are a few things I would say. First of all, you will be a changed and different person on the other side of holding your son or daughter.
It is, in my opinion, the single fastest rate of change, the single biggest change that anyone experiences in their life after they themselves are born—welcoming your first child. I try to carve out time for family dinner. We have many Sunday mornings—maker mornings—with 2 of my sons, where we block out an hour or 2 and build something at home.
And so, it comes down to rituals and discipline around making time and space. Anything important in life, in my view, is a product of clear goals and good habits. And so I think if you have a clear goal around how you want to be as a parent, and then habits that help you build towards that, I think that's a very important ingredient.
And then the other is making kids' interests your own. I'm terrible at basketball. My oldest son is an incredible basketball player. I am so proud of him. I go and watch him play, and he does things I could never do. Not only can't I do that, I could never do that.
And so I follow the playoffs. I've learned about the sport. I've learned about the best players. I've learned about coaching so that I can try to enjoy and support him in this interest more fully than I otherwise would be. And so I think for each of our 4, it's about being aware of what gets the synapses going for them, what they light up about, and then making that interest my own.
Would you say that's the same for your partner? Do you need to have aligned interests in a partnership, in a marriage, or is it good to have different ones?
I mentioned I'm married to my high school sweetheart. We will have been together for almost 30 years, and I'm not that old. I think a great marriage is a partnership. A partnership means you are working in pursuit of, in service of, some shared set of goals.
You asked about interests. I think having shared interests in what you are pursuing as a partnership is deeply important. Those goals include happy kids who grow into adults who can enjoy their lives and contribute meaningfully to those around them, building a set of values in one's family that are aligned with your own, and ensuring, as part of that partnership, that the other member in it themself thrives and fully realizes themself.
Their interests may be different, but there can be a shared interest in enabling each other to become the best that you're able to become.
Final one for you, but I do like it. What's the kindest thing that anyone's ever done for you?
I feel such gratitude to my parents. I'm sorry if it's a straight-down-the-fairway answer. My father was a career cardiologist. My mom's a quite talented quilt maker. Neither of them were in engineering or technology, and they saw that when I got ahold of my first computer, I just lit up.
Not really understanding what computers were about, my mom was good with them in the '80s, but it was not at all clear where they would go. But they could see that I was obsessed with them. They supported that interest to the hilt.
I remember going with my mom and dad to buy an early Power Mac. My dad was pushing: “Would you be able to do more if we had more memory in it?” I think I would be. And it's like, “Well, we should get more.” I was like, “Is this real life?”
My mom would take me out of school 1 day a year, and we would go to Ken's House of Pancakes and get breakfast. She would make up a doctor's appointment or something for me, and then we'd go to Macworld. I would get to spend the day at Macworld, which for me was like nirvana.
And so I feel such gratitude to them for seeing in me that interest and how I lit up about this thing that was unfamiliar to them, but that they then pushed and enabled. And, of course, there's a direct line from that to 18 wonderful years at Google, starting Sierra, and to today.
They must be very proud of you.
I think they are. I know that they are.
The thing that strikes me from this show is—I don't mean to be sycophantic; it's just—what a good person you are.
Oh.
Do you know what? I interview a lot of people, and they're brilliant, and they're intellectually brilliant.
You obviously are that. But what strikes me is what a genuinely good person you are, which is really very tangible. I really can’t thank you enough for doing this. You’ve been an incredible guest.
Thank you so much, Harry. I really appreciate it.