为何提出正确问题,是 AI 时代最重要的能力
- 传闻中的 200亿美元 NVIDIA 合作,从第一次通话到资金到账只用了约 3周。 Jonathan Ross 找到 Jensen Huang,原本是想买约 100,000块 GPU,部署 Groq 已经搭建好的 GPU-LPU 混合系统;Jensen 看到成果后,反而认为不如让所有 NVIDIA 客户都能用上。Ross 纠正了“走投无路才卖身”的叙事:Groq 这次并不缺现金,协议价值只比上一轮估值高出略多于 2x,而且 Groq 本可以按许可交易的金额融资。
- 技术核心是,LPU 与 GPU 是互补关系,而不是替代关系。 “18轮卡车还是最后一公里的面包车,你会选哪个?答案是两者都要”:受算力约束的矩阵乘法交给 GPU,受内存吞吐约束的交给 LPU,因为不存在“唯一完美的架构”,两者结合才能“击穿瓶颈”。很多人误判在于把 prefill 与生成拆给不同硬件,但这不是正确的切分方式,因为“生成 token 才是难点”。
- 速度让模型变得更聪明,而不只是更快——Ross 用自己在 Google 参与创造的 TPU 上运行 AlphaGo 证明了这一点。 Ross 给出的近似 Elo 是:AlphaGo 在 GPU 上约 3,200,Lee Sedol 约 3,550,同时提醒自己可能把首位数字记错了;对话中 TPU 的结果被描述为约 3,900 或 2,900。同一个模型只有在硬件支持更深层搜索后,才找到了概率约为 1/10,000 的第 37手棋。“思考得更快,就能思考得更聪明”,在 AI-to-AI 的代理流量中,速度还会以指数方式放大。
- 就在 3、4年前,快速推理还不受待见,Groq 内部也不例外。 有人一边说它没有价值一边离职,客户会问“为什么我需要一个比我阅读速度更快的 LLM”,Ross 还曾两次被团队劝退 LLM 机会,包括 GitHub CEO 直接打电话来询问用于代码补全的芯片。最终奏效的办法是让人亲自试用:一段 LLM 在 Groq 上运行的 X 病毒视频,做到了多年第一性原理论证都做不到的事。
- 正如 Senra 通过 Ross 的融资故事和凯恩斯选美竞赛所框定的那样,资本已经不再是 AI 里的制胜押注。 过去,创投机构抱团下注有其合理性:融资最多的创业公司看起来可能更占优势;但“历史上第一次,创业公司不再缺现金……投入更多资金不再构成优势,可人们仍然表现得好像它是”。反差在于,典型的西海岸 VC 放弃了 Groq,东海岸 crossover 基金却投中了 Senra 所称、规模几乎是 NVIDIA 此前最大交易 3x 的一笔交易。
- AI 时代真正决定成败的能力不是回答问题,而是提出问题。 “信息时代的成功,取决于能否回答问题;AI 时代的成功,将取决于能否提出正确的问题。” 每个人都要从 IC 转向“AI 的领导者”;Ross 认为学校课程应围绕真实社区问题展开,让学生把问题拆解成一组可以交给 AI 的提问。
- 代码的边际成本“正趋近于零”,软件创造将向更多创始人开放。 Ross 的 EA 不会写代码,却已经能搭建旅行应用;软件创作正像读写能力一样,开始向过去缺乏技术能力、但可能拥有良好品味的人开放。他预计,不依赖大型团队的个人创始人也能做出有价值的公司。
- Ross 目前刻意维持的不满,来自全世界的算力短缺。 “如果因为算力不够,治愈癌症要多花 1年,那就是我的错。” 这背后的经历是:Groq 曾在仍未有产品时距离现金耗尽只剩 3周,靠“Groq bonds”——以薪资换股权——留住了团队;80%的员工参与,其中约一半降到了法定最低工资。
1. 从电话到汇款只用 3周:传闻中的 200亿美元 NVIDIA 交易如何落地
- Ross 的说法是,Groq 已经完成 GPU 与 LPU 的整合,于是找到 Jensen,想买约 100,000块 GPU 自己部署;Jensen 看到了这套成果,认为不如把它提供给所有 NVIDIA 客户。“第一次提出这个想法的那通电话,距离资金到账大约只有 3周。”Senra 感叹:“所以 Jensen 行动很快。”Ross 回答:“当然,领先别人就靠这个。”
- Ross 用物流作比喻:18轮卡车和最后一公里面包车,“答案是两者都要”。LLM 处理一个 token 时,会进行多次矩阵乘法,有时受算力约束,适合 GPU;有时受内存吞吐约束,适合 LPU。“瓶颈到处都是……不存在唯一完美的架构”,所以把两者配对,才能“击穿瓶颈”。
- Ross 驳斥资金告急的叙事:距离死亡只剩 3周的经历发生在“很多年前”,这一次“我们没问题”。他说,协议价值只比上一轮估值高出略多于 2x,而且 Groq 有能力按许可交易的金额融资。
2. AI 使用 AI:速度、微支付与可追问的每日简报
- 速度之所以会在今天形成复合效应,是因为“人类可以等 1、2秒……AI 却会一直等在那里,因为它生成 token 的速度快得多”。一个代理可以把任务交给另一个代理——“AI 很擅长使用 AI,这才是真正的 agentic”——于是任务呈指数级扩散,链条上的每一环都在意延迟。
- Ross 认为,代理支付目前还处在早期:“支付系统还没有真正为这个场景准备好,但如果能实现微支付,支付笔数就会暴增。”他举例说,自己做 Signal/WhatsApp 代理的业余项目时,需要电话号码,就得先在 Twilio 证明自己是人类;如果代理拥有一笔授权预算,AI 本可以在预算范围内自行支付,无需他介入。
- Ross 的业余项目会反过来塑造工作方式,其中一个是个性化的“总统每日简报”:它从长篇文字逐渐变成标题加追问。“我通过 AI 学习时,读的不是一段静态内容,而是在和它互动”,就像玩 20 Questions。Senra 也提到,Spotify 联席 CEO Gustav Söderström 做过类似的个人播客:过滤愤怒诱饵和政治内容,同时呈现他所关注的人正在讨论什么。
- Senra 从 Ross 的推文中提炼出的核心命题是:“信息时代的成功,取决于能否回答问题;AI 时代的成功,将取决于能否提出正确的问题。”所有人都要从 IC 转向“AI 的领导者”;学校过去训练我们记忆答案,但现在“你只要问 AI,它知道答案。你只需要想出正确的问题。这是根本性转变”。
3. 领导力首先意味着有人追随,然后找到适合自己的形态
- Ross 的第一个原则来自 Jon Levy 的一本书:“没有追随者,就不是领导者。”领导力和投资一样,有创投、债权、种子、crossover、PE 等无穷多种形态。新创始人的错误,是照搬别人给的建议,执行一套“并不适合自己”的方法。
- Ross 很了解自己:和 Senra 两档播客里的那些控制狂不同,他从 18岁起就没拿过驾照——“我不需要控制驾驶,我想控制的是思考”——因此会招聘那些在大多数公司环境里可能表现糟糕、但能够自主行动的人。对职业早期的人来说,应该去一家领导风格能教给你、且你真正用得上的公司。
- 他坦言:“刚开始创业时,我是世界上最差的领导者之一。”这让 Groq 浪费了 3、4年。他把任务交给无法自主运作的人,结果所有事情都停滞了;等他事后发号施令时,这些指令对他而言“太不自然了,以至于他们不接受”。
- 解决办法,是把目标简化到一枚挑战币上就能写下:每个 Groq 员工都带着一枚写有“每秒2,500万 tokens”的硬币,同时只设定极少约束。“给一个人的约束越少,他就越有自由用解决方案让你感到惊喜。”Senra 联想到 Skunk Works 的 Kelly Johnson:“极致表现往往来自一个残酷清晰的优先级。”Ross 进一步强调,团队只有在能够以一种好的方式让你感到意外时才可能创新,而这意味着你不能把目标限制得过死。
4. NVIDIA 内部经验:砍掉一对一,以及信心带来的解锁
- NVIDIA 是“你能见到的最不政治化的大型组织”,Ross 认为原因在于 Jensen 从不只对某一个人说某一件事。一对一会产生不同解读和小圈子,大型会议则能形成统一信息。他的规则是:讨论某人的指责性邮件,必须把被讨论的人抄送进去——“否则你就是在允许政治发生”。
- Ross 从 Jensen 身上学到的第二课,是自己“花太多心思玩 3D 国际象棋”,而 Jensen 只问:“客户需要什么?那就为他们把它做出来,其他事情自然会跟上。”这也意味着,不要向客户兜售你自己都不相信他们需要的东西。Senra 补充说,Jensen 曾把管理约 60名直接下属比作管理 AI 代理:每个人都在自己的领域比他更聪明,这本身就是一套管理范式。
- Ross 的信心故事发生在公司约 35人时:他跟随一位管理 2,000人组织的领导者,提前在心里做出每个决定,再看对方如何决策,结果一一吻合。“我的决定没有改变,但我的领导方式改变了。”人们会追随那些以信心传达出来的决定。至于规模,他说,一个 450人的创意组织在某些方面“更像是在管理 5,000人”;“人越优秀,就越难管理。”
5. 旅鼠、西海岸怀疑者与凯恩斯选美竞赛
- Ross 观察到的融资社会学是:“典型的西海岸 VC 更像旅鼠”——一个人拒绝后,所有人都会跟着拒绝;典型的东海岸 VC 则“都觉得自己比对方聪明”,会各自做分析。最终,Groq 由东海岸 crossover 基金出资;Senra 兴奋地指出,NVIDIA 这笔交易规模几乎是此前最大交易的 3x,发生在西海岸 VC 错过 Groq 之后。
- Senra 借用凯恩斯选美竞赛来解释这种行为:下注的不是最漂亮的模型,而是被下注最多的模型。过去,跟随其他投资人的选择似乎是理性的,因为最多人下注的创业公司看起来确实可能更有优势。但 Ross 认为,“历史上第一次,创业公司不再缺现金……投入更多资金不再构成优势,可人们仍然表现得好像投入更多钱就能给那家创业公司带来优势。”
6. 更快就是更聪明:AlphaGo 的证明与交易真正的起点
- 需要把功劳归给正确的人:芯片组合的想法来自 COO Sunny,这正是 Ross 自主性原则的体现;相关工作做了大概 3、4个月,“可能略长一些”。很多人误以为应该把 prefill,也就是读取,与生成拆到不同硬件上,但“生成 token 才是难点。读取比写入容易,AI 也是如此。”Groq 不介意把成果展示给 NVIDIA,因为它本来就想成为 GPU 客户。
- Jensen 之所以立即行动,是因为 LPU 就像“让现有模型瞬间接入宽带”。与现实中的宽带转型不同,这种提速不需要任何人重建网站,性能提升可以直接落地。
- Ross 从自己在 Google 参与 TPU 的经历中得出更深层的结论:DeepMind 在世界冠军围棋比赛前 30天发来邮件,AlphaGo 在 GPU 测试局中输掉后,团队把它移植到 TPU 上。Ross 给出的近似 Elo 是:AlphaGo 在 GPU 上约 3,200,Lee Sedol 约 3,550,同时提醒自己可能把首位数字记错了;对话中 TPU 的结果被描述为约 3,900 或 2,900。同一个模型获得了更多算力;GPU 没有找到第 37手棋这一万分之一概率的落子,是因为它在搜索链条中太深。
- Ross 也承认:“作为 TPU 的创造者,我必须承认 GPU 现在更强了。”胜负手在生态系统,但 LPU 可以让模型更快地搜索得更深。“思考得更快,就能思考得更聪明。”
7. 现实商、决定胜负的主游戏,以及把变革管理当成全部工作
- Groq 招人看重的是“现实商”(reality quotient),而不是 IQ:“有很多非常聪明的人,如果现实拍着他们的肩膀,他们也认不出来。”这种能力的极致表现,是找到决定胜负的主游戏:MySpace 最大化注册账户数,Facebook 最大化月活用户;“如果你最大化月活用户,就会击败那个最大化注册账户数的人。”
- 每秒2,500万 tokens 的目标,就是把这场决定胜负的主游戏变得人人可见:芯片速度、软件、电力成本、数据中心、制造和供应链,每个人都能把自己的工作与这个目标连接起来。
- Ross 从工程师转向创始人的关键转变是:“我的工作就是全职做变革管理。而变革管理的第一原则,是让它看起来不像变革。”只要员工的注意力锚定在目标上,方法变化就会被体验成一切照旧;领导者的职责,是提供足够背景,让他们相信“自己的工作其实没有改变”。
8. 运气回报,以及“我打算这么做”的修复
- Jim Collins 的观点是,最好的公司并不会获得更多运气,而是更善于抓住运气。Ross 用自己的反面案例说明这一点:GitHub CEO 曾打电话说需要用于 LLM 代码补全的芯片,因为 GPU 买不到;他的团队说“不行,做不到”,他还在两次机会中都让团队说服了自己,“尽管我内心深处其实知道我们应该做”。他反问:“如果我们当时成了 Microsoft 为 OpenAI 运行 LLM 的第一家推理引擎,情况会好多少?”
- 第三次遇到类似机会时,他亲自算了一遍,尽管所有人都认为不可能,Groq 最终“恰好达到了那些性能数字”。而在 3、4年前,快速推理还极不受欢迎,甚至有人因此离职。Ross 对那些茫然表情的解释是:“当人们不理解某件事的第一性原理,只是因为炒作才参与进来时,他们的理解不足以看出你做的事情为何不同。”
- 营销上的突破和 ChatGPT 时刻如出一辙。Ross 曾在 ChatGPT 发布约 3个月前看过 Anthropic 的演示,但现场反应平淡;只有当答案与自己直接相关时,魔法才会出现。Groq 把速度放到线上,有人在 X 上发了视频,Ross 在挪威做演讲时发现查询变慢了,因为使用量已经暴增:“我们突然病毒式传播了。”
- 管理上的修复来自 David Marquet 的《Turn the Ship Around》:意图式领导。问“我应该做这个吗?”会引来悲观意见;说“我打算这么做”,通常不会得到反馈,除非真的存在问题——就像那艘潜艇上的船员终于会说:“等等,舱门还开着。”第三次遇到机会时,Ross 说的是“我打算这么做”,结果“大家没有再说‘我们做不到’,而是纷纷加入进来,说‘我们应该这样做’”。
9. Groq bonds:让所有人都把手放在方向盘上
- Groq 曾在仍处于产品前阶段时,距离现金耗尽只剩 3周;当时公司正在开发一种此前从未尝试过的编译器,目标是消除人类手写 kernel 的需要。Ross 审阅拟议裁员名单后得出结论:“如果按这份名单裁员,我们就完了。”唯一的办法,是降低现金消耗,而不是裁掉员工。
- 于是,他召开全员会议,现场挂起二战时期的战争债券海报,并推出“Groq bonds”,即以薪资换股权。80%的员工参与,约一半员工把薪资降到法定最低工资;原本年薪数十万美元的工程师降到“50美元、60,000美元……是真正的痛苦”。这项计划节省了超过 3周的现金消耗,可能更接近 2个月;Groq 最终在只剩 3周现金时完成融资。
- 预期中的人员流失并没有发生:流失率低于 10%,或许更接近 5%。Ross 对此的概括是:“让所有人都把手放在方向盘上。”坐在车里的乘客会害怕蜿蜒道路,但驾驶者拥有更多控制感,也更愿意承担风险。
10. 按负面特征招聘,尽早锁定胜局,像 Jordan 一样挑衅
- Groq 维护一份带版本号的“people spec”:“如果你不把自己想找的人写下来,你就不可能按这个标准招聘。”候选特征包括运气回报和“诗性设计”——“诗歌就是语义密度,每个词都重要”——同时每个正面特征都有一个负面孪生项:挥霍运气、极繁主义设计。“你真正招聘的,是为了避开这些负面特征。”因为一个人可能把破坏性特质带进整个团队。Ross 最大的区分是:培养人才要展示正面特征,筛选人才要排除负面特征——这是“完全不同的思维模式”。他是在观察一位擅长清除问题的 HR 负责人后学到这一点的。
- 一项受重视的特质,是把损失厌恶转化为生产力。架构会议上,有人会说:“如果这么做,芯片速度会快一倍。”所有人都无动于衷,因为他们听到的是“下一代芯片”;Ross 听到的却是:“如果这代芯片不这么做,它的速度就会只有本来可能达到的一半。”他要招聘的是那些会“提前把胜利记到账上”的人。
- 对于 Senra 那期讲 Michael Jordan 的节目,Ross 怀疑 Jordan 沉迷下注和挑衅是有意提高赌注:公开输球会极其丢脸,从而逼出超人表现;创业者在成功前就宣布成功,使用的是同一种机制。Senra 引用了 Tim Grover 的说法:“一旦你告诉别人你会多么彻底地击败他,你就必须真的去做到。”
11. 制造不满、免费代码,以及教孩子提问
- Ross 在一群成功人士中观察到,有些拥有数亿美元财富的创业者仍然对财富不满,另一些人则对自己过去做出的产品不满。每个人都有某种驱动力来自不满。“如果你想持续推动事情向前,就必须拥有一种始终不满的性格。”Ross 现在的不满是算力短缺:“如果因为算力不够,治愈癌症要多花 1年,那就是我的错。”Senra 举了另一个例子:Edwin Land 会在白板上记录因每延误 1天而死于车灯眩光的人数。
- 乐观的一面是,“代码配给”正在结束——“边际成本正趋近于零”。软件创作正从技术专家控制的事情,转向更广泛的可及能力。Ross 的 EA 现在已经能搭建在线旅行应用;“很多人将获得创建软件的机会,他们过去可能永远不具备技术能力,但他们本来就有良好品味,也知道什么是好产品。”他预计,不需要大型团队的个人创始人也能做出有价值的公司。
- Ross 给父母的建议是:“不要再教孩子回答问题,开始教他们提出问题。”课程体系应围绕真实社区问题重建,比如许可证办理、地方活动,让学生写出真正有用的应用。如果学生只需上网查答案,或让 AI 直接解决问题,“你就没有教给他们真正需要的东西”;但如果给他们一个必须自己提出问题、再让 AI 解决的问题,“那你就教会了他们”。
Let’s start with this rumored $20 billion partnership that you have with NVIDIA. Can you talk about the structure of the deal and how it came about?
The most interesting part about it is that the call where the idea was first floated was about 3 weeks before the money was in the bank.
Oh, so Jensen moves fast.
Of course. That’s how you stay ahead.
So, how did it come about?
We had been working on integrating GPUs and LPUs together. The best way to describe why this helps is: if you were building out a logistics network for the United States and I told you that you could have either 18-wheelers or vans for last-mile delivery, which one would you pick? The answer is both, right?
GPUs and LPUs combined ended up giving better performance across the performance curves. We had implemented it, and we had gone to Jensen asking if we could buy about 100,000 GPUs because we were going to deploy them ourselves. Jensen saw what we had done and thought maybe it would be better to make this available to all of their customers.
You and I had this conversation at NVIDIA GTC, and you were talking about the fact that these technologies are very complementary. Can you explain a little bit more about that?
When you’re processing an LLM token, what’s happening is you’re doing all these different matrix multiplies. Some of them are more compute-constrained, and some of them are more memory-throughput-constrained. The ones that are more compute-constrained, we put on the GPU, and the ones that are more memory-throughput-constrained, we put on the LPU.
The bottlenecks are all over the place. There are all sorts of different bottlenecks. There is no one strategy, and there is no one perfect architecture. So, the realization was that you put these two things together and you defeat the bottlenecks across all of the different matmuls.
There’s another thing that I love that you said when we had this conversation: when AI is talking to other AI, speed is becoming more and more important.
A human can wait a second or two to get a response when they type a command into a computer. AI is just sitting there waiting because it produces these tokens so much faster. It thinks so much faster. Now you bring in LPUs, and the speed is so much faster that it just becomes all about how you move as fast as you can.
AI is really good at using AI. That’s really what agentic is, right? Humans benefit from using AI, and so does AI. Just like you would do research about me before I show up on your show, AI is going to kick off a job doing research on different tools it’s going to use while it’s using this other tool. So, it kicks it off to another AI. You get this exponential growth.
I’m going to go on a tangent just for a second, and then we’ll come back to this. I talked to a lot of founders about this recently, but the whole point is that when agents are making payments, the amount of payments that are going to be made is going to skyrocket. Do you have any insight into that?
Yeah, I think it’s still early, and one of the limiters is that payments aren’t really built for this yet. But if you can make micropayments, the number of payments is going to skyrocket.
I did a little hobby project, and for the hobby project I needed a couple of different phone numbers so I could have a bunch of different agents on with me on Signal and WhatsApp. I had to go to Twilio, and I had to prove I was a human in order to get the number and all this stuff, and it was just this big pain to get it done.
If, on the other hand, I could have just allocated a budget to the AI and it could have just used that budget, it would have spent it, and I would never have known. It would just have been within the budget.
Tell me more about the hobbies and the side projects, because this has come up in a couple of our previous conversations that you and I have had.
I like to do cutting-edge stuff on my personal computer. I don’t have access to the work codebase for things where there’s risk. I’ll spin up a server in GCP or AWS, and I’ll just start building stuff.
I’ve built everything from apps that tell you what airplanes to take if you’re traveling, what routes, and which routes have the best seats, all the way to apps that do certain mathematical things, like a daily brief, and all of this stuff. Very simple stuff. At work, I use it very extensively, but I always start with a hobby project before I bring it to work.
What’s the daily brief?
Every morning, I get an email that tells me what’s going on in the world based on what I’m interested in and how I interact with it, and just a whole bunch of research. It’s very much like the presidential daily brief, except personalized for me.
And is this in text form?
Yes, you read it. It’s text with links that I can click through to learn more. The big shift I made on the daily brief, though, was that I started off by getting a whole bunch of text and reading it. Then I realized, “Oh, it’s AI. It’s done a whole bunch of research. It has the context. Why don’t I have it just summarize everything? Just give me a bunch of headlines, and I can ask follow-up questions.”
I just spent time and recorded an episode with Gustav Söderström, who’s the co-CEO of Spotify right now, and he did something very similar because he tries to avoid any kind of feeds. The whole thesis behind the organizing principle of Spotify is “time well spent.” He’s like, “Well, me just scrolling and getting rage-baited on X isn’t beneficial, but there is information I want to know.”
There are a couple of different places that his agents go out and essentially summarize stuff he might be interested in. He’s like, “Hey, don’t include rage bait. Don’t include politics.” But he also heavily emphasizes, “These are the people that I’m interested in. What are the people that I’m interested in speaking about? If the same people that I’m interested in are speaking about the same topic, I want to know about that.”
Then he can read it, but now he turned it into this personal podcast, and he actually listens to it in Spotify. It’s essentially his daily briefing, but he listens to it in 5, 10, or 15 minutes.
For me, the interactivity of it is important because it’s sort of like—have you ever played the game 20 Questions as a kid?
Yeah.
So, you can figure out within 20 questions what’s in someone’s mind. Questions allow you to distill down into what you actually care about. If you have a podcast, it’s a static form. You hear it, but if you can interact with it, then you can just get to the information you most want.
As I’m learning through AI, I’m not reading a static piece of content. I’m interacting with it. I’m glad you mentioned the thing about questions because I have a series of quotes that I’ve saved from you, and I love your tweets, by the way.
You said, “Success in the information age was about being able to answer questions. Success in the AI age will be about being able to ask the right questions.” Can you expound on that?
Yeah, and this also goes toward a shift: people are moving from being ICs, or individual contributors, to all being leaders—but leaders of AI. What really good leaders do is they don’t do the work themselves. They don’t have the answer themselves. They’re just asking the question.
They’re just considering everything they’re hearing, and then they ask the question that no one else asked, or that everyone is thinking but is afraid to ask. So, with AI, because it can go off and solve all of these problems for you, it can do the research report. The question that you ask determines what you get, and that determines the output.
In the information age, we were all trained to just answer questions. That’s what school is about: “Remember this, remember that, remember this.” With AI, you just ask AI. It knows. You just have to think of the right question. It’s a fundamental shift.
I want to go to something you texted me about your views on leadership. What is your description of leadership?
I got this from a Jon Levy book, but it didn’t detail it much. The first principle of leadership is you have followers. Duh, right? It’s that simple. You’re not a leader unless you have followers.
But when you think about leadership as having followers, it’s also like thinking of investing as making money. There are a lot of ways to be an investor. You can be a venture investor.
You can give debt, or you can take equity. You can be seed stage, Series A, growth, crossover, or do convertible notes. There’s private equity, and there are so many different ways to be an investor—public markets, right? It’s the same with leadership.
With leadership, the mistake I often see with new founders is they’re like, “How do I be a leader?” The problem is they don’t realize that there are an infinite number of ways to be a leader. So, they go off and listen to someone, get all this advice, and try to execute on it. It’s not true to them.
One of the things that was very different for me as a leader: most of the people that I hear on your podcast are control freaks. They want things done their way.
On Founders Podcast or this one?
Well, both, actually.
And so, had I tried to be an absolute control freak, that wouldn’t have worked for me. I’m one of these weird people where I can go to a restaurant and tell the waiter, “Just bring me whatever you think is best.” I don’t even have a driver’s license. I haven’t had a driver’s license since I was 18.
Why don’t you have a driver’s license?
Because I don’t feel the need to drive. I’d rather think. I don’t need to control driving. I want to control thinking. I want to focus. I want to be on my phone. I want to be doing something useful. That’s been the case since I was 19 years old.
I’m happy to delegate things in a way that others are not. One of the things that’s different about my leadership is that when I hire people, I hire very autonomous people who often go off and execute on their own. They would be terrible in most corporate environments, but I also can’t hire the same kinds of people. That was true to me. Had I gone the other way, I wouldn’t have been very successful. If other people go this way, they’re not going to be successful.
Then again, there are a lot of things that are very similar. Most of the best Silicon Valley leaders lead from a place of inspiring their people, as opposed to trying to get people to be afraid. But that is a form of leadership. There are plenty of leaders out there who are successful because they cause people to become afraid. You just have to pick the form of leadership that works for you.
When you’re picking where you’re going to work early in your career, you should probably work somewhere where you’re going to learn the lessons that are good for you. If you’re more of a person who can show appreciation and gratitude, you probably should not work for someone who inspires fear, because you’re not going to learn any lessons that you can use yourself.
Man, this is so important. This actually fires me up. Have you ever spent any time with Tobi Lütke?
I haven’t.
Okay. The conversation that I had with him on the show, I don’t know, 5 months ago—I still think about it every few days. In fact, we go through the past episodes and I’m constantly finding new insights from that. That’s why you see all these clips that we’re putting out on X from things that we did 6 months ago.
One of the things I love that Tobi said in the conversation was, “There’s not one right way to do things. There are probably 100 ways that could accomplish your goal. You have to do the one that’s based on you.” This is why you see all these different tech companies that are so different from each other. Apple has complete silos. Google has everyone accessing the codebase. They’re just completely different ways taken to the absolute extreme.
We just had Dana White on the show, and his whole thing was, “The first thing to do is know yourself.” So, I’m going to ask you a question. When you figured out your leadership style—wait one second—you have to really know who you are and what fits you. The second thing is what you actually want to do in life. Once those are your 2 biggest questions, then you just wake up and get after it. Once you figure that out, you just wake up and attack step 2.
So, when did you figure out that your leadership style wasn’t the typical way and that you needed these autonomous people? By the way, if you ever started another company again, I would have to imagine it would just be you and a bunch of agents.
It would probably be AI. I think this is going to be a shortcut for founders in the future, and we can get into this, but the first thing that you have to do as a founder is go from the technical thing that you know how to do and that you can add value with to learning how to manage people. For me, that probably cost Groq 3 to 4 years.
Say more about this.
I was a terrible leader. I was one of the world’s worst leaders when I started. I gave people a little too much latitude because I’m more of a delegator, but I entrusted people who probably shouldn’t have been entrusted with that level of autonomy. I didn’t hire people who could operate autonomously, but I was naturally someone who would delegate and give autonomy.
What ended up happening was things would just grind to a halt because they wouldn’t know what to do, I wasn’t telling them what to do, and they were used to being told what to do. Finally, I would get so frustrated that I would go in there and tell them what to do, but it was so unnatural to me that they didn’t accept it.
Why is telling people what to do unnatural to you?
I work through questions. I like to set high-level direction. For me, my leadership style was to come up with a goal that was so simple that I could put it on a challenge coin and give it to everyone. It took me a while to get there, but everyone at Groq had a challenge coin that said “25 million tokens per second” and had a little graph of it going up. Everyone knew that was the thing to do.
It’s sort of like when you ask an AI agent to do something. The fewer constraints you give it, the more freedom it has to solve your problem. I liked to work with incredibly creative people who would come back with surprises.
Say that part again. Say that part again. This is important.
The fewer constraints that you give someone, the more freedom they have to solve the problem and the more freedom they have to surprise you with the solution. If you want to run a highly creative and innovative organization, then what you really want to do is minimize the number of constraints. But you also have to give them the things that matter.
If you aren’t able to very crisply distill what you’re trying to accomplish, then either you’re going to over-constrain or under-constrain the people. When you’re trying to do something as a founder, you are inherently trying to disrupt an old industry. There’s a moat. You’re trying to do something differently. If you’re not doing something differently, what’s the point? There are already these well-established, well-funded companies.
If you can bring a team together that is also disrupting and innovative, and it’s not just you, then it goes from being Superman to being the Avengers. That was just my natural leadership instinct. Had I been more of a command-and-control person, I should have doubled down on command and control. You have to just do what’s natural.
Have you spent any time studying Kelly Johnson, the guy who did Skunk Works at Lockheed?
Not too much.
Okay. He has this great quote that reminded me of what you just said when you did the challenge coin for Groq. He says, “Extreme performance often comes from one brutally clear priority.”
Yes, yes. You see this because if you don’t give people a clear enough but under-constrained objective that allows them to surprise you in a good way with the result, then you are not giving them the ability to innovate. I want to double down on that for a second. The only way for your team to innovate without you being the innovator is that they must be able to surprise you in a good way, which means you must not over-constrain the goal.
And that—you’re saying it took you 3 to 4 years after the founding of Groq to figure out?
That was in basic management, dealing with people. Little things like one of the great lessons that I’ve learned at NVIDIA. Actually, let’s take a lesson from NVIDIA, because now I’m there. Jensen is world-class. One of the things that I observed, because I’ve worked at other tech companies, is that there’s no politics. It’s the least political large organization you will ever see.
I learned this lesson at Groq, but I didn’t take it to the extreme. Seeing it in the extreme shows me just how valuable it is. There is no circumstance at NVIDIA where Jensen has one-on-ones with people and tells them one thing.
I learned this at Groq because what would happen is I would have a conversation with one person, and then I would have a conversation with another person. Both of them heard very different things, would talk to each other, and come to very different conclusions. But when I had a room full of people and said something to them, it’s amazing how they all heard the same thing.
When you are leading groups of people, if you want to reduce the amount of politics and people going off and forming side cliques and all this, stop having one-on-ones. Have big meetings with everyone you want to tell something to, and tell them all at once.
And don't allow anyone to send you an email. Copy everyone on the email. The moment someone sends you something—if someone says, “Hey, this person's screwing up on this thing”—copy that person on the email and let them jump in. Otherwise, you're allowing politics to happen.
What else have you learned from Jensen?
I got way too cute trying to play 3D chess. Jensen is very much just, “What does the customer need?” I want to develop trust with the customer. I want to always tell the customer things that are true and that I believe and can support. If I have a thing that isn't what the customer wants, I'm not going to sell it to them. I'm going to sell things to customers that they actually need and that I believe they need. I'm not going to think, “How do I build moats? How do I do any of this?” It's just, “What does the customer need? Just build that for them, and everything else follows.”
I've done a few episodes of Founders Podcast on Jensen. One of them was “How Jensen Works,” which essentially strips away all the biographical information that's in the book The NVIDIA Way. I think there are about 19 main ideas that I cover in that podcast.
One thing we just mentioned earlier is, if you're a founder today, what does founding a company look like going into the future? It might just be you and a co-founder and 10,000 AI agents. Jensen has this great line where people are scared of managing AI agents that might be smarter than them. He's like, “I already do that.” He has, I don't know, maybe 60 direct reports. He's like, “Every single one of them is smarter in their domain than I am, and I have no problem orchestrating them and managing them.” I thought that was a great metaphor.
One of the things that really good founders do goes back to asking questions. A really good founder is able to do this: even though it isn't their domain, someone comes to them and says something, and they ask a question. That person is like, “Oh crap, I didn't think of that.” Right? Just over and over and over again. It's a skill you can hone.
This is in every book on Bezos.
Yeah. But I think this is universal. I think any good founder is able to do this. Again, going back to those early stages of going from a non-founder to a founder, you're going to end up hiring people, and they're going to be like, “No, no, I'm the expert in this area. Just trust me. You're a kid, you know, just trust me.” And you have to learn confidence.
One of the things that happened for me that was very helpful early on in being a founder was that, back when we were maybe 35 people, I got to shadow someone who was running an organization of 2,000 people. It was funny because there was no NDA in place or anything, but every meeting we went into, he's like, “Oh, this guy's got an NDA. Don't worry, you can say anything in front of him.” And I was just like, “Okay.”
We would go into every meeting, and I was sitting there silently thinking, “What would I do?” At the end of each one, when he would make his decision, it was exactly what I would have done. What I hadn't realized before that moment was that I didn't have the confidence. Part of the problem with being a leader was that I needed to have confidence in a decision so that other people would have the confidence to execute.
Many people have too much confidence. Some people have too little confidence. If you have too little confidence, you're probably the kind of person who thinks through things a lot more, but you still need to get to a point where you act with confidence. When I realized I was making the same decisions as this very experienced founder or CEO, I started acting with confidence, and people started following my direction much more. I didn't change my decisions, but it changed my leadership.
How many employees did you have at Groq when you did the partnership with NVIDIA?
About 450, but it was more difficult to manage than a typical group of 450. In the military, depending on your rank, you're allowed to have a certain number of people reporting to you. The higher your rank, the more people you can have as an officer. But when you have scientists reporting to you, the number is actually much smaller—dramatically smaller.
Because I had such a creative organization, it was much harder to manage them. The problems manifested. So it probably was more like managing a group of 5,000 people than 450 in many ways. In other ways, it was like managing an even smaller group because innovations would just happen on their own. The better the people, the harder they are to manage.
Can you explain the state Groq was in?
Yeah.
I think you were at one point close to running out of money. Is this not accurate?
Early on at Groq—and this is a lesson for founders, because if you're doing a capital-intensive business, you're going to need to raise money—one of the things we went through was that we had raised money from some VCs who fell out of favor with other VCs, and others didn't want to co-invest. So every time we tried to raise, we struggled. There was also a little bit of bimodality in how VCs behave.
Can you say more about that?
The way I like to put it is that typical—not all, but typical—West Coast VCs are more like lemmings, and typical East Coast VCs all think they're smarter than each other. So when you try to raise from the West Coast, if one VC puts money in, all the others want to put money in. In New York, one VC investing means nothing. They're going to run their own analysis; they really do not care what other VCs are doing. The flip of that is, if you're on the West Coast and one VC passes, they're going to go tell every other VC, and like lemmings, they're all going to pass as well.
We had this problem where all the VCs on the West Coast didn't want to invest in us. We had very few of the typical VCs invested in us at the end. We had a bunch of crossover funds from the East Coast invested.
Hold on. That's actually kind of hilarious. The biggest deal NVIDIA ever does by almost 3x, and the West Coast VCs missed it. They all chose other things that were safer. So there's this thing called the Keynesian beauty contest. Have you heard of it?
No, I don't think so.
Okay. John Maynard Keynes, the economist. The idea runs in parallel to VC, and when you see this, you can start to understand some of the behavior—some of the lemming behavior—because it's actually a good idea to follow other investors.
Imagine I give you a magazine with a bunch of models in it, and your job is to bet on models and say which one is the most beautiful. But the determiner of the most beautiful model is not who's most beautiful; it's who has the most money put on them. For that model, pro rata to how much money you put in, you get all of the money that's been bet on all the models. If your model doesn't win, you get nothing. You lose all your money, and it goes to the winner.
If that's the case, if someone has a lot of money, they can put it down on any model they want, whether they're beautiful or not. A bunch of people made other bets. It doesn't matter. They just come in and make a big bet.
Now, when you're watching some of these wagers that are happening in Silicon Valley, with people putting in big bets, it's because that used to be what won. What's changed, though, is that unlike the Keynesian beauty contest, where the winner is the one with the most money, in reality there's a point at which you get enough money and you don't need more.
For the first time in history, startups are not starved for cash. They have all they need and more. So now everyone's getting funded to the level that they need, and putting more money in is not an advantage. But people are still acting as if putting more money in gives that startup an advantage.
Okay. So can you tell the story that you told when we were together at GTC about the drastic change in Groq's fortune and how fast that happened?
Yeah. Going back to what you asked about almost running out of money, we were about 3 weeks from running out of money at one point.
But that was many years ago.
That was many years ago.
You weren't close to running out of money this time.
No, no, no, no. We were fine by then.
You were fine, but the last valuation you raised at compared to the value of this agreement was drastically different.
It was really only a little over 2x, so that wasn't a huge jump. We also had the ability to raise at the amount that we did the licensing for.
Well, let's just go back to what we talked about on stage, where it's just like how fast you guys had this insight and then how fast the trajectory of Groq changed.
It was that 3-week period from the time when we presented and asked to buy GPUs until not only the deal was done, but money had been wired.
But how many months before were you working on this?
It was probably 3 or 4, maybe a little bit longer. But what had happened was I didn't initially think that it was going to be that big of a deal, so I didn't suggest doing it. In fact, it was Sunny, my COO. This goes back to the whole autonomy thing.
This is the story I want.
Yeah. He had the idea of trying to put our chips together.
Explain what you mean by putting our chips together.
The LPU and GPU, as mentioned, are better at different parts of what's called the decoder layer of an LLM.
The GPU is better at the attention portion, and the LPU is better at applying the weights, which is the thing that gets trained as opposed to the memory.
What we realized—and this is what most people get wrong when they're trying to do this themselves—is that they'll take the reading of tokens, or what's called prefill, and they'll do that on one piece of hardware. Then they'll put the generation of tokens on another piece of hardware. But the generation of tokens is the hard part. That's the thinking. Reading is easier than writing, right? And that's true for AI as well.
What we figured out was—and again, this was all a bunch of people. It wasn't any one person; it was a group that all innovated—once he brought up the idea, “Why don't we put them together?” Because they're different bottlenecks, the team figured out, “This part goes here, this part goes there.” We implemented it, and it worked.
We also weren't afraid to show it to NVIDIA because we wanted to become a customer of NVIDIA. We wanted to buy GPUs. So what ended up happening was we went, we presented it, it made a lot of sense, and the deal happened.
Okay. So from Jensen's point of view, once he sees this—
Yeah.
Right. Why does he decide to act so quickly?
This is the nature of a successful entrepreneur. You move quickly. You don't wait, right? There's an opportunity cost to waiting. Technology is not a business where you can wait a year.
What I'm getting at is, why is that so important to his business?
Right now, when you go to use AI, it may feel somewhat fast, but that's because you're not used to using it much faster. When you first used the internet, it felt fast compared to mailing things around. But when you got broadband, you realized, “Oh my gosh, this is so much better. I'm never going to go back.”
The difference is that broadband actually needed people to make their websites faster for it to be usable. If the servers are slow, you don't get a benefit. So it took a while to roll it out, make it good, and get video that could stream and all that.
The difference is, you put these LPUs into a system and all of a sudden the generation of tokens gets faster. It's like getting broadband instantly on these existing models. Rather than having to wait a minute to get an answer, you can get an answer in 10 seconds, and that really starts to compound.
Let me walk through an example of why it's not just speed, but it's also quality. The other thing that I did was create the Google TPU. At Google, there was a time when I'd already moved over to Google X. I was no longer working on the TPU at that point. We had already done it, and someone else from the TPU team who was at Google X came by and showed me this email from DeepMind saying, “Hey, we've got this competition. We think we're going to lose. There's a prize purse. Is your chip as fast as we've heard?”
We're like, “Yes.” It's like Ghostbusters: when someone asks you, “Are you a god?” you say, “Yes.” Someone asks you, “Is your chip as fast as I've heard?” You're like, “Yes.”
So we replied back, “Yes,” and they're like, “Great. The competition's in 30 days. We're going to play the world champion. Go.” We played our test games, and we lost. We needed to win. So they had no choice but to port over to that TPU chip. We did it, and a bunch of interesting things happened.
I think—do you know what an Elo score is?
Yeah.
Okay. For those who don't know what an Elo score is, it's sort of your ranking in chess or Go. Something like a 200-point advantage is insurmountable. The probability of you winning is basically zero.
AlphaGo running on GPUs had an Elo score of about 3,200. Lee Sedol was about 3,550. I might be getting the first digit wrong. It might be 2,000 instead of 3,000, but it was more than 200.
And then, when put on TPUs, it actually went to something like 3,900 or something ridiculous—or 2,900, whatever the first digit was. It jumped dramatically.
I think he didn't expect he was going to lose, but he lost badly. It was the exact same model. What changed was that the ability to compute more made the results smarter.
The way that these models work is like the idea in Thinking, Fast and Slow by Daniel Kahneman.
Yeah, I read that book.
What AI does is, if I have 270 possible moves, which is what you have on a Go board, the AI is going to rank those moves and say, “This is the best move, this is the next-best move,” and so on.
What happens is, you virtually play that best move, and then you virtually play the counter move, and then you virtually play the next one. Then you see how the game unfolds. What you can also do is try that second-best move. Occasionally, that second-best move, when you play it out, actually turns out to be the right move in this context. It's not the one that you would normally do, but in the context it's better, and you can see that as you play it out.
In the second game, there was this famous move called Move 37, which was creative. It was original. It actually wasn't completely original. It was a 1-in-10,000-game move. It had been in the canon of games we were trained on, but when we went back and played it on GPUs, it never found that move because it was too deep in the chain.
Over time, these GPUs have gotten as good as and better than TPUs. As the creator of the TPU, I have to admit that GPUs are now better. This is the benefit of having an entire industry behind you, along with the ecosystem and everything.
But at the time, TPUs had some novel innovations. Now bring in the LPU, and you can go deeper faster. You can search faster. So you can actually make a model smarter by making it faster.
The realization is that now that you've got this ability to reflect, think deeply, and change the outcome based on your thinking, being able to think faster makes you think smarter. That's the advantage of pairing the LPU and the GPU.
Let's go into a couple of your ideas that you've discovered in the 10 years you were running Groq. What is your idea about the reality quotient?
One of the things that we hired for at Groq was reality quotient, which is different from intelligence quotient. The way to think of it is, there are plenty of really smart people who wouldn't recognize reality if it tapped them on the shoulder.
You've got to say more about that.
You just know these people who will construct these very elaborate stories in their minds that are completely disconnected from reality, right? Then there are some people who are incredibly street-smart but couldn't do basic arithmetic or anything like that.
Reality quotient often starts off as being able to recognize reality, but in its most extreme form, it's the ability to choose the dominant game that's being played. What most really successful founders and entrepreneurs do is recognize that everyone else is playing one game, and they realize that if you play this higher-level game, you win.
A simple example: MySpace was focused on the number of accounts signed up. Facebook focused on monthly active users. That was the dominant game, right? If you have monthly active users, that's more important than accounts signed up. If you maximize the monthly active users, you're going to beat someone who's maximizing accounts signed up. You're playing a better game.
As a founder, you're often able to do this better than other people. Your job leading those people is to try and help them connect their activities to that dominant game.
When running Groq, I said our goal was to get to 25 million tokens per second of capacity in our data centers. That meant everyone had a different way to contribute to that. It meant they could make the chip faster. It meant they could make the software faster. It meant we could deploy more. It meant we could get our power costs down so that we could deploy more chips for lower OpEx. It meant getting more data centers. It meant fabbing more chips. It meant finding ways to optimize the supply chain so we could put orders in faster and fill things faster for customers.
So everyone could connect what they were doing to that one dominant game. The challenging part of being a founder is that you often see that and tell it to people, and everyone wants to stick with the old way of doing it or care more about the process.
Moving from being an engineer to being a founder, the thing that finally clicked for me was that if I was going to do something disruptive, my job was full-time change management.
And the first principle of change management is to make it feel like it isn’t a change.
Why is that so important?
People do not like change. No human being likes change. The difference between people who appear to like change and people who don’t like change is often that they’re looking at different things. For the one who is fine with the change, nothing changed.
If I’m playing this dominant strategy here and that doesn’t change, I’m just trying to maximize the number of tokens per second that I’ve deployed. When my approach changes, nothing really changed. But if you’re down here thinking, “How do I make this chip faster?” and not thinking about the software, then you might view something that changes the chip as a change.
Your change-management duty is to give people enough context so that they see that their job hasn’t really changed. What they’re trying to accomplish is the same thing, and there is no change.
Tell me about “Return on Luck.”
I read Great by Choice by Jim Collins a while ago, and it had this chapter in it that really resonated with me. The thesis is that the most successful companies don’t have more lucky events; they just seize on that luck better than other companies.
I read this very early on as a founder, and I started to notice that it was true. There was a really good example of this when LLMs first started to become a thing. I remember getting a phone call from the CEO of GitHub, basically saying, “I need a bunch of GPUs. We’ve now gotten LLMs to do code completion. Even though we’re part of Microsoft and everything, we just can’t get GPUs. Could we use your chips, your LPUs?”
I went to the team, and I’m like, “We’ve got an opportunity. We can do this.” They were all like, “Nope, not going to work. Can’t run it on these chips.” I’m like, “No, it looks like it’s the ideal thing to run on our chips. It looks almost perfect.” They’re like, “No, it doesn’t work. There are all these things that are in GPUs that we don’t have.”
I’m like, “Yeah, but none of those are important for LLMs.” They were looking at the wrong things. I let them convince me that we shouldn’t pursue it, even though, in my bones, I knew that we should.
This happened another time, where there was another opportunity to deploy an LLM, and I let them talk me out of it. The third time, I thought, “No, I’m going to do it myself.” I went through the whole thing, did the arithmetic, and determined the performance. Everyone disagreed with me about whether it was possible. I’m like, “No, look at this.” In the end, we ended up hitting exactly those performance numbers.
The thing was, they were looking at the wrong things. They were looking at all the reasons why it couldn’t be done rather than why it could be done. We talked ourselves out of it. I had multiple lucky opportunities that I didn’t seize.
How much better off would we have been had we been the original inference engine running LLMs at Microsoft for OpenAI? That would have been a very different outcome. We lost a little bit, but we were still ahead of the curve when we realized that fast inference was going to be a thing.
I remember very early on, we would talk to potential customers. We even had a video where we sped up what it looked like, and everyone who looked at it was like, “Why do I need an LLM to be faster than I can read?”
You’re not going to be doing the reading.
You’re not going to be doing the reading, but also, that’s not how the internet works. Are you okay with a web page showing up?
I hate that.
Yeah, because eyes don’t move that way. Eyes move all over. You need the entire thing there. You’re going to look at it, and even before you read it, oftentimes you’ll have a sense of, “This isn’t what I needed,” and you’ll start typing your question without reading everything that came out.
So I realized that fast inference was going to matter.
No one else did.
Even within Groq—
What do you mean, no one else did?
Even within Groq, there was a lot of pushback. We had a lot of turnover at this point.
How many years ago was this?
This was probably 3 or 4 years ago.
Okay.
A lot of people were leaving, saying there was no point to fast inference. It wasn’t going to add any value to the ecosystem. Even though you can draw very simple parallels, like dial-up versus broadband, no one could connect it.
I’m going to interrupt you real quick. Can you say more about this? Now everybody’s just talking about fast inference. It’s like everything, but I wasn’t paying attention to this 4 years ago. I had the other stuff; I just wasn’t paying attention.
Can you talk about the difference? I think this is one of the most important parts of your company’s story: just how contrarian—maybe that’s not even the right word—but how out of favor your main idea was. People were like, “No, it’s not important.”
Well, everything was out of favor. Everything we did was considered a bad idea. But if you don’t do things differently, you have no advantage, right?
Why do you think so many people 4 years ago just didn’t understand?
When people don’t understand the first principle of something and they’re getting involved in it because it’s hype, they don’t understand enough to understand why what you’re doing is different.
So that had been a disorienting experience for you?
Yeah.
To keep saying this.
What eventually worked was—and this goes to a little bit of marketing we came up with—we realized that there was no possible way, no matter what we showed people, for them to accept that fast inference was going to be helpful unless we let them try it.
I remember this example from Eric Schmidt. He was involved in this thing, SCSP or whatever, and they showed off Anthropic, an LLM from Anthropic, about 3 months before the ChatGPT moment. I remember sitting there seeing it, seeing demos of this AI answering questions in the audience, and no one reacting. I’m like, “How is no one reacting to this?”
Now compare that to the ChatGPT moment, where everyone reacted. What was the difference? When people asked their question and got an answer to their question that was specific to them, that was magical. Seeing text show up for someone else’s answer wasn’t magical.
I realized that the only way we were going to get people to understand the value of speed was if we implemented this and put it on the internet. So we did. What ended up happening was that we put it online, and I remember I was doing a little bit of a world tour trying to find customers.
I was in Norway, doing a presentation, and I noticed that when I was doing queries using some of the open-source models, it just felt a little slow to me. Not that slow, but a little slower than usual. I’m like, “What’s going on here?” Norway’s farther away than the servers, but I had tested it earlier and it was fine.
I checked in, and our usage had skyrocketed. Someone had posted on X a video of an LLM running on Groq that was just running super fast, and it was viral. All of a sudden, everyone started creating applications using it, posting those, and it was just such eye candy when people saw it that everyone started creating their own. We just went viral.
Say more about this experience that you had, though. I’m really interested in how you’re coming at it from first principles. You’re saying these people are getting involved in it because it’s hype and it’s the thing being spread around at the moment.
What was that experience like for those several years when you were just going through this? You’re trying to explain why this is going to be important, and you’re just hitting blank stares or brick wall after brick wall.
There’s this common theme where a lot of really good innovators are innovators because they experienced the problem before others did. Remember my experience of AlphaGo on TPUs and being able to outperform the world’s best Go player only because of the hardware that we switched to.
Yes, I was able to get this return on luck more than others. I’m more able to say, “Okay, there’s an opportunity. I’m going to go for it.” But I was also exposed to the opportunities first. And you need both.
You have to be in a position to see the future ahead of others. This is the common saying, right? The future is already here; it’s just not evenly distributed. I was in a situation where I got to see the future because I was really willing to seize luck and double down on it when everyone else was like, “No, let’s not pursue this opportunity.” I’m like, “Yes, this is the opportunity.”
Those 2 things are what work together.
When you’re saying, “Yes, this is the opportunity,” would you describe the response you were getting as opposition or indifference?
One of the biggest shifts in my leadership came from another book, Turn the Ship Around! by David Marquet.
I read that too.
Okay.
I technically read books for a living.
I don’t know.
I adopted that very heavily in my leadership once I read it because it worked really well with my autonomous leadership style. The basic idea in intent-based leadership is that if I ask someone, “Should I do something?” they have opinions. Most people will be pessimistic and give you negative opinions.
On the other hand, if you express intent-based leadership and say, “I intend to do this,” people don’t tend to offer their opinion. But if it’s very wrong and there’s a reason, they will push back.
The example is the submarine commander who took over, I think, the USS Santa Fe. I think it was the worst in nuclear readiness in the nuclear submarine fleet. In a year or 2, he got it to number 1 in readiness, and all he did was shift from command and control to intent-based leadership.
The quintessential example would be to say, “Dive the boat.” There had been incidents where a submarine had dived while the hatch was open, and no one wanted to push back because the commander was command-and-controlling, and they were used to doing whatever the commander said. But when people would use this intent-based leadership and say, “I intend to move the boat down to 500 feet,” then all of a sudden someone would say, “Wait, the hatch is open.”
They’re involved in it now.
They’re involved in it. Everyone says, “I intend to do this. I intend to do this.” It gives everyone an opportunity to say what it is that they’re doing so people can hear it, but you’re not asking for an opinion.
The issue was that, earlier on, I was getting opinions from people, and they were stopping me.
Going back to the 3 examples in Return on Luck.
Yes.
Okay.
And if I had just said, “I intend to do this.”
Is that what you did in the third example?
Yeah. I literally put a presentation together and said, “We’re going to get to this particular speed per chip.” Rather than people going, “We can’t do this,” they all jumped in and said, “This is how we do it.”
It’s a very small change in phrasing, but it makes all the difference in your ability to move forward. You’re not inviting friction, but people would still give feedback when it was really important. When there was a real problem, they would raise it.
How does intent-based leadership tie to the question? If you were getting opposition or indifference, where were you going with that?
I was inviting pessimism by asking for people’s opinions.
Were you asking potential customers? These are teammates.
Teammates. Teammates.
Okay.
Almost everything that is difficult is difficult because you can’t go to one extreme or the other. You actually have to decide, in the context, whether you’re going this way or that way.
One of the difficult things is getting feedback. You hear this from leaders all the time: early in their careers, they get way too much pushback; later in their careers, they don’t get enough feedback. How do you balance it so that you’re getting real feedback versus just getting unnecessary pushback?
Part of that is this subtlety of phrasing: “I intend to do this,” as opposed to asking for an opinion.
Let’s go back to this time when you were 3 weeks away from running out of money.
Yeah.
You came up with this idea of bonds—of Groq bonds. Before I even get there, something just popped into my mind as you were speaking earlier about your leadership style. Basically, you’re going to tell them what you want to do. You have this organizing principle, but you’re not going to tell them how to do it. You’re going to let them surprise you.
Yeah.
That’s Phil Knight in Shoe Dog. He says that over and over and over again.
Then I’m reading about Groq bonds, and it sounds very similar to some of the things that Phil Knight had to do because Nike was so close before it IPOed. It had to IPO out of necessity because it kept running out of money or coming very close to it. He actually converted some of the loans that he got from his employees into equity and wound up doing very well for them.
So, explain this idea that you had for Groq bonds.
We were going to run out of money, and the leadership team that I had at the time was starting to put together a list of layoffs—who we were going to lay off. When I started reviewing the list, it became very clear to me that if we did that layoff, we were dead.
Why?
We were already struggling to keep up with what we needed to implement because this wasn’t even pre-product-market fit. This was before the product worked. We had to write a very special compiler that had never been written before, one that didn’t require human beings to write what are called kernels. It had never succeeded before. No one had ever done this.
Our architecture didn’t work with kernels the way everything else worked. So, we had to get to this critical-mass point before our product would even work. We hadn’t done that yet, and we were talking about cutting people who were critical for that.
When I realized that layoffs weren’t going to solve the problem, it was just a simple bit of burn math. We were going to run out of money, but we weren’t going to have the talent we needed to succeed, and we were going to have these other costs. I realized we had to reduce our burn without reducing our people, and the only answer was to get people to take a salary cut.
So, we had an all-hands meeting, put up World War II-looking pictures of war bonds, and called it Groq bonds. It wasn’t technically a bond. It was an exchange of salary for equity, and we expected that we were going to have pretty high attrition. We actually didn’t.
80% of the employees participated, and I think about half went to the statutory minimum salary by law. Remember, engineers get paid hundreds of thousands of dollars. These folks cut their salaries down to $50,000 or $60,000, whatever the statutory minimum was. Real pain.
We saved more than 3 weeks’ worth of runway. I think it was closer to 2 months. We had 3 weeks of money left when we raised. Had we not done this, we would have gone out of business.
The thing that’s interesting—I go back to this all the time because we’ve had a lot of close misses where we had to keep the team together—is a phrase I have: “Put everyone’s hands on the steering wheel.”
When people are passengers in a car, they’re more nervous about a windy road or a scary road. But when they’re the driver, they feel more in control. They’re more willing to take a risk.
By doing this, we put everyone’s hands on the steering wheel. They were participating in saving our runway, and we had less than 10% attrition. It might have been closer to 5% when we announced Groq bonds, which was actually probably better than our attrition rate before then.
I love this idea. The other side of looking for ways not to fire people is hiring. You also have some interesting lessons you learned in the decade you were building Groq about hiring that are pretty counterintuitive, which is true of a lot of the people who appear on the show. They have very counterintuitive ideas about hiring.
I was very good at hiring people who were incredibly smart and talented, but a lot of the people we brought in caused organizational problems. I’ve already alluded to that.
The reason was that I’m pretty clever, and when I meet someone, I can come up with a reason why I should hire them. I think a lot of people do this. They convince themselves, “I should hire this person. They’re great because of this. They have this experience. They have this attribute. I’m going to hire this person.”
So, we have this thing we call a people spec, or a data rubric. Very much like you have a product spec, we had a people spec. It had version numbers, and we would change it.
If you don’t write down what you’re looking for in people, you’re not going to hire that. You’re not going to be consistent. So, we framed the people spec in positives—things that you look for, like return on luck.
Give me a couple more examples of what the positives on that people spec were.
Poetic design. Poetry is semantic density. It’s when you say so much in so few words.
Which is actually really important to you.
Hugely important.
Yeah. This phrase—I think it’s on the Groq blog—was “Make every word count.” I think you’d repeat it over and over again.
“Every word matters.”
Or “Every word matters” there.
Yeah. The smallest possible, most minimal expression of the thing you’re trying to achieve is the most poetic. That’s not just in words; it’s also in design. Something is poetic even if it’s not words, and so that’s another one.
Each of these has a negative version. The opposite of return on luck would be “squanders luck.”
The opposite of poetic design would be maximalist design—just throw every feature in. Some of these products are like, “Where am I supposed to click?”
Mhm.
It’s really easy to spot people who fit some of the positives and not realize they have some of the negatives. What you’re really hiring for is to avoid those negatives, because if one person comes in with that negative, they’re bringing it into the whole team.
The biggest flip in my hiring was when I went from looking for positives, which is what you do when you’re trying to grow talent, to looking for negatives, which is what you do when you’re trying to select talent.
Explain that.
When I’m trying to help someone grow and improve, I want to show them the path. There’s a famous example of how to increase the amount of money given to a charity. It’s not making people realize how great the charity is, and it’s not making them feel good about giving to charity. It’s about telling them where to send the money.
If you tell them how—if you give them a skill or a technique—then they can very often learn it and do it. When you’re trying to grow people, show them the positive. Don’t say, “Hey, don’t squander luck.” Show them what return on luck looks like: There was this opportunity once that everyone said no to, and we said yes, and it made us successful.
When you’re hiring, you’re really looking to vet people, and you’re trying to say no to things. That’s a very different motion. Some people are really good at growing, and some people are really good at hiring. But you have to separate those two into very different mental modes.
The reason I noticed this at Groq was that we hired a head of HR who was very good at noticing problems with people and getting them out. As I observed her doing that, I realized I’d been hiring all wrong.
I think the way you described this to me was that you essentially inverted it, and now you’re hiring for loss bias.
This is one of the attributes, and I think it’s an important one. Humans have a natural loss bias. People attach a mathematical number to it: A loss is six times more painful than a gain.
You see this when someone will invest money, lose 20%, and it will be very painful, but they didn’t invest in something that grew 100%, and that hurts them less than losing 20% of the money—even though not getting the gain, the opportunity cost, is much higher.
There’s a personality trait in people that I call booking the win early. We would be in an architecture meeting, and someone would say, “If we do this, the chip will be twice as fast.” I’d look around the room, and no one seemed that excited about doing it. What’s going on here?
I started to realize that everyone was hearing, “If we do this, the chip will be twice as fast. Let’s put that in the next chip.” I was hearing, “If we don’t do that in this chip, the chip’s going to be half as fast as it could be.”
As soon as I heard that something could be done, I would book it. I would immediately assume that if I didn’t do it, I’d lost this thing. So, as I started to hire, I would look for other people who had the same sort of book-the-win-early attitude. The moment they hear something’s possible, they book it, and they’re like, “I don’t want to lose that thing.”
A lot of the most successful entrepreneurs manufacture their own discontent.
I want to get there in 1 second. Yeah, but I think this hiring for loss bias, and applying it not only to the talent but also to the meetings you’re having in product design and the things that would make your product better, is actually really important.
You mentioned before that you learned from an episode of Founders on Michael Jordan one way to do this, where he would challenge his teammates to bets. Why was that an interesting idea to you?
When I heard your episode on Michael Jordan, I was thinking, “Yeah, he is very intentionally throwing his keys over the fence so he has to go fetch them.”
Michael Jordan was a very aggressive competitor. He would make bets on everything. Could you throw a quarter and hit a target closer, or something? Just weird stuff like that. He would do it nonstop.
Most people are afraid to taunt someone else—a competitor—because if they lose, they’re going to feel really bad. Remember that loss bias is heavy. If I’m like, “We’re going to go play basketball. I’m going to wipe the floor with you,” and then I lose, that’s humiliating.
What I suspect Michael Jordan was doing was very intentionally taunting the other players so that a loss would be humiliating, forcing himself to perform at superhuman levels. He was just doing it over and over again.
Most people are so afraid of putting themselves out there and suffering the negative outcome that they won’t get their hopes up. They’ll actually keep their sights much lower. But entrepreneurs start a company and say, “Of course I’m going to be successful. I’m going to tell everyone, I’m going to go raise money, I’m going to put my reputation on the line, and I’m going to be forced to perform.”
Michael Jordan’s trainer, Tim Grover, wrote the book that I did that episode on. The way he describes this is exactly what you’re saying: “Once you tell somebody how bad you’re going to fuck them up, you have to actually go and do that.”
What Tim Grover realized when he was studying Jordan’s career was that Jordan intentionally heaped more pressure on himself. The more pressure he put on himself, the better he performed and the higher he rose throughout his career.
I think this is also tied to something that you and I have talked about, which you called manufactured discontent.
There’s a book on the counter. We were talking in the kitchen earlier, before we started recording, about David Ogilvy, who’s one of my heroes. He calls this divine discontent. You’ll find that the best entrepreneurs, the best athletes—anybody who reaches the top of their profession—they don’t rest on their laurels, and they don’t sleep on wins.
There’s another book right next to that: The new biography of Steve Jobs just came out. Steve Jobs demonstrated this concept perfectly. He was just like, “Well, you made this great product. Now what?” He believed that if you make something wonderful, the only thing to do is do it again. Don’t think about it; just the next day, go on and make another great product, or keep doing this.
Essentially, they’re telling you the journey is the reward. So, talk about your idea of manufactured discontent.
I was having a conversation with a bunch of entrepreneurs and a bunch of people in other fields. Some of them had made a lot of money, and some of them hadn’t. Everyone in the discussion was incredibly successful.
What we started to realize was that even though some of these entrepreneurs had made hundreds of millions of dollars, they were comparing themselves to others, and they never had to work again in their life. But because they were unhappy with their wealth, they had a reason to continue, start another company, and do more.
Meanwhile, the other folks who were very successful were quite happy with their wealth, but what they were unhappy with was their previous work product—a previous piece of writing or something like that. Because everyone in this room was successful, what we identified was that everyone had something they were discontent about that drove them.
I started looking at my own life, and there were a lot of periods where there was genuine discontent because we hadn’t had product-market fit. But then, once we had product-market fit at Groq, I was unhappy with the scale. Then I was unhappy with other elements, and I just kept finding things to be unhappy with.
Most people can be quite content with the status quo, and they’re not going to keep pushing to innovate. You have to have a personality where you are constantly discontent if you’re going to keep pushing things forward.
What are you discontent about today?
At the moment, I’m discontent with the lack of compute in the world. AI is revolutionary. It’s going to change everything for people. There are pros and cons, but the pros are massive.
There are going to be medical discoveries, and if it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault. Every single person who dies from cancer, every single person who becomes old and infirm and dies—there could come a point, we don’t know, where AI comes up with ways to slow aging.
All of that feels like it’s on my shoulders, and I need to perform. I need to make sure that the world has more compute.
I love that idea of saying, “Every day that we miss out on this mission, there’s a real cost to it.”
Edwin Land, founder of Polaroid and Steve Jobs’s hero—he’s one of my favorite entrepreneurs of all time. I won’t shut up about him. Way before he invented the Polaroid camera, he was actually trying to invent new ways to reduce headlight glare, because in the early days of the automobile, there were so many people dying because of the oncoming headlights of cars.
What he did, very similar to what you did, was have an organizing principle. You said, “We have 25. We need to get to 25 million tokens.”
He would put on the whiteboard, “300 people died today because of this.” If it takes us an extra week, that's 2,100 extra people. I don't know what the number is, but it's something like that.
I do think you're in a perfect position. This show is a love letter to capitalism, and I think we should end on optimism. I think the best entrepreneurs in the world are default optimistic and default aggressive at the same time.
Can you give me an overview of what you actually think is coming with AI and, as a result of AI, some of the most optimistic things that you could say? You see what's going on right now: everybody's upset. It's super unpopular. People want to blow up data centers; they want to attack certain people inventing the technology. I don't think we've done a good enough job of telling a more positive story.
I think that goes back to people perceiving a change. I had a recent post that got a lot of negative feedback, in which I said that there's effectively been code rationing. As a software engineer, the default is that code is expensive to write, so I'm going to be very careful about what code I write. I'm not going to create a feature unless I'm absolutely sure that it's the right feature to create. I'm not going to implement something until I've got it figured out.
What we've seen from agile software development is that when you take the risk and implement something and get feedback, you end up getting better results. But there's still a very natural predilection to say no to things. There's the concept of a no engineer: it's someone's job to say no to things. They're the ones in the meeting who say, “We can't do this. We shouldn't do this.”
What I'm seeing is that code is becoming almost free. Its marginal cost is approaching zero, and it's shifting the way that things are done for professional engineers. You just implement the thing, experience it, and say, “Reimplement it in this different way based on my experience.”
The other shift is accessibility. It's very much like literature and literacy. There was a time when scribes were the only people who could read and write, and they controlled access to the written word. Then we got much simpler reading and writing systems—alphabets as opposed to ideographs and hieroglyphs—and all of a sudden, many people could learn to read. Then we got an education system, and everyone could learn to read and everyone could read and write. All of a sudden, it became about the quality of the written word, not just the written word. But everyone had access.
My EA creates software applications now. When I go on a trip, she creates a little app that I can click through. It tells me what the weather is going to be, updates live, pulls information from sources, and gives me all my phone numbers and all sorts of extra information. That would have been impossible for an individual who didn't know how to write code before.
What I think is going to happen is that a lot of people are going to get access to creating software to solve problems who would never have had the technical capabilities before, but who would have had good taste and known what good is. There's just going to be an enormous number of founders. Unlike in the past, when you didn't have access to the capital or the talent, I think you're going to see individual founders without large teams creating very valuable companies that solve real problems for people.
I love that framing. We'll end on this: one of my favorite quotes of yours was, “Looking forward to a year of massive upleveling for anyone who wants it.”
Anyone who wants to learn can now learn a subject. You just ask questions. The problem with traditional education is that it was force-fed to you. It wasn't interesting. If you're going to learn something, it needs to be interesting.
The ability to ask questions in the moment when you want to learn something is going to fundamentally change education. This goes back to what I said earlier: the AI age is going to be about asking questions.
A lot of people ask me what they're going to do for their kids, and my answer is: stop teaching them to answer questions and start teaching them to ask questions. Curricula should be revamped around a problem that actually matters for the community. Maybe you need to fix the way that permitting is done in the city. Maybe you need a way to improve how you get the word out about some sort of event that's occurring.
Have the students write actual applications that are useful for the community they're in and solve real problems, and then have them ask questions. When you create homework or a test for kids, if they can look up the answer online or ask AI to solve it, you haven't taught them what they need for the next stage. But if you give them a problem where they have to ask the questions and get AI to solve it, then you have.
Thanks for the time, man. Glad you did this.
Thanks.