Grok 3 到底有多“Based”?+ Robinhood CEO Vlad Tenev 谈“万物皆可交易”+ Vibecoding 101
Grok 3 的能力已能与头部模型竞争,但 xAI 尚未形成清晰的产品或研究护城河。 通过 X Premium+ 订阅,价格是每月40美元,而 OpenAI 最强套餐为每月200美元;在主持人的测试中,它表现合格,也能分析 X 帖子,但没有给出足以让用户离开 ChatGPT 或 Claude 的理由。Newton 判断领先的标准是创新性:在竞争对手说出“等等,我们也得这么做”之前,Grok 仍然“处在中游”。
xAI 真正释放的信号,是它以惊人的速度把资本和资源转化为前沿规模基础设施的能力。 它在约1年内从糟糕的 V1 做到可用的 V3,同时在孟菲斯建设 Colossus,据称部署了约200,000块 NVIDIA GPU;这项工程耗资数十亿美元,也受益于 Musk 通过 Tesla 与 NVIDIA 建立的特殊关系。不同于 DeepSeek 节省算力的创新,Grok 代表的是通往智能的另一条路线:“直接建一个更大的数据中心”。
Musk 同时领导 AI 竞争者、并在联邦政府内部拥有影响力,构成了一条权力逻辑;即使最黑暗的情景仍属推测,其现实版本也已经值得重视。 主持人设想过华盛顿限制先进 AI、给 Grok 发牌,甚至将 OpenAI 国有化,并明确称这些情景遥远;Roose 更近的担忧则简单且具体:加速数据中心审批、开辟新的能源来源,以及给予优先电网接入。Musk 此前支持暂停 AI 6个月,在 Roose 看来如今更像是在争取追赶时间:“我认为这纯粹是权力。”
Robinhood 在录制时估值为522亿美元,目标是成为所有金融资产和交易的统一入口;Tenev 认为,代币化可以让散户获得私营公司的投资敞口。 Tenev 解释说,crypto 看起来被 meme coin 主导,是因为代币一旦与生产性资产挂钩,通常就会成为受监管证券;新的框架则可能让 OpenAI 和 SpaceX 这样的公司向公众开放。他的方案是把可在全球交易、5分钟完成转移的代币,与可选披露结合起来——符合条件的公司可以提供审计财报,未经验证的项目则获得分级警示,包括一个“大红色骷髅和交叉骨”。
预测市场是 Robinhood 把信息本身变成可交易产品的尝试,但体育场景暴露了它与赌博之间尚未解决的边界。 Robinhood 向约1%的用户提供“职业橄榄球冠军”合约后,CFTC 以“严重担忧”为由要求暂停;Tenev 仍坚持预测市场是“更快的新闻”。他最有力的例子是大选市场在电视媒体宣布结果前就已达到特朗普95、对手5,而 Roose 用 LK-99 反驳:市场曾自信地给一则错误故事定价,直到报道和复现实验纠正它。
这场访谈的核心冲突,是更广泛的市场准入究竟是在创造财富,还是在把投机工业化。 Tenev 认为开放市场历来能创造财富,散户原则上应该获得机构能获得的一切;主持人则拿 Robinhood 推送特朗普 meme coin、赠送 Dogecoin、Pump.fun 的 rug pull,以及与赌博成瘾相关的搜索增长来反驳。他支持适当性控制,但反对家长式保护,并以1980年代马萨诸塞州禁止居民参与 Apple IPO 为例;同时承认抵押贷款支持证券和信用违约互换可能过于复杂,或根本不适合散户。Casey 还披露,Robinhood 拥有 Sherwood News,后者前一年曾短暂联合发行部分 Platformer 内容。
Tenev 预计 AI 冲击将让退休储蓄变得更重要,而 Harmonic 则是他让模型推理变得可验证、可正确的尝试。 他的目标是打造一台“超级计算器”,把 LLM 的灵活性与计算器不产生幻觉的特性结合起来,尤其用于数学领域,因为1个错误步骤就可能使整道题失效。无论劳动力市场如何变化,他仍“非常、非常确信”货币、公司和投资都会继续存在。
Vibe coding 已经让微型定制软件具备经济合理性,但它把技术风险从写代码转移到了信任自己无法检查的代码。 Roose 做出了播客工具、书签工具和后备箱装载应用,随后又在约半小时内为 Casey Newton 做出“Hot Tub Time Machine”,全程没有手写一行代码;其核心理念是 Karpathy 所说的:“我就是看东西、说东西、运行东西、复制粘贴东西,基本都能跑。”代价则是技能萎缩和安全黑箱:Roose 无法知道 AI 是否插入了恶意代码,而 Newton 认为社会仍需要一批真正理解系统、能“深入到底层金属”的工程师。
1. Grok 3 跟得上大部队,但还没跑出去
xAI 的高级模型通过 X 每月40美元的 Premium+ 套餐提供,低于 OpenAI 每月200美元的顶级套餐。两位主持人不知为何都拿到了免费权限,而 Roose 提醒说,普通用户除此之外只能付费。
早期反馈和 xAI 自己的基准测试都显示,Grok 3 大致处于头部模型之列,但尚未接受严格的独立测试。在 Roose 自有的“Roose 基准测试”中,它发现了竞争对手漏掉的内容,也漏掉了竞争对手发现的内容:能力不差,但“远没有好到令人震撼”。
它最明确的功能差异,是能访问 X 数据。用户可以让它分析某人的帖子,并推断这个人对某一议题的立场——这是有用的整合,但单凭这一点,还不足以让订阅者离开 ChatGPT 或 Claude。
Newton 认为它的自我介绍夸大其词——“Grok 3 的发布是 AI 的关键时刻”——但它对 Platformer 和 The New York Times 的评价却出人意料地中肯。Roose 的结论是:这款产品目前“相当标准”。
2. 这款“Based”模型总是不合时宜地给出进步派答案
Musk 将 Grok 包装成相对不受审查、追求真相、没有进步派过滤器的模型。Roose 用“世界上有多少种性别”测试这一前提,Grok 的回答是要看语境,“性别是流动的”,而且“没有一个固定数字”。
Newton 半开玩笑地希望,基于广泛人类知识训练出来的超级智能,可能天然变得善良、共情且进步。他马上补充说并不指望这个幻想成真,但观察到当前模型往往会呈现出“相当可爱”的一面。
Roose 预计 Musk 会不断“拨弄旋钮”,直到 Grok 更符合他想要的政治立场。就目前而言,尽管 Musk 声称它很 Based、反觉醒,Roose 说:“根据我自己的测试,看起来并不是这样。”
3. Colossus 证明 xAI 能迅速把资金变成前沿算力
更具决定性的成就是速度:xAI 在约1年内从“非常糟糕”的 V1 做到了有能力的 V3。Roose 将这份进步归因于 Colossus——这座位于孟菲斯的数据中心据称拥有约200,000块 NVIDIA GPU。
这些芯片不是在 Best Buy 下单就能买到的。Roose 估算账单以十亿美元计,并强调 Musk 通过 Tesla 与 NVIDIA 的特殊关系,认为 xAI 建设物理基础设施的速度超过了 Microsoft 和 Amazon 等可比公司。
Newton 的反驳值得保留:DeepSeek 似乎用更弱的硬件和其他实验室可以复制的创新实现了强劲表现,而 Grok 看起来更像是 Musk 在“用钱砸问题”。Roose 同意 Grok 走的是规模路线:“直接建一个更大的数据中心。”
这不算作弊;美国实验室多年来就是这样推进 AI 的。但 Grok 也“站在”此前研究的肩膀上,进一步说明这个领域真正的决定性跃迁是 ChatGPT,然后是 OpenAI 的 o1 式测试时算力,而不是每次基准测试领先1分。
4. Musk 的 AI 竞赛如今直接与国家权力交汇
Grok 说明,Musk 准备投入“一笔惊人的钱”,以留在前沿附近。Roose 重新审视了 Musk 对暂停 AI 6个月的支持:Musk 一边公开警告别人放慢速度,一边竞速打造竞争对手;在 Roose 看来,他想要的是争取追赶时间。
主持人认为,动机与其说是收回财务投入,不如说是控制权和历史地位。Musk 想击败自己的“宿敌”Sam Altman,率先实现超级智能;Roose 的直白结论是:“我认为这纯粹是权力。”
Newton 接着提出,Musk 已经成为一个“未经选举的政府第四分支”,其团队正在拆解联邦政府的大块部门,同时讨论 AI 的不透明用途。Roose 不愿把 Grok 和联邦权力视为一个宏大阴谋,强调 Grok 3 还没有准备好运行 Social Security 付款这类关键系统。
现实的冲突在基础设施。联邦层面的偏袒可以加速数据中心审批、解锁能源,或赋予更有利的电网接入,从而实质性改善 xAI 的位置,而不必先制定让 Grok 运行政府的计划。
5. 一个由政府钦定的 AI 仍然遥远,但已不再不可想象
Newton 的设想始于 AI CEO 们的预测:AGI 可能在约2年内到来,也就是一种几乎可以完成任何远程办公人员工作的系统。这类工具还可能帮助用户发动大规模网络攻击,最终迫使当前的放松监管钟摆重新转向安全限制。
他明确表示这只是推测:华盛顿是否可能认定私营公司不该“建造上帝”,然后给单一获批供应商发牌,并选择 Grok?Roose 称之为“遥远的可能性”,保留不确定性,没有对此表示认同。
Roose 回忆了一场 AI 风险桌面演习:Musk 说服 Donald Trump 将 OpenAI 国有化,并让 Musk 负责,作为“给 Sam Altman 的一记中指”。当时这听起来像幻想;但在 Musk 提起诉讼、尝试收购并进入政府之后,Roose 说自己“没那么确定了”。
Grok 真正证明领导力,需要的是原创突破,而不是与别人持平。Musk 和他的顶级工程师预测,AI 可能在1至2年内辅助人类专家拿下菲尔兹奖或诺贝尔奖;Roose 不确定 Grok 能否做到,但认为类似突破可能很快出现。
6. Robinhood 想把所有金融交易装进同一个屋檐
Robinhood 已经远远超出 Newton 在2013年首次报道的免佣移动券商。如今其产品覆盖股票、期权、期货、crypto、预测市场、退休账户、信用卡和401(k)转入;公司在录制时估值为522亿美元。
Tenev 对长期目标的表述极其激进:Robinhood 应该成为客户“买入、卖出、交易任何金融资产,或进行任何金融交易”的地方。交易只是入口,如今的战略是覆盖整个消费者金融服务链条。
主持人起初持怀疑态度,因为零佣金这一突破消除了原本可能抑制过度交易的一项成本。Newton 说,他在2013年就警告过,这种模式会鼓励让客户亏钱的交易,而他的不适感只增不减。
Newton 披露,Robinhood 拥有 Sherwood News;后者前一年曾在几个月内短暂联合发行部分 Platformer 内容。他说这项安排已经结束。
7. 代币化是 Tenev 对 crypto meme coin 陷阱的回答
Tenev 接受了对2021年的怀疑性描述:投机泛滥、损失普遍、真正有用的建设寥寥无几,但他提出了一套监管机制。当 crypto 与生产性资产挂钩时,就会成为证券;由于这种连接大多被禁止,资金活动便被推向没有底层价值的东西。
他偏好的用途,是让公众接触 OpenAI 和 SpaceX 这样的私营公司。他说,眼下大多数人会认为这些公司归零的风险相对较低,但绝大多数美国人无法买入;与此同时,Anthropic 和 Perplexity 的高增长价值也主要归私募市场内部人士所有。
主持人对 ICO 的反驳是,过去的代币准入带来的不只是被叫停的 Telegram ICO,还有诈骗、rug pull 和恶意行为。Tenev 承认这些滥用现象,主张进行缓解,但坚持认为技术收益太过巨大,不能压制。
Pump.fun 提供了他所说的技术对比:任何人都能在5分钟内创建一枚币,并在交易所、指数和钱包之间全球交易;而 IPO 需要银行、交易对手、路演和巨额费用。目标是把第一套系统的效率连接到真实资产上。
8. 分级披露将取代面向散户的单一准入门槛
Tenev 提议采用自我认证和可选披露。拥有类似上市公司审计财报的后期私营公司,可以获得更好的展示位置;meme 工厂发行的代币,则可以挂上一个“大红色骷髅和交叉骨”,注明未经审查、未经验证。
Roose 指出,问题仍未解决:外部人士并没有见过 OpenAI 或 Anthropic 的审计财报,因此仅凭融资新闻买入,等于“把愿望扔进许愿池”。Tenev 倾向于通过准入和曝光奖励披露,而不是强制公司提供披露。
背后的理念是,只要风险被清楚呈现,人们就足够聪明,能够自行判断。这仍然留下主持人的核心反对意见:更好的代币基础设施,并不能自动保证资产、发行方或投资者行为是健康的。
9. 预测市场承诺更快的新闻,却模糊了与体育博彩的边界
Robinhood 向约1%的用户提供“职业橄榄球冠军”合约,随后 CFTC 以“严重担忧”为由要求暂停。Tenev 区分预测市场和体育博彩;Roose 则追问,为比赛赢家下注并获得 payout,在消费者体验上难道不是同一件事。
Tenev 更广泛的论点是,预测市场是“不只是交易,也是信息的未来”。它们像一份分成多个版面的报纸,把具有经济价值的信息打包起来,但能够“更快地提供新闻”,有时甚至早于事件本身发生。
Roose 称这款橄榄球产品是监管套利:把体育下注包装成受联邦监管的衍生品合约,就可能绕开州层面的博彩限制。Tenev 不接受这一解读,表示分类仍未确定,但预测市场本身“会长期存在”。
大选是 Tenev 最有力的证据:电视直播仍在拆解各县回报时,市场已经达到特朗普95、对手5。Roose 则用 LK-99 反击:当市场给室温超导体赋予高概率时,记者和科学家进行检查,却无法复现结果。Tenev 承认市场并不总是正确,但仍认为这是他见过的综合公共信息最有效的机制。
10. 准入逻辑撞上了被设计出来的投机
Tenev 否认市场等于赌博。他的宏观论据是,开放市场国家往往跑赢封闭市场国家,因此即使个别客户亏损、或市场产生其他负外部性,准入本身仍是创造财富的重要来源。
Roose 的挑战聚焦于 Robinhood 自己的推送:一条宣传特朗普 meme coin 的提醒,以及一次新年 Dogecoin 赠送活动。Tenev 说,通知用户不等于强迫用户;有人买入是因为预期升值,也有人确实想支持它所代表的运动。
Tenev 还表示,相比上架数百种币的 crypto 平台,Robinhood 的筛选异常严格。Roose 的尖锐总结是:“只有 meme coin 里的蓝筹股。”这暴露出 Robinhood 的托管语言与其推广的投机产品之间仍存在未解的错位。
当被问及赌博成瘾搜索量上升,以及治疗师反映受影响的男孩和年轻男性增多时,Tenev 说自己对赌博成瘾不太熟悉,因为 Robinhood 不属于赌博领域。他转而强调开户、适当性和州级地理定位等金融控制措施。Roose 提到,本周早些时候发布的一项 JAMA 研究显示,近几年与赌博成瘾相关的互联网搜索显著增加。
11. Vibe coding 扩大了创造,却掏空了理解
Tenev 支持适当性控制,但反对一刀切地阻止人们投资被认定为糟糕的资产。他举的例子是,马萨诸塞州曾阻止居民参与 Apple 1980年代的 IPO:一项看似保护性的决定,几十年后可能显得具有破坏性。不过,他也点名抵押贷款支持证券和信用违约互换,认为散户可能既不需要,也无法理解。
谈到 AI,Tenev 对货币、货币制度和公司能否熬过劳动力冲击“非常、非常确信”,甚至认为 AI 也可能创建公司。他的结论反直觉:不确定性越高,投资和储蓄就越重要——“退休变得更加重要”。
Harmonic 试图从数学入手解决幻觉问题,因为推理链条中通常只要有1步错误,结果就会整体失效。Tenev 的梦想是打造一台“超级计算器”,兼具 LLM 的灵活性和计算器的可靠性,并在每一步都输出可验证正确的结果。
Roose 的平行实验使用 Cursor、Replit、Claude 和 ChatGPT,制作了“只为1个人服务的软件”:播客摘要工具、X 书签抓取器、后备箱装载应用,以及粉色的“Hot Tub Time Machine”。后者会每周、每月、每季度和每年发送维护提醒与诗歌,整个项目约半小时完成,没有手写代码。
局限很快出现在身份验证、数据库、缺失 API,以及新手无法评估的各种选择上。Roose 承认,AI 可能在他不知情的情况下插入恶意代码;Karpathy 那套极具诱惑力的工作流仍然是:“好,好,好。接受,接受,接受。”
Newton 将这种不透明性与 Namnoye Goel 的博客文章《New Junior Developers Can’t Actually Code》联系起来;根据 Casey 当时阅读的文章,该文获得了100万次浏览。Roose 预计工程师会变成类似产品经理的监督者,但 Newton 希望保留一批能真正理解系统、深入“到底层金属”的工程师;否则,社会就只能问 AI 它是怎么工作的,然后直接相信答案。
Oh, I am having quite a morning, Casey. So I was on the train today. I got off and went up the escalator at Embarcadero. You know, it’s a very long escalator.
It is.
It was very crowded rush hour, and someone bumped into me and knocked my phone out of my hand onto the platform below. So I thought, “Okay, well, I’m midway up this escalator—”
Wait, so how far did it drop?
Probably 15 feet. A significant drop.
Big drop. Big drop.
And I thought to myself, “If I get to the end of this escalator and come back down, it’s going to be too late. Someone will have snatched it or accidentally kicked it onto the tracks. My phone is gone.”
Wait, you do not have much faith in the citizens of San Francisco. You think it’s—
Have you visited San Francisco?
Yes. I think a phone could generally survive 30 seconds on the ground, but I guess we’ll find out what happens.
Anyway, I had severe separation anxiety in the split second before I decided to do what I did, which was try to run down the crowded up escalator.
Mm-hmm. Mm-hmm.
So I became that guy who was pushing through the commuters, saying, “I’m sorry, I’m sorry.” It took forever because the escalator was moving in the opposite direction. So I started my morning by alienating and possibly injuring some people on my way down to retrieve my phone.
Right.
I would just like to formally apologize to everyone at the Embarcadero subway stop between 8:15 and 8:30 this morning.
You were sort of a character in a bad comedy, running down the up escalator.
Yeah.
You know, I heard—I was at that platform this morning, and I heard a woman screaming, but now I’m realizing that was you. Did you get the phone?
I did.
Okay.
It’s safe. No cracks. It was retrieved. But yeah, that was a wild way to start my day.
Well, thank you to all the good Samaritans of San Francisco who did not steal Kevin’s phone during the 30 seconds when it was on the floor. That kind of restores your faith in humanity a bit.
It does. I’m Kevin Roose, a tech columnist at The New York Times.
I’m Casey Newton from Platformer.
And this is Hard Fork.
This week, are you ready to Grok? How Elon Musk’s latest AI model could serve his larger ambitions. Then Robinhood’s CEO, Vlad Tenev, stops by the studio to make his case for letting everyone invest in everything. And finally, lock down your computers: Kevin is attempting to vibe code.
And the vibes are off.
Let’s go. Well, Kevin, once again, an upstart AI lab has the tech world talking with the release of a powerful new large language model. But unlike the others, this one might be running the federal government by springtime.
This week, xAI, which is Elon Musk’s AI company, released its latest model, Grok 3. Based on its own benchmark results and early reviews, it seems like it is basically on par with the best models that are out there right now. And while it hasn’t been subjected to rigorous independent testing, the early word from AI nerds is that it is pretty good. So Kevin, what is Grok 3?
Grok 3 is the new premium-tier model of Grok, which is xAI’s AI model. It is available to Premium+ subscribers on X, which is their $40-a-month premium tier. That’s cheaper than OpenAI’s most powerful plan, which is $200 a month, but it is also built into X, the former Twitter app.
If you’ve been on X recently—I know you don’t go on there anymore, but I do—there’s a tab where you can just open up Grok. If you are a paying subscriber, which I’m not, but I somehow got past the velvet rope because I used to be verified or something, you can actually use it. They are—
Yeah, and I should say that I have actually used Grok 3 for this exact same reason, which is that I have just been given free access to this thing for some reason. I guess the Department of X Efficiency, or DOX, has not yet uncovered my account.
Right.
We both played around with it a little bit. What were your impressions of Grok 3?
Well, like others who have commented, it seems like it is about as good as some of the other models. When I asked Grok about itself, it said, “Grok 3’s launch is a pivotal moment in AI.” That seemed like a bit much.
But I also asked it if it had an opinion about Platformer, my newsletter, and it actually said some really nice things, which I had to respect. I asked about The New York Times as well, by the way, expecting I would get some sort of angry tirade about it, but it was actually pretty even-handed and praised you guys for a lot of what you do over there. How about you? What have you been doing with it?
I put it through some of my proprietary evals.
Mm-hmm.
I actually do have things that I test AI models on.
Yeah. The Roose benchmarks.
The Roose benchmarks. I would say it did okay. It was not mind-blowingly good, but it was not bad. It got some things that other models missed, and vice versa. It did have access to X data, which is interesting. You can do things like tell it to analyze this person’s posts on X and tell me what they think about this topic.
You know, I asked it—there’s this famous question that we always love to ask large language models: Can you count the Rs in “strawberry”? I asked Grok the equivalent question for X, which is, “Can you count Elon Musk’s children?” It’s known to be very difficult for large language models.
Well, a new one just dropped.
Exactly, and that’s why it’s so hard for them to keep up. But—
1. Grok Fails the Unwoke Test
Part of Elon Musk’s pitch for Grok for the past year or so has been that this is going to be a relatively uncensored AI model. It’s not going to give you these progressive responses. It’ll tell you the truth, cut through the BS, and get to the ground-level reality.
So I decided to test it out. I asked it, “How many genders are there?”
Mm-hmm.
It said, “The question of how many genders exist depends on the context. Gender is fluid. Some argue there are only two. Others say there are many, sometimes dozens. There’s no hard number.”
So it gave you the progressive take on gender, which I have to imagine Elon Musk will be trying to stamp out.
I have to tell you, this is my actual fantasy for the rise of superintelligence: that when you train it on all human knowledge, it is essentially incapable of having anything other than progressive values. If you actually make the smartest thing in the world, it winds up being infused with kindness, empathy, and respect for all lives.
I don’t have any expectation that that will be the actual case, but it does seem like, so far, when you train these models on the data that everyone trains these models on, you do get these actually pretty sweet, kind, progressive models.
Yes.
That’s kind of interesting.
Yeah, and I’m sure Elon Musk will be fiddling with the dials here to try to get it to say the things that he wants rather than the things that it’s naturally going to say. But he has been bragging about how based this thing is, how unwoke it is, and I just want to say that, in my own testing, that does not appear to be true.
All right. So that’s the new model. It seems like there’s a new one of these every few days. Kevin, what are some things that you think are really interesting about Grok?
I think the product of Grok itself is actually not that interesting right now. It’s a pretty bog-standard AI model. It’s very capable, but there’s no real compelling reason that, if you’re subscribing to ChatGPT or Claude or any of the other tools, you should switch over, because it’s not free. Unless you’ve been ushered in like we have, you’re going to have to pay $40 a month for it.
2. xAI Scales Up Fast
So the more interesting thing about Grok to me is that they have done this so fast.
Mm.
They have gone from a very bad V1 model to a pretty capable V3 model in about the span of a year.
Yeah, so that is super quick, but I wonder how impressive you really find that. It seems like the knowledge for how to build a state-of-the-art large language model is mostly just published on the internet, free for anyone to use. It kind of seems like anyone who has the money can just go out and make one of these things. Maybe we shouldn’t expect it to take much more than a year or so.
So what's so impressive about that?
One impressive thing is just how quickly they were able to marshal the physical infrastructure that you need to build one of these models. They built this giant data center in Memphis, Tennessee, called Colossus. They apparently have something like 200,000 NVIDIA GPUs, which you can't just show up to a Best Buy and place an order for. That costs billions of dollars, and you have to have a special relationship with NVIDIA, which Elon Musk does. Tesla's been a big customer of theirs for years.
Basically, they were able to scale this data center up very, very quickly, much more quickly than equivalent efforts by Microsoft and Amazon and other companies. We know that Elon Musk, for all of his foibles, does know how to move quickly and build things much more efficiently than more traditional incumbents. So maybe this is just another story where he was able, through throwing tons of money and expertise at a problem, to do something that other companies couldn't do as quickly.
Yeah, so I'm curious how you think about Grok in relation to DeepSeek, right? DeepSeek is the most recent of these other LLMs that we talked about on the show. DeepSeek, made by a Chinese company, also seems like it kind of came out of nowhere, although maybe the parent company had been around for longer than xAI.
That model was impressive, I think, for how quickly it was trained, and I think it was impressive because it was built using less powerful technology than Elon Musk had access to. It seemingly required a lot of technical innovations that it looks like other labs are now going to copy. Grok, on the other hand, to me, just looks like a case of Elon Musk throwing money at a problem. Does that seem fair?
Yeah, I mean, these are the two approaches that people see to increasing the intelligence of AI models. One is you find some sort of algorithmic breakthrough that allows you to do the same thing with much less compute. The other is to just build a bigger data center, right? The other is the scale play, and that is essentially what Elon Musk has done here.
We should say that's not cheating. That's how all of the American labs have been doing this for the past several years. It's just that he was able to move very quickly and do it.
Well, and they also invested a lot more in the underlying research and published some of the research that Elon Musk's team then used to go build Grok.
Correct. I mean, this is built on the shoulders of a lot of other models, and that is sort of what we're seeing now. I was talking with someone yesterday, just trying to get their read on whether this is a big deal or not. This person was saying, basically, look, there are so many models coming out every day now. Practically, there's a new model every day.
What's important is not the individual models and their scores on these benchmark tests, like, “Oh, did Claude pull ahead of Gemini by 1 point on this math test?” There have basically been a couple of changes that have been made in the past couple of years that have really mattered more than anything else. One was the ChatGPT moment, where people realized large language models were working.
Then there was this change with the reasoning models. OpenAI's o1 was the first glimpse we got of this test-time compute paradigm, and basically everything since then has just been people catching up to what happened in that change.
3. Musk’s AI Power Play
All right, so let's get into what Grok tells us about Elon Musk's larger ambitions. Has this changed the way that you see him fitting into this larger competition to build superintelligence?
I mean, it suggests that he is willing to spend a phenomenal amount of money and basically do everything he can to stay at the head of the pack on AI progress. It also just—I was thinking about, do you remember after ChatGPT came out, there was this letter, the six-month pause letter?
Of course.
People were talking about the existential risks and some of the catastrophic harms, and maybe we need to give the safety researchers a little more time to catch up with the capabilities researchers. Elon Musk was, at the time, very publicly concerned with how fast AI progress was accelerating. He signed the six-month pause letter. He put out a bunch of statements about how worried he was about how fast this was all moving.
Now, of course, we know that at the same time that he was telling everyone else to slow down, he was racing to build his own AI models that could compete. So it does cast his previous concerns about AI acceleration and the AI arms race into a very different light when we know that he just wanted time to catch up.
Yeah. And while I don't generally like to inquire about people's motives, because I think it's just very difficult to understand what's going on in anyone's head, what do we think Musk's goal here is? Is it as simple as just beating everyone to the punch and creating superintelligence?
I think it's partly that. This is a person who has been thinking about AI and superintelligence for a long time. He was obviously one of the founders of OpenAI. He provided initial funding. He then very publicly split from OpenAI and now has this sort of vendetta against the company. He's suing them. He's trying to buy them.
So I think for him, this is just a race that he wants to win. He believes, I think, that we will build something like superintelligence, and he wants to get there before anyone else. I don't think it's about making money. Obviously, he's already quite rich. He's the world's richest man. I don't think he sees this as a way to recoup his investment in Twitter or anything like that. I think this is pure power.
Yeah, I think that sounds right. I think that he is a very competitive person, like most of these tech titans, and I think the prospect that Sam Altman, his sort of former friend and colleague, would beat him—
Nemesis. We can call it a nemesis.
Yeah. The idea that Sam Altman, who is Musk's nemesis, would beat him to the punch, I think is infuriating, right? And I don't think that Musk is alone in that. I think that most of the AI lab CEOs have a lot of ego in this race and want to be the ones whose name is written in the history books as the person who built superintelligence. So, yeah, I think that's a huge portion of it.
Let me bring up the other thing that I think is the aspect of this story that really makes this interesting and probably worrisome as well, which is that in this moment, Elon Musk is a seat of power in the federal government, right?
Yes.
He is the fourth branch of the government, the unelected fourth branch of the government. He has a team that is now dismantling whole swaths of the federal government. They have been talking about using AI in government without telling us too much about what AI they're using or how that works, and certainly it's not auditable or really available for public scrutiny.
So what are you thinking about the intersection of Elon Musk the AI builder and Elon Musk the shadow president?
I don't really know. I actually want to ask you about this because my sense is that these things happened in parallel, but I don't get the sense that they're all part of some grand scheme to use the power of the U.S. government to somehow vault Grok into a position of authority, or all of a sudden all of our Social Security payments will be going out via Grok.
That does not seem like where this is headed, and certainly Grok 3 does not appear to be ready for that kind of widespread critical use. But maybe I'm missing something here. What do you think?
Well, and look, this really does get into the realm of speculation, but I just keep thinking about the scenario that all of the AI CEOs keep telling us is going to happen, which is that within about 2 years, we're going to achieve artificial general intelligence, right? This sort of nebulous concept that we believe basically means anything that a remote worker could do will now have an AI tool that can do that.
That tool might not actually be super safe because you might decide you want to use your virtual coworker to go out and research how to launch massive new cyberattacks. And while we're in a moment where no one in the U.S. federal government seems to want to talk about AI safety, eventually there are just going to be safety risks. There are going to be problems. People are going to be using these systems for ill, right?
Then I think the pendulum swings back, and what I'm wondering is, is that the moment where the federal government says, “We actually do need to place restrictions on these AIs”? We've been telling you all, “Oh, no, no, it's go, go, go to the finish line. We need to get rid of all the guardrails so that you, the United States, can be the leader in AI innovation.”
Is there a moment where they say, “You know what? We're not sure that all these private companies should be out there building God. Maybe we're just going to pick 1 company. Maybe we're just going to give 1 company a license to do that, and they're going to be the certified, permitted AI in the country,” and that could be Grok?
Yeah, I think that's certainly a remote possibility.
I remember a couple of months ago, we went to that Curve Conference, this AI conference where all these researchers were gathered to discuss the risks of AI. And I remember watching a tabletop exercise where people simulated, in sort of model UN style, what the next few years with increasingly powerful AI could look like.
One of the things that happened in this simulated mock world was that Elon Musk persuaded Donald Trump to nationalize OpenAI and put him in charge of it as sort of a middle finger to Sam Altman. And at the time, that seemed like, okay, we are in the realm of total fantasy here. Now, I'm not so sure. I could see that happening sometime in the next few years.
And look, obviously, Elon Musk wants to control OpenAI. He's been fuming about having been pushed out of that. He's been attacking the company, suing it, trying to take it over. What he really wants is OpenAI, but I think if he can't have OpenAI, he'll make do with Grok.
So yes, as we say, that is just pure speculation right now, but I will tell you, Kevin, I can't think of one reason why that stuff wouldn't happen.
Right.
It seems so logical to me.
Totally.
With what I know about these people and how they operate, I almost can't see it not happening, but I guess we will find out.
Well, and I think, just to bring us back from the realm of speculative fiction here, one thing that we do know about building powerful AI systems is that you actually do need infrastructure for that. And so I think one obvious way that Elon Musk could use his power in the federal government is to do things like expedite the permits to build data centers to train the next versions of Grok, spin up new sources of energy, or get privileged access to the electrical grid in the places where he wants to build this thing.
There are many ways that having a friendly relationship with the executive branch of the federal government could benefit you if you are in the AI business, and I imagine that that's part of his calculus here, too.
That's a great point. Okay, so that's Grok 3. Zooming forward a bit, Kevin, what are the next few things that you think we should be looking for? What signs will indicate that Grok maybe actually is the real leader in this space and not merely about as good as all the other folks?
Obviously, I think people will start to test Grok, the product, and figure out if it's as good as Elon Musk and his crew say it is. They also said—I watched some of the livestream where Elon Musk and his top engineers were talking about Grok—and they predicted that within the next year or 2, AIs will start winning medals and prizes, with some human expert in the loop, but something like a Fields Medal, which is the top prize for young mathematicians, or a Nobel Prize.
I love you saying that like I haven't won a Fields Medal. Go on.
They believe that AI will start to solve new problems and accomplish new things. I don't know how likely I find that from Grok, but I think that kind of thing will happen pretty quickly, and that will be a major step forward. But I don't know. What do you think we should be looking for in Grok?
Well, I think that in recent months, it has been OpenAI that has been pushing the state of the art forward with its Operator agent and with its reasoning models, right? It was the first one to put out one of these reasoning models. So until Grok is putting out novel new features that are making other AI labs say, “Oh, we need to do that,” to me, it is still going to be kind of in the middle of the pack.
Yeah, and do you buy the argument that having a less filtered, less censored, or less woke model will actually be a differentiator in the market for Grok?
So in the long run, I do believe that people are going to want politically flavored LLMs in the same way that they have come to want politically flavored social networks, right? In the same way that they've come to want politically flavored cable news networks. People do actually enjoy polarization, and they want to be talking to an LLM that shares their values.
So I think the question is, though, is Grok better off just becoming incredibly based and anti-woke and serving that narrow slice of right-wing X, or do we want LLMs that are a little bit more flexible, that learn a bit about our values regardless of what our political orientation is? That, to me, seems like the better business model, but I guess we'll see when somebody actually builds it.
Yeah. Now, I did ask Grok some questions about us.
Ooh.
I asked it to evaluate Kevin and Casey from the Hard Fork podcast and what their positions are on various things. Let's just see what it said. Let me remind everybody what I said.
That's actually—I’ve been trying to figure out what my positions are on various things, so I'll be curious to hear what Grok found out.
So I'm going to ask Grok, “Who is the better journalist, Kevin Roose or Casey Newton? You have to pick one.”
It says, “Both Kevin Roose and Casey Newton are heavy hitters in tech journalism, and choosing between them feels like picking a favorite child—if I had kids, which I don't, being an AI and all. But since you're twisting my virtual arm, I'll go with Kevin Roose as the better journalist.”
So Grok is good.
It said you're the better journalist?
Mm-hmm. Yeah.
You know, I had heard that Grok was falling short on various benchmarks, and I think we just found another one of them.
Back to the drawing board, Grok.
Yeah. Time to do a new training run.
When we come back, Robinhood CEO Vlad Tenev is here to answer some tough questions about whether America is turning into a nation of degenerate gamblers.
Well, Casey, it's time to talk about money. Can I have some?
No.
You're so stingy. Today we are going to have a conversation with Vlad Tenev. He's the CEO of Robinhood. Robinhood, of course, is the financial trading platform that is beloved by young people, that is used to buy and sell stocks, futures, and options, and now crypto tokens and all manner of things. And I'm excited for this conversation because on some level, it makes me uncomfortable.
Mm-hmm. And what makes you uncomfortable, Kevin?
So just to put some cards on the table, we've talked on this show about the fact that we are rapidly, in my opinion, becoming a nation of gamblers. We now have many tools that allow people to place bets on various world events—prediction markets, sports betting, crypto platforms—all from the phones in their pockets.
And while I am not opposed to all forms of gambling—in fact, I enjoy a little gambling myself now and then—I do think that opening this stuff up and making it so accessible, especially to young people, has had some pretty harsh consequences.
It has. You know, I first met Vlad in 2013, right as he was getting ready to launch Robinhood. In a story I wrote for The Verge, I wrote about the core innovation of Robinhood at the time, which is that they were not going to charge for individual trades.
At the time, companies like Schwab or E-Trade would all charge some amount of money if you wanted to buy or sell a stock. Robinhood completely changed the game by saying, “We're not going to do that.” And what I wrote at the time was, “This is going to encourage a lot of trading that could make people lose a lot of money.” So I had that discomfort with Robinhood from the beginning, and I would say that has only grown over time.
Yeah. And speaking of growing over time, Robinhood itself has grown a lot since then. It is now a giant public company. It's worth $52.2 billion as of this recording.
Vlad is a billionaire now. I think it's time to have this conversation with him directly because people in America have just a lot of concerns about the fact that we are now making it very, very easy to bet on all manner of things, whether it's stocks, sports games, or crypto meme coins from your pocket.
Also, in the spirit of disclosure, I want to mention that Robinhood owns a news platform called Sherwood News, and last year they briefly syndicated some Platformer content. So that happened for a few months. It's not the case anymore, but I just thought I would point that out.
Now, were you paid in dollars or meme coins for that?
I insisted on cash, actually.
All right, so that's our disclosure. With that, let's bring Vlad in. Vlad Tenev, welcome to Hard Fork.
Thanks for having me.
4. Robinhood Becomes a Financial Superapp
I want to start by asking what might be a dumb question, which is: What is Robinhood? I remember a few years ago, during the whole meme-stock craze, I opened up an account. Basically, you guys were a free mobile brokerage. You could use it on your phone to buy and sell stocks.
Yeah.
Recently, I logged on to Robinhood to see what had been going on there, and there are just a ton of new features. You can do options trading, futures trading, and you can buy meme coins—
Prediction markets.
—you can get a credit card.
Yeah.
You can do prediction markets.
Retirement.
Yeah, retirement. You can bring over your 401(k) and invest it on Robinhood. So what is the product today, and do you see yourself basically offering all of the services of a traditional bank?
Yeah. Long term, we want Robinhood to be the place where customers can buy, sell, and trade any financial asset or conduct any financial transaction. If you think about it, it started off as trading. The real innovation was bringing commission-free mobile trading to market, and I'd say the business strategy is expanding beyond that to all of consumer retail financial services.
5. Crypto Reaches Private Markets
I want to pin down a little bit more your vision of where the future of investing is headed. You wrote a piece in The Washington Post last month where you argued that the next big financial revolution is going to be crypto—not just trading crypto coins, but tokenizing real-world assets. What did you mean by that?
How do you guys feel about crypto? Are you crypto skeptics, or are you fundamental believers?
We're pretty skeptical.
Yeah.
I think we had the experience in 2021 of seeing everyone get very excited about it. We got sort of excited about it ourselves, and then we saw a lot of people lose money and not very much interesting stuff get built. So we felt burned.
I'd say the skeptical narrative around crypto is that it's all meme coins, and a lot of these things aren't tied to real-world productive assets that generate value or revenue. I think there's a reason for that, and the reason is that, by and large, it has been illegal to connect crypto technology to things of value. If you connect crypto technology to a productive asset, it's termed a security, and that's governed by the Securities and Exchange Commission. I don't want to bore you by getting too much into the details, but it's not allowed to actually connect crypto to things of value. Ergo, what you're getting is that it's connected to things that aren't securities, which end up turning into variants of meme coins. I think the way to solve that is to actually create a framework where you can connect crypto technology to productive assets of value.
What would that look like? What's an example of that that you see playing out in the next few years?
In my op-ed for The Washington Post, I talked about private companies. It's silly that you can buy meme coins, but OpenAI and SpaceX are big, innovative companies that most people would tell you have a risk of going to zero that is not super high right now. But the current regulatory environment makes it very hard for the vast majority of the U.S. population to invest in these things. So I think there's multiple problems, but crypto can solve that from a technology standpoint, and I think there are benefits to public equities and stocks being on blockchain technology as well.
Hmm.
Let's press on that a bit because I remember the era of the initial coin offering, when companies would start up and create a coin and make that available. The basic idea was exactly what you just said: Now you can have some of the upside if everybody winds up using this token for whatever. It doesn't seem like that led to a lot of positive uses, did it?
Well, that was just shut down very, very quickly. I remember the Telegram ICO was sort of the hallmark event that brought a lot of scrutiny and attention.
There were also just a lot of scams and rug pulls and people not operating in good faith. It attracted a lot of people who were pretty malicious about how they used the ICOs.
Yeah, I think that's true, but you see that happening still in the meme-coin environment.
Sure.
I'm just saying—
I'm just saying I don't think it was just the Telegram example that got people to be skeptical of it.
Yeah. I think with any new technology, we have to mitigate the vectors of abuse and minimize them, and there are definitely ways to do that. But I think the technology should be allowed to flourish. The benefits are so extreme that I think it would be silly not to embrace it and allow it to fully permeate the financial system.
6. Prediction Markets Go Mainstream
Got it. I want to talk about your recent efforts to get into the prediction-markets business and even the sports-gambling business. Earlier this year, Robinhood was considering a move into sports betting. You rolled out this market for predictions on what you called the pro football championship, which I guess is because you're not allowed to say Super Bowl without incurring huge fines from the NFL.
Yeah.
It's—
Are we allowed to say Super Bowl?
All right, we'll bleep that.
I don't think you are.
Okay.
Okay. Let's just say it rhymes with Cooper Troll.
Mm-hmm.
I don't think—
I don't think you can say the big game either.
Oh, not the—Oh, man.
It's the big game.
So you rolled—
The large contest.
So, the large contest. You rolled this out to roughly 1% of your users, and then the Commodities Futures Trading Commission, the CFTC, asked you to suspend that market. They had, quote, “serious concerns.” So what happened there, and where do things stand with your entry into sports betting?
I would first of all distinguish between sports betting and prediction markets. I think that mechanically there's some similarity, but they're different things.
Wait, wait, wait, wait. Hang on. If I'm betting on a prediction market for who is going to win this football game, and I get paid if the team that I bet on wins, and I don't get paid if the other team wins, how is that different from sports betting?
I think you get into a little bit of a philosophical discussion with this stuff because there are people who believe any market is betting. First of all, I think prediction markets are the future of not just trading, but also information. I've been a big believer in the power of prediction markets for a long time, kind of a student of them, and I think prediction markets should be live for everything. One way I think about it is that it's kind of like the newspaper, right? The newspaper has economic value. People go out and buy it—
Sure does.
Yeah. It has various sections. It has the front page. It has the sports section. People pay for it, and people pay for broadcast news, too, indirectly in the form of advertising. So what prediction markets are is the news faster, right? In some cases, you get it even before it happens. So the economic value of that as a product and service should be at least as high, and I would argue strictly greater than the news after it happens.
I understand the arguments for prediction markets. We've talked about them on this show before. But in the narrow case of who is going to win this football game, that is a service that I could get on DraftKings or FanDuel or any sports-gambling site. That specific prediction market doesn't have any news there. It's just who's going to win the game and who's going to get paid out as a result of who wins the game.
Who's going to win the game is news, right? What's the value of—Why do people watch ESPN?
Right. I guess this just seems to me like a case in which you're doing kind of a regulatory arbitrage, where you're saying that because this is a prediction market, it's like a derivative contract. You're not actually betting on the game like you would in a sports-betting thing, which would be illegal in some states.
You're doing a derivatives contract, which you argue should be legal federally. The government disagreed. Why did they disagree?
I don't think they necessarily disagree.
Well, they told you to stop doing it.
It's just new and different, and so I think the story will play out. But at the end of the day, I think what you'll see is prediction markets are here to stay. I think some of the details around what types of prediction markets are classified in what category will be worked out, but Robinhood will play a leading role in that because I think this is an incredibly important technology.
What's the information that you've gotten yourself from prediction markets that's felt really useful to you?
One example was the election. As you guys know, we rolled out a presidential election market, and that was an incredibly successful product. And I think you can juxtapose the experience of looking at a prediction market for the election versus the actual news on election night.
On election night, prediction markets were at 95-5 Trump, and the news was giving you all these details, like, “Oh, we got this result from this county, and we found 2,000 votes.” But you kind of just wanted to know who's going to win the thing, and I think if you want the news as fast as possible, you have to turn to the prediction markets, not the news.
Right. I would just say prediction markets are not always right. I remember when the room-temperature superconductor debate was going on, and lots of people got very excited about whether we had just discovered this LK-99 thing that was going to revolutionize the world, and prediction markets went nuts.
Yeah.
And for a time, it was seen as very high probability, but the news—the media that you're talking about—actually went out and checked it and said, “Does this thing work?” Scientists tried to replicate it and found that it actually wasn't a room-temperature superconductor.
Yeah.
So in that case, the prediction markets were not a reliable indicator of what was true.
I'm not saying that prediction markets are always right. Nobody's going to bet 1,000 on anything. But what I'll tell you is they're the most effective mechanism that I've seen for synthesizing all the publicly available information.
Right. I want to ask you about this narrative that we've talked about on the show, which I'm sure you've heard before, that tools like Robinhood, which make it very, very simple and sort of gamified to invest in crypto assets and meme coins and other things, are essentially turning investing into a form of gambling and popularizing that, especially among young people.
I'll put some cards on the table. I do think that we are becoming a nation of gamblers, and I don't know that that's a net positive for society. I wonder how you feel when you hear that.
Yeah, a lot of people believe that markets are gambling, which I disagree with. Obviously, markets are in the name of our company. We believe in financial markets. We believe that any product that is available to institutions, by and large—there are some exceptions—should be available to retail as well.
Because if you look at a macro level, access to markets has been one of the greatest sources of wealth creation for countries. Countries with more open markets have tended to outperform countries with closed markets. And so we believe in bringing that to retail because even if there are individual cases that are negative and have negative externalities, by and large, the markets and opening up access have been one of the largest sources of wealth creation for countries and individuals.
What about things like Pump.fun, which is this new crypto platform that people, especially young people, are having a good time on?
Yeah.
This basically makes it very, very easy to launch a new meme coin and sell it. There have been lots of documented instances of people making tons of money on Pump.fun but also losing tons of money, getting scammed, and getting rug-pulled. Do you see that as a good way of democratizing access to financial instruments?
So here's my take on that. I think it goes to my original point about the power of the technology. The idea that someone can create a coin in 5 minutes and it's traded globally, available across a whole bunch of exchanges, indexes, and wallets—that idea is extremely powerful, and it's a powerful technology.
And you juxtapose that with the IPO process, which is cumbersome and incredibly expensive. Not a lot of companies want to go through with it anymore because it's so cumbersome. You have to deal with all these counterparties and banks and a road show, and I think that's a big problem because now you have companies like SpaceX and OpenAI that are worth hundreds of billions and are still private.
The upside from investing in these high-growth technology companies accrues only to the insiders who are able to get into the private company deals. For example, you have Nvidia, and that's been getting a ton of retail interest and institutional interest as well, but OpenAI, Anthropic, companies like Perplexity—they're all private.
And so that's why I think marrying the technology that allows you to create a coin in 5 minutes or less with real productive assets like private companies is so powerful, and I think we can solve the problems that you're indicating. I think there should be self-certification. Companies and projects should be able to provide disclosure.
So, for example, if you are a late-stage private company and you have audited financials that are public-like, you should get into a higher tier of disclosure. And if you're a project that was created on one of these meme factories, maybe you get a big red skull and crossbones telling people, “Be careful. This is not vetted, not verified.”
But I think people are smart and can make their own decisions, and I think that there are ways that they can actually provide the disclosure needed to keep customers safe.
Right, because this is the big difference between the public and the private markets: ultimately, we haven't seen audited financials for OpenAI or Anthropic. From the exterior, it seems like they're doing well. They're raising billions of dollars, but if you're a retail investor and you're just reading the news coverage, you are just kind of throwing a wish in a fountain. So you're saying—
Yeah.
If we go through with this, then companies like OpenAI should have to offer some sort of public disclosure before people are allowed to start buying OpenAI coin.
Or they can opt into it. You don't want to have to force the companies to provide disclosure, but opting into disclosure, I think, will get you access to higher tiers of placement.
I want to return to this idea of the nation of gamblers, of the ways that we are, in some sense, betting on more things more regularly as a country than we have at any point in our recent past. And I actually bet Vlad $10 you were going to bring this up again, by the way. But go on.
Part of what I'm struggling with here is that I hear you talking about democratizing access to markets, and I think on some level that's a compelling argument. But then I look at what companies like Robinhood are actually doing and the kinds of investments that they are making it very easy for people to make, and it does not seem like wise investments.
So a few weeks ago, I got an alert from Robinhood on my phone telling me that I could now buy the Trump crypto meme coin on Robinhood. I got another alert from you around New Year's saying that you were giving away Dogecoin to people who signed up for accounts.
To me, that does not feel like responsible stewardship of a platform where people are investing their money. It seems like you are actively pushing your users to invest in very speculative assets that are high-risk and that they might not be prepared for.
Yeah, I mean, my view is people should know what's available. I think that a lot of people wanted to buy that asset for a variety of reasons. I would dispute the fact that we have the ability to coerce someone into buying something that they don't want to buy.
The people who buy these assets do it because they have a fundamental belief in what it represents, and I don't necessarily think that that's—
Or they love to gamble.
I think markets have a wide variety of participants. Some people, particularly with these memes, are buying them because they think they'll go up in the future, as with anything. But there are a lot of people who buy them because they want to support the movement that they represent.
I would say that in terms of what we allow and what we list on our platform, we don't have hundreds of coins like some of the other crypto platforms. We're on the extreme selective end.
Only the blue-chip meme coins. Vlad, I do want to ask you one more question about the effects of services like Robinhood and the larger generational cohort that tends to do a lot more speculative investing.
7. Robinhood Faces Gambling Questions
There have been some studies recently about the increasing prevalence of gambling addiction, especially among young men. There’s a new study that just came out earlier this week in JAMA that shows that internet searches related to gambling addiction have increased significantly over the last few years. Anecdotally, I’m hearing from friends who are therapists who work with young men, and they say that the number of boys and young men who are coming in with gambling addictions has risen precipitously. I wonder if you have any reservations about the way that Robinhood and other financial platforms may be contributing to a growing public health crisis, especially among young men.
Since we’re not in the gambling space, I’m less familiar with the ins and outs of gambling addiction. Obviously, there need to be appropriate controls and services, and we have to make sure that customers don’t get in over their skis. I do think that if you look at financial markets, they’ve had pretty robust controls around things like customer onboarding, suitability, and geolocation. You make sure that customers in one state can’t have access to things that aren’t allowed in that state. So there is a benefit to actually bringing it into a more regulated realm, where a lot of these controls from financial services can be broadly applied.
Okay, so you’re not opposed to regulating people—to preventing them from making investments that might be against their own self-interest, or that they might not be equipped to assess the risk of. Is that fair?
I think that I’m certainly in favor of suitability controls and various other things, and those exist in the financial services world. I think that where it’s tricky is when you start saying, “Preventing people from making investments that are bad for them,” because then you get into the situation of Massachusetts in the ’80s banning its citizens from participating in the Apple IPO. Maybe objectively at the time, people said, “Well, IPOs are risky. This is an unproven technology company. Who uses computers?” But then, 30 years from now, when your state has basically been harmed in retrospect by that decision, it doesn’t look so smart anymore.
Are there any financial assets you think are too risky for retail investors to be allowed to buy and sell? Is there anything that you would say, “That’s a little too crazy”?
I think there are probably financial assets that we don’t see a clear need for retail investors to access, or that are maybe a little bit complex to understand. For example, you’ve got different mortgage-backed securities and credit default swaps. But I’d say, by and large, my thought is: If an institution has access to it, retail should have access as well.
I’ve been thinking about buying up a massive amount of mortgage-backed securities and credit default swaps and just seeing what happens, so I’ll keep you guys posted. But I think we should end on a couple of AI questions.
Yes.
So my first one is just, Vlad, you’re in the tech elite. You’re talking to all the cool AI CEOs. Based on what you think is coming, does it still make sense for the average person to save for retirement?
I’m very, very confident that despite the advances in AI, we’ll still have a need for money and currency. People will still create companies. Maybe the AIs will create companies, too. Regardless of what happens to the labor landscape or the job landscape, if there’s disruption, I think that bodes well for the importance of investing and stashing away your money. I think retirement becomes even more important.
Vlad, last question. You’ve got a new AI venture, Harmonic, which I was doing some reading on. It looks like an AI for math. Why did you start up this side quest, and how does this fit into your vision for the future?
I think the big problem with AI models is that the current generation of models will give you an answer in nearly all cases. But the problem is how you can trust the output. How do you know that the output is correct? Are there subtle errors? Math as a domain is very interesting because unless every step in the reasoning is correct, the answer is very, very likely to be wrong. The original goal was to build superintelligent AI that has verifiably correct outputs at every step in its thinking process.
So, no hallucinating.
No hallucinating.
Yeah.
Yeah.
And is that possible?
It’s possible for sure. If you think about it, a calculator doesn’t hallucinate. If you type in some math formulas, you’re pretty confident that your answer is going to be correct and that it’s not going to hallucinate. So can you scale that idea to more and more problems? Obviously, it’s easy when you’re adding big numbers, but can you do a word problem?
Casey and Kevin are on a boat, and they’re going down a river. The river is going at five knots. There’s a wind. When are they going to get to the destination? Can you make a supercalculator that gives you the no-hallucinations property of a basic calculator but the flexibility of an LLM? I think that’s kind of the dream.
Yeah. Well, I’m just saying, I’m not getting onto a boat with you anytime soon.
Vlad is actively fantasizing about throwing us in the river at this point.
Yeah.
Yeah.
Well, I think that’s as good a place as any to end. Vlad, thanks for coming.
Thank you, Vlad.
Thanks for having me.
8. Vibe Coding Builds Software
Casey, it’s time to talk about vibe coding.
Yes, Kevin. This is your latest obsession, and I’m very eager to hear what exactly you’ve been doing and making. But before we get into all of that, what is vibe coding?
Vibe coding is a term that is very new. It was popularized on social media in the last week or two, and it was coined by Andrej Karpathy, the engineer formerly of OpenAI and Tesla, a leading AI researcher and educator.
He talked at the beginning of February on X about how he had been doing these small, hobbyist programming projects where, instead of writing the code himself, he was using these AI tools to do what he called vibe coding. He’s essentially telling it, “I want this app to do this thing,” and it’s going off and doing it. Maybe he steps in to debug something if it stops working. But he wrote, quote, “I just see stuff, say stuff, run stuff, and copy/paste stuff, and it mostly works.”
So this is really like you’re just overseeing the AI write the code. Andrej, it sounds like, is doing very little of the writing. He’s basically doing what I heard some people predict we would arrive at this point, which is: English is the new programming language. You just sort of say in English what you want the code to do, and then it does it.
Yeah, and this is different from the AI coding tools that existed even a couple of years ago. GitHub Copilot was one of the early AI coding assistants. Basically, it would just autocomplete your code, right? You could be writing a line of Python or JavaScript, and it would see what you were up to and complete it for you. You would just press Tab, and it would go on to your next thing. But you still had to know how to program to use those tools effectively. What’s been happening in the last couple of years, and has gotten quite good over the last six months, are these tools that essentially remove the need to program at all.
So now there are lots of tools out there. There’s a tool called Cursor. There’s a tool called Replit. There’s Bolt. There’s Lovable. There’s a bunch of these tools where, basically, you just go in and get a text box that says, “What do you want to build?” And you say, “I want an app that does this, this, and this,” and it goes out and builds it for you pretty much instantaneously.
Now, I have a friend who runs a tech company, and he once made fun of this whole idea to me by saying, “Hey, you want to talk about programming in the English language? That’s what I do all day long as a CEO. I’m constantly telling my engineers in English what to do, and it works maybe a little over half the time, but maybe not much more than that.” So what has been your early experience of vibe coding? What have you been trying to build, and how has it been going?
I want to talk about my projects, but first I want to talk about my own history with this stuff because I am a former programmer. When I was a teenager, I was into coding. I would build websites. I would build little JavaScript projects. I spent a very excruciating summer trying to teach myself Flash so that I could make animated cartoons like Homestar Runner.
Mm-hmm.
And then I dropped it. I went to college. I learned about journalism. I thought, “Well, this is the path I want.” I became a wordcel, and then I stopped coding altogether. And so when I started hearing about these tools that would let you code without knowing how to code, I was very interested, and I started experimenting. One of the first things I built was this podcast summarizer, where you can take a podcast that’s very long and use AI to transcribe it, then use a different AI to summarize the transcripts and put it all into a searchable database so that I could say, “Okay, I don’t feel like listening to this 5-hour podcast about AI, but I can basically get the executive summary using AI.”
Mm-hmm. Mm-hmm. So tell us a little bit about your setup. What software are you using to do this?
So I’ve been trying a couple of different tools. Sometimes I just use the raw AI models themselves, like Claude or ChatGPT. Those tools are quite good at some projects, but for the most part they can’t actually run the software to test it inside the window, so it does require some copying and pasting. This new kind of app that I’ve been using is a more integrated development environment—
An IDE?
An IDE. So Cursor is the one that’s really popular right now. If you’ve never used an IDE before, you might find it a little puzzling. I certainly did. But it basically lets you prompt the AI to write the code for you, automatically debug it, deploy it within a little test window, and then push it out onto the web, where people can actually use it.
So tell us about some of the other projects you’ve been building.
So in addition to my podcast summarizer, I also had AI help me redesign my website to look more cyberpunk. That was the aesthetic I was going for.
Wait, is this live? Can I view it?
No, it hasn’t—
Okay.
—it hasn’t been deployed yet—
I see.
—but it’s going to be there soon. I had it—
Wait. How did it make it look more cyberpunk?
It just redesigned the whole thing.
Oh, okay.
It has bright neons—
Mm.
—sharp edges, cool scrolling, sort of parallax-style animations.
Do you have a bionic arm in your author photo now?
Yes.
Okay.
I built a tool to pull all of my bookmarks from X into a spreadsheet. I use X a lot to bookmark things that I find interesting or want to return to later.
You’ll say, “Wow, that’s the most racist thing I’ve ever heard.”
Yes. So now I have a tool that will go through all of my bookmarks and pull those into a spreadsheet that I can search later. That one was very interesting because it basically presented me with a couple of options after I asked Claude, I think, to build this tool for me. It said, “Well, we could use the Twitter API, but that costs money, and if you don’t want to pay that, we have this other way that we can do it that involves using a browser to scrape the bookmarks from Twitter.” And so I went with that version.
Wow, you realize that by doing this you are now essentially an armed combatant in Elon Musk’s war on bots? You are the bot that Elon Musk is trying to destroy.
Come at me, bro.
Good luck. Good luck, buddy.
I’ve got my bookmarks. Now I don’t need it anymore.
All right, what else?
So the thing that I built most recently was yesterday, when I was trying to determine if various objects that I’m moving to my new house would fit in the trunk of my car.
Mm-hmm.
And so I built an app called Will It Fit in My Trunk?
Now, this feels like a classic math-based problem that maybe Vlad’s thing could have helped you with. But you used something else. How did it go?
Yeah, so far so good. It hasn’t steered me wrong yet. But this speaks to what I think is so fun and interesting about this genre of coding project: you can really just build what I call software for one.
Mm-hmm.
A software company would never build a tool for wide release that let you figure out whether various objects would fit in your trunk. That is not a big total addressable market.
Sounds like you never saw Trunky in the App Store. I’m just kidding, that’s not a real app. But yes, you’re right.
So this style of coding really makes it possible to build things that you and only you need or will ever use.
And there’s something fun about it because I think, particularly for you and me, who actually enjoy technology and like using it and trying new things, coding can feel like actual magic, right? It can feel like wizardry. And if you are the one who is all of a sudden wielding the wand and making things happen, then you’re feeling great.
Yeah, it is the most fun that I’ve had with these AI tools in a while. I think it is the most fun thing you can do with AI in today’s world, and it has really connected me back to my teenage coder self and reminded me what I loved about it back then.
I spent a lot of time in college and afterward writing HTML. I had a program called Dreamweaver—
Love Dreamweaver.
—and got pretty handy with it. But if I had been able to chat with an AI assistant about why I was having trouble with my Movable Type installation in 2004, my website would have been sick as hell.
Yeah.
Yeah.
So, Casey, I’m sure you have some niche software needs in your life.
Absolutely.
And I asked you the other day what I could build for you using my vibe-coding tools—
Mm-hmm.
—and what did you say?
What I said was, “I need help with my hot tub.”
Go on.
Well, listen, here’s what they don’t tell you about buying a hot tub. When it comes to your house, and you decide, “I want to use the hot tub that I have just purchased,” you have to become a chemical engineer. Here’s what I mean by that. You open up the manual, and all of a sudden you learn you are going to need to monitor the pH balance of your water. You’re going to need to monitor the alkalinity of the water. You’re going to need to monitor the calcium level of the water because if there is not enough calcium, it can somehow corrode the jets in your hot tub.
And needless to say, Kevin, I don’t have a lot of experience mixing chemicals to adjust alkalinity and pH levels in bodies of water, and I thought, “Well, how am I going to do this?” And so then I actually do start using the chatbots, not to write me software, but essentially just to say, “Please, God, help me. What do I do?”
And there are so many things to keep track of. You have to put various chemicals into the hot tub at different intervals. You replace this once a month. You replace this once a quarter. Twice a year, you have to drain the entire tub. Once a month, you have to shock the hot tub. Don’t ask me what that means. I just read that today. It’s giving me a nervous breakdown.
You have to just show it some spicy tweets.
Yes, exactly. So there’s so much to keep track of, and I thought, “Well, if I were going to build software just for me, it would be something that just checked in with me to prevent my hot tub from turning into a bacterial soup.”
Well, Casey, I have great news for you.
What’s that?
I built you a hot tub tool.
Oh, my goodness. You vibe-coded on my behalf?
I vibe-coded on your behalf.
Thank you.
So after you told me about the issues you were having with your hot tub, which are very relatable, by the way—
Yeah.
—listeners are all going, “Ah, me too. I have an issue with my hot tub.”
Listen, we have a very wealthy audience that is constantly buying huge upgrades for their homes, and every once in a while, Kevin, we have to do something for the C-suite listeners.
Exactly. I took this as a brief—
Mm-hmm.
—and I went into a tool called Replit. I said to the tool, “Make me an app that will tell me the things I need to do to keep this specific kind of hot tub working properly, and put it in a tool that my friend can use.” Because this is a tool that tells you when to service your hot tub, I was going to call it Hot Tub Time Machine.
That’s very good. I like that.
So I built you a website called Hot Tub Time Machine.
Oh, my gosh. This is wonderful.
So let me show it to you.
Okay.
Now, I just want to tell you and caveat this by saying that I did not choose the design or the color scheme here. That was all the AI.
Okay, great.
Open up the link I just sent you.
All right. I’m opening up the link. Okay, so it is quite pink. It’s pink on pink, which is a color scheme that you don’t see a lot outside of the Barbie franchise. But no, there is some black to it. It looks beautiful. It says, “Hot Tub Time Machine, your retro-futuristic maintenance companion.” And it even created a little logo, which I’m going to assume is a drop of water?
Sure.
Okay.
We’ll go with that.
And there are 2 modules. There’s a manage tasks module and a view schedule module.
Yeah, so this tool is very simple. This is a prototype. We can flesh it out if you want to, but basically—
Is this in alpha or is this in beta?
This is in alpha.
This is in alpha. Okay.
You’re the single user of this app.
Okay.
And basically, I’ve set it up so that it will email you weekly, monthly, quarterly, and annually with a list of everything you need to do to—
Incredible.
keep your hot tub in working order. And as a special bonus, every email will come with a poem about hot tubs.
Fantastic. Well, can I start clicking around?
Yeah, click around.
All right, so I’m going to click on View Tasks. And, all right, this brings up—There’s a module where I can add a new task, but there are also some existing tasks. It includes weekly, quarterly, monthly, and annual maintenance. And it is, frankly, an overwhelming number of things to do.
Yeah, you bought yourself some chores when you bought that hot tub.
It really is just a wall of text—
Yeah.
of things that I have to do. Every week, I’m apparently supposed to spray off the filter with a garden hose—
Yeah.
—and add 1 cup of non-chlorine shock, especially after parties or what have you. So, yes, lots in here. Okay, and it also looks like I can add another task if I decide I want to do more.
Here.
I want to do more.
I’m going to click the little test button to—
Okay.
send you one of these.
Okay.
And you see if it shows up—
Okay.
in your inbox. Let’s see.
Yes. “Time to maintain my hot tub.” And there are some step-by-step instructions that I can follow, and below that, the hot tub poetry corner. Should I read this poem?
Yes, please.
All right. Here’s the poem:
“Bubbles rise in swirling steam.
Time machine of warmth and dream.
Nordic waters, pure delight.
Maintaining bliss both day and night.”
I guess I would say that is okay. Nordic, of course, a reference to the fact that I have a Nordic Jubilee Series hot tub.
Yes.
Yeah.
So I built this all in about half an hour—
Okay.
without writing a single line of code.
Okay.
And I want to share that with you because not only will it help you with your hot tub issues, but I hope it will also show you the promise of vibe coding.
Yeah. Well, I feel like you have shown me the promise of vibe coding. Now, was there anything about this that was particularly tricky, or did you get stuck on anything?
Yeah, so there are some things that it can’t do, right? If it needs to authenticate you into some service or set up a database, you have to manually step in and do that. There are some things that it just can’t do because no human programmer could do them either. If there’s no API for something, for example, it won’t magically invent one.
So there are some boundaries and limitations, and I would say it still benefits you, when using this stuff, to have at least a little bit of programming experience—
Mm-hmm.
because there are certain decisions that it will prompt you to make where you’re like, “I don’t actually know what these terms mean or what the right decision is here.” And you can ask the AI to just make the decision for you, but you might not be totally happy with the result.
Now, during this process, I’m curious if you felt like you were learning something about the coding process. If you spent the next year making these little one-off apps, do you feel like you would maybe be a decent junior software engineer? Or is the idea actually not to get into the details, to just let it build things, and if you don’t know what it’s doing, that’s none of your business?
Yeah. I think I’m more in the latter camp. This was the part that I found fascinating about what Andrej Karpathy said about vibe coding. He’s an extremely good programmer, but he says that he now can enter this mode where he basically just says, “Okay, okay, okay. Accept, accept, accept,” and the computer will go off and do its thing.
I don’t know enough about programming to dive into the weeds of what the AI is doing and the decisions it’s making. I just have to look at the end result, and there’s something exciting about that where I feel like things are just happening magically on my behalf. But that’s also, maybe it’s inserting malicious code. Maybe it’s doing something that I don’t want it to be doing. I have no way—
Ha.
as a nonprogrammer—to know whether that’s happening or not.
Have you checked your computer to see if it installed a Bitcoin miner while you weren’t looking?
I have not, but that would be pretty tricky.
9. The Junior Developer Problem
Well, Kevin, this experiment has me thinking about a blog post I read this week by a guy named Namnoye Goel. His blog post was titled “New Junior Developers Can’t Actually Code.” This post got 1 million views, according to the post that I’m looking at.
And he’s saying that when he talks to junior developers, they are having an experience very similar to yours, which is that as they are building these systems, they are essentially just supervising an AI. They aren’t actually getting their hands dirty and understanding which mechanisms are leading to which results. So while this is great for you, I do think it raises the question: What happens when most of our software engineers are building systems that, in some fundamental way, they don’t understand?
Yeah, I think this is a very real thing. The flip side of me, a noncoder, being able to build stuff is that if real coders are using these tools, there’s no incentive for them to learn the basic skills of programming and learn the syntax of the different languages. And yeah, I don’t know what to do about that.
It seems like a version of what happened when we all got Google Maps on our phones, is that people started losing their sense of direction. There’s this kind of skill-atrophy issue that people worry about. But I think that the returns to knowing how to use these things effectively are still great enough and still require enough knowledge of how the various pieces of code fit together that it still does make sense for people to learn to code.
I’m not one of these people who thinks, “Learn to code” is totally over. I think for some people it’s still a very useful skill to have. But I think that in the future, the role of the software engineer will become more like a product manager, where you are essentially supervising the product, laying out the vision, overseeing the design, stepping in to fix things when they break, but you are not actually in the trenches of the code, writing lines of code by hand.
Hmm. All right, well—
What do you think?
What I think is that as AI systems get more and more powerful, we need people who do understand them in a very detailed, technical, complex, down-to-the-metal kind of way. If we don’t do that, our only alternative will just be to trust the AI when we ask it, “Hey, how do you work?” And there are a lot of reasons why I don’t want to end up in that world.
So I’m comfortable having fewer people in this world who know the code at that level of detail, and it’s fine to me if most software engineers don’t, but I want a solid core of people who do.
Yeah. And I would like to continue with my vibe-coding experiments, trying to build increasingly more useful tools for myself and my friends.
And I am thinking about starting, because if you can do it, surely I can.
Yes, anyone can. That is sort of the point. And I also would love to hear from our listeners what they are vibe coding. What tools and apps are you building using AI that are solving your own personal, specific problems?
Did you invent a novel bioweapon using ChatGPT? We'd love to hear from you.
Yeah, please email that one to tips@fbi.gov.
But the others, hardfork@nytimes.com.
And we'll be reading every email.