电子邮件的未来:Superhuman CTO谈你的收件箱才是真正的AI Agent(而不是ChatGPT)—Loïc Houssier
Superhuman的核心押注是:收件箱,而不是通用聊天机器人,将成为知识工作的主动式AI Agent。Loïc Houssier认为,ChatGPT需要用户先提供上下文,而Superhuman可以从工作已经发生的地方直接行动:电子邮件、Coda摄取的知识,以及Grammarly跨应用存在的能力。“我们就在你工作的地方”(We are where you work),这场战略竞争本质上是“被动响应”对“主动出击”。
产品路线图依靠小步、 high-confidence 的自动化推进,而不是显眼的“AI功能和闪光特效”。自动分类、邮件串摘要、两天后跟进提醒、预先写好的草稿、可用时间回复和新闻简报合成,都服务于Superhuman的核心约束:绝不能拖慢高要求用户。Ask AI如今专门对抗“Agent偷懒”,通过每次约40封邮件的分页持续搜索,直到找到那根针。
Superhuman想成为AI执行助理,但最后20%的判断仍是最难的部分。自动草拟系统正在学习拒绝供应商、将求职者转给HR等重复行为;日程安排测试版则会把3个可用时间直接插入草稿。然而,“80%”的准确率仍可能制造过多丢弃工作,日历数据也无法可靠判断何时值得打断日程去响应VIP,或某段关系是否在上周发生了变化。
1支由3名工程师组成的AI团队,通过窄工具、模型路由和异常严格的端到端评测支撑产品。不同工具分别检查可用时间、关系、搜索和写作;Sonnet尤其擅长Agent交接并避免偷懒式交接,测试中的OpenAI模型在这方面较弱,而Gemini最近也成为另一项强劲选择。校准测试包括Rahul要求找回5年前一封写明桌子木材的承包商邮件:“在我们搞定这个查询之前,他不会满意。”
质量优先于推理成本,这为高端定价和可预测产能的供应商创造了空间。Superhuman会先从“最高端、最贵的模型”开始,体验跑通后再优化;Baseten的盒装产能模式比波动的按Token计费更有利于预测3个月和6个月的成本。一些客户表示,宁愿每月支付$200使用最好的模型,也不接受“半吊子的垃圾”(half crap),因为节省1个高管小时的价值可能是订阅费的10倍。
运营数据表明AI正在提高工程产出,但Loïc明确拒绝把全部功劳归给AI。数据来自员工自报:约80%的人会给PR打标签;在被标记的PR中,约80%–90%涉及AI,其中约90%的使用效果为正面。人均每周PR数量从Q1约4个升至Q2的5个,再到Q3的6个。Superhuman还配套提供1小时预算审批、24小时安全审查、约50名工程师、约100,000名付费用户,以及平均任职约4年的资深远程团队。
随着实现成本下降,产品判断和工程基本功反而更有价值。Loïc把AI编程比作从汇编转向C:这只是又增加了一层抽象,仍要求理解内存、服务器和底层机器。他的判断没有保留:AI“会更快地区分优秀工程师和糟糕工程师”(will separate faster the good engineers from the bad engineers);创业公司如今可以在2周内复制技术,差异化则转向工作流、延迟、界面质量和对用户的同理心。
1. DocuSign让Loïc明白,分发和合规可能压过漂亮技术
Loïc最初学习数学,差点攻读3年制应用数学与密码学博士,之后转向进攻性安全。他在潜艇行业工作了约1年,也因此失去了作为安全研究员时所拥有的技术权威。面对鱼雷、雷达和核发动机领域的专家,他学会“把自我放进兜里”(put my ego in my pocket),通过提问来领导团队,并逐渐习惯帮助那些在各自领域远比自己懂行的人。
他在巴黎创办的电子签名公司技术实力很强,也正在成为法国行业领导者,但DocuSign使用的是另一套商业语言。欧洲团队向CIO销售技术能力,DocuSign则向HR负责人和其他职能部门销售商业价值。这种反差让合作、最终被收购,比继续单打独斗更具吸引力。
这笔交易需要6个月的业务剥离,因为法国财政部阻止了整体出售:这家公司还为法国国防部处理安全与身份认证工作。团队、源代码、系统,乃至数据中心基础设施,都必须拆分并复制,DocuSign才能收购签名业务。Loïc回头给出的建议极其直接:“疯狂。别这么干。”
Swyx提出“DocuSign的人都在做什么?”这个问题,揭开了隐藏的规模经济:本地数据中心、不同的欧洲签名标准、FedRAMP、各国数据驻留要求,以及密钥一旦被盗就会失效的本地部署设备。从外部看像是人员臃肿的配置,实际上可能是进入市场所需的基础设施;日本市场还要求理解印章文化,而不是想当然地套用西式签名。
他对法国更广泛的结论是:优质教育和工程能力还不够,还需要韧性、创业心态、产品驱动增长和全球野心。欧洲创业公司应考虑从一开始就面向英语世界和全球市场,而不是先从法国或意大利起步。
2. Superhuman只在能让高要求用户更快时加入AI
Loïc的产品原则不是上线“AI功能和闪光特效”;每项功能都必须让那些本来就有极高要求的用户更快。延迟和打扰都是产品失败,因此智能必须在正确的时点出现,同时不破坏Superhuman的速度。
早期功能刻意不炫技:给推介和营销邮件分类、总结长邮件串,并识别仍需回复的消息。邮件两天没有得到回复后,Superhuman可以重新推送这条线程;下一代功能则直接准备好跟进邮件,把“糟了,对,我本来想提醒大家的”变成按一下Send。
同样的演进也覆盖可用时间请求和任务委派。草稿可以建议抄送合适的执行助理,正在开发的日程安排测试版则会插入可直接使用的时间。每一步都在移除一个小决策,而不是要求用户进入单独的AI工作流。
Loïc如今会自动归档约30–40份新闻简报,周五再询问本周收到的全部内容,以及哪些值得关注。这种行为变化和功能本身同样重要:用户从ChatGPT那里学到,查询一个知识库可能比维护“一份邮件待办清单”更容易。
3. Ask AI的设计目标是完成任务,而不是返回选项
Swyx最初是在Superhuman搜索栏里误触发功能,后来开始主动提出更宽泛、尚未成形的问题。如今他用它找合同,Alessio用它找电话号码。Loïc则找回了6个月前一次会议中的PowerPoint链接,估计上下文搜索节省了约30分钟。
Loïc把“Agent偷懒”定义为一种明确的质量缺陷。当他说明天找15分钟审阅一份文件时,系统应该直接选择并预订一个时间,而不是给出几个选项后把工作交还给他。一个称职的人类行政助理听到的是“找到它,做完它”,Agent也应以此为标准接受评估。
深度检索无法把整个邮箱塞进一个上下文窗口,因此Ask AI使用语义分页。它可能先检查40封最有可能相关的邮件,排除后再检查下一批40封,不断扩大语义邻近范围,直到找到正确结果——即使有数百封邮件在语义上都很接近。
编排层由可用时间、联系人关系、草拟和其他动作的窄工具组成。Agent先制定计划,为每一步选择最佳工具,再执行;不存在一个“几乎什么都能做的万能大工具”。
4. 模型选择不如能暴露不同失败模式的评测重要
据Loïc介绍,只有3名工程师负责这套AI栈,但他们已经在框架、模型和内部代理层之间反复迭代。Sonnet在Agent交接和避免偷懒式交接方面尤其出色,测试中的OpenAI模型在这方面较弱,而Gemini刚刚加入比较,成为另一个有潜力的选项。
Superhuman最初的评测是朴素的“查询、回答、查询、回答”模式。后来它演变为按维度组织的标准查询集,覆盖Agent交接、在“海量邮件”中深度搜索、日期理解以及其他反复出现的弱点。每个端到端结果都按这些维度评分,而不是笼统判断好或坏。
Rahul提供了对抗式现实:他每天可能收到500–1,000封邮件,并要求找出5年前一段承包商讨论,确认一张桌子使用的木材。日期问题又增加了另一类失败,因为模型很难正确理解相对于今天的“上个季度”等表达。
Loïc区分了普通可靠性和Superhuman用户期待的“档次”。Toyota可以是高质量产品,但Audi或Porsche买家期待的是另一种东西;Superhuman的高要求用户迫使团队在细节上投入不成比例的精力,而低价、低强度的产品可能会容忍这些问题。
5. 收件箱正在成为可计算的权威记录系统
Swyx提出的挑衅式问题——“我上个月在Waymo上花了多少时间?”——把收据变成了一个会计数据库。Agent筛选Waymo邮件,提取行程时长,再通过生成代码汇总,因为“LLM不擅长数学”。Superhuman当时正与Anthropic讨论更广泛的代码执行组件。
Superhuman仍依赖Gmail和Outlook提供垃圾邮件识别等基础设施,而不是变成另一台邮件服务器。Loïc认为,重建客户公司系统里已经具备的能力,用户价值很低:“如果我们直接接入并把它做得更好”,在邮箱之上的那一层就足够了。
速度要求本地副本,因为每次交互都必须在100毫秒以内完成。安装时会下载最近30天的邮件用于离线运行,之后继续保留新增历史;一名使用2年的客户因此可能在本地存有2年的邮件。移动端曾使用Realm,但现在已基本停止维护,团队正考虑SQLite等替代方案。
AI检索则在Turbopuffer中对约5年的历史邮件增加嵌入和混合搜索。推理分布在OpenAI、Anthropic、Gemini,以及由Baseten托管的Llama或BERT分类器等开放模型和供应商之间;具体路由取决于使用场景和当前模型质量。
6. 离线推理首先是质量问题,其次才是成本问题
设备端创业公司往往主打更低的推理费用,但Loïc表示,Superhuman用户“要的是质量,而且愿意为质量付费”。他感兴趣的是离线语义搜索:产品本来就能在无网络时运行,但一旦远程推理消失,AI能力就会明显变弱。
限制不仅在智能,也在占用空间。Superhuman已经因为本地缓存邮件而消耗存储和内存;Swyx举出的DeepSeek-V3.1例子约有600B参数,说明直接把前沿规模模型装进设备并不现实。任何本地能力都必须适配现有占用,不能进一步恶化用户对应用体积的抱怨。
Superhuman更希望掌握端到端体验,就像其移动应用使用Swift和Kotlin,而不是React Native,以获得原生感。但Loïc也表示,他“很希望设备供应商能做得更好”;目前iOS方面的表现一直不尽如人意,这给专门供应缺失层的公司留下了空间。
7. 高端AI经济学偏好先证明价值,再优化Token
Loïc的排序非常明确:先使用最好、最贵的模型,建立所需质量,再降低成本。“如果它成功了,那就很好,即使它很贵。”成功会带来足够规模,使优化值得投入,而不是让一项尚未验证的功能一开始就背负过早的约束。
Baseten固定容量、盒装式的模型部署,比纯粹按Token计费的无服务器推理更容易控制支出。在更大的、“接近IPO阶段”的业务框架下,Superhuman越来越需要可信的3个月和6个月预测;即使采用率波动,固定容量也能给CFO一个有边界的成本区间。
Token预测仍然混乱,因为Superhuman处理的邮件从极短回复到超长线程都有。团队必须估算邮件中位长度、平均消耗量和采用率,模型里“总会有一些魔法”。Swyx把企业采购重新表述为每万亿Token的价格,而不是个人开发者关注的每百万Token价格。
Alessio问,如果给每封收到的邮件都生成草稿,会不会摧毁$40订阅的经济模型。Loïc的回答基于价值:一些客户明确要求最好的模型,并表示愿意每月支付$200,因为CEO或VC的1小时时间价值可能是订阅费的10倍以上。
8. AI执行助理受阻于判断,而不是草拟
内部自动草拟仍在测试,正在学习重复性行为:Loïc会礼貌拒绝每天收到的数百封AI工具推介,并把求职者转给HR。Swyx认为,把既有模板与AI个性化结合起来,可能无需接近AGI就能解决实际的“最后一公里”。
覆盖率是尚未解决的产品门槛。如果80%的草稿有用,但用户还必须丢弃20%,自动化可能带来烦躁,而不是杠杆。Loïc问,能够接受的比例究竟是90/10,还是80/20。真正相关的指标是节省了多少注意力,而不是生成了多少文字。
日程安排暴露出更深层的缺口。Superhuman可以草拟3个午餐时间,但Swyx希望自己完全不参与;人类EA知道哪个VIP值得调整繁忙日程,也可能知道上周的谈话已经改变了双方关系。仅凭交互次数无法编码这种判断。
Loïc仍将AI EA称为“目标”,同时让人类判断保留在流程中,可能由用户的EA参与。Swyx提议收购一家人类助理公司,观察工作流并生成专有训练数据;Loïc同意这可能是最好的学习路径,但也称其运营承诺“强度很高”。
9. 电子邮件可能从文字行走向语音和环境式沟通
Loïc的3个孩子用手机聊天,通过WhatsApp与家人联系,也通过Snapchat、TikTok或Instagram沟通;他说一切都在变成语音,而他们越来越想要一段关于文章的视频,而不是文章本身。他更大的判断是,人类过去写作,是因为口头故事缺乏存储能力;如今音频和视频无处不在,这一历史约束正在减弱:“还有什么必要写作?”
放到电子邮件上,问题就变成:1年后用户还会不会打字。Rahul也许会通过说话宣布一项功能,收件人在通勤时听到的是他的真实声音,而不是模仿他“声音和语气”的文字。机会在于重新设计沟通,而不是给今天的消息行表格外挂一个聊天机器人。
浏览器之争遵循同一逻辑。Swyx认为,掌握浏览器可以最大化上下文;Loïc则预计浏览器会变得更薄,甚至消失在操作系统中。Raycast已经替他取代书签和导航,浏览器最终可能只剩下网页视图、本地存储和扩展。
10. 跨应用上下文是Superhuman对ChatGPT的回答
Coda Packs是用于摄取数据的集成,Superhuman在技术上也已经拥有聚合公司知识的摄取管道。Grammarly可以知道用户正在Google Docs中工作、撰写LinkedIn帖子,或在Jira、Salesforce和LinkedIn之间切换,然后准备一封邮件——不过Loïc明确指出,知道这些并不意味着Grammarly会使用全部数据。
一旦这些资产汇合,系统就能把紧接在邮件之前发生的工作上下文加入消息。收购仅在3个月前完成,因此Loïc将这种融合描述为未来优势,而不是已经完成的能力。
ChatGPT则不会天然知道用户刚刚离开了哪个应用;它需要等待用户粘贴上下文。因此Loïc直白地将OpenAI和Superhuman描述为竞争对手。Superhuman想要的优势是无处不在:“相较于等待你提问的ChatGPT,我们可以更主动。”
Swyx质疑Loïc关于知识图谱系统“还没到位”的说法,因为Superhuman自己并没有尝试过;Loïc立即承认了这一点。他真正担心的是分类法:人物和公司显而易见,但项目、任务、计划、层级以及领域特有功能,在不同用户之间差异极大。
Obsidian用户可能建立一套完全主观的图谱,这让通用生产力本体变得困难。Loïc预计Superhuman会接入专门的图谱引擎,而不是自己构建:持续记住一切类似“Jarvis”,每当新邮件或交互数据出现,就需要大量重新计算。
11. AI提高产出,但组织质量决定这种提升能否复利
Q1,Loïc取消了采购摩擦:实验工具在1小时内获得预算批准,安全审查在24小时内完成。采用率迅速上升,在前端工作上的表现尤其好,但当时在Go后端和Swift上的效果较弱;v0赢得了内部交互式产品原型的自由市场。
Q2聚焦测量和用例发现。探索陌生代码区域的时间,在Claude Code帮助下从约1天降至约30分钟;在他的例子中,他还使用了Warp。工程师为每个PR标记是否使用AI,以及AI是否带来帮助;约80%的人提交了反馈,其中约80%–90%使用了AI,而报告中的约90%使用效果为正面。
人均每周PR数量从Q1约4个升至Q2的5个,再到Q3的6个。Loïc强调,AI只是原因之一,技术战略、组织设计和更清晰的优先级同样重要;PR数量也只是产出的代理指标,不是完整的生产力衡量标准。
Superhuman仍然精简:约50名工程师服务约100,000名付费用户,平均任职时间接近4年。其3人AI团队分布在Patagonia和加拿大;收购完成后,邮件业务作为拥有独立损益表的“复合型创业公司”运营,Rahul仍负责领导,而Shishir的管理团队实际上充当其董事会。
Alessio的反驳值得保留:工程速度加快,可能压垮市场、客服和产品一致性;发布更多功能,也可能意味着更多Bug。因此Loïc拒绝陷入“我们能做更多,那就做更多”的惯性循环,即使大公司为扩张提供了更多产能。
在劳动力问题上,Loïc不接受从每周4个PR增加到6个,就应自动带来50%薪酬增长的观点。公平薪酬仍然必要,但资深工程师也能从低摩擦团队中交付有用工作获得满足:“我基本上可以成为最好的自己。”
他最后的判断毫不含糊:人们仍然应该学习编程。AI更像从汇编转向C,而不是废除底层系统知识;优秀工程师会变得“惊人”,而“糟糕、懒惰”的工程师可能把生成结果误认为理解,反而变得更差。
因此,商业护城河转向产品能力。如果创业公司能在2周内大致复制技术,差异化就来自对用户、工作流、视觉质量和延迟的理解。Superhuman正在美洲招聘远程、以产品为导向的工程师,包括那些关注延迟、因为用户能直接感受到延迟的后端工程师。
As a human species, we started to write because we didn't have enough storage for the stories we were telling each other, so we had to write to store those stories. Now all the content can be stored on YouTube, TikTok, or whatever. What's even the need to write? What's the need? Because everything can be vocal.
Being a bit more grounded, what does it mean for the future of the user experience for email and communication? Will people still type, or will they just talk to emails and want to hear an email? This is where it becomes interesting because Rahul, as a CEO, maybe next year he doesn't want to write to you with the new feature. Maybe he wants to talk to you. Then the way you will receive our marketing campaign about the new features is that you'll be in your car commuting, listening to Rahul talking about that.
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.
I just realized I have the tough job of always pronouncing names.
I know, man. You’ve got to prep. You go on YouTube, namepronunciation.com.
Loïc Houssier, welcome.
Wow. I'm impressed.
Did I get it right?
I know. You got it right.
Yeah. Okay, all right, cool.
You got it right. I'm surprised. Usually I make the joke, like, “Yeah, you know what? You can say it the way you want and everything,” but you nailed it, so I'm impressed. Thanks for having me, guys.
Yeah, of course.
Thanks for coming by. So you're CTO of Superhuman Mail, which is the new name for Superhuman. I've been using Superhuman for a long time. I think I had one of Rahul's personal onboarding sessions back in the day. We're here to talk about all things AI engineering, but you also have a lot of history in Productboard, Firstbase, DocuSign, and nuclear submarines.
Yes, yes. That's kind of the fun icebreaker that I give to people sometimes. They're like, “Two truths and a lie.”
Yeah.
“I went into a submarine,” and people are like, “Nah, no way.” But I did. I did. I spent one year working around submarines.
The trajectory is a bit weird. You were an engineer, and then you were sort of chief of staff on some submarine thing. And then you went back to engineering.
1. From Math to Submarines
I started studying math. I'm a math graduate. I was about to do a PhD in math and applied math in cryptography—crypto before crypto, to some extent. It was cool for a moment, and then I was like, “No way I spend 3 years of my life on the same topic.”
But in the same lab, there were a bunch of people doing security, offensive-security-type stuff, and I was like, “That's what I want to do.” So I was basically an engineer and a security researcher in that lab. I did that in a pretty big corporation. I was in telco and then in the defense industry.
In the defense industry, they have this nice career framework. You're young, high-potential-ish, in quotes, so they want you to do different types of jobs and have a spiral career so that at some point you eventually reach the C-level. They gave me the opportunity to be out of the tech industry for a year, and I went to a harbor. I was there as a financial controller and a process-improvement type of person, basically helping people do a better job.
That was interesting because I had no clue about torpedo systems, radar systems, or even the nuclear engine inside a submarine. Still, I had to help people take a step back from what they were doing. That was really fun because I came from Paris with my tie, my suit, and my ego. I was used to driving people through my technical legitimacy in the security space, and all of a sudden I didn't have any technical legitimacy at all, but I still had my ego.
So I put my ego in my pocket and basically drove by questioning people: “How does that work? Help me. I don't get it.” Just by questioning, I built a new skill: getting curious, understanding how people work, and being comfortable facing people who are way smarter than me and know their fields better, but probably having a way to ask questions to help them identify gaps or productivity gaps, for example. That was cool, but I missed the tech, so I moved back to the tech industry after basically 2 years.
What are some of the other highlights or stories from your other experiences? DocuSign is another product that we all use.
2. The DocuSign Carveout
No, DocuSign was cool. DocuSign was cool because it was an acquisition.
I have an interest in DocuSign, yeah.
Yeah. I was CTO of a small company in Paris, and we were a typical European company, Alessio: very focused on the tech, not very focused on the marketing. We were one of the biggest signature companies in Europe, but it was a very fragmented market. We were winning in France and starting to expand, and DocuSign came and said, “Guys, we need to do a partnership and everything.”
Pretty soon, they understood that the European market was tough, and the technology behind DocuSign wasn't sufficient because of the lack of standards, compliance, and everything. So pretty soon, they were like, “It's with us or against us.”
Mm-hmm.
But the way they were explaining the value, I was like, “Holy cow. We're not talking the same language. We're doing the same job. We're selling the same type of software. But we're talking to CIOs from a technical standpoint, while they're talking to the head of HR and heads of functions and selling them the value.”
It was pretty easy for us to understand: “Whoa, whoa, whoa. That's not the way to sell a product. Better to partner with them.” So they did an acquisition, but it wasn't a full acquisition. It was a security-oriented company with 2 business lines: 1 doing signatures, which was the one DocuSign was interested in, and the other doing strong authentication—PKI stuff, SSL certificates, those types of things.
We were working for the Department of Defense in France, so we had the Ministry of Finance in France basically saying, “No way, no-go. You cannot sell.” We had to do a carve-out, which is one of the funniest acquisition types you can do. You have your team, and you need to divide everything into 2: your team, your systems, your source code, and all of that. Even your data center—you have to replicate it and get rid of all the shared systems and everything.
We did that for something like 6 months to be able to sell the carved-out company to DocuSign. Crazy. Don't do that.
Are you still involved at all with the French startup ecosystem? I'm curious how you've seen things evolve since then.
3. France Targets the World
Yeah, it's pretty interesting. I've seen a change. Now that I'm getting some gray hair and have some experience, I try to give back to some extent, so I spend more time helping the ecosystem there.
But it's funny to see the difference. We live in a small bubble, and it's crazy to see how even other tech scenes are different. The grit—just the grit to get shit done and move forward and everything. They have great education, great engineers, and all of that, but not the mindset of creating things.
Mm-hmm.
There aren't that many entrepreneurs. It's changing. We've had some successes in Europe, especially in AI. There is some cool stuff happening. But still, the way to think about product-led growth—Superhuman nailed it—the way to structure your organization to scale fast, and the level of ambition as well.
Maybe don't target France or Italy to start with. Target English and the world from the get-go. That would be something to think about. I'm doing that quite a bit, and it's highly rewarding. It's pretty cool.
There's a common question that people have about DocuSign that I'm just going to indulge.
Sure.
What do all the people do at DocuSign?
I love it.
You know this is a meme, right?
No, no, no. It's—
Why do you need so many people?
Well, you have signing.
Yeah.
Why do you need 3,000 engineers?
It sounds crazy, but you want to nail Europe. You need a different product. You need a different team to run your local data centers because of the compliance.
You cannot just run your data centers from the US, so you need a local team there. And by the way, the way to do a digital signature in Europe is totally different. The stack itself is different.
Oh.
The way to make a digital signature is different; the standards and the ways are not the same. So you need a dedicated team to maintain that thing. The same way, some people want to have DocuSign on-premises, so you need a team building an appliance to basically plug and play, like, “Okay, you have your DocuSign appliance doing—”
There’s a DocuSign box?
There’s a DocuSign box.
Wow.
It was an acquisition made in Tel Aviv at the time. Wonderful people were building a security appliance where you check the box and, poof, the keys disappear. If someone steals your box, no one can sign in your name.
You’re kidding me. Oh my God.
No, some banks—
What if there’s an earthquake?
That’s a good question. They are mounted, so there’s earthquake mitigation associated with this. Just that, but apply it to FedRAMP.
Now they make more money.
Dedicated teams—
Wow.
Dedicated data centers. And we need DocuSign to run in Canada because of data residency, or we need the same in Australia. Okay, cool. And now you have something even different. We want Japan as a market. But Japan is not signature; it’s hanko. It’s kind of like a stamp.
So you need a team to understand how the Japanese market is thinking about even processing an agreement. It’s totally different. Then you have verticalization, with different verticals and everything. I mean, it’s a good business, it’s well run, and people are not coasting there. So there’s a lot of work, and it’s very interesting to see it from the inside, because when you see those memes, you’re like—
Right.
“Yeah, I know,” but, damn, I see the people. I see what they’re doing.
Yeah, you’re the VP of Engineering, so you know.
It’s like—
You actually know. Yes.
Yeah.
Yeah, I just wanted to get that. Obviously, we—
I hope it’s providing some—
This episode is not about DocuSign, but we have to ask.
Oh, yeah.
No, of course. Of course. Totally legit.
Yeah, let’s talk about Superhuman. So you joined in January 2025.
Yes.
Just give people a lay of the land of Superhuman AI. I think a lot of people who are listening are familiar with the email client.
Yeah.
I think the AI stuff is generally new.
Yes.
So maybe you can give the canonical definition of what you want to do with AI in Superhuman, and then we’ll dig into it.
4. AI Accelerates Email
The main driver is how you can put AI in the product to accelerate people’s productivity. It’s not just to do AI things with sparkles and everything. We don’t care that much about that. Our people have pretty high expectations, and they don’t want to slow down, so you cannot add latency.
Everything that we do is done in a way to improve people’s productivity, including AI. The first thing that we started to do was auto-label emails.
Mm-hmm.
Is it a pitch? Is it marketing? It’s a typical classification that you could do. People can say, “Okay, everything that is a pitch, I will look at that on a Friday.” During my typical days, I don’t look at it. That was one of the first things.
Summaries are another example. You have a long thread, and you want to know what the thread is about that someone shared with you. You have a quick summary. Nothing that was groundbreaking, but it was well thought out—just adding things that make sense at the right time.
Mm-hmm.
Another example is that we now automatically detect whether one of your emails requires an answer. If there’s no answer after 2 days, it pops up: “Hey, you need to send another email to the person because you didn’t have an answer.”
That was the first step. The second step was, “You know what? The draft is already ready. You can just hit Send.” It’s very subtle, but it’s like adding, “Oh, damn, shoot, yes, I wanted to remind people to give me an answer,” and the draft is already there. Pretty cool. Sent.
Now it’s detecting, “Oh, this is a request for your availability.” Or you have an executive admin who is doing that for you. Your draft is like, “Hey, let me CC the right person,” and boom—it’s ready, it’s sent, it’s done.
And then there’s the typical chatbot, because more and more of the use cases we see involve people using AI inside Superhuman to query their emails. A good example is that, as tech people, we receive a bunch of Substacks, a bunch of newsletters. Some are great, and sometimes the content is just mediocre.
I probably have 30 or 40 subscriptions, because everyone has something interesting to say at some point. Now I don’t read them. I auto-archive those, and every Friday I just ask Ask AI, which is the name of the feature.
Mm-hmm.
I ask my email, “Tell me about the summary of all the Substacks that I received this week. What should I pay attention to?” Then I can deep-dive into the places where I want to pay attention.
This is always thought through in a way to accelerate the pace and try not to be in your way. Hopefully. Feel free to ping me if that’s not the case.
Yeah, I would say I don’t know if this is a recent change, but I feel like I’ve started using Ask AI a lot more. I’ve been a Superhuman user for many years, and you’ve had it for a while, but somehow this year it kicked up a notch. I don’t know if anything changed in the product, because I wasn’t using it before, or if it’s just me trying it again.
That’s a good question.
Yeah.
I think people are more and more used to the muscle of querying things because of ChatGPT.
Yeah, yeah. So the general consumer behavior is—
Yes, exactly.
Yeah.
I mean, now every single product has a chatbot where you can ask questions, so it’s becoming more and more natural to ask questions compared to managing a to-do list of emails.
And agentic search as well. Previously, I was like, “Oh, you have to embed my documents, and then it’s just going to retrieve.” That’s not what I want. But agentic search, where you can actually figure out what I mean by my question when it’s half-formed, expand it, and then answer it—it’s actually really good.
Yeah, and we spend a lot of time on the quality of the answers.
And this is a framework. It’s not LangChain, right? It’s your own framework.
Yes. I mean, we’ve done a lot of iteration.
Yeah.
There are a lot of subtleties and multiple pieces there, and multiple different models—
Mm-hmm.
—based on what they’re really good at. But where we’ve spent quite some time lately is around quality and making sure that we are generally good across different dimensions, especially for typical queries, and optimizing for them.
Yeah.
One thing we try to solve for is agent laziness. Through this chatbot, one of my use cases is that I receive a Slack message saying, “Hey, Loïc, can you review this document, please?” It’s a tech strategy document that I need to review. I take the link, go to Ask AI, and paste it in, saying, “Hey, find me 15 minutes tomorrow. I need to review this document.”
Typically, I don’t need the agent to say, “I found this slot and this slot and this slot. Which one do you prefer?” I just asked for 15 minutes. “Find it, do it.”
When I had an admin, and I asked her on Slack, “Find me 15 minutes,” she didn’t ask me whether I needed it in the morning or the afternoon. She just did it. We’re working on this agent laziness, because the handoff to the user was losing time.
Working on making things happen faster—we spend a lot of time on this. That’s why you might have felt that the overall quality is better.
Yeah. My old joke was that, because the way you trigger it is that you actually type it in the search bar, when I was trying to do a normal search, it would sometimes accidentally trigger Ask AI. My joke was that most of my AI usage was just accidental because I actually wanted to search. But then I started using it more, and the kinds of questions you ask change.
Yeah.
Yeah, I use it to find people’s phone numbers—
Stuff like that. Just like, “Hey, what’s…”
I use it to find my contracts because I have so many contracts, right? From all my sponsors and venue things.
Yeah.
Yeah.
Yeah, 1 of the use cases that blew my mind: I was at a conference, and they shared a PowerPoint link with me 6 months ago. I couldn’t find the deck because I wanted to reuse some of the content. I couldn’t find it for whatever reason, so I asked Ask AI, “I’m pretty sure they shared a PowerPoint link or something like this. Can you find it?” It surfaced the content there in the link. I probably saved 30 minutes of searching through my emails, so it was pretty cool.
It’s to you, so it’s—
Yes.
Because there’s no way you can fit all your email into a context window.
Yeah.
Right?
No.
Anything else that’s more complicated—
So, we had to do some pagination because, let’s say I’m doing that: “I’m pretty sure I attended a conference where they shared a link with me.” In my case, I don’t do plenty of conferences, but someone like Rahul, my CEO, is basically attending a conference every 3 weeks or something. I’m not kidding.
That is his job.
That is his job. No, I mean—
And he’s fantastic at it.
And damn, I’m learning so much from him. But clearly, depending on the use case, you have more than 30, even hundreds of emails that can be semantically close to your answer. So you need to go through that.
We had to implement a paginated search: semantic search for the first 40, then deep search—“Not that one. Okay, next 40, next 40.” I’m using this agentic loop: while you haven’t found the answer, continue, even extending the semantic-search proximity until you find the right one, because it might be buried on page 2 of the search results, technically.
How did you design the tools to get to the agent? Just give people an overview of the framework and what it looks like. How are you structuring these interactions? Is there just 1 Superhuman agent that does everything, or do you have separate ones?
5. Small Tools Build Agents
No. We have separate tools, clearly. Even the agent—I would call it a set of tools. There’s a bunch of tools: a tool to detect your availability, a tool to understand who the people you interact with are, and a tool to write an email.
Every single action is very tool-specific, so it’s not 1 magic, big tool that can do pretty much everything. It’s a set of small tools that are used within the agentic framework. There’s a first step that’s, “Hey, what is the best tool to do this?”—kind of building a plan. For each step, what is the tool, and then making the calls.
Yeah.
I think now the tools-versus-skills distinction that Anthropic talked about is the hardest thing: how much you want to put in each tool, and then there’s the MCP discussion. I’m curious how you evaluate the tools, too. When you build them, how do you think about how to name them and how to write the description? How much work have you had to do to nail it?
I don’t think we spend that much time on it. Again, I’ll defer to my 3 engineers working on it, which is interesting. We can talk about the amount of people you need to work on those stacks when you want to be serious. I have fantastic people, so I feel blessed.
Most of the time was spent trying the different agentic frameworks and trying to understand the different models—the ones that solve which types of problems—because every single model is good for something. Sonnet was really great for agent handoff. The laziness was really great. The OpenAI version of it was not that good. Now we have Gemini coming into the room—last week, poof. Okay, that one is cool as well.
So I think everyone has built a way to switch easily from 1 model to the other.
Routers, so model routers.
Everyone has an LLM proxy to some extent, and an agent proxy to implement different stuff. That’s becoming interesting because the way to tweak and tune them is different. It’s still easy to switch from 1 agentic framework to the other, but at some point, I think it will be harder and harder, and the stickiness of them will be tricky. But to answer your question, we didn’t spend that much time on the tools themselves, I believe.
How do you think about evals? Are you evaluating 1 email draft at a time, or are you evaluating a longer workflow? Run us through it: when you’re testing Gemini, how do you decide what it’s good at and what it’s not good at? What’s the eval structure?
At first, we had a relatively naive approach: query, answer, query, answer, with a set of queries. Over time, we evolved into thinking more about the different dimensions that we wanted to target. Agent handoff is a very typical type of problem space that you want to make sure you select the right model for.
Typically, we get a bunch of queries targeting hard handoffs that we’ve identified by brute force or whatever, trying to target a set of what we call canonical queries along that dimension—the specific problem space of agent handoff. But there’s more. There’s deep search: a shit ton of emails, and you want to find that needle in the haystack. That’s a different type of category, so you need to have canonical queries targeting that type of dimension.
Every user will have their own way to question their own data set, and we cannot replicate every single data set of every person. The good thing is that we have a bunch of users, like Rahul or myself, who receive a shit ton of emails. Pardon my French, by the way. I don’t know if that’s okay for the show.
No, you’re good.
But he receives probably 500 to 1,000 emails a day.
He’s still part of the onboarding. It’s like, “I will send an email to Rahul, and he will reply.” I’m sure it’s not actually him.
Sometimes it’s him.
Yeah.
He’s reading pretty much everything. I don’t know how he’s doing it, but he is really paying attention, especially to the tone and why something is going sideways. He really associates the brand and tone of the people talking to the company with himself, which is kind of bringing us to the next level as well.
Yeah.
Thinking about all those dimensions is really key. Even if you have an eval tool, the way you structure your different queries to target those dimensions is important. Then we have those specific queries, the route queries, typically.
The one we joke about—and one of the first that we used to calibrate our quality—was a weird story. About 5 years ago, he did some refurbishing in his house, and he had this table made of a specific type of wood. He was discussing it with the contractor, and he wanted Ask AI to find that email and the type of wood that was discussed in the thread with that guy 5 years ago.
Until we nailed that query, he was not satisfied with the deep-search approach. This is where we were like, “Holy damn. Okay, so that’s a different set.” But we’re also talking about dates. Another dimension is dates: what is last quarter compared to today?
Large language models are not really good with dates, so how do you manage that? We have specific queries for that. So we were like, “Oh, okay, there are dimensions that we need to care about.”
Now we structure all the evals end to end: what is the query, and whatever happens there, there’s an answer. Was there a good agent handoff? Were the dates nailed or not? And so on. It’s pretty intensive in terms of brainpower put into quality. Again, because Superhuman is a high-perceived-quality type of product, we had to invest that amount of time there.
Yeah, high real quality. It’s not just perceived.
No, but I think this is important, because what is quality?
I don’t know. Yeah.
The feeling.
If I buy a car that’s a Toyota, it’s good quality, and I get the quality for my bucks. If I buy an Audi or Porsche, I expect a different grade. Maybe it’s grade. The grade is different, and it’s high grade, but there are high expectations, so there’s a high amount of time—
Yeah.
—I spend on quality.
Yeah. In product management, there’s this concept of the high-expectations user.
And Rahul was 1 example of those. I was just wondering: who are the most outlier, extreme people? How are they using AI in their email? Just in general, what are the most extreme examples that you’ve come across, obviously, because that’s how you work?
Oh, that's a good question.
For example, you had, “How much time do I spend in Waymo last month?” right?
Yeah.
Which basically turns your email into an accounting system, because it's a source of truth. I don't know if I would do that in Superhuman. Is it reliable?
It is reliable.
Wow.
When you think about the amount of work, we're working right now with Anthropic to basically build, on the fly, a small component of Lambdas that will build the code to do the aggregation. This is an easy example.
So it's like a code execution thing.
Yes, it's a code execution piece. But this one is relatively simple because you just have to have the agent extract from the email: select the emails from Waymo, from the Waymo receipt, extract the time—the duration of the trip—and then do the aggregation. But that's not easy. That aggregation is not easy.
Yeah.
LLMs are not good at math.
Yes.
There was some support for it, and right now we're discussing extending this approach to more. It's interesting—
Are you operating on the email file itself, or is there a fundamental—Is it like a row in a database and you're just writing a SQL query?
When we ingest the data—
Just the email address, yeah.
So we ingest the data.
Yeah, okay.
6. The Email Data Stack
We ingest the data. We rely on Gmail and Outlook, of course, because they're doing some great stuff that we don't want to do, like spam detection—
And Superhuman will never do it.
And probably.
Probably never do it.
Probably.
Which is being an IMAP server or—
Exactly.
Yeah.
Do I want to do that? Probably not. Maybe—
You know, HEY Mail did it.
Yeah. Other than that, is it something we want to spend time on? Is it really valuable for our end users? I'm not sure. They live in a closed system. They will live in a different company.
Yeah. Outlook.
Yeah.
Yeah, yeah.
They have Outlook and Gmail; it's already there. If we can just plug in and make that better, I mean, it's good there.
I mean, in some cases, Superhuman was the original wrapper company. If people think about GPT wrappers, this is the Gmail wrapper—the Gmail wrapper. At first it was a LinkedIn wrapper, and now it's a Gmail wrapper.
Yeah. I think more of it than Gmail itself, so. It's very true. It's very true. That said, you can question what an SMTP server really is. It's—
It's a server that conforms to a spec with some database. Maybe not even.
Maybe not even.
Yeah.
Maybe not even. I mean, they're doing way more stuff. Especially Gmail—the search capability is, of course, crazy good and all of that, but—
Yeah.
To do what you do, you need a server-side clone of my Gmail, and then you also need a local cache.
Yeah.
We need a local cache. We work offline. That was one of the things we did initially, besides the UX and the speed. We have everything local. One of the reasons is that we want to be fast, and every interaction should be under 100 milliseconds.
Yeah.
With the network, you cannot—you just can't. So everything needs to be local. Yes, we have a copy of emails locally on the device, and it works in the enterprise world because—
SQLite?
Interestingly, for mobile, it used to be Realm.
Yeah, Realm DB. Yeah, yeah. Is it a Facebook tech?
MongoDB.
MongoDB.
It was acquired by MongoDB. But it's now somewhat sunsetted, so we need to find a different way to do things.
Oh.
It might be SQLite. But on-device, everything is stored locally. That was the old search, where we basically had a database with rows of emails. For everything that is AI, we have all the vector embeddings and all of that. So we have a hybrid search, and we use—I don't know if we can name brands, but we use Turbopuffer on the backend to store 5 years of history.
Yeah.
It's stable infrastructure. They do things pretty well. It's fast.
Yeah.
So—
I think Turbopuffer is relatively public with their customer list, so I don't know.
No. Yeah. No, I think they are—
We'll let the PR department figure it out.
They talked about it, anyway, but it's stable infrastructure. They do things pretty well. It's fast.
I'll briefly comment that I know any number of local-first database companies that would love to work with you. If you're saying that you're on the market for a Realm replacement, they will come and talk to you.
I'm more than happy. My AI and mobile teams are really looking for something—
They will love nothing more than to be Superhuman's database. Okay, I want to just focus on the AI side, right?
Sure.
People want to know where their inference is running, what you're sending over, what the provider can see—that kind of stuff.
It depends.
Yeah.
It depends on the use case and on the type of model we want to use. There's some stuff we run with inference companies and open models. There's some stuff that we run with OpenAI and Anthropic. So it's pretty diverse.
Mm-hmm.
It changes based on the quality of the models. We're a GCP shop.
So lots of credits for Gemini?
Yes. We have an incentive to probably spend some dollars there.
Yeah, I mean, it's nice that they're also a leading model anyway, so you're not actually compromising—
And they are doing some pretty good stuff there.
Yeah.
But we use Baseten to run, I would say, some Llama and some BERT models for classification. We're probably having some discovery discussions with some YC companies about models on-device as well, because—
Yes, they work offline.
Yes. Interestingly, those companies started doing on-device mostly for cost reduction. That was their pitch: “We'll reduce your cost.” We don't care that much. Our users want quality, and they're okay to pay for that quality. But we want to solve for offline. If you're offline, semantic search doesn't work as well. So we're discussing with the companies—
What are your design constraints for offline inference? For example, DeepSeek-V3.1 would be like 600 billion parameters. I don't think you want to take up 600 gigs.
People are somewhat complaining about our footprint—
Yeah.
It's both in memory and on the device because we store local emails. When you install Superhuman, we download the last 30 days of emails so that we can search when you're offline, at least for the last 30 days. But we keep that history. It starts at 30 days, and if you've been a customer for 2 years, technically we optimize for 2 years of email on your device. So that's interesting.
On the local model, any thoughts on whether every app is going to have its own model versus having a device model that people run?
Oof. It's a lot of space. What would you prefer?
I'm curious. Would you rather have the user take care of the inference and rely on that, or do you want to own the whole experience?
Superhuman will want to own the full experience. We're pretty picky in the way things are happening. But at the same time, if we talk about mobile, you want the mobile experience to feel like your device. We're basically not doing React Native. We're doing Swift and Kotlin because we want the app to feel like the user experience in general on iOS or Android. But for the models, that's a good question. I would love the device provider to be better.
Right.
I mean, we can question local devices—
Right.
iOS has done some work there, but it has been underwhelming so far. They're still working at it, and that's why we have YC companies that are spending time there and doing some cool stuff.
Yeah. Amazing. Interesting question about Baseten. They're a very different cloud inference provider for open models compared to, let's say, Fireworks AI and Together AI. The general pitch is that they don't charge by token; they charge by box, effectively. Anything else that's interesting about working with them versus the other inference providers that you buy?
They're easy to work with.
Yeah.
I mean, that's—when you're a startup, you want to move fast. They're really easy to work with. They know what they're doing.
So the priority is what? Cost? Speed?
For us, it's quality.
Yeah.
So it's quality and speed.
It's all open models. It's all the same quality.
We would always start with the highest and most expensive model to get the right quality.
Yeah.
When the quality is nailed, then we can spend time trying to optimize.
Right. But all these providers—Baseten, Fireworks AI, Together AI—they all have access to the same models.
Yeah.
So unless they quantize heavily—which all of them say they don't—
In that case, the fact that it's a box means you control your costs much better.
Yeah. So it's fixed capacity.
It's fixed capacity. So when I discuss with my CFO, when it's token-based, the exercise is much more involved, trying to understand how we set adoption and all of that.
But that's serverless—usage-based serverless. It scales up, scales down.
Sure.
Fair, but the cost control is becoming a thing.
Yeah.
It was a thing before the acquisition. Now that we're part of a bigger umbrella, understanding your cost structure and being able to make projections that are closer to reality is more important. Like all pre-IPO-ish companies, you want to really understand where you'll be in 3 months, 6 months from a cost standpoint. So Baseten for that is pretty cool because you have more latitude to stay within the bracket of a box, basically.
I was thinking about this. A lot of people think about cost in terms of dollars per 1 million tokens. Right? I think that that is actually amateur thinking. That's only the kind of pricing you care about if you're a solo developer. But once you're at such a large scale like you guys, and something I learned at Cognition, you should actually care about price per 1 trillion tokens, because we spend multiple trillions per month. When you unlock that scale, you unlock different ways to spend that aren't serverless, token-based pricing. So basically, I think Baseten makes a lot of sense on a price-per-trillion basis.
Yeah. I didn't look at it that way, so it's pretty interesting. But no, that's fair. I mean, we've built so many different models trying to understand the cost per 1 million tokens, and then you have to infer what the average number of tokens is, because we treat every single email. There are really short emails and very long emails. You have to understand your data, what the median is and all of that, to make your projection, and there's always some magic.
The reality is, you don't have the time to—I mean, I'm an advocate of, let's move fast. If it's successful, it's great, even if it's expensive. Rather than trying to optimize the cost too early, just go with something that you control and that's fast, and you'll have time. I mean, it's a good problem to have. Success is a good problem.
Yeah.
When do you think it's going to break from a cost perspective? Say you were to draft every single email that I get. I'm sure you would lose money on the $40 a month.
Yes and no. I think it's a matter of how much more productive we make you. We have some customers who told us that initially, when we were talking about the different models and everything, they said, "Take the better model. I'm ready to pay $200 a month, but get me the best model."
Yeah.
They said, "I don't want half-crap because it's—
Right.
—less expensive. So always give me the best."
And these are all high-value CEOs and VCs.
Right. Yeah.
One hour of their time is worth—
Way more.
Ten times the amount of the subscription.
So why isn't there a $200-a-month subscription?
That's a good question.
Yeah.
I'm not in charge of the pricing and packaging.
But wait. An example would be: what's one thing that you would like to do that you cannot do with today's models, even though you try pushing quality? Your customers are telling you, "Actually, we really want this," or maybe Rahul's telling you that he really wants this.
I don't know.
Yeah.
I don't know. I think we have the means. We have the means to do pretty much everything that we want to do. It's a matter of executing and doing it right.
The way I'll put it is, if you can articulate what you cannot do today that you think you should be able to do, and your customers would pay you for it, the models will make it happen. But the problem that you have, and the problem that I have with Cognition, is we cannot articulate what it is. We will know if it's better, but only once it exists.
7. Email Enters the Voice Era
No, that's a good framing. The other piece that I think is pretty tricky is that there's a transformation happening in the user experience. Even the way we're thinking about the user interface right now is totally switching. The way we think about emails right now is still some sort of to-do list. It's a table, to some extent, with rows. What would it be like in a year? People will be interacting more and more with their systems through a conversational aspect.
Mm.
I see my kids. My kids don't type on their phones; they talk. I have 3 kids, and all of them talk on their phones.
What age?
One is working, one is in college, and one is in middle school.
Okay. On WhatsApp?
WhatsApp, because they're European and they need to talk with the family. The reality is Snapchat, TikTok, Instagram—whatever. They communicate over Instagram. I'm like, "That's not an image tool," or something.
I feel like a boomer.
Yeah, I am. I definitely am.
What's interesting is that—and we can debate this—we, as a human species, started to write because we didn't have enough storage for the stories that we were telling each other, so we had to write to store those stories. Now all the content can be stored on YouTube, TikTok, or whatever. What's even the need to write? What's the need? Everything can be vocal.
I see kids now: everything is vocal. They don't read articles. They want a TikTok video talking about the article. Coming back—and I'm sorry, I'm getting very high here—but being a bit more grounded, what does that mean for the future of the user experience for email and communication? Will people still type, or will they just talk to emails and want to hear an email?
This is where it becomes interesting, because Rahul, as a CEO, maybe next year he doesn't want to write to you with the new feature. Maybe he wants to talk to you. The way you will receive our marketing campaign about the new features will be, well, he'll discuss it with you or talk to you with his voice. It won't just be voice and tone in terms of writing; you'll really be in your car commuting, listening to Rahul talking about that.
Coming back to what cannot be done right now, I think the main problem is nailing the new user experience. With OpenAI, now you can do stuff with emails. They're trying to do some stuff there. All those chatbots try to be basically the new OS, to some extent. How do you interact with those new apps? What is an app even in this new world? That's what's really interesting, and that's why I'm glad to work with Rahul, because the guy is so freaking visionary. If there's one company to nail it, there aren't a lot, and I believe Superhuman is one of them.
Yeah. I think the inbox is the ultimate private data source. I feel like even when I see all these companies that are like, "Talk to your AI clone to get advice," or things like that, so many times, man, I'm just writing the same thing over and over.
How many founders email me asking for help with XYZ task?
Yep.
And the answer is almost always the same. There should almost be a way for Superhuman to be the advisor on my behalf. You should be able to predict what I will respond to this email.
It's called auto-draft for response. We're still testing it internally. Especially—sorry to cut you off—but the same is true for me. How many companies are reaching out to pitch me on AI frameworks, AI tooling, or whatever?
Right.
My answer is—although I don't answer because I receive hundreds of them—honestly, “Thank you, don't have the time,” and everything. That's what's cool. Because I want to be polite, it's automatically generated for me because they learn that I usually don't care. That's my answer.
Or if someone is pitching me on working with us, or applying to work with us, my answer is usually, “Please reach out to HR. I'm CCing HR,” and everything.
Yeah.
So now we're almost able to understand how you typically reply. But if it's covering only 80% of your use cases and you need to discard 20%, where's the cost-benefit value? Is it annoying to have 20% where you're like, “Ah, discard. I want to write it myself”? Is it good? What's the limit—90/10, 80/20?
I think it's AI plus the snippets that you have. I have snippets for a bunch of things, like vendors. I have this super-long snippet: “Thank you so much for reaching out about your company.”
Yeah.
“Sounds like a great product.” It goes on, and then the response is, “Thank you so much for your thoughtful response.” And I'm like, “Great. Get it out of the way.” But I feel like if you could use that—
Yes.
—plus AI to do the small—
Yes.
—kind of last-mile—
Yeah.
—thing, I think that would be enough.
Yeah.
You don't really need AGI. I'm excited for it.
Q1, Q1, Q2, something like this.
I mean, dude, I pay $200 a month to OpenAI and Anthropic. I'll give you $200 a month if you make me not write the same thing over and over. Deal.
I think, more generally, what he's trying to get at—and what Superhuman is starting from a very good basis for, but is not there yet—is kind of like an AI EA. I don't know if this comes up a lot. Well, I have people I work with who read my emails and respond for me. They have memory, and they know my normal preferences. They have human judgment, which LLMs don't have. Is that something you would want to build, or do you think you want to leave that to others?
8. Superhuman Builds an AI Assistant
That's the goal. When we really kick off the revamp of our AI world and what AI means for Superhuman, Rahul did a pretty good pitch on it. There was a pretty nice video, I think it was in March, for the launch of the new AI.
That's the vision: you have an EA. Most of the people using Superhuman are C-suite people, founders, and all of that, so pretty fast, they need someone to help them with their emails. We want to do most of that job. We're getting there. But that's the goal.
The first thing is answering with your availability. Right now we can do it; right now it's in beta. In my emails, when someone internally asks, “Hello, can we meet next week for lunch?” I automatically have 3 slots proposed in a draft, and I can just send the draft prepared for me.
Yeah.
It's still up to you to decide whether or not you want to send the draft.
That's the thing. I don't want to be involved.
And this is where your EA will always be better than an LLM—
Yeah.
—because she knows the types of people you're okay having lunch with. Or maybe they have the context because—
Sometimes you're busy, but you're like, “Oh, VIP, I will move this.” You know what I mean? I need to—
Exactly.
And your calendar isn't going to know.
We're getting closer because we know how much time you interacted with that person. But how much time you interacted with them doesn't mean that maybe last week you had a bad discussion with them, and now you're not friends anymore for whatever reason. Your EA would know.
There will always be limitations to this. That's why we want 2 people to always be in the loop. Maybe it's your EA that's in the loop.
Yeah. It's so helpful when I'm not in the loop. We can batch it, and I have my once-a-day call with the EA. But obviously that will happen. You know what? Some ways that other people are pursuing this—Notion's trying to go after it, right?
Yeah.
And they have Notion Mail, Notion Calendar, and obviously they really care about AI. Some other people are doing this interesting thing where they buy an EA company, a company that already does virtual assistance, and then monitor what they do.
First of all, Superhuman can provide me an EA who is a human, and then slowly replace parts of it with AI.
Mm.
I'm curious what you think about that. That's a more aggressive approach if you really want to—
I mean, that's probably the best way to understand how an EA works and the type of work that they're doing.
Yeah. Make your own data.
Yeah. I mean, that's intense. That's intense, but sure.
You have the money.
And you pretty fast understand what types of workflows you want to automate first. So having that data would be obviously pretty interesting.
Yeah.
One of his portfolio companies bought a sort of—
Law firm.
Yeah, yeah. Do you think that's an accurate description, or am I glorifying it too much?
Yeah.
No, it's an accurate description. It just behaves as a law firm, though.
Right. Just treat it as a law firm, and then internally start to optimize.
You now have so many customers that you might need a lot of EAs to do it for everybody. But I'm curious: I think the memory is kind of the killer feature of the EA. It's understanding in real time.
I'm curious, now that you're within Superhuman—the company now called Superhuman Mail—do you feel like there are a lot of advantages to being email plus documents, plus being embedded in everything? Do you feel like that helps close some of these gaps?
9. Context Escapes the Inbox
Yeah. For example, Coda is an interesting piece of software. Coda is like a Notion equivalent.
Yeah, we used it at Amazon.
Yep.
It's a pretty good one, and a lot of enterprise companies are starting to use Coda more and more because of its flexibility. Coda has this concept of Coda Packs, which are integrations—glorified integrations, if I can say it that way—but they're ingesting the data. The data is there.
We technically have an ingestion pipeline that can aggregate all the knowledge about you in the company, which is great. And now, if you add Grammarly, Grammarly is ubiquitous. Grammarly knows that you're in Google Docs. Grammarly knows that you're crafting a post on LinkedIn. Technically, they can know. It doesn't mean that they use the data—
Right. Yeah, of course. Yeah.
—but they're everywhere. When you're getting into your email, I know that you're currently on Jira with that context. All of a sudden, I can pop up some of the context. I know that you're writing to that person—oh, it's about this. I can expand and augment your email because I know where you're coming from.
The data will be there through Coda. Grammarly knows basically where you're switching from Google Docs to Salesforce to LinkedIn, and now you're writing an email. So we have this augmented context, much more precise compared to something like ChatGPT, for example. They don't know where you are because you're switching windows. You're coming from Salesforce to ChatGPT; they don't know where you were. They wait for you to pass the content to get the context.
If you're Grammarly, I know where you're coming from. So when everything has converged—we've only been acquired 3 months ago.
Yeah.
But when everything has converged, from a contextualization standpoint and a knowledge standpoint, we know way more. So we'll be way more accurate in the way we help you.
Maybe predicting a fourth acquisition, but wouldn't it make sense to have your own browser?
That's a good question. I think there's much more to be done in the productivity space before, I would say, solving a browser, and everyone is trying to do a browser.
Yeah. Atlassian, Perplexity, OpenAI.
I'm still sad that Arc is no longer being developed because of Dia, but Dia has been stopped.
They're rebuilding Arc in Dia.
Yeah. But it feels very unstable now. More and more people are basically saying, “Okay, let's go back to Firefox.”
Well—
More and more people are doing that because there are so many browsers. You want to wait for the war to be done and have the clear winner.
No. I disagree. You should—
I'm in Atlas.
You should go all in. What are you using?
I use Atlas. Yeah.
Atlas, yeah. I'm also on Atlas now.
Oh.
Yeah.
Interesting. I'm still on Arc because I'm—
It still doesn't have profiles.
Yeah.
That's the biggest issue.
So, based on the different emails or logins I have, I switch between Atlas, Chrome, and Arc.
Yeah. Interesting.
Yeah.
My personal one is on Chrome.
But I'm just saying, if that context matters to you, right? With Coda and all those things and Grammarly, all this, you might as well have your browser.
No one—
I—
No one will get upset at you for saying, “Oh, we have a browser.” It'll be like, “Yeah, it makes sense.”
Or it'll be like, “Oh no, one more?”
But it's the Superhuman one, and that's a good brand.
That's interesting. I foresee the browser disappearing completely. I'm like—
Hmm.
Ooh. Okay, that's the title.
My main, central piece of software that I use in my productivity tool is Raycast.
Ah.
Yeah.
I'm a Mac user, so I use Raycast.
Raycast.
For the people who don't know Raycast, it's basically a way better Spotlight on Mac. I don't need bookmarks in my browser anymore. What is a browser doing besides providing you a view on a website? Nothing. So even to some extent, Raycast should just be a web view, because what I do with Raycast is—
Then you're turning Raycast into a browser.
Is that a browser if it's just rendering HTML?
Yeah.
Okay, so—
Right.
If—
Everything is a browser.
So, yeah, if it's only rendering HTML—
What else do you want? Do you want JavaScript? Do you want—
I don't know.
Local storage? You want what?
Yeah. Local storage is one.
Extensions.
You need a browser to have your local extension.
Hmm.
But to have your local storage, that is pretty massive, like Superhuman. But what's left? Everything that was making a browser a browser before, which was bookmarks, basically the history that you had, maybe cookies and everything—what's left? If you get rid of that, it's just a view, a web view to some extent.
Yeah. It's a clean application platform with an open app store. You know, there's a Marc Andreessen line: “Well, the operating system is just a poorly debugged set of device drivers for the browser.” He said, “The browser is the actual application interface.”
Oh.
From the person who made the browser.
Yeah.
Makes sense.
Of course. Yeah.
Yeah. I think the browser will become thinner and thinner. I believe it will become thinner and thinner, but it will disappear, or it will just be—
Yeah.
Embedded in the OS eventually.
Yeah.
So.
One more technical thing, and then we can go to organizational things. You mentioned understanding the person. One part of memory is just the knowledge graph, and one part of the knowledge graph that really matters is the entities that I deal with, right? I've dealt with him for 4 years, and we have that context, and basically what's possible today in Superhuman and maybe what is possible in the future, right?
Mm-hmm.
Do you, for example, use a graph database or something like that?
Not yet, and it's interesting because you're mentioning what's missing right now. I think that this knowledge-graph-oriented database—I'm not there yet, to some extent.
But have you actually tried it, or are you just saying that?
No, we didn't try.
Yeah, that's the thing.
Not right now.
It's not fair to say they're not there yet if you haven't tried.
Correct.
Yeah.
Correct.
Even from a taxonomy standpoint, when you think about those entities—
Yeah.
What are those?
Yeah.
If you are verticalized, say—
People, companies.
Yes. But then you start talking about projects, but is it a project? Is it a task? Is it an initiative? Is there a hierarchical aspect to those? How deep is the tree?
Yeah, yeah. These are all valid questions.
I think it's very—
And Superhuman's history is Rapportive, where the person is the core of the—
Correct.
—the universe.
No, but there are some obvious entities.
Yeah.
But if you want things to be really personalized, these entities are very, very subjective. I'm a user of Obsidian.
Yeah.
I'm a note-taking nerd, and for the people who don't use Obsidian, it's—
Another local-first app. Yeah.
It's another local-first app in which you build your own workflows and where you basically, through templates, define your own entities that make sense for you. There are no 2 graphs that are similar, even if you're using the note app, say, for the same thing. So trying to infer a generic knowledge graph that can be reused with dedicated entities—people, tasks, projects, and everything—is harder than it seems.
Oh, yeah.
Interestingly, we were thinking about it when I was at Productboard. At Productboard, we have roadmaps for so many tools. Based on that, you can probably infer some taxonomy about what a SaaS product is, but even trying to generalize this into a tree that can be repeatable for people is hard.
There is some common stuff: authentication, authorization, billing, user management, dashboards, whatever. Every SaaS company has this. But when you enter the domain of the company, it's totally different because their features and their surface area are very different. So even there, trying to form the knowledge that you have and abstract the entities that will be the same for everyone—
Yeah.
—is not easy. So it means that for each user, you need to have an unoptimized graph that is subjective and dependent on the person. You need to build the graph based on just the data, and you don't have a real way to optimize for it. But you're fair. You're right. We didn't try. But also because—
Many people have failed. It's fine.
And I don't even foresee a path where that can be surfaced into more productivity gain. At the end of the day, what is the problem you're trying to solve? It's super nice from a technology standpoint and even from a thinking-process standpoint: What is the ultimate data model for a productivity nerd and all that? But what are you improving from an experience standpoint? Is it the accuracy of the draft that I'm preparing for you?
I want my AI EA to remember everything I've talked about, everything I've done, everything I've talked to everyone about, every conversation I've had.
Yeah, but then it's Jarvis, and it's almost AGI to some extent. So—
Yeah.
You have the context that no one else has.
Yeah, but then there's the amount of compute. You need to recompute your graph every time you receive new stuff and everything, so it's an interesting space.
I think—to your point, we probably won’t be the one solving for that as an endpoint solution. I think there are companies that should focus on this and—
Yeah.
…be like, “Hey, I’m the engine that will ingest everything that you’re doing. We build a graph, and the graph will be the best graph ever. For each account or tenant, we’ll build a graph for you.”
Yeah.
That would be great. But is it something for turbopuffer? Is it something for those vector database companies to solve for? Maybe. I don’t know.
Yeah. For what it’s worth, I’m actually dating someone who’s doing Upside, and they’re mining emails for CRM population and building a knowledge graph from emails.
Interesting.
Basically, they’re happy that you’re not doing it because—
I’d love to have an intro.
Because obviously, if you do it, then you’re a very serious competitor.
No, but I think it’s not easy.
Yeah.
So I would love to discuss—
Sure, sure, sure.
…but I think we would probably be more a consumer of the outcome rather than the builder of that layer.
Yeah. I think the other big consumer, obviously, would be OpenAI.
Of course.
They clearly want to eat everything inside ChatGPT.
I mean, this is a cool exit strategy for such a company.
For them, yeah.
Of course.
I mean, do you want to build a Superhuman app inside ChatGPT? I feel like the answer is no, right?
The answer is that ChatGPT, or OpenAI, and Superhuman are competitors. This is what we fight against, to some extent. We have a different approach, I think, but especially this ubiquitous Grammarly presence—we are everywhere and in everything. I think we want to be more proactive because we are where you work. We can be more proactive compared to ChatGPT, which is waiting for you to do things to help you do the thing.
Mm-hmm.
So there’s reactive versus proactive. I think we’re more on the proactive side. That’s the competition. I would say it’s for notes, but when Rahul is questioning the quality of our AI queries on Superhuman, he’s comparing us to Gemini and OpenAI. That’s the competition we’re fighting against.
Yeah. Speaking of Gemini, the chat app obviously has privileged access to all of Google. So they can also check emails—
Privileged access, and the search engine is crazy good. But—
Break them up.
Rahul, break them up.
All right. Yeah, yeah.
Awesome. On a broader side, you mentioned you only have 3 people working on AI. What’s the coding AI adoption at Superhuman on the engineering team?
10. AI Reshapes Engineering
Interestingly, our path was—we started to really think about it in Q1, with a bunch of people using some stuff and everything. We didn’t have any data, just anecdotal feedback and all of that. The first thing we did was cut the red tape.
Mm-hmm.
“Hey, folks, free for all. I’ll approve the budget in 1 hour. You can try anything you want, and deal with the security team—24-hour turnaround to get things approved from a security standpoint, because you don’t want to do some—
Right.
…crazy things.”
Huge. Q1 was everyone trying everything. It was really interesting to see how things were working super well on the front end, a bit less on the back end. We’re a Go shop on the back end. Everyone working on iOS and Swift was like, “Eh, not that good at the time.” But there was huge adoption in terms of tooling. Also, on the product side, a lot of v0, I would say—
For Next.js?
No, v0.
Or you just use it for anything.
We just use it for prototyping.
Ah.
It’s to be as close as possible because we have a founder who is very picky and wants to review the design. A design in Figma is great, but when you can click and do real stuff, it’s so much better, and Figma isn’t there just yet.
Oh, Figma has Figma Make. We interviewed Dylan.
Yeah. Sure.
It’s getting better, I would say.
It’s getting better.
Yeah, yeah, yeah.
But as PMs, they use v0 or tooling like this because it’s—
Not Lovable?
Superhuman is like v0. v0 is a standard. Again, it was free for all.
Yeah, yeah.
Try whatever you want and everything.
Free market, right?
So, free market. And free market—
v0 won.
…v0 won. Always winning, and it’s still a free market.
Q2 was more about, okay, let’s try to understand where this is working and where this is not working. So we compiled a huge list of wins and areas where, like, to do this—not good. To onboard a new code area, amazing. I used to spend a full day understanding all the entry points and dependencies in the code stack that I didn’t know. Now I need 30 minutes with Claude Code, and I understand how things are working.
Even for me, I’m not in the code anymore, but instead of asking my engineers, “How are we managing the refresh tokens with Gmail?” I just use Claude Code, and I’m using Warp.
Yeah.
I’m—
Warp?
Warp.
Yeah.
Warp is good. But anyway, with Warp and Claude Code, I can understand how this shit is working, and boom, boom, boom, boom, boom. I’m providing the links to the right files, explaining the high-level concept to you, and I don’t waste my engineers’ time just answering a question.
So that was Q2, and we started measuring. For every PR, we have to put a label: I used AI, or I didn’t use AI, and if I used AI, it was productive or it was not. So we’re trying to understand the lay of the land.
Roughly, I think we have 80% of people who are really flagging the PR. Out of that 80%, probably 80%–90%, I would say, is AI usage. It’s all declarative—we’re not plugging in any tool to measure the real number of tokens and everything. And out of those 90%, again, 90% had a positive impact. But it’s not always in the code. It might just be discovery, understanding the lay of the land, or stuff like this.
So 81%. So 90 times that.
So technically, yes, it’s 90—90 of 80. But by inference, if I caricature it, I would say 80% of usage is happy usage.
So roughly 80% of lines of code written in Superhuman. But probably more—
It’s not always lines of code, because—
Probably more than that. Yes.
Yeah, because of PRs.
Because of PRs.
What is the discovery? Most of the time you spend is not writing code; it’s trying to understand what you need to solve for, and this is the part that has been reduced.
In terms of real KPI—and AI is not the only reason why we’ve accelerated—in Q1, we were roughly at 4 PRs per engineer per week. That was in Q1. In Q2, we were closer to 5 PRs per engineer per week, and in Q3, we’re closer to 6. So the global throughput—and again, PRs per engineer per week, we can debate—
Right. Yeah, yeah, yeah.
That’s a throughput measure, and it increased quite a lot. But again, AI is only a piece of it: technical strategy, clarity of what you want to do, organization. There’s a lot associated with that. So we feel pretty good.
One question that a lot of the AI leadership people I talk to have is, “Am I supposed to ask more of my engineering team now? Am I supposed to hire fewer people? Should we ship more as a company?”
I think the thing about AI is that you can do a lot more, but most companies are not built to do a lot more. Especially if you ship 100 more features, you don’t really have marketing to market 100 more features. You don’t have support to learn 100 more features. How do you think about structuring teams and the expectations around it?
That’s interesting because Superhuman historically was very lean in terms of organization. Superhuman, like—we have 50—
50. Crazy.
…50 engineers.
And your user base is roughly—
Uh—
…how many million?
Yeah, less than that. Paying users, probably 100,000. Something like this.
So it's still relatively small.
You're still supporting a lot.
But, yeah. I would say it's a small, pretty senior team, and the average tenure is probably 4 years, so long tenure. We're fully remote as well, which is interesting. My AI team is distributed between Patagonia and Canada, so we have access to a different pool of the right people—not trying to compete in the Bay, because people want to go to Anthropic and OpenAI.
Mm-hmm.
And those guys—
They pay too much money.
Yeah. I mean, it's obviously not the same competition.
Yeah.
So we find people where they are, and people who don't want to move to the Bay and all of that. There are some great people there. We have relatively small teams, and we increase the capacity. We try not to move too fast because we're qualitative. There's kind of a vicious circle: “Oh, we can do more, let's do more.” But all of a sudden, the number of bugs coming in is also growing. So we try to be conscious.
Now we're working under Grammarly—the new Superhuman—so there's also an incentive to invest a bit more, because it's a product that is working. Shishir is really willing to implement a model called the compound startup. We're still a startup within Grammarly. We have our own P&L, and we still have Rahul as a founder. The only difference between now and before is that our board is Shishir and the exec team at Grammarly/Superhuman.
But we want more people. We obviously want Superhuman to have more reach and do a bit more. So now we're kind of scaling that, and we're adding more capacity. AI is helping, of course, but it's also helping with onboarding and a lot of that. We're adding some capacity.
Yeah. I think the mainstream pushback on it is: “Hey, you used to pay me X to do 4 PRs a week, so am I getting paid 50% more? Then I gotta ship 6 PRs a week.” That's why there's a lot of pushback around AI from people as well: “Hey, look, I'm using this and you're getting more out of it, but I'm not getting more out of it.”
I strongly disagree with that statement.
I would disagree too.
Yeah, I disagree too. But I'm saying that when you listen to people outside of our bubble, there's a lot of discussion around—
Yeah.
—you know, where the value is accruing.
So basically, if you only look at it as paying for output, was the previous payment wrong or was the current payment wrong? One of them is wrong.
Right.
That's an interesting point. The way I see it is that engineers are well paid. We are a very fortunate, I would say, part of the population. Our salaries are probably pretty good and part of the top 5% in the country, or even in the world. When we talk about Maslow's pyramid, engineers, at some point when they're pretty senior, don't rush for $10K or $20K more. If we talk about millions, sure, but that's the 1% of the 1%.
Mm-hmm.
For the rest of the population like us, I think the joy and the dopamine come from what you ship. Having the ability to ship more value and have more customers, and being happy with what you do—you end your day and feel like, “Damn, that was a good day.”
So I think the discussion is not about the money itself. It's like, “Oh, damn, I'm in an environment where I ship fast. I can have all the tools that I request within 24 hours. I can basically be the best version of myself, and I have fun in a good team.” You don't have a lot of attrition when you have an environment like this.
Sure, money—you need to pay people a fair amount. But if you're just fair, people tend to stay if you have the right environment. Helping them go from 4 PRs a week to 6, they're like, “Shoot, I'm so much better than at the beginning of the year. That's so cool.” And you don't have that everywhere.
Yeah. I'm with you. I'm curious to see more of the scores evolve. Awesome. Any parting thoughts?
Just generally, what's your take on AI in the software industry? You've been in this for 2 decades. Do you think people should still learn to code? Do you think the junior developer is screwed? Any opinions on those common topics?
Yes, of course. You need to learn to code. I see this as kind of the switch from assembly to C.
Yeah, it's a higher level.
It's just another level of abstraction. But at the end of the day, you still need to understand how a computer is working. You need to understand how memory is working, like swaps and all of these things happening on a server, how a server is working—serverless, in quotes. It's always a server belonging to someone. You need to understand the fundamentals to be good with AI.
I do believe that AI will do only one thing: separate good engineers from bad engineers faster. If you're a good engineer and you're using AI well, you will be an amazing engineer. If you're a poor, lazy engineer and you don't want to understand the things that you're doing, AI will make you even worse because you'll have the feeling that you get it, but you won't be looking behind the magic, behind the curtain, to understand how things work. So I think AI is a blessing for our job.
Awesome.
Great.
Any final calls to action—hiring people, things you want people to do, like trying the product and giving you feedback on?
Of course, try the product. Of course, complain to me if things aren't great. Yes, we're hiring. We're hiring product engineers—people that have a strong appetite for the user experience, because I do believe that in a world where the technical moat isn't much of a moat anymore, because startups can build something close to what you're building in two weeks, the difference is how you think about the user, the flow, and all of that. People that have this appetite for a nice interface, a beautiful product that people love—this is the type of engineers we want. Good engineers, that's a baseline, of course, but with this spike into the user experience. Even if you're a backend engineer, but you care about latency because it's having an impact on the end user and all of that, this is the type of engineers we're looking for. We don't care where you are. You can be in Patagonia, as I said, or up north in Canada. We try to limit things to the Americas, basically. We're looking for bright, gritty people that want to have fun. We're seriously fun.
Cool. Thanks for joining us, man. This was fun.
That was cool.
Thank you.
Thanks for having me.