[BidClub_]
Latent Space · · 78 分钟

Snipd:用于学习的 AI 播客应用——与 CEO Kevin Ben-Smith 对谈

swyxKevin Ben-Smith

YouTube
TL;DR
  • Snipd 通过证伪最初的产品假设找到了方向:用户喜欢类似 TikTok 的发现流,但他们仍会听完整节目,还疯狂创建 Snips。 4人组成的纯技术团队由此从社交切片转向学习系统,捕捉、总结并保存播客知识;在耳机上连按3次,就能把刚听到的内容转成笔记。这些行为证据为业务明确了任务:阻止听众在条件反射般点开下一集前遗忘“99%”的内容。

  • 如今的切入点更像是覆盖在音频之上的结构化知识层,已处理超过100万档播客。 节目可以获得文字稿、带说话人分离和姓名标注的音频、嘉宾简介、AI章节、书籍提取、作者关联节目、带时间戳来源的聊天,以及保留原文和音频的 Snips。动态广告会打乱固定时间戳,因此 Snipd 构建了模糊音频匹配系统——“基本上就是播客版 Shazam”——让用户实际收到的任何音频文件都能重新对齐。

  • Snipd 的模型策略同时考虑毛利率和延迟:对成本敏感的转录、说话人分离模型自行部署,其余高阶任务则根据最低所需智能水平,在 OpenAI、Gemini 和 Perplexity 之间路由。 Perplexity 搜索的报价约为每1,000次5美元,明显高于普通 LLM 调用;Kevin 欣赏 Claude 3.5 Sonnet 的措辞、个性,但由于成本较高,团队会在许多工作负载中选择其他模型。可预测的预处理任务持续运行,突发的用户请求则留在 API 上,Kevin 将其比作“AI 领域的 AWS”。

  • 产品的护城河在于生产可靠性,而不是演示的低门槛:节目聊天15分钟就能做出原型,但要做到“99%的时间”正常运行,需要提示词格式、前端渲染、启发式规则和无数正则表达式。 Snipd 在创业公司的“凭感觉评估”之外,还采用 LLM-as-judge:先用便宜模型生成5个候选结果,再让更强的模型挑选,用于名言、书籍和说话人识别。目前托管式闭源模型仍占优,因为迭代速度比不确定的开源模型节省更重要。

  • 消费级 AI 的机会在于隐入工作流,因为“你已经可以问聊天框了”并不等于打开应用后,个性化摘要已经按用户需求生成并呈现出来。 Kevin 不断回到“待完成的任务”:点击文字稿中的一个词就跳到对应音频,这种简单交互可能比模型的新奇程度更重要。他的终点判断是,“AI 是未来的电力”——最终 Snipd 应该只是一个默认具备智能的播客应用。

  • 语音是 Snipd 设想中的学习闭环:节目结束时,用一段2至3分钟的对话迫使听众选出一个收获,并把它连接到行动。 每月超过5亿播客听众已经具备收听习惯,因此 Snipd 不必像 Duolingo 那样靠猫头鹰、连续打卡和通知制造新习惯。Kevin 认为这可能让单集节目的价值提升“10倍”,但也明确表示具体交互仍需实验。

  • 扩张方向覆盖内容、发现、视频和创作者,但分发激励将决定 Snipd 能拥有整条技术栈中的多少部分。 Kevin 希望加入原生有声书、YouTube、AI 生成内容、可与用户对话的推荐系统,以及“可后台播放的视频”;这类视频即使约90%的消费伴随其他活动,其主要优势仍在发现。swyx 的反提案是创作者工具——坐下、录制、按一个按钮、完成——因为 YouTube 已经是最强的播客平台;迁移困难,以及托管在 Substack 的节目无法通过 Apple Watch 播放,暴露了机会与平台风险。

摘要 · 为研究而整理的核心内容

1. 需求,而不是黑客松奖杯,造就了 Snipd

  • Kevin 在苏黎世联邦理工学院学习数学和经济学,主攻量化金融。但回头看,最能说明他与这条路匹配度不高的信号是:“我从没在业余时间读过一篇关于这个领域的学术论文。”一位朋友发给他机器学习课程的讲义;一个周末后,他的反应是:“我靠,这就是了。我爱上它了。”

  • 他不断在银行里寻找应用机器学习的理由,最终意识到,要追逐这条路就必须“彻底转向”。Kevin 辞职,随后在一家苏黎世早期创业公司负责了5年 AI 团队,交付银行销售模型,并把晦涩的交易入账文本转成易读的商户描述。

  • 那位朋友后来与他一起参加 HackZurich,做了一个播客自然语言搜索工具。他们在2个半小时的 Joe Rogan 节目里搜索“抽大麻”,准确找到了 Elon Musk 的对应片段。获奖提供了“启动能量”,但更强的验证来自其他参赛者立刻追问:“我能用吗?”并说出了相邻的需求。

  • Snipd 依然异常精简:团队只有4名技术人员,2人负责后端和 AI,2人负责前端。这支团队同时维护 iOS、Android 和 Apple Watch 应用。产品采用免费增值模式,Kevin 还提供了一个可免费使用1个月高级版的链接。

2. TikTok 式假设失败,暴露了真正要完成的任务

  • Snipd 的首个版本约在3年半前上线,当时 ChatGPT 和 Whisper 尚未出现。最初的概念是一个社交化、类似 TikTok 的信息流:部分用户收听完整节目,把精彩片段剪成 Snips,其他人则通过这些片段发现播客,或把它们当成达到某种目的的工具。

  • 创始人原以为吸引用户观看 Snips 容易,说服用户创建 Snips 困难,因此激进地优化创建环节。现实恰好相反:用户喜欢通过片段发现播客,但仍想听长音频——“他们疯狂地创建 Snips”。于是,Snipd 将重点转向捕捉、留存和学习。

  • swyx 曾通过更早的“个人混音播客”亲身经历过这种痛点:记下时间戳和 URL,之后下载 MP3,手动剪出5至10分钟的片段,录制评论,再重新发布。Kevin 的重新定义是:尽管播客是“全球最大的知识来源之一”,播客应用却仍然只是“改装过的音乐播放器”。

3. 每一集节目都变成可导航的知识对象

  • swyx 的基础要求没有妥协:Snipd 仍必须是完整的播客播放器,本质上是 Overcast 或 Apple Podcasts 的超集。他起初不喜欢 Snipd 的主动加入式队列和下载机制,但后来发现,节目中可见的 Snip 数量及其位置,成了衡量听众兴趣的有效代理指标。

  • 对已启用功能的节目,Snipd 会转录音频、分离并标注说话人,为每位嘉宾生成简短传记和头像,创建带标题和描述的章节,并提供可点击的文字稿。选中文字即可跳转到对应时间戳;Kevin 常用这个简单交互说明,消费级价值并不等同于 AI 的复杂程度。

  • 书籍提取不止于识别书名。LLM 读取文字稿,Perplexity 及其他编排工具负责获取封面、作者和简介;随后 Snipd 会展示其他包含该作者的节目。早期模型过度选择 Sam Altman、Elon Musk 等被频繁提及的人物,因此团队收紧了产品标准,只保留真正作为嘉宾出现的人。

  • 听众可以对节目进行关键词搜索或聊天,询问某个主题何时出现,索要要点,并通过时间戳引用回到音频。Snips 会把总结后的洞见与文字稿、声音一起保存;用户也能通过 Watch 创建 Snips,但 Kevin 表示,托管在 Substack 的播客无法在那里播放——即使通过 Apple Podcasts 也不行,而 Substack 在被联系后似乎并不在意。

4. 动态广告迫使 Snipd 打造“播客版 Shazam”

  • Kevin 唯一披露的处理规模数字是超过100万档播客。一些节目会自动处理,另一些则在付费用户提出请求后进入管线。转录和说话人分离先生成语音区块,之后由 LLM 编排识别嘉宾,并把姓名分配给相应区块。

  • 动态广告让普通时间戳变得不可靠。每次播放请求都可能返回不同的 MP3,广告会根据 IP 地址等因素插入;如果 Snipd 转录的是另一个版本,插入点之后的每个逐词时间戳都会发生偏移,章节、文字稿导航、聊天引用和 Snips 全部失效。

  • 因此,Snipd 会在播放过程中将听众收到的音频与文字稿重新同步。匹配接近音频字节层,而不是依赖部分转录;它采用模糊匹配,不要求完全相等。Kevin 的简洁描述是:“我们基本上造了一个播客版 Shazam”,这是为了让主产品正常运转而做的副项目。

5. 技术栈只为每项任务购买所需的智能

  • 后端约90%使用部署在 Google Cloud Platform 上的 Python。移动端以 Flutter 和 Dart 共用一套代码,Flutter 无法覆盖的部分则使用原生开发,例如 Apple Watch。被问到 Flutter 是否是个好决定时,Kevin 给出的限定回答是:“到目前为止,是的。”团队没有离开 Flutter 的计划。

  • Snipd 起步时 GPT-3.5 Turbo 尚未出现,最初自行运行并微调开源模型。转录使用过 wav2vec 2.0;它结合 transformer、连续音频和自监督学习,让 Kevin 相信音频会沿着文本走过的路径发展——即使当时理想的产品能力还不存在。

  • 如今 Snipd 使用 Whisper,并继续自行部署部分成本敏感型转录和说话人分离模型;大多数下游工作则调用 OpenAI 和 Gemini API。需要联网搜索的任务交给 Perplexity;Kevin 报价约为每1,000次查询5美元,并表示 Google 搜索的 grounding 并没有便宜到足以替换已经运行良好的系统。

  • 统领团队的规则不是“选最好的模型”,而是先判断任务需要什么级别的智能,再在该级别购买性价比最优的模型。Claude 3.5 Sonnet 是 Kevin 在措辞、个性、编程和头脑风暴方面最喜欢的模型,但成本仍然偏高,因此许多工作负载会选择其他模型。“鉴于我们处理大量内容,价格确实是我们会考虑的因素。”

6. 播客结构让说话人分离变得可行,但并未解决

  • Snipd 受益于一个边界清晰的领域:播客通常拥有高质量、受控的录音室音频和稳定的背景噪声。Kevin 将其与面向会议的语音产品对比,后者必须处理任意房间和环境,能够使用的结构化启发式规则和领域内数据都稀缺得多。

  • 一个启发式规则是播出时长:在1小时的节目里,一个只出现30秒的声音大概率既不是主持人,也不是嘉宾,而是插入的广告商。底层系统会对语音片段做 embedding,并按说话人聚类;Snipd 修改了聚类方式,还应用了通用说话人分离服务无法提供的播客专属规则。

  • LLM 将文字稿含义与声学聚类结合起来,识别说话人并重新校准交接点。swyx 质疑 LLM 可能引入错误,不适合充当精确权威;Kevin 直言承认,Snipd 的说话人分离“也并不完美”,但他说,最近的节目相比1年前处理的节目已经有了明显改善。

  • 不过,在一集难度很高、嘉宾众多的 TechMeme Ride Home 节目中,swyx 仍认为 Snipd 的展示和识别效果优于 Descript。两人都预计,如今这套专用管线最终会被一个直接摄入原始音频的多模态模型取代;Kevin 测试过 Gemini 1.5 Flash,但表示其成本仍远高于自部署管线。

7. 生产质量藏在正则表达式、评审模型和“凭感觉评估”里

  • 节目聊天展示了原型与产品之间的距离。Kevin 说,把文字稿粘贴到 ChatGPT,让它生成带时间戳的回答,15分钟就能做出演示;但要做到“99%的时间”正确运行,就需要精心格式化上下文、设计提示词、处理引用、渲染前端,并从格式错误的响应中恢复。

  • Snipd 用“无数个正则表达式”修复格式错误,避免它们最终变成难看的 UI。再接一个 LLM 可能听起来更符合 AI 直觉,但聊天必须立即流式输出;现有 API 很难把不完整的输出再接入另一条实时纠错流。swyx 的结论是,这套看似不优雅的机制就是普通的“现实世界工程”。

  • Kevin 将公司的大量评估文化概括为“凭感觉评估”。创业公司可以容忍偶尔的粗糙,因为用户获得快速迭代作为交换;而 Spotify 级别的组织,可能会让同样一半的功能经历6个月的法律和组织审查。这里交换的是速度,并不意味着失败已经消失。

  • Snipd Wrapped 暴露了长尾问题:模型通常能选出一句有代表性的精彩引语,但偶尔会返回一段莫名其妙的乏味内容。解决办法是先用更便宜的模型生成5个候选,再让强得多的 LLM 评审并选出一个。Snipd 在验证识别出的书籍和说话人时,也采用同样的挑战者模式。

8. 当聊天框消失,消费级 AI 才会赢

  • swyx 正式提出的功能需求是可通过用户自定义提示词来定制摘要。Kevin 指出,如果每位用户都要重新生成一份摘要,长上下文个性化仍然很贵;不过,随着模型成本下降,“一切都个性化”最终应该会成为可能。

  • Kevin 当下给出的技术答案是正确的:打开节目聊天,要求它用指定风格生成摘要。但随后他接受了更深层的需求——swyx 在移动中收听,不想主动发起聊天。真正的产品应该在打开时就提供已经生成、结构化并以视觉方式呈现的偏好信息。

  • 这一区别构成了 Kevin 的消费级 AI 论点:开发者必须“走出聊天框”。智能可能已经存在,但机会在于找到人们与它互动的自然界面。他反复提醒做 AI 的创业者:“这里真正要完成的任务是什么?”

  • 终点是隐形基础设施。Kevin 的比喻是:“AI 是未来的电力。”没人会把麦克风或手机营销成“支持电力”,未来“AI 驱动”也会同样显得多余。Snipd 现在使用这个说法,是因为新奇感能吸引注意力;但它真正想成为的,只是更好的播客产品。

9. 语音可以把节目结尾变成学习触发器

  • Kevin 会通过现实生活中触发用户打开产品的场景来评估消费级产品:旅行会触发 Airbnb,但学语言没有自然的每日时刻。Duolingo 是制造出这种习惯的例外,它靠通知、连续打卡、排行榜和“猫头鹰梗”完成了这一点;Kevin 称它为这场游戏的“GOAT”。

  • Snipd 已经继承了一个触发点——独处、走路、开车、锻炼,或其他开始听播客的场景——但还没有处理所学内容的触发点。Kevin 自己的强制机制是每集节目结束后只选一个收获,即使其中有10个想法都很有价值,因为优先级排序会检验相关性,也更可能促成行动。

  • 他设想的语音产品会在播放结束时启动:不再自动进入下一集,而是由 AI 伴侣进行2至3分钟的复盘对话。每月超过5亿播客听众已经具备前置习惯。如果用户觉得这点微小投入能实质改善记忆、应用和思考,Kevin 认为 Snipd 可能让他们获得的价值提升“10倍”。

  • 具体实现仍有意保持开放;swyx 起初抗拒把它描述成另一个“与你的播客聊天”功能,但后来认识到其真正的使用场景在收听之后。Kevin 也预计语音克隆会逐渐普及:早期 NeurIPS 研讨会上曾引发严肃伦理争论的能力,如今已进入大众工具;未来,一个获授权的“AI swyx”或许可以讨论 Latent Space 节目。

10. YouTube 将发现问题变成平台问题

  • Kevin 的内容路线图不止于 RSS 播客:用户已经可以添加 YouTube 视频和上传有声书,而他希望加入原生有声书,以及把 Deep Research 与 NotebookLM 式播客生成结合起来的 AI 节目。swyx 则推动他优先考虑视频播客而非有声书,因为在他看来,YouTube 已经是“最好的播客平台”。

  • 发现机制也应从不透明的个性化推荐,转向用户可以直接对话的算法。Kevin 举的例子是对信息流说:“我知道我他妈特别喜欢”猫视频,但接下来2小时给我展示 AI 内容。TikTok 可能没有动力放弃对互动时长的优化,而与用户目标一致的学习产品有。

  • Kevin 采用了 Spotify 首席产品官 Gustav Söderström 提出的“可后台播放的视频”这一说法:即使在 YouTube 上,约90%的播客消费仍伴随其他活动,但画面有助于呈现片段、幻灯片、演示和主持人与观众的连接。他加入视频的最大理由是发现——视频推荐更有吸引力,但用户不必全程投入注意力。

  • swyx 认为,除听众和 Snip 制作者外,播客创作者是第三类利益相关者。Riverside 最接近这一方向,Descript 则掌握编辑环节,但两者都没有提供他想要的工作流:“坐下,录制,按一下按钮,完成。”机会同时伴随着真实的迁移阻力:尽管他很喜欢 Snipd,仍花了4至6个月才迁移,因为 OPML 无法携带听到一半的节目、排名和积累下来的状态。

swyx

Hey, I’m here in New York with Kevin Ben-Smith of Snipd. Welcome.

Kevin Ben-Smith

Hi. Hi. Amazing to be here.

swyx

Yeah. This is our first-ever outdoor podcast recording, I think. It’s quite a location for the first time, I’d say.

Kevin Ben-Smith

I was actually unsure because it’s cold. I checked the temperature; it’s 1°C.

swyx

It’s not that bad with the sun.

Kevin Ben-Smith

No, it’s quite nice.

swyx

Yeah. Yeah. Especially with our beautiful tea.

Kevin Ben-Smith

With the tea.

swyx

Yeah. Perfect. We’re going to talk about Snipd. I’m a Snipd user. Apart from Twitter, it’s the number-one-used app on my phone. When I wake up in the morning, I open Snipd and see what’s new. In terms of time spent or usage on my phone, I think it’s number 1 or number 2.

I had to talk about it because we’re in an AI podcast. We have to talk about AI podcasts. But before we get there, we just finished the AI Engineer Summit, and you came for the 2 days. How was it?

Kevin Ben-Smith

It was quite incredible. For me, the most valuable part was being in the same room with like-minded people who are building the future and seeing the future. Especially when it comes to AI agents, I often have conversations with friends who aren’t in the AI world, and it happens so quickly that it sounds like you’re talking in science fiction. It’s just crazy talk.

It was so refreshing to talk with so many other people who already see these things and be inspired by them, rather than always feeling like, “Okay, I think I’m just crazy, and this will never happen.” It really is happening, and for me it was very valuable.

swyx

So day 2 was more relevant for you than day 1?

Kevin Ben-Smith

Yeah, day 2 was the engineering track. That was definitely the most valuable for me, also as a practitioner myself. There were 1 or 2 talks about voice AI and AI agents with voice, which was quite fascinating. I also spoke with the speakers afterward, and they were very open. There’s this sharing attitude that I think is generally quite prevalent in the AI community. I learned a lot of practical things that I can now take away with me.

swyx

Yeah. On my side, I watched only about half of the talks because I was running around. I think people saw me toward the end; I was kind of collapsing. I was on the floor toward the end because I needed to get some rest. But I’m excited to watch the voice AI talks myself.

Kevin Ben-Smith

Yeah, do that. From my side, thanks a lot for organizing this conference and bringing everyone together. Do you have anything like this in Switzerland?

swyx

The short answer is no. I have to say, the AI community in Zurich, especially where we’re based, is quite good and growing. It’s especially driven by ETH, the technical university there, and all of the big companies that have AI teams there.

Google has its biggest tech hub outside the U.S. in Zurich. Meta is doing a lot with Reality Labs. Apple has a secret AI team. OpenAI is there, and SwiftKey just announced that they’re coming to Zurich. So there’s a lot happening.

Yeah, I think the most recent notable move was that the entire vision team from Google—Lucas Beyer and all the other authors of SigLIP—left Google to join OpenAI. I thought that was a big move, for a whole team to move all at once and at the same time.

I’ve been to Zurich, and it just feels expensive. It’s a great city with a great university, but I don’t see it as a business hub. Is it a business hub? I guess it is, right?

Kevin Ben-Smith

Historically, it’s a finance hub.

swyx

A finance hub?

Kevin Ben-Smith

Yeah. There are some large banks there, especially UBS, the largest wealth manager in the world. But it’s really becoming more of a tech hub now, with all of the big tech companies there.

swyx

And research-wise, is it all ETH, or are there other things?

Kevin Ben-Smith

Yeah, it’s all driven by ETH, and then there’s the university EPFL in Lausanne, which is also doing a lot. But it’s really ETH.

swyx

Otherwise, it’s a beautiful city. I can recommend that anyone come visit Zurich. Let me know; I’d be happy to show you around. Of course, you have nature so close, the mountains so close, and beautiful lakes. I think that’s what makes it such a livable city.

The cost isn’t cheap, but we’re in New York City right now, and I paid $8 for a coffee this morning. The coffee is cheaper in Zurich than in New York City.

Okay, let’s talk about Snipd. What is Snipd? Then we’ll talk about your origin story, but let’s get it crisp. What is Snipd?

Kevin Ben-Smith

I always see 2 definitions of Snipd. I’ll give you 1 really simple, straightforward one and then a second, more nuanced one, which I think will be valuable for the rest of our conversation.

The simplest way to put it is that we’re an AI-powered podcast app. If you listen to podcasts, we’re providing this AI-enhanced experience. But from a more nuanced perspective, we have a big focus on people like your audience who listen to podcasts to learn something new. Your audience wants to learn about AI—what’s happening, what’s the latest research, what’s going on—and we want to provide a spoken-audio platform where you can do that most effectively. AI is basically the way we can achieve that.

swyx

Means to an end.

Kevin Ben-Smith

Exactly.

swyx

When you started, was it always meant to be AI, or was it more about social sharing?

Kevin Ben-Smith

The first version that we ever released was about 3.5 years ago. This was before ChatGPT.

swyx

Before Whisper?

Kevin Ben-Smith

Yeah, before Whisper. A lot of the features that we now have in the app weren’t really possible yet back then. But from the beginning, we always had a focus on knowledge. That’s the reason why our team listens to podcasts.

We did have a different approach. The idea in the very beginning was that the name is Snipd, and you can create what we call Snips, which are basically small snippets or clips from a podcast. We envisioned a social platform, sort of like TikTok, where some people would listen to full episodes, snip certain best parts of them, and post those in a feed. Other users would consume this feed of Snips and use it as a discovery tool or as a means to an end.

You would have both people who create Snips and people who listen to Snips. Our big hypothesis in the beginning was that it would be easy to get people to listen to these Snips but super difficult to get them to create them. So we focused a lot of our effort on making it as seamless and easy as possible to create a Snip.

swyx

It’s similar to TikTok. You need CapCut for there to be videos on TikTok.

Kevin Ben-Smith

Exactly. For Snipd, whenever you hear an amazing insight or a great moment, you just triple-tap your headphones. Our AI then saves the moment that you just listened to and summarizes it to create a note. That’s basically a Snip.

We built all of this, launched it, and found the exact opposite. People used Snips to discover podcasts, but they really loved listening to long-form podcasts. They were creating Snips like crazy. This was definitely one of those aha moments when we realized that we should double down on knowledge and learning—helping you learn most effectively and capture the knowledge that you listen to, so you can actually do something with it.

We live in a world where there’s so much content. We consume and consume and consume, and it’s so easy to finish one podcast and immediately start listening to the next one. Five minutes later, you’ve forgotten 99% of what you actually just learned.

swyx

You don’t notice it, and most people don’t notice it, but this is my fourth podcast. My third podcast was a personal mixtape podcast where I manually snipped sections of podcasts that I liked, added my own commentary on top of them, and published them as small episodes.

They would be 5- to 10-minute snips of something that I thought was a good story or a good insight, and then I added my own commentary and published it as a separate podcast.

Kevin Ben-Smith

It’s cool. Is that still live?

swyx

It’s still live, but it’s not active. You can go back and find it if you’re curious enough. You’ll see it.

Kevin Ben-Smith

Nice, nice. You have to show me later.

swyx

It was very manual. My process would be: I’d hear something interesting, note down the timestamp and the URL of the podcast, and put it in my note-taking app. I used to use Overcast, so it would just link to the Overcast page. Then, whenever I felt like publishing, I would take one of those items, download the MP3, cut out the clip, record my intro and outro, and publish it as a podcast.

But now, with Snipd, I can just double-click or triple-tap.

Kevin Ben-Smith

Those are very similar stories to what we hear from our users. It’s normal that you’re doing something else while listening to your podcast. A lot of our users are driving, working out, or walking their dog. In those moments, when you hear something amazing, it’s difficult to write it down. You have to take out your phone.

Some people take a screenshot, write down the timestamp, and then later have to go back and try to find it again. Of course, you can’t find it anymore because there’s no search—there’s no Command-F. These were all issues that we encountered ourselves as users, and given that our background was in AI, we realized, “Wait, this should not be the case.”

Podcast apps today are basically repurposed music players, but we look at podcasts as one of the largest sources of knowledge in the world. Once you have that different angle, together with everything that AI is now enabling, you realize, “Hey, this is not the way podcast apps should be.”

swyx

Yeah, I agree. You mentioned something there: you said your background’s in AI. First of all, who’s on the team, and what do you mean by your background being in AI? Those are 2 very different questions.

Kevin Ben-Smith

Maybe starting with my backstory: it actually goes back, let’s say, 12 years or something like that. I moved to Zurich to study at ETH, and I studied something completely different. I studied mathematics and economics, basically specializing in quant finance.

swyx

Okay, well, all right.

Kevin Ben-Smith

So, yeah, there were all of these mathematical models for asset pricing, derivative pricing, and quantitative trading. For me, what fascinated me most was mathematical modeling, mathematics, and statistics, but I was never really that passionate about the finance side of things.

swyx

Really? Oh, okay. Yeah, I mean, we’re different there.

Kevin Ben-Smith

One symptom that I notice now, looking back, is that during that time, I think I never read an academic paper about the subject in my free time.

Then, toward the end of my studies, I was already working for a big bank. One of my best friends comes to me and says, “Hey, I just took this course. You have to do this. You have to take this lecture.”

I’m like, “What is it about?” He says, “It’s called machine learning.” And I’m like, “What kind of stupid name is that?” He sent me the slides, and over a weekend I went through all of them. I just knew: “Freaking hell, this is it. I’m in love.”

swyx

Wow. Yeah. Okay.

Kevin Ben-Smith

Over the course of the next 12 months, I really got into it. I started reading all about it, reading blog posts, and building my own models.

swyx

Was this course by a famous person at a famous university? Was it a Coursera thing?

Kevin Ben-Smith

No, this was an ETH course.

swyx

Oh, it was ETH. A professor at ETH? Did he teach in English, by the way?

Kevin Ben-Smith

Yeah, yeah, yeah.

swyx

Okay. So these slides are available somewhere?

Kevin Ben-Smith

Yeah, definitely. Though now they’re quite outdated.

swyx

Yeah, sure, sure.

swyx

Reflecting on the finance thing for a bit, I used to be a trader, on the sell side and buy side. I was an options trader first, and then I was more of a quantitative hedge fund analyst. We never really used machine learning. It was more like a little bit of statistical modeling, where you fit your regression.

swyx

No, I mean, that’s what it is. Or you solve partial differential equations and then use numerical methods to solve them. That’s for your degree. That’s not really what you do at work, right? Unless I don’t know what you do at work.

Kevin Ben-Smith

In my job, no. No, we weren’t solving the partial differential equations.

swyx

You learn all this in school and then you use it.

Kevin Ben-Smith

Let’s put it like this: in some things, yeah, I did code algorithms that would do it, but they were basically the most basic algorithms, and then you just slightly improved them. You tweak them here and there. It wasn’t like starting from scratch with a new partial differential equation.

swyx

No. Yeah, I mean, that’s real life, right? Most of it is kind of boring, or you’re using established things because they’re established because they tackle the most important topics.

swyx

Yeah, portfolio management was more interesting for me. We were sort of the first to combine social data with quantitative trading, and I think now it’s very common.

swyx

Then you went deep on machine learning. What happened next? You quit your job?

Kevin Ben-Smith

Yeah.

swyx

Wow.

Kevin Ben-Smith

I quit my job because I started using it at the bank as well. I desperately tried to find any kind of excuse to use it here or there, but it was clear to me that if I wanted to do this, I had to make a real cut.

So I quit my job and joined an early-stage tech startup in Zurich, where I built up the AI team over 5 years.

swyx

Wow.

Kevin Ben-Smith

We built various machine-learning systems for banks, from models for sales teams to identify which clients would like which product, what to sell to them, and for what reasons, all the way to doing a lot with bank transactions.

One of the most fun projects for me was an NLP model that would take the booking text of a transaction, such as a credit-card transaction, and prettify it. They had all of these numbers, abbreviations, and whatnot in there. Sometimes you’d look at it and think, “What is this?” The model would just change it to something like “CVS.”

swyx

Would you have hallucinations?

Kevin Ben-Smith

No, no, no. The way everything was set up, it wasn’t yet a fully end-to-end innovative neural network like what you would use today.

swyx

Okay, okay, awesome. And then when did you go full-time on Snipd?

Kevin Ben-Smith

That was afterward. The friend who got me into machine learning also got me interested in startups. He’s had a big impact on my life. His background is also in AI and data science.

The 2 of us would just jam on startup ideas every now and then. We had a couple of ideas, but because we were working full-time, we thought, “We could participate in HackZurich. It’s just a weekend. Let’s try out an idea, hack something together, and see how it works.”

The idea was that we’d be able to search through podcast episodes within a podcast. We did that, and long story short, we managed to build something that made us realize, “Hey, this actually works.” You could find things again in podcasts through natural-language search.

We pitched it on stage, and we actually won the hackathon, which was cool. I think we also had a good pitch and a good example. We used the famous Joe Rogan episode with Elon Musk where Elon Musk smokes a joint. It’s a 2½-hour episode, so we were on stage and searched for “smoking weed.” It would find that exact moment and play it, with Elon Musk coming on and smoking.

swyx

Was it video as well?

Kevin Ben-Smith

No, it was completely based on audio, but we did have the video for the presentation, which, of course, had an amazing effect. That gave us a lot of activation energy, but it wasn’t actually about winning the hackathon.

The interesting thing that happened was that after we pitched on stage, several of the other participants came up to us—many of them—and started saying, “Can I use this? I have this issue.” Some also told us about other problems that were very adjacent to this, asking whether they could use it for those as well.

That was the moment when we realized it wasn’t just us having these issues with podcasts and getting the most out of this knowledge.

swyx

Yeah, there are other people.

Kevin Ben-Smith

That was, I guess, 4 years ago or something like that. Then we decided to quit our jobs and start this whole Snipd thing.

swyx

How big is the team now?

Kevin Ben-Smith

We’re just 4 people. We’re all technical: 2 on the backend side, with the AI and all of the other backend things, and 2 on the frontend side, building the app.

swyx

Which is mostly Android and iOS.

Kevin Ben-Smith

Yeah, it’s iOS and Android. We also have a watch app for Apple, but it’s mostly iOS.

swyx

The watch thing is very funny because in the Latent Space community, most of us have been slowly adopting Snipd. You came to me about a year ago and introduced Snipd to me. I was like, “I don’t know. I’m very sticky to Overcast.” Then, slowly, we switched. Why watch?

Kevin Ben-Smith

It goes back to the fact that a lot of our users do something else while listening to a podcast, right? Giving them the ability to capture this knowledge even though they’re doing something else at the same time is one of the killer features.

Maybe at some point I should give a bit more of an overview of all the features that we have.

swyx

Sure.

Kevin Ben-Smith

This is one of the killer features, and one big use case that people use it for is running. If you’re a big runner, a big jogger, or you cycle really, really competitively, a lot of people don’t want to take their phone with them when they go running.

You load everything onto the watch, so you can download episodes. If you have an Apple Watch with internet access and a SIM card, you can also stream directly.

swyx

That's also possible.

Kevin Ben-Smith

Of course, it's basically very limited to just listening and snipping, and then you can see all of your snips later on your phone.

swyx

Let me tell you about this error I just got: “Error playing episode. Substack, the host of this podcast, does not allow this podcast to be played on an Apple Watch.”

Kevin Ben-Smith

Yeah, that's a very beautiful thing. We found out that all of the podcasts hosted on Substack cannot be played on an Apple Watch.

swyx

What is this restriction? What?

Kevin Ben-Smith

Don't ask me. We tried to reach out to Substack. We tried to reach out to some of the bigger podcasters who host their podcasts on Substack to let them know.

swyx

Uh-huh.

Kevin Ben-Smith

Substack doesn't seem to care. This is not specific to our app. You can also check out the Apple Podcasts app.

swyx

Yeah.

Kevin Ben-Smith

It's the same problem. It's just that we actually identified it, and we tell the user what's going on.

swyx

I will say, we host our podcast on Substack, but they're not very serious about their podcasting tools. I've told them before; I've been very upfront with them, so I don't feel like I'm on them in any way. It's kind of sad because otherwise it's a perfect creator platform, but the way that they treat podcasting as an afterthought, I think it's really disappointing.

Kevin Ben-Smith

Maybe, given that you mentioned all these features, I can give a bit of a better overview of what we have.

swyx

Okay, I'll tell you my version. You can correct me, right?

First of all, I think the main job is for it to be a podcast-listening app. It should basically be a complete superset of what you normally get on Overcast or Apple Podcasts or anything like that. You pull your show list from Listen Notes. How do you find shows? I type in anything, and you find them, right?

Kevin Ben-Smith

Yeah, we have a search engine powered by Listen Notes, but in the meantime, we have a huge database of, like, 99% of all podcasts out there ourselves.

swyx

Huge database.

Kevin Ben-Smith

Like, 99% of all podcasts out there.

swyx

What I noticed is that the default experience is that you do not automatically download shows. That's one very big difference for you guys versus other apps, where, if I'm subscribed to something, it automatically downloads, and I already have the MP3 downloaded overnight.

For me, I have to actively put it onto my queue, and then it automatically downloads. Initially, I didn't like that. I think I maybe told you that. I was like, “Oh, this is a feature that I don't like,” because it means that I have to choose to listen to it in order to download it. Is this opt-in? Is this between opt-in and opt-out?

So, I opt in to every episode that I listen to. Then you open it, and it depends on whether or not you have the AI stuff enabled, but the default experience is no AI stuff enabled. You can listen to it, and you can see the snips—the number of snips and where people snip during the episode—which roughly correlates to interest level. Obviously, you can snip there.

I think that's the default experience. I think snipping's really cool. I use it to share a lot on our Discord. We have tons and tons of people sharing snips and stuff, and tweeting stuff is also a nice, pleasant experience. But the real features come when you actually turn on the AI stuff.

Kevin Ben-Smith

I think that was a good basic overview. Maybe I can add a bit to it with the AI features that we have.

One thing that we do every time a new podcast episode comes out is transcribe the episode, do speaker diarization, identify the speaker names for each guest, extract a mini bio of the guest, and try to find a picture of the guest online and add it. We break the podcast down into chapters—AI-generated chapters—with a title and a quick description for each chapter.

swyx

That one's very handy.

Kevin Ben-Smith

We identify all the books that get mentioned on a podcast.

swyx

I don't use that one.

Kevin Ben-Smith

It depends on the podcast. There are some podcasts where the guests often recommend an amazing book. You can also find that again later on.

swyx

So, literally, you search for the word “book” or, like, “I just read blah blah blah”?

Kevin Ben-Smith

No, it's all LLM-based. We have an LLM that goes through the entire transcript and identifies whether a user mentions a book. Then we use the Perplexity API, together with various other LLM orchestration, to go out there on the internet, find everything there is to know about the book, find the cover, find out who the author is, and get a quick description of it.

For the author, we then check which other episodes the author appeared on.

swyx

Yeah, that is killer. For me, if there's an interesting book, the first thing I do is listen to a podcast episode with the writer because they usually give a really great overview already on a podcast. Sometimes the podcast is with the person as a guest. Sometimes the podcast is about the person without them there. Do you pick up both?

Kevin Ben-Smith

Yes, we pick up both in our latest models, but what we currently show you in the app—the goal is to only show you the guest, to separate that. In the future, we want to show the other things more, but that's—

swyx

For what it's worth, I don't mind. If I like somebody, I'll just learn about them regardless of whether they're there or not.

Kevin Ben-Smith

Yeah, I mean, yes and no. We've seen that there are some personalities where this can break down. The best examples for me are Sam Altman and Elon Musk. They're just mentioned on every second podcast, and it picks them up even though they're not on there.

swyx

I see.

Kevin Ben-Smith

We updated our algorithms and improved that a lot, and now it's gotten much better at only picking someone up if they're a guest.

To come back to the features, there are 2 more important features. We have the ability to chat with an episode.

swyx

Yes. Of course.

Kevin Ben-Smith

You can do the old-style searching through a transcript with keyword search, but I think for me, this is how you used to do search and extract knowledge in the past.

swyx

Old school.

Kevin Ben-Smith

The AI way is basically an LLM. You can ask the LLM, “Hey, when do they talk about topic X?” If you're interested in only a certain part of the episode, you can ask it to give you a quick overview of the episode or the key takeaways. Afterwards, you can also ask it to create a note for you. This is really very open-ended.

Finally, there's the snipping feature that we mentioned. Whenever you hear an amazing idea, you can triple-tap your headphones or click a button in the app, and the AI summarizes the insight you just heard and saves that together with the original transcript and audio in your knowledge library.

swyx

I also noticed that you skipped dynamic content.

Kevin Ben-Smith

We don't skip it automatically.

swyx

Oh, sorry, you detect—

Kevin Ben-Smith

But we detect it, yeah. That's one of the things that most people don't actually know. The way ads get inserted into most podcasts is that every time you listen to a podcast, you actually get access to a different audio file. On the server, a different ad is inserted into the MP3 file automatically.

swyx

Yeah, based on IP.

Kevin Ben-Smith

Exactly. What that means is that if we transcribe an episode and have a transcript with timestamps—word-specific timestamps—and you suddenly get a different audio file, all the timestamps are messed up. That's a huge issue, and for that we actually had to build another algorithm that dynamically, on the fly, resyncs the audio that you're listening to with the transcript that we have.

swyx

That's a fascinating problem in and of itself. Do you sync by matching up the sound waves, or do you sync by matching up words? Basically, do you do partial transcription?

Kevin Ben-Smith

We're not matching up words. It's happening basically at a byte level.

swyx

Matching?

Kevin Ben-Smith

Yeah, byte-level matching.

swyx

Okay.

Kevin Ben-Smith

It relies on there being exact matches at some point. Actually, we're not doing exact matches; we're doing fuzzy matches to identify the moment. We basically built Shazam for podcasts, just as a little side project to solve this issue.

swyx

Yeah, yeah. Actually, fun fact: apparently the Shazam algorithm is open. They published a paper and talked about it.

Kevin Ben-Smith

Yeah, I haven't really dived into the paper. I thought it was kind of interesting that basically no one else has built Shazam.

swyx

Yeah, I mean, the one thing is the algorithm. If you now talk about Shazam, the other thing is also having the database behind it and having the user mindset that if they have this problem, they come to you, right?

I'm very interested in the tech stack. There's a big data pipeline. If you share what the tech stack is, what are the most interesting or challenging pieces of it?

Kevin Ben-Smith

The general tech stack is that our entire backend—or 90% of our backend—is written in Python. We're hosting everything on Google Cloud Platform. Our front end is written with—well, we're using the Flutter framework.

swyx

Ah.

Kevin Ben-Smith

So, it's written in Dart and then compiled natively. We have 1 codebase that handles both Android and iOS.

swyx

You think that was a good decision?

Kevin Ben-Smith

It's something that a lot of people are exploring. So up until now, yes.

swyx

Okay.

Kevin Ben-Smith

Look, it has its pros and cons. Earlier, I mentioned that we have an Apple Watch app.

Yeah.

I mean, there's no Flutter for that, right? So you build native, and then, of course, you have to sync these things together. I'm not the front-end engineer, so I'm just relaying this information, but our front-end engineers are very happy with it. It's enabled us to be quite fast and be on both platforms from the very beginning.

When I talk with people and they hear that we are using Flutter, they usually think, "It's not performant. It's super janky," and everything. Then they use our app, and they're always super surprised. Or if they've already used our app and I show it to them, they're like, "What?"

swyx

Yeah. So there is actually a lot that you can do. There are a few concerns, right? One, it's Google, so when are they going to abandon it? Two, they're optimized for Android first, so iOS is a second thought. You can feel that it is not a native iOS app. But you guys put a lot of care into it.

Maybe three, from my point of view, as a JavaScript guy, React Native was supposed to be that dream, and I think that it hasn't really fulfilled that dream. Maybe Expo is trying to do that, but, again, it does not feel as productive as Flutter. I spent a week on Flutter and Dart, and I'm an investor in FlutterFlow, which is the low-code Flutter startup that's doing very, very well. I think a lot of people are still Flutter skeptics.

Kevin Ben-Smith

Yeah.

swyx

Wait, so are you moving away from Flutter?

Kevin Ben-Smith

No, we don't have plans to do that.

swyx

You're just saying about the watch-out. Okay, let's go back to the stack. That was just to give people a bit of an overview. I think the more interesting things are, of course, on the AI side.

Kevin Ben-Smith

As I mentioned earlier, when we started out, it was before ChatGPT, before the ChatGPT moment, before there was the GPT-3.5 Turbo API. So in the beginning, we were actually running everything ourselves: open-source models, trying to fine-tune them.

swyx

What did you use before Whisper for transcription?

Kevin Ben-Smith

Yeah, we were using wav2vec 2.0.

swyx

I see. It was the Google one, right?

Kevin Ben-Smith

No, it was the Facebook one. That was actually one of the papers that, when it came out, was one of the reasons why I said we should try to start a startup in the audio space. Before that, I had been following the NLP space quite closely. As I mentioned earlier, we did some stuff at the startup I was working at before. Wav2vec 2.0 was the first paper that I had at least seen where the whole transformer architecture moved over to audio.

swyx

Yeah.

Kevin Ben-Smith

A bit more generally, it was the first time that I saw the transformer architecture being applied to continuous data instead of discrete tokens. It worked amazingly. The transformer architecture plus self-supervised learning—these 2 things moved over. For me, it was like, "Hey, this is now going to take off similarly to how the text space has taken off."

With these 2 things in place, even if some features that we want to build are not possible yet, they will be possible in the near term with this trajectory. So that's a little side note.

In the meantime, we're using Whisper. We're still hosting some of the models ourselves. For example, the whole transcription and speaker-diarization pipeline needs to be as cheap as possible. We're doing this at scale, where we have a lot of audio.

swyx

What numbers can you disclose? Just to give people an idea, because it's a lot.

Kevin Ben-Smith

We have more than 1 million podcasts that we've already processed.

swyx

When you say 1 million, processing is basically that you have some kind of list of podcasts that you auto-process, and others where a paying member can choose to press the button and then transcribe it, right? Is that the rough idea?

Kevin Ben-Smith

Yeah, exactly. If you press that button, or we auto-transcribe it, first we do the transcription and the speaker diarization. Basically, you identify speech blocks that belong to the same speaker. This is then all orchestrated with an LLM to identify which speech block belongs to which speaker.

Together with that, as I mentioned earlier, we identify the guest name and the bio. So all of that comes together with an LLM to assign speaker names to each block.

Most of the rest of the pipeline we've now migrated to LLM APIs. We mainly use OpenAI and Google models—the Gemini models and the OpenAI models—and we use some Perplexity, basically, for those things where we need web search.

swyx

That's something I'm still hoping for, especially from OpenAI: that they will also provide us an API. No way. Basically, for us as a consumer, the more providers there are, the more competition there is, and that will lead to better results and lower costs over time.

Kevin Ben-Smith

I don't see Perplexity as expensive. If you use the web search, the price is like $5 per 1,000 queries, which is affordable. But if you compare that to just a normal LLM call, it's much more expensive.

swyx

Okay. Have you tried Exa?

Kevin Ben-Smith

We've looked into it, but we haven't really tried it.

swyx

We started with Perplexity, and it works well. If I remember correctly, Exa is also a bit more expensive. I don't know. They seem focused on search as a search API, whereas Perplexity is maybe more of a consumer business with higher margins. I'll put it like this: Perplexity is trying to be a product; Exa is trying to be infrastructure. That's my distinction there.

The other thing I will mention is that Google has a search grounding feature.

Kevin Ben-Smith

We've also tried that. We didn't go into too much detail in really comparing it quality-wise, because we already had the Perplexity one, and it's working. I think the price there is actually higher than Perplexity.

swyx

Really? Google should cut their prices.

Kevin Ben-Smith

Maybe it was the same price. I don't want to say something incorrect, but it wasn't cheaper. It wasn't compelling. Then there was no reason to switch.

In general, for us, given that we work with a lot of content, price is actually something that we do look at. For us, it's not just about taking the best model for every task, but really identifying what kind of intelligence level you need and then getting the best price for that, to be able to really scale this and let our users use these features with as many podcasts as possible.

swyx

Yeah. I wanted to double-click on diarization. It's something that I don't think people do very well. I'm a Bee user. I don't have it right now, but they were supposed to speak, and they dropped out at the last minute. We've had the Bee AI guys on the podcast before, and it's not great yet. Do you use just pyannote, the default stuff, or do you find any tricks for diarization?

Kevin Ben-Smith

We do use the open-source packages, but we've tweaked them a bit here and there. For example, if you mention the Bee AI guys, I actually listened to the podcast episode. It was super nice, thank you. When you started talking about speaker diarization, I just had to think about their use case. With all of the different environments, it can basically be anything. It's completely out of domain; there's no data for this.

I was feeling for them, because our advantage is that we're working with very high-quality audio. It's very controlled, usually recorded in a studio. This is quite an exception, I guess.

swyx

It is kind of a studio. It's pretty quiet. There's consistent background noise, which you can edit out.

Kevin Ben-Smith

Yeah. There's New York.

swyx

It's nice. It's a character.

Kevin Ben-Smith

That, of course, helps us. Another thing that helps us is that we know certain structural aspects of the podcast. For example, how often does someone speak? If there's a 1-hour episode and someone speaks for 30 seconds, that person is most probably not the guest and not the host. It's probably some ad, like some speaker from an ad.

swyx

Okay.

Kevin Ben-Smith

So we have certain heuristics that we can use and leverage to improve things. In the past, we've also changed the clustering algorithm. Basically, how a lot of this speaker diarization works is you create an embedding for the speech that's happening, and then you try to somehow cluster these embeddings and find out: this is all one speaker; this is all another speaker.

There, we've also tweaked a couple of things where we again used heuristics that we could apply from knowing how podcasts function. That's also actually where I was feeling so much for the Bee AI guys, because all of these heuristics are probably almost impossible for them to use. It can just be any situation, anything.

Another thing is that we actually combine it with LLMs: the transcript, LLMs, and the speaker diarization. We bring all of these together to recalibrate some of the switching points—when does the speaker stop, and when does the next one start?

But the LLMs can add errors as well. I wouldn't feel safe using them to be so precise. At the end of the day, just to avoid giving the wrong impression, the speaker diarization we're doing isn't perfect either.

swyx

I basically don't really notice it. I use it for search.

Kevin Ben-Smith

Yeah, it's not perfect yet, but it's gotten quite good. Especially if you take a latest episode and compare it to an episode that came out a year ago, we've improved it quite a bit.

swyx

Well, it's beautifully presented. I love that I can click on the transcript and it goes to the timestamp. It's so simple, but it should exist.

Kevin Ben-Smith

Yeah, I agree. I agree.

swyx

I'm loading a 2-hour episode of The TechMeme Ride Home, where there are a lot of different guests calling in, and you've identified the guest names.

Kevin Ben-Smith

Indeed. These are all LLM-based.

swyx

Yeah, it's really nice. The speaker names—I would say I'm a power user of all these tools—you've done a better job than Descript.

Kevin Ben-Smith

Okay, well—

swyx

Descript has so much funding. They had OpenAI invested in them, and they still suck. So, keep going. You're doing great.

Kevin Ben-Smith

Thanks, thanks. I would say that, especially for anyone listening who's interested in building a consumer app with AI, if your background is in AI and you love working with AI and doing all of that, the most important thing is to keep reminding yourself of what the job to be done actually is here. What does the consumer actually want?

For example, we're delighted by the ability to click on this word and have it jump there. This is not rocket science. You don't have to be Andrej Karpathy to come up with that and build it. I think that's something that's super important to keep in mind.

swyx

Yeah, amazing. There are so many features; it's so packed. There are quotes that you pick up, summarization—and, by the way, I'm going to use this as my official feature request. I want to customize how it's summarized. I want custom prompts, because your summarization is good, but I have different preferences.

Kevin Ben-Smith

Yeah, I completely get your feature request, and I think it just shows that people have asked for it. Maybe, in general, as a way of thinking about the future, I think everything will be personalized. This isn't specific to us.

Today, we're still in a phase where the cost of LLMs—at least if you're working with long context windows like we are—has to be taken into consideration. There are a lot of tokens in an entire podcast, so if we regenerated everything for every single user, it would get expensive. In the future, the cost will continue to go down, and then it will just be personalized.

That being said, you can already do this today. If you go to the player screen and open up the chat, you can ask for a summary in your style.

swyx

Yeah, okay. I mean, I listen to consume, you know. I've never really used this feature. I think that's me being a slow adopter.

Kevin Ben-Smith

No, no—I mean, when does the conversation start?

swyx

Okay. I mean, you can just type anything.

Kevin Ben-Smith

I think what you're describing is maybe an interesting topic to talk about. I told you, “Look, we have this chat; you can just ask for it.” This is how ChatGPT works today, but if you're building a consumer app, you have to move beyond the chat box. People don't always want to type out what they want.

Your feature request, even though it's theoretically already possible, is actually saying, “I just want to open up the app, and it should be there in a nicely formatted, beautiful way, so I can read or consume it without any issues.” I think that's generally where a lot of the opportunities lie in the market right now if you want to build a consumer app: taking the capability and intelligence, but figuring out the best user interface—the best way for a user to engage with that intelligence naturally.

swyx

This is something I've been thinking about as AI that's not in your face. Right now, we like to say that Notion has Notion AI, and there's a little thing there, or some other platform has the sparkle or magic-wand emoji: “That's our AI feature. Use this.” A lot of people don't like it. It should just become invisible, kind of like invisible AI.

Kevin Ben-Smith

100%. The way I see it is that AI is the electricity of the future. We don't talk about how this microphone uses electricity or how this phone uses electricity. You don't think about it that way; it's just in there. It's not an electricity-enabled product. It's just a product.

It will be the same with AI. Right now, it's still something you use to market your product. We do the same thing because it's still something people recognize as new. But at some point, it will just be a podcast app, and it will be normal that it has AI in it.

swyx

I noticed you do something interesting in your chat where you source the timestamps. Is that part of the prompt, or is there a separate pipeline that adds the sources?

Kevin Ben-Smith

This is actually part of the prompt. It's all prompt engineering: figuring out how to provide the context—we provide the entire transcript—and then getting the model to respond correctly in a certain format, and rendering that on the front end.

swyx

This is one of those examples where it's so easy to create a quick demo. You can just go to ChatGPT, paste this thing in, say, “Do this,” and 15 minutes later you're done. But getting it to production level, so that it actually works 99% of the time, is where the difference lies.

Kevin Ben-Smith

For this specific feature, we also have countless regular expressions. They're there to correct certain things the LLM does because it doesn't always adhere to the format correctly. Then it looks super ugly on the front end.

swyx

Why don't you use an LLM for that? That's sort of the AI-native way. Who uses regular expressions anymore?

Kevin Ben-Smith

With the chat, for user experience, it's very important to have streaming. Otherwise, you have to wait so long until your message arrives. We're streaming the text live, just like ChatGPT.

If you're streaming the text and something is incorrect, it's currently not easy to pipe that stream into another stream and get the corrected stream back.

swyx

Yeah, yeah, yeah. Stream it into another stream, get the corrected stream back—that would be amazing. I don't know; maybe you can answer that. Do you know of any way to do it?

Kevin Ben-Smith

There's no API that does this. You can't stream it in.

swyx

If you own the models, you can take whatever token sequence has been emitted and start loading that into the next one, if you fully own the models.

Kevin Ben-Smith

I don't know. It's probably not worth it. What do you think is better?

swyx

I think most engineers who are new to AI research and benchmarking don't know how much regular-expression work goes into normal benchmarks. It's just this ugly list of 100 different matches for whatever criteria you're looking for.

Kevin Ben-Smith

Yeah. No, it's very cool. I think it's an example of real-world engineering.

swyx

Do you have tooling that you're proud of that you developed for yourself? Is it just a test script?

Kevin Ben-Smith

I think it's a bit more. Vibe evals was a term that came up in one of the talks—I think it might have been the first day of the conference. A lot of the talks were about evals, which are so important. For us, it's a bit more like vibe evals.

That's also part of being a startup: we can take risks. We can accept the cost of something sometimes failing a little bit or being a little off, and our users know that. They appreciate that, in return, we're moving fast, iterating, and building amazing things.

With Spotify, or something like that, half of our features would probably be in a 6-month review through legal—or whatever—before they could ship.

swyx

Let's just say Spotify is not very good at podcasting. I have a documented dislike of its podcast features. Overall, they're not very well integrated.

Any other LLM-focused engineering challenges or problems that you want to highlight?

Kevin Ben-Smith

I think it's not unique to us, but it goes again in the direction of handling the uncertainty of LLMs. At the end of last year, we did a sort of Snipd Wrapped, and one of the things we thought would be fun was to do something with an LLM and the snips that a user has.

Three, let's say, unique LLM features were that we assigned a personality to you based on the snips that you had. It was all just a bit of a fun, playful way—

swyx

I’m going to look at mine. I forgot mine already.

Kevin Ben-Smith

I don’t know whether it’s still in the Discord. We all took screenshots of it.

swyx

Ah, okay.

Kevin Ben-Smith

It’s in the Discord. The second one was a learning scorecard, where we identified the topics that you snipped on the most, and you got a little score for that. The third one was a quote that stood out. The quote is actually a very good example: we would run that for a user, and most of the time it was an interesting quote, but every now and then it was a super-boring quote. You’d think, “Why did you select that? Come on.”

For that, the solution was to say, “Give me 5 candidates.” It accepted 5 quotes as candidates, and then we piped them into a different model as a judge—an LLM as a judge. We used a much better model.

swyx

Okay.

Kevin Ben-Smith

With the initial model, as I mentioned earlier, we do have to look at the cost because we have so much text going into it. We use a slightly cheaper model there, but the judge can be a really good model that chooses 1 out of 5.

swyx

This is a practical example. I can’t find it. Bad search in Discord. So, do you recommend having a much smarter model as a judge?

Kevin Ben-Smith

Yeah.

swyx

And that works for you?

Kevin Ben-Smith

Yeah.

swyx

Interesting. I think this year I’m very interested in LLM-as-a-judge being developed more as a concept. For things like Snipd Wrapped, it’s fine. It’s entertaining, and there’s no right answer.

Kevin Ben-Smith

We also use the same concept for our books feature, where we identify the books that were mentioned. Ninety percent of the time, it works perfectly out of the box in 1 shot, but every now and then it starts identifying books that weren’t really mentioned, books that aren’t books, or it starts making up books. We basically have another LLM challenge it. We do the same thing with the speakers, now that I think about it. I think it’s a great technique.

swyx

Interesting. You run up a lot of costs. You mentioned costs: you moved from self-hosting a lot of models to the big lab models—OpenAI and Google. Anthropic?

Kevin Ben-Smith

No, we love Claude. In my opinion, Claude is the best when it comes to the way it formulates things.

swyx

The personality.

Kevin Ben-Smith

The personality. I actually really love it, but the cost is still high.

swyx

You tried Haiku, but you have to have Sonnet?

Kevin Ben-Smith

With Haiku, we haven’t experimented too much. We obviously work a lot with 3.5 Sonnet. For coding, in Cursor, and in general for brainstorming, we use it a lot. I think it’s a great brainstorming partner. But with a lot of things that we’ve done, we opted for different models.

swyx

What I’m trying to drive at is: how much cheaper can you get if you go from closed models to open models? Maybe it’s 0% cheaper. Maybe it’s 5% cheaper. Or maybe it’s 50% cheaper. Do you have a sense?

Kevin Ben-Smith

It’s very difficult to judge that. I don’t really have a sense, but I can give you a couple of thoughts that have gone through our minds over time. We realize that, given that we have a couple of tasks where so many tokens are going in, at some point it will make sense to offload some of that to an open-source model.

But going back to the fact that we’re a startup, we’re not an AI lab or whatever, the most important thing for us is to iterate fast. We need to learn from our users, improve the product, and maintain the velocity of those iterations. For that, the closed models hosted by OpenAI, Google, and Anthropic are just unbeatable because it’s simply an API call. You don’t need to worry about so much complexity behind it. That’s the biggest reason why we’re not doing more in this space.

There are other considerations for the future. We have 2 different usage patterns for LLMs. One is the preprocessing of a podcast episode: the initial processing, including the transcription, speaker diarization, and chapterization. We do that once, and the usage pattern is quite predictable because we know how many podcasts get released and when. We can have a certain capacity, and we’re running that 24/7. It’s 1 big queue running 24/7.

swyx

What’s the queue job runner? Is it Django, just the Python one?

Kevin Ben-Smith

No, that’s just our own. We have it in our database, and the backend talks to the database, picking up jobs and writing them back.

swyx

I’m just curious about orchestration and queues.

Kevin Ben-Smith

We also have a lot of other orchestration where we use Google Pub/Sub.

swyx

Okay.

Kevin Ben-Smith

The other usage pattern is when, for example, a user action triggers an LLM call. It has to be real-time, and there can be moments when usage spikes, followed by moments when there’s very little usage. For that, LLM API calls are perfect because you don’t need to worry about scaling up, scaling down, or handling those issues.

swyx

Serverless versus serverful.

Kevin Ben-Smith

Yeah, exactly.

swyx

I see OpenAI and all of these other providers as the—well, I guess—the Amazon, sorry, AWS, of AI. It’s similar to how, before AWS, you would have to have your own servers, buy new servers, or get rid of servers. With AWS, it became much easier to ramp things up and down.

Kevin Ben-Smith

Yeah, and this is taking it even to the next level for AI.

swyx

I’m a big believer in this. Basically, it’s intelligence on demand. We’re probably not using it enough in our daily lives to do things. We should be able to spin up 100 things at once, go through them, and then stop. I feel like we’re still trying to figure out how to use LLMs in our lives effectively.

Kevin Ben-Smith

Yeah, 100%. I think that goes back to the whole opportunity for a startup. It’s not about letting the big labs handle the challenge of more intelligence. It’s about the existing intelligence: how do you integrate it? How do you actually incorporate it into your life?

swyx

It’s AI engineering. Okay, cool, cool, cool. The 1 other thing I wanted to touch on was multimodality in frontier models. Dwarkesh had an interesting application of Gemini recently, where he fed raw audio in and got diarized transcription out—or timestamps out. I think that will come.

Basically, what we’re saying here is another wave of transformers eating things. Right now, models are pretty much single-modality things. You have Whisper, you have a pipeline, and everything.

Kevin Ben-Smith

No, no, no. We only feed the raw files.

swyx

Do you think that would be realistic for you?

Kevin Ben-Smith

I 100% agree. Basically, everything that we talked about earlier—the speaker diarization, the heuristics, and everything else—in the future would just be put into 1 big multimodal LLM, and it would output everything that you want.

I’ve also experimented with that, just with Gemini 1.5 Flash, for fun. The big difference right now is still the cost difference: doing speaker diarization this way or doing transcription this way is much more expensive than the pipeline we’ve built.

swyx

I need to figure out what that cost is because, in my mind, Gemini 2.0 Flash is so cheap. Maybe it’s not cheap enough for you.

Kevin Ben-Smith

No, I mean, if you compare it to Whisper and speaker diarization, especially self-hosting it—

swyx

Yeah, yeah. Okay.

Kevin Ben-Smith

But we will get there, right? This is just a question of time. As soon as that happens, we’ll be the 1st ones to switch.

swyx

Awesome. Anything else that you’re eyeing on the horizon as you think about a feature or incorporating some new AI functionality into the app?

Kevin Ben-Smith

There are so many areas that we’re thinking about. Our challenge is more about choosing.

swyx

Choosing.

Kevin Ben-Smith

Looking at the next couple of years, there are 4 big areas that interest us. One is content. Right now, it’s podcasts. You mentioned that you can also upload audiobooks and YouTube videos.

swyx

YouTube. I actually use the YouTube one a fair amount.

Kevin Ben-Smith

In the future, we want to have audiobooks natively in the app, and we want to enable AI-generated content. Think of Deep Research and NotebookLM podcast generation—put those together. That should be in our app.

The second area is discovery. In general, discovery as a paradigm in all apps will undergo a change thanks to AI.

swyx

I noticed that you don’t have—so, you have download counts and most snips, right? Something like that?

Kevin Ben-Smith

Yeah. On the discovery side, we want to do much more. Before Elon bought Twitter, there was a lot of talk about bringing your own algorithm to Twitter. That was Jack Dorsey’s big thing; he talked about it a lot. I actually think this is coming, but with a bit of a twist.

I think what AI will enable is not that you bring your own algorithm, but that you will be able to talk and communicate with the algorithm. You can just tell the algorithm, “Hey, you keep showing me cat videos. I know I freaking love them, and that’s why you keep showing them to me. But for the next 2 hours, I really want to get more into AI stuff. Do not show me cat videos.” Then it will just adapt.

Of course, the question is that big platforms like TikTok do not have the incentive to offer that.

swyx

Exactly. That’s what I was going to say.

Kevin Ben-Smith

But we actually are driven by helping you learn, get the most out of things, and achieve your goals. For us, it’s very much in our incentive to say, “Hey, you should be able to guide it.”

That was a long way of saying that I think this will happen a lot in recommendation. I think collaborative filtering will be the first step for RecSys, and then some other fancy stuff.

Maybe to go back to the question that you had before, these were the first 2 areas. The other 2 are voice as an interface and voice AI.

swyx

How does this exist?

Kevin Ben-Smith

Maybe I can first tell you a bit about why I find it so interesting for us. Historically, there has been so much talk about voice as an interface, and it always fell flat. The reason why I’m excited about it this time around is that, with any consumer app, I like to ask myself: What is the moment in my life, or the trigger in my life, that gets me to open this app and start using it?

For example, take Airbnb. The trigger is, “Ah, you want to travel,” and then you open up the app. For apps that do not already have this natural trigger in your life, it’s very difficult for a consumer app to get a user to use it.

swyx

You need a hook.

Kevin Ben-Smith

There’s basically only 1 super-successful app that has been able to do that without this natural trigger, and that is Duolingo.

swyx

Ah.

Kevin Ben-Smith

Everyone wants to learn a language, but you don’t have this natural moment during your day when you think, “Ah, now I need to open up this app.”

swyx

The notifications.

Kevin Ben-Smith

Exactly. The owl memes.

swyx

Exactly.

Kevin Ben-Smith

They gamified it super successfully and made it super beautiful. They are the GOATs in this game. But what’s much easier is when there is already this trigger, and then you don’t have to do all of the streaks and leaderboards.

Now, if you look at what we are doing and our goal of getting people to really maximize what they get out of their listening, we’re interested in a couple of features where we know we can 10x the value that people get out of a podcast. But we need them to do something for that. There is friction involved because it’s all about learning, right? It’s about thinking for yourself.

swyx

Apply the knowledge.

Kevin Ben-Smith

Exactly. You have to be forced to think about what was actually the main takeaway for you from this episode. There’s something that I like doing myself for every episode that I listen to: I try to boil it down to 1 single takeaway. Even though there might have been 10 amazing things, you pick 1—the most important one.

This is an active process, a forcing function in your brain to challenge all of the insights and really come up with the 1 thing that is applicable to you and your life, and what you might want to do with it. It also helps you turn it into action.

This is basically a feature that we’re interested in, but you have to get the user to use it. If this is all text-based, then we’re basically playing the same game as Duolingo, where at some point you’re going to get a notification from Snipd saying, “Hey, swyx, come on. You know you should do this.” Maybe there’s a blue owl.

But if you have voice, you can hook into the existing habits that the user already has. You already have this habit of listening to a podcast. You’re already doing that. Once an episode ends, instead of just jumping into the next episode, you can now have your AI companion come on and have a quick conversation. You can go through these things.

How that looks in detail, we still need to figure out. But this paradigm of staying in the flow relates a bit to what you were saying about AI that is invisible. You’re staying in the flow of what you’re already doing, but now we can insert a completely new experience that helps you get the most out of your listening.

swyx

I think your framing of this is very powerful. I think this is where you are a products person more than an engineer, because an engineer would just be like, “Oh, it’s just chat with your podcast. It’s chat with a PDF, chat with a podcast. Okay, cool.”

But you’re framing it in a different light that actually makes sense to me now, as opposed to previously. I don’t chat with my podcast. Why? I just listen to the podcast, right? But for you, it’s more about retention and learning and all that. You’re very serious about it. That’s why you started the company, and you’re focused on that.

I would admit that I’m still stuck in that consume, consume, consume mentality. I know it’s not good, but this is my default. That’s why I was a little bit lost when you were saying all the things about Duolingo and the trigger. My trigger for listening to the podcast is that I’m by myself. That’s my trigger.

But you’re saying the trigger is not about listening to the podcast. The trigger is remembering, retaining, and processing the podcast I just listened to.

Kevin Ben-Smith

No, what I meant is that you already have this trigger that gets you to start listening to a podcast. You already have that, and so do, I don’t know, millions of people. There are more than half a billion monthly active podcast listeners.

But you do not have this trigger, as you just said yourself, that gets you to regularly process this information. Voice, for me, is the ability to hook into your existing trigger. The trigger I was talking about is that your podcast ends, and you’re still listening. We just continue.

This can be 2 or 3 minutes. I’m not saying it has to be a 60-minute process. I think 2 or 3 minutes can happen completely naturally. If we manage to do that and you start noticing as a user, “I’m freaking out—I’m just now spending 3 minutes with this AI companion, but I’m taking this much more away,” then we’ve won.

Retention is 1 thing, but you start to take what you’ve learned and apply it to what’s important to you—your thinking.

swyx

A lot of people rely on Anki notes, flashcards, and all that to do this. But making the notes is also a chore, and I think this could be very interesting.

I’m just noticing that it’s kind of a different usage mode. You already talked about this: the name Snipd is very snip-centric. I originally resisted adopting Snipd because of that. But now you observe that people are listening to long-form episodes and talking at the end.

The ideal implementation of this is that I browse through a bunch of snips from the things I’m subscribed to. I listen to the snips, I talk with it, and then maybe it double-clicks on the podcast and finds other timestamps that are relevant to the thing that I want to talk about. I was just thinking about that. I don’t know if that’s interesting.

What are your thoughts on voice cloning? I’ve had my voice cloned, and people have talked to me through an AI version of me. Is that too creepy?

Kevin Ben-Smith

I don’t think it’s too creepy in the future. With a lot of these things, society is going through a change. Things seem quite weird now, but in the future they’ll seem normal.

I think voice cloning has already become much more normalized. I remember I was at the NeurIPS conference—I think it was in San Diego?

swyx

No, LA.

Kevin Ben-Smith

LA. It was the Florida one.

swyx

Yeah, yeah, yeah, yeah. Florida.

Kevin Ben-Smith

Everyone says that was peak NeurIPS. I remember there was a talk or workshop by Lyrebird. They actually got acquired by Discord later. They were showing off their technology and there was a huge discussion afterward about all of the moral and ethical implications. It really felt like this would never be accepted by society.

You look at it now: you have ElevenLabs, and anyone can just clone their voice. No one really talks about it as if the world is going to end. I think society will get used to it.

In our case, there are some interesting applications where we’d also be super interested in working together with creators, like podcast creators, to play around with this concept.

swyx

I think that would be super cool if someone could come on to Snipd, go to the Latent Space podcast, and start chatting with AI swyx.

Kevin Ben-Smith

Yeah. No, I think we’ll be there. We want to—obviously, I think as an AI podcast, we should be the first consumers of these things.

swyx

Yeah. I would say that one observation I’ve made about podcasting—this is just the general state of the market, and you can ask me your questions, things you want to ask about podcasters—is that we’re focusing a lot more on YouTube this year. YouTube is the best podcasting platform. It is not MP3s. It is not Apple Podcasts. It is not Spotify. It’s YouTube.

It’s just the social layer of recommendations and the existing habit that people have of logging on to YouTube and getting that. That’s my observation. You can riff on that.

Kevin Ben-Smith

The only thing I would say is that when you were listing your priorities, you said audiobooks first over YouTube, and I would suggest, if I were you—

swyx

Yeah, as in YouTube video podcasts.

Kevin Ben-Smith

I mean, it’s obvious that video podcasts are here to stay.

swyx

Not just here to stay—bigger.

Kevin Ben-Smith

What I want to do with Snipd is obviously also add video to the platform.

swyx

Oh, yeah.

Kevin Ben-Smith

The way I see video is, I do believe—I like this concept of backgroundable video. I didn’t come up with this concept. It was actually Gustav Söderström, the CPO of Spotify.

swyx

Exactly. Exactly.

Kevin Ben-Smith

When I speak with people, it remains true that they listen to podcasts when they do something else at the same time. That’s around 90% of their consumption, also if they listen on YouTube.

But every now and then, it’s nice to have the video. It’s nice if you’re, for example, just watching a clip. It’s nice if they sometimes mention something, like they show some slides or something where you need to have the visual with it. It helps you connect much more with the host as a listener.

But the biggest benefit I see with video is discovery. I think that is also why YouTube has become the biggest podcast player out there, because they have the discovery. Discovery in video is just so much easier, so much better, and so much more engaging.

swyx

For consumers?

Kevin Ben-Smith

Yeah, for consumers.

swyx

Okay. I think that you almost have three different audiences. The vast majority of people for you are the people listening to podcasts, right?

Kevin Ben-Smith

Of course.

swyx

Then there’s a second layer of people who create Snips, right? Who add extra data and annotation value to your platform. By the way, we use the Snip count as a proxy for popularity, because we have download counts, but platforms like Spotify rehost our MP3 file, so we don’t get any download count for Spotify.

Snip count is active: I opt in to listen to you, and I shared this. Those are really, really good metrics.

But the third audience that you haven’t really touched is the podcast creators, like myself. For me, discovery from that point of view—not from your point of view—discovery for me is that I want to be discovered, and I think YouTube is still there. Twitter, obviously, for me, Substack, Hacker News. I try really hard to rank on Hacker News.

When TikTok took this very seriously, they prioritized the creators of the content. For you, the creator of the content was the Snips. But there may be a world for you in which you prioritize the creators of the podcasts.

Kevin Ben-Smith

Yeah, interesting observation. What are some of your ideas or thoughts? Do you have some specific examples?

swyx

Riverside is the closest that has come to it. Descript is number 2. Descript bought a Riverside competitor, and as far as I can tell, it hasn’t been very successful.

Descript has a very, very good niche and a very, very good editing angle, and then just hasn’t done anything interesting since then. Although Underlord is good, it’s not great. Your chapterization is better than Descript’s. Again, they should be able to beat you. They’re not.

Riverside is good also—very, very good. We actually recently started a second series of podcasts within Latent Space that is YouTube-only, because you only find it on YouTube. It’s also shorter. This is a 1.5- to 2-hour thing; the other one is remote-only, 30 minutes, chop-chop. Send it all into Riverside.

Riverside is pretty good for that. Not great. It doesn’t do good thumbnails. It doesn’t do the editing very well; it’s still a little bit rough. It has an auto-editor where whoever’s actively speaking is focused on, and then sometimes it goes back to the multi-speaker view. That kind of stuff. People like that.

But the Shorts are still not great. I still need to manually download them and then republish them to YouTube. I still need to pick the Shorts. They mostly suck. There are still a lot of rough edges there.

Ideally, as a creator, you know what I want. You definitely know what I want. I sit down, record, press the button, and I’m done.

swyx

Yeah. I think you guys could do it.

Kevin Ben-Smith

Okay. So if I can translate that, for you it’s really about simplifying the creation process of the podcast.

swyx

Yeah. And I’ll tell you what: this will increase the quality, because the reason that most podcasts or YouTube videos are made by people who don’t have life experience, who are not that important in the world, and who aren’t doing important jobs.

What you actually want to enable is CEOs to each make their own podcasts. They’re busy. They’re not going to sit there and figure out Riverside.

A lot of the reason that people like Latent Space is that it takes an idiot like me, who could be doing a lot more with my life, making a lot more money, and having a real job somewhere else, and I just choose to do this because I like it. Otherwise, they’ll never get access to me or to the people that I have access to.

That’s my pitch.

Kevin Ben-Smith

Cool. Anything else that you normally want to talk to podcasters about?

I think we’ve covered everything.

I guess, as a last message, go try out Snipd. It’s a freemium version, so you can use and try out everything for free. I’m also happy to provide you with a link that you can add to the show notes to try out the premium version for free for a month if people want to do that. Give it a shot.

swyx

I would say, yeah, thanks for coming on. After you demoed it to me, I did not convert for another 4 to 6 months because I found it very challenging to switch over.

I think that’s the main thing. You basically have OPML import, right? But there’s no way to import all the existing half-listened-to episodes or my rankings or whatever.

For listeners who are interested, I have a blog post where I talked about my switch. Just treat it as a chance to clean house.

Kevin Ben-Smith

That is a good point.

swyx

Yeah. Do new things and refocus your start. Restart 2025.

Great. Well, thank you for working on Snipd. Thank you for coming on. We usually spend a lot of time talking to big companies, venture startups, B2B SaaS, and that kind of stuff.

But I think your journey as a small team building a B2C consumer app is the kind of stuff that we also like to feature, because a lot of people want to build what you’re doing. They don’t see role models that are successful, confident, and having success in this market, which is very challenging.

Thanks for sharing some of your thoughts.

Kevin Ben-Smith

Thanks. Thanks for having me, and thank you for creating an amazing podcast and an amazing conference as well.

swyx

Thank you.