[BidClub_]
Latent Space · · 61 分钟

RAG之后的检索:混合搜索、Agent与数据库设计——Turbopuffer 的 Simon Eskildsen

Alessio FanelliswyxSimon Hørup Eskildsen

YouTube
TL;DR
  • Turbopuffer 押注 AI 将创造一个15年一遇的数据库窗口,因为模型可以围绕知识进行推理,却无法以完整保真度存储所有知识。 Simon Hørup Eskildsen 的品类判断是,每家公司都会把大型数据集接入 AI,从而催生一个外部事实源:“我们不可能把所有东西都压缩进几TB的权重里。”

  • 公司的起点,是一个极其具体、残酷的成本缺口:Readwise 整套基础设施每月约5,000美元,而一个有用的推荐功能预计就要花30,000美元。 Eskildsen 推断,如果成本只有原来的十分之一,Readwise 很可能已经上线该功能;于是他从餐巾纸算式出发,设计了一套以对象存储为先的搜索引擎,而不是先做广泛客户调研。“这件事一直萦绕着我。”

  • Turbopuffer 的架构窗口,直到云端 NVMe SSD 在2017年前后出现、S3在2020年12月实现一致性、以及2024年末 S3 支持 compare-and-swap 后才真正打开。 持久状态放在对象存储,热数据上浮到 NVMe 和 DRAM,系统不再需要独立的共识层;Eskildsen 的运行原则很直接:“我不想让状态存在两个系统里。”

  • Cursor 和 Notion 提供了早期商业验证:Cursor 用1-2周完成迁移,并将成本削减95%。 当 Notion 需要降低俄勒冈州跨区域约14毫秒的公共交换路径延迟时,Turbopuffer 花约5,000美元购买暗光纤,承担出口流量费用并调优 TCP。Eskildsen 认为,买还是自建如今“已经不太是我们能不能做,而是我们有没有时间做”。

  • Agent 式检索把查询量从一次 RAG 查找扩展为单个 Agent 发起的多路并发搜索,迫使数据库重新适配经济模型。 随着客户并行执行语义、全文和正则搜索,Turbopuffer 正将查询价格下调约5倍。主持人的总结是“所有工作负载都是混合型的”,而 Eskildsen 转述 Sualeh 的说法:检索就是“缓存算力”。

  • Turbopuffer 能实现盈利,部分原因在于早期基础设施账单直接压在 Eskildsen 的信用卡上,迫使团队在获得机构融资前就从第一性原理出发优化成本。 当前定价仍由存储、写入和查询组成,只是中间靠“胶带和唾沫”勉强粘着;部署方式则覆盖 SaaS、单租户和客户 VPC。Eskildsen 与 Lachy Groom 的融资约定同样反常:如果年底前仍看不到产品市场匹配,“我们就把所有钱都退还给你。”

  • 短期上行空间来自从向量搜索拓展到全文搜索,并将数据集规模做到1,000亿条,同时不丢掉创业公司的专注度。 ANN v3 在约40毫秒 P50、200毫秒 P99 下搜索1,000亿个向量;ANN v4 已在推进,ANN v5 正在规划,FTS v3 的功能也在逐步发布。更长期的查询计划可能包括 OLAP、日志、时序和图,但 Eskildsen 认为最可能后悔的是“试图做得太多”。

  • Turbopuffer 的执行模式依赖极其挑剔的招聘,而不是不断堆积员工数量。 每名候选人起初都默认被拒,除非面试官愿意“举起双拳”争取录用;团队寻找的 P99 工程师,能够识别餐巾纸算式与现实之间10倍的差距,再让软件向物理极限靠拢。

摘要 · 为研究而整理的核心内容

1. AI 缺失的那一层,是可搜索的外部记忆

  • Eskildsen 将今天的 Turbopuffer 严格定义为搜索引擎:提供向量搜索和全文搜索,明显需要更多能力的工作负载可能应该交给其他系统。更大的目标,是成为非结构化数据的搜索引擎,而不是又一个套着 AI 标签的通用数据库。

  • 他的前提始于压缩极限:模型可以吸收“数艾字节、数艾字节”的训练数据,并编码出理解世界的推理方式,但“我们不可能把所有东西都压缩进几TB的权重里”。因此,AI 需要一个外部系统,以“完整保真度和真实性”保存知识。

  • Eskildsen 认为,一家大型数据库公司需要满足3个前提:最终触达每家公司的工作负载;竞争对手难以改造复制的存储架构;以及一条路径,能够实现客户可能对已存储数据提出的几乎所有查询计划。Oracle 抓住了一个时代,约15年后,Snowflake 和 Databricks 又抓住了另一个时代——或许更久;把大型数据集接入 AI,可能定义下一个时代。

2. 一个30,000美元的功能,先于客户暴露了市场

  • 在 Shopify 近10年的经历,让 Eskildsen 学会在极端流量下扩展数据库,包括应对接近每秒100万次请求的事件。最让他头疼的系统,是2015年前后的自托管 Elasticsearch:项目受其限制,而要暴露 Shopify 所需的 Lucene 行为也非常困难。

  • 离开 Shopify 后,他实践所谓的“天使工程”,在朋友的公司做约3个月的项目,包括 Readwise、Replicate 和 Causal。在 Readwise,他名义上的工作是改进 Postgres——“归根结底就是调 autovacuum”——直到 ChatGPT 时刻到来,团队开始考虑用文章嵌入做推荐。

  • 原型的效果好得近乎令人不安:Readwise 一位联合创始人的推荐结果里出现了关于生孩子的文章,而 Eskildsen 当时还不知道这件事。但 Readwise 整套基础设施每月成本约5,000美元,为这个单一功能嵌入并索引文章,算下来却接近30,000美元。

  • Readwise 暂停了这个功能,等成本下降,但 Eskildsen 无法放下:“这件事一直萦绕着我。”他手头唯一的市场数据是:如果成本降到十分之一,公司很可能已经上线该功能。于是他开始研究向量索引和云基础设施,而不是先炮制一套宏观叙事。

3. 对象存储成了数据库本身,而不只是备份

  • Eskildsen 的餐巾纸算式是:把持久数据放进对象存储,将活跃部分拉入 NVMe SSD,再把最热的一小部分提升到 DRAM。在他的简化例子里,S3 中1TB每月成本约200美元,可能只有5%-10%需要驻留 NVMe,更少部分需要 DRAM,从而大幅降低“膨胀”存储数据的成本。

  • 代价也很明确:每次写入可能需要几百毫秒,第一次查询可能要半秒。因此,Turbopuffer 不适合高事务型工作负载;Eskildsen 后来表示,其写入延迟约100毫秒。他从未假设搜索是纯粹的读密集型场景,因为 Readwise 的内容变化可能带来比实际搜索更多的写入。

  • 他的第一个向量设计几乎是刻意保持原始:把聚类元数据放在 clusters.json 文件里,每个聚类存放在独立对象中,先取回最近的聚类,再在本地计算邻居。这样大约只需要两次存储往返,而不是一长串相互依赖的读取。

  • 更深层的系统原则,是在轮次之间尽量少做决策、同时发起大规模并发:一次并行发送大约1,000个 S3 请求,处理结果后再重复,且最多不超过约3轮。Eskildsen 认为,以这种方式使用时,NVMe 的带宽可以接近 DRAM,只相差一个较小倍数;对象存储则可以把网卡跑满。

4. 3次云基础设施升级,让这套架构成为可能

  • 关键时间线很重要:云端 NVMe SSD 在2017年前后出现;S3 在2020年12月实现一致性;S3 则直到2024年末才获得 compare-and-swap。三者合在一起,让团队可以围绕对象存储构建数据库,而不必维护独立的底层数据库、ZooKeeper 或类似的共识层。

  • compare-and-swap 允许多个节点下载 metadata.json,修改后仅在期间没有其他节点改动原文件的情况下写回;发生冲突时直接重试。Turbopuffer 刚开始时,Google Cloud Storage 已经提供了这一原语,多少带有运气成分:Eskildsen 选择 GCP,是因为 Shopify 使用它,而且他认识其加拿大团队。

  • 因此,Turbopuffer 选择了“all in”:即使关闭所有服务器,也不会丢失数据。Eskildsen 和联合创始人 Justine 更偏好这种方式,而不是双状态运行,因为他们最糟糕的值班经历,都与系统失去同步有关。当被问及为什么选择光纤而不是 ZooKeeper 时,他回答:“当然更愿意。我不想让状态存在两个系统里。”

  • Notion 是 AWS 客户,当时希望降低延迟,这一信念也因此付出了代价。由于云服务商的区域在地理上分离,俄勒冈州的流量要经过西雅图,延迟约14毫秒。于是 Turbopuffer 在俄勒冈州的 AWS 和 GCP 区域之间购买暗光纤,经由波特兰交换中心路由,成本约5,000美元;团队自行承担出口流量费用,并接受单线路设计,而行业通常会部署多条冗余线路。

5. Cursor 和 Notion 把架构转化成了产品市场匹配

  • Turbopuffer 的上线版本刻意保持简陋:Eskildsen 独自工作到夏天结束后,在一台8核机器上的 tmux 里运行一个 Rust 二进制文件。部署意味着盯着请求日志,在业务安静的时刻按下 Control-C——这是他从 Shopify 带来的规则:基础设施只有在展现出“至少一点 PMF 的迹象”后,才值得获得更多复杂度。

  • Cursor 联合创始人 Arvid 先发起了一段简短交流,内容是 QPS、成本和增长预测。后来 Sualeh 提议在太平洋时间凌晨约4:00通话,Eskildsen 从东海岸接了电话,意识到自己需要见这支团队后,便赶到旧金山;当时 Cursor 的 Postgres 正处于宕机状态,他临时建议对 autovacuum 进行调优。

  • Cursor 在接下来1-2周完成迁移,Turbopuffer 将其成本降低95%;Eskildsen 认为,这修复了 Cursor 的单用户经济模型。他招来了 Justine——“我在 Shopify 共事过的最优秀工程师”——两人在接下来1-2个月里确保数据库永远不会成为 Cursor 的问题。

  • Notion 的内部工程师此前独立画出了几乎相同的存储架构,后来才发现 Turbopuffer 已经把它做了出来。Eskildsen 对这笔采购的解释颇具启发性:AI 已经把买还是自建从“我们能不能做”变成了“我们有没有时间做”。一个表现得像团队延伸的供应商,买到的是速度。

6. 混合检索之所以能存活,是因为不同查询揭示不同事实

  • Cursor 使用自有嵌入模型,将完整代码库切块并生成嵌入;据称,这在某项特定评测中带来了25%的提升,而且在更大的代码仓库上尤其有效。它的 Agent 使用语义搜索寻找相似或功能相关的代码,同时也使用 grep;两种机制谁都没有取代谁。

  • 主持人追问“既然有 grep,RAG 是否已经死了”,最后落到更广泛的结论:工作负载本来就是混合的。语义搜索、词法搜索和正则表达式,分别回答不同问题。Eskildsen 不愿预测宏观未来——“事实证明,那是巨大的时间浪费”——而是专注于收集具体的客户案例。

  • Cursor 也把外部数据库视为一道安全边界:其私有嵌入模型提高了逆向还原的难度,文件路径经过混淆,存放在 Turbopuffer 存储桶中的客户数据则使用 Cursor 自有的加密密钥加密。Eskildsen 认同这些是任何外部数据库都应采取的合理做法,并非 Turbopuffer 特有的让步。

  • 按 Eskildsen 的回忆,Sualeh 的框架是:检索就是“缓存算力”。在某个特定时刻,模型聚焦于某个特定上下文,而搜索提供一个针对该状态定制的中间层。Eskildsen 不愿预测其价值将如何随时间变化,但当前工作负载表明,它对特定查询确实重要。

7. Agent 把一次检索变成一阵并发搜索

  • Eskildsen 将经典 RAG 与8,000 token 的上下文窗口、以及一次必须精打细算的检索联系在一起。Agent 则把搜索当作工具调用,反复查询并改变工作状态,同时由模型负责推理。

  • 架构变化在于单个用户会话内部的并发,而不只是跨用户批处理:“一个 Agent 驱动多个。”Notion 每轮往返都会发起 Eskildsen 所称的“荒谬数量”的查询;Cursor 的 Agent 也越来越并行。目标与 Turbopuffer 的内部机制一致:用大量搜索命中热数据集,同时尽量减少串行轮次。

  • 主持人提到 Cognition 会并行执行8次快速上下文搜索,并追问 Agent 如何避免把同一个请求发出8次。他们的答案是查询多样性:混合检索提供根本不同的搜索模式,而不是对同一个语义请求做表面变化。

  • 搜索次数增加后,单位经济模型也随之变化。Turbopuffer 正将查询价格下调约5倍,未来还可能进一步下调,以支持这类突发并发。Eskildsen 表示,写入量相对读取仍然极高,但如果客户全面采用 Agent 式并行,二者比例可能发生变化。

8. 透明融资强化了第一性原理经济学

  • 最初的定价“非常凭感觉”:Eskildsen 先估算物理成本,再加上一点利润。当 Cursor 的使用量加速增长时,它的账单仍低于 Turbopuffer 的 GCP 账单,于是他和 Justine 不断优化,努力在云端负债压在个人信用卡上的同时,实现约5%的利润率。

  • 这种压力帮助 Turbopuffer 实现盈利,也让其 VC “大为懊恼”。当前定价仍拆分为存储、写入和查询,但 Eskildsen 称,这是原有结构靠“胶带和唾沫”粘起来的版本,后续还会继续调整。客户可以选择 SaaS、专用单租户集群,或部署在自有 VPC 内的 BYOC。

  • 在竞争对手准备发布产品的同时,Eskildsen 进行融资。他选择 Lachy Groom,而不是数据库专业投资人,因为他可以不做准备就打电话,直接坦率沟通:如果年底前没有 PMF,“我们就把所有钱都退还给你。”面对陌生游戏,他的规则很简单:“我就把牌摊开来打。”

  • Groom 缺乏数据库专业知识,反而成了优势而非障碍:创始人和员工提供深度,Groom 则帮助寻找候选人和客户,但从不假装自己懂得更多。接受这笔支票,也意味着 Eskildsen 有意让公司“成为我人生旅程的一部分”;一旦员工和投资人依赖于公司,他就会全力投入。

9. Turbopuffer 只有在搜索赢得下一幕后才会扩大边界

  • 第一幕是向量搜索,第二幕是全文搜索。Turbopuffer 声称,在 Common Crawl 规模数据集上处理异常长、由 LLM 生成的查询时,性能优于 Lucene;同时,它也在补齐成熟词法搜索引擎用户期待的大量功能,并吸引传统搜索产品的客户迁移。

  • 即便是极短的人类查询,全文搜索仍然有价值:在 Command-K 中输入“si”,嵌入可能会把结果引向西班牙语或意大利语里表示“是”的词,而字面前缀搜索则可能找出一份以“These are all the reasons I hate Simon”开头的文档。混合搜索同时映射语义和用户的精确意图。

  • 规模是另一项短期优先任务。ANN v3 在约40毫秒 P50、200毫秒 P99 下搜索1,000亿个向量;ANN v4 正在推进,ANN v5 正在规划;全文搜索则将逐步改进,最终推进到 FTS v3。Eskildsen 还希望打造一个具备 phpMyAdmin 实用性的数据库控制台,而不是继续堆积一个运行了2年的创业公司仪表盘。

  • 长期来看,一家大型数据库必须支持聚合、连接以及几乎所有查询计划。可能的下一幕包括更简单的 OLAP、链路追踪、日志、时序和图,全部建立在 Turbopuffer 底层键值系统之上;Simon 提到一份报告称,Cursor 将约20TB从 Postgres 迁出,以推迟分片。但今天,搜索必须仍然是客户采用它的首要理由:“到年底,我们最可能后悔的,就是试图做得太多。”

Simon Hørup Eskildsen

I don't think I've said this publicly before, but I just called Lachy and was like, "Look, Lachy, if this doesn't have PMF by the end of the year, we'll just return all the money to you. You and I don't want to work on this unless it's really working. We want to give it the best shot this year, and we're really going to go for it. We're going to hire a bunch of people, and we're just going to be honest with everyone." When I don't know how to play a game, I just play with open cards. Lockey was the only person who didn't freak out. He was like, "I've never heard anyone say that before."

Alessio Fanelli

Everyone, welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.

swyx

Hello, hello. We're recording in the Kernel studio for the first time. Very excited.

Alessio Fanelli

Today we're joined by Simon Eskildsen of turbopuffer. Welcome.

Turbopuffer has really gone on a huge tear. I do have to mention that you're one of my newest members of the Danish Octopus Mafia, where a lot of legendary programmers have come out of it, like Bjarne Stroustrup, Rasmus Lerdorf, Anders Hejlsberg, and the V8 and Google Maps teams. You're mostly Canadian now, but isn't it interesting that there's such a strong Danish presence?

Simon Hørup Eskildsen

Yeah, I was writing a post not that long ago about the influences. I grew up in Denmark, right? I left when I was 18 to go to Canada to work at Shopify, and I would still say that I feel more Danish than Canadian. This is also the weird accent. I can't say "th." My wife is also Canadian.

I think one of the things in Denmark is that there's such a ruthless pragmatism, and there's also a big focus on aesthetics. People really care about what things look like. Canada has a lot of attributes, and the US has a lot of attributes, but I think there have been lots of great things to carry with me. I don't know what's in the water in Aarhus, though.

I don't know that I could be considered part of the Octopus Mafia quite yet compared to the phenomenal individuals you just mentioned. Barroso is also Danish-Canadian.

Alessio Fanelli

Okay.

swyx

I don't know where he lives now, but he's the PHP guy.

Alessio Fanelli

Yeah, and obviously Tobi Lütke moved to Canada as well. This is an interesting talent move.

swyx

I think I would love to get from you the definition of turbopuffer, because you could be a vector database, which is maybe a better word now in some circles. You could be a search engine. Let's just start there, and then we'll maybe run through the history of how you got to this point.

Simon Hørup Eskildsen

For sure. Yeah, so turbopuffer is, at this point in time, a search engine, right? We do full-text search and vector search, and that's really what we specialize in. If you're trying to do much more than that, this might not be the right place yet, but turbopuffer is all about search.

The other way that I think about it is that we can take all of the world's knowledge—all the exabytes and exabytes of data that there are—and we can use those tokens to train a model, but we can't compress all of that into a few terabytes of weights, right? We can compress into a few terabytes of weights how to reason with the world and how to make sense of the knowledge, but we have to somehow connect it to something external that actually holds that in full fidelity and truth.

That's the thing that we intend to become, right? That's a very holier-than-thou kind of phrasing, but being the search engine for unstructured data is the focus of turbopuffer at this point in time.

Let's break that down. Some people might say, "Didn't Elasticsearch already do this?" Other people might say, "Is this search on my data? Is this closer to RAG than to a public search thing?" How do you segment the different types of search?

Simon Hørup Eskildsen

The way that I generally think about this is that there's a lot of database companies, and I think if you want to build a really big database company, you need a couple of ingredients to be in the air, which only happens roughly every 15 years.

You need a new workload. You basically need the ambition that every single company on Earth is going to have data in your database multiple times. You look at a company like Oracle, right? I don't think you can find a company on Earth with a digital presence that doesn't somehow have some data in an Oracle database. I think at this point that's also true for Snowflake and Databricks, right? Fifteen years later—or even more than that—there's not a company on Earth that doesn't directly or indirectly consume Snowflake or Databricks, or any of the big analytics databases.

I think we're in that kind of moment now, right? I don't think you're going to find a company over the next few years that doesn't directly or indirectly have all their data available for search and connected to AI. You need that new workload. You need something to be happening where there's a new workload that causes that to happen, and that new workload is connecting very large amounts of data to AI.

The second thing you need to build a big database company is some new underlying change in the storage architecture that isn't possible with the databases that have come before you. If you look at Snowflake and Databricks, commoditized massive fleets of HDDs—that was not possible in the '90s, right? It just wasn't in the air, so we didn't build these systems. S3 and so on wasn't around.

I think the architecture that's now possible, and that wasn't possible 15 years ago, is to go all in on NVMe SSDs. It requires a particular type of architecture for the database that is difficult to retrofit onto the databases that are already there, including the ones you just mentioned.

The second thing is to go all in on object storage, more so than we could have done 15 years ago. We don't have a consensus layer; we don't really have anything. In fact, you could turn off all the servers that turbopuffer has, and we would not lose any data because we're completely all in on object storage. This means that our architecture is just so simple.

The third thing you need to do to build a big database company is that, over time, you have to implement more or less every query plan on the data. What that means is that you can't just get stuck in "This is the one thing that a database does." It has to be ever-evolving, because when someone has data in the database, they eventually expect to be able to ask it more or less every question. You have to do that to get the storage architecture to the limit of what it's capable of. Those are the 3 conditions.

swyx

I just wanted to get a little bit of the motivation. You left Shopify, where you were a principal engineer, an infrastructure guy. You also had Kernel Labs inside Shopify, right? Then you consulted for Readwise, and that kind of gave you the idea. I want you to tell that story. Maybe you've told it before, but just introduce people to the new workload and the aha moment for turbopuffer.

Simon Hørup Eskildsen

For sure. Yeah, I spent almost a decade at Shopify. I was on the infrastructure team from the fairly early days, around 2013. At the time, it felt like it was growing so quickly, and all the metrics were doubling year on year. Compared to what companies are contending with today, it was very cute growth. Some companies are seeing that month over month. Of course, Shopify has been compounding for a very long time now.

I spent a decade doing that, and the majority of it was just making sure the site was up today and making sure it was up a year from now. A lot of that was really just the Kardashians driving very large amounts of data to Shopify as they were rotating through all the merch and building out their businesses. We just needed to make sure we could handle that, right? Sometimes these were events with 1 million requests per second.

We had our own data centers back in the day, and we were moving to the cloud. There was so much sharding work and all of that that we were doing. I spent a decade just scaling databases, because that's fundamentally what's most difficult to scale about these sites.

The database that was the most difficult for me to scale during that time, and the most aggravating to be on call for, was Elasticsearch. It was very difficult to deal with, and I saw a lot of projects that were just being held back in their ambition by using it. I mean, self-hosted—self-hosted because this was 2015, right? So it's a very particular vintage. It's probably better at a lot of these things now.

It was difficult to contend with, and I just think about it: it's an inverted index. It should be good at these kinds of queries and do all of this, and we often couldn't get it to do exactly what we needed it to do, or basically get Lucene to expose raw Lucene for what we needed it to do.

Simon Hørup Eskildsen

That was something we did on the side and panic-scaled when we needed to, but it wasn't a particular focus of mine. So I left, and when I left, I wasn't sure exactly what I wanted to do. I'd spent about a decade inside the same company; I'd grown up there. I started working there when I was 18.

Speaker 2

You only do Rails.

Simon Hørup Eskildsen

Yeah, Rails.

Speaker 2

He's a Rails guy.

Simon Hørup Eskildsen

I love Rails. So good.

Speaker 2

We all wish we could still work in Rails.

Simon Hørup Eskildsen

I know.

Speaker 2

I know, but I tried learning Ruby. It's just too much—too many options to do the same thing. It's me.

Alessio Fanelli

You know, there's a way to do it.

Simon Hørup Eskildsen

I love it. I don't know that I would use it now, given Cloud Code, Cursor, and everything, but still, I guess if I'm just sitting down and writing a teaspoonful of code, that's how I think.

But anyway, I left, and I wasn't sure. I talked to a couple of companies, and I thought, "I need to see a little bit more of the world here to know what I'm going to focus on next." So what I decided was that I was going to—I called it "angel engineering"—hop around in my friends' companies in 3-month increments and just help them out with something. I vested a bit of equity and solved some interesting infrastructure problems.

I worked with a bunch of companies at the time. Readwise was one of them, Replicate was one of them, and Causal—I don't know if you've tried this, but it's a spreadsheet engine where you can do distributions. They sold recently. We even used that in FP&A at Turbopuffer. So, a bunch of companies like this, and it was super fun.

When the ChatGPT moment happened, I was with Readwise for a stint. We were preparing for the Reader launch, which is where you queue articles and read them later. I was just getting their Postgres up to snuff, which basically boiled down to tuning autovacuum. I was doing that, and then this happened, and we thought, "Maybe we should build a little recommendation engine and some features to try to hook in the LLMs." They weren't that good yet, but it was clear there was something there.

So I built a small recommendation engine. We thought, "Okay, let's take the articles that you recently read, embed all the articles, and then do recommendations." It was good enough that when I ran it on one of the co-founders of Readwise, I found that it was recommending articles about having a child. I thought, "Oh my God, I didn't know they were having a child." I wasn't sure what to do with that information, but the recommendation engine was good enough to suggest articles about it. The recommendations actually worked really well.

But this was a company that was spending maybe 5 grand a month in total on all of its infrastructure.

Speaker 2

[Gasps]

Simon Hørup Eskildsen

When I did the napkin math on running the embeddings of all the articles, putting them into a vector index, and putting it in production, it was going to be 30 grand a month. That just wasn't tenable. Readwise is a proudly bootstrapped company, and paying 30 grand for infrastructure for 1 feature versus 5 grand for everything else just wasn't tenable.

So it went into the bucket of, "This is useful, it's pretty good, but let's return to it when the cost comes down."

swyx

Did you say it grows by feature? So, for 5 to 30, is that by the number of articles? What's the scaling factor?

Simon Hørup Eskildsen

It scales by the number of articles that you embed.

Simon Hørup Eskildsen

But what I meant by that is 5 grand for all the other infrastructure—Heroku dynos, Postgres, and everything else—

swyx

Then the storage is 30?

Simon Hørup Eskildsen

Yeah, and then 30 grand for 1 feature, which is figuring out what other articles are related to this one. So it was just too much to power everything. Their budget would have been maybe a few thousand dollars, which still would have been a lot. We put it in the bucket of, "Okay, we're going to do that later. We'll wait for the cost to come down."

And that haunted me. I couldn't stop thinking about it. I thought, "Okay, there's clearly some latent demand here. If the cost had been a tenth, they would have shipped it." This was really the only data point that I had. I didn't go out and talk to anyone else.

So I started reading. I couldn't help myself. I didn't know what a vector index was, and I barely knew how to generate the vectors. There was a lot of hype about this in early 2023. There was a lot of hype about vector databases; they were raising a lot of money. I didn't know anything about it. I was trying these little models and fine-tuning them. I was just trying to get a lay of the land.

I have a GitHub repository called Napkin Math. On Napkin Math, there are rows of numbers like, "This is how much bandwidth—this is how many—you can do 25 GB/s on average to DRAM. You can do 5 GB/s of writes to an SSD." All of these numbers, right? S3, how much bandwidth you can drive per connection.

I was sitting down thinking, "Why hasn't everyone built a database where you just put everything on object storage? Then you pop it into NVMe when you use the data, and you pop it into DRAM if you're querying it a lot." It seemed fairly obvious.

The only real downside to going all in on object storage is that every write will take a couple hundred milliseconds of latency. But from there, it's really all upside. You do the first query, and it takes half a second.

It occurred to me that the architecture is really good for that. It's really good for object storage, and it's really good for NVMe SSDs. You couldn't have done that 10 years ago, going back to what we were talking about before.

You really have to build a database where you have as few round trips as possible. This is how CPUs work today. It's how NVMe SSDs work. It's how S3 works: you want to have a very large number of outstanding requests.

Basically, you go to S3 and make 1,000 requests to ask for data in 1 round trip. You wait for that, make a new decision, do it again, and try to do that a maximum of 3 times. But no databases were designed that way.

With NVMe SSDs, you can drive bandwidth within a very low multiple of DRAM bandwidth if you use them that way. The same is true with S3. You can fully max out the network card, which generally isn't maxed out, and get very good bandwidth. But no one had built a database like that.

So I thought, "Can't you just take all the vectors, plot them in the proverbial coordinate system, get the clusters, put a file on S3 called clusters.json, and then put another file there for every cluster—cluster-1.json, cluster-2.json?" That's 2 round trips. You get the clusters, find the closest clusters, and then download the cluster files for the closest N.

Your nearest neighbors locally.

Simon Hørup Eskildsen

Yes. And then you would build this file. It's ultra-simplistic, but it's not far off from what the first version of Turbopuffer was. Why hasn't anyone done that?

In that moment, from a workload perspective, you're thinking this is going to be a read-heavy thing because they're doing recommendations. Is the fact that writes are so expensive now? Or with AI, are you actually not writing that much?

Simon Hørup Eskildsen

At that point, I hadn't really thought too much about it. Well, no, actually, it was always clear to me that there were going to be a lot of writes because, at Shopify, the search clusters were doing, I don't know, tens or hundreds of QPS, right? You usually have to have a human sit and type in, but we did—I don't know how many updates there were per second. I'm sure it was in the millions, right, into the cluster.

So I always knew there was a 10:1 to 100:1 read-to-write ratio. In the Readwise use case, there would probably be a lot fewer reads than writes, too. There was just a lot of churn in the amount of stuff going through versus the number of queries.

I wasn't thinking too much about that. I was mostly thinking about the fundamentally cheapest way to build a database in the cloud today, using the primitives that were available. And this is it.

You have 1 machine and, let's say, 1 TB of data in S3. You pay $200 a month for that, and maybe 5% to 10% of that data needs to be in NVMe SSDs, with less than that in DRAM. You're paying very little to inflate the data.

swyx

By the way, when you say no one else has done that, would you consider Neon to be on a similar path in terms of being S3-first and separating compute and storage?

Simon Hørup Eskildsen

Yeah, I think what I meant by that is just building a completely new database. I don't know if we were the first. I had just looked at Napkin Math and thought, "This seems really obvious." So I'm sure 100 people came up with it at the same time. It's like the light bulb and every invention ever, right? It was just in the air.

I think Neon was first to it, and they're trying to retrofit it onto Postgres. They built this whole architecture where you have an in-memory layer and then sort of mmap back to S3. I think it was very novel at the time to do it for OLTP, but I hadn't seen a database that was truly all in—not retrofitting it.

Simon Hørup Eskildsen

A database built purely for this. No consensus layer, even using compare-and-swap on object storage to do consensus. I hadn't seen anyone go that all in. I'm sure there's someone who did that before us. I don't know. I was just looking at the napkin math.

swyx

And when you say “consensus layer,” are you strongly relying on S3's strong consistency? You are. Okay. So that is your consistency layer?

Simon Hørup Eskildsen

It is the consistency layer. I think this is also something that most people don't realize: S3 only became consistent in December 2020.

swyx

I remember this coming out during COVID, and people were like, “Oh, it was just a free upgrade.”

Simon Hørup Eskildsen

Yeah.

swyx

They just announced it. “We have consistency, guys.” And everyone was like, “Okay, cool.”

Simon Hørup Eskildsen

I'm sure they had it in production for a while. They were probably just like, “It's done,” and people were like, “Okay, cool.” But that's a big moment, right?

swyx

NVMe SSDs were also not in the cloud until around 2017, right? You just sort of had NVMe SSDs in 2017, and people were like, “Okay, cool. There's 1 SKU that does this. Whatever.” It takes a few years.

Simon Hørup Eskildsen

And then S3 became consistent in 2020. So now you don't have to have this big foundation database, or ZooKeeper, or whatever, sitting there contending with the keys. That's what Snowflake and others have slowly been forgoing.

swyx

Slowly been foregone.

Simon Hørup Eskildsen

Exactly. Just gone, right? So you push it to the—whatever, however many hundreds of people they have working on S3—and it's solved. Compare-and-swap wasn't in S3 at that point in time.

swyx

By the way, I don't know what that is, so maybe you want to explain it.

Simon Hørup Eskildsen

Yes. Compare-and-swap is basically this: imagine that you have a database, and it might be really nice to have a file called metadata.json. That file could say things like, “These keys are here, and this file means that.” There's a lot of metadata that you have to operate in a database, but that's the simplest way to do it.

Now you might have a lot of servers that want to change the metadata. They might have written a file and want the metadata to contain that file. If you have 100 nodes contending with this metadata.json, compare-and-swap allows you to download the file, make the modifications, and then write it only if it hasn't changed while you were making the modification. If it has changed, you retry. You just have these retry loops.

If you have 100 nodes doing that, it's going to be really slow, but it will converge over time. That primitive wasn't available in S3 until late 2024, but it was available in GCP.

The real story is certainly not that I sat down and big-brained it and said, “Okay, we're going to start on GCS. S3 is going to get it later.” It really wasn't that. We got really lucky. We started on GCP because Shopify ran on GCP, so that was the platform I was most familiar with. I knew the Canadian team there because I'd worked with them at Shopify, so it was natural for us to start there.

When we started building the database, we really thought we had to build a consensus layer, like having ZooKeeper or something to do this. But then we discovered compare-and-swap. I was like, “Oh, we can kick the can. We'll just do metadata in JSON. It's fine. It's probably fine.” We just kept kicking the can until we had very strong conviction in the idea.

Then we hinged the company on the fact that S3 would probably get this. It started getting really painful in mid-2024 because we were closing deals with Notion, which was running in AWS. We were like, “Trust us. You really want us to run this in GCP?” And they were like, “No, I don't know about that. We're running everything in AWS.”

The latency across the clouds was so large, and we had so much conviction that we bought dark fiber between the AWS and GCP regions in Oregon, at the internet exchange. GCP was like, “We've never seen a startup do this. What's going on here?” We were just like, “No, we don't want to do this.” We were tuning TCP windows—everything—to get the latency down because we had such high conviction in not doing a metadata layer on S3.

So those were the 3 conditions: compare-and-swap to do metadata, which wasn't in S3 until late 2024; S3 becoming consistent, which didn't happen until December 2020; and NVMe SSDs, which didn't land in the cloud until 2017.

Speaker 2

In some ways, it's a very big cloud success story that you were able to put this all together. But doing things like buying dark fiber is actually something I've never heard of.

Simon Hørup Eskildsen

It's very common when you're a big company, right? You like connecting your own data center or whatever. But if you're buying in Ashburn, Virginia—US East—the GCP and AWS data centers are within 1 millisecond of each other on the public exchanges.

In Oregon, uniquely, the GCP data center sits a couple hundred kilometers east of Portland, while the AWS region sits in Portland. The network exchange they go through is in Seattle, so it's a full 14 milliseconds or something like that. We were like, “Okay, we have to go through an exchange in Portland.”

Alessio Fanelli

That's cool. Can you imagine talking to the GCP rep and it's like, "No, we're going to buy because we know we're going to turn. We're going to turn from you guys and go to AWS in like 6 months. But in the meantime, we'll do this."

Simon Hørup Eskildsen

I mean, like they, you know—

Alessio Fanelli

This workload still runs on GCP for what it's worth, right? Because it was so reliable. It was never about moving off GCP. It was honestly just about giving Notion the latency they deserved. We didn't want them to have to care about any of this.

Alessio Fanelli

Yeah, whatever needs to be done.

Alessio Fanelli

Yeah, I got it. And you'd rather do this than run your ZooKeeper, I guess?

Simon Hørup Eskildsen

Way rather. It doesn't have state. I don't want state in 2 systems. All of that was informed by Justine, my co-founder, and me having been on call for so long. The worst outages are the ones where you have state in multiple places that's not syncing up.

It really came from a very pure source of pain: imagining what we would be okay being woken up at 3:00 a.m. about. Having something in ZooKeeper was not one of them.

Speaker 2

You're talking to a company like Notion. Do they care, or do they just care about the latency?

Simon Hørup Eskildsen

They just cared about latency.

Speaker 2

The latency costs, that's it?

Simon Hørup Eskildsen

They just cared about latency, right? We absorbed the cost. We were like, “We have high conviction in this. At some point, we can move them to AWS.” So we thought, “We'll buy the fiber. It doesn't matter.”

It's $5,000, and usually when you buy fiber, you buy multiple lines. We were like, “We can only afford 1.” But we would just test it to make sure that when it went over the public internet, it was smooth. So we did a lot of that.

Speaker 2

Yeah, whatever needs to be done. And what were the actual workloads? Because when you think about AI, 14 milliseconds really doesn't matter in the scheme of a model generation.

Simon Hørup Eskildsen

This workload still runs on GCP, for what it's worth, because it was so reliable. It was never about moving off GCP. It was honestly just about giving Notion the latency they deserved. We didn't want them to have to care about any of this.

They were also like, “Egress is going to be bad.” I was like, “Okay, screw it. We're just going to VPC-peer with you in AWS. We'll eat the cost.”

Speaker 2

Which is—I mean, Notion is a database company. They could have done this themselves. They do a lot of database engineering themselves. How do you even get in the door? Just talk through that.

Simon Hørup Eskildsen

The last time I was in San Francisco, I was talking to one of the engineers who was one of our champions at Notion. They were just trying to make sure that the per-user cost matched the economics they needed.

Speaker 2

Uh-huh.

Simon Hørup Eskildsen

The way I think about it is, I have to earn a return on whatever the cloud charges me, and then my customers have to earn a return on that. It's very simple, right? There has to be gross margin all the way up, and that's how you build the product.

So our customers have to make the right set of trade-offs that Turbopuffer makes, and if they're happy with that, that's great.

Speaker 2

Do you feel like you're competing with build internally versus buy, or buy versus buy?

Simon Hørup Eskildsen

Yeah, sorry. This was all to build up to your question. One of the Notion engineers told me that they'd sat down and probably drawn out on a napkin, “Why hasn't anyone built this?” Then they saw Turbopuffer and were like, “Well, it's literally that.”

Simon Hørup Eskildsen

AI has also changed the buy-versus-build equation. It’s not really about, “Can we build it?” It’s about, “Do we have time to build it?” And I think they felt like, okay, if this is a team that can do that and feels enough like an extension of our team, then we can go a lot faster, which would be very, very good for them.

They put us through the test, right? We had some very, very long nights to do that POC, and they were really our second big customer after Cursor, which also involved a lot of late nights, right?

Speaker 2

Yeah, should we go into that story? The sort of Cursor story? They credit you a lot for working very closely with them. I just want to hear it. I’ve heard this story from Sualeh’s point of view, but I’m curious what it looks like from your side.

Simon Hørup Eskildsen

I actually haven’t heard it from Sualeh’s point of view, so maybe you can now cross-reference it. The way that I remember it was that the day after we launched—which was just, you know, I’d worked the whole summer on the first version—Justine wasn’t part of it yet because I didn’t tell anyone that summer that I was working on this. I was just locked in on building it, because it’s very easy otherwise to confuse talking about something with actually doing it. I thought, “I’m not going to do that. I’m just going to do the thing.”

I launched it, and at this point Turbopuffer was a Rust binary running on a single 8-core machine in a tmux instance. Deploying it was like looking at the request log and then Command-C-ing it, or Control-C-ing it, and saying, “Okay, there’s no request. Let’s upgrade the binary.” It was literally the scrappiest thing you could imagine. It was on purpose, because at Shopify we did that all the time. We ran things in tmux all the time to begin with, before something had at least the inkling of product-market fit. It was like, “Okay, is anyone going to hear about this?”

One of the Cursor co-founders, Arvid, reached out. The Cursor team are all IOI/IMO contenders, right? They just speak in bullet points and facts. It was this amazing email exchange: “This is how many QPS we have. This is what we’re paying. This is where we’re going.” We were just conversing in bullet points.

I tried to get a call with them a few times, but they were really riding the PMF bull in late 2023. One time, Sualeh emailed me at—I think it was 4:00 a.m. Pacific time—saying, “Hey, are you open for a call now?” I’m on the East Coast, and it was 7:00 a.m., so I said, “Yeah, great, sure, whatever.” We started talking, and I didn’t know anything about sales. Something just compelled me: I had to go see this team. There was something there.

So I went to San Francisco and went to their office. The way that I remember it is that Postgres was down when I showed up at the office. Did Sualeh tell you this?

Speaker 2

No.

Simon Hørup Eskildsen

Okay. Postgres was down, and it was like they were distracting me with that. I was trying my best to see if I could help in any way. I knew a little bit about databases. Back to tuning autovacuum: “I think you have to tune autovacuum, Sualeh.” We talked about that, and then that evening we talked about what it would look like if they worked with us.

I just said, “Look, we’re all in. We’ll do whatever you tell us.” They migrated everything over the next week or two, and we reduced our costs by 95%, which I think kind of fixed their per-user economics. It solved a lot of other things.

This was also when I asked Justine to come on as my co-founder. She was the best engineer I ever worked with at Shopify. She lived 2 blocks away, and we were just like, “Okay, we’re going to get this done.” And we did.

We helped them migrate, and we worked like hell over the next month or two to make sure that we were never an issue. That was the Cursor story.

Speaker 2

Is code a different workload from normal text? Is it just text? Is it the same thing?

Simon Hørup Eskildsen

Yeah, Cursor’s workload is basically that they embed the entire codebase, right? They chunk it up in whatever way they do. They have their own embedding model, which they’ve been public about, and they’ve found on their evals that there’s one particular workload where it’s a 25% improvement. They have a bunch of blog posts about it.

I think it works best on larger codebases, but they’ve trained their own embedding model to do this. If you use the Cursor agent, you’ll see it do searches. They’ve also been public about how they’ve post-trained their model to be very good at semantic search as well.

That’s how they use it. It’s very good at queries like, “Can you find me other code that’s similar to this?” or “Can you find me code that does this?” They also use grep—

Alessio Fanelli

Yeah, of course. It’s been a big topic of discussion: is RAG dead because grep? You know.

Speaker 0

We see demand—

You need semantic search in every part, yes.

Speaker 0

We see demand, and I like case studies. I don’t like just doing thought pieces on where this is going and trying to be all macroeconomic about AI. That’s turned out to be a giant waste of time, because no one can really predict any of this.

I just collect case studies. Cursor has done a great job talking about what they’re doing, and I hope some of the other coding labs that use turbopuffer will do the same. It does seem to make a difference for particular queries. We can also do text, and we can also do regex.

I should also say that Cursor’s security posture with turbopuffer is exceptional, right? They have their own embedding model, which makes it very difficult to reverse-engineer. They obfuscate the file paths. It’s very difficult to learn anything about a codebase by looking at it.

The other thing they do is encrypt it with their encryption keys in turbopuffer’s bucket. It’s really, really well designed.

Alessio Fanelli

Is this extra stuff they did to work with you because you’re not part of Cursor?

Speaker 0

Exactly.

Alessio Fanelli

And this is just best practice when working with any database, not just you guys?

Okay, yeah, that makes sense. I think, for me, the learning is that all workloads are hybrid. You want the semantic, you want the text, you want the regex, you want SQL. I don’t know, but it’s silly to be all in on one particular query pattern.

Speaker 0

I really like the way that Sualeh at Cursor talks about it, although I’m going to butcher it here. I’m a database scalability person. I don’t know anything about training models other than what the internet tells me.

The way he describes it is just like cache compute, right? You have a point in time where you’re looking at some particular context, focused on some chunk, and you say, “This is the layer of the neural net at this point in time.” That seems fundamentally really useful—to do cache compute like that.

I’m not sure how the value of that will change over time, but there seems to be a lot of value in it.

swyx

Maybe talk a bit about the evolution of the workload. Even search, maybe 2 years ago, was 1 search at the start of an LLM query to build the context. Now you have agentic search, however you want to call it, where the model is both writing and changing the code, and it’s searching it again later.

What are some of the new types of workloads, or changes you’ve had to make to your architecture for them?

Speaker 0

I think you’re right. When I think of RAG, I think, “Hey, there’s an 8,000-token context window, and you better make it count.” Search was a way to do that.

Now, everything is moving toward the agent. Just let the agent do its thing, right? Back to the thing before: the LLM is very good at reasoning with the data, and so we’re just a tool call, right? That’s increasingly what we see our customers doing.

What we’re seeing more demand for from our customers now is a lot of concurrency, right? Notion does a ridiculous number of queries in every round trip just because they can. When I use the Cursor agent, I also see them doing more concurrency than I’ve ever seen before.

Similar to how we designed the database to drive as much concurrency in every round trip as possible, that’s also what the agents are doing. That’s new. It means there’s an enormous number of queries all at once to the dataset while it’s warm, in as few turns as possible.

Can I clarify 1 thing on that?

Speaker 0

Yes.

Alessio Fanelli

Are they batching multiple users, or is 1 user driving multiple queries?

Speaker 0

1 user driving multiple—1 agent driving multiple.

Parallel searching a bunch of things.

Speaker 0

Exactly. Yeah, yeah.

Cognition also did this for the fast-context things, like 8 in parallel at once.

Speaker 0

Yes.

Alessio Fanelli

An interesting problem is, well, how do you make sure you have enough diversity so you’re not making the same request 8 times?

I think that’s probably also where the hybrid comes in, because that’s another way to diversify.

Speaker 0

It’s a completely different way to do the search. That’s a big change, right? Before, it was really just 1 call, and then the LLM took however many seconds to return. But now we see an enormous number of queries. We’ve tried to reduce query pricing. This is probably the first time I’m saying that, but query pricing is being reduced by 5×, and we’ll probably try to reduce it even more to accommodate these workloads of doing very large amounts of queries.

That’s 1 thing that’s changed. I think the write-to-read ratio is still very high, right? There’s still an enormous amount of writes per read, but we’re probably starting to see that change if people really lean into this pattern.

swyx

Can we talk a little bit about the pricing? I’m curious, because traditionally a database would charge on storage, but now you have token generation that is so expensive, where the actual value of a good search query is much higher because they’re saving inference time down the line. How do you structure that? What are people receptive to on the other side, too?

Speaker 0

Yeah, the turbopuffer pricing in the beginning was just very simple. The pricing on these search engines before turbopuffer was very serverful, right? It was like, “Here’s the VM, here’s the per-hour cost.” Great. I just sat down with a piece of paper and said, “If turbopuffer is really good, this is probably what it would cost with a little bit of margin.” That was the first pricing of turbopuffer. I just sat down and was like, “Okay, this is probably the storage and whatever,” on a piece of paper.

swyx

It was vibe pricing.

Speaker 0

It was very vibe-priced, and I got it wrong.

swyx

Oh.

Speaker 0

Well, I didn’t get it wrong, but turbopuffer wasn’t at-first-principles pricing, right? When Cursor came on turbopuffer, I didn’t know any VCs. I didn’t know anything about raising money or anything like that. I just saw that my GCP bill was a lot higher than the Cursor bill. Justin and I were just like, “Well, we have to optimize it.”

To the chagrin of the VCs now, it means that we’re profitable because we had so much pricing pressure in the beginning, because it was running on my credit card. Justin and I had spent tens of thousands of dollars on compute bills, spinning off the company, very bad Canadian lawyers, and things to get all of this done because we just didn’t know. If you’re steeped in San Francisco, you just know: “Okay, you go out and raise a pre-seed round.” I never heard the word “pre-seed” at this point in time.

swyx

You had Cursor, you had Notion, and you had no funding.

Speaker 0

With Cursor, we had no funding. By the time we had Notion, Lachy was here. So it was really just—we vibe-priced it 100% from first principles, but it was not performing at first principles. We did everything we could to optimize it in the beginning so that at least we could have a 5% margin or something.

I wasn’t freaking out because Cursor’s bill was also going like this as they were growing. So my liability and my credit limit were actively calling my bank. He was like, “I need a bigger credit limit.”

Anyway, that was the beginning. The pricing was storage, writes, and query, right? The pricing we have today is basically just that pricing with duct tape and spit to try to approach a margin on the physical underlying hardware. This year, you’re going to see more and more pricing changes from us.

swyx

How much does stuff like VPC peering matter? You’re working in AWS land, where egress is charged and all that.

Speaker 0

We probably don’t. We have an enterprise plan that just has a base fee because we haven’t had time to figure out SKU pricing for all of this. You can run turbopuffer either in SaaS, right? That’s what Cursor does. You can run it in a single-tenant cluster, so it’s just you. That’s what Notion does. And then you can run it in BYOC, where everything is inside the customer’s VPC. That’s what, for example, Entropic does.

swyx

What I’m hearing is that this is probably the best CRO job for somebody who can come in and help you with this.

Speaker 0

turbopuffer hired—I don’t know what number this was—but we had a full-time CFO as the 12th hire or something at turbopuffer. I hear a lot of companies, and I don’t know how they do it. They have 100 employees and not a CFO. Having a CFO is like—

swyx

You’re out of business, man. You know?

Speaker 0

It’s so good. Money Mike just handles the money and a lot of the business stuff. He came in and helped with a lot of the operational side of the business. So, COO-CFO, somewhere in between.

swyx

Just a quick mention of Lachy, because I’m curious. I’ve met Lachy, and he’s obviously a very good investor in Physical Intelligence. Call it a generalist super angel, right? He invests in everything. I always wonder: is there something appealing about focusing on developer tooling and focusing on databases, going, “I’ve invested for 20 years in databases,” versus being a Lachy, where he can maybe connect you to all the customers that you need?

Speaker 0

This is an excellent question. No one’s asked me this. Why Lachy? There were a couple of people we were talking to at the time, and when we were raising, we were almost a little—we were a bit distressed because 1 of our peers had just launched something that was very similar to turbopuffer.

Someone gave me the advice at the time: just choose the person where you feel like you can pick up the phone, not prepare anything, and be completely honest. I don’t think I’ve said this publicly before, but I just called Lachy and was like, “Look, Lachy, if this doesn’t have PMF by the end of the year, we’ll just return all the money to you. I just don’t want to work on this unless it’s really working. So we want to give it the best shot this year, and we’re really going to go for it. We’re going to hire a bunch of people, and we’re just going to be honest with everyone.”

When I don’t know how to play a game, I just play with open cards. Lachy was the only person who didn’t freak out. He was like, “I’ve never heard anyone say that before.”

I didn’t even know what a seed or pre-seed round was, probably even at this time. I was just very honest with him. I asked him, “Lachy, have you ever invested in a database company?” He was just like, “No.” At the time, I was like, “Am I dumb?” But I think there was something that really drew me to Lachy. He’s so authentic and honest, and I just felt like I could say everything openly. That was a perfect match at the time, and honestly, it still is. He was just like, “Okay, that’s great. This is the most honest, ridiculous thing I’ve ever heard anyone say to me.”

Alessio Fanelli

A competitor launch? This may not work out?

Speaker 0

It was more just: if this doesn’t work out, I’m going to close up shop by the end of the year, right? I don’t know. Maybe it’s common. I don’t know. He told me it was uncommon. I don’t know. That’s why we chose him.

He’s been phenomenal. The other people we were talking to at the time were database experts. They knew a lot about databases, and Lachy didn’t. This turned out to be a phenomenal asset, right? Justin and I know a lot about databases. The people we hire know a lot about databases. What we needed was someone who didn’t know a lot about databases, didn’t pretend to know a lot about databases, and just wanted to help us with candidates and customers. And he did.

I have a list of the investors I have a relationship with, and Lachy has performed excellently in the number of sub-bullets of what we can attribute back to him. Just absolutely incredible. When people talk about no ego and just the best thing for the founder, I don’t think anyone—even my lawyer—is like, “Yeah, Lachy is the most friendly person you will find.”

Simon Hørup Eskildsen

Okay, this is the most glowing recommendation I’ve ever heard.

Speaker 0

He deserves it. He’s very special.

Alessio Fanelli

Yeah. Okay, amazing. Since you mentioned candidates, maybe we can talk about team building, especially in SF. It feels like it’s easier to start a company than to join a company. I’m curious about your experience, especially not being in SF full-time and doing something that is very low-level and technical.

Speaker 0

Yeah, joining versus starting. I never thought that I would be a founder. Turbopuffer started as a blog post, then it became a project, then it almost accidentally became a company, and now it feels like it’s becoming a bigger company. That was never the intention. The intentions were very pure. It was just, “Why hasn’t anyone done this?” And, “I want to be the first person to do it.”

I think some founders have this idea: “I could never work for anyone else.”

Simon Hørup Eskildsen

I really don't feel that way. I want to see this happen, and I want to see it happen with some people that I really enjoy working with. I want to have fun doing it. This has all felt very natural in that sense. So it was never a question of joining versus founding. It was just that this found me at the right moment.

Alessio Fanelli

Well, I think there's an argument that you should have joined Cursor, right? So I'm curious how you evaluated, “Okay, I should actually go raise money and make this a company,” versus, “This is a company that's growing like crazy, and it's an interesting technical problem. I should just build it within Cursor.” Then they don't have to encrypt all this stuff, and they don't have to obfuscate things. Was that on your mind at all?

Simon Hørup Eskildsen

Before taking the small check from Lachy, I did have a hard look at myself in the mirror: “Okay, do I really want to do this?” Because if I take the money, I really have to do it, right? The way I think about it is that you kind of need to be up enough to want to go all the way. That was the conversation where I was like, “Okay, this is going to be part of my life's journey: to build this company and do it in the best way that I possibly can.” If I ask people to join me and ask people to get on the cap table, then I have an ultimate responsibility to give it everything.

I don't think it occurs to me that everyone takes it that seriously, and maybe I take it too seriously. I don't know. But that was a very intentional moment, and then it was very clear: “Okay, I'm going to do this, and I'm going to give it everything.”

swyx

A lot of people don't take you this seriously. But—

Let's talk about this concept of the P99 engineer. People are 10x-ing, everyone's saying maybe engineers are out of a job. I don't know, but you definitely see a P99 engineer, and I was wondering if you'd talk about it.

Simon Hørup Eskildsen

Yeah, so the P99 engineer was just a term that we started using internally to talk about candidates and talk about how we wanted to build the company. Everyone else is like, “We want a talent-dense company.” I think that's almost become trite at this point. What I credit the Cursor founders a lot with is that they just arrived there from first principles: “We just need a talent-dense team.”

I think I've seen some teams that weren't talent-dense in a counterfactual run, which, if you've been in a large company, you will just see. It will logically happen at a large company. That was super important to me and Justin, and it's very difficult to maintain. So we needed wording for it.

I have a document called “Traits of the P99 Engineer.” It's a bullet-point list, and I look at that list after every single interview that I do and in every single recap that we do. Every recap ends with some version of: “I'm going to reject this candidate completely, regardless of what the discourse was, because I want to see people fight for this person.”

The default should not be, “We're going to hire this person.” The default should be, “We're definitely not hiring this person.” If everyone is like, “Maybe, throw a punch,” then this is not the right—

Do you ever feel like, if there's one, there must be at least one champion who's like, “Yes, I will put my career on the line for this”? I see what you mean. “Career on the line” may be a better way to say it.

Simon Hørup Eskildsen

Yeah, I would say so. Someone needs to have both fists up and be like, “I'd fight.” Right? And if one person says that, then okay, let's do it, right?

Yeah.

Simon Hørup Eskildsen

It doesn't have to be absolutely everyone, right? The interviews are always designed so that you're checking for different attributes. If someone is knocking it out of the park in every single attribute, that's fairly rare. But that's really important.

The traits of the P99 engineer—there are lots of them. There's also the traits of the P999 engineer and the P9999 engineer. This is a long list.

Alessio Fanelli

Okay.

Simon Hørup Eskildsen

I'll give you some samples of what we look for. I think the P99 engineer has some history of having bent their trajectory or something to their will, right? Some moment where they just made the computer do what it needed to do. There's something like that, and it will occur to them at some point in their career, hopefully multiple times.

Alessio Fanelli

Give me an example of one of your engineers that—

Simon Hørup Eskildsen

I'll give an example. We launched this thing called ANN v3. We're also working on v4 and v5 right now, but ANN v3 can search 100 billion vectors with a P50 of around 40 milliseconds and a P99 of 200 milliseconds. Maybe other people have done this. I'm sure Google and others have done this, but we haven't seen anyone, at least not in a public, consumable SaaS, that can do this.

That was an engineer—the chief architect of Turbopuffer, Nathan—who more or less just bent the software. It was not capable of this, and he just made it capable for a very particular workload in a 6-to-8-week period with the help of a lot of the team. There have been numerous examples of that at Turbopuffer, but that's really bending the software and x86 to your will. It was incredible to watch. You want to see some moments like that.

Isn't that P999?

Simon Hørup Eskildsen

I think—

What's it called, P999? That was only in 2019.

Simon Hørup Eskildsen

So that is too high for P999. Nathan is—Nathan is like, “Yeah, there's a lot of nines after that P.” I think that's one trait.

Another trait is that the P99 spends a lot of time looking at maps. Generally, it's their preferred UX. They just love looking at maps. Have you ever seen someone who just sits on their phone and scrolls around in a map? Or do you not look at maps a lot? You guys don't look at maps?

I guess I'm not feeling that. I know, but—

Simon Hørup Eskildsen

You just disqualified yourself. What about trains? Do you like trains?

I mean, they're—

Simon Hørup Eskildsen

Not enough.

Yeah, okay.

Simon Hørup Eskildsen

This is just my nice autism, is what I call it.

I love looking at maps. It's my preferred UX, and I like lots of—

Like a lot of random places?

Simon Hørup Eskildsen

So, like, you know—

Yes, okay. There you go.

Simon Hørup Eskildsen

So instead of random places, how do you explore the maps? No, it's just a joke.

Unless you're just obsessed by something and you like studying a thing.

Simon Hørup Eskildsen

The origin of this was that, at some point, I read an interview with some IOI gold medalist, and the question was, “What do you do in your spare time?” The answer was just, “I like looking at maps.” I was like, “I feel so seen.” I just love scrolling around. It's like, “Oh, Canada is so big. Where's Baffin Island?” I don't know, and I love it.

Anyway, one trait of the P99 is that they're obsessive. You'll find traits of that. We do multiple interviews at Turbopuffer that just try to screen for some of these things. There are lots of others, but these are the kinds of traits that we look for.

swyx

I'll tell you, some people listen for some of my DevRel stuff. I do think about DevRel as maps. You draw a map for people. Maps show you what is commonly agreed to be the geographical features, what a boundary is, and they also show you what is not there.

I think a lot of developer tools companies try to tell you they can do everything. But let's be real: your 3 landmarks are here, here, and here. Everyone comes here, here, and here. You draw a map, and then you draw a journey through the map, and to me, that's what developer relations looks like. So I do think about things that way.

Simon Hørup Eskildsen

I think the P99 thinks in trade-offs, right? The P99 is very clear about, “Hey, Turbopuffer, you can't run a high-transaction workload on Turbopuffer, right? The write latency is 100 milliseconds.” That's a clear trade-off.

I think the P99 is very good at articulating the trade-offs in every decision, which is exactly what the map is in your case, right?

swyx

Yeah, it's my world. It's my world.

How do you reconcile some of these things when you're saying you bend the world—the computer—versus the trade-offs? Sometimes it's like, “Well, these are the trade-offs,” but the P999 is like, actually, there's not a real trade-off because we can make something that nobody has ever made before and actually make it work.

Simon Hørup Eskildsen

The way I think about bending your trajectory to your will is, if you sit down and do the napkin math, you're just like, “Okay, if I have 100 machines, they have this many terabytes of disk, they have this bandwidth, whatever,” right? You sit down and do the high-school napkin math on how many QPS we should be able to drive to it, similar to how I did the vibe pricing, right?

If you can sit down and do that, and then you observe the real system and see, “Oh, we're off by 10x,” bending your trajectory to your will is just making the software get closer and closer to that first-principles line. The P99 might even be able to cross the line by finding even more optimizations than from first principles.

So bending the software to your will is about that. A 100-millisecond P99 to S3—I mean, now you're talking about someone really high-agency who goes to Seattle, finds the S3 team, and is like, “How are we going to make this 10?” It's not quite what we talk about, right? But, yeah.

Alessio Fanelli

What’s the future of Turbopuffer?

Simon Hørup Eskildsen

Turbopuffer started out—Act 1 of Turbopuffer was vector search. That’s all we did to begin with. Act 2 of Turbopuffer is and was full-text search. Turbopuffer today has a fairly state-of-the-art full-text search engine. We beat Lucene on some queries, in particular very long queries that we’ve optimized for because those are the text-search queries we see today. They’re generated by LLMs, they’re augmented by LLMs, and we see them on web-scale datasets, right? Like someone searching for a very long text string on all of Common Crawl. We’ve beaten Lucene on some of those benchmarks, and we expect to continue to beat Lucene on more and more queries.

That’s the performance and scale. Turbopuffer does phenomenally now at full-text search performance and scale. What we work on now is more and more features for full-text search. People expect a lot of features with full-text search, and full-text search is still very valuable, right? If you go in and you press Command-K and you search for “si,” an embedding-based search might be like, “Oh, this is something agreeable,” because that’s “yes”—that’s “sí” in Spanish, right?

Speaker 2

And it works in Italian, too.

Simon Hørup Eskildsen

But in full-text search, that’s the prefix of maybe a document like, you know, “These are all the reasons I hate Simon,” right? That’s a completely different thing. So that augmentation to how the human brain works on mapping data to a user is very important, but it’s a lot of features. That feature growth is what we’re firmly focused on. You will see us adding to the changelog every month—more and more full-text search features.

We’re fully compatible, and we’re seeing people move from some of the traditional search engines onto Turbopuffer for that. That’s a big focus of Turbopuffer this year. The other focus of Turbopuffer this year is scale. We’re seeing more and more companies that want to search basically Common Crawl-level types of datasets, both internally at companies and externally, and query, like, 100 billion vectors or 100 billion documents at once. This is tricky, and we want to make it cheaper and faster. That’s a big focus for Turbopuffer this year.

We just released ANN v3, which we talked about before. We’re working on ANN v4, and we’re also already planning what we’re going to do with ANN v5, right? On full-text search, we’re working on a lot of features. Many of these features will be FTS v3, but it will roll out incrementally. Those are some of the really big features.

The other thing is our dashboard. Have any of you ever logged into the Turbopuffer dashboard? There’s not very much there. It almost looks like a founder 2 years ago just sat down and wrote enough of a dashboard that there was at least something there, and then other people just sort of added stuff on for the next 2—the following 2 years—and then, at some point, the scale and other things had to catch up.

But adding things like, “I want phpMyAdmin back.” Do you guys remember? I guess it was so good, right? I think that software-hardware integration between the console, the dashboard, and the database itself—I’m really excited for that. There are lots of other things that are going to come out in the next 2 years. We talked a bit about some pricing and things like that, but those would be some of the big hitters right now.

Speaker 2

You talked about Acts 2, 4, and 5. I mean, I just have to ask: Yes, this is stuff that you’re working on this year, but I’m sure in your mind you already have the next phase that you’re already thinking about.

Simon Hørup Eskildsen

Yes.

Speaker 2

Act 4?

Simon Hørup Eskildsen

Yeah.

Speaker 2

Act 5?

Simon Hørup Eskildsen

What I’ll say about the other candidates is, you don’t have to decide yet. But, you know, I’ll just say that if you want to build a big database company, the database over time has to implement more or less every query plan. When you have your data in a database, you expect it to, over time, not just search, but also—hey, I want to aggregate this column, I want to join this data—all of that.

But when you’re a startup, your only mode is really just focus. So you have to lay out the action and not get overeager. I think we’ve seen some of our peers get very overeager and overextend themselves. What I keep telling the team—I was just having breakfast this morning with our CTO and chief architect, and we were talking about what we’re most likely to regret at the end of the year—is having tried to do too much.

Act 3 candidates could be a bunch of simpler OLAP queries. It could be leaning ourselves a little bit more into seeing some people who want to do traces and logging and things like that—some very simple use cases. It could be that. It could be maybe some time series. Some people are trying to do that, right? There are lots of different things that you can do with Turbopuffer.

But for now, if you’re trying to do something other than search on Turbopuffer as the primary use case, you probably shouldn’t. We see some customers that are like, “Oh, Cursor moved, like, 20 terabytes of Postgres data into Turbopuffer because it’s there. It works, and these particular query plans we know work well.” So they just moved it all to defer sharding.

We look for patterns like that in what future acts of Turbopuffer are going to be before firmly doubling down on them. But today, if you’re using Turbopuffer, it should be because search is very important to you. We might do a lot of auxiliary queries to that, but that should not be the main reason to go to Turbopuffer at this point in time.

Speaker 2

Yeah. You didn’t mention one thing I was looking for: graph-type queries, like graph-database graph queries. Can you basically trivially replicate this with what you already have?

Simon Hørup Eskildsen

We see some people doing that, right? Because—

Speaker 2

You have parallel queries, and it’s the same thing.

Simon Hørup Eskildsen

Exactly. So we see some people doing that, right? Under the hood, Turbopuffer is just a KV, right? Then we expose things on top of it. So we are seeing people do that.

I think our roadmap is very much just the database that connects AI to a very large amount of data. That’s the path. To do that in the right order—which is what a good startup is around—is figuring out what the order to do things in is. Our customers are P99, and they will tell us what they care most about next. Some of them are doing graphs now. If they need more graph-database features, they’ll be banging on our door, and we’ll prioritize accordingly.

Alessio Fanelli

Tea. All right, give us the tea. This is Yabukita Kamairicha from The Green Tea Shop.

Simon Hørup Eskildsen

That’s right. We were just talking beforehand about caffeine, I think. Especially when I’m on a trip like this to San Francisco, I consume a lot of caffeine. But this is my preferred caffeine. It’s this green tea. I have an Airtable with 200 teas that I’ve tried over time, over the past 15 years, and this one is my favorite.

When you drink tea, there are like 6 different types of tea. I like green tea in particular. I generally prefer Chinese green tea, and I don’t really like Japanese green tea, but this little prefecture somewhere in Japan has specialized in— it’s Japanese, but doing it the Chinese way—and it’s just phenomenal.

The interesting thing about the tea world is that you can find this particular tea. There are probably hundreds of places that sell it, but they all go to a different family, right, on whatever mountain they have these Camellia sinensis bushes on. This Japanese woman in Toronto from The Green Tea Shop—I don’t know, she just found a really good family, because that’s the best one.

The best time of year to get this is in a few months, when they do the spring harvest. Now it’s kind of old. I just love the spring for the fresh tea. So I hope you enjoy it, but it’s not the right time of year. It’s out of season.

Alessio Fanelli

Yeah. I actually didn’t even know tea had seasons. This is how unsophisticated I am. But I think that it ties in with loving maps, being obsessed, and being P99 in everything that you do. Yeah, but that’s great. Awesome. Well, as we were saying, we have instant hot water at Kernel, so any tea lover can come by.

Simon Eskildsen

I have a little tea kit where I bring a little thermometer—a ThermoWorks thermometer. Last Friday, when we did demos, I had this thing where, if there weren’t enough demos, I filled the remaining time talking about something completely ridiculous as an incentive for people to actually demo. Last time I spent 20 minutes walking through my Airtable and going through my entire tea travel kit, including the temperature monitor.

Because, yeah, you will show up. There’s only a boiler; you can’t get it to the right temperature. You need this at 80°. Anyway, sorry.

Yeah, we have an electric kettle with the temperature thing at home.

swyx

I would watch this. You should start a company YouTube channel, but it doesn’t have anything about search. It just has tea. On the other hand, I don’t think I could talk, but something that I started doing—do you two know Sam Lambert of Coding at Scale?

Of course.

swyx

Very outspoken guy.

I love the guy, and just last week we went on X Live and sat and shot the shit for an hour.

swyx

I think we'll probably…

Do that again. Yeah, so probably come up there. I don't know what we'll call it. Maybe P99 Live or the P99 pod or something like that.

swyx

P pod. [laughter] Cool. Well, thank you so much for your time here. I know you have to go, but this has been a blast, and you're clearly very passionate and charismatic. So I bet you'll get some P99 engineers on this podcast.

Yeah. Thank you so much for having me. It was a pleasure.

RAG之后的检索:混合搜索、Agent与数据库设计——Turbopuffer 的 Simon Eskildsen — 文字稿与摘要 | BidClub