[BidClub_]
Latent Space · · 61 min

Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer

Alessio FanelliswyxSimon Hørup Eskildsen

YouTube
TL;DR
  • Turbopuffer is betting that AI creates a once-in-15-years database opening because models can reason over knowledge but cannot store it all in full fidelity. Simon Hørup Eskildsen’s category thesis is that every company will connect large datasets to AI, creating demand for an external source of truth: “We can’t compress all of that into a few terabytes of weights.”

  • The company began with a brutally specific cost gap: Readwise spent about $5,000 monthly on its entire infrastructure, while one useful recommendation feature projected at $30,000. Eskildsen inferred that it would have shipped at one-tenth the price, then designed an object-storage-first search engine from napkin math rather than broad customer research. “That haunted me.”

  • Turbopuffer’s architectural window opened only after cloud NVMe SSDs arrived around 2017, S3 became consistent in December 2020, and compare-and-swap reached S3 in late 2024. Durable state lives in object storage, hot data rises into NVMe and DRAM, and there is no separate consensus system; Eskildsen’s operating principle is blunt: “I don’t want state in two systems.”

  • Cursor and Notion supplied the early commercial proof, with Cursor migrating in one or two weeks and cutting its cost by 95%. When Notion needed lower latency across Oregon’s roughly 14-millisecond public-exchange path, Turbopuffer bought about $5,000 of dark fiber, absorbed egress and tuned TCP. The buy-versus-build shift, Eskildsen argues, is now “not really about can we build it? It’s about do we have time to build it?”

  • Agentic retrieval is expanding query volume from one RAG lookup into many concurrent searches by one agent, forcing database economics to adjust. Turbopuffer is reducing query pricing roughly 5× as customers parallelize semantic, full-text and regex searches against warm datasets. The hosts’ synthesis was that “all workloads are hybrid,” while Eskildsen relays Sualeh’s framing of retrieval as “cache compute.”

  • The business reached profitability partly because early infrastructure bills sat on Eskildsen’s credit card and forced first-principles optimization before institutional funding. Pricing remains storage-plus-writes-plus-queries with “duct tape and spit,” while deployment spans SaaS, single-tenant and customer-VPC configurations. His fundraising compact with Lachy Groom was equally unusual: if product-market fit failed to appear by year-end, “we’ll just return all the money to you.”

  • Near-term upside rests on moving from vector search into full-text search and 100-billion-item datasets without losing startup focus. ANN v3 searches 100 billion vectors at roughly 40-millisecond P50 and 200-millisecond P99; ANN v4 is underway, ANN v5 is being planned, and FTS v3 features are rolling out incrementally. Longer-term query plans may include OLAP, logging, time series and graphs, but Eskildsen says the likeliest regret is “having tried to do too much.”

  • Turbopuffer’s execution model depends on unusually selective hiring rather than headcount accumulation. Every candidate begins as a rejection unless an interviewer is prepared to “have both fists up” and fight for the hire; the sought-after P99 engineer can identify a 10× gap between napkin math and reality, then bend the software toward the physical limit.

Digest · the substance, structured for research

1. AI’s missing layer is a searchable external memory

  • Eskildsen defines Turbopuffer narrowly today as a search engine: it provides vector search and full-text search, while workloads requiring substantially more may belong elsewhere. The larger ambition is to become the search engine for unstructured data rather than another generic database wearing an AI label.

  • His premise starts with compression limits: models can absorb “exabytes and exabytes” of training data and encode ways to reason about the world, but “we can’t compress all of that into a few terabytes of weights.” AI therefore needs an external system holding knowledge “in full fidelity and truth.”

  • Eskildsen sees three prerequisites for a major database company: a workload that eventually touches every company, a storage architecture incumbents cannot easily retrofit, and a path toward implementing nearly every query plan customers might ask of stored data. Oracle captured one era; Snowflake and Databricks captured another roughly 15 years later—or more; connecting large datasets to AI could define the next.

2. A $30,000 feature exposed the market before customers did

  • Nearly a decade at Shopify taught Eskildsen to scale databases under extreme traffic, including events approaching one million requests per second. The most aggravating system to operate was self-hosted Elasticsearch circa 2015: projects were constrained by it, and exposing the Lucene behavior Shopify needed proved difficult.

  • After leaving, he practiced what he called “angel engineering,” taking roughly three-month assignments at friends’ companies including Readwise, Replicate and Causal. At Readwise, his ostensible job was improving Postgres—“basically boils down to tuning autovacuum”—when the ChatGPT moment suggested embedding articles for recommendations.

  • The prototype worked almost uncomfortably well: recommendations for one Readwise co-founder surfaced articles about having a child before Eskildsen knew the news. Yet Readwise’s entire infrastructure cost around $5,000 monthly, while embedding and indexing its articles for this single feature penciled out near $30,000.

  • Readwise shelved the feature until costs fell, but Eskildsen could not: “That haunted me.” His only market datum was that the company likely would have shipped at one-tenth the cost, so he began learning vector indexes and cloud primitives instead of manufacturing a broad macro thesis.

3. Object storage became the database, not merely its backup

  • Eskildsen’s napkin math suggested putting durable data in object storage, pulling active portions into NVMe SSDs and promoting the hottest subset into DRAM. In his simplified example, a terabyte in S3 cost about $200 monthly, perhaps 5%–10% needed NVMe residency, and still less needed DRAM—dramatically reducing the cost of “inflating” stored data.

  • The tradeoff is explicit: every write may take a couple hundred milliseconds, and a first query might take half a second. Turbopuffer is therefore unsuitable for high-transaction workloads; later Eskildsen states its write latency is around 100 milliseconds. He never assumed search was purely read-heavy, noting that Readwise could create more writes through content churn than actual searches.

  • His first vector design was almost deliberately primitive: store cluster metadata in a clusters.json file, keep each cluster in its own object, fetch the nearest clusters, then calculate neighbors locally. That creates roughly two storage round trips rather than a long sequence of dependent reads.

  • The deeper systems principle is massive concurrency with few decisions between rounds: issue perhaps 1,000 S3 requests together, process the result, then repeat no more than about three times. Used this way, Eskildsen argues, NVMe can approach DRAM bandwidth within a low multiple, while object storage can saturate the network card.

4. Three cloud upgrades made the architecture newly possible

  • The enabling chronology matters: cloud NVMe SSDs appeared around 2017; S3 became consistent in December 2020; and S3 gained compare-and-swap only in late 2024. Together they made it possible to build a database around object storage without maintaining a separate foundation database, ZooKeeper or similar consensus layer.

  • Compare-and-swap lets many nodes download metadata.json, modify it, and write it back only if nobody changed the original meanwhile; conflicts simply retry. Google Cloud Storage already offered the primitive when Turbopuffer began, largely by luck: Eskildsen chose GCP because Shopify used it and he knew its Canadian team.

  • Turbopuffer consequently went “all in”: even turning off every server would not lose data. Eskildsen and co-founder Justine preferred this to dual-state operations because their worst on-call experiences involved systems falling out of sync. Asked why he would choose fiber over ZooKeeper, he answered, “Way rather. I don’t want state in two systems.”

  • That conviction became painful when Notion, an AWS customer, wanted lower latency. Oregon traffic took a roughly 14-millisecond route through Seattle because the providers’ regions were geographically separated, so Turbopuffer bought dark fiber between the AWS and GCP regions in Oregon, routing through the Portland exchange for about $5,000, absorbed egress and accepted a single line when multiple redundant circuits were customary.

5. Cursor and Notion converted architecture into product-market fit

  • The launch was intentionally skeletal: after working alone through the summer, Eskildsen shipped a Rust binary on one eight-core machine inside tmux. Deployment meant watching request logs and pressing Control-C during a quiet moment—a Shopify-derived rule that infrastructure should earn sophistication only after showing “at least the inkling of PMF.”

  • Cursor co-founder Arvid initiated a terse exchange of QPS, costs and growth projections. When Sualeh later proposed a call around 4:00 a.m. Pacific, Eskildsen accepted from the East Coast, sensed he needed to meet the team, and arrived in San Francisco while Cursor’s Postgres was down—prompting an impromptu recommendation to tune autovacuum.

  • Cursor migrated over the following week or two, and Turbopuffer reduced its cost by 95%, which Eskildsen believes repaired its per-user economics. He recruited Justine, “the best engineer” he had worked with at Shopify, and the pair spent the next month or two ensuring the database never became Cursor’s problem.

  • Notion’s internal engineer had independently sketched essentially the same storage architecture, then discovered Turbopuffer had built it. Eskildsen’s explanation of the purchase is revealing: AI has changed buy versus build from “can we build it?” to “do we have time to build it?” A vendor that behaves like an extension of the team buys speed.

6. Hybrid retrieval survives because different queries reveal different truths

  • Cursor chunks and embeds entire codebases using its own embedding model, reportedly producing a 25% improvement on one specific evaluation and working especially well on larger repositories. Its agent uses semantic searches to find similar or functionally related code, but it also uses grep; neither mechanism eliminates the other.

  • The hosts press the “is RAG dead because grep?” argument, then land on the broader lesson that workloads are hybrid: semantic search, lexical search and regex serve different questions. Eskildsen avoids predicting the macro future—“That has turned out to be a giant waste of time”—and instead collects concrete customer case studies.

  • Cursor also treats the external database as a security boundary: its private embedding model makes reversal harder, file paths are obfuscated, and customer data in Turbopuffer’s bucket is encrypted with Cursor’s own encryption keys. Eskildsen agrees these are sound practices for any external database, not concessions unique to his company.

  • Sualeh’s framing, as Eskildsen recalls it, is that retrieval is “cache compute”: at a particular moment, the model is focused on a particular context, and search supplies an intermediate layer tailored to that state. Eskildsen will not predict how its value changes over time, but current workloads show it matters for specific queries.

7. Agents turn one retrieval into a burst of concurrent searches

  • Eskildsen associates classic RAG with an 8,000-token context window and one retrieval that had to count. Agents instead treat search as a tool call, repeatedly querying and changing their working state while the model handles reasoning.

  • The architectural shift is concurrency within one user session, not merely batching across users: “One agent driving multiple.” Notion launches what Eskildsen calls a ridiculous number of queries per round trip, Cursor’s agent is increasingly parallel, and the objective mirrors Turbopuffer’s internals—hit a warm dataset with many searches while minimizing sequential turns.

  • The hosts cite Cognition doing eight fast-context searches in parallel and ask how an agent avoids issuing the same request eight times. Their answer is query diversity, with hybrid retrieval supplying fundamentally different search modes rather than cosmetic variations on one semantic request.

  • More searches change unit economics. Turbopuffer is reducing query pricing about 5×, with further reductions possible, to support these bursts. Eskildsen says write volume remains extremely high relative to reads, but expects the ratio could shift if customers lean fully into agentic parallelism.

8. Open-card financing reinforced first-principles economics

  • Initial pricing was “very vibe priced”: Eskildsen estimated physical costs and added a little margin. When Cursor’s usage accelerated, its invoice remained below Turbopuffer’s GCP bill, so he and Justine optimized relentlessly to reach even a roughly 5% margin while the cloud liability expanded on his personal credit card.

  • That pressure helped make Turbopuffer profitable, “to the chagrin” of its VCs. Current pricing still decomposes into storage, writes and queries, but Eskildsen calls it the original structure held together with “duct tape and spit”; more changes are planned. Customers can choose SaaS, a dedicated single-tenant cluster, or BYOC inside their own VPC.

  • While raising amid a competing launch, Eskildsen chose Lachy Groom over database-specialist investors because he could call without preparation and speak plainly: if PMF did not arrive by year-end, “we’ll just return all the money to you.” His rule when unfamiliar with a game is simple: “I just play with open cards.”

  • Groom’s lack of database expertise became useful rather than disqualifying: the founders and employees supplied that depth, while Groom helped with candidates and customers without pretending otherwise. Accepting the check also marked Eskildsen’s deliberate commitment to make the company “part of my life’s journey” and give it everything once employees and investors depended on him.

9. Turbopuffer will broaden only after search earns the next act

  • Act One was vector search; Act Two is full-text search. Turbopuffer claims to beat Lucene on some unusually long, LLM-generated queries over Common Crawl-scale datasets, while adding the large feature surface users expect from mature lexical engines and attracting migrations from traditional search products.

  • Full text remains valuable even for tiny human queries: typing “si” into Command-K might lead embeddings toward the Spanish or Italian word for “yes,” while literal prefix search could surface a document beginning “These are all the reasons I hate Simon.” Hybrid search maps both meaning and exact user intent.

  • Scale is the other near-term priority. ANN v3 searches 100 billion vectors at roughly 40-millisecond P50 and 200-millisecond P99; ANN v4 is underway and ANN v5 is being planned, while full-text improvements will roll incrementally toward FTS v3. Eskildsen also wants a database console with the practical usefulness of phpMyAdmin rather than the startup dashboard accumulated over two years.

  • Long term, a major database must support aggregation, joins and nearly every query plan. Possible next acts include simpler OLAP, traces, logging, time series and graphs atop Turbopuffer’s underlying key-value system; Simon cites a report that Cursor moved roughly 20 terabytes from Postgres to defer sharding. Yet search must remain the primary reason to adopt it today: “What we’re most likely to regret at the end of the year is having tried to do too much.”

Simon Hørup Eskildsen

I don't think I've said this publicly before, but I just called Lachy and was like, "Look, Lachy, if this doesn't have PMF by the end of the year, we'll just return all the money to you. You and I don't want to work on this unless it's really working. We want to give it the best shot this year, and we're really going to go for it. We're going to hire a bunch of people, and we're just going to be honest with everyone." When I don't know how to play a game, I just play with open cards. Lockey was the only person who didn't freak out. He was like, "I've never heard anyone say that before."

Alessio Fanelli

Everyone, welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.

swyx

Hello, hello. We're recording in the Kernel studio for the first time. Very excited.

Alessio Fanelli

Today we're joined by Simon Eskildsen of turbopuffer. Welcome.

Turbopuffer has really gone on a huge tear. I do have to mention that you're one of my newest members of the Danish Octopus Mafia, where a lot of legendary programmers have come out of it, like Bjarne Stroustrup, Rasmus Lerdorf, Anders Hejlsberg, and the V8 and Google Maps teams. You're mostly Canadian now, but isn't it interesting that there's such a strong Danish presence?

Simon Hørup Eskildsen

Yeah, I was writing a post not that long ago about the influences. I grew up in Denmark, right? I left when I was 18 to go to Canada to work at Shopify, and I would still say that I feel more Danish than Canadian. This is also the weird accent. I can't say "th." My wife is also Canadian.

I think one of the things in Denmark is that there's such a ruthless pragmatism, and there's also a big focus on aesthetics. People really care about what things look like. Canada has a lot of attributes, and the US has a lot of attributes, but I think there have been lots of great things to carry with me. I don't know what's in the water in Aarhus, though.

I don't know that I could be considered part of the Octopus Mafia quite yet compared to the phenomenal individuals you just mentioned. Barroso is also Danish-Canadian.

Alessio Fanelli

Okay.

swyx

I don't know where he lives now, but he's the PHP guy.

Alessio Fanelli

Yeah, and obviously Tobi Lütke moved to Canada as well. This is an interesting talent move.

swyx

I think I would love to get from you the definition of turbopuffer, because you could be a vector database, which is maybe a better word now in some circles. You could be a search engine. Let's just start there, and then we'll maybe run through the history of how you got to this point.

Simon Hørup Eskildsen

For sure. Yeah, so turbopuffer is, at this point in time, a search engine, right? We do full-text search and vector search, and that's really what we specialize in. If you're trying to do much more than that, this might not be the right place yet, but turbopuffer is all about search.

The other way that I think about it is that we can take all of the world's knowledge—all the exabytes and exabytes of data that there are—and we can use those tokens to train a model, but we can't compress all of that into a few terabytes of weights, right? We can compress into a few terabytes of weights how to reason with the world and how to make sense of the knowledge, but we have to somehow connect it to something external that actually holds that in full fidelity and truth.

That's the thing that we intend to become, right? That's a very holier-than-thou kind of phrasing, but being the search engine for unstructured data is the focus of turbopuffer at this point in time.

Let's break that down. Some people might say, "Didn't Elasticsearch already do this?" Other people might say, "Is this search on my data? Is this closer to RAG than to a public search thing?" How do you segment the different types of search?

Simon Hørup Eskildsen

The way that I generally think about this is that there's a lot of database companies, and I think if you want to build a really big database company, you need a couple of ingredients to be in the air, which only happens roughly every 15 years.

You need a new workload. You basically need the ambition that every single company on Earth is going to have data in your database multiple times. You look at a company like Oracle, right? I don't think you can find a company on Earth with a digital presence that doesn't somehow have some data in an Oracle database. I think at this point that's also true for Snowflake and Databricks, right? Fifteen years later—or even more than that—there's not a company on Earth that doesn't directly or indirectly consume Snowflake or Databricks, or any of the big analytics databases.

I think we're in that kind of moment now, right? I don't think you're going to find a company over the next few years that doesn't directly or indirectly have all their data available for search and connected to AI. You need that new workload. You need something to be happening where there's a new workload that causes that to happen, and that new workload is connecting very large amounts of data to AI.

The second thing you need to build a big database company is some new underlying change in the storage architecture that isn't possible with the databases that have come before you. If you look at Snowflake and Databricks, commoditized massive fleets of HDDs—that was not possible in the '90s, right? It just wasn't in the air, so we didn't build these systems. S3 and so on wasn't around.

I think the architecture that's now possible, and that wasn't possible 15 years ago, is to go all in on NVMe SSDs. It requires a particular type of architecture for the database that is difficult to retrofit onto the databases that are already there, including the ones you just mentioned.

The second thing is to go all in on object storage, more so than we could have done 15 years ago. We don't have a consensus layer; we don't really have anything. In fact, you could turn off all the servers that turbopuffer has, and we would not lose any data because we're completely all in on object storage. This means that our architecture is just so simple.

The third thing you need to do to build a big database company is that, over time, you have to implement more or less every query plan on the data. What that means is that you can't just get stuck in "This is the one thing that a database does." It has to be ever-evolving, because when someone has data in the database, they eventually expect to be able to ask it more or less every question. You have to do that to get the storage architecture to the limit of what it's capable of. Those are the 3 conditions.

swyx

I just wanted to get a little bit of the motivation. You left Shopify, where you were a principal engineer, an infrastructure guy. You also had Kernel Labs inside Shopify, right? Then you consulted for Readwise, and that kind of gave you the idea. I want you to tell that story. Maybe you've told it before, but just introduce people to the new workload and the aha moment for turbopuffer.

Simon Hørup Eskildsen

For sure. Yeah, I spent almost a decade at Shopify. I was on the infrastructure team from the fairly early days, around 2013. At the time, it felt like it was growing so quickly, and all the metrics were doubling year on year. Compared to what companies are contending with today, it was very cute growth. Some companies are seeing that month over month. Of course, Shopify has been compounding for a very long time now.

I spent a decade doing that, and the majority of it was just making sure the site was up today and making sure it was up a year from now. A lot of that was really just the Kardashians driving very large amounts of data to Shopify as they were rotating through all the merch and building out their businesses. We just needed to make sure we could handle that, right? Sometimes these were events with 1 million requests per second.

We had our own data centers back in the day, and we were moving to the cloud. There was so much sharding work and all of that that we were doing. I spent a decade just scaling databases, because that's fundamentally what's most difficult to scale about these sites.

The database that was the most difficult for me to scale during that time, and the most aggravating to be on call for, was Elasticsearch. It was very difficult to deal with, and I saw a lot of projects that were just being held back in their ambition by using it. I mean, self-hosted—self-hosted because this was 2015, right? So it's a very particular vintage. It's probably better at a lot of these things now.

It was difficult to contend with, and I just think about it: it's an inverted index. It should be good at these kinds of queries and do all of this, and we often couldn't get it to do exactly what we needed it to do, or basically get Lucene to expose raw Lucene for what we needed it to do.

Simon Hørup Eskildsen

That was something we did on the side and panic-scaled when we needed to, but it wasn't a particular focus of mine. So I left, and when I left, I wasn't sure exactly what I wanted to do. I'd spent about a decade inside the same company; I'd grown up there. I started working there when I was 18.

Speaker 2

You only do Rails.

Simon Hørup Eskildsen

Yeah, Rails.

Speaker 2

He's a Rails guy.

Simon Hørup Eskildsen

I love Rails. So good.

Speaker 2

We all wish we could still work in Rails.

Simon Hørup Eskildsen

I know.

Speaker 2

I know, but I tried learning Ruby. It's just too much—too many options to do the same thing. It's me.

Alessio Fanelli

You know, there's a way to do it.

Simon Hørup Eskildsen

I love it. I don't know that I would use it now, given Cloud Code, Cursor, and everything, but still, I guess if I'm just sitting down and writing a teaspoonful of code, that's how I think.

But anyway, I left, and I wasn't sure. I talked to a couple of companies, and I thought, "I need to see a little bit more of the world here to know what I'm going to focus on next." So what I decided was that I was going to—I called it "angel engineering"—hop around in my friends' companies in 3-month increments and just help them out with something. I vested a bit of equity and solved some interesting infrastructure problems.

I worked with a bunch of companies at the time. Readwise was one of them, Replicate was one of them, and Causal—I don't know if you've tried this, but it's a spreadsheet engine where you can do distributions. They sold recently. We even used that in FP&A at Turbopuffer. So, a bunch of companies like this, and it was super fun.

When the ChatGPT moment happened, I was with Readwise for a stint. We were preparing for the Reader launch, which is where you queue articles and read them later. I was just getting their Postgres up to snuff, which basically boiled down to tuning autovacuum. I was doing that, and then this happened, and we thought, "Maybe we should build a little recommendation engine and some features to try to hook in the LLMs." They weren't that good yet, but it was clear there was something there.

So I built a small recommendation engine. We thought, "Okay, let's take the articles that you recently read, embed all the articles, and then do recommendations." It was good enough that when I ran it on one of the co-founders of Readwise, I found that it was recommending articles about having a child. I thought, "Oh my God, I didn't know they were having a child." I wasn't sure what to do with that information, but the recommendation engine was good enough to suggest articles about it. The recommendations actually worked really well.

But this was a company that was spending maybe 5 grand a month in total on all of its infrastructure.

Speaker 2

[Gasps]

Simon Hørup Eskildsen

When I did the napkin math on running the embeddings of all the articles, putting them into a vector index, and putting it in production, it was going to be 30 grand a month. That just wasn't tenable. Readwise is a proudly bootstrapped company, and paying 30 grand for infrastructure for 1 feature versus 5 grand for everything else just wasn't tenable.

So it went into the bucket of, "This is useful, it's pretty good, but let's return to it when the cost comes down."

swyx

Did you say it grows by feature? So, for 5 to 30, is that by the number of articles? What's the scaling factor?

Simon Hørup Eskildsen

It scales by the number of articles that you embed.

Simon Hørup Eskildsen

But what I meant by that is 5 grand for all the other infrastructure—Heroku dynos, Postgres, and everything else—

swyx

Then the storage is 30?

Simon Hørup Eskildsen

Yeah, and then 30 grand for 1 feature, which is figuring out what other articles are related to this one. So it was just too much to power everything. Their budget would have been maybe a few thousand dollars, which still would have been a lot. We put it in the bucket of, "Okay, we're going to do that later. We'll wait for the cost to come down."

And that haunted me. I couldn't stop thinking about it. I thought, "Okay, there's clearly some latent demand here. If the cost had been a tenth, they would have shipped it." This was really the only data point that I had. I didn't go out and talk to anyone else.

So I started reading. I couldn't help myself. I didn't know what a vector index was, and I barely knew how to generate the vectors. There was a lot of hype about this in early 2023. There was a lot of hype about vector databases; they were raising a lot of money. I didn't know anything about it. I was trying these little models and fine-tuning them. I was just trying to get a lay of the land.

I have a GitHub repository called Napkin Math. On Napkin Math, there are rows of numbers like, "This is how much bandwidth—this is how many—you can do 25 GB/s on average to DRAM. You can do 5 GB/s of writes to an SSD." All of these numbers, right? S3, how much bandwidth you can drive per connection.

I was sitting down thinking, "Why hasn't everyone built a database where you just put everything on object storage? Then you pop it into NVMe when you use the data, and you pop it into DRAM if you're querying it a lot." It seemed fairly obvious.

The only real downside to going all in on object storage is that every write will take a couple hundred milliseconds of latency. But from there, it's really all upside. You do the first query, and it takes half a second.

It occurred to me that the architecture is really good for that. It's really good for object storage, and it's really good for NVMe SSDs. You couldn't have done that 10 years ago, going back to what we were talking about before.

You really have to build a database where you have as few round trips as possible. This is how CPUs work today. It's how NVMe SSDs work. It's how S3 works: you want to have a very large number of outstanding requests.

Basically, you go to S3 and make 1,000 requests to ask for data in 1 round trip. You wait for that, make a new decision, do it again, and try to do that a maximum of 3 times. But no databases were designed that way.

With NVMe SSDs, you can drive bandwidth within a very low multiple of DRAM bandwidth if you use them that way. The same is true with S3. You can fully max out the network card, which generally isn't maxed out, and get very good bandwidth. But no one had built a database like that.

So I thought, "Can't you just take all the vectors, plot them in the proverbial coordinate system, get the clusters, put a file on S3 called clusters.json, and then put another file there for every cluster—cluster-1.json, cluster-2.json?" That's 2 round trips. You get the clusters, find the closest clusters, and then download the cluster files for the closest N.

Your nearest neighbors locally.

Simon Hørup Eskildsen

Yes. And then you would build this file. It's ultra-simplistic, but it's not far off from what the first version of Turbopuffer was. Why hasn't anyone done that?

In that moment, from a workload perspective, you're thinking this is going to be a read-heavy thing because they're doing recommendations. Is the fact that writes are so expensive now? Or with AI, are you actually not writing that much?

Simon Hørup Eskildsen

At that point, I hadn't really thought too much about it. Well, no, actually, it was always clear to me that there were going to be a lot of writes because, at Shopify, the search clusters were doing, I don't know, tens or hundreds of QPS, right? You usually have to have a human sit and type in, but we did—I don't know how many updates there were per second. I'm sure it was in the millions, right, into the cluster.

So I always knew there was a 10:1 to 100:1 read-to-write ratio. In the Readwise use case, there would probably be a lot fewer reads than writes, too. There was just a lot of churn in the amount of stuff going through versus the number of queries.

I wasn't thinking too much about that. I was mostly thinking about the fundamentally cheapest way to build a database in the cloud today, using the primitives that were available. And this is it.

You have 1 machine and, let's say, 1 TB of data in S3. You pay $200 a month for that, and maybe 5% to 10% of that data needs to be in NVMe SSDs, with less than that in DRAM. You're paying very little to inflate the data.

swyx

By the way, when you say no one else has done that, would you consider Neon to be on a similar path in terms of being S3-first and separating compute and storage?

Simon Hørup Eskildsen

Yeah, I think what I meant by that is just building a completely new database. I don't know if we were the first. I had just looked at Napkin Math and thought, "This seems really obvious." So I'm sure 100 people came up with it at the same time. It's like the light bulb and every invention ever, right? It was just in the air.

I think Neon was first to it, and they're trying to retrofit it onto Postgres. They built this whole architecture where you have an in-memory layer and then sort of mmap back to S3. I think it was very novel at the time to do it for OLTP, but I hadn't seen a database that was truly all in—not retrofitting it.

Simon Hørup Eskildsen

A database built purely for this. No consensus layer, even using compare-and-swap on object storage to do consensus. I hadn't seen anyone go that all in. I'm sure there's someone who did that before us. I don't know. I was just looking at the napkin math.

swyx

And when you say “consensus layer,” are you strongly relying on S3's strong consistency? You are. Okay. So that is your consistency layer?

Simon Hørup Eskildsen

It is the consistency layer. I think this is also something that most people don't realize: S3 only became consistent in December 2020.

swyx

I remember this coming out during COVID, and people were like, “Oh, it was just a free upgrade.”

Simon Hørup Eskildsen

Yeah.

swyx

They just announced it. “We have consistency, guys.” And everyone was like, “Okay, cool.”

Simon Hørup Eskildsen

I'm sure they had it in production for a while. They were probably just like, “It's done,” and people were like, “Okay, cool.” But that's a big moment, right?

swyx

NVMe SSDs were also not in the cloud until around 2017, right? You just sort of had NVMe SSDs in 2017, and people were like, “Okay, cool. There's 1 SKU that does this. Whatever.” It takes a few years.

Simon Hørup Eskildsen

And then S3 became consistent in 2020. So now you don't have to have this big foundation database, or ZooKeeper, or whatever, sitting there contending with the keys. That's what Snowflake and others have slowly been forgoing.

swyx

Slowly been foregone.

Simon Hørup Eskildsen

Exactly. Just gone, right? So you push it to the—whatever, however many hundreds of people they have working on S3—and it's solved. Compare-and-swap wasn't in S3 at that point in time.

swyx

By the way, I don't know what that is, so maybe you want to explain it.

Simon Hørup Eskildsen

Yes. Compare-and-swap is basically this: imagine that you have a database, and it might be really nice to have a file called metadata.json. That file could say things like, “These keys are here, and this file means that.” There's a lot of metadata that you have to operate in a database, but that's the simplest way to do it.

Now you might have a lot of servers that want to change the metadata. They might have written a file and want the metadata to contain that file. If you have 100 nodes contending with this metadata.json, compare-and-swap allows you to download the file, make the modifications, and then write it only if it hasn't changed while you were making the modification. If it has changed, you retry. You just have these retry loops.

If you have 100 nodes doing that, it's going to be really slow, but it will converge over time. That primitive wasn't available in S3 until late 2024, but it was available in GCP.

The real story is certainly not that I sat down and big-brained it and said, “Okay, we're going to start on GCS. S3 is going to get it later.” It really wasn't that. We got really lucky. We started on GCP because Shopify ran on GCP, so that was the platform I was most familiar with. I knew the Canadian team there because I'd worked with them at Shopify, so it was natural for us to start there.

When we started building the database, we really thought we had to build a consensus layer, like having ZooKeeper or something to do this. But then we discovered compare-and-swap. I was like, “Oh, we can kick the can. We'll just do metadata in JSON. It's fine. It's probably fine.” We just kept kicking the can until we had very strong conviction in the idea.

Then we hinged the company on the fact that S3 would probably get this. It started getting really painful in mid-2024 because we were closing deals with Notion, which was running in AWS. We were like, “Trust us. You really want us to run this in GCP?” And they were like, “No, I don't know about that. We're running everything in AWS.”

The latency across the clouds was so large, and we had so much conviction that we bought dark fiber between the AWS and GCP regions in Oregon, at the internet exchange. GCP was like, “We've never seen a startup do this. What's going on here?” We were just like, “No, we don't want to do this.” We were tuning TCP windows—everything—to get the latency down because we had such high conviction in not doing a metadata layer on S3.

So those were the 3 conditions: compare-and-swap to do metadata, which wasn't in S3 until late 2024; S3 becoming consistent, which didn't happen until December 2020; and NVMe SSDs, which didn't land in the cloud until 2017.

Speaker 2

In some ways, it's a very big cloud success story that you were able to put this all together. But doing things like buying dark fiber is actually something I've never heard of.

Simon Hørup Eskildsen

It's very common when you're a big company, right? You like connecting your own data center or whatever. But if you're buying in Ashburn, Virginia—US East—the GCP and AWS data centers are within 1 millisecond of each other on the public exchanges.

In Oregon, uniquely, the GCP data center sits a couple hundred kilometers east of Portland, while the AWS region sits in Portland. The network exchange they go through is in Seattle, so it's a full 14 milliseconds or something like that. We were like, “Okay, we have to go through an exchange in Portland.”

Alessio Fanelli

That's cool. Can you imagine talking to the GCP rep and it's like, "No, we're going to buy because we know we're going to turn. We're going to turn from you guys and go to AWS in like 6 months. But in the meantime, we'll do this."

Simon Hørup Eskildsen

I mean, like they, you know—

Alessio Fanelli

This workload still runs on GCP for what it's worth, right? Because it was so reliable. It was never about moving off GCP. It was honestly just about giving Notion the latency they deserved. We didn't want them to have to care about any of this.

Alessio Fanelli

Yeah, whatever needs to be done.

Alessio Fanelli

Yeah, I got it. And you'd rather do this than run your ZooKeeper, I guess?

Simon Hørup Eskildsen

Way rather. It doesn't have state. I don't want state in 2 systems. All of that was informed by Justine, my co-founder, and me having been on call for so long. The worst outages are the ones where you have state in multiple places that's not syncing up.

It really came from a very pure source of pain: imagining what we would be okay being woken up at 3:00 a.m. about. Having something in ZooKeeper was not one of them.

Speaker 2

You're talking to a company like Notion. Do they care, or do they just care about the latency?

Simon Hørup Eskildsen

They just cared about latency.

Speaker 2

The latency costs, that's it?

Simon Hørup Eskildsen

They just cared about latency, right? We absorbed the cost. We were like, “We have high conviction in this. At some point, we can move them to AWS.” So we thought, “We'll buy the fiber. It doesn't matter.”

It's $5,000, and usually when you buy fiber, you buy multiple lines. We were like, “We can only afford 1.” But we would just test it to make sure that when it went over the public internet, it was smooth. So we did a lot of that.

Speaker 2

Yeah, whatever needs to be done. And what were the actual workloads? Because when you think about AI, 14 milliseconds really doesn't matter in the scheme of a model generation.

Simon Hørup Eskildsen

This workload still runs on GCP, for what it's worth, because it was so reliable. It was never about moving off GCP. It was honestly just about giving Notion the latency they deserved. We didn't want them to have to care about any of this.

They were also like, “Egress is going to be bad.” I was like, “Okay, screw it. We're just going to VPC-peer with you in AWS. We'll eat the cost.”

Speaker 2

Which is—I mean, Notion is a database company. They could have done this themselves. They do a lot of database engineering themselves. How do you even get in the door? Just talk through that.

Simon Hørup Eskildsen

The last time I was in San Francisco, I was talking to one of the engineers who was one of our champions at Notion. They were just trying to make sure that the per-user cost matched the economics they needed.

Speaker 2

Uh-huh.

Simon Hørup Eskildsen

The way I think about it is, I have to earn a return on whatever the cloud charges me, and then my customers have to earn a return on that. It's very simple, right? There has to be gross margin all the way up, and that's how you build the product.

So our customers have to make the right set of trade-offs that Turbopuffer makes, and if they're happy with that, that's great.

Speaker 2

Do you feel like you're competing with build internally versus buy, or buy versus buy?

Simon Hørup Eskildsen

Yeah, sorry. This was all to build up to your question. One of the Notion engineers told me that they'd sat down and probably drawn out on a napkin, “Why hasn't anyone built this?” Then they saw Turbopuffer and were like, “Well, it's literally that.”

Simon Hørup Eskildsen

AI has also changed the buy-versus-build equation. It’s not really about, “Can we build it?” It’s about, “Do we have time to build it?” And I think they felt like, okay, if this is a team that can do that and feels enough like an extension of our team, then we can go a lot faster, which would be very, very good for them.

They put us through the test, right? We had some very, very long nights to do that POC, and they were really our second big customer after Cursor, which also involved a lot of late nights, right?

Speaker 2

Yeah, should we go into that story? The sort of Cursor story? They credit you a lot for working very closely with them. I just want to hear it. I’ve heard this story from Sualeh’s point of view, but I’m curious what it looks like from your side.

Simon Hørup Eskildsen

I actually haven’t heard it from Sualeh’s point of view, so maybe you can now cross-reference it. The way that I remember it was that the day after we launched—which was just, you know, I’d worked the whole summer on the first version—Justine wasn’t part of it yet because I didn’t tell anyone that summer that I was working on this. I was just locked in on building it, because it’s very easy otherwise to confuse talking about something with actually doing it. I thought, “I’m not going to do that. I’m just going to do the thing.”

I launched it, and at this point Turbopuffer was a Rust binary running on a single 8-core machine in a tmux instance. Deploying it was like looking at the request log and then Command-C-ing it, or Control-C-ing it, and saying, “Okay, there’s no request. Let’s upgrade the binary.” It was literally the scrappiest thing you could imagine. It was on purpose, because at Shopify we did that all the time. We ran things in tmux all the time to begin with, before something had at least the inkling of product-market fit. It was like, “Okay, is anyone going to hear about this?”

One of the Cursor co-founders, Arvid, reached out. The Cursor team are all IOI/IMO contenders, right? They just speak in bullet points and facts. It was this amazing email exchange: “This is how many QPS we have. This is what we’re paying. This is where we’re going.” We were just conversing in bullet points.

I tried to get a call with them a few times, but they were really riding the PMF bull in late 2023. One time, Sualeh emailed me at—I think it was 4:00 a.m. Pacific time—saying, “Hey, are you open for a call now?” I’m on the East Coast, and it was 7:00 a.m., so I said, “Yeah, great, sure, whatever.” We started talking, and I didn’t know anything about sales. Something just compelled me: I had to go see this team. There was something there.

So I went to San Francisco and went to their office. The way that I remember it is that Postgres was down when I showed up at the office. Did Sualeh tell you this?

Speaker 2

No.

Simon Hørup Eskildsen

Okay. Postgres was down, and it was like they were distracting me with that. I was trying my best to see if I could help in any way. I knew a little bit about databases. Back to tuning autovacuum: “I think you have to tune autovacuum, Sualeh.” We talked about that, and then that evening we talked about what it would look like if they worked with us.

I just said, “Look, we’re all in. We’ll do whatever you tell us.” They migrated everything over the next week or two, and we reduced our costs by 95%, which I think kind of fixed their per-user economics. It solved a lot of other things.

This was also when I asked Justine to come on as my co-founder. She was the best engineer I ever worked with at Shopify. She lived 2 blocks away, and we were just like, “Okay, we’re going to get this done.” And we did.

We helped them migrate, and we worked like hell over the next month or two to make sure that we were never an issue. That was the Cursor story.

Speaker 2

Is code a different workload from normal text? Is it just text? Is it the same thing?

Simon Hørup Eskildsen

Yeah, Cursor’s workload is basically that they embed the entire codebase, right? They chunk it up in whatever way they do. They have their own embedding model, which they’ve been public about, and they’ve found on their evals that there’s one particular workload where it’s a 25% improvement. They have a bunch of blog posts about it.

I think it works best on larger codebases, but they’ve trained their own embedding model to do this. If you use the Cursor agent, you’ll see it do searches. They’ve also been public about how they’ve post-trained their model to be very good at semantic search as well.

That’s how they use it. It’s very good at queries like, “Can you find me other code that’s similar to this?” or “Can you find me code that does this?” They also use grep—

Alessio Fanelli

Yeah, of course. It’s been a big topic of discussion: is RAG dead because grep? You know.

Speaker 0

We see demand—

You need semantic search in every part, yes.

Speaker 0

We see demand, and I like case studies. I don’t like just doing thought pieces on where this is going and trying to be all macroeconomic about AI. That’s turned out to be a giant waste of time, because no one can really predict any of this.

I just collect case studies. Cursor has done a great job talking about what they’re doing, and I hope some of the other coding labs that use turbopuffer will do the same. It does seem to make a difference for particular queries. We can also do text, and we can also do regex.

I should also say that Cursor’s security posture with turbopuffer is exceptional, right? They have their own embedding model, which makes it very difficult to reverse-engineer. They obfuscate the file paths. It’s very difficult to learn anything about a codebase by looking at it.

The other thing they do is encrypt it with their encryption keys in turbopuffer’s bucket. It’s really, really well designed.

Alessio Fanelli

Is this extra stuff they did to work with you because you’re not part of Cursor?

Speaker 0

Exactly.

Alessio Fanelli

And this is just best practice when working with any database, not just you guys?

Okay, yeah, that makes sense. I think, for me, the learning is that all workloads are hybrid. You want the semantic, you want the text, you want the regex, you want SQL. I don’t know, but it’s silly to be all in on one particular query pattern.

Speaker 0

I really like the way that Sualeh at Cursor talks about it, although I’m going to butcher it here. I’m a database scalability person. I don’t know anything about training models other than what the internet tells me.

The way he describes it is just like cache compute, right? You have a point in time where you’re looking at some particular context, focused on some chunk, and you say, “This is the layer of the neural net at this point in time.” That seems fundamentally really useful—to do cache compute like that.

I’m not sure how the value of that will change over time, but there seems to be a lot of value in it.

swyx

Maybe talk a bit about the evolution of the workload. Even search, maybe 2 years ago, was 1 search at the start of an LLM query to build the context. Now you have agentic search, however you want to call it, where the model is both writing and changing the code, and it’s searching it again later.

What are some of the new types of workloads, or changes you’ve had to make to your architecture for them?

Speaker 0

I think you’re right. When I think of RAG, I think, “Hey, there’s an 8,000-token context window, and you better make it count.” Search was a way to do that.

Now, everything is moving toward the agent. Just let the agent do its thing, right? Back to the thing before: the LLM is very good at reasoning with the data, and so we’re just a tool call, right? That’s increasingly what we see our customers doing.

What we’re seeing more demand for from our customers now is a lot of concurrency, right? Notion does a ridiculous number of queries in every round trip just because they can. When I use the Cursor agent, I also see them doing more concurrency than I’ve ever seen before.

Similar to how we designed the database to drive as much concurrency in every round trip as possible, that’s also what the agents are doing. That’s new. It means there’s an enormous number of queries all at once to the dataset while it’s warm, in as few turns as possible.

Can I clarify 1 thing on that?

Speaker 0

Yes.

Alessio Fanelli

Are they batching multiple users, or is 1 user driving multiple queries?

Speaker 0

1 user driving multiple—1 agent driving multiple.

Parallel searching a bunch of things.

Speaker 0

Exactly. Yeah, yeah.

Cognition also did this for the fast-context things, like 8 in parallel at once.

Speaker 0

Yes.

Alessio Fanelli

An interesting problem is, well, how do you make sure you have enough diversity so you’re not making the same request 8 times?

I think that’s probably also where the hybrid comes in, because that’s another way to diversify.

Speaker 0

It’s a completely different way to do the search. That’s a big change, right? Before, it was really just 1 call, and then the LLM took however many seconds to return. But now we see an enormous number of queries. We’ve tried to reduce query pricing. This is probably the first time I’m saying that, but query pricing is being reduced by 5×, and we’ll probably try to reduce it even more to accommodate these workloads of doing very large amounts of queries.

That’s 1 thing that’s changed. I think the write-to-read ratio is still very high, right? There’s still an enormous amount of writes per read, but we’re probably starting to see that change if people really lean into this pattern.

swyx

Can we talk a little bit about the pricing? I’m curious, because traditionally a database would charge on storage, but now you have token generation that is so expensive, where the actual value of a good search query is much higher because they’re saving inference time down the line. How do you structure that? What are people receptive to on the other side, too?

Speaker 0

Yeah, the turbopuffer pricing in the beginning was just very simple. The pricing on these search engines before turbopuffer was very serverful, right? It was like, “Here’s the VM, here’s the per-hour cost.” Great. I just sat down with a piece of paper and said, “If turbopuffer is really good, this is probably what it would cost with a little bit of margin.” That was the first pricing of turbopuffer. I just sat down and was like, “Okay, this is probably the storage and whatever,” on a piece of paper.

swyx

It was vibe pricing.

Speaker 0

It was very vibe-priced, and I got it wrong.

swyx

Oh.

Speaker 0

Well, I didn’t get it wrong, but turbopuffer wasn’t at-first-principles pricing, right? When Cursor came on turbopuffer, I didn’t know any VCs. I didn’t know anything about raising money or anything like that. I just saw that my GCP bill was a lot higher than the Cursor bill. Justin and I were just like, “Well, we have to optimize it.”

To the chagrin of the VCs now, it means that we’re profitable because we had so much pricing pressure in the beginning, because it was running on my credit card. Justin and I had spent tens of thousands of dollars on compute bills, spinning off the company, very bad Canadian lawyers, and things to get all of this done because we just didn’t know. If you’re steeped in San Francisco, you just know: “Okay, you go out and raise a pre-seed round.” I never heard the word “pre-seed” at this point in time.

swyx

You had Cursor, you had Notion, and you had no funding.

Speaker 0

With Cursor, we had no funding. By the time we had Notion, Lachy was here. So it was really just—we vibe-priced it 100% from first principles, but it was not performing at first principles. We did everything we could to optimize it in the beginning so that at least we could have a 5% margin or something.

I wasn’t freaking out because Cursor’s bill was also going like this as they were growing. So my liability and my credit limit were actively calling my bank. He was like, “I need a bigger credit limit.”

Anyway, that was the beginning. The pricing was storage, writes, and query, right? The pricing we have today is basically just that pricing with duct tape and spit to try to approach a margin on the physical underlying hardware. This year, you’re going to see more and more pricing changes from us.

swyx

How much does stuff like VPC peering matter? You’re working in AWS land, where egress is charged and all that.

Speaker 0

We probably don’t. We have an enterprise plan that just has a base fee because we haven’t had time to figure out SKU pricing for all of this. You can run turbopuffer either in SaaS, right? That’s what Cursor does. You can run it in a single-tenant cluster, so it’s just you. That’s what Notion does. And then you can run it in BYOC, where everything is inside the customer’s VPC. That’s what, for example, Entropic does.

swyx

What I’m hearing is that this is probably the best CRO job for somebody who can come in and help you with this.

Speaker 0

turbopuffer hired—I don’t know what number this was—but we had a full-time CFO as the 12th hire or something at turbopuffer. I hear a lot of companies, and I don’t know how they do it. They have 100 employees and not a CFO. Having a CFO is like—

swyx

You’re out of business, man. You know?

Speaker 0

It’s so good. Money Mike just handles the money and a lot of the business stuff. He came in and helped with a lot of the operational side of the business. So, COO-CFO, somewhere in between.

swyx

Just a quick mention of Lachy, because I’m curious. I’ve met Lachy, and he’s obviously a very good investor in Physical Intelligence. Call it a generalist super angel, right? He invests in everything. I always wonder: is there something appealing about focusing on developer tooling and focusing on databases, going, “I’ve invested for 20 years in databases,” versus being a Lachy, where he can maybe connect you to all the customers that you need?

Speaker 0

This is an excellent question. No one’s asked me this. Why Lachy? There were a couple of people we were talking to at the time, and when we were raising, we were almost a little—we were a bit distressed because 1 of our peers had just launched something that was very similar to turbopuffer.

Someone gave me the advice at the time: just choose the person where you feel like you can pick up the phone, not prepare anything, and be completely honest. I don’t think I’ve said this publicly before, but I just called Lachy and was like, “Look, Lachy, if this doesn’t have PMF by the end of the year, we’ll just return all the money to you. I just don’t want to work on this unless it’s really working. So we want to give it the best shot this year, and we’re really going to go for it. We’re going to hire a bunch of people, and we’re just going to be honest with everyone.”

When I don’t know how to play a game, I just play with open cards. Lachy was the only person who didn’t freak out. He was like, “I’ve never heard anyone say that before.”

I didn’t even know what a seed or pre-seed round was, probably even at this time. I was just very honest with him. I asked him, “Lachy, have you ever invested in a database company?” He was just like, “No.” At the time, I was like, “Am I dumb?” But I think there was something that really drew me to Lachy. He’s so authentic and honest, and I just felt like I could say everything openly. That was a perfect match at the time, and honestly, it still is. He was just like, “Okay, that’s great. This is the most honest, ridiculous thing I’ve ever heard anyone say to me.”

Alessio Fanelli

A competitor launch? This may not work out?

Speaker 0

It was more just: if this doesn’t work out, I’m going to close up shop by the end of the year, right? I don’t know. Maybe it’s common. I don’t know. He told me it was uncommon. I don’t know. That’s why we chose him.

He’s been phenomenal. The other people we were talking to at the time were database experts. They knew a lot about databases, and Lachy didn’t. This turned out to be a phenomenal asset, right? Justin and I know a lot about databases. The people we hire know a lot about databases. What we needed was someone who didn’t know a lot about databases, didn’t pretend to know a lot about databases, and just wanted to help us with candidates and customers. And he did.

I have a list of the investors I have a relationship with, and Lachy has performed excellently in the number of sub-bullets of what we can attribute back to him. Just absolutely incredible. When people talk about no ego and just the best thing for the founder, I don’t think anyone—even my lawyer—is like, “Yeah, Lachy is the most friendly person you will find.”

Simon Hørup Eskildsen

Okay, this is the most glowing recommendation I’ve ever heard.

Speaker 0

He deserves it. He’s very special.

Alessio Fanelli

Yeah. Okay, amazing. Since you mentioned candidates, maybe we can talk about team building, especially in SF. It feels like it’s easier to start a company than to join a company. I’m curious about your experience, especially not being in SF full-time and doing something that is very low-level and technical.

Speaker 0

Yeah, joining versus starting. I never thought that I would be a founder. Turbopuffer started as a blog post, then it became a project, then it almost accidentally became a company, and now it feels like it’s becoming a bigger company. That was never the intention. The intentions were very pure. It was just, “Why hasn’t anyone done this?” And, “I want to be the first person to do it.”

I think some founders have this idea: “I could never work for anyone else.”

Simon Hørup Eskildsen

I really don't feel that way. I want to see this happen, and I want to see it happen with some people that I really enjoy working with. I want to have fun doing it. This has all felt very natural in that sense. So it was never a question of joining versus founding. It was just that this found me at the right moment.

Alessio Fanelli

Well, I think there's an argument that you should have joined Cursor, right? So I'm curious how you evaluated, “Okay, I should actually go raise money and make this a company,” versus, “This is a company that's growing like crazy, and it's an interesting technical problem. I should just build it within Cursor.” Then they don't have to encrypt all this stuff, and they don't have to obfuscate things. Was that on your mind at all?

Simon Hørup Eskildsen

Before taking the small check from Lachy, I did have a hard look at myself in the mirror: “Okay, do I really want to do this?” Because if I take the money, I really have to do it, right? The way I think about it is that you kind of need to be up enough to want to go all the way. That was the conversation where I was like, “Okay, this is going to be part of my life's journey: to build this company and do it in the best way that I possibly can.” If I ask people to join me and ask people to get on the cap table, then I have an ultimate responsibility to give it everything.

I don't think it occurs to me that everyone takes it that seriously, and maybe I take it too seriously. I don't know. But that was a very intentional moment, and then it was very clear: “Okay, I'm going to do this, and I'm going to give it everything.”

swyx

A lot of people don't take you this seriously. But—

Let's talk about this concept of the P99 engineer. People are 10x-ing, everyone's saying maybe engineers are out of a job. I don't know, but you definitely see a P99 engineer, and I was wondering if you'd talk about it.

Simon Hørup Eskildsen

Yeah, so the P99 engineer was just a term that we started using internally to talk about candidates and talk about how we wanted to build the company. Everyone else is like, “We want a talent-dense company.” I think that's almost become trite at this point. What I credit the Cursor founders a lot with is that they just arrived there from first principles: “We just need a talent-dense team.”

I think I've seen some teams that weren't talent-dense in a counterfactual run, which, if you've been in a large company, you will just see. It will logically happen at a large company. That was super important to me and Justin, and it's very difficult to maintain. So we needed wording for it.

I have a document called “Traits of the P99 Engineer.” It's a bullet-point list, and I look at that list after every single interview that I do and in every single recap that we do. Every recap ends with some version of: “I'm going to reject this candidate completely, regardless of what the discourse was, because I want to see people fight for this person.”

The default should not be, “We're going to hire this person.” The default should be, “We're definitely not hiring this person.” If everyone is like, “Maybe, throw a punch,” then this is not the right—

Do you ever feel like, if there's one, there must be at least one champion who's like, “Yes, I will put my career on the line for this”? I see what you mean. “Career on the line” may be a better way to say it.

Simon Hørup Eskildsen

Yeah, I would say so. Someone needs to have both fists up and be like, “I'd fight.” Right? And if one person says that, then okay, let's do it, right?

Yeah.

Simon Hørup Eskildsen

It doesn't have to be absolutely everyone, right? The interviews are always designed so that you're checking for different attributes. If someone is knocking it out of the park in every single attribute, that's fairly rare. But that's really important.

The traits of the P99 engineer—there are lots of them. There's also the traits of the P999 engineer and the P9999 engineer. This is a long list.

Alessio Fanelli

Okay.

Simon Hørup Eskildsen

I'll give you some samples of what we look for. I think the P99 engineer has some history of having bent their trajectory or something to their will, right? Some moment where they just made the computer do what it needed to do. There's something like that, and it will occur to them at some point in their career, hopefully multiple times.

Alessio Fanelli

Give me an example of one of your engineers that—

Simon Hørup Eskildsen

I'll give an example. We launched this thing called ANN v3. We're also working on v4 and v5 right now, but ANN v3 can search 100 billion vectors with a P50 of around 40 milliseconds and a P99 of 200 milliseconds. Maybe other people have done this. I'm sure Google and others have done this, but we haven't seen anyone, at least not in a public, consumable SaaS, that can do this.

That was an engineer—the chief architect of Turbopuffer, Nathan—who more or less just bent the software. It was not capable of this, and he just made it capable for a very particular workload in a 6-to-8-week period with the help of a lot of the team. There have been numerous examples of that at Turbopuffer, but that's really bending the software and x86 to your will. It was incredible to watch. You want to see some moments like that.

Isn't that P999?

Simon Hørup Eskildsen

I think—

What's it called, P999? That was only in 2019.

Simon Hørup Eskildsen

So that is too high for P999. Nathan is—Nathan is like, “Yeah, there's a lot of nines after that P.” I think that's one trait.

Another trait is that the P99 spends a lot of time looking at maps. Generally, it's their preferred UX. They just love looking at maps. Have you ever seen someone who just sits on their phone and scrolls around in a map? Or do you not look at maps a lot? You guys don't look at maps?

I guess I'm not feeling that. I know, but—

Simon Hørup Eskildsen

You just disqualified yourself. What about trains? Do you like trains?

I mean, they're—

Simon Hørup Eskildsen

Not enough.

Yeah, okay.

Simon Hørup Eskildsen

This is just my nice autism, is what I call it.

I love looking at maps. It's my preferred UX, and I like lots of—

Like a lot of random places?

Simon Hørup Eskildsen

So, like, you know—

Yes, okay. There you go.

Simon Hørup Eskildsen

So instead of random places, how do you explore the maps? No, it's just a joke.

Unless you're just obsessed by something and you like studying a thing.

Simon Hørup Eskildsen

The origin of this was that, at some point, I read an interview with some IOI gold medalist, and the question was, “What do you do in your spare time?” The answer was just, “I like looking at maps.” I was like, “I feel so seen.” I just love scrolling around. It's like, “Oh, Canada is so big. Where's Baffin Island?” I don't know, and I love it.

Anyway, one trait of the P99 is that they're obsessive. You'll find traits of that. We do multiple interviews at Turbopuffer that just try to screen for some of these things. There are lots of others, but these are the kinds of traits that we look for.

swyx

I'll tell you, some people listen for some of my DevRel stuff. I do think about DevRel as maps. You draw a map for people. Maps show you what is commonly agreed to be the geographical features, what a boundary is, and they also show you what is not there.

I think a lot of developer tools companies try to tell you they can do everything. But let's be real: your 3 landmarks are here, here, and here. Everyone comes here, here, and here. You draw a map, and then you draw a journey through the map, and to me, that's what developer relations looks like. So I do think about things that way.

Simon Hørup Eskildsen

I think the P99 thinks in trade-offs, right? The P99 is very clear about, “Hey, Turbopuffer, you can't run a high-transaction workload on Turbopuffer, right? The write latency is 100 milliseconds.” That's a clear trade-off.

I think the P99 is very good at articulating the trade-offs in every decision, which is exactly what the map is in your case, right?

swyx

Yeah, it's my world. It's my world.

How do you reconcile some of these things when you're saying you bend the world—the computer—versus the trade-offs? Sometimes it's like, “Well, these are the trade-offs,” but the P999 is like, actually, there's not a real trade-off because we can make something that nobody has ever made before and actually make it work.

Simon Hørup Eskildsen

The way I think about bending your trajectory to your will is, if you sit down and do the napkin math, you're just like, “Okay, if I have 100 machines, they have this many terabytes of disk, they have this bandwidth, whatever,” right? You sit down and do the high-school napkin math on how many QPS we should be able to drive to it, similar to how I did the vibe pricing, right?

If you can sit down and do that, and then you observe the real system and see, “Oh, we're off by 10x,” bending your trajectory to your will is just making the software get closer and closer to that first-principles line. The P99 might even be able to cross the line by finding even more optimizations than from first principles.

So bending the software to your will is about that. A 100-millisecond P99 to S3—I mean, now you're talking about someone really high-agency who goes to Seattle, finds the S3 team, and is like, “How are we going to make this 10?” It's not quite what we talk about, right? But, yeah.

Alessio Fanelli

What’s the future of Turbopuffer?

Simon Hørup Eskildsen

Turbopuffer started out—Act 1 of Turbopuffer was vector search. That’s all we did to begin with. Act 2 of Turbopuffer is and was full-text search. Turbopuffer today has a fairly state-of-the-art full-text search engine. We beat Lucene on some queries, in particular very long queries that we’ve optimized for because those are the text-search queries we see today. They’re generated by LLMs, they’re augmented by LLMs, and we see them on web-scale datasets, right? Like someone searching for a very long text string on all of Common Crawl. We’ve beaten Lucene on some of those benchmarks, and we expect to continue to beat Lucene on more and more queries.

That’s the performance and scale. Turbopuffer does phenomenally now at full-text search performance and scale. What we work on now is more and more features for full-text search. People expect a lot of features with full-text search, and full-text search is still very valuable, right? If you go in and you press Command-K and you search for “si,” an embedding-based search might be like, “Oh, this is something agreeable,” because that’s “yes”—that’s “sí” in Spanish, right?

Speaker 2

And it works in Italian, too.

Simon Hørup Eskildsen

But in full-text search, that’s the prefix of maybe a document like, you know, “These are all the reasons I hate Simon,” right? That’s a completely different thing. So that augmentation to how the human brain works on mapping data to a user is very important, but it’s a lot of features. That feature growth is what we’re firmly focused on. You will see us adding to the changelog every month—more and more full-text search features.

We’re fully compatible, and we’re seeing people move from some of the traditional search engines onto Turbopuffer for that. That’s a big focus of Turbopuffer this year. The other focus of Turbopuffer this year is scale. We’re seeing more and more companies that want to search basically Common Crawl-level types of datasets, both internally at companies and externally, and query, like, 100 billion vectors or 100 billion documents at once. This is tricky, and we want to make it cheaper and faster. That’s a big focus for Turbopuffer this year.

We just released ANN v3, which we talked about before. We’re working on ANN v4, and we’re also already planning what we’re going to do with ANN v5, right? On full-text search, we’re working on a lot of features. Many of these features will be FTS v3, but it will roll out incrementally. Those are some of the really big features.

The other thing is our dashboard. Have any of you ever logged into the Turbopuffer dashboard? There’s not very much there. It almost looks like a founder 2 years ago just sat down and wrote enough of a dashboard that there was at least something there, and then other people just sort of added stuff on for the next 2—the following 2 years—and then, at some point, the scale and other things had to catch up.

But adding things like, “I want phpMyAdmin back.” Do you guys remember? I guess it was so good, right? I think that software-hardware integration between the console, the dashboard, and the database itself—I’m really excited for that. There are lots of other things that are going to come out in the next 2 years. We talked a bit about some pricing and things like that, but those would be some of the big hitters right now.

Speaker 2

You talked about Acts 2, 4, and 5. I mean, I just have to ask: Yes, this is stuff that you’re working on this year, but I’m sure in your mind you already have the next phase that you’re already thinking about.

Simon Hørup Eskildsen

Yes.

Speaker 2

Act 4?

Simon Hørup Eskildsen

Yeah.

Speaker 2

Act 5?

Simon Hørup Eskildsen

What I’ll say about the other candidates is, you don’t have to decide yet. But, you know, I’ll just say that if you want to build a big database company, the database over time has to implement more or less every query plan. When you have your data in a database, you expect it to, over time, not just search, but also—hey, I want to aggregate this column, I want to join this data—all of that.

But when you’re a startup, your only mode is really just focus. So you have to lay out the action and not get overeager. I think we’ve seen some of our peers get very overeager and overextend themselves. What I keep telling the team—I was just having breakfast this morning with our CTO and chief architect, and we were talking about what we’re most likely to regret at the end of the year—is having tried to do too much.

Act 3 candidates could be a bunch of simpler OLAP queries. It could be leaning ourselves a little bit more into seeing some people who want to do traces and logging and things like that—some very simple use cases. It could be that. It could be maybe some time series. Some people are trying to do that, right? There are lots of different things that you can do with Turbopuffer.

But for now, if you’re trying to do something other than search on Turbopuffer as the primary use case, you probably shouldn’t. We see some customers that are like, “Oh, Cursor moved, like, 20 terabytes of Postgres data into Turbopuffer because it’s there. It works, and these particular query plans we know work well.” So they just moved it all to defer sharding.

We look for patterns like that in what future acts of Turbopuffer are going to be before firmly doubling down on them. But today, if you’re using Turbopuffer, it should be because search is very important to you. We might do a lot of auxiliary queries to that, but that should not be the main reason to go to Turbopuffer at this point in time.

Speaker 2

Yeah. You didn’t mention one thing I was looking for: graph-type queries, like graph-database graph queries. Can you basically trivially replicate this with what you already have?

Simon Hørup Eskildsen

We see some people doing that, right? Because—

Speaker 2

You have parallel queries, and it’s the same thing.

Simon Hørup Eskildsen

Exactly. So we see some people doing that, right? Under the hood, Turbopuffer is just a KV, right? Then we expose things on top of it. So we are seeing people do that.

I think our roadmap is very much just the database that connects AI to a very large amount of data. That’s the path. To do that in the right order—which is what a good startup is around—is figuring out what the order to do things in is. Our customers are P99, and they will tell us what they care most about next. Some of them are doing graphs now. If they need more graph-database features, they’ll be banging on our door, and we’ll prioritize accordingly.

Alessio Fanelli

Tea. All right, give us the tea. This is Yabukita Kamairicha from The Green Tea Shop.

Simon Hørup Eskildsen

That’s right. We were just talking beforehand about caffeine, I think. Especially when I’m on a trip like this to San Francisco, I consume a lot of caffeine. But this is my preferred caffeine. It’s this green tea. I have an Airtable with 200 teas that I’ve tried over time, over the past 15 years, and this one is my favorite.

When you drink tea, there are like 6 different types of tea. I like green tea in particular. I generally prefer Chinese green tea, and I don’t really like Japanese green tea, but this little prefecture somewhere in Japan has specialized in— it’s Japanese, but doing it the Chinese way—and it’s just phenomenal.

The interesting thing about the tea world is that you can find this particular tea. There are probably hundreds of places that sell it, but they all go to a different family, right, on whatever mountain they have these Camellia sinensis bushes on. This Japanese woman in Toronto from The Green Tea Shop—I don’t know, she just found a really good family, because that’s the best one.

The best time of year to get this is in a few months, when they do the spring harvest. Now it’s kind of old. I just love the spring for the fresh tea. So I hope you enjoy it, but it’s not the right time of year. It’s out of season.

Alessio Fanelli

Yeah. I actually didn’t even know tea had seasons. This is how unsophisticated I am. But I think that it ties in with loving maps, being obsessed, and being P99 in everything that you do. Yeah, but that’s great. Awesome. Well, as we were saying, we have instant hot water at Kernel, so any tea lover can come by.

Simon Eskildsen

I have a little tea kit where I bring a little thermometer—a ThermoWorks thermometer. Last Friday, when we did demos, I had this thing where, if there weren’t enough demos, I filled the remaining time talking about something completely ridiculous as an incentive for people to actually demo. Last time I spent 20 minutes walking through my Airtable and going through my entire tea travel kit, including the temperature monitor.

Because, yeah, you will show up. There’s only a boiler; you can’t get it to the right temperature. You need this at 80°. Anyway, sorry.

Yeah, we have an electric kettle with the temperature thing at home.

swyx

I would watch this. You should start a company YouTube channel, but it doesn’t have anything about search. It just has tea. On the other hand, I don’t think I could talk, but something that I started doing—do you two know Sam Lambert of Coding at Scale?

Of course.

swyx

Very outspoken guy.

I love the guy, and just last week we went on X Live and sat and shot the shit for an hour.

swyx

I think we'll probably…

Do that again. Yeah, so probably come up there. I don't know what we'll call it. Maybe P99 Live or the P99 pod or something like that.

swyx

P pod. [laughter] Cool. Well, thank you so much for your time here. I know you have to go, but this has been a blast, and you're clearly very passionate and charismatic. So I bet you'll get some P99 engineers on this podcast.

Yeah. Thank you so much for having me. It was a pleasure.

Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer | BidClub