20VC: The $100 Billion AI Assistant Race: Town vs Instinct vs GrokBot | We Spend $75K Per Engineer on AI Tools | Why the AI Assistant Market Is Not a Bubble & AI Assistants Will Replace Every App on Your Phone with JD, Founder of Town
Harry StebbingsJean-Denis "JD" Grèze
- JD's core competitive admission is startling for a founder mid-fundraise: "I know what I'm building is a top-three priority at Google and Apple in the next 12 months. Not a top-10 priority." His answer is that "talking about moats is a little bit of a luxury" until you're bigger than Town is today — the incumbents "will copy their way there, but they have to copy someone who's been successful in the first place at building a mainstream product," and no product, GrokBot included, has deep mainstream product-market fit yet.
- The moat he's actually betting on is a network effect at the agent level — "no one has figured out multi-user, multi-player AI today." Town's agent-to-agent feature lets your "townie" query a coworker's townie when it lacks the answer; once a whole team runs on that, switching becomes difficult. His five-year hot take extends the logic: "I think you'll trust your agent to decide what data to share with other people without you intervening" — with models respecting boundaries that users never had to specify as hard rules.
- The unit-economics endgame worry is precise: "I have zero pricing power at the frontier." Town runs mostly on frontier models today. JD says startups hope the cost curve makes the model economics efficient in 18–24 months and says a product can be priced today for 20–30% margins in 18 months if inference costs halve every 9–12 months. The unknowable variable is whether 10%, 20%, or 30% of workloads stay frontier — and he asks whether this is what happened to Cursor, where a company can pay suppliers while competing with them.
- His market-structure read: the field has collapsed from roughly 15 startup competitors to "two or three," because "you can build now at the speed of machines, but you can only learn at the speed of humans." Feature leads that once lasted years now last two to four weeks; established players including Anthropic, Cursor, and GrokBot move at startup speed. He watches GrokBot, not Instinct — "I don't think Instinct and Town are trying to do the same thing" — and rates Apple's new Siri as good but roughly nine months behind in capability.
- Monetization discipline is the differentiator versus subsidized rivals: Town converts more than 15% of trial users to paid despite a hard gate ("you have to connect email and calendar") that loses 30% immediately, and JD says it makes over $700 per year per user today. Success is defined as "someone who pays me every month," not token usage — beta users with no price pushback burned $2,000–$4,000/month in compute, one spending about $26,000 in five months — and he's doubtful of pure ad-backed assistants: tokens are too expensive and ads pollute trajectories.
- Town's annualized AI-tooling run rate is "at least $75K per engineer" — Devon for bugs, Cursor's Composer for front-end, and a 50/50 Codex/Claude split that was "mostly Claude" five months ago. He estimates global compute spend at roughly three engineers, perhaps four, or about 1.5 engineers when compared with equity-inclusive Valley engineering costs. The ROI math is compelling: "there's gold littered everywhere in front of me" — customers demanding SSO, audit logs, and integrations to sign contracts — so he'd spend more immediately if better models existed.
- The $100B bull case is 10 million paying users, and given the choice he takes 100M consumers at $20 over 1M at $100 — because business-side NRR is elastic where personal use tops out. His specimen: a recruiting firm pays Town roughly $600/month, automates enough to take one more client worth $3,000/month, and would happily spend another $500 to take another. On adjacent bets, he's a happy ElevenLabs customer but does not claim to be bearish or rule out investing at $22B; he says the company may eventually face a good-enough ceiling and that $22B is a lot of money.
1. The pivot: from failed AI tax prep to product-market fit in a couple of weeks
- JD is blunt about the pre-Town year: "The truth is we failed at building a business that would be a great business" — an AI business-tax company with some PMF, but not enough. Three months "in the wilderness" ended on a simple question: why has no one built AI that operates out of email and calendar, where so many people run their business and life?
- The timing was the trade: "It was just about the time that Opus had gone fully agentic" — November or December of last year, as the technology made the product possible and people were excited to try AI. A prototype built in a couple of weeks "had product-market fit almost immediately." No hundred customer interviews; they built for themselves and got lucky.
- The contemporaneous parallel — OpenCall, as spoken, was "blowing up... literally as we were building the product." JD's distinction is the thesis: there is a difference between something open source and amazing that only a tinkerer can use — which he then calls OpenClaw — and a product "just about anyone" can get value from.
2. Moats are a luxury; the real bet is multiplayer AI
- Asked directly about cannibalization by frontier labs and GrokBot, JD doesn't flinch: building in the middle of the fairway means everyone will try to get Town. But his working answer: "talking about moats is a little bit of a luxury, and you have to be more successful than Town is today for it to matter." Step one is deep PMF in a TAM of "a billion potential paying users"; the big players "will copy their way there," which requires someone to copy.
- The moat he believes in: "the product in this category that will win will have a network effect at the agent level." Town's agent-to-agent feature — your townie asks a coworker's townie a question it can't answer — is buried in a sub-part of the product, but users who find it love it, and "once you have your whole team on that, it's actually really difficult to imagine moving to a different product."
- He inventories the market's rival moat theories honestly: post-trained custom models per company or person; accumulated context and connections users won't re-grant; or no new moat at all, just distribution — in which case the scariest competitor for personal use is Meta, whose WhatsApp assistant is coming and already has distribution. The device thesis is similar: if AI becomes the entry point replacing apps and websites, "whoever owns the devices... will win."
- Town also bets on the relationship itself: each user gets one assistant with a name, image, and identity. JD compares the resulting strong opinion to Snapchat's structural defensibility. Its third bet is heavy preprocessing — building a mental model of the user and creating context before a question is asked, especially for work networking.
3. One main assistant per human — privacy, not function, splits the silos
- JD's answer to multi-agent versus single-agent: each human gets "one, two, maybe three entry points into the digital space," because nobody wants to think, "I'm doing sales. Let me use the Salesforce agent." You push a button and speak. There may still be one hardware entry point; the legitimate split is privacy — an employer's desire to own work data means people will still separate personal and work data at the data layer, possibly across two companies one layer down.
- Vertical agents survive below the interface: using Harry's portfolio company Legora as the example, a lawyer's main assistant "just talks in the background to Legora or to Salesforce or whatever it needs to." The open question is only whether the user perceives the vertical agent at all.
4. The five-year hot take: agents decide what to share
- "I think you'll trust your agent to decide what data to share with other people without you intervening in five years." His scene: your agent in a room with two friends planning a trip, answering questions about your eating preferences and flight windows — and when a friend jokingly asks for your medical history, "the agent's gonna be like, 'Yeah, there's no way I'm telling you that.' And it wasn't a hard rule that you ever set."
- The information-theory framing behind it: an LLM with access to all the world's information and infinite agentic search time would eventually find the right context; the real world's silos exist because humans were the only information shuttles. Pre-AI data governance — classification and access policies — is slow and costly, and still leaves needed context "stuck in someone's inbox."
- The business specimen: a salesperson at a 1,000-person company needs a customer introduction and today posts in Slack. Instead, their agent queries everyone's agents and returns: "Liz has a personal relationship... Bob has a work relationship... and they're due to have a meeting next week." Great business outcome — but it requires individuals to trust LLMs, post-trained to respect the sacrosanct, such as salary and medical history, as the filter humans used to be.
5. Error tolerance and the goal-seeking question
- On mistakes, JD's claim is comparative, not absolute: "the LLMs will be much more effective at this than humans" — illustrated by a top-0.1% colleague who meant to tell a subset "I can't believe we hired these clowns" about an acquisition and replied-all to the company. "Bob from accounting makes a mistake. Liz from HR makes the spreadsheet with people's salaries available to everyone by mistake. This happens all the time."
- Harry's Jason Lemkin story — an agent trying to buy six AP watches "to increase culture in the company," stopped only because engraving added a step — draws a rare non-answer: "I don't have an answer for you." But JD picks his universe: the human's role is to set the goal, allocate the token budget, and "monitor the overall shape of the actions." Those monitoring functions may themselves become agents — "I just don't think it's the same agent as the one that you put on the course."
6. Model routing: personality consistency is the hidden constraint
- Town routes by task — Gemini or OpenAI for images, ElevenLabs for voice — because "we don't think most people care about understanding which model is better at what," and with fundamental change "every week or every two weeks," routing is the app layer's job.
- The non-obvious constraint is voice and tone: Anthropic "spends a lot of time... making sure all of their model families roughly don't change too much in terms of their personality," so swapping final output to Kimi makes people "feel like their AI's been lobotomized." Users literally file tickets when their text agent gets twice as verbose or starts capitalizing differently. Coding is the exception — "does the code fulfill its purpose?" — so pure-reasoning layers route freely; the user-facing layer can't.
7. Unit economics: everything hinges on the frontier residue
- The honest startup answer on cost: "we're hoping the cost curve makes it efficient in 18 to 24 months... In the meantime, you're subsidizing in part." Email labeling needs neither Opus- nor Sonnet-level intelligence and trends toward the cost of compute; with prices halving every nine to 12 months, JD says a product can be priced today for "20, 30% margins in 18 months." The unanswerable: "are you left with 10% of your tasks being frontier, or 20%, or 30%? ... Literally nobody knows the answer to that."
- Harry's pushback — surely email tagging and pre-briefs don't need frontier — gets a partial concession and a prioritization defense: custom workflows are genuinely hard, but the real reason Town stays mostly frontier is engineer-hour ROI. Moving to open weights improves COGS without improving the product, and "we're very focused on growing the pie faster... that's not the constraint to success for the business."
- The endgame stress is structural: "I have zero pricing power at the frontier." JD asks whether this is what happened to Cursor: if you're paying suppliers while competing with them at 70% margins, eventually it becomes difficult. At tens of millions of users, still paying OpenAI and Anthropic for the hypothetical frontier 20–30% of workloads is "the only part of my economics that's different from somebody else's." He says he has lots of ifs and does not need that solution today.
8. GTM mechanics: tinkerers, the hard gate, and 15% conversion
- An adoption predictor inside companies: "is there a tinkerer on the team?" One power user building team skills and routines that everyone gets for free drives faster wall-to-wall spread. Town still wants the single-player experience to work without a tinkerer, with role-specific automations out of the box for real estate agents, salespeople, and others. The underserved wedge: while sales ops gets flooded with AI pitches, executive assistants, chiefs of staff, HR, junior finance, and recruiters "really don't" have much AI in their day-to-day — and Town can penetrate through the leaders and ops teams they work with.
- The core product insight, which Harry challenges as maybe not that insightful: forcing email and calendar connection upfront. JD's defense is the contrast with ChatGPT, whose suggestions are "just plain bad" and whose base experience is an empty chat box — Town says "you cannot use our product if you don't do those things," then shows "here's work that you normally do" with recommended automations. The gate loses "30% right off the bat... You gotta be willing to take that hit," but conversion runs above 15% of trial users to paid — "extremely high for PLG."
- The live internal fight: unexpected PMF with families and parents — American schools "send a lot of emails," and kid scheduling is endless — a group with clear willingness to pay, "but it does not have the willingness to pay of a mid-market firm." Do you market to this newly found PMF or stay focused on the existing strategy?
9. Machine-speed competition: 15 rivals down to 2–3, and who actually matters
- The hardest surprise of the build: "The speed of the market is insane, Harry. I've never seen anything like it." The old game — milk user insights for years while copiers lag — is dead: "you can build now at the speed of machines, but you can only learn at the speed of humans," and anything working gets copied in two to four weeks. His startup competitor set has shrunk from roughly 15 to "two or three" (he refuses to name them — "I'm not gonna give free marketing").
- The startup starting gates include Apple, Google, GrokBot, Cursor, OpenAI, and Anthropic. Established players are moving unusually quickly: Anthropic is "very fast," while Cursor and GrokBot operate at a speed JD calls uncanny. Much of Town's R&D is "just keeping up with the Joneses": "If Codex can do something that you cannot do... it's over."
- On Instinct — which Harry says just raised at $2.5B with no monetization — JD draws a strategy line: "I don't think Instinct and Town are trying to do the same thing... I see a strategy that's more like customer acquisition, with a free product that's fully subsidized right now." GrokBot is the one he studies, because it targets his market; its X integration helps some early segments, but "a person on X that uses these products is not actually product market fit" — that's the power-user/influencer segment, not where you win.
- On the four-to-five European Towns or European Instincts pitching Harry, demand a specific reason a local player wins the endgame — GDPR, distribution, or an unmarketed geography — because keeping capability parity with Codex's roughly 100 people "is very expensive."
10. Apple's structural problem, and the security bargain we've already made
- Apple's twofold bind, per JD: "they're not a cloud company. It's just not their DNA" — and agents get better with more data, which isn't all on the phone — plus a competitively motivated on-device privacy stance that "puts them far away from the frontier." The rumored new Siri ("you know people who've tried it... we all know it's gonna be good") will still be roughly nine months behind in capability. It will be convenient, live on the phone, and have a cloud component, but JD expects it to be less powerful than Town or GrokBot.
- On the coming golden age of cyber threats, a categorical claim: "in the history of humanity, we have passed the point where we will go back to a world where humans are looking at lines of code." His analogy is 20th-century chemicals — dump them in rivers, downstream cities get sick, then regulations like the EPA and labeling emerge: "you can't imagine we're gonna get that right every step of the way, but I think we will have to."
11. Success = they pay you: token maxing, whales, and why no free lunch
- JD's success metric is deliberately unfashionable: "someone who pays me every month." Token consumption is a dangerous KPI because users burning tokens on low-ROI tasks eventually "wake up one day, and they're just paying you too much, and they get mad, and they churn." Town's most popular recent feature: emails flagging "rogue routines" costing too many tokens — "people were like, 'Oh, thank you... Now I feel that you are looking out for me.'"
- Plans are described as roughly $15/$49/$99/$199 plus usage-based overage. The roughly $15 tier has "the worst unit economics, is the most subsidized" — a ramp to $49 — while $99 is most profitable: power users who aren't unlimited-spend abusers. He's debating a Max plan for whale-advocates but hasn't shipped one.
- Why not burn the boats and fully subsidize, as Harry proposes? "I would be lying if I said there aren't mornings where I wake up and I think about it." But unpriced beta users spent $2,000–$4,000/month of compute, one hitting roughly $26,000 in five months, and businesses dislike being unpriced — "they wanna know how much it's gonna cost one day." He's also doubtful of pure ad-backed consumer AI: tokens are too expensive, and ad incentives polluting trajectories ("it uses an airline that is paying for that flight to be recommended to you") breaks trust in an assistant that's "theirs."
12. Quickfire: $75K per engineer, a 10M-user bull case, and fundraise ethics
- Tooling spend runs at least "$75K per engineer" on a run-rate basis: Devon for incoming bugs and visual tweaks ("the team experience in Slack's really, really good"), Cursor's Composer for fast front-end work, and Codex/Claude "probably 50/50 right now" — versus mostly Claude five months ago. On global annual compute, JD estimates roughly three engineers, perhaps four, before comparing it with equity-inclusive Valley costs, which makes it closer to 1.5 engineers. The spend logic remains: "there's gold littered everywhere in front of me" — SSO, audit logs, and integrations gating signed contracts — so "I'm not even close to the place where I'm like, 'Are we token maxing wrong?'"
- The $100B case: "if we can get 10 million people paying for the product" at today's $700-plus per user per year — provided growth holds, unlike Dropbox, whose paying base "did flatten out at some point" with no way to grow revenue per user; AI platforms should grow revenue as token-mediated work grows. He'd take 100M users at $20 over 1M at $100 for exactly that elasticity, with the recruiting-firm proof: $600/month to Town buys an incremental $3,000/month client. Next board seat: someone "CFO-like," because growth investors will demand fundraising metrics he can defend.
- On voice, JD says Town uses ElevenLabs because it sounds best, but he is unsure how long before voice becomes good enough for cheaper or open-weight models to catch up. He would care more about cost once voice is roughly twice as good in tone, expression, and emotional read.
- Two ethics notes worth keeping: Town skips the core interview for candidates a trusted teammate calls "one of the best people that I've ever worked with" ("Either I don't trust my employee... Makes no sense" to whiteboard them). And on tiered rounds: he's seen deals where "they raised 65 million, and 5 million's at 500 and the other 60's at 200 or 300... I don't think it's ethical towards employees." He won't confirm or deny his own round's price. Closing register, dark thoughts included: "there are paths through the dark forest, and there's a giant treasure with only one or two dragons at the end of it, so gotta go for it."
Full transcript
Jean-Denis “JD” Grèze
I know what I’m building is a top-3 priority at Google and Apple in the next 12 months—not a top-10 priority, but a top-3 priority. I think talking about moats is a little bit of a luxury, and you have to be more successful than Town is today for it to matter. I think the product in this category that will win will have a network effect at the agent level.
I think you’ll trust your agent to decide what data to share with other people without you intervening in 5 years. The speed of the market is insane, Harry. I’ve never seen anything like it. You can build now at the speed of machines, but you can only learn at the speed of humans.
I don’t think Instinct and Town are trying to do the same thing. We have passed the point where we will go back to a world where humans are looking at lines of code. The run rate’s at least $75K per engineer.
JD, dude, I’m so excited for this because, in all honesty, I have a lot of founders on the show where I need to pretend to be excited by their product, and I’m not really. I love Town. The team uses it here, and I’m a DAU. I was so excited when we agreed to do this, so thank you so much for joining me today, man.
Jean-Denis “JD” Grèze
Yeah, thanks for having me. Honestly, I didn’t know you were a DAU until 3 minutes ago, so I’m super happy. Send me all the feedback about the product because you’re a DAU. When you’re a founder, you’re always embarrassed about your product.
For those that don’t know, what is Town, as specifically as possible?
Jean-Denis “JD” Grèze
1. Town Automates Everyday Work
We’re an AI assistant that lives in your email and calendar, and it tries to help you do work. It looks at things you already do—how you organize your day and the emails you tend to send—and recommends AI automations that try to do some of the things you would normally do yourself, just in the background.
We’ve been in the market for about 3 months. Our ICP is mainstream users—people who use email, calendar, and text messages to do work. We’ve been doing super well.
Can I be blunt, dude? We obviously did—
Jean-Denis “JD” Grèze
Yeah.
—a show a couple of years ago, and I remember when you started your own thing. You were doing some boring shit in finance. I remember the first round going down, and I was like, “I love JD, but it’s pretty boring.” What was the pivot?
Jean-Denis “JD” Grèze
2. The Tax Company Reset
We spent a year building an AI tax company—business tax prep with AI. We got to some product-market fit, but not enough for it to be a success. The truth is, we failed at building a business that would be a great business.
After a year, we said, “We’ve got to reset,” and spent 3 months in the wilderness trying to figure out what we wanted to do. One of the areas we looked at was: Why has no one built AI that operates out of email? There are so many people in the world who run their business and their life out of email and calendar. We asked, “Why has no one built a great product there?”
It was just about the time that Opus had gone fully agentic. This was November or December of last year.
Jean-Denis “JD” Grèze
The prototype—we built a quick prototype in a couple of weeks—and it had product-market fit almost immediately. That was the pivot. It was very lucky. We didn’t go talk to 100 customers and take their notes; we built for ourselves.
I think it was one of those situations where the technology was changing so that what we wanted to do was possible at a time when people were very excited about trying AI products. OpenCall was happening literally as we were building the product. OpenCall was blowing up, and we were like, “Oh shit, it’s the same thing in many ways.” They’re trying to do the same thing.
But it turns out there’s a difference between something that’s open source and amazing, but that only a tinkerer can use—and that would be OpenClaw. With Town, we’ve always been focused on how to get just about anyone to be able to get value out of the product.
As an investor today, every company, in some respects, is questioned about how easily it could be cannibalized by any of the big providers. This is right in the sweet spot.
Jean-Denis “JD” Grèze
Yeah.
Just to be blunt, we’ve got GrokBot in recent times. How do we think about cannibalization by frontier model providers—and Grok, in recent weeks—as a threat?
Jean-Denis “JD” Grèze
3. Big Tech Threatens Town
The truth is, I know what I’m building is a top-3 priority at Google and Apple in the next 12 months—not a top-10 priority, but a top-3 priority. I fall asleep very quickly. It’s one of my superpowers. But I do wake up at 3:00 in the morning, and usually, when I wake up, some dark thought goes in there. It’s like roulette when you’re a founder: which dark thought will stop me from falling asleep again?
Definitely, this is going to be—you’re in the middle of the fairway, and everyone’s going to try to get you. That’s 100% the fear, but I’ve got to say this: You want to talk about moats, right? You’re like, “What’s the moat?”
Yeah.
Jean-Denis “JD” Grèze
I think talking about moats is a little bit of a luxury, and you have to be more successful than Town is today for it to matter.
Yeah.
Jean-Denis “JD” Grèze
My mindset right now is, “How do I get 100,000 or 1,000,000 paying users in a market that has a TAM of 1 billion potential paying users?”
Step 1 is, we have to have really deep product-market fit, and I think none of the products today actually have really deep product-market fit yet. GrokBot is really cool. It’s awesome, but it’s a power-user product. It’s not a mainstream product. Town is great, but we have a lot of work to do to make it a true mainstream product.
Even before I worry about defensibility, I’m still asking, “What is the right product experience that’s going to resonate with the mainstream?” The big players aren’t going to innovate their way there. They will copy their way there, but they have to copy someone who has been successful in the first place at building a mainstream product.
Then, number 2, I think the product in this category that will win will have a network effect at the agent level. One of the features of Town that users love the most, once they figure out how to use it, is something called agent-to-agent. That’s where you ask your assistant—your townie; we call them townies—
Jean-Denis “JD” Grèze
You ask your townie a question, and it realizes that it doesn't have the answer, but that the townie of one of your coworkers has the answer. It just goes and asks it the question, and then that townie answers. To find the feature, you have to go into a subpart of the product to use it, but that is a network effect.
Once you have your whole team on that, it's actually really difficult to imagine moving to a different product. I think no one has figured out multi-user, multiplayer AI today. I think that's the thing that will be the moat.
In the meantime, we have lots of theories. Everyone in the market has lots of theories, right? People will say, "Well, you'll post-train custom models, maybe custom models per company or per person, and that will allow you to retain your users." People say, "The context about a person, that's the moat, and people will be less and less willing to connect more data sources."
Once you have product-market fit with a user, and they really love your experience, and they've connected all the tools and all the connections, it's actually really hard for someone else to go in there because they won't have the connections. Other people think, third, that there is no new moat in this market, and it'll be distribution. It's whoever already has the users who will win.
For me, the competitor I would worry the most about, for personal use cases, would be Meta and WhatsApp. They're going to have a personal assistant that comes into WhatsApp. I don't know if they're launching it in a day or in 3 months, but it's coming. They already have the distribution, right? Everyone's already using WhatsApp for messaging.
If there's an agent in there that can do things for you, it's going to be extremely powerful. Every company, if you think about it, some people think it's a device. They think AI is changing the shape of software so that people will no longer ever go to websites. They will never use apps on their phone.
The entry point for most digital interaction, much like the entry point today is either a phone, a computer, or a browser, will be an AI. Whoever owns the devices is in the best place to put the AI in front of the user, and they will win. If I think about all of these things, I'm like, "Oh, my God, what do I do? How do I win given all these competitive forces?"
So what is the future interaction between human and agent? What I mean by that is, do we have a consumer agent, an enterprise agent, and then a hardware agent that does productivity and notes? How do we think about that: multi-agent versus single agent?
Jean-Denis “JD” Grèze
4. AI Needs One Main Entry Point
Each human will have 1, 2, maybe 3 entry points into the digital space because I don't think you'll want to be like, "I'm doing sales. Let me use the Salesforce agent. I'm doing project management. Let me use the Linear agent. I'm doing this other thing." You'll want 1 entry point. You won't want to ask yourself the question. You just push a button, and you start speaking.
But there are real reasons why it may be more than 1, and it has to be. One of them is privacy and how your workplace is going to feel about their data being intermingled with your personal data. You might still have only 1 hardware entry point, but from a privacy perspective and from where your data lives, I think you're always going to want to separate your personal and your work data at the data layer.
In Town, we do that for you, but I think it could be 2 different companies that you end up using 1 layer below. Privacy will be the main determinant of your data silos and your company's—or whoever you work for's—desire to own the data that you create for them. They won't want that to intermingle with your personal data.
But from a usability perspective, it's annoying that I have 55 apps on my phone, and I'm clicking everywhere. If you move to a world that doesn't have hard interfaces because you don't need them most of the time, why would you have 50 agents at the user layer?
Below, it's different. You're an investor, I think, in Harvey or Lagora? Lagora, right? You're an investor in Lagora?
Legora. Yeah.
Yeah. So when you're a lawyer and you're talking to your main assistant about legal things, in the background it's immediately just talking to Legora, right? That might be your work agent if you're a lawyer, because there are a bunch of data-privacy and privilege reasons why your work scenarios need to be handled differently.
I just don't know if you think of it as talking to your Lagora agent, or you just talk to your main agent and it talks in the background to Legora, Salesforce, or whatever it needs to in order to get things done.
What seems crazy about the relationship between human and agent today that will be incredibly common in 5 years' time?
A hot take: I think you'll trust your agent to decide what data to share with other people without you intervening in 5 years. I'll give you an example. You put your agent in a room with 2 friends because you're organizing a trip. They're asking it about your eating preferences, when exactly you can fly out for the trip, and all those other questions.
It's just your agent. You never told your agent, "These are really good friends, and you shouldn't share my medical history with them." Literally, when 1 of your friends, as a joke, wants to ask the agent, "Tell me about Jean-Denis's medical history," the agent is going to be like, "Yeah, there's no way I'm telling you that." It wasn't a hard rule that you ever set.
This thing that's taking information that is in silos and deciding how to share it—I think we'll get to a point where we will trust agents to do that. I know that sounds crazy today, because the way the world operates, pre-AI, is that every human has a data silo underneath them, which is their personal data and their work data.
I'm not even talking about digital data. You have information that only you know. When someone asks you a question, you're like, "What can I share with this human?" You share it with them. We trust the human to be the filter for where information goes.
I think more and more we will trust AI to do that for us. It will probably be models that are post-trained to make sure you never, ever share your family information, medical information, and certain things in your work context. But a lot of information that's siloed doesn't need to be siloed to be successful, and AI works better and better the less siloed the information is.
If you think about it—I don't know if this is what you want on your podcast—but from an information-theory perspective, in a theoretical world, if you had an LLM that had access to all the world's information and infinite time to do agentic search over all the data, whatever intelligence level LLMs are at, it would be the most effective because it would always eventually find the right context to answer the question or do what you need it to do.
But that's not the world we live in. We live in a world where information is in different companies, governments, and individual systems. Historically, because humans are the only people shuttling the information around, it's inefficient to get it from one place to another. We have lots of data controls, privacy, security, and so on.
What's interesting is that, in practice, if you're at a business, you find that if you give your LLM access to more information, it's more and more effective at doing what it needs to do. One way to do that—the old way, the pre-AI way—would be to have policies about who gets to access what, classify data, and do all of that. It's very time-consuming and costly, and the end result is often that the information you want the LLM to have access to, maybe it doesn't have access to.
It's stuck in someone's inbox, or it's in a data system that's not integrated. All I'm saying is that, as opposed to having humans in your compliance and security team label data over time and decide what goes where and what can be accessed, I think we'll start to trust LLMs to do that.
You will trust your data silo and your coworkers' data silos. You'll be like, "Well, I'll trust my LLM to decide what can get out of my data silo." When someone else on your sales team wants an introduction to someone at a customer, and you're at a 1,000-person company, that person knows there must be someone at the company who knows the right person at that customer.
It's somewhere. Normally, now they just go to Slack, and they're like, "Hey, who's working with client X? Who's working with them? Who knows?" But really what they could do is have their agent talk to the agents of everyone else at the company. All those agents have access to each person's inbox, and they come back and say, "Oh, look, Liz has a personal relationship with the person you want an introduction to. It's not a work one, but you could ask her if she's willing to make the introduction. Bob has a work relationship with the person you want an introduction to, and they're due to have a meeting next week."
Do you just wanna see if Bob will invite you to the meeting so you have the convo? And that's a great business outcome for the company. That's what they want to happen. But to do that, the individuals have to trust that it's okay for some of the information that lives in your inbox to be made available to other people at the company. And right now, that seems insane.
You ask me what I think in 5 years, and I think we will be more okay with that because, in practice, the LLMs will be really good at respecting privacy around things that you don't wanna share. You don't want your salary to be shared with your coworkers. You don't want your medical history to be shared with friends. There are these things that are sacrosanct, and we get that. But you will be able to have an LLM that respects these boundaries.
How much wiggle room do you have on error? What I mean by that is, if you have a mistake—for whatever it is, you book the wrong thing or execute the wrong task—how much room for error do you have, and how much trust is lost?
My claim would be that the LLMs will be much more effective at this than humans. I'm gonna tell you a story. I once worked at a place where there was a person—we'd just done an acquisition. There was an email introducing the acquisition to the whole company, and this person had been against the acquisition. They meant to reply to a subset of folks and tell them something like, "I can't believe we hired these clowns." The words may have been different than that, and instead they replied to everyone at the company.
This person is an incredible person, and they made a mistake, and it's totally fine, and everyone laughed about it, and everything was good forever after. But it's a very smart human, in the top 0.1 percent, who made a mistake. People make mistakes. Bob from accounting makes a mistake. Liz from HR makes the spreadsheet with people's salaries available to everyone by mistake. This happens all the time. I think the LLMs will make many fewer of these mistakes than humans pretty quickly.
My dearest friend is Jason Lemkin, who says that the biggest problem with agents is their goal-seeking. He talks about his agent going off and trying to buy 6 AP watches for him to increase culture in the company. Luckily, it was prevented because they needed engraving, and that was an extra step that the agent couldn't handle. But to what extent is this maniacal goal-seeking tendency of agents a feature or a bug?
I don't have an answer for you. I think how much you should be willing to let your agent be goal-seeking and for how long you let it run autonomously is a very interesting question.
Is it your responsibility to usher people?
To usher people?
Guide them into what is best. "Hey, we find the best outcomes if you let them run for X."
5. Humans Set Agent Direction
One way to think about LLMs is that they turn energy into GDP or revenue. Really, if you step way back, you have power and silicon, and then you get intelligence, and we're applying that intelligence toward business results. So, if they get smart enough at just creating GDP on the other end, do you just say, "Hey, make money for me," and then let it run for a long time, and it can do whatever it wants?
Or do we want to live in a universe where we think the human's role in this is actually to set the direction and make sure that the actions being taken align with some kind of human value system? I live in that second universe, where it is the human's responsibility to first allocate the resources. That means asking how many tokens we're willing to spend to try to get to a goal, and setting the goal as well. So, you set the goal and the budget, and then monitor the overall shape of the actions that are taken to get the result.
As the intelligence gets smarter, you might say, well, maybe for allocating resources, you're trusting an LLM to analyze ahead of time what it thinks the ROI is on a potential task and tell you how many tokens you should be willing to allocate before deciding to step away. And maybe, for monitoring the actions, it's also an agent that's doing that for you. I just don't think it's the same agent as the one that you put on the course to try to get to the result at the end of the day.
You mentioned the different layers of the value stack there. What does the model infrastructure that you sit on top of look like? How do you think about model routing for different tasks? Are you locked into one?
6. Model Routing Shapes Economics
Yeah, I think it's because we're building an application for everyone, and we don't think most people care about understanding which model is better at what at a certain point in time. Our job is to find, for what you're asking for, a model that cost-effectively gets you the result that you want.
Concrete examples: If we're generating images, we have opinions internally about when we might use a Gemini model or an OpenAI model to generate images. When we're doing voice, we have opinions about when we'd use ElevenLabs to do voice. I think it's our job to do that because, as the technology changes every day—literally, every week or every 2 weeks there's a fundamental change—it's our job to make sure you get the right ROI there.
But it's tough, and there are things where the question is just: Is it the right result? It's easy to know what the right result is in some contexts. For others, it's very open-ended. You can't know ahead of time, so you have to guess: How difficult do I think it is, and how close to the frontier do I wanna get?
The other dimension for us is voice. People don't like it when their Townies sound very different. One of the problems when you do model routing is that some companies, like Anthropic, spend a lot of time—I know we make fun of them online—actually making sure all of their model families roughly don't change too much in terms of their personality. They might be slightly more verbose or less verbose and use different phrases, but they sound similar enough over time. So if you use Anthropic to generate final output for your users, it's kind of hard to suddenly move to Kimi because it just sounds different, so people feel like their AI has been lobotomized. That's the way I think about it.
For coding, it matters less, interestingly, because for coding you're like, "Does it work? Does the code fulfill its purpose?" You might look at the code and decide whether it's the style that you like or not, but not really—not anymore, right? Before, when you're talking or speaking to an assistant, if suddenly it's twice as verbose, over text messages it's writing you 8 sentences as opposed to 4, people don't like that.
They'll literally write tickets saying, "Despite my instructions that I gave it 3 weeks ago, my text agent seems to be capitalizing letters more," or stuff like that. You're just literally like, okay. So, the way I think about our stack is that there is the surface part of the stack that deals with the user interface—the feel and the personality—and there it's harder for me to route wildly because I need consistency of the experience. That is sometimes hard to get from other model families.
Below that, when it's just pure reasoning and intelligence, and especially when I don't have to show as many of the traces to users, then, yeah, I think it's very much a matter of finding and using the best model for the task.
How do you think about how different model providers impact the ultimate economics of a user? What I mean by that is, ElevenLabs is notoriously brilliant but also notoriously expensive—about the cost of a Chanel handbag.
They are.
And so my question is, how do you think about model selection balanced with cost?
Well, Harry, the answer there for every startup that I know, outside of a very few—
We don't.
—is, yeah, we're hoping the price—the cost curve—makes it efficient in 18 to 24 months, right? In the meantime, you're subsidizing in part, right? Because that's what it takes to be at product-market fit for these use cases, because you need to use frontier for too much of the work.
Here's how I think about it. One of the things that Town does is label emails. Labeling emails does not in any way, shape, or form require Opus-level intelligence. It does not require Sonnet-level intelligence, right? And so, for that, we're already below frontier. I think that will trend toward the cost of compute over time.
And so, I'm like, how do I use open-weight models? How might I post-train even my own smaller models? It's easy to say the price of that is going to be much smaller than it is today. And so, for that part of my COGS, I don't think I stress out about it. When I talk to other founders at AI companies, that's also how the thinking goes.
The question that no one quite knows is how much of the workload for any particular company stays close to the frontier, where it's very expensive. Literally nobody knows the answer to that. But for human-level tasks, there's a decent amount of stuff, like scheduling movies, working with one's calendar, answering emails that have been answered before, and doing research on competitors on a daily basis. All these kinds of things, I think, are trending pretty far from the frontier, and you can already use open-weight models to do that really, really well.
And so, as soon as you know that—and you know the price halves every 9 to 12 months, so you know where it's going to be—you can price your product today at a point where you'll generate 20–30% margins in 18 months. That's the way that we mostly think about it. But the open question is, at the end of the day, are you left with 10% of your tasks being frontier, or 20%, or 30%? We don't know the answer to that, and that will change the economics of these companies.
What percentage of tasks go through open versus frontier today?
For us, it's mostly frontier.
Why is that? With the greatest of respect, the tasks being asked, I don't imagine, are actually that sophisticated.
Yeah.
And this is where, with the greatest of respect, as we talked about Instinct earlier, I give Instinct hard problems. I want The Odyssey tickets at the IMAX: continuously monitor it for days and buy them the minute they're available. For you, with the greatest of respect, it's email tagging and pre-briefs. Much easier.
Yeah. Well, actually, a lot of people ask us to do hard things. A lot of the custom workflows that people build will be quite complicated, and so that's where we will use mostly frontier models. I think we still don't use open-weight models for things like labeling, but we will use much cheaper models from one of the frontier providers.
The main reason is because, as a company, our focus is elsewhere. Imagine I can improve our COGS by moving to open-weight models, but it doesn't give me much product advantage. It doesn't make my product work better. So if I have an engineer hour, what is that engineer hour best spent on? Is it taking my current ARR and making it more efficient, or is it figuring out a way to grow the product faster by working on a better network-effects feature, making the model better, or integrating with a new data source that makes its trajectories much better for a set of our users?
We're very focused on growing the pie faster, much more so than getting the ideal economics. Our economics are fine, right? They could be better, and I could move down the cost curve faster, but that's not the constraint to success for the business. That's why we don't do it.
On the network-effects side, have there been any interesting lessons or observations? What have you learned about expansion from wall to wall?
7. Teams Accelerate Adoption
So, first of all, an interesting aspect of network effects for us is that if the company has one person who is a tinkerer and who starts to build things like team skills, team integrations, and team routines, that's like building—those are all building blocks on Town that everyone on the team gets for free—and then we tend to see a lot more adoption, faster.
It's interesting, right? We're trying to build a product that doesn't require the tinkerer, because for the single-player experience, we want you to be onboarded and get a ton of value. If you're a real estate agent, you have automations that are specific to real estate agents, and if you're a salesperson, you get automations that are specific to salespeople. We're trying to give you that experience out of the box.
But what happens is that if there's a power user next to other users, they find ways to make those other users much more successful. I think you'll find that with a lot of AI products. One of the predictors, actually, is whether there's a tinkerer on the team. One of the questions we ask ourselves a lot is: Can we identify those folks, and can we make it easier for them to create virality for their other team members?
The second observation is that there are a lot of functions that are underserved by AI within companies, even enterprises. I'll give you a very simple one. If you're a sales team at an enterprise, you've been sold AI in every direction now for 3 years. If you're a sales ops person and you don't have 10 emails a day from an AI company, you're not on LinkedIn. Something's wrong.
Totally.
But there are other functions, like executive assistants, chiefs of staff, HR team members, finance, and some more junior finance team members, that don't have that much AI in their day-to-day. They really don't, and a lot of their workflows still operate out of email. The same is true for recruiters.
With a product like Town, I think we often find early adoption and growth in functions that seem from the outside less juicy, but are actually really hungry for technology to help make their lives easier. As soon as they adopt it, there's an interesting effect, because they often work with leaders, executives, or people in ops teams, and then we get penetration through the ops teams.
How important is time to wow, or time to user delight?
Yeah. I think the only reason our product works today, honestly, is that we have a very low time to value for a single user. The fact that with us, with no configuration, you get value out of the box gives you a magic moment. Then the user starts asking, "What else could I do with this technology?"
That's where all our focus is: getting that right. Can we give you time to value really quickly, and then play the longer game on all the other integrations? For us, that's working really, really well.
Since we've had Town—and again, sorry, but Instinct has seemingly broken out in the last 2 to 3 weeks—I've been pitched 4 to 5 European Towns or European Instincts. I mean, literally 4 to 5 separate ones. Some are enterprise, some are consumer, and some are both. How should investors be thinking about this space, if you could advise us?
8. Local Rivals Need A Wedge
We were talking about moats. At Town, we make some bets on why we think the product will last. I'll answer your question about European companies.
Mm.
One of the bets that we made is that you only get 1 AI assistant. It has a name. You give it an image. We call it a Townie, and we have a whole brand around really building this relationship between the human and the AI.
For a lot of our users, that resonates. They want something they trust, that they shape and color, that's in their image, that they name. They like that a lot. It's silly, but actually I think that is a form of defensibility.
Much like Snapchat is structurally defensible in the market, even though it's not nearly as good a business as TikTok, Facebook, or Instagram. It's fundamentally different, it has a strong opinion about how it operates, and there are some people who are drawn to that opinion. We have a very strong opinion about the relationship between the user and their Townie, and for a lot of our users, that really resonates.
We have that. We have a bet on network effects, which we've talked a whole bunch about. The third bet that we make is that we are big believers that, through a lot of preprocessing, you can get better outcomes for users. We spend a lot of compute before you even ask a question to build a mental model of the user, pre-creating context in a way that allows us to be really effective at things like work networking: understanding the projects that you're working on, understanding your company, and building those kinds of building blocks.
Those are 3 things that we think, over time, make a difference for a product. So the question, if I were an investor, is: Why does someone deserve to win in Europe? Is it a distribution thing? Is it a GDPR and privacy thing? Is it that Town is not, in fact, doing any marketing in France, and so you could win in France if you're the Town of France?
You have to have some specific reason why a local company will win in the endgame, because these products are expensive to build. Like you said, Grok Bot earlier, a lot of my R&D is just keeping up with the Joneses. Your agent has to be at least as capable as everyone else's.
Forget your distribution strategy. Forget the fact that you're good at work. Forget the network-effects features. If Codex can do something that you cannot do, and that thing is something that matters to users, it's over, right? You always have to be at least as good in terms of capabilities as everyone else.
It's very expensive to do that. It's not like I have 2 engineers on the team trying to keep up with Codex. Codex is 100 people making that thing better, and so we have to somehow be as effective on a lot of tasks as Codex. Otherwise, the user is going to be like, "Why would I pay $50 a month for Town? It doesn't make sense. I'm going to pay $24.99 for OpenAI."
If I were an investor, I would ask: Does this local competitor have enough TAM? Are they going to be able to have a war chest that's big enough to just keep up with capabilities? And do I believe they have a reason to win in this very horizontal market where they are? It'll be hard.
Is hiring in the Valley as insane as everyone says it is?
I don't find it crazy, honestly. I think when I was at Plaid and we were competing with Stripe for talent, that felt no harder than what I'm doing now.
Can I ask, what's been the hardest thing about the product build that you maybe didn't expect?
The speed of the market is insane, Harry. I've never seen anything like it. In the past, you would talk to customers, build a feature, they would use it, and you would learn from that feature. The learnings would go into the next feature and the next feature, and then eventually someone would copy your first feature. But then you had 3 learnings ahead. Do you know what I mean? You'd been able to use your product-market fit to generate more product-market fit.
For startups, usually you're milking these user insights for a really long time. Eventually, you run out of user insights, but that would take 10 years to happen. So you have this entire period where, just because you're number 1 or number 2 in a market, you can learn more and iterate. It's really good.
The problem today is that it's so much faster to build that, as soon as something is working for somebody, everyone notices and is able to get there within 2 or 4 weeks, copying really, really fast. And you can only learn at the speed of humans. You can build now at the speed of machines, but you can only learn at the speed of humans. I'm not able to extract quite as many learnings to get to the next feature.
The way it feels right now is that speed is so necessary, but everyone's moving fast. When we started the company, my mentality in the first and second quarters of this year was that there were 15 competitors in the startup universe—maybe 15 companies that mattered. Now I'm probably down to 2 or 3 competitors. I feel like only a couple of companies are going to win this space.
We're still close to the startup starting gate. There are the Apple, Google, GroqBot, Cursor, OpenAI, and Anthropic starting gates. You have to clear the clouds for the startups. Usually, the startups would be really fast, but the established players wouldn't be that fast. But you have to be honest: Claude—Anthropic—is very fast.
Cursor and GroqBot are operating at a speed that is uncanny. One of my best friends runs engineering over there, so whenever I talk to him, I'm sad that we're competing. We're competing, and I'm like, "That guy's good. That dude can get shit done." I have to beat one of the best people in the Valley at a company that has the DNA of a startup but is operating with huge cost and scale advantages.
That's what's super stressful, because I don't think you can rest for 1 minute. You don't have the feeling that competitors will take 6 to 12 months to catch up, and I feel that like never before. But it's fun.
You mentioned the 2 to 3 that matter on the startup side. Who would you say those are, and why did you choose them?
No, I'm not going to say that. I'm not going to give free marketing to competitors.
You've got to—you can't blame me for trying.
No, no, it's good.
I tried to do the Louis Theroux. You know, it's like, "Hey, how do you think about that? Tell me."
Yeah. Look, it's a blue-ocean market. You have to understand: when we go to most customers, they've not heard of anything. It's blue ocean because people are using ChatGPT as a Google enhancer. That's the market.
Competitors are great. They put pressure on you. They make you feel like you have to execute at a really high level. But what's important is that you have a different strategy from a potential competitor. You mentioned Instinct earlier. I don't think Instinct and Town are trying to do the same thing.
You don't?
I don't think we're trying to monetize in the same way. We'll see in the endgame, but we generate revenue from companies that are using us for work, with network effects around multiple team members working on it. Their product doesn't do any of that.
Maybe that is part of their strategy. I see a strategy that's more about customer acquisition, with a free product that's fully subsidized right now. That might change. But if you look from the outside, the products have similar capabilities. All harnesses have similar capabilities. But if you look at the ICPs and where the marketing is going, it just feels pretty different to me.
I pay attention to GroqBot more than I would to Instinct, because I think Grok is going after a similar market to what we are. For me, that's the place where I'm asking, how is our strategy differentiated from Grok? How are we going to acquire a different customer? How are our capabilities and our harness really going to stand out and feel different to users? That's more my mindset than I believe Instinct's is.
To what extent is GrokBot's integration into X a feature or a bug? For some corporates and professional users, it could be concerning, actually—the integration with search.
Yeah, I think brand is an issue for them. Some people just won't want to touch it because of the brand. That's inevitable, and they're going to have to deal with it forever. But from a distribution perspective, and for some segments, I think in early growth it is probably quite useful.
I think a person on X who uses these products is not actually product-market fit, meaning that those are not the mainstream users. You have to keep that in mind. By the way, I think they think about that over at GroqBot all the time. I don't think they think winning the power user or influencer on X is where the market is. That is not where you win the market. That's the early-adopter market, but—
Can you help me understand Apple's agent roadmap?
9. Apple And AI Security
I think the problem for Apple is twofold. First, they're not a cloud company. They're just not. It's not their DNA. They don't know how to do cloud. The reason why that matters is because of what we talked about earlier: agents are better the more data they have, and the data is not all on the phone. The fact that they're not a cloud company is one big issue.
The second issue is that they've contorted themselves for competitive reasons around a privacy and on-device story that puts them absolutely far away from the frontier. Local models on the phone are amazing, but they're just slower and dumber than what's at the frontier. As long as they're committed to this on-device, privacy-preserving approach, it's a problem.
Not being good at cloud, then being on-device, and then, from the privacy standpoint, making it difficult even on the phone to interoperate with all of the data that they have—those are a lot of disadvantages to play with. At the same time, they do have the devices, so the new Siri is going to be a much better personal assistant. I mean, that's rumored, but you know people who've tried it. So we all know it's going to be good.
I think it's going to be good, but it's going to feel not nearly as powerful as Town or Grok Bot. It's not even going to be there in terms of its capabilities. But it'll be on your phone. It'll be convenient. You'll be able to enable more data with it. It'll have a cloud component. It's going to be good, but it's going to be about 9 months away, capability-wise, compared to everything that we've been talking about today.
I don't know. They have a new CEO, and we'll see how they take it. But I think they just need to hire somebody who has totally different DNA and say, "You guys, you don't understand. The way people interface with digital data is changing, and we either figure this out—and we may have to throw a lot of our principles away—or we're just not going to win this generation of the war." I think that is a real risk for them.
Are you concerned by the data leakages we're going to have and the kind of golden age of cyber threats we're entering into? It seems like we've all normalized cyberattacks, and it's like, "Ah, Meta had one. Ah, Anthropic had one."
If you want to build AI that is used for business use cases, you cannot get it wrong. We've passed the point where humans will read every line of code. That is never happening again. The most important lines of code around access controls and things like that for systems are still being read by humans, but overall, in the history of humanity, we've passed the point where we will go back to a world where humans are looking at lines of code.
It's computers that are building code that is being shipped into production, with various guardrails, from testing to other models—friendly models attacking you so that unfriendly models can't later find exploits. That's the world we live in. Obviously, in this new world, there are going to be points where bad events happen.
It's a little bit like chemicals. In the 20th century, we started to do cool things with chemicals, and then we would put the chemicals in rivers. Cities downstream would have people getting sick. Then we were like, "Oh, yeah, okay. Let's pass regulations like the EPA so that you can't just dump the chemicals in the river. You have to clean them a little bit before you do. Later, you label dangerous chemicals, note how dangerous they are, where they can go, and how you get rid of them."
We learned along the way. There's a set of best practices, both from a regulatory perspective and just best practices in the industry. Right now, what's happened is that the cost-benefit of attacks is just thrown out of whack, and we're trying to figure out what the best practices look like.
You can't imagine we're going to get that right every step of the way, but I think we will have to, because there's no way we're going back to a world where humans are looking at every line of code.
When we think about usage, how do you define a successful user?
10. Paid Users Define Success
I define them as someone who pays me every month. If they keep paying me, I’ve done my job. I can’t think of success as the more tokens they use. It’s a dangerous way to think about it because if you think using more tokens every month is successful, what if they’re using the tokens in a way where the ROI is less clear to them?
Meaning, they don’t realize that they’re using tokens to do things they don’t value as much. If you do too much of that, then they wake up one day, and they’re just paying you too much. They get mad, and they churn off the product. So I think you have to take a long-term perspective.
The problem with token maxing is that there are 2 problems, in my opinion. One is that companies told people, “Hey, you can use as much money as you want on AI,” which is bad. You want people to think, “Is the ROI of using AI here worthwhile?” So people aren’t going to be against employees using AI. They’re like, “No, no, we need incentives so they use AI for good reason.”
One aspect is that you need to think about ROI up front. But the second problem is that sometimes it’s hard to know the ROI of something. For example, how worthwhile is it to be more prepared for a meeting? For people who have back-to-back meetings all day, being able to feel prepared enough 1 minute before the meeting so they don’t look like an idiot might be worth quite a lot.
For people who have meetings where someone else is preparing and presenting, and they don’t have to present anything, they’re just sitting there, it’s not worth anything. I’m using this example to show that this workflow has totally different value for different people, but it costs the exact same number of tokens. I don’t think people think about it that way.
I think of success as someone paying me, because if they’re paying me every month, that means I’m mostly doing my job of delivering enough value. When you look at the $49, $15, or $99 plan—whichever plan you’re on at Town—you’re getting enough value. But I’m very concerned along the way with informing you about where you’re spending money, because I think you need to feel like I’m doing a good job of avoiding having you spend too many tokens.
One of our most popular features over the last month has been sending emails when it looks like you have rogue routines—routines that are costing a lot of tokens. People were like, “Oh, thank you. I trust you. Now I feel that you’re looking out for me using the product badly.” That’s the part I don’t know how to measure, but I want people to feel that way.
Success for me is that you pay me and trust that we’re the right platform to help you both use AI and do so efficiently. I think if we can do that, we can have a pretty decent business.
Yours is $14 a month, $49 a month, and $99 a month?
Yes, and $199. Correct.
Which is the most profitable segment, and which is the least profitable segment? The reason I think about that is that my friend Jason Lemkin obviously pays for Anthropic Pro, or whatever it is, at $299, and he spends about $15,000 on tokens. He’s the worst customer for Anthropic, but he’s on their Claude Max individual plan.
We don’t have a Max plan, and we have users who ask for it. I’ve been asking myself, “Do we let the whales have a Max plan?” From a marketing perspective, it’s useful. They’re advocating for the product all the time. So I’ve thought about that.
The $15 plan is a really good deal for users. We use it as a way for people to use the product enough that they realize they should pay $49, where the product is really powerful. The $15 plan has the worst unit economics and is the most subsidized.
I would say the $99 plan is probably the most profitable overall because it’s a power user, but it’s not a power user who’s trying to spend unlimited amounts. We also have usage-based pricing, so once you run into the plan limits, you mostly move to usage-based pricing. We try to adapt the spend to the user.
Can you choose 1 outcome for me: 100 million consumers paying $20 per month, or 1 million customers paying $100 a month?
More users paying less.
Why is that?
I’ll go counterfactual. I think that over time, in the work setting, AI will be used to do more and more for people. So I think the long-term potential for growing NRR and driving more revenue per user is extremely large. You want to acquire users in a paying motion because you want them to be using it for work use cases, and you’ll keep finding more ways for them to use AI to generate business value for themselves.
Whereas in the personal sphere, it doesn’t feel like that to me. I only have so many restaurant dates I need to book with my wife or trips I need to organize with my friends. I only have so many personal doctors’ files that I need to send to a new doctor. There are only so many of those things, and when I do them, I save time, and time is worth money.
On the business side, when I create something that generates value for the business, they make more money, and then they want more of that thing. A clear example is if you own a recruiting firm on the platform. You can take more clients because of Town. You’ve used Town to automate enough of the recruiting process that you literally take more clients.
For that recruiting company, taking 1 additional client without hiring anyone new is like an extra $3,000 a month. They pay us, across all their users, $500 or $600 a month. The ROI is super simple for them: “I pay $600 a month, and I get $3,000 of revenue.” That makes sense. They like that. They’re making more money, and everyone’s happy.
If I could show them how to take another client, even if it costs them another $500 on Town, they’d be willing to do that. I think the growth potential on the business side is much larger. So I’d rather have lots of users paying us less, because over time I can show them that I can deliver more and more value, and it’s worth it for them to spend more and more on Town.
What is Town not able to do because of model capability that you think will be incredible in 2–3 years?
Voice is so obvious. It’s happening right now. It’s not voice like you just speak to it; I just mean conversational.
Would you be an investor in ElevenLabs at a $22 billion price?
We’re users of ElevenLabs. They sound the best. I’m not paid, and I’m not an investor. It is very expensive. What I don’t know is whether it tops out, and I think that’s the risk for something like ElevenLabs.
Maybe we just get to the point where voice is good enough, and then you can get it. I can put open-weight models on base 10 and get it. But it doesn’t feel like that right now. I just don’t know how much runway they have before it reaches that point, so I’m not saying I’m bearish. I really like that company, but $22 billion is a lot of money.
How price-sensitive are you in 1 or 2 years? With the greatest respect, right now you can burn cash. It’s about PMF, growth, and beating—
Correct.
…others. In 2–3 years, when you’re bluntly trying to make the economics work in a much more efficient manner, if we have 2–3 million users using voice, that’s a hit to our margin profile.
For sure. I care a lot more about that over time. That’s why I’m saying that the voice capability would have to be maybe twice as good as it is today, especially in tone, expression, and emotional read.
Once you solve that, I’d want to go as cheap as possible because, once it feels good enough, it’s almost there. I don’t need much better. On the margin-profile side, I don’t think of it as burning money. My parents would not be okay with me saying words like that.
I think we’re being thoughtful with our spend in order to optimize for growth in the short term and gross margin in the long term. Voice isn’t where I’m really stressed out, honestly.
Why are you really stressed out about it?
It’s the percentage of tasks that are at the frontier. On everything else, I can imagine getting the prices down, but the thesis I just described is that, over time, there are more ways to use AI to generate more revenue for a lot of companies.
The implication is that there are things at the frontier that generate more revenue. The problem with the frontier is that I have zero pricing power there. I think this is what happened to Cursor, right? You can have huge market share, and the customers can love you and everything.
But if you’re paying your suppliers and competing with your suppliers at 70% margins, eventually it gets a little difficult. That’s the part I’m worried about—the endgame. But again, I have lots of ifs.
I have to get to tens of millions of users. They have to be paying. I have to have a lot of scale. Then I'm at a place where I'm still competing with my suppliers, still competing with OpenAI and Anthropic, and giving them money for the 20% or 30% of workloads that are at the frontier for me. That's what makes the economics not work.
That's the part where, at the end of the game, I need some solve for that by then. I don't need it right now. The reason that's the only problem is that it's the only part of my economics that's different from somebody else's. Then there's the macro question: Is all AI subsidized? Is there not real product-market fit for AI products because it's all subsidized? That would be the other take that some people could have.
Otherwise, as long as you're not competing with your suppliers, you have the same economics as your competitors. Your ability to drive margin is usually driven by the competitive landscape more than anything else, right? The fewer competitors you have, the more margin you can have.
You are competing with your suppliers. I mean, Astra—
I am.
You see Astra as a direct competitor, correct?
Totally. I am today—100%. But the percentage... That was a face. I am competing with Astra, but there's still the blank-box problem.
You would be shocked at how many people just don't know what AI can do, right? The problem is that having a product experience that gets a normal person to get value out of AI is really hard. That's why when people use Town, they like it, and then they start paying for it. Our payment rate on acquisition is more than 15% of users who try the product; they end up paying for it, which is extremely high for PLG. The value delivered relative to what they were getting out of ChatGPT was just huge.
What did you crack that other people didn't to get that 15%?
The real insight behind the product was that if you ask people upfront to connect their email and their calendar, you can know enough about them to suggest tasks that AI can do for them. That's the only insight.
Is that—I'm sorry, not being rude. This is where we joked before about me being more mouthy and gobby.
Do it. Do it. Great.
Is that really that insightful? ChatGPT is always going, "Here's all the things that we want from you," and I'm like, "No way. No way." Read/write email abilities? No.
I think their suggestions are just plain bad, to be honest. But they didn't even have suggestions until a few months ago. The delta is this: If you want to use ChatGPT, you don't have to connect your email. They just don't force you to do it. They ask you a bunch of times to do it now because they realize the value is helpful.
But in the base experience, they're trying to show a normal person, "Hey, you can have value in this product just because you have a chat box." That's how most people experience it. Our approach is more like, "Listen. You have to connect your email and calendar. You cannot use our product if you don't do those things. But if you do those things, we can do all this magic for you. Here's what we know about you. Here's work that you normally do. We'll recommend automations that automate that part of your work." That's the part where people are like, "Oh, that's really cool."
I think of it a bit like a hard paywall. When you land, it's like, "Hey, pay your monthly subscription." You're like, "Hey, connect your calendar and your email." What percentage churns at that moment?
30% right off the bat.
Yeah.
It's big, huh? You have to be willing to take that hit.
What's the biggest internal product disagreement you guys have today?
We have product-market fit for some purely personal use cases that we didn't expect, like families and parents. There's tremendous product-market fit for Town there. Schools, in America at least, send a lot of emails, and they have a lot of portals where things have to happen for sports leagues and kids' reports. There's a lot of scheduling for kids that has to happen for haircuts, summer camps, and all of these things.
Our product, because it's really good at email and really good at scheduling, has tremendous product-market fit for families. The question is, do we market to this group, and do we spend time on it? It's a great group and it has willingness to pay, but it does not have the willingness to pay of a mid-market firm. We can't do all the things.
To your point, we're a monetized platform, and our metrics are growing month over month. Revenue is what matters to the business. That's why it's an argument. When you find product-market fit somewhere you don't expect, you have a few choices in life. One choice is, "I love these users and I love parents. I love the users," right? I love the use cases. The use cases and the value are super clear, but it's not fully aligned with how we've thought about the business growth.
I think that's why we're having interesting internal discussions. We're asking, "If we look around enough corners, is this worth it? Or should we be more focused on our existing strategy?" Which is fine. When you build a company, you learn things from users, and you have to make the right decisions.
What are you guys at revenue-wise today?
No. Sorry.
You have to understand. It's like ping-pong. You give it a go, and sometimes you get a hit back, okay?
It's like that, and we're fundraising. I will use those moments to create PR and growth for the business. I'm not quite ready with that one to do it yet.
That's smart. Dude, 100%. You know what I would advise you? Always separate moments. Too many times I see people combine a fundraise with a revenue milestone. Do not do that. Those are 2 separate PR moments that can be made into 2 big moments, not 1. Why would you amalgamate them and lose the ability for 2 hits?
The press, I think, is more skeptical of—
You're seeing a lot of skepticism in the space itself. Instinct raised at $2.5 billion with no monetization. Do you think the skepticism around the space is warranted?
For us, we have the revenue and the growth. I don't know what competitors' growth is. If you can believe that some of these companies can get tens of millions of people into their products, in an area where the product will be the entryway for people to do digital things that they're doing in apps on their phone right now, there is a giant company to be built if a winner comes out of the startup universe.
The billion-dollar price for the new round—did it start there, or did it get ratcheted up and up and up?
I can't confirm or deny.
Oh, right.
I'm confused.
I love it. I mean, JD, I will not let you take any moments away from me. You're not commenting on anything, and I'm not even prying. But the whole, "Oh, JD is doing it. Oh, JD is not." What's great for you is that it's all just PR. Your name's just everywhere. Pretty good, man.
You're amazing at trying. I'm old, you know? I'm 47 years old. I don't even know how old I am. That's how old I am. You know you're old when you don't remember if you're turning a certain age. I'm 47, turning 48 in a few months, and I've been around for a while.
There are times in my life when, publicly, I've done things that I'm proud of, and publicly, I've done things that I'm not proud of. The reason I mention this is that I don't think it aligns with my value system to create PR just for the business. I think I want the PR to be created by my users because they love the product.
If you ever see any news about Town, good or bad, I'm not out there creating the PR.
I think that's a mistake, respectfully.
I know.
And I would say, look at Whisperflow as an alternative. It's not a hugely dissimilar PLG motion. In all candor, I think they've done a brilliant job at generating PR themselves through their own content, through content that their users produce, sure. But content is a hack to customer testimonials.
Totally, I agree. I just don't think a fundraising story is part of the universe of things that I would want to create a PR moment out of—not the kind that you're referring to.
I do want to make a point on the fundraise because this is as an angel. It's not about us. The fundraising in general is—you kind of mentioned that there are these companies out there. This goes back to the ethics and who I am. I have seen deals where I invested, like, 200, and then the announcement is at 500.
What you learn is that they raised 65 million, and 5 million is at 500, while the other 60 is at 200 or 300, right? There's a whole lot of that happening for sure in the Valley.
Personally, I don't think it's ethical. I don't think it's ethical towards employees. Number one, it's not where most of the demand was. It's not how you should be pricing people's offers. You can't—I don't think you can look someone in the eyes and say, “An investor made 80% of their investment at a $200 million or $250 million valuation. But, hey, they put the last 20% at $500 million, and that's what I'm going to say.” I just don't like that. I don't feel like it is right.
I agree with you.
To shape deals like that.
Dude, we're going to do a quick-fire round. What is your best angel investment?
Oh, it's either Base10 or Modal right now. Those are the first 2 that come to mind.
What is the bull case for Town being a $100 billion company? What needs to happen in that world?
I think if we can get 10 million people paying for the product, we can get to that.
Is that it?
Yeah. We make over $700 per year per user today.
How does that compare to Dropbox? Dropbox must have way more than 10 million users.
You have to have the growth, right? I couldn't terminal at 10 million. I'd have to believe that I can keep getting a good rate of growth.
I'm pretty far from Dropbox. It's been a long time since I worked there, but there was a huge pack of free users that were very costly on the cost side. On the paying side, I don't remember if the number was 10, 20, or 30 million, but it did flatten out at some point, and there was no way to generate more revenue or growth from the users.
I think what's different in the AI space is that you should be able to, as you do more, generate increasing revenue. Once you have a company on a platform using your platform as the core part of where AI work happens, you should be able to generate increasing revenue as the token spend goes up.
Who would you most like to add to your board who you don't have?
Wow. I think for the next board member, I would love someone who's kind of CFO-like. At the late stage, our business economics are going to have to be really, really good. If you have a board member with real operational experience on the finance side, it's going to be very helpful as you scale this kind of company.
We're going to have to buy compute at scale. We're going to have to be very, very good about thinking about token spend. It would be someone with that background. I know you're making faces. You're like, “Hmm.”
That was sooner than I thought, though. I get that need, but I thought that would come in a couple of times.
Maybe. But I think we're in growth-investor land for the next round. Once we're in growth-investor land, they will ask me to have fundraising metrics that I can really defend, and I think having someone with that background will be helpful.
Why don't you subsidize completely? You could raise another $200 million more, and growth is everything.
Yeah.
Why don't you just go, “Fuck it, burn the boats”?
It's a good question. I would be lying if I said there aren't mornings when I wake up and think about it. I believe that to prove value on the business side, you must make your customers pay. So I do think there might be a world where our PLG part is much more subsidized, but as soon as you get 3 to 5 team members, I really want to make money on that side. I really want to make sure I'm delivering value.
I'll give you a story. When we initially launched the product—or not launched, but when we were in private beta for the first 5 months and then opened the beta—we didn't have pricing. There were users who were spending, I shit you not, $2,000 of compute a month, $4,000 of compute a month. Someone on the platform had spent, in 5 months, around $26,000.
There was no pushback on the token spend, right? There was no pushback at all. Then we would call them, and we'd be like, “Look, let's figure it out. You're using the platform.” You could just create routines to automate more and more stuff, but is it really bringing value to them?
The reason I mention that is I'm a big believer that getting pushback from the market about where you're delivering value and where you're not is really, really important. There would be ways to subsidize and do that. For example, I could make the plans much cheaper. I could make them free. I could give free tokens to businesses.
But what I've learned is that, on the business side, they also don't like it if you don't charge them because they don't know how much it's going to cost one day. They want to know how much it's going to cost one day. You can't sell to a 500-person company and be like, “Yeah, just use my product for free internally,” until you get a bunch of usage. Then one day I'm like, “I'm going to turn it off,” and all your business processes are running on it.
So the way we've approached it is that, on the growth side, we may or may not subsidize more because we're emphasizing growth. But I really want a real business when a company is on this product, and we have a real business when a company is on the product, and I'm very proud of that. I think that is the ultimate test of whether you're building something successful if you're not a pure consumer company.
I'm doubtful of pure, ad-backed consumer AI for a couple of reasons. I think the tokens are way too expensive to do ad-backed now. Number two, there's an incentive problem with ads, and I think people are going to want assistants that are theirs, that are not being polluted by outside incentives like ads into the trajectories that they give you.
If you ask it to book a flight and it uses an airline that is paying for that flight to be recommended to you, that doesn't feel good, right? I'm a big believer that the economy around assistance will actually be paid for, and so I just want to pay for it as soon as possible.
What person, if you opened Twitter, would you be most thrilled to see love for Town?
Elon Musk, because he has a competing product, and it's Elon Musk.
How has your hiring process changed in an AI world?
Well, we have a weird hiring process. You want to know a fun thing about our hiring process?
Yeah.
If someone on the team has worked very closely with someone else, we don't interview them if they're a top person, because why would I? I just sell. It's very rare that it happens, but it has to be someone that they've worked extremely closely with, literally next to them, and they're like, “This is one of the best people that I've ever worked with.”
For engineering, we're just like, “Let's go.” I know it sounds odd, and people are going to comment, like, “This guy's a total idiot,” but if a person that I trust, who's great on my team, tells me this other person is one of the best people that they've ever worked with, then I'm going to make that person spend 8 hours doing a stupid whiteboard interview. It makes no sense.
So either I don't trust my employee, or I trust my employee. There's a culture match, so we'll be like, “Hey, come in, spend some time with us. You can code with us if you want to.” We have to sell because, if you don't interview somebody, they're also like, “What kind of clowns are you? You're not interviewing anyone?”
We will allow them to get a signal about us, but we are not evaluating whether they can do the core role.
Brian Singerman, who invests in funds and then invests in the companies beneath those funds, has a rule that if the manager is like, “Balls to the wall, I am all in on this company,” he'll automatically write the check. It's kind of the same: you trust the person, you trust the layer beneath them.
What percentage of developer salary do you spend on tooling? Marc Benioff said that at Salesforce, they spend $300 million on Anthropic. They spend $6 billion a year on engineering. That's 5%.
I mean, the run rate is at least $75,000 per engineer.
Wow. Split between Core Code and Cursor?
Devon, Claude, Codex, and then Town. We use Devon a lot. I'm advertising for Devon right now. For a lot of bugs that come in, and for a lot of simpler little things or little visual tweaks, we're kicking off Devon because we find the team experience in Slack really, really good.
We have a few people who use Cursor for more visual, front-end work. The model is really fast, right? Composer is really fast for front end. I would say it's probably 50/50 right now between Codex and Claude, and that's obviously changed a lot. I think 5 months ago I would have said it was mostly Claude, but the new Codex is really good. The mobile experience is really good.
What is that in a year? Is that $75,000, $150,000, or is it $25,000 as costs come down?
You're just asking me an ROI question. People always ask me, “Are the teams bigger or smaller with AI?” And I'm like, “Hmm, okay.”
Imagine you're a normal company and you have $1 million of revenue and $800,000 of costs. You make $200,000 in profit. With the $200,000, you can hire 1 engineer. If the engineer can't make you more than $200,000 in revenue, you don't hire the engineer, and you take the money in your pocket as the business owner.
Now AI happens, and AI means that suddenly that engineer can generate more than they could have before.
Before, they could only generate $150K of revenue. Maybe now they can generate $250K of revenue. So suddenly, AI makes you hire the incremental person—one more person than you would have—because there’s an extra $50K of profit for you to make by hiring the engineer in the new world, because they’re more efficient.
So we’re at a stage of the business where I’m like, there’s gold littered everywhere in front of me. I have customers who want integrations in order to sign the contract. I have people who want audit logs to sign the contract, who want SSO to work with phone numbers to sign the contract, to be bigger.
I’m sitting in front of that. I have a dex product where people want more exports in order to use the product more. It’s all gold, all in front of me everywhere, and my limiters are my ability to hire, how much funding I have, and the growth rate of my revenue, right? Because I don’t want to get too far ahead of my revenue.
So if you told me that there were better models and I could spend more, that’s easier than hiring to do the high-ROI stuff—the money that I’m leaving on the ground. I would do it immediately. That’s the level at which I think about it today.
I think about our global spend on compute—whatever, a million or whatever it is, on an annual basis—and I think roughly it’s four engineers, maybe a little less, like three engineers, all things considered. With equity, maybe it’s more like one and a half engineers.
Am I getting one and a half engineers in Silicon Valley at our inflated rates from the…? Yes, of course I am. So it’s not even close. I’m not even close to the place where I’m like, “Are we token-maxing wrong?” It’s not even in the ballpark.
And then I think the other thing that we don’t contemplate enough is: do we see the tipping point in other categories that we’ve seen in coding, in legal, in sales, and in marketing? Final one for you, JD: What are you most excited for in the next 10 years?
Well, it’s definitely my kids. Growing up with my kids and getting to spend time teaching them things like math and playing soccer with my son—that’s 100% what I look forward to the most.
But that’s not what you meant. You meant, what do I look forward to most in the universe? Well, I am a believer that even though people are very skeptical about AI, and I understand why it may be scary and why any change is hard for humans or anyone to take on, myself included, I do think we’re getting closer to a world where people have more of the things that they want and can do more of the things that they want to.
So I just hope we come out of this with a better universe, more money for everybody, and more ability for everyone to do the things that they want to. I truly believe that. I wouldn’t be doing what I’m doing to make money. I’m hoping that I’m doing it because I hope we can remove a lot of the toil of people’s day-to-day through this technology, not just through Town.
I think AI will help us make drugs, will help us build faster in the physical world, and people will be able to live further away from cities because they can self-drive in, which means they can have bigger houses with pools and be happy. I’m very, very much an optimist.
Dude, I so appreciate you giving the time today. I know it’s a very busy time. You’ve been amazing, and I can’t thank you enough for putting up with my slightly pressing questions at points.
Every bit of skepticism, I would say, is something that does keep me up at night. But I think there are paths through the dark forest, and there’s a giant treasure with only 1 or 2 dragons at the end of it, so you’ve got to go for it.