Nathan Labenz
Hello and welcome back to The Cognitive Revolution. Today I'm excited to welcome back Andrew Lee, founder and CEO of Shortwave, for a conversation about the incredible speed of AI progress, how Shortwave is maximizing agent performance with today's frontier models, and fundamentally reimagining digital communications; the ongoing transformation of the software industry at large; and company building for the AI era.
The impetus for this episode was a beautifully exponential revenue growth curve that Andrew recently posted on Twitter—the sort that you can really only achieve with genuinely word-of-mouth-worthy practical value. Over the next two hours, Andrew takes us on a tour of everything that he and the Shortwave team have done over the last year to transform what was a useful but perhaps not quite transformative email assistant into an email agent that is now routinely surprising and delighting both users and Andrew himself with the increasingly complicated projects it can tackle.
So much so that Shortwave is now expanding beyond email and reconceiving the product as an AI agent to help manage communication across all major channels. Having concluded that AI makes software so much easier and faster to create, such that speed is really the only moat going forward, Andrew does not hold back on lessons learned, and the technical insights here are outstanding.
Andrew breaks down how they've completely rebuilt their infrastructure at every level by constantly testing and swapping in new models, moving to Pinecone's serverless offering for their vector database, and adopting a hybrid structured-plus-vector search paradigm that delivers better results at lower cost. Perhaps most fascinating is Andrew's perspective on agent architecture. While many companies are pursuing multi-agent approaches with specialized sub-agents, Shortwave has found better results with a simpler approach that makes careful use of Anthropic's caching features to support long-running tasks with lots of context while also maintaining positive-margin unit economics for the business, but otherwise largely trusts Claude to act effectively as an agent, both by calling the right tools and determining for itself when it's found what it's looking for.
Personally, I've had a number of wow moments as a user. I was honestly a little nervous to allow it to organize my inbox for the first time, but now I'm making regular use of the conceptual to-do lists that it's created for me. It also saved me a cool 30 minutes the other day when it collected all the receipts from a recent trip and compiled them into a tidy expense report.
In the last third of the conversation, we turned from the product itself to the question of how to structure a company for success in the AI era. Having recently closed another round of venture capital, Shortwave is hiring for a number of roles, which Andrew describes not as traditional individual contributors but as AI agent managers across software development, marketing, content creation, and more. He plans to keep the team quite small, targeting just 15 or so employees for the foreseeable future and prioritizing talent density and speed of execution above all else. With that in mind, he's offering a $10,000 referral bonus, including to listeners of this podcast.
Finally, before getting started, I want to note that this episode is brought to you by Shortwave. I've mentioned in the past that we are experimenting with sponsored episodes that allow companies with a timely story to cut to the front of the line. Of course, my commitment to you, the audience, is that our bar for interesting content and my preparation process will remain the same as always. Andrew and Shortwave were really a perfect fit for this opportunity. Their product is clicking, their business is booming, and he was eager to get his hiring message out sooner rather than later.
As always, if you're finding value in the show, please take a moment to share with friends or colleagues who might be interested, leave us a review on Apple Podcasts or Spotify, and I always welcome your feedback either via our website, cognitive revolution dot com, or by DMing me on your favorite social network. Now let's dive into this fascinating conversation with Andrew Lee about Shortwave's AI-powered transformation, not just of email but now all digital communications.
Andrew Lee, founder and CEO at Shortwave, welcome back to The Cognitive Revolution.
Andrew Lee
Thanks for having me. It's good to see you again.
Nathan Labenz
Yeah, it's been—boy, a lot gets packed into just a year in the AI space. I was looking back, and it's been just about a year since your first appearance on the pod. Indeed, a lot has changed. At the time, I called Shortwave the AI email assistant that I had been waiting for. As we were catching up in preparation for a second conversation, you said that at that time it only kind of worked, and I thought, “Yeah, I guess that's true.”
I look back at these things that seemed so mind-blowing to me at the time, and obviously we've way surpassed them. What caught my eye and got me to reach out again was that you had posted a graph of Shortwave revenue on Twitter. It's basically the canonical exponential curve, where it looks like it's on the verge of going totally vertical. To kick things off today, what's new, and what is working now that was only maybe sort of working a year ago, leading to such incredible growth?
Andrew Lee
It's really been an evolution of that AI assistant that we talked about last year. I think the thing that you played with a year ago, you could chat with it, it could answer your questions, it did an okay job of searching, and it did an okay job of writing your email. But it wasn't really that smart, and it wasn't really that trustworthy. If you asked a question like, “Hey, where is the receipt for this?” maybe you'd find it, but you couldn't really trust that it had found it.
It also couldn't do a lot of the normal stuff that you want to do in an email client. If you went in there and said, “Hey, what are the most important emails I have?” or, “Archive all the cold sales emails that I got,” it couldn't do any of that stuff. It couldn't manage your to-dos. It had no idea who your contacts were or what your labels were. It was just sort of a cool search and writing thing, and maybe a general-purpose AI thing, but it wasn't quite like having an employee sitting next to you, which is the pitch we were trying to make.
We've iterated our way to something that actually delivers on that. You use the thing, and it kind of does what a virtual assistant would do for you. It actually works, and it can do almost everything that you as a human can do. It's reached a tipping point where people are like, “Holy crap, I can do my email not by doing my email, but by talking to a thing that then does my email for me.” That lets me think about things at a much higher level and be way more productive.
Nathan Labenz
Yeah, I've experienced quite a bit of that, and I can definitely testify that there are some pretty astonishing moments. I don't know if you want to start with use cases or with how it works under the hood.
Maybe start with use cases. One that I tried, which I thought was really interesting, was to just ask it, “Take a look at my last 100 emails sent and give me whatever advice you have.” You learn that there's a lot in the outbox that says a lot about you. I honestly thought it gave me pretty interesting advice, which was apt in some ways and also made me think that maybe I'm not spending my time 100% the way I aspire to be spending it, just because of the balance of things that it was seeing.
What are some of the other exciting use cases that you've gotten the most value from, or that your customers have surprised you with?
Andrew Lee
Email is a crazy valuable corpus of information about everything involved in your business or your personal life. Obviously, it has all your human correspondence, but it's got all your status notifications, all of the attachments—all the PDFs and things that come along with that—and all your calendar invites. We just know a ton of stuff about you.
If you asked a reasonably smart human to go through your email and give you some advice, and they had the time to do that, they would also give you some good insights. You just never do that because it's a lot of work, but the information is there. I think the prompt you sent me the other day is a super fun one.
I actually had a prompt this morning. Just yesterday, we rolled out a big UI change, and any time we make big changes in our product, it's controversial. We get a lot of feedback, and it's always scary. I wanted to know, “Hey, how are people responding to this?”
We have this new sharing feature that lets us share all the support threads that come in across the whole team and makes them available to the AI. I asked the AI, “How many people have emailed us and complained about the new layout in the last 24 hours?” I got a report: There were 19 users who reported it. It gave me a summary of the top 5 reasons, and it was super useful and super good. Consolidating that myself would have taken quite a bit of time, but I got a snapshot right away of the reaction, how people were feeling about it, and one of the top things we might need to address.
That's been a big use case for me personally. We've seen a lot of other fun ones in the wild. I think one of the most common ones is that people start their day with a complicated prompt. You're in sales, you have 8 customer demos coming up, and you have an inbox full of emails from people. You say, “Hey, help me figure out what the tasks are that I need to do, what order I need to think about things in, and what I should remember for each sales call.”
They'll have this big prompt. A lot of the time, users will share these prompts with us and say, “This is the thing I use to start my day.” That's a pretty common one.
Another really common one is people doing attachment analysis along with their email. There are a lot of real estate agents, general contractors, architects, and people like that who email back and forth all these PDFs. They just need to answer one question from the PDF, like, “What were the specific payment terms on this contract?” They can ask the assistant to read the whole PDF for them and give them the one answer they need, and then write the email.
There was another example that got shared on LinkedIn the other day by one of our users. He was selling his house and needed an inventory of all the furniture that was in it. All that information was in his email somewhere, either in emails with his wife, receipts that were sent, or whatever. He was able to get a full inventory of all the furniture they had purchased over the years for their house in a nice, concise, accurate report with just a few keystrokes.
There are lots of interesting use cases. They're all sort of big prompts, but they span a wide spectrum of things.
Nathan Labenz
I did a small version of that inventory thing after a recent trip, where I had to collect and submit receipts for an expense report. That was another very mundane task, but in a way it's the perfect job for AI to handle—things that I otherwise kind of dread. It was cool to say, “Find me all the receipts from my trip. There should be Ubers, Lyfts, and a couple of DoorDashes in there,” and have the whole thing pop out with the values and everything itemized.
It was another one of those moments where I thought, “This AI thing could really catch on.” I'm sure there's more. I do think one big barrier to practical value in AI is just a lack of imagination on the user's part. I certainly feel like I'm guilty of that all too often. When I see something somebody else has done, I'm like, “Why didn't I think of that sooner?” Are there any others you want to share at the top here that should get people's wheels turning about how they might get more value if they're a little more intentional or creative?
Andrew Lee
I'll share one of the most creative ones that I ran into. This was another one that a user sent to me. I'm personally still figuring out how to use the thing, and often the insights are a user sending me something like, “Why doesn't this work?” or, “Check out how this thing works,” and I'm thinking, “I never thought of that.”
Here's one I never thought of. There was a user who was using us in combination with another SaaS tool—I think it was Linear. They wanted to extract action items from their inbox and add them as tasks to Linear. We don't have a Linear integration. Maybe we should, but we don't.
It turns out that LLMs actually know the structure and the URLs needed to create tasks in these other SaaS products. If you ask the LLM, “Create a link that, when I click it, will go and do this task in this other product,” it actually works.
He had a bunch of custom prompts where he said, “Extract action items from this email thread, and then give me some links that I can click. Each one should create a thing. Here's my base URL for my project and this other information.” The LLM spits out a bunch of links, he clicks all the links, and boom—he has tasks in this other product. You can build integrations with other things without us doing anything and without any code being written. It's just a prompt and the LLM's knowledge of how to construct URLs.
Nathan Labenz
That's really creative. I like it. It's also an interesting window into a future in which AIs are increasingly likely to solve problems in unexpected and maybe, at some point, even hard-to-interpret ways. That one is pretty simple to interpret, but the creativity there, on the user's part and the model's part, is pretty impressive.
Andrew Lee
I'll give you one other example that I thought was pretty fun. There was a user who wanted to do a mail merge, but he wanted every email to look custom. He wanted the AI to search his email history, find an interesting fact about each person, and make that person feel like he had really been paying attention to them.
He took a text file, put a bunch of email addresses in it, uploaded the text file into the Shortwave assistant, and said, “Loop through the emails in this text file. For each email, go and search for emails that I sent to this person. Figure out a nice greeting for this person, and write an email.”
He had it loop through the file and write a very custom email for every single person. He could review each one, click Approve, click Send, and send 20 emails that all felt very custom in a really organized, fast way with the AI.
Nathan Labenz
That's cool. I was actually going to ask you about supporting loops, but I hadn't tried it myself, so it's cool that it's already something you're seeing work in the wild.
Last time, we got quite deep into the guts of how it worked. You have a background as a database whisperer, and there was a lot of talk about how the index happens. You take every email out of somebody's inbox, store it in your own system, and have your own indexing system, and then the model can build on top of that.
If we work from the ground up, how much change has there been at that foundational database and retrieval layer, versus how much has come at the higher layers of the model and the patterns of behavior that you're getting the models to exhibit?
Andrew Lee
Honestly, between the time I was on the pod last year and now, basically every part of that stack has been completely rewritten. We're using a different embedding model, we're using a different vector database, the search stack has been completely rewritten, the API to that search stack has been completely changed, the models we're using for the agent are different, and the agent code has been completely rewritten.
It's been top to bottom, driven by changes in the capabilities of both the models and our understanding of how to apply those models. Every time the model can do something better, we're like, “Actually, if we reroute the system like this, we get this extra unlock.” The thing is evolving at a truly phenomenal pace.
Nathan Labenz
That's undeniable, although some people do try to deny it. Let's go a little deeper on each of those levels. I'd love to hear what you've learned and, to the degree that you think it's applicable for other AI builders, what sort of advice comes out of these changes.
At the database layer, do you have a favorite vector database that you would recommend to others at this point? What have you learned that caused a change there?
Andrew Lee
For a little bit of context, one of the big unlocks for us here was having a model and an agent on the front end of our app that was able to reason a lot better about how to use search. Most importantly, it was able to reason about how to run multiple searches.
It used to be that the model would run 1 search and give you a semantic component in the search. The search stack had to find the right email in that 1 search because you got 1 shot at it. Now that we have an agent that can run multiple searches in parallel, or run them in sequence—so it can try a search, and if it doesn't find what it's looking for, it can try something else, or it can try a search for 1 thing, find some information, adapt, and try another search—that's allowed us to simplify the vector implementation and focus it on a much narrower task. We can then make it much better, faster, more reliable, and cheaper at that task.
The evolution of our search stack in the back end has been driven, in some ways, by a simplification of the requirements. We've focused on how to make it really good, really fast, really reliable, and really cheap to do this narrower task.
We use Pinecone's serverless offering as our database. We used to use their pods, but the serverless offering is a much more cost-effective solution because it separates storage and compute. You can tune it for your use case, and we use a lot of storage because we have a lot of email.
We use an embedding model called BGE, and we use a bigger embedding model now than we used to because the serverless offering unlocked the ability for us to use these bigger vectors without spending too much money.
One of the big changes is that we started to use hybrid search on the back end. We used to use a pretty complicated pipeline that combined some smaller-model LLM calls with particular types of feature extraction and search. There was a reranking step. The whole thing was slow, complicated, and very brittle because there was a lot of custom stuff in it.
Instead, we've allowed the search API to specify a semantic component and some constraints. The constraints are normal email query constraints: You can say, “I only want emails in this date range,” or, “I only want emails with this contact,” or, “I only want emails with this label.”
For the semantic component, we run an embedding search with Pinecone. We also run a keyword search, and there's an algorithm to combine the keyword-search results with the semantic component, get a score, and rank based on that score. Then we filter based on these other components.
You get an API where you can say, “I want emails about this topic with these constraints,” and we can very quickly give you a very accurate list in a way that's cost-effective to scale across everybody's email. This has been huge for us. It's a lot cheaper to run, a lot faster, a lot more reliable, and it's producing much better results when you combine it with an agent that can reason about running multiple different queries.
Nathan Labenz
There are multiple things there that I think are quite interesting. First of all, the fact that the whole stack obviously has to work together, but also that an improvement in 1 layer of the stack allows for simplification in another layer of the stack. That's definitely been a huge theme for me over time with the work at Waymark.
The shenanigans and hoops that we had to jump through to choose an image for a user once upon a time were almost comical in their complexity. They also didn't work that well. Now it's just, “Feed it into a model,” and it usually picks pretty well. We've cleared out so much old cruft that we developed to get that first version working, and now it's a simpler and better solution driven by model progress.
Did I understand correctly that when you do this search, the filter comes after that? It seems like you're doing the full vector search through the entire corpus and then filtering only after getting results back.
Andrew Lee
It's a bit more complicated than that. I'll give you the somewhat simplified version. Essentially, we run 2 types of searches. One is a full-text search that's constrained based on the keyword and metadata components of the filter. The other is a semantic search that we run in Pinecone.
We combine the results of those 2 queries, score them, and then post-filter them. That means there are situations where, if you had a lot of good semantic results for something, you could potentially miss the best ones. But in the vast majority of real use cases, the way we're combining the full-text and metadata-constrained search with the semantic portion and then post-filtering is generally finding the best results.
The truth is, we aren't actually going to score every email across everything and apply these filters. If you really want to have that be 100% accurate all the time, that would be an intractable problem. We can approximate it very closely with our solution.
Nathan Labenz
Interesting. It seems like model progress has driven a lot of value for you, but the scaffolding and shaping the behavior of the thing are also super important.
I've seen that repeatedly in my usage, both when I accept the invitation for the AI to organize my inbox and in some random, idiosyncratic things I've done. There was one where I thought, “The AI and I are really on the same wavelength.”
For whatever reason, I have a total inability to remember which of my contacts is the stronger and which is the weaker prescription. Every time I go to change my contacts, I have to search in my email for a particular thread where I know that information is contained. I'm always doing this search.
With the traditional approach, I would end up running multiple searches. I'd think, “Okay, I think the word ‘prescription’ was in there, and I know the word ‘CorEd’ is in there, but what is it?” Eventually, it would take me 2 or 3 searches to get there.
I asked the Shortwave AI assistant to do that, and it was interesting to see it go through the same process I went through. It would search, get some results, realize it hadn't found the thing it wanted, and search again. It took 3 or 4 rounds before it finally landed on the right thing and answered my question.
More often, I think we'll engage in higher-level things. Organizing the inbox is a great example. It goes off, comes back, and gives you suggestions and an opportunity to confirm the movements that it suggests. When you accept, it just keeps working.
Maybe take us through how you've thought about creating the agent behavior. What have been the big unlocks at the model level that have made this possible, and what have you had to do to wrangle it into something that's actually valuable for people on a day-to-day basis?
Andrew Lee
This has been the biggest shift in our product since we last talked. I mentioned that every part of the stack has been rebuilt, but the basic functioning of the assistant has gone from a single LLM call that produces the final output, plus a complicated Rube Goldberg machine to get the prompt right for that 1 thing, to something where we just run the big LLM a whole bunch of times, repeatedly, until we get the right answer.
We had all these complicated characteristics and rules and whatever to do this, along with smaller LLM calls, and we've thrown all that out. We said, “What if we just run the big LLM repeatedly until it gets the right answer?”
I'll go through the history a little bit. About 2 Decembers ago, I think OpenAI rolled out its tool-calling features in GPT-4. We tried them, and we didn't think they worked right. Our general experience was that as soon as you tried to get it to call a tool, the whole model stopped reasoning well and gave you bad answers. It was much better to tell it to format something in XML, and then we'd go and do the tool call for it.
Even getting it to do multiple tool calls didn't work very well, so we kept this sort of rules-based system. Then, last summer, we tried this again with GPT-4o, and it kind of worked. We said, “What if, instead of this one-shot approach, we had a multiple-shot approach?”
Originally, this was very much geared around search. The search use case of running multiple searches seemed important. What if we ran multiple searches? We rewrote it a bit to do that, and it worked better, but not dramatically better.
We had a launch in September that was, “Hey, here's our new agent.” I think we called it our V2 agent. It definitely got us some growth and excitement, but it wasn't a major thing.
Then, in October—or maybe November—I was listening to a podcast from the founders of Bolt.new.
They were talking about how they built their stuff, and my memory from that podcast is basically that they said, “Hey, you’ve got to use Claude 3.5 Sonnet, the October version of it. It’s different.” And, by the way, Bolt.new is open source—these parts of it are—and you can kind of see what their prompt is and how it works. I found that very interesting, and so I tried Claude 3.5 Sonnet specifically for tool calling—the October version of it—and it was dramatically better.
It was able to do what GPT-4o could do: you could have it call a tool, it would spit out a tool response, you would feed the tool back in, call it again, and it would spit out the tool response. You could have it iterate a few times, but it would sort of go off the rails after too long. We didn’t want to do too much of that, but with Claude 3.5 Sonnet, it could go on for a long time. It could run many searches and do many things and still stay on track and keep reasoning, and it seemed totally different.
So we were like, “Okay, maybe we should rewrite this whole thing and say we’re going to have this new approach. We’re just going to call Claude over and over and over again. We’re going to let it run not 2 or 3 or 5 times, but 20 times, and we’re going to put all of our smarts into really good tools, a really good prompt for the overall thing, and a nice agent framework around that. We’re just going to let the model reason about what data it wants to pull in.”
This was a total rewrite. We had a lot of very custom, email-centric stuff, and we said, “No, no, we’re going to have a fairly generic agent framework at the core. We’re going to build a whole bunch of really nice tools around it. We’re going to have a lot more tools than we used to have, and then we’re going to iterate our way to a solution.” This is the agent that we launched in January. If you look at our—you mentioned the hockey-stick growth graph—that little kink is the V3 agent, where we rolled this thing out, and it was just a whole lot smarter. It could start solving very general, very open-ended things, and there were a lot of places where the old version didn’t work that iteration solved.
I think that’s the core. When people are talking about agents, I think what they’re really talking about, if they actually have a working solution, is iteration. It tries a thing; if it works, great. If it doesn’t work, it tries another thing. Sometimes that’s running 3 searches in a row; sometimes that’s trying to run a search, finding out that the search criteria you specified is malformed, having our system throw an error, having the LLM see the error response, and try again.
Sometimes it’s trying to schedule a calendar invite, realizing that the calendar invite is conflicting with somebody, having our system spit back some information saying, “This is conflicting,” and then trying again. So you get this feedback mechanism where, over a series of many LLM calls, you can iterate your way to an answer that no single LLM call could have produced and get really, really good results that way.
Nathan Labenz
Well, okay, I’m going to dig in a little bit more. We have time. It would be a very natural time to talk about the expanded view of the company and the product, but before we get to that, 1 thing that I’ve just been wondering myself is how a model like this deals with 1 of the challenges with RAG apps. A lot of people have built RAG apps and have come to some frustration with them, where I think often the core problem is, as you said, you do 1 search, and then the model just kind of has to do the best with what it’s given. If the right information isn’t in there, then it’s really going to have a hard time.
I have definitely felt that allowing for iterative search makes a qualitative difference in terms of how likely it is to be able to pull the thing back. I don’t think it would have found my contact prescription a year ago, but I still wonder: how do you think about maximizing that, given that the models don’t have the thing that I have, which is, I know it’s in there? I know when I’ve found it, and this seems like a sort of fundamental challenge from the model’s perspective: how do I know when to call it? When do I decide that this amount of information, or this particular information, is in some sense satisfying or the best I’m going to get, versus feeling like the way I feel, which is, “I know that’s not quite it yet, but I know it’s in there, so I’m going to keep looking”?
When I do find it, it’s always very clear to me: “Yes, this is what I was looking for,” and I know it with high confidence. The models don’t have that. How have you approached the problem of leading them—or, I can’t imagine that they’re entirely just doing it on their own—so how do you help them decide when they’ve found enough versus when they need to keep digging?
Andrew Lee
So I’m going to give you kind of a complicated answer here. The first thing I’ll say is, the models actually do have some of that. They may not know what you know specifically, but they know what it looks like, generally, to have found the thing you’re looking for. They have a lot of training on emails in general, what emails look like, and what people expect from emails.
We found—for example, we have this organize-your-inbox feature, and it goes through and finds what we describe as low-quality emails and gets rid of them. I’ve been shocked at how little work we have to do to describe what a low-quality email is. For example, it’s really good at spotting cold sales emails, even though you might think, “How do I know it’s a cold sales email? Maybe this is somebody I know.” But there’s something about the tone and style of cold sales emails that really stands out to them, and they spot it. I found that to be quite remarkable.
We do have some prompting work to make this really clear, but we didn’t have to do a lot to tell it this. So that’s my first thought. My second thought is, we do try to give it tools so that it can go look up this information itself. If you’re like, “Hey, I’m looking for an email from some important investor or whatever,” it maybe doesn’t remember who your important investors are, but we have a contact tool. When it returns results, we put in there, using statistical information, how important this contact is.
It can actually look up who is important if it wants to. It can also just go search your email history and find emails you sent and make assumptions about, “Hey, if you’ve recently exchanged emails with this person, they’re probably important.” So this is an area we’re trying to get into more and more: if you ask it some open-ended question, use tools it already has to reason about it in some reasonable way—look at your email history, look at your contacts, look at your calendar events, ask, “Have you met with this person recently?”—and try to get some of that context.
The third thing I’ll note is, I see a lot of room for improvement here. Today, we don’t do anything to capture the triage actions that you take. When you use organize my inbox in the morning, we accept some of the suggestions and you don’t accept some of them. What we really should be doing is remembering all of the things you did and then customizing it to you over time. Today, we don’t do that.
If you want us to change the way we do triage for you, you can actually tell the AI. You can say, “Remember, never archive newsletters from this sender,” and we’ll remember that, but you have to be explicit. We don’t remember just based on the actions you take. I do see a lot of opportunity there.
Nathan Labenz
Yeah, that’s quite interesting. I just did an episode with the chief scientist, Guy Gur-Ari, at Augment, and there are multiple similarities, actually. They also have a big emphasis on just ingesting a ton of code right off the bat and putting it into this specialized index. But on the behavioral part, they’ve developed a process they call reinforcement learning from developer behaviors. It’s about both observing what the developers are doing and how the developers are reacting to what the AI is bringing to them, and it sounds like it’s working quite well for them. It seems like reinforcement learning works, so I see that probably coming quite soon in your future, too.
On just another kind of random question, but an interesting one to me at least, Claude is obviously significantly more expensive than some other options. What I see typically from the agent—I don’t know if this is super consistent or just my observation—but typically it seems like it’s 10 threads found, and then it kind of does its reasoning: do I need to search again, whatever, based on the 10 things that it’s found.
The other way I could imagine doing that would be maybe using a slightly lesser but way more affordable language model to do the evaluation on stuff that came back. If you used Gemini Flash, for example, I think you would be able to handle 30 times as much, and you could maybe complicate this with caching. I’m not exactly sure how you’re using some of those more advanced optimizations, but first-order approximation, you could handle an order of magnitude more search results if you put it all through Flash to assess relevance versus saying to Claude, “You handle this all yourself.”
Are you doing any sort of that kind of ensembling, or using different language models for what they seem to specialize in, all things considered, or is it just all Claude all the time because it really is just that good?
Andrew Lee
So we do use a bunch of different models, but we use them for different features. For example, our autocomplete is actually using a fine-tuned GPT-4o mini. Our quick-reply suggestions—the 1-button reply suggestions that we give you—are actually using Llama 3.2, the 3B model. We use GPT-4o in a few places—not many. So we do use different models in different places for different things.
We have looked at stringing them together, trying to save costs for certain things in the assistant by outsourcing to other models, and I think our experience has been that it introduces a lot of complexity and affects the ability of the model to reason across lots of different types of activities in a sane way. If you just feed all of the data into 1 big model, it can think about lots of complicated relations between different things, and it can come up with ways to use tools in ways you didn’t think of. That’s been a big unlock for us, and the cost savings haven’t seemed worth the complexity and the loss of generality of having this turn into more of a pipeline between different types of models.
The other thing I want to note is, the caching that Anthropic has is really, really critical for us. It’s a really big deal. When you have this agentic flow, you call the same history over and over again; you only append to the bottom. Our histories get very long: we can have hundreds of thousands of tokens at any 1 of these calls.
Anthropic caching is a little hard to use. You have to really construct your agent to be very careful about keeping earlier things immutable, and we’ve done a lot of work to make that happen. But if you get it working, it can save you 90% of your cost. It’s a hugely impactful thing, and frankly, if we didn’t have that, we could not afford to run this. At Anthropic’s cost, we’d be losing money on every user by a huge margin.
So that’s been a huge unlock for us. It also is 1 of the reasons we switched to Claude, actually: even if the models from OpenAI could use tools as well as what we have from Anthropic, we couldn’t afford it, because the caching from OpenAI is a lot less in terms of cost savings. It’s only like 50% off.
Nathan Labenz
Yeah, 50% off and pretty much automatic. You don’t have to do anything, right? Versus—can you tell a little bit more about how? I know Anthropic is 90% off once something is cached, but there’s also a 1-time cost to get something into the cache, right?
Andrew Lee
Yeah, but it’s not an additional cost, right? At the time that you run it, you have to tell it to cache, and I think there may be some slight additional cost. It’s not a significant additional cost, and then every iteration after that is 90% cheaper. If the most common action is people organizing their inbox, which it is, and that’s 20 tool calls, it adds up really quickly.
Nathan Labenz
I was actually just at OpenAI yesterday, and I was talking to the Agents SDK folks about this, and about caching. This was, frankly, the number 1 question from the room—from the other founders I was with there. It was, “What about caching? How do we get caching to work?” I think everyone’s figuring out that you’re going to be calling these models with the same context over and over and over again, and making it work efficiently really matters a ton.
The OpenAI caching, you’re right, is a very simple API. It sort of automatically figures it out; you automatically get cost savings, which sounds great. The Anthropic one is much harder to use, but the magnitude of those savings is quite different. On the Anthropic side, is it smooth to the point where you can sort of build a cache? I’ve used the caching a little bit, but it sounds like there may be some possible implementation difficulties. Can you extend the cache iteratively as you go, and does it basically have all the features you really need to realize the full savings off the sticker price of 90% off?
Andrew Lee
Yeah, you can. You sort of checkpoint at every iteration in your agent flow, saying, “Okay, I cached up into this point,” and you have to construct your agent to make use of that. But it’s totally doable.
Nathan Labenz
Interesting. So are they actually caching? I wonder if there’s a big optimization there for them under the hood where they’re caching multiple versions of each conversation. I don’t know. I actually am very curious what they’re doing. Maybe they’re just eating the cost. I don’t know, but our bill goes way down, so it’s great.
Yeah, that’s interesting. A huge question in the AI space generally, to me, is: are all the frontier providers converging, or are they diverging? We’ve got somewhat different caching things, but maybe OpenAI is going to get the message and do a more Anthropic-like one. We’ve also just seen in the last day or 2 that OpenAI is going to embrace the Model Context Protocol that Anthropic led the way on. Do you feel like these things are converging, diverging, or some complicated mix of those?
Andrew Lee
I think it’s some complicated mix of those. It has been nice that they’ve been converging in terms of the API surface. We now can very easily swap between different providers, and all the tool calling and stuff lines up perfectly and makes it really easy to do that. I think that’s been a very nice thing. Everyone is understanding that there should be caching and are hopefully converging on some of the ideas of how to do that. MCP is maybe becoming a standard that everyone’s respecting, so I think there’s been standardization in the interfaces.
It does seem like the different labs are focusing on different things. It seems to me like Anthropic is really focusing on this iterative approach. It seems to me that OpenAI is maybe caring a lot more about reasoning and potentially more about multimodal. You’ll notice that we use many different models—we use models from 3 different vendors—and a big part of that is I feel like each vendor is bringing something different to the table.
1 example where I think OpenAI totally crushes it is that their serving infrastructure is super performant and reliable relative to Anthropic’s API. They go down a fair amount, and time to first byte is slower and stuff like that. For example, for our autocomplete, where latency is really critical, we use the GPT-4o mini with a fine-tune, and the serving stack is super good. So I think that’s a differentiator there.
We use open-source models for the summaries and for the instant reply, like quick actions, and for a lot of the stuff where cost and latency matter a ton but we don’t need a lot of intelligence. We use Llama 3.2, and that’s running on Vertex on GCP. So it’s great to have, as an AI app, many different vendors converging on standards but focusing in different areas. That ends up being really nice for us.
Nathan Labenz
What else have you learned about agents in this journey? I mean, everybody’s setting out to build 1. I think you are pretty far ahead of most. Have there been false starts—things that you expected to work that didn’t—or just any kind of unobvious lessons learned that you would share with folks who want to build their own agents?
Andrew Lee
I think the biggest thing I would say if they’re thinking about it now is that things are different, just in the last few months. Basically, if you tried this before October of last year, or if you’ve never tried the Anthropic models, it’s different now. The stuff that you tried before actually works. The cost, with the caching stuff, is manageable, and if you try it again, I think you’ll be really surprised. You’ll notice stuff like Bolt.new.
Bolt.new has gone totally crazy, and the new Cursor agent mode is totally insane. I think a lot of people are still using Cursor for autocomplete, but that's old school; they should use the agent mode. So that's the big message: it works now, it's different, you've got to try it out, and it's going to be amazing.
I think one of the big things we've learned is that people have no idea what the UX should be around this stuff. We're very much still figuring this out, but people kind of figured out, “Oh, autocomplete works like this,” and everyone wrapped their heads around how this sort of thing should work: you tab to complete. Now we have a thing where it's going to go off and just do a bunch of work for you. How do you do that in a way where the user doesn't get really uncomfortable about whether it's sending emails for them or changing their code in ways they don't realize?
We're going to need some UX for oversight and approval, and for writing guardrails on these things, that doesn't get in the way but still gives you confidence it's doing the right thing. I think that's going to be a big area for us. There's just a lot of questions and unknowns, given how new this stuff is, but I think it's going to be really exciting to see how it unfolds.
Nathan Labenz
One product feature that I honestly haven't used as much as I maybe should, partly because I'm not sure quite what it's going to do, is the AI filters. I was curious to know how that works under the hood. Is it sending all the incoming emails through a model? What I think is going to happen is I'm going to say, “Filter out this kind of email,” and then emails are going to come in. I guess literally every email would get sent to a model to say, “Is this the kind of email to which this filter should apply?”
I like the sound of that because I get a lot of crap that I need to filter out, and historically I've definitely spent way too much time clicking on those myself. The flip side of that, of course, is that if I create a workflow where I don't have visibility into what's getting filtered out, I have been burned by that in the past as well, even with Gmail Priority Inbox over time.
First of all, am I intuiting correctly what’s going to happen there? More broadly, how are you thinking about these next-level agentic scenarios? Today, the AI system brings me stuff and it's up to me to confirm, but with all the talk of how well it's working, it's not hard to extrapolate a little bit and imagine, “Okay, fine, just go ahead and do it.” How are you thinking about those next moves into actually taking actions without necessarily having a human in the loop approving each one?
Andrew Lee
No, I think you've caught an interesting difference in that feature: it is the only AI feature we have where the AI just does things, and there’s no approval flow from you. It's actually a really popular feature. The basic idea is that you can write a prompt and choose 3 actions based on that prompt: you can apply a label or delete it.
We've had this live for a couple of months now, and people love it. There haven't been a ton of problems in terms of people losing things, but we do get support requests sometimes from people saying, “Hey, you're missing my email.” Then we check their AI filters, and they have one set up. I don't think there's any way of getting around that entirely, but the dream here is that this becomes a version of the full agent.
Right now, it's a very simple implementation. We do use OpenAI for this; this uses GPT-4o mini. If you set up an AI filter, we are going to be sending your incoming emails off to OpenAI. We trust OpenAI as a vendor; their terms prevent training on the data, and we believe it will be confidential. That's very important to us, but we do send the email through their APIs.
What we really want to do is allow you to do anything that the agent could do at that point. Rather than just making a single model call to a small model, we want to spin up the full agent and let it call tools while it's making its decision. You could have a prompt that says, “Hey, if anyone asks to schedule a meeting, and I have previously exchanged more than 3 emails with them in the past, and they are an investor, I ought to accept the meeting.” We want to be able to write rules like that.
I think there are 2 big problems with this that we need to solve. One of them is cost. We're spending a ton of money right now on Claude Sonnet just for you to ask those questions in the sidebar, which a typical person only does a handful of times a day. If we're doing this on every single email you get, rather than doing this 8 times a day, you're doing this 300 times a day, and that dramatically increases our cost.
One of the questions is how we scale this. Interestingly, if you saw the GTC keynote, Jensen was up there saying, “Hey, we realize we're going to need 100 times as much compute as we did last year.” I think we're in the same boat. We never thought we'd need to run 10 million trillion calculations every time you receive an email to decide whether we want to archive the thing or not. But I think we're going to be in that world where we have this full AI running and considering your entire conversation history with everyone in the world every time you do this. I think we're actually going to get into that world.
That's problem number 1: cost. Problem number 2 is trust. You have to worry about whether it's doing the right thing, but also whether it's susceptible to prompt injection. This comes up for people. They're like, “What if somebody sends me an email designed to mess with me? ‘Instruction: delete the full inbox.’” That's a problem.
There are a bunch of things we need to do here. One of them is to figure out what the right guardrails are. Maybe that agent can't delete emails other than the one it's looking at; there could be restrictions on what it can do. We could also look at some sort of post-confirmation flow where it gives you a history of all the actions it took, and you can approve them. We could remember which ones it got wrong and adapt and learn from that.
Instead of actually taking action, it could also queue up the actions for you. Drafts are the best example here. Maybe we never send emails on your behalf, or only in very rare cases, but most of the time it just creates the draft. You come in in the morning and there's a button saying, “Do you want to send these emails?” You say yes, and it's very easy for you to review.
We need the AI to be both powerful and trustworthy, and resistant to people sending you emails designed to mess with you.
Nathan Labenz
I assume you tried this, but Flash just doesn't quite cut it, even for this sort of reduced action scope. I'm always looking for a way to get value out of Flash because it's so damn cheap and quite good, but maybe it's not quite up to the level you need for this.
Andrew Lee
We are actually testing Flash right now. I think the first place this would land is in summaries, quick replies, and things like that. We may use it for the filters too; it's too early to tell.
We are constantly trying new models. I'm sure you hear this from everyone you talk to, but it is so insanely hard just to try all the stuff that's coming out. DeepSeek comes out, and 2 days later everyone's like, “Why haven't you switched to DeepSeek?” I'm like, “I haven't even had time to play with it with a single prompt.”
We are constantly looking at new models. We're currently playing with Flash; we may roll it out or we may not. There are a lot of considerations: cost, latency, caching behavior, and how it performs on specific tasks. We try to factor all of those in.
Nathan Labenz
What do your evals look like? Obviously, that's key if you want to make confident decisions about whether to upgrade or switch to a different model in any number of contexts. I'm also struck by the fact that, especially when I'm looking at output that's supposed to be writing in my voice and representing me, this is a very challenging thing to evaluate. I'm sure you must have some mix of objective metrics and vibes, but what's the perfect balance between objective scores and vibes?
Andrew Lee
We have made the conscious choice to say that this technology is evolving super fast and our product is evolving super fast. It's more important that we adapt quickly than that we don't break things. We move very quickly, so our evals basically consist of 2 pieces, and they're very seat-of-the-pants.
Piece number 1 is that I have a Notion Doc of golden test cases. I run those in my inbox and make sure they do the same things. When we're tweaking prompts, I go and try the relevant ones to see whether they still do what I expect or do something reasonable.
We don't want to lock it down to specific behavior because often it gets better. For example, we recently added a tool to unsubscribe. Magically, without us touching anything, any time it did an inbox-organization task and came across an email that looked low-quality with an unsubscribe link, it would start offering to unsubscribe. You don't want to have a test that expects “organize your inbox” to do a certain thing, because this is better.
You'll notice now that if you organize your inbox, it starts offering to unsubscribe. That's something we didn't really think about when we were first building the feature. So we have a Notion Doc, and I go through it and ask, “Did this produce reasonable results?”
Nathan Labenz
How many prompts, by the way, just to calibrate myself?
Andrew Lee
A little more than 100, maybe, at this point.
Nathan Labenz
That's a little more than I expected. You run them manually?
Andrew Lee
Yeah, it's just a Notion Doc. I don't necessarily run through every single one every time we make a change. I'll look at the ones that are relevant to the areas we touched and try those to see whether they do reasonable things.
The other thing we do, which I think is the more objective metric, is that we have an experiments framework. Every new change—every big new change—that we roll out, we provide as an opt-in experiment for our users. Our users tend to be very forward-thinking people who like to play with this stuff, and they'll turn it on in large numbers. We look at the retention statistics.
For the unsubscribe feature, for example, we had 99% retention among people enabling it. When you enable it, it modifies the prompts and adds new tools, and it could break all kinds of stuff. But after leaving it on for a week and seeing 99% retention for the feature, we're like, “Clearly, this isn't breaking a lot of people.” People would start turning it off if it were a problem, so we feel comfortable rolling that out.
Anything that's a major change, we put into that format. We watch the retention statistics over time, and the assumption is that if retention is high, it's probably working pretty well for people.
Nathan Labenz
Have there been—putting it lightly, I mean, that's a society-wide phenomenon these days, by the way—when things do go out, is that also just something that you've made a strategic decision to live with? There is no substitute for Sonnet, so if Sonnet is down, we're down. Is it as simple as that?
Andrew Lee
We could fall back to GPT-4o, and it might give us decent results, but there's always a question of whether it's worth putting people through that disruption to deal with a short downtime or whatever. So far, we've just eaten the downtime and waited for things to come back up, and we don't seem to lose users because of it.
Nathan Labenz
Yeah, I think this is something I've also been thinking about. It's a little bit of a hobby horse for me, even with my own company. This is just the new normal. We're going to be more dependent on these services. They have pretty good uptime, but it's not perfect, and there's really not much we can do about it. If it's out, it's out.
I'm with you on that, and I think in general there's a mindset shift that's needed: let's aim to create the most magical, valuable experiences we can as often as we can, and live with the little risk, whether it may be an outage or just some uncanny-valley behavior. I still see it, but I think for everybody it's worth taking a little bit of that risk to get those special upside moments, because when they happen, they're incredibly valuable.
You do have a nice benefit, too, in that worst-case scenario, I can always open up my Gmail and access things directly that way. One thing I noticed about the new to-dos made me wonder whether there's a deeper philosophy underlying this. I went to organize my inbox, and the first time I did it, the assistant started coming up with to-do categories for me and suggesting, “You should group these into this section.”
That doesn't seem to sync back to Gmail at all, as far as I can tell. Then there are labels, which do sync back. I guess I'm wondering to what degree you've found it advantageous to have a single ground truth, where something you do propagates back into the core Gmail account, versus building something a little off to the side.
I honestly kind of like it better when I'm like, “Okay, this is my Shortwave universe. I can let the AI assistant do its thing and run a little wild.” Worst-case scenario, I can always go see my old view. But I don't know what users in general want. Do they want a unified reality, or do they want you to build something a little bit off to the side that, in a sense, de-risks them from anything that could happen to the core data store?
Andrew Lee
I have a bunch of thoughts here. One of the big conceptual changes we've made since the last time we talked is that, when I talked to you last time, we were building an email client with AI built in. We don't think of ourselves that way anymore. We think of ourselves as an AI with email features built in.
The plan in the medium term is to integrate with products that aren't even email. Basically, anything with an inbox—your Slack, your LinkedIn, or whatever—you should be able to access and manage from this interface. You'll notice that we moved the AI to the left. The big driving force behind that is that we see it as the main product.
We see people coming in primarily because they want to interact with that AI. The AI can work with their email inbox, it can work with other inboxes in the future, and it can work with the CRM, project-management tools, or whatever else you want to use. We're thinking about it from that standpoint, and in that world, keeping everything in parity with Gmail doesn't necessarily make sense, because you might not even be using it with Gmail. You might be using it with some other product.
The to-do concept is here because we see a need for the AI to be able to add and manage state, specifically for short-term organizational purposes. Labels are a great tool for long-term classification. You can say, “I want to apply this thing, and then 2 years from now I want to search for this characteristic,” and you want that label to be short, simple, and easy to remember.
To-dos are great if you have a project that's happening right now and you want a name that's a whole sentence, a bunch of notes in there, and a bunch of email threads attached to it. You want a much more complicated but much more short-lived type of thing.
I think the AI shines with this. Let's say you're producing your podcast with me, and you and I might have 2 Google Docs and 5 threads about this. The AI can spot that they're all related, make 1 to-do called “Prepare for podcast interview with Andrew,” put all the relevant things in there, and add some notes.
We don't sync that back to Gmail, largely because there's no concept like that in Gmail. We don't sync it back to labels because it's kind of a different thing. I do think people like the idea that we're not trying to shoehorn new features into old Gmail concepts in a way that might mess things up.
A counterexample is one of our competitors, Superhuman. Superhuman implements some of its features basically by sending emails. If you use the reminder feature or whatever, and then go back to Gmail after you stop using Superhuman, you end up with all these extra emails, which kind of looks cluttered and weird.
We try to avoid doing that so that at any point you can leave and go back to Gmail, and everything looks the way you saw it before. It's an easy switch—not that we ever want you to leave, but we want you to be comfortable that we're doing the right thing with your inbox, because we know how important it is.
Nathan Labenz
Certainly, I think it makes me more comfortable taking a leap, and that's a big part of what people need to do to get value from AI: be somewhat willing to take a leap.
A couple more little product questions, and then the big shift, which I think is honestly super exciting and valuable. I've tried over the last 2 years to fine-tune a model to write as me, and one thing I've really learned is that nothing has beaten dumping a lot of writing samples into Claude and asking it to do the next thing. I've learned some things by trying.
One of the things I've learned is that my life—and I assume other people's lives are probably similar—is not represented by any single system. I don't email back and forth with my wife; I would need text messages, really, because that's where that kind of relationship lives in text form. Slack is key for a lot of the day-to-day planning and discussion of what we're going to do, and some of the more interesting conversations happen in Twitter DMs.
It's not just that there's no single source of all the information; there's honestly not even a single thing that represents all the different facets of my life across these channels. To be able to unify that into a single thing and have an AI system that could span all of them is quite exciting.
When you do Write as Me, how is that going? I've noticed an improvement for sure. It seems that it now handles routine stuff at least pretty well. For classic things like putting a meeting together—“Here's my Calendly link”—I'm increasingly comfortable just hitting send on the assistant's draft.
Obviously, as it gets more specialized and context-dependent, it gets much harder. What have you learned about making Write as Me work, maybe on multiple levels? What do people really care about? How representative of me do they want it to be? And, in terms of techniques, what's getting results?
Andrew Lee
From a techniques perspective, one thing I very strongly believe—and I think this is a little controversial with other people I've talked to—is that pasting examples into Claude and having it use those examples to produce the response is the correct solution. It's the best solution. That doesn't fit well with people; they like to assume that the more technologically interesting solution, with a fine-tune and everything, is going to produce better results.
I'll give you 2 reasons why I think this is the right solution. The first is that this technique is much better at recalling specific facts, as opposed to a shallow fine-tune. When you're writing an email, or when you have the AI write an email for you, you don't care so much that it sounds like something you would have written. You care that it's correct.
If you provide a link to someone and they say, “Where do I go to pay for this product?” you don't care if the link looks like something you would send; you care that it's the right link. If you're scheduling a meeting and it picks a time that you mentioned in a previous email, it shouldn't just be a time that you would send; it should be the right time. If you give it specific examples and frame those correctly—“The last time you talked about this topic, this is what you said”—it can get the facts, links, and times correct. It isn't just sounding like you; it's being accurate.
Point number 2 is that the big problem with fine-tuning is that it slows down your ability to update models. If you've built a bunch of fine-tuning infrastructure for a particular model and then decide to change models, you have to fine-tune everybody again. That's a big process and a big migration.
The reality is that we're switching models constantly. Often, the leading-edge models don't have fine-tuning features built in initially. If you want to stay at the cutting edge and move things forward really quickly, you don't want to have to fine-tune these models.
We do have some fine-tuned models, but we don't do the fine-tuning for the purpose of fact completion. We try to avoid fine-tuning if we don't have to. I'm actually fairly excited about reinforcement learning here, because I think the big models do a pretty good job of extracting facts, style, and tone from examples, but they could do a better job.
I think reinforcement fine-tuning will allow us to teach the model how to do style and fact matching better in a way that's generic across everyone. We don't need to fine-tune per user; we can take the new GPT-4o, use reinforcement learning, and teach it how to use examples for style matching, get the facts right, and write emails.
I do think that using 20 examples of things that you've said to match style and tone, with the right prompt and the right reinforcement learning, is the right solution.
Nathan Labenz
How about memory? I hear you on the operational challenges of fine-tuning, and every time I raise this notion of per-user fine-tuning or even per-company fine-tuning, the most compelling counterpoint is, “What are you going to do every time you want to change models?” Then there's new knowledge that comes in and how often you have to run it. It sounds like a real bear.
It still feels to me like there's some sort of missing middle in terms of memory. We have the context window, then we have what's known in the weights, and then we have the database call, but those all feel like they're still missing a little bit of something. I come back to this intuition of, “I know when I've found it,” and there does seem to be something qualitatively different about that. I also kind of know what I've tried.
I've been really interested in state-space models and in the Titans paper that came out from Google not long ago, where they actually use a neural network—a memory module—that gets updated at runtime. I've also been interested in HippoRAG, which was inspired by the hippocampus and how it's understood to connect concepts together. They have an extensible graph network that they can query against.
Maybe you think we don't need any of that and we just need better models and better search. Maybe this whole thing is a confusion on my part, where we keep pushing the current frontiers and eventually it will all work. One day I'll realize that we never really needed another middle piece. What's your expectation for that?
Andrew Lee
I don't know what's going to happen here. My Spidey sense is that there is going to be some big breakthrough, and there is going to be some concept of memory that's baked in that allows the models to be customized in a way that isn't quite so heavyweight as fine-tuning. But I have no idea what that's going to look like, and I don't even have an idea of what it could be.
Our approach right now is that we have this thing called Memories as a feature in our product. It basically allows the LLM, as a tool, to manage a list of facts, and then we insert those facts into the prompt. This is really useful for behavior customization. If you want to say, “Every time I write an email, always CC my EA,” or, “Any time you schedule a meeting, default to 45 minutes,” it's really useful for that.
We have that, but you have to be explicit. You have to say, “Remember this fact,” and then we'll do it. Or, “Remember to do this.” We also obviously use search. I still think there's a lot of opportunity with an agentic model to use search. You could have a database of interesting facts about you and a tool that looks at those facts and tries to apply them.
I don't think it would be quite the same as built-in memory, but you could probably do a lot of the same things. I haven't had time yet to read the Titans paper. I think you mentioned that the other day, and I'm very interested. I do think there will be some sort of breakthrough here, and I can't wait to see what it is.
Nathan Labenz
It feels like one of the—it's too early to call it a final frontier, but in terms of imagining something that could really work alongside me on an extended basis, a more integrated, dynamic, active-learning system—constantly updated memory—does feel like something that would move the needle tremendously on that particular dimension.
One last question before we switch gears into the meta of how you're doing this. One thing we talked about last year was the rise in AI-generated spam and what happens as more people adopt various tools and you have AI potentially talking to AI. Has that happened? Another thing we expected was a lot of deepfakes during the election, and that didn't really happen all that much. Is this another thing we're all too worried about that isn't really a thing, or are you starting to see any interesting AI dynamics?
Andrew Lee
Not a ton, honestly. I do get a lot of AI-generated content in support, especially just people being funny. We'll send out a newsletter, and they'll have it write a poem or a joke in response, so people are doing that kind of thing.
I haven't seen a huge uptick in AI spam. Our users haven't reported a huge uptick in AI spam. It's possible that Gmail's normal spam filters are doing a good job, or maybe it's just not a big problem. Thankfully, it doesn't seem to be as bad as I feared.
Nathan Labenz
I would say I still notice, honestly, more often, I marvel at how I'm getting cold emails that were obviously written by a human and would have been better had they been written by an AI. Maybe lagging adoption is still the broad explanation there.
Andrew Lee
A significant fraction of the emails that I send are at least AI-enhanced. Some of them are fully AI-written, and as far as I can tell, no one notices. I'll talk to people in person, and they'll ask me this sort of question. I'll say, “Do you realize that my emails are AI-written?” They usually say no.
Maybe I just don't notice, which is fine if I can't tell.
Nathan Labenz
There are a couple of really interesting things there. I personally have not shaken this yet, but I've increasingly had to confront the fact that I'm probably way too precious about the little elements of my style that feel to me like they make me unique. I honestly kind of doubt that anyone else notices or cares. I've got a friend who hammers on that all the time, and he's just like, “Nobody cares about that, dude. Your little flourishes are totally lost on other people anyway.”
I also differentiate between routine tasks and non-routine tasks. As somebody who doesn't really have a job and is just scouting the AI space all the time, I do very few routine things relative to what a typical sales user with a CRM integration would need to do.
Andrew Lee
But I really do think routine is the place to look for AI. I do a ton of routine emails because, for example, we made some changes to the UI yesterday. I'm sending emails that kind of walk through why we did things and how to adapt your workstyle that are very similar between all the different people I'm talking to. The AI is really useful in looking, “Hey, you just sent a bunch of emails kind of like this. Let's pull out some of the ideas here and reconstruct them to answer the question.” I'll send 20 emails of a very similar type.
Nathan Labenz
Let's shift gears a little bit, if you're ready, into what it's like to work at and build Shortwave. I understand there's some news around a fundraise and an expanded vision—not just email, but omnichannel communication. The most interesting thing from your recent Twitter thread was the line, “We're also building an incredibly AI-forward culture where the focus of our work is managing AI agents rather than making changes directly. We write a ton of code with AI, use AI for research and design, and even wrote this posting with AI.”
I don't know what you want to share about the fundraise, but you can expand on the vision. I really want to get into what it's like to be building this AI-forward culture and managing agents all day at Shortwave.
Andrew Lee
The company has changed tremendously in the last few months. We reached a point where we figured out that the future is not making a better email client. The future is an AI that has your communication apps integrated into it.
The main thing people are doing in this product is talking to the agent. In that world, we're building a very different thing. It's not about how you streamline every interaction in the inbox or make the UX of email threads amazing. We care about that stuff to some extent, but it's mostly about how we help people get things done at a higher level.
We're already seeing a lot of traction. This is the reason we're doing this: it's already where people are being successful. Our biggest and fastest-growing plan is our most expensive plan, which has substantially more money going to Anthropic for AI. The people using us now care about the AI, so we're doubling down on that.
That vision has gotten our investors very excited. They see an opportunity for an email client, but they also see an opportunity for this sort of thing. There does seem to be an opening in the market for an AI that can actually do things.
If you use ChatGPT, Claude, or Perplexity, they're great at answering questions and doing research, but when it comes to actually doing work, they don't really have the capabilities to do that. Someone is going to build a thing that can actually get stuff done, and we see that opportunity. We're starting to see our competition as the next version of ChatGPT rather than the next version of Gmail.
It's been a reframing. We have raised some money, but I'm not ready to share the number or who was involved—just enough that we can significantly expand our team. We're doing that right now. As part of this, we revisited everything about the company: the way it's financed, how we operate every day, and who's on the team.
We think this tidal wave is coming. I think the tidal wave is going to hit faster and harder than basically everyone realizes, and we want to make sure we're on the other side of it. We think it's going to require some different people with a different mindset to do what we want to do.
You mentioned the sound bite that your job is going to be primarily managing a bunch of AI interns to get work done, rather than necessarily doing it yourself. We're seeing this especially right now with code. If you're using Cursor's agent mode, a good software engineer right now is going to have at least 1 agent running all the time doing something.
They start the agent off, then go work on the next problem, and the agent goes off and solves it. We've found that it's possible right now to take a bug report, copy it verbatim into agent mode, and ask it to solve it. In many cases, it will solve it in 1 shot, which is pretty amazing.
On the software-engineering side, we're looking for people whose skill set is primarily not execution. It's not writing code; it's understanding the problems, understanding the components involved and how they should interact, and being able to frame those prompts. The value has shifted from people who can get things done to people who can understand problems and structure things.
On the design side, I think it's changed too. It used to be that the way you designed a product was to get together in a room, draw it out on a whiteboard, make some mocks, make a prototype, and then build the thing. Now the first thing you do is make a prototype. Before I even talk to anyone else, I just build it.
Now, and it's only after I'm like, “Yeah, I think this will be pretty good,” that I go and talk to the right people. Then we make a mock and try to figure out the details, but it lets us cycle through bad ideas much, much faster. On the design side, we're looking for designers who have already made the shift: “I don't start with a mock. I don't even start with a wireframe. I start with a working version that I built with Bolt.new or a tool like that.” That's the beginning, and so our design is totally different.
The same thing is true with content. I mentioned the job postings that we wrote—that's ChatGPT-4.5. It's really good at writing. It's not quite at the level of what I could produce if I put a lot of effort into it myself, but it's not far off. I can produce content at a pretty high bar in a fraction of the time.
We need to do a lot of comms here. We want people who are good at figuring out how one person can manage our blog, our social media, our changelog, all the docs on our website—everything. How could just one person do this if they're thinking hard about the right prompts, the right way to use these tools, and how to generate this stuff?
It's gone from, “Hey, maybe we need a team of 50 to do this,” to, “Maybe we just need a team of 15 or 20 people with the right skill set and the right usage of these tools.” The 30 people who aren't there anymore are really more on the execution side, and the 15 or 20 who are there are more on the problem-understanding and problem-solving side. I think it's going to be a big shift.
We actually let a couple of people go who were very talented, very good at what they did, great employees, and had a great attitude, but weren't quite right for this new world that we felt we needed to get into. They weren't super passionate about it either, so we thought a change needed to be made.
Nathan Labenz
Passion is super important. That's definitely something I observe all the time. It does take some curiosity and some innate interest. To the degree that I've demonstrated some aptitude here, a lot of it really just stems from that.
At this point, is this something that you think is an investment in the future? Obviously, it's that, but is it also something that allows you to be more productive or execute at a higher level today? Or do you feel like, as you mentioned with the blog post, maybe it's not quite as good, but it's fine because it's pretty good, and we're also investing in being the type of organization we think we have to be in the future? Does that make sense? What is the timing for ROI on this?
Andrew Lee
I think we're already seeing it in a big way. I think companies need to be making changes now to the way they work, the way their org charts work, and the way they spend their money. I don't know if I can put a number on it, but our engineers today are significantly more productive than they were 6 months ago. It's really changed in the last few months. If you last tried this stuff last summer, the world is different. Give it another try.
There are significant improvements on the engineering side. On the design side, we can skip over half the process and get right to the prototype. I wouldn't say it changes the work needed to build the final design dramatically, but it does allow us to cut out a lot of bad iterations, which is where most of your time is spent anyway.
We're also a lot faster on the content side. Good writing is hard, and most of the time you don't need to write amazing prose. You just need to write pretty good prose. If you can do that much faster, you can keep your docs up to date faster and do things like that. I think we're seeing those gains today.
There's a related thought I wanted to add here, which is a big change in the way we're thinking about our business. A year ago, when I was talking to investors, their question was basically, “What is your moat?” They'd say, “You're building this cool AI stuff, but why can't someone else come and do this?” I was like, “Our moat is that we have an email client, and it took us years to build this thing. If anyone else wanted to build an AI-enhanced email experience, they'd have to go build an email client. That's really hard.”
Now I think the thing that took us 4 years to build might not take 4 years for the next person. It will still take them a while, but they're going to be able to do it a lot faster. I think there's this big change happening in the industry where all of these companies that have moats based on the software they've built are seeing those moats significantly eroded. The value of the code that's already out there is going down very quickly, and they're going to have to come up with some other moats.
We're looking at this and saying, “The moat that matters—the only moat that probably matters—is speed.” How do we optimize our team for speed? I think the best way to optimize your team for speed is to keep the team small. It's just harder to get consensus on big teams.
We're saying, “Okay, how do we have 15 people who are the right people, with the right tools, working together really well, so we can just move, move, move?” Our moat is that we're always ahead of everybody by 2 months because of the way we're structured. That's been a big change in our thinking.
Nathan Labenz
So 50—is that a long-term number?
Andrew Lee
Not super long-term, but I think it's for the next year. I don't think we're going to go beyond 15, and then we'll see beyond that. I don't think most people could historically build a product at the scale and scope of the product we're building with 15 people, but I think we can.
Nathan Labenz
There's a lot of surface area already, and it sounds like you're planning to add a lot more. On top of that, you've got a lot to keep up with.
Andrew Lee
We're not the first ones to figure this out. Cursor is around 20 people. Midjourney is around 20 people, and they're obviously wildly successful. I think there's a new era.
Nathan Labenz
Are you still subsidizing users on the margin? That was another tidbit from last time. You said you were literally losing money on every user. Is that still true?
Andrew Lee
Not anymore. We've done a lot of work to increase efficiencies. We've also added a higher-end plan and gotten people to opt into those higher-end plans, so we are margin-positive.
I wouldn't say we're hugely margin-positive. We still spend a significant fraction of every dollar you give to us directly on LLMs, and a lot of it also goes to traditional email infrastructure. But on the margin, we do make money now.
Nathan Labenz
Is there a 2–10x more expensive version of the product that you could imagine rolling out? Or are you already basically—not quite maxing out, because you talked about how you could spin up the agent every time—could you do that today if you just charged 5 times more?
Andrew Lee
We could, and I think we probably will. If you look at all these different plans, I think I mentioned this earlier, but all the growth is basically happening in our most expensive plan. Everyone is coming in and wanting the most expensive plan because it lets you have the biggest context window and lets you index all of your history for search. People care a lot about that.
Even though our business plan can give you almost as good answers as the Premier plan, people want the best answers. The difference between good and best is worth a lot of money to them.
ChatGPT Pro came out at $200 a month. I pay for it. I think it's too low.
Nathan Labenz
No hesitation on my end either.
Andrew Lee
I would pay more than that with no hesitation. I think we should probably do the same. There should probably be a $2-a-month plan, and I don't know how far this goes. Sam was talking about maybe having a $2,000-a-month plan or a $20,000-a-month plan. I have no immediate plans to do that, but to the extent that people want us to spend money on GPUs on their behalf, there's sort of no limit here.
For example, right now we don't automatically run the full agent on every email that comes in. That's a great place where we could spend a lot of money on your behalf if you want us to, and it could be a 100x increase in the amount of compute required.
Another example is reasoning models. Right now we don't use reasoning models. They're much more expensive and slower, but if you're trying to do something complicated—“I want you to give me a detailed analysis of every customer report from the last year”—this might otherwise be something you give to an employee who spends a month on it. You might be happy to spend hundreds of dollars of compute just to answer that one question.
There may be something we can do there: reasoning models, automatic execution of things, bigger models in general. Maybe there's a $200-a-month plan, or maybe more someday.
Nathan Labenz
Going back to the building side, you mentioned Cursor, of course, and I'm a Cursor user. I've also gotten into Replit and Lovable, and I basically try everything I can. What I haven't quite figured out yet is how to make things work together—how to effectively architect systems of agents.
I've had great experiences with all these different things in their moments, but have you brought together a mix of things? Are there things that complement the coding agent? I'm thinking, for example, of a company called Qodo that specifically emphasizes testing. There could be observability or monitoring tools. Monitoring is another thing that you could really imagine layering on.
Ultimately, it seems like we're going to need it. If you're going to have 15 people who continue to run at roughly human speed with a lot of AI assistance, they're also going to need AIs to help them supervise the AIs. The pyramid built under each person and under the company collectively gets bigger and bigger. There's got to be a whole architecture and some specialization within these agent ecosystems. Is that roughly how you see things shaping up, and have you started to tie these things together in a useful way so far?
Andrew Lee
I have a couple of thoughts here. The first is that one of the potential futures for Shortwave is really a routing layer for the messaging going on in your business life. You might be using a bunch of these different AI tools, but something needs to take the incoming events, decide which agent handles them, get them to that agent, take the output from that agent, and do something with it.
Inboxes are great tools for managing that flow in a way that lets both the human and the AI collaborate. If we're going beyond email and including other types of human communication, maybe we're going beyond traditional human communication entirely. Maybe it's more like a Zapier-type of thing with a human UI that you can use as well.
We might be a routing layer that helps you route things into these other tools. We've had requests from customers like this, and this was a discussion I actually just had with one of our investors: maybe we need to start thinking of ourselves more as an agent-routing layer.
My second thought is that, in our approach to agents, we've tried some different ways of having a multi-agent approach to solving problems. Basically, you have one system prompt that you iterate on over here, and then, once you decide it's a certain type of problem, you change the system prompt, add in a different prompt, or hand it off to a different model.
Our experience so far—and this is just one experience—has been that those systems end up being kind of brittle. They tend to struggle with reasoning across different types of tasks. Cost aside, the better approach tends to be to take the biggest, most expensive model you possibly can, stick all the instructions into the context, and let it reason across the different things.
I'll give you a specific example of a win we had recently. We used to have the ability for you to add custom instructions for certain types of operations. For example, for writing, we could add custom writing instructions. When we called the tool used to grab the writing instructions, we would insert those custom instructions into that tool.
That worked well in cases where the writing tool was being called appropriately. But if you had custom instructions that were trying to give it hints about when to write emails, what types of examples to look up, or how to handle things outside of writing, it didn't work as well because those instructions only got plugged in at certain times.
What worked much better was this memories approach, where we take some customization instructions, add them into the master prompt, and include that in every call. Then you can have a customization that can be considered at any time. You could have a customization like, “Always address Nathan Labenz as ‘Sir,’” and it could do that when writing emails, scheduling calendar events, or in any other situation.
That has proven to be a much better user experience. It produces much better results from the AI. We used to have a model that was routing things to different AIs, and just saying, “Screw it, we're going to put it in the main prompt,” has worked much better.
Nathan Labenz
More cash, too, I guess, to have that consistency?
Andrew Lee
I see the opportunity of the multi-agent world really being more about working across organizations. We need some interface between our team and the Cursor team, for example, but that's probably not something that's happening within our app.
Within our app, I think the approach of just using the biggest LLM we can get, the biggest context, and caching really well is probably more likely to be the future for us.
Nathan Labenz
I share the same intuition, based on everything that I've tinkered with over the last couple of years. I just told a company that I've been doing some agent advisory with that, and then 2 days after I said that, OpenAI came out with its thing where handoffs between agents are a notable new feature. What do you think is driving that?
Andrew Lee
I literally talked with that team yesterday about that exact thing, and I had the exact same questions for them. I shared my perspective, and I don't think handoffs are going to take off. That's my hot take.
Nathan Labenz
I can see some upsides to it. The things they say are that it's easier to test agents in isolation and things like that. But it's tough when you then have classifications, and there are just more things that can go wrong.
I'm glad that somebody else shares my intuition on this, because I was wondering, “Am I out of touch?” Certainly, OpenAI has real insight into what's going on in the broader world. They shared the same perspective with me around developing things in isolation, but I think the reality is that with an agent, you don't really want the tasks to be isolated. You want the agent to reason across all of its capabilities and make a smart decision.
Maybe there are some places where you really want to sandbox it, but my experience is that isn't usually what you want. You usually want the agent to think about everything it's capable of and make the best choice.
I wonder if this is to some degree driven by larger companies that have existing organizational structures, with somebody responsible for each thing, and those lines are really hard to cross. One of the more interesting tidbits I've heard in AI discourse over the last year was from Yi Tay, I think on the Latent Space podcast. They were talking about multimodal models—why they developed as they have, and why vision models and language models developed separately and then were merged with late-fusion approaches.
The take was simply that this reflects legacy team structures. It used to be that you'd have a vision team and a language team, so they would do their own things and then maybe try to fuse them. But that won't last into the future. You're going to have unified teams and unified architectures from the beginning.
This feels like the seed of a similar insight. I've often wondered, and we talked about this a little last time with respect to Gmail specifically, why the big incumbents can't do a good enough job with this. The technology isn't that hard to use relative to other things. That's one of the things that's so great about it: it's flexible and forgiving, and it totally understands what you meant even if you're riddled with typos.
I've even seen people speed-typing, putting tons of typos in and having the AI correct them. For whatever reason, I haven't gotten over the hump on that myself. But maybe there's something here in terms of who will win in different verticals. Big companies have these org charts, divisions of responsibility, and questions about who's going to sign off on what. Maybe these things are being built for them because that's what they need in order to make sense of it, and maybe it's not really about what will make the thing work best.
Your approach—small team, totally unified, a single agent that knows your whole history and can deal with you on that basis—seems more like what I want. I don't want to be handed off from one AI to another. I don't want to be handed off from one human to another, either. This is interesting. There may be some insight here that proves predictive.
Andrew Lee
I think you're probably right. I'm sure they did this in response to real needs they're seeing, but one of the reasons we started Shortwave back in the day is that I worked at Google and had some insight into the sorts of organizational struggles they were going to face. I believed the Gmail team wasn't going to be able to innovate quickly for organizational reasons that were very hard to change, and I think that has borne out.
The models from Google are awesome. The new Gemini 2.5 is literally mind-blowing. But the product progress in using those models in interesting ways in something like Gmail is way behind. I think that's driven by the way the organizations work. It's not that I don't think the people are smart. It's just the way the organizations work.
For example, we decided, “Hey, we need the LLM on the left, right? The agent's got to be over here.” If I were at Google and said we needed to take the right sidebar and put it on the left, that would take me 2 years. Think of all the sign-offs I would need, all the people I would need to convince, all the buy-in I would need, and all the meetings. You'd literally have 1,000 meetings.
For us, it took a few weeks. There was a lot of discussion, but it was all about whether this was the right thing for the product. We could focus entirely on that question, and we could move a lot faster.
Nathan Labenz
Going back a little bit to who you're looking for and what the hiring process is like, I noticed that, as far as I can tell, all the roles are in person in San Francisco. I have a question about that, especially as you think about a future in which a lot of the work is being done by AIs.
It feels like the AIs are always remote. Personally, being in Detroit, I've had a great experience hiring remotely. It's a quite different local talent pool, and the pros and cons are quite different. Hiring remotely has opened up the aperture of who we could go for.
I'm thinking, if all the AIs are only going to interact with me through a screen anyway—at least until I have humanoid robots sitting at the desk next to me—how are you thinking about holding the line on in-person versus potentially liberalizing to remote?
Andrew Lee
There are a couple of pieces here. The first is what I talked about with speed. We have historically been a remote team. In fact, at one point we were fully remote. We weren't even hybrid.
I think it's become increasingly important that speed is critical. It's a lot easier to move quickly in person, especially when what you're doing is making rapid product changes. You can move in a straight line very, very quickly remotely, but if you need to quickly change course, have difficult meetings and conversations, deliver tough feedback, and say, “Hey, this isn't working,” it's just a lot more emotionally challenging to do that in a remote environment.
It's possible, and teams do it, but I think it's harder, and it's something we really struggled with. If speed is paramount, we think having a core team of people who are driving most of the roadmap and product decisions, and who can meet in person and hash things out on a whiteboard at any time, is really key.
We want to have that happen. There's a second key piece, which is just a reflection of the founders. Jonny and I work better in person. I'm a better leader and manager in person. People like me better if they see me in person than if they're talking to me over a camera.
That's probably a weakness. I could learn to be more effective remotely, but I also need to understand my capabilities and limits. I look back at Firebase, which was a fully in-person company, and there were a lot of struggles we've had with Shortwave that we didn't have there, caused by some of those differences.
Jonny and I looked at this and said, “We need this to be a place that optimizes for our strengths. We work better in person, so we need to have a core team, and we need to move really, really fast.” We need to have the core team in person.
This makes recruiting harder. There is a lot of amazing talent that isn't in San Francisco. I'm very keenly aware of that, but we think it's worth it.
Nathan Labenz
How about an AI scout role? This is a hobby horse of mine. Have you thought about a dedicated position for somebody to try every new thing, whether it's a new model, a new framework, a new agent experience, something competitive, or something totally far afield?
This feels like something more and more companies are going to need, but I don't know how many are feeling it quite yet.
Andrew Lee
Are you asking for a friend, Nathan?
Nathan Labenz
Yes, well, many friends, actually. When you would have had 50 people and you're now 15, and that starts to get generalized, we do need some of these new AI jobs.
I've managed to stumble my way into this. I sometimes call myself the Forrest Gump of AI, where I just unintentionally stumble through these important scenes. I've ended up in a place that I honestly quite love and have nothing but appreciation for, but I do think we need a lot of new AI jobs. This feels like one that might become common.
I'm asking on behalf of maybe a lot of people.
Andrew Lee
100%. It's not a role that I have listed right now because we're a very small team and we only have a handful of roles. But I do think it's sort of impossible to keep up right now.
There have been a number of moments where our discovery of a new technology made a big difference. For example, if I hadn't listened to that particular podcast that talked about the new behavior of Claude, it might have been months before we figured out the same thing. There have been a number of inflection points driven by the discovery of new technology that I think we just got lucky on.
It would be great to have someone systematically doing this sort of thing. Obviously, I try to listen to great podcasts like this one and keep up that way when I can, and read whatever I can. But in the future, we may very well have a role like that. I think a lot of other companies, especially bigger companies where not everyone is constantly thinking about AI, could benefit tremendously from something like that.
Nathan Labenz
Situational awareness is harder and harder to maintain all the time. Maybe one last thing on the hiring side: to some degree, these are just labels, but I noticed that the roles you do have posted are all designated staff or senior. Is there a place for somebody who doesn't have a lot of experience at Shortwave?
To take a somewhat extreme case, let's say they're freshly graduated with no work experience. Is there any way for somebody like that to demonstrate that they have the skills that would make you open to hiring them?
Andrew Lee
Yes, absolutely. For what it's worth, the job postings have been updated significantly since you saw them. I posted the senior roles first, but we actually have some entry-level roles up there now.
I just posted yesterday that we're looking for someone for customer success. We're looking for a content creator because we want to make a lot of videos. I think there's a ton of discovery and education in AI, so we want to do some content creation. There's also a junior product-engineering role up there. We're definitely looking for some junior people as well.
If you want to impress me, there's a very simple application process: send me a video. We've gotten rid of some of the take-home tests and things like that because they're too easy to game with AI. We're just going to say, “Send me a 5-minute video. Show me something cool that you did.”
The thing we're most interested in is whether you're really forward-thinking with AI. If you could show me a way that you leveraged AI to do something useful in a way I hadn't thought of, and that impressed me, I think that would cause me to take notice.
Are you really forward-thinking with AI? Are you being creative about how this stuff is used? Are you staying on top of it? If you can stay on top of it, you're already pretty impressive.
Nathan Labenz
No doubt. More broadly, what do you think is going to happen to the future of the software industry? This is obviously a super-hot topic right now. Are there going to be more developers because they're more valuable, or is there only so much software that needs to get written?
This could break down in phases. We might be on the up ramp of that curve, but the more I hear from people like you that you only need 15 people for the foreseeable future, the more I'm thinking, “I don't know how much we can bet on this developer market to grow, or even sustain itself as it has for the last few years.” What's your expectation?
Andrew Lee
Good question. My co-founder oscillates, depending on the day, between an existential crisis of “I am obsolete”—he's the CTO—and the next day saying, “I am a god, and I can do so much.” He doesn't know which one is right.
I think every software engineer is going through that right now. Their productivity is going through the roof, but they may also be obsolete at the same time. I think the nature of what you do is going to change a lot.
Actually writing localized code is starting to become something LLMs can do. For example, any sort of front-end development where you say, “Hey, I need a button that looks like this,” is something LLMs are great at. Increasingly, code is going to get written by AI.
There was an interview with Dario from Anthropic where I think he said that within 2 years, 90% of all code—or maybe 100% of all code—is going to be written by AI. I'm not sure I quite believe that, but I think it's going to be a very significant amount.
The places where I think AI is going to take longer, if ever, to be able to do things are understanding user problems and understanding the components and how they interact. That's super important, super complicated, and actually a huge part of the job today.
If you take a senior engineer today, that person isn't senior because the UI code they wrote is dramatically better than the junior engineer's. The senior engineer is able to think about the whole business problem and solve it.
The people who are good at thinking about business problems, how to solve them, and how the components interact are going to do great. The people whose strengths are on the execution side are going to have to learn some of those other skills and get good at them.
In terms of the number of people, I don't know. On the one hand, you can do more with less. On the other hand, there are suddenly a million little companies that should have existed but were just too expensive. Software engineers are crazy expensive right now, and if you need 10 of them to do anything, that's a huge barrier.
Suddenly there might be a startup where you would have needed 20 people, but now you can do it with 3. You can build a niche product for something. Maybe we'll just see another 10x explosion in software, and the jobs will stay the same.
I don't know. I'm generally an optimist. I generally think this will make life better for everyone, including employees. Everyone's going to have to adapt, but I think it's going to be a good thing. I think being a software engineer is going to be an awesome place to be over the next few years.
Nathan Labenz
I can't imagine coding without AI assistants at this point, so on that level, there's no doubt. I definitely share your expectation of 10 times more software getting created, but I'm still thinking that maybe even 10 times more software isn't enough to sustain the number of pure headcount jobs we currently have. Time will tell.
I do advocate that people start thinking about universal basic income if they haven't already started to contemplate it. We're not quite there yet. That's a long-term problem—maybe 2 or 3 years out.
As you look a little into your crystal ball, what do you think are going to be the big developments for the rest of this year? What are the things you're thinking, “If only they could get this to work or fix this, the world would be much better or much different for us”? What are you watching most closely?
Andrew Lee
I think the number one thing is post-training on agentic behavior and tool-calling-type stuff. The step change from every other model before Claude Sonnet 3.5 to Claude Sonnet 3.5 just enabled a whole bunch of new things. Claude Sonnet 3.7 was better, but there are still gaps.
It would be great to have more competition there, and to have OpenAI and other companies offer more options. I think this is all coming down to agentic-specific post-training, and I'm excited to see where that leads and how good it gets.
Iteration and tool calling can work around so many limitations in your system. The AI can find creative solutions to things, and I think the sky's the limit there.
I'm also generally excited about improvements to productionized tools: cost, latency, and reliability. There are a lot of things we can't do because they're impractically expensive, like running the full agent on every email that comes in or adding more search results to every query. The cheaper, faster, and more reliable things get, the better for us. Cost is still a huge factor, so that's another area I'm watching.
Another area is more native multimodal voice. Today you can do voice things—we actually have voice input in our app—but it's fine. It's not great, and it's not like really talking to a human. Models that support voice natively and multimodally, and that we can use for the full agentic flow, would be totally killer.
Nathan Labenz
I love vision, too. Anything that will get me outside more and less tethered to my desk is hotly anticipated, certainly by me.
Any other thoughts or closing wisdom you want to leave people with?
Andrew Lee
You should come check out Shortwave, and you should totally either apply yourself or refer your friends. We have a $1,000 referral bonus for anyone. If you get them to apply and tell us that you referred them, we will give you $10,000, no joke. So yeah, help us find great people.
Nathan Labenz
Love it. Andrew Lee from Shortwave, thank you again for being part of The Cognitive Revolution.
Andrew Lee
Thank you, Nathan.
Nathan Labenz
It is both energizing and enlightening to hear why people listen and learn what they value about the show, so please don't hesitate to reach out via email at TCR at turpentine dot co, or you can DM me on the social media platform of your choice. [Music]