Hello and welcome back to the Cognitive Revolution. Today's guest, Logan Kilpatrick, needs no introduction. This is his fifth appearance on the show and his tireless work in support of AI application developers previously at OpenAI and now for the last year and change at Google is legendary. In this conversation, with the benefit of at least a little time to process, we're looking back and digesting the overwhelming volume of major new AI models and products that Google and others have recently released. Logan describes his personal experience at Google as their AI usage has grown some 50x from 10 trillion tokens per month just a year ago after he started to 500 trillion tokens per month today, which is notably more than 50,000 tokens per month for every person living on planet Earth. Logan also shares his perspective on Google's incredible organizational transformation from what once was described as a sleeping giant to now an indisputable top tier AI powerhouse with assets headlined by the strongest overall compute infrastructure of any company. top tier and paro frontier models including Gemini 2.5 Pro, highly original and viral products like Notebook LM, gamechanging applications in medicine and science that are starting to ship to trusted users and what I have always and still consider to be the deepest bench of AI research talent and the most diversified well-rounded research agenda to be found anywhere in the world. He also offers thoughtful analysis on whether we'll continue to see convergence among leading AI companies or more divergence as the lowhanging fruit gets picked. Why he believes that startups still have unprecedented opportunities despite big tech's advantages. The implications of anthropic cutting off windsurf after they partnered with OpenAI. How the blinding speed of Google's latest diffusion language model could bring about yet another revolution in software. and why despite all the AI capabilities advances he's seen and helped to popularize, he's still betting that humans will continue to matter and taking a relationshipcentric approach to his work. Speaking of relationships, perhaps the highest alpha part of this episode was when I asked Logan for advice for those who want to break into the early access programs and other support structures that he and people in similar positions can provide. I won't spoil his response here, but it did include his personal email and an invitation to reach out. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends or leave us a review. We always welcome your feedback, too, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. With that, I hope you once again enjoy hearing from Logan Kilpatrick of Google Deep Mind.
Nathan Labenz
Logan Kilpatrick, everybody knows who you are. Welcome back to The Cognitive Revolution.
Logan Kilpatrick
Thank you, Nathan. This is the world record for the most number of times. I feel like you should just make me an independent recurring segment on some regular cadence, because we get to do this a lot, and it’s wonderful to be back.
Nathan Labenz
It’s only your calendar that would prevent that from happening, so be careful what you wish for.
We were joking about titling this podcast “The Decade of the Week of May 15 to May 22, 2025.” Holy moly, we’ve got just an absolute avalanche. I’ve been saying for a long time that my grip on all AI news is slipping, and with this moment, I think it’s officially slipped for everybody. We’ve come a real long way since GPT-4, and a ton of stuff is happening.
I want to run it down, but I also realize at this point we can’t even be comprehensive. I also want to take some strategic opportunities to zoom out a little bit and get your bigger-picture perspective on some things.
The first one in that vein is that over the last however many months, we’ve seen several waves of leading AI companies launching very similar things in pretty short periods of time. This has happened with reasoning models, and most recently it has happened again with coding agents. On a feature level, we’re also seeing quite a bit of connecting to your Gmail, connecting to your Google Docs, and different kinds of context retrieval.
Some of that’s obvious, but some of it is pretty core research-driven, right? Getting the models to reason. How do you understand why the different leading companies seem to have such a similar development trajectory and also launch timelines?
Logan Kilpatrick
Yeah, that’s a great question. I think there are a couple of dimensions to this. One, I think on the research side, there are true innovations that, when people go out and talk about them, become clear in hindsight why something should be done. I think reasoning is maybe that story. Obviously, DeepMind has been working on that reasoning stuff for a long time.
I think some of the particular techniques just became clear: let’s make this level of order-of-magnitude investment, and things ended up working out pretty well. Then you bake in all the other stuff that we had been trying that was independent and different, and you start to see some really interesting things.
I think some of this is people lighting the path, and this is what’s awesome about the ecosystem: other people light a path, you get to benefit from the path they’ve lit, and then you go and bake those innovations into what you’re doing, plus benefit from the bets you were making that were independent of that. I think that’s exciting on the research side.
I think on the product side, AI, both on the model side and the product side, is the most competitive ecosystem in the entire world right now. There is not a more competitive ecosystem, with the amount of money, talent, intellectual capital, speed of execution, and so on, than there is in this AI ecosystem right now.
I think there are a lot of competitive people who are really good at what they do, and they don’t want to be pushed behind by their competitors. There’s this feeling that you have to stay on par with what everyone else is doing.
That actually ends up being a tension I feel as somebody who builds products in the AI ecosystem: the tension between doing what you think the long-term future is versus not trying to look like you’re behind in the present moment. Finding that balance point between the long-term bets that are distinct and actually going to give you a differentiated perspective over time, and the short-term requirement that we just need to have parity, is a tough challenge.
I feel this on the developer side as well, because we obviously provide compatibility layers for people who use other model providers to come and use Gemini. There are always thoughts about how much we invest in that to get feature parity while also developing next-generation API capabilities and model capabilities. It’s a tough trade-off.
As for the dates and timing, I think some of it is intentional. I think a lot of it actually just ends up being relatively happenstance, that things launch around the same times.
It is always fun to see the conspiracies of X, Y, and Z companies just sitting on something and then, with 1 hour’s notice, all deciding to put something out. Anyone who’s worked in any company larger than 20 people knows that’s not actually possible. The amount of operational overhead and burden to do something like that is enormous. Companies are not that nimble. Even your favorite AI lab is not nimble enough to have that level of reaction.
Nathan Labenz
Well, I do have to say Google and DeepMind, and your team specifically, have been pretty nimble, right?
A year ago, and definitely 2 years ago, the outside view of Google was “sleeping giant” turned sclerotic—everybody managing their own little fiefdoms. I don’t know to what degree that was really true. Now the narrative is totally flipped: the giant is wide awake and really remarkably keeping pace with even much younger and smaller companies.
One of the most interesting data points shared at the Google I/O event was 500 trillion tokens per month now being processed across Google’s services. We’re now into the next month, so at the rate of that curve, it might be literally 2 times more already. I don’t know if you’re watching the dials that closely to know, but how has that happened at Google culturally?
What has shifted, or what has your experience been internally—not just going through that curve, but rallying everybody to actually support all the different work that has gone into supporting that curve?
Logan Kilpatrick
Yeah. One of the interesting parts about this story is that, at the core, it’s a people and organizational story. I think the challenge is that it’s just not—that’s not selling front-page New York Times articles, or whatever your favorite analogy is.
Historically, if you look at how Google was set up to do a bunch of this work, I think the reality was that it wasn’t set up for this moment. Google was structured with many different teams doing AI work, and for a lot of the right reasons, actually, because they were pursuing fundamentally different goals in some sense.
Google Brain, as an example, historically had a very wide breadth of truly different research. That’s where a lot of the Transformer and other things came out of—that very varied breadth of research.
At the same time, you had Google Research, which was doing more applied things in some cases and trying to upstream a lot of that into other parts of Google. Of course, Brain also did that sometimes.
Then you had DeepMind, where Demis and the team had a strong opinion about how they thought—and how they think—we’ll get to AGI.
So they were pursuing a very specific research direction, and you've seen a lot of the things that have come out of DeepMind over the last 6 or 7 years around that. I think in a world where there wasn't a clear winner as far as what the technology, at least in the short term, could be to help scale us closer to systems that get closer to AGI, it made sense for Google to have those bets across the board, with very different structural organizations and approaches to doing this. But I think it became clear at a certain point that one of those things was working in the short term, and we should rally resources and get everyone on the same page. I think that happened in the middle of 2023, when Google Brain and part of Google Research actually merged with DeepMind.
I think that was the start of the story of Google putting itself in a position to be successful for the next 10 years with this technology. The challenging part is that large human systems are extremely complex. It's not just that there are a lot of humans involved; I think it's very easy to forget the level of complexity and chaos in human systems, regardless of what you want the outcome to be. I think the DeepMind team, Demis, and the folks on the leadership team there have done a great job of reinventing the culture and all that stuff in order to make an organization out of two very different organizations, bring them together under a single roof, and actually set up the team structure so that Google could be successful building these models and going through the process.
The large-scale training runs don't take a minute, so your iteration cycle isn't super quick. Your iteration cycle for releasing models and doing all the end-to-end work isn't super quick. It took time to actually get the iteration cycle going, and obviously OpenAI and others—or OpenAI specifically—had been doing that iteration cycle a little bit more leading up to some of these moments than I think we had been doing at the time.
As you do that iteration cycle, and at the same time all of a sudden everyone wants AI, the 500 million tokens a month is a 50× increase over the previous year. At the same time that you're setting up the right organizational structure and doing the iteration loop to make sure that you're actually training the world's best models, you also need to scale up hardware. That doesn't happen instantaneously either.
We need TPUs to do research. We need TPUs in order to do inference. The amount of TPUs that you need isn't just sitting there idly waiting. There's work involved and timelines involved and all that.
All things considered, given the constraints, I think we're in an incredibly good position. The thing that gets me most excited is the slope of improvement across all those dimensions. How can we work better as a team and get everyone on the same page? How do we keep making sure that the breadth of research upstreams back into the main Gemini models? How do we make sure we have the world's best infrastructure? How do we make sure that we get that iteration cycle for releasing new models down, and sort of learn the hard lessons and develop rigor around that?
I think we're doing all those things, which has been incredibly exciting to see. I think the last comment I'll make is that DeepMind has also transitioned, and I think we talked about this before, from being an organization that did foundational research to now actually building products. I think that's the last step of this organizational journey: How do we build the Gemini app? How do we think about what we do for developers? How does that actually influence how we train models and what that iteration cycle looks like?
All of that stuff has now happened and the work has been done, and I think we're putting the pieces in the right places. I think now, for the next 3 to 5 years, we get to see the outcome of making good and hard decisions to put things in the right place.
Nathan Labenz
Yeah. If I had to summarize that—and maybe contrast, and I won't ask you to contrast—there's another big AI research organization out there that has a sort of fragmented structure, tons of compute, and researchers who have been pursuing lots of different directions for a number of years. If anybody hasn't already identified it, that is Meta, and the contrast has been pretty strong.
One possible explanation would just be that leadership at DeepMind saw that, yeah, we're getting close. This actually seems like it might be a thing now, and so it's worth going through all that trouble of reorganizing. I'm sure there are many things Demis would rather do than reorganize an organization and redraw lines of reporting—who reports to whom and whatever. But if you're close and you're getting on that wartime footing, so to speak—not that I ever wanted to see an AI war—but it's worth it.
You haven't seen that same kind of thing at Meta and maybe a few other companies, and you also haven't seen the integration of the work. I mean, they do have Meta AI in their apps, but clearly it's not on the same level. They also haven't done the reasoning thing. I'm sure they had Meta researchers at the same San Francisco parties getting those same very few big hints that, hey, this seems to be sort of working, that everybody else seemed to say, “Okay, we better make sure we're on that train,” and thus far they kind of haven't. So I guess for me, the takeaway there was maybe just the importance of conviction in leadership to do whatever it takes and push hard on small hints that seem credible. That seems to maybe matter a lot right now.
Logan Kilpatrick
Yeah. The other piece that I'll add is that I think incentives matter a lot, and for Google, the incentives in the DNA matter a lot. Google has been—and Sundar has said this many times, and I think he's spot-on—an AI company since Sundar took over in 2016, or whatever it was. People joke that Google was sitting on the Transformer and didn't use it. The Transformer was powering Google Search at multibillion-user scale in a bunch of different ways. So the technology was being used. It wasn't in the same incarnation as the current generative AI stuff, but it was being used at that level of scale.
Building models, deploying them, building that infrastructure, and that iteration process had been in Google's DNA. It obviously needed to be reformulated a little bit for the current team that's doing that across Google.
The other piece of this is just the incentives as well. If you look across Google's products, Google, organizationally and in terms of what the future of our products looks like, is so incentivized to make great models because the great models that we make—and this is an interesting thread that we should talk about—are present across all of our products. You're writing in Docs, you're doing things in Sheets, you're in Waymo, you're doing stuff on YouTube with video, you're a Cloud enterprise customer and you're doing something—all of those use cases end up benefiting from this.
It's not just an add-on thing. It is fundamental to the success of those products. So I think there's an interesting angle to this around how incentivized Google is to be successful. I think we're highly incentivized, and it's in the DNA of what the company's been doing for the last 10 years. I think those two things as well—if you don't have them, it makes this moment probably a lot more painful than it would have to be otherwise.
Hey, we'll continue our interview in a moment after a word from our sponsors. In business, they say you can have better, cheaper, or faster, but you only get to pick two. But what if you could have all three at the same time? That's exactly what cohhere Thompson Reuters and Specialized bikes have since they upgraded to the next generation of the cloud. Oracle cloud infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs where you can run any workload in a high availability, consistently high performance environment and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper, OCI costs up to 50% less for compute, 70% less for storage, and 80% less for networking. And better in test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is the cloud built for AI and all of your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com/cognitive. That's oracle.com/cognitive. Build the future of multi-agent software with agency. Agncy. The agency is an open-source collective building the internet of agents. It's a collaboration layer where AI agents can discover, connect, and work across frameworks. For developers, this means standardized agent discovery tools, seamless protocols for inter agent communication and modular components to compose and scale multi- aent workflows. Join Crew AI, Langchain, Llama Index, Browserbase, Cisco, and dozens more. The agency is dropping code, specs, and services, all with no strings attached. Build with other engineers who care about highquality multi-agent software. Visit agency.org and add your support. That's agncy.org.
Nathan Labenz
That 500 trillion tokens per month is 50,000 tokens per month for every human being on the face of the earth, which is a pretty crazy number.
Logan Kilpatrick
That is crazy. It’s grown a lot faster than I expected it might. Obviously, there are other providers out there processing a lot of tokens, too. So we’re now getting into the regime where—we should acknowledge that not everybody’s using it—we’re starting to get some significant inference numbers on a per-capita basis.
Nathan Labenz
As you look ahead, do you think we’re going to continue to see this sort of convergence, where the leaders will mostly be doing stuff that’s measurable on the same bar charts? Or do you think we’ll start to see more divergence, which could mean different form factors and significantly different strengths and weaknesses? Who knows what divergence might look like, but what’s your expectation there?
Logan Kilpatrick
Yeah, I would guess we see more divergence, to be honest. I was actually just at a dinner a couple of nights ago, talking to some founders, and a bunch of them were clearly betting on the fact that there’ll be model convergence. I think it depends on what sort of—what level of abstraction you want, as far as what things are going to converge.
My general sense is that the low-hanging fruit has been captured. So now it’s about what structural advantages you have as a business to train LLMs. I think Google has a really important infrastructure advantage in the ecosystem, along with a bunch of other things like that, where I think you’ll actually see those things shine through. I think getting to this point was not unexpected. Getting to the next level is not going to be easy for a lot of people to do, and that’s where the world-class teams and folks who are really making an order-of-magnitude bet on this are going to see the advantages.
Intuitively, through the lens that it only gets harder from this point, AI innovation does not become easier after this point. Making these models better is going to be more difficult. I would guess a bunch of teams start to—and it’ll be interesting to see what the size of the labs is. Maybe all the labs will keep doing everything, but I think there really will be opportunities to focus on a specific area.
I’m just conjecturing here, but you could imagine Anthropic deciding, “Hey, we actually just want to be the world’s best coding-model company. That’s the only thing we care about, and that’s what success looks like.” I don’t think this is going to be true because they have a very broad mission of what they want to do, and it doesn’t seem like it’s specifically code. But you could imagine that there are companies doing some angle of this where they end up deciding, “Hey, there’s actually value in diverging from this path of something super general, because we could build a really great business by doing that.”
I think it’s, again, at odds with some of the broad missions that these companies have, but I do think there’s real value in that. You can really go deep and start to understand how to build a long-term business and company around some of those things. Maybe the big labs won’t do that, but there are obviously more people training foundation models than just the big labs. A lot of those companies are going down the path of, “Let’s find something that we’re really good at. Let’s build a differentiated perspective on how to solve this problem, whether from a model perspective or an infrastructure perspective.” And I think that makes a lot of sense, honestly.
Nathan Labenz
Yeah, that’s interesting. I don’t know. I should be at least somewhat deferential—you probably know better than I do—but I just look at how fast these core foundation models from the leaders are getting better, and I’m like, I would not want to be a Tier 2 foundation-model trainer in today’s world. Especially if it’s only getting harder from here, and you’re saying that from the DeepMind position, I’m like, boy, that sounds really hard from any other position.
I guess it won’t be for you to sign on to this in this moment, I don’t think. But, transparently, my position is that I just think the big tech companies are going to win anything and everything that they want to win. You’re kind of seeing this bleeding into the application layer as well, right?
I’m interested to hear—I know you’ve been very focused on supporting developers directly via AI Studio and the APIs. You’re also, I’m sure, in regular dialogue with folks like Cursor and Windsurf and anybody who might use Gemini 2.5 Pro as a coding model. But now we’ve also seen all 3 of the big frontier developers in this last wave put out a coding agent, too, right?
How are they feeling, and how are you talking to them about the fact that they’re using it? You want them to use the model, but you also now have a competing product on the market against them, right?
Logan Kilpatrick
Yeah. Well, a couple of things. One, I think the product that we do have—you’re referencing Jules, right?
Nathan Labenz
Yeah.
Logan Kilpatrick
At least for us, Jules is definitely super early. I’m super excited. It’s a great team inside Google that’s working on it, but obviously the level of adoption that some of these other AI coding products have is very much on a different level.
Nathan Labenz
500 million ARR—they just said today.
Logan Kilpatrick
Yeah, I saw that tweet, which is exciting for them. I was talking to someone last night, and the comment that I made—which continues to be true, and I don’t know if I’ve said it on another episode that you and I have talked about—is that there’s no better time in human history than right now to be building a startup. Truly, if you’re building a startup to build language models and compete against all the big labs, you better be very well-capitalized to do that, because that’s a very difficult problem.
If you’re building out the application layer, it’s never been easier. The time to build software, the opportunity to explore new ideas, and the pace at which this current AI moment enables you to potentially scale monetization, build really retentive user products, and build a profitable business—all of these things have never been easier. The barrier to doing all those things has never been lower than it is right now.
As somebody who fundamentally believes in developers changing the world, I think that’s the coolest opportunity ever. Sure, the big tech companies will hopefully be successful as well, and will sell infrastructure and do some things at the application layer. But the real opportunity is that there are a million and 1 different problems to be solved, and some of these large, billion-user products solve things in a really general way.
The cool thing is that you can really go deep for some specific user segment and solve their problem in a unique way. The cost to do that—from building a startup and writing the software to do it—has never been lower. In the startup world, there are a thousand and 1 different AI tools that you get to leverage in order to get to that place.
Larger companies, just because of the level of security, privacy, and enterprise requirements, often don’t use a lot of those tools. It’s a different, bespoke set of tools, and the pace of tooling innovation for large companies often happens a little bit slower than it does for the startup ecosystem. You have all of these entrenched speed advantages, and then you couple in the idea that everyone’s going to have a bunch of agents building stuff for them in the future.
I continue to be super excited for people, even in the coding space at the application layer, who are building stuff. There are so many cool things to be built.
Nathan Labenz
The importance of speed—or the criticality of the advantage of speed for startups—I think is definitely extreme now. It’s always been true, I suppose, but it seems like it’s taken on extreme importance now.
I recently talked to Andrew Lee from Shortwave, who’s building a Gmail on top of Gmail, but also a Gmail competitor. He said, “After really soul-searching deeply, we came to the conclusion that our only advantage is speed.” Focus is the other piece.
Logan Kilpatrick
I think this goes to big companies, and I feel this as well. There’s lots of tension for me because the cool thing about Google is that there are a million and 1 innovative things happening. The challenge is, how do you actually balance that and take action based on the million and 1 innovative things that are happening? It’s a real burden.
The nice thing for startups is that you don’t have a million and 1 innovative things happening. You can just go and do 1 thing. I have a lot of envy for folks who have that because you just don’t need to make a lot of decisions. You can really focus on solving the problem at hand.
I think the speed of execution and the ability to focus on just a single thing is a blessing, so take advantage of that as much as possible.
Nathan Labenz
What do you make of this Windsurf news lately? The brief story is that they agreed to a deal with OpenAI. They had been using Claude as their primary model, and then Anthropic cut them off by virtue of having agreed to this deal with OpenAI. On the face of it, that's all pretty reasonable, but if I am a coding-agent company or whatever that's thinking, “What's my long-term prospect here?” right now, the speed advantage goes to the startups because the big tech companies have been so friendly, I guess, to the rest of the ecosystem as to put the models out before they've implemented them in their own products, in many cases.
But it's not too hard to imagine that flipping, right? If Google wanted to say, “Okay, Gemini 3, we're going to deploy it in our own coding agent, Gmail, and Docs, and then a few months later we'll put it in the API,” that would definitely flip the speed advantage on its head. Do you think startup founders should be worried about that?
Logan Kilpatrick
Yeah, that's an interesting question. I think, on the Anthropic piece, I do think that the byline of Anthropic wanting to sort of invest in who they think will be long-term partners and getting compute to those customers—I think, actually, as somebody who's spent a bunch of time thinking about how we get compute to the right teams that are building products, I have empathy for that argument. I think that could make sense. It's totally defensible from a simple business-strategy perspective, right? Don't support your competitors. I think that's totally defensible in many, certainly in normal business contexts.
As far as our strategy, the great thing for builders is that Google Cloud is the 5th-largest enterprise business in the entire world, and the mandate of Cloud is to bring this infrastructure to the rest of the world—to bring Google's infrastructure to the rest of the world—so that people can build world-class startups and not need to rebuild the level of infrastructure that Google built in order to scale the internet to where it is today. It's such a core and foundational part of the business that I would find it hard to believe the strategy shifting from shipping across our own services, but also shipping across Cloud services.
Interestingly, oftentimes today it's actually even more extreme than the picture that you painted. The external developers often have an even larger speed advantage from a model perspective because, if you think about who the customer of models inside Google is, it's teams that are building billion-user products and teams that are building 150-million-user products. I was just talking to someone about some of the features and products that exist inside Google Workspace, and some of the ones that I've never even thought about have 150 million monthly active users, which is crazy.
That user persona—if you go and talk to enterprise users of LLMs—they don't move the LLMs as quickly. They don't switch models as quickly because, even though they're inside of Google, there's still all of the normal constraints of building a large user product. You don't want the behavior to shift, and you have to do a ton of evals. You have to do all these things, and all of that requires time.
When you have a small product, it's very easy to quickly switch models, and you don't really have to think about it that much. But I think for teams inside of Google, they do have to think about that, and the responsibility to the users that we have is to be really thoughtful about that. So, again, I think the time horizon would shift so dramatically as far as getting these models out the door if the strategy became, “We have to sort of force internal teams to use these and then deploy them before we give them to external developers,” that they'd be so far off from where we are today that I have a hard time imagining that would make sense.
Also, from the business perspective, it's important for us to give LLMs to developers because it's a core part of the Google Cloud business, which is, again, a huge business for Google.
Nathan Labenz
I know Daniel Cocatello. I know you guys—I don't know how well, necessarily—but you overlapped at OpenAI. I'm sure you're familiar with his AI 2027 scenario. One of the interesting things in that is that he projects that, basically, over these next 2 years, model developers are going to start widening the gap.
I saw Roon not too long ago on Twitter. Somebody asked, “How much ahead is what you have internally versus what we see externally?” And he said, “You guys have no idea how good you have it. It's 2 months. You're on the bleeding edge, just behind where we are internally.” But Daniel's projection is that this will change, and that for multiple reasons—including wanting to use the models intensively for their own ML research automation, dreams of recursive self-improvement and takeoff, and who knows what—the gap is going to widen.
He has pretty aggressive scenarios in mind there. He thinks that, basically, the developers are going to start to hold the best models back for themselves. The public will kind of satisfice more often, and really insane stuff will be held very closely and known to few people. It sounds like you don't buy that scenario, basically, or at least don't see any signs of that happening at Google.
Logan Kilpatrick
I think there are 2 dimensions. First, it's hard to get signal on how good models are. Evals are just such a difficult problem, so you often don't really have an intuition as to whether this could be the right model long-term if you don't release it to the world. I think that's just one of many pressures on the idea that you should get the model out the door.
I would underscore the momentum—sort of a quasi-momentum war—that happens. It's really important to project what the external momentum looks like from an AI perspective because, ultimately, I think if you go and talk to developers and people who are building companies, that's actually a really large influence on who they end up building on.
I also think there's a bunch of other threads to this, including that switching stuff is hard. There are so many layers that I have a hard time buying that we're not going to deliver models to the world in the same way that we're doing it now. I think there are just many levels of motivation and game theory that tell me that won't be true.
It's also interesting to think about how you actually capture the most value from this technology from an economic perspective. Maybe it's not us. The cool thing about developers is that, obviously, Google has large distribution, but you get this really wide aperture of distribution across so many different things.
Maybe the economic model looks slightly different, where the unit of intelligence is a token, assuming the models are way better and can do all this economically productive stuff, and you charge people on a per-million-token basis. I could buy that—that changes in the future, and the economic model looks different from how developers do it today.
But I still think fundamentally you would want to build a great business giving that to other people, because how you're going to use this model looks very different from how other people are potentially going to use it. You could build a great business doing that by releasing it to the world.
Nathan Labenz
Yeah, the importance of feedback definitely is not to be missed. That also connects back to the “what's going on at Meta” line of conversation—which, for the record, you're not commenting on, but I'm just tangentially mentioning. They're not getting nearly as much of that, right, as the companies that currently have the best models on the market are getting.
So, yeah, that's interesting. How much do you use other companies' models? Do you go and do your own different vibe checks across different providers? What's your model diet?
Logan Kilpatrick
Yeah, I play around with a bunch of stuff. I think it's interesting. It's fun to see how—I'm also, independent of my job, somebody who loves technology and loves seeing cool AI products—so I spend a lot of time playing around with all the coding models. I think that's probably the thing that I experiment with the most.
But there's tons of cool stuff happening in the audio space right now. We launched our native audio model at I/O, which was one of the threads, and it's available in NotebookLM and a bunch of other products, as well as for developers. It's been really interesting to see that as an emergent space that people are building products and services in.
Nathan Labenz
I've been spending a bunch of time playing around with the products and models. ElevenLabs just landed a new model to do something similar—native, really robust, natural-sounding audio. It's been super cool. Across whatever dimension you want, there's tons of fun stuff to play with.
I still try ChatGPT occasionally to play around with it and see what that experience has evolved into. It's fun to be somebody who likes using this technology. It is interesting, though. I'm not one of those people who sends the same query to 3 different models, examines the differences between them, and does it in 3 different tabs.
These are people I have a lot of respect for who are running companies and all this stuff, and I'm like, "That seems like a pretty cryptic way to be doing that type of experimentation." That has led me to think there's probably an interesting product to build there, where people are really trying to understand the nuances and differences between models and engage with multiple answers. I think the multiple-pieces-of-content thing is a really interesting thread to pull on.
You can imagine that in the future and in product experiences. But I'm not at the level where I have Claude, Grok, ChatGPT, and Gemini open at all times, asking my question in 3 or 4 places. I don't do that all the time by any means, but I do it on occasion.
My general philosophy is always to try to be doing 2 things at once: 1 being whatever the object-level task is, and 2 being learning about AI's ability to help me with that object-level task. If I have any sort of contract review, that would be a great example. If I'm going to take on some advisory agreement or whatever, they'll send me the contract.
I'll send it to at least 3 AIs. If none of them have an issue with it, I'll just sign it without even reading it myself. Usually, they are pretty consistent, but that's also an interesting opportunity to see just how they're presenting things a little differently. Claude is typically the shortest and the least formatted.
I don't know. It is hard to characterize. How would you characterize it? Gemini 2.5, for me—which you just launched a new version of onstage at the AI Engineer World's Fair—I can't claim any deep familiarity with the new one because it just came out yesterday. But the Gemini 2.5 Pro class of models, I think we're on the third date-stamped version, right?
It was one of those hair-raising moments for me because the command of the context window that it has was just so incredible. I dumped a research codebase into it that had 400,000 to 500,000 tokens. No other provider, at least with the level of access that I have, even supports that length of context.
To see the command that it had of it was incredible. I was literally going back and forth debugging problems in AI Studio. Don't tell anybody, because they might cut me off for this behavior, but it's rewriting whole files for me. Then I'm saying, "I got this bug. Please fix it," and it's rewriting another version of these long scripts for me with 500,000 tokens of context.
That was like, "Wow, this feels like that step change." I wonder what other step changes you would highlight that people might not be fully aware of, or what more subtle vibes and tone differences you think distinguish Gemini from other options.
Logan Kilpatrick
Yeah, this whole model-behavior piece has been really interesting to see. Folks have a strong reaction to default personalities, and I think we're very early in coming up with a rigorous point of view as far as how to make a default personality that works well through certain lenses.
If you look at the requirements for building the Gemini model, the baseline Gemini model is used across so many different products, even inside Google. Those products have such varied points of view and products and users that they're building for, so it becomes really difficult to come up with a default personality.
I know Claude and Anthropic have done a ton of stuff with trying to make the model personality feel distinct, and they have a point of view about what that should look like. I think this goes back to the advantage for startups. Anthropic gets to do that because the consumer product is relatively small compared with 1-billion-user products.
It has been interesting to see us take a more middle-of-the-road approach—not try to have too much of a personality, but also make sure that the model can have that if that's what the product people want to build. The best example of this is the Gemini app. The Gemini app probably wants to actually have a personality and do some of that stuff.
I think the challenge becomes how you maintain that. I've seen that time and time again: as you change the models, the personality changes dramatically. If you're intimately conversing with these models, it feels like the person or model you were talking to before is now gone and has been replaced by something else.
I think that's a pretty jarring experience for today's model-iteration process. There is some interesting stuff to happen to make that not be the case. I actually have a tweet queued up that I need to put out about long-context capabilities, because far and away, 2.5 Pro in its current iteration has a really large gap.
I'm in my era of tweeting things live right now, just because it's top of mind. I'll put this out, and then I'll send you in the chat the tweet that I just put out. It's far and away one of the gaps in model performance right now. Long context is clearly one of those gaps. I'll put it in the chat for us.
Nathan Labenz
I just opened Twitter, and there it was—8 seconds ago. This is showing OpenAI's MRCR, which is a long-context eval that OpenAI built. You can see the delta between the models. Far and away, the latest version of 2.5 Pro is on the order of 20% better.
Logan Kilpatrick
This is 8 needles, which is the hardest version. It's not just single-needle context, which is retrieving 1 thing from the context window. Even the Gemini 1.5 Pro model from over a year ago was close to 100% accurate. That was basically a solved problem.
The problem is exponential decay as soon as you start adding more needles. To see this level of progress from a model perspective, with the ability to process and find 8 distinct items, is pretty remarkable—especially given that there wasn't a whole lot of long-context innovation in the last year.
I think it's this combination. I've had this conversation—I was just with Jack Ray yesterday, who leads our reasoning team and our reasoning efforts and originally worked on long context. I've talked to him a lot about this fusion of long context and reasoning.
Really, it's reasoning that enables you to use the full context window that's available. It's cool to see that happen. It's cool to see that happen finally.
Nathan Labenz
Yeah, you can feel it. Of course, benchmarks and practical use are not the same thing. I don't know what I would have said about Gemini 1.5 in terms of whether it felt like it had that depth of command.
I did a few things where we put whole books through it and asked it to find relevant quotes, and it could do that pretty well. But this new thing, if people haven't tried it recently, is worth exploring. Video is one of the best use cases, I think.
I've been doing this, and it's a fun experiment. Take a long video you've watched, then go and ask a bunch of questions. The challenge is that you perhaps already have to watch the video, but that use case tends to shine, which is really interesting.
Logan Kilpatrick
We've also, interestingly, seen a shift in what the distribution of request sizes looks like because of how good the context window has gotten. Historically, we'd been like, "Why does no one really use long context that much?" We were definitely ahead of the technology and the curve from that perspective.
Now, with 2.5 Pro, long-context usage is dramatically higher than it has historically been. It's been awesome to see people coming around and building on it. I think it gets closer to this future of the whole long-context-versus-RAG discussion.
Historically, you could have brushed it off a little bit because the model wasn't really that good and no one was really doing it. But I think the future is going to look more like people putting more and more stuff into the context window. Of course, they'll still need RAG in some cases, but it's awesome to see that corner turning from a long-context perspective.
Nathan Labenz
Yeah, "turn your hyperparameters up" is one of my current mantras. Another good one, for people who want to get a qualitative sense of the command that the new models have of long context: I have a simple Colab notebook for extracting emails from the Gmail API, where I just use “‘from:me’”—basically, just email sent by me—as my simple search.
Nathan Labenz
So filter out all the crap I’m not engaging with, but just threads that I’ve sent a message to, and pull all of those. You can go back, depending on your volume of email, pretty far and get a pretty robust picture of who you are that still fits into 1 million tokens. Then you can start to get a sense for what the model understands of you from those 1 million tokens, and it’s pretty impressive. I can share that Colab notebook if anybody wants to mess around with it.
I put it in that format so that you can do it without your data ever having to leave Google. It’s just going from your Gmail to your own Google Drive via the Colab notebook, so it was the most secure way I could think to make it that I could share with you. I don’t want to have your email. That’s the last thing I need.
Nathan Labenz
Okay, you’ve mentioned a few times being with people—dinners, talking to founders. I guess there are 2 angles on that. One is, what are you looking for in those groups?
Everybody who’s building wants to be in the inner circle of early access programs, the trusted tester rosters, and all that kind of stuff. How do people get into those programs? How do they get into the trusted tester sets? Also, can you enable the Veo 3 API for me, please? How much is networking—how you’re keeping up—versus other ways of keeping up?
Logan Kilpatrick
Yeah, that’s a good question. I think this looks different depending on what you’re doing. For me, maybe—I’m not sure how well this will track across people—but for folks who are listening this far into the conversation, I assume you’re an AI enthusiast, really, after the first 5 minutes. I love that.
So send me an email. Honestly, we have a super-robust early access program. We’d love feedback from people building interesting things. If you’re building something interesting, email me at lkilpatrick@google.com. Send me an email. I’d love to hear about what you’re building, and I’d love to get you into the early access program.
It’s not some big thing. Some stuff is more secret, and some stuff is less secret. We really just love to work closely with developers, get feedback, and be as open and collaborative with people as possible. So email us.
Hopefully, the Veo 3 API is a work in progress. We don’t have one available at the moment that we can onboard people to externally. We’re setting up a bunch of things and also working on how we can make it so that the model’s order of magnitude of scale for the API product, relative to putting it into a consumer product with a high price point, is just different. It’s a lot of different dimensions.
We’re working on ways to make sure the model can actually work at the scale of demand we’re going to see from an API perspective. It’s incredible to see the audio really bringing video to life. Historically, I’d been pretty skeptical of a lot of the video models. It was cool to see the video generated, but the practical use cases were hard for me because the amount of work it would take to do something meaningful with that video was pretty substantial.
I’m curious, for Waymark—I’ve used the product, but I haven’t used it recently—how important audio has been as part of that story. I think it really brings the video to life for me now that it has audio, and the audio feels like it’s actually native to what the video was meant to be saying.
Nathan Labenz
Yeah, it’s incredible. First of all, I’ve been thinking recently that we’re quite lucky. I don’t think anybody planned this, but there was a lot of hand-wringing about deepfakes and fake voices, cloned voices making calls, including a little bit from yours truly around the election that didn’t really come to pass. I think maybe that was mostly because the models weren’t quite there yet.
We’re fortunate that it’s landing early in a cycle where, hopefully, by the time the next election comes around, we’ll have enough reps, people will have built up cultural immunity to it, and there’ll be more guardrails and whatever, such that hopefully we’ll be able to deal with it. But it is getting to the point now where I genuinely don’t always know if something is AI video or real video.
For Waymark in particular, audio has been really important. We have traditionally, and still do, take the approach of just having a voice-over track. We mostly make TV commercials, and we mostly partner with big media companies. YouTube ads are a natural part of that. These are all sound-on environments.
Anytime you see a TV commercial, there’s usually a person talking to you, and then there are visuals, and there might be on-screen text and images and what have you. That’s our usual approach. We use a mix of providers, but ElevenLabs has certainly been a very important provider for us, and their voice quality just continues to climb.
With Veo 3, it kind of opens up a new dimension. In the past, we mostly used images that the businesses have. The next step is, well, if we can bring those images to life by doing image-to-video with even a Veo 2, then that just makes the whole thing more dynamic.
Quality is really important there. I would say Veo 2 mostly hits the mark, but sometimes has a little bit of weird stuff. Honestly, it is pretty damn good. But with Veo 3 now, it’s like, oh, you could even rethink the form factor a little bit. You could imagine having the voice-over talk a bit, but then also flipping over to a clip and having that thing present in a different voice.
It definitely opens up the space of possibilities for us in terms of the sorts of stories we can try to tell. We’re mostly telling small local business stories, but there are a lot of different ways to tell them. We’ve been relatively narrow in that space over time just because the technology could only do so much.
This is the kind of thing that our creative team sees, and it’s up to them to figure out exactly what they would want to make out of this new thing now that you can have all kinds of different voices showing up in a real context like that. Honestly, I think we’re still wrapping our heads around it and are also somewhat limited by the fact that we’re still just testing it in the actual top-tier Gemini app.
My personal AI spend is up to about $1,000 a month, which is also an interesting thing. I’ve been saying that for a while, but I hadn’t actually gotten there. Now I’m pretty much there between OpenAI, Claude, and Gemini, all at the top level, plus 20 other things that I’ve accumulated. It’s amazing to be spending $1,000 a month on AI subscriptions, but some of these things you’ve got to have.
Logan Kilpatrick
Yeah. Do you think you’re getting that level of value out of them? I assume, given the position that you’re in, that some of them are duplicative because you want to test all the different stuff. But if you were to remove the duplicative ones and just had whatever the best was across a bunch of different categories, do you feel like you’re getting that level of productivity boost relative to what you’re spending today?
Nathan Labenz
No question. If I weren’t committed to testing everything and having the earliest point of view on things that I can get, I think I could get a very similar productivity boost for much less. But the productivity boost is still dramatically higher than what I’m paying. No doubt, the acceleration of all sorts of different work is tremendous.
Last week, I was traveling a bit and ended up coding 2 different apps on Replit, just with the agent doing almost everything for me. It’s starting to feel like delegating work to other humans, much more so than a few years ago when we talked about prompt engineering. The original prompt engineering, right, is setting things up so that a natural completion of what you provided would be what you wanted.
Now I literally don’t think that much about the fact that this is even AI. It’s more just, “Here are some product notes,” and if it messes up, then I’m like, “Why did you mess up? Did I mislead you or something?” I have to think a little harder. But I’m really struck by how the communication to the AIs now feels much more natural and much more high-level. No doubt, the boost is tremendous.
Logan Kilpatrick
This is the eval that I think is one of the most exciting ones to me: relative to the amount of money that you’re spending, how much value is that creating? It’s hard to measure, I know, so it is somewhat theoretical. But that is what I think, long-term, as we move away from all the regular academic benchmarks being saturated, et cetera, is the economic productivity that’s created by some of these systems.
You need a broad sort of mandate in order to do something like that, just because there are so many possibilities—an infinite number of possibilities. But it is really interesting to think about, and I do think it’s a cool north star to drive up the amount of value you can create in the world in a very positive way through a $20 subscription.
The value you get today from a $20 subscription relative to what it’s going to be in 5 years, I think, is actually materially different. It’ll be cool to see that play out.
Nathan Labenz
Yeah. I think another dynamic that’s going to be interesting to watch is what, if any, stable equilibrium we ever arrive at. I think right now one of the reasons that there’s so much surplus for me is that we’re not yet in equilibrium.
Nathan Labenz
And so, to a certain degree, I have superpowers that other people don't have, which they could have, but they don't because they're not aware of them or they just haven't developed the habits. A lot of it is honestly just thinking to do it in the moment: go use the AI instead of doing it manually, for whatever version of it you might be considering.
Back in the holidays late last year, there was a project—and I wasn't involved in the business side of this at all—but somebody basically came to me and said, “Hey, I've got an audio production project. This company typically has a big network, and they want to do a ton of local radio ads, and I thought of you. Maybe you could do it.” They usually would pay a couple hundred bucks per location, per version of the ad. This would end up being in the six figures, but they were wondering if I could do it for less.
The discount that we were able to provide to this company relative to what they were used to spending was probably around 75%. Still, the revenue per hour that I actually spent on it was probably around $3,000 an hour—not all of which came to me, by the way. But that sort of disequilibrium, I think, doesn't last forever. A lot of people will figure that stuff out over time. So I do wonder, in that project in particular and in general, whether I'm still being compensated based on assumptions that have not fully taken on board the fact that productivity can—and in some places has—significantly jumped. I think maybe it'll happen in the future. I don't know.
Logan Kilpatrick
Yeah, I think there are still going to be those edges in the future, which is interesting. If anything, I think the pace of innovation is going to go up. Going back to my comment from before, that doesn't mean it's not going to be more difficult, but I do think the pace of innovation is going to continue going up and to the right. Because of that, I think there are going to be a lot of discontinuities and opportunities. Being on the frontier is likely to be disproportionately rewarded because you're using all the tools and stuff like that, which is super interesting to see play out.
But there are so many edges and so many opportunities that are left as the frontier keeps moving forward that, even if you're just showing up today and saying, “I'm not on the frontier,” there are probably 50 things that you could go and explore that end up being super, super interesting and a force multiplier.
Speaking of things that are on the frontier and super interesting, let's talk a little bit about agents. Obviously, everybody's talking about agents in all sorts of different ways. Here's my horseshoe theory of agents: I've found that the latest things, whether it's Claude Code or Jules or any of these more agentic models that take multiple steps and do bigger mini-projects for you, feel much more like the original ChatGPT to me, in that the mode of interacting with them is very turn-based.
What's happening is that the turns are getting bigger. The output is getting bigger and, hopefully, more valuable—hopefully more accurate—in order to be able to do all that stuff and succeed. But you're still on a one-off basis. You're still on the hook as a human for figuring out: Did it do what I wanted? Did I ask it the right thing? Is this actually working for me at all or not? And how do I proceed based on what it did? I have to evaluate that on a step-by-step basis.
And then, in the middle, is where I think people are actually getting scalable automation value. They're not letting the AI choose its own adventure. They're not just saying, “Here's 50 tools and a goal. Go,” which can sometimes create these magic moments, but often doesn't do what you want.
In the middle, it's a much more structured paradigm, whether it's LangChain or whatever, that's like: “We're going to break this thing down into its constituent parts. We're going to have 8 different prompts for the 8 different steps. There might be a couple of little forks or double-back points in there. So we'll give the AI some discretion to choose exactly what route it's going to follow, but it's a pretty on-rails sort of system.” Those seem to be the things, from what I've seen, where people are actually getting to the point where the reliability is high enough that they no longer have to look at the output on a task-by-task basis.
So how would you coach people as they think about the spectrum from the original chatbots that are now familiar to workflows, agents, agentic systems, and autonomous systems? How do you see that spectrum, and where should people be? I'm sure you've got lots of thoughts.
Logan Kilpatrick
Yeah, my take right now is that, with reasoning, it's become very clear that a lot of the scaffolding will move into that layer. You'll send a request and provide a bunch of scaffolding to the model in the reasoning step. Today, it'll have access to search, code execution, a code sandbox, tools, and function calling, but models in general are on this trajectory to become agents out of the box, which is really interesting.
They'll have all these capabilities baked in, and the thing will be able to do a lot of things. Of course, there will be limits to what it does because you don't build everything into it, but by default it will have access to do a lot of things like that. You can imagine having a bunch of other hosted tools and things like that, which then sort of gets the data flywheel spinning, as far as actually being able to build and train the models to go and do that.
Then you can imagine some of those trajectories that you're describing, with the flows that the model goes through and the way that it tries to solve problems. All of that ends up also being upstream of the model. So I do think the models are on that path to be systems and agents out of the box.
But the practical reality is that there will still always be a need for scaffolding. I think it's this balance of how you make the current version of the product that you want work in a way that likely needs to use scaffolding, but you don't build it in a way that ends up being a one-way door. As soon as the model can do that thing, you'd have to fundamentally rework it. Maybe the coding models are good enough that rewriting everything from scratch actually won't be that difficult, and it'll all be fine.
Historically, if you have a larger product, it ends up being really hard. I think this is actually a transition that I've talked to a lot of companies and products about. They're in the middle of this AI 2.0, LLM 2.0 transition moment, where they had actually built a lot of the original tooling around the fact that models weren't good at a lot of things. So they had all this additional scaffolding, all these additional layers and systems, and it was actually a pretty complex system to make LLMs work in production at scale.
Now that the models have become so good and can do a lot of these things natively, you can actually remove a lot of that complexity. Again, depending on the complexity of what you built, that can actually be really, really difficult. I think the folks who built the scaffolding and the complex system did the right thing because they wanted to make that product experience work. They probably benefited from the fact that they were AI-native and powered by it from the beginning, and hopefully won a bunch of customers and business.
But I think you also need to make sure that you can continue to adapt, because I think the models will be able to do more and more and hopefully take on more and more of that burden. I've had an increasing number of conversations with people who are in that boat of going through that transition right now, and it's been specifically because of reasoning that this has become possible for a lot of people.
Nathan Labenz
Yeah. Long context obviously goes hand in hand with that. We've experienced that at Waymark, especially in image processing. I've told this story repeatedly, too, so just again, super briefly: it's kind of crazy how hard I had to work at one point in time just to have any minimal understanding of what a random user-uploaded image was. Now it's like, “Feed 100 images into Gemini Flash,” and it'll just tell you which ones to use. It's really simple.
So what was once a highly scaffolded workflow—and had to be, in order to get to the reliability point—now, for us, is basically just a prompt. That does seem like that sort of cycle will repeat. That sounds like basically what you're describing. Just making sure you're ready to rip out the scaffolding and convert it to a prompt as that moment starts to hit for whatever you're building is the recommendation.
Logan Kilpatrick
Yeah. I had a conversation with Josh Woodward, who runs the Gemini app and Google Labs, and he was saying how this played out for them in NotebookLM as well. Originally, to make those NotebookLM Audio Overviews happen, it was a 14-step process, and there were all these different handoffs and steps in the loop, most of them powered by Gemini.
Today, it's a 4-step process. It's dramatically simplified the level of complexity because the models are just so good at doing a lot of those things now that they don't need to have an entire bespoke system built around writing the transcripts for the Audio Overviews, which is really cool.
You actually feel that in the product experience in some ways, too, where the product experience has become a lot faster, and there are a lot of other things that are possible because you don't need 14 different independent LLM calls that all sort of have to happen in sequence.
Logan Kilpatrick
So it’s been cool to see the product experience actually benefit in a lot of ways from this level of simplicity that’s come as the models have gotten better.
Nathan Labenz
Any other things you think are really interesting, hidden gems, underappreciated, or just strong trends in the agent space? A2A is something I’ve been looking into and honestly haven’t really been able to wrap my head around yet.
Logan Kilpatrick
Yeah, I’ve got a non-agent thing that’s interesting and I’m happy to talk about, but I think on the agent side, at least for A2A, the quick mental model is that there are just parts of the agent-building ecosystem that MCP doesn’t solve. And I think A2A is trying to solve some of those. One of the examples is the auth model and things like that.
So there are parts of the story from putting agents into production at scale that still need to be solved. And I think it’s an open question where MCP is going to go long term. Is it going to do a bunch of those things? Is it going to leave space for other frameworks or standards to solve some of those problems? I don’t think we know yet, so I’m watching closely and interested to see what happens.
Nathan Labenz
Okay. What else is on your mind?
Logan Kilpatrick
Diffusion. Did you see the demo of Gemini?
Nathan Labenz
I did have it in here. I didn’t get to it, but yes. Did you get to play around with it yet?
Logan Kilpatrick
I haven’t used it.
Nathan Labenz
You could. I’d love to have it on that list as well.
Logan Kilpatrick
I’ll get you on the list. Unbelievably, first of all, it does make, in some intuitive sense, a lot more sense to me than the autoregressive model. When I reflect on my own pattern of thinking, I feel like what I’m doing is much more fuzzy and high-level first, and then it gets segmented down into parts. Then I try to do those parts, and at some level I’m writing sentences token by token.
That resonates far more than trying to sit down and write the whole thing linearly from the first token to the last, even with a reasoning model or a scratchpad place to mess around. So, yeah, would it surprise me if, in 2 years, the diffusion paradigm has won because this sort of coarse-to-fine structure turns out to be better? Not really. And damn, is it fast. It’s unbelievably fast. Unbelievably fast.
I’m really excited. I think even if there’s a world where the next-token-prediction paradigm continues, just for people who want to build products that have that level of speed, maybe it’s—I don’t know—it’s unclear at this point what the performance-trade-off characteristics will be. Is the cost going to be the same? All those things. So there are a bunch of open questions.
But assuming you could build product experiences for a similar cost to what they are today, with similar model quality, there are a lot of really interesting product experiences to be built if you have that level of speed. I think that’s actually the thing that could enable this personal generative UI experience to really happen. If the tokens can actually be generated that quickly, rendering on a screen in the blink of a human eye would be really, really cool to see.
So I’m super excited, and I’ll get you on the list for access to that. I think, even if it doesn’t end up working out, it’s just a good reminder that we need to be pushing in different directions, because there are other paradigms that I think could work. Maybe it’s not next-token prediction, and there are a bunch of properties of some of these things, like editing, which is becoming more and more common for a lot of these use cases, that the diffusion model seems to be really well suited for.
Nathan Labenz
Yeah, I suspect in the end it could be quite a bit better for a lot of use cases, too. I mean, it just seems so natural. I always say the transformer is not the end of history. Obviously, there is an attention mechanism in a lot of these diffusion models, too. So attention is still part of what we need.
What do you think we’re missing right now from AGI? Memory is one often-cited candidate. What’s on your list?
Logan Kilpatrick
Memory is definitely one of those. I think AGI is going to end up being much more of a product experience. If I have a hypothesis about how people are going to end up having the AGI moment, my assumption right now—and we’ll see if this plays out—is that someone is going to release a model that ends up being really good. It’s not going to be this thing where everyone says, “We’ve clearly built whatever your definition of AGI is.”
Which is also the problem: everyone now has a different definition of AGI. It’d be easy if we all had the same definition, but we don’t. So that’s the other problem: it’s not going to happen that way. I think it is going to be a product experience. Someone is going to weave together the right components at the product level with a model that’s really smart.
Maybe—I don’t know—the delta in how smart the model needs to be relative to today for this experience to actually work could be just that long context is 50% better and reasoning is 50% better, and then you somehow figure out a way for memory to work. The memory piece is actually a completely different engineering, neuroscience, and human psychology problem: How do you surface the right things at the right time?
I think someone’s going to build that experience, and people are going to say that the feeling of this thing is going to be AGI. Again, it’s really a product experience enabled by a model, but the model itself isn’t able to do all those things. It’s what happens when you take the model and build everything around it, and do it in a really thoughtful way, that people are going to say is sort of the AGI moment for a lot of folks.
So that’s my guess right now. And again, I think the models are doing more and more of this stuff, and you could imagine maybe the models are doing the memory stuff themselves and that gets trained into the model. I think that’s very far out there, but in the short term, it’s definitely going to be a product experience that gets us to AGI. That’s not what I think the AGI narrative is—it’s so model-driven right now. I just don’t think that’s actually how people are going to feel and experience what ends up happening.
Nathan Labenz
Really good memory work is coming out of Google, as you might expect. We recently did an episode on the Titans architecture, and there’s already been a follow-up to that. It’s now up to 10 million tokens of memory and probably able to go beyond that, but they’ve demonstrated up to 10 million tokens with pretty strong memory performance. So, yeah, it could be coming sooner rather than later, but it is kind of a distinct module, right? Obviously, our brains have many modules, too.
Okay, maybe last question, then I’ll let you hit on anything else you want. One of the striking moments from I/O was when Sergey was asked in a little fireside chat what he thinks the future of the web is going to look like in 5 to 10 years. He almost spit out his coffee at that moment, where he was like, “The future of the web?” He’s like, “I don’t think we know what the future of the world is going to look like in 5 to 10 years.”
And that’s a striking reminder that even the people pushing the frontiers of this technology don’t have a crystal ball and don’t really know what exactly we’re getting into. So I wonder what your expectation for the future of your life and your job is in the next, let’s say, 2 to 5 years. Are we going to get a drop-in Logan replacement? I mean, NotebookLM is going to replace me. Do you think you’re on the chopping block in the next 2 to 5 years for AI replacement as well, or how do you see this shaping up?
Logan Kilpatrick
Yeah, I was sitting in the front row of that fireside chat next to Cory, who’s our CTO in DeepMind, and Emanuel Tropa, who drives a bunch of our infrastructure stuff. It was fun to see their reactions as well to some of the conversation.
I have such a fundamentally human-centric view of the world. Even in today’s world, as somebody who builds AI and thinks all the AI products are cool, I write everything I do personally. All of the work that I do, every email that I write, every tweet that I write, is written through my head, and I have, in probably 95% of cases, zero AI assistance involved in that process.
It’s because I have conviction in my worldview and because I have conviction in my tone. I think maybe you can make a loose approximation of someone, but the reality is that I want to be the entity that has agency over the things that come out around who I am. I think that’s fundamentally important.
I think people will end up having this fundamental question for themselves: Who do they want representing them? Even if I have this digital twin that knows all the things, and maybe it could make loose approximations, and I could say, “Yeah, that seems reasonable. I could potentially see myself saying something like that,” do I actually want that thing going and saying those things on my behalf? Probably not. That is a very foreign concept compared to what humans do today.
I think maybe the only exception—not a notable exception to this, but an example against this—is people who run companies. I can imagine that, if you have a large company, you’re like, “Oh, some team or some person is representing Google, as an example.” Someone’s saying, “Maybe I wouldn’t have said it that way, or I wouldn’t have phrased it that way,” and they’re still representing Google as a whole, but they’re sort of an independent agent on behalf of it.
I think that, unless you’ve had that experience, it’s still fundamentally different from the human experience of, “I want some other entity representing me.” It’s not clear to me that people are really going to want that experience. I personally don’t want that experience right now, and that’s my personal opinion. That’s the decision that I’m making.
But I do think it’ll be interesting to see where the balance is, as far as how much people do that. This goes back to—and I have a bunch of these random convictions on my personal website—I think the value of humanity, using you as an example, Nathan, is that this podcast is exponentially more valuable in a world where AI can generate human-sounding things, analyze content, and put together research reports.
The reality is that the next-token prediction coming out of all of those systems isn’t the next-token prediction coming out of your brain. Or, if you want to use that example instead, the diffusion of thoughts coming out of your brain—that’s what I care about. I care about your perspective because you’re another human, and we have shared lived experiences, we’ve done stuff together in person, and all that stuff.
I think there are places where you won’t care about that, maybe because of the type of content. There are certain dimensions where that won’t matter as much. But I really do fundamentally believe that humans are interested in what other humans have to say. When I think about someone sending me AI content that was written or generated by AI, I just care a little less. I’m not really that interested.
I can kind of tell they’re not willing to put in the craft and the time to do something. Why am I that interested in it? Again, there are exceptions to this. Software is a great example: if someone builds great software, do I care whether or not a human wrote it? Not really. Maybe in some cases I’d appreciate it more if a human did it, and maybe not in some cases.
It is interesting. There’ll be a spectrum. I’m not worried. I think the way that people do work will shift in some capacity. I think the value of having a differentiated perspective is also just going to be incredibly beneficial in a world where intelligence isn’t the sort of limiting factor in a lot of ways.
All that to say, I’m excited for another 500 to 600 podcast episodes from you over the next 2 to 5 years.
Nathan Labenz
Yeah, thank you. We’ll see how many I can tick off.
I think there’s a lot to appreciate in your thoughts there, and I’m with you on some portion of it. There’s this whole notion that if people don’t have jobs, they’ll have no meaning, and I’m not on that train. I definitely am on the idea that one of the great things about AI could be that it would allow people to make much more of an effort, or put much more of an emphasis, on making connections with each other.
At the same time, I’m like, I don’t know. LLMs are getting awfully good, and they can handle any topic on demand. I’ve noticed that in myself. NotebookLM isn’t a huge fraction of my listening yet, but it is starting to eat away at my listening to other podcasts.
There are times when I want to know about a specific thing, nobody’s done a podcast on it yet, and NotebookLM will. It has that background knowledge, so even if it’s a little worse in some ways—and maybe it is—the expressiveness of the voice and all that is getting pretty good, too. It hits a spark in some other ways that really matter.
I think the fundamental point here is actually a great example that Sundar gave. This is a great comparison between search and AI chat product experiences. Everyone 2 years ago was saying, “Now that ChatGPT has however many hundreds of millions of monthly active users, search is going to go away.” Yet Google searches are growing, the number of queries on search is growing, and the search business is still growing.
It’s because they’re actually, in some sense, solving fundamentally different problems. That NotebookLM example you gave—“I want on-demand entertainment about this very specific topic, and maybe no one has created that type of content before”—is kind of a different use case.
Maybe this doesn’t fully track across all podcasts, but there are lots of podcasts that I listen to where I’m like, “I just want to hear what this person has to say about this.” That’s the thing I actually care about, and I’m willing to listen to them talk about whatever. But if I’m trying to learn about some very discrete task that I know nothing about, the chance that one of my top 5 favorite podcasts has talked about that is probably slim.
You need some other mechanism in order to do that. Maybe that’s not fully true, so I’m curious what your reaction is to that. But I think that is going to play out across a lot of other domains and dimensions, where this thing is actually net additive.
It puts pressure on things in some capacity because there’s a limited amount of time in the human day, but it doesn’t end up being as disruptive as it would look on paper.
Nathan Labenz
Yeah, I hope not. That time limit is a very hard constraint as it stands right now. I’ve recently started to increase my listening speed; it had been 2x by default.
Actually, YouTube just increased its mobile maximum speed. These are the kinds of things that really move the needle for me. It used to be capped at 2x, and I don’t even know what the maximum is now, but you can go well beyond 2x. Now I can listen to things at 2.5x by default and save myself another 6 minutes.
That saves me 6 minutes on a 1-hour piece of content. This is the way I’m trying to pack more and more in. I don’t think I can go too much farther down that path. At some point, in terms of competition for time and attention, you hit some sort of fundamental limit.
Maybe we get Neuralink working, and that’s the next big unlock, where the bandwidth increases so dramatically that all bets are off. I do think there’s a credible line of thought, honestly, that upgrading human cognition in deeply integrated ways is going to be necessary.
That’s sort of Elon’s brief pitch for why he founded Neuralink: to be able to go along for the ride with the AIs. Especially when you see these diffusion models. The autoregressive ones are already faster; they can write. Gemini can write much faster than I can read, and then the diffusion models are another order of magnitude faster.
The speed of it all is just going to be another wild thing to contend with. Now you’ve got these video models. I just saw one that is generating video dynamically in real time, and you can interact with it.
I share a lot of the excitement and enthusiasm, and I’m also like, I don’t know that I can compete with all this stuff. It just seems like it’s going to get really, really good at everything, be ubiquitous, and be so personalized to each individual user, kind of knowing what they know and what they don’t need to know.
How do I compete as somebody who’s a one-size-fits-all with not a huge audience? There are enough people out there that I certainly can’t customize the podcast for each one. And when AI can do that, I think the point is what people want: the Nathan experience, I think, is the point.
Logan Kilpatrick
That’s at least my fundamental bet and conviction. In the long term, in a world where I could spin up 1,000 podcasts that look something similar to yours, people want your perspective. There’s value in that, even though it’s not the highest-order optimization of the delivery of the content or whatever.
That’s my bet, so we’ll see if that ends up being true. But I have conviction in that bet. So, hopefully it’ll turn out right.
Nathan Labenz
Yeah, I hope you’re right. Well, I know many people want the Logan experience, and I know we’re over time, so I appreciate you sharing so much time and information with us today. I look forward to doing it again in the not-too-distant future.
Logan Kilpatrick
This was great, Nathan. Thank you for having me. It’s fun to chat, and hopefully I’ll see you in person again soon.
Nathan Labenz
Cool. Veo 3 diffusion model. Put me on your list.
Logan Kilpatrick
I will.
And with that, Logan Kilpatrick, thank you for being part of the cognitive revolution. It is both energizing and enlightening to hear why people listen and learn what they value about the show. So, please don't hesitate to reach out via email at tcrturpentine.co or you can DM me on the social media platform of your choice. [Music]