[BidClub_]
The Cognitive Revolution · · 57 min

Breaking: Gemini 2.0 Flash Goes Live - Inside Google DeepMind's Latest Release with Logan Kilpatrick

Nathan LabenzLogan Kilpatrick

YouTube
TL;DR
  • Google put the updated Gemini 2.0 Flash into production at 10 cents per million input tokens and 40 cents per million output tokens, while previewing Flash-Lite and experimental 2.0 Pro. Paid Flash defaults to 2,000 requests and 4 million tokens per minute, with higher tiers reaching 10,000 requests and 10 million tokens per minute. Logan Kilpatrick expects DeepMind’s tighter integration with products to accelerate both research and product progress by making the organization “end to end” from model creation through deployment.

  • Gemini’s strongest commercial wedge is cheap coding intelligence for text-to-app products such as Bolt.new and Cursor, with Lovable and v0 as possible adopters. Kilpatrick contrasted startups spending roughly $40,000-$50,000 monthly on LLMs with an estimated Flash bill of $1,000 or less—a “40x cost reduction”—and argued that falling costs particularly empower developers without tier-one VC backing. Google’s higher-end ambition remains categorical: “We’re going to have the world’s best coding model at Google,” with Pro and reasoning expected to carry that effort.

  • Flash-Lite primarily protects Google’s low-price positioning, while Pro is designed for capability-first experimentation. Flash-Lite preserves the 1.5 Flash price point of 7.5 cents per million tokens and does not support expensive features such as native image or audio generation; Nathan Labenz questioned whether moving from 7.5 to 10 cents could truly break a business. Kilpatrick largely agreed, conceding that continuity also prevents “Google’s raising the price for developers” from becoming the story. Pro’s price had not yet been announced because it was experimental, but Kilpatrick expected it to follow prior Pro models and be much more expensive.

  • The multimodal Live API points toward persistent AI “co-presence,” but today’s economics and architecture still impose a 10-minute session limit. Usage reports span coding help and blind users trying to navigate daily life. Separately, Labenz described using Advanced Voice Mode to guide his family through Mario 64. Kilpatrick expects iteration on cost, memory, state and context before always-present assistants can scale, but sees the direction as “very clear.”

  • Long context may become substantially more useful when paired with reasoning rather than immediate answer generation. A 1 million- or 2 million-token window can answer questions about a few relevant facts, Kilpatrick said, but synthesizing “a thousand different things” remains difficult; reasoning could let models work through the material and eventually move information via tools. Native image and audio output adds another vector: specialist image models may remain higher-quality in some domains, but Gemini can trade some quality for world knowledge—unlocking cases where prior models “weren’t smart; they were just good at generating images.”

  • Kilpatrick’s startup map centers on vision-language systems, reasoning-enabled agents, agent-facing internet infrastructure and better evals. He sees domain-specific computer-vision stacks as “up for grabs,” predicts reasoning will make more currently brittle agent products work over the next two years, and expects websites to need new protections, attribution and value-capture mechanisms once non-human visitors become normal. The evaluation problem remains difficult: benchmark sprawl obscures model choice, while even sophisticated teams struggle to turn taste into repeatable tests—“most of the problems in life end up being eval problems.”

Digest · the substance, structured for research

1. DeepMind now spans the model-to-product loop

  • Kilpatrick joined Google 10 or 11 months earlier and saw close DeepMind collaboration from day one. The organization has since moved from fundamental research toward production, with the Gemini app moving over and, within the prior three months, AI Studio and the Gemini API joining the same broader effort.

  • His description of the new structure is unusually integrated: DeepMind now “end to end does the research, creates models, and then actually brings them to products inside of Google.” Removing friction between researchers and product development should make it easier to extract capabilities from new models.

  • For outsiders uninterested in Google reorganizations, Kilpatrick offered a concrete test: the change should produce “an acceleration of model progress” and “an acceleration of product progress.” Labenz’s outside view was that both already appeared to be speeding up.

2. Cheap coding intelligence broadens the software-creation market

  • Kilpatrick’s favorite emerging use case is text-to-app creation. Bolt.new had just added Gemini, Cursor was using 2.0 Flash, and he hoped Lovable and other builders would follow as non-programmers increasingly create basic domain-specific software from prompts.

  • Labenz called the shift a “software Supernova”: people without coding skills can now make credible basic full-stack applications, even if enterprise platforms remain beyond them. The important transition is from coding assistance to software creation by people who never intended to become developers.

  • The cost contrast drives the opportunity. Kilpatrick had seen startups reporting roughly $40,000-$50,000 in monthly LLM usage; he estimated comparable Flash usage might cost $1,000 or less, “like a 40x cost reduction.” He described continuing to push Flash’s capabilities while limiting dramatic price increases.

  • Yet cheap intelligence does not help every company equally. Well-funded YC startups may barely notice another cost reduction, Kilpatrick argued, whereas individual developers can pursue use cases previously reserved for teams with millions in financing.

3. Real-time multimodality is becoming AI co-presence

  • The multimodal Live API supports collaborative, real-time interaction across video, audio and text, pointing toward what Kilpatrick calls “co-presence”—an assistant able to see what a user sees and interact alongside them. Usage reports span coding help and blind people navigating ordinary life.

  • Separately, Labenz described using Advanced Voice Mode with his children during Mario 64: when they could not find a star, he showed or described the level and asked where to go. Even his one-year-old had begun reaching for the phone to “talk to AI.”

  • Kilpatrick cautioned that the present system is still early. Sessions are limited to 10 minutes because continuous presence is expensive, and scaling it requires better memory, state and context across everything a user has previously shown or discussed.

  • The analogy is text LLMs, which moved from “kind of a toy demo” to billion-user-scale applications in roughly a year and a half. He expects co-presence to follow an iterative path—and sees repeated claims that AI products are merely low-value wrappers as continuing to be wrong.

4. Reasoning may unlock long context’s real value

  • Long context already works in production, and Kilpatrick believes 200,000 versus 1 million or 2 million tokens can matter. The limitation is attention: asking about a few items works well, but combining “a thousand different things” across the full window remains difficult.

  • His emerging hypothesis is that reasoning provides the missing layer. A model that can think through a huge context—and eventually use tools to bring information in and out—may turn long context from impressive capacity into a practical enabler.

  • Kilpatrick proposed a direct experiment in AI Studio: run the same long-context task through 2.0 Pro and a reasoning model in compare mode. His intuition, explicitly not a settled result, is that “the extra reasoning steps actually make a difference.”

  • Gemini still consumes video, audio and images natively, while native image and audio output was available to early testers. Imagen 3 was scheduled for the API the day after recording; dedicated generators may retain a quality edge in some domains, but Gemini’s world knowledge could support uses where prettier, less knowledgeable models fail.

5. Vision-language models can collapse vertical computer-vision stacks

  • Labenz asked whether Flash’s low cost could support passive monitoring: factory-floor safety, policy violations or fall detection for seniors who dislike wearable alarms. These workloads could consume the enormous token volumes behind Kilpatrick’s earlier suggestion to spend “a dollar a day on Flash.”

  • Drawing on his past work as a machine-learning engineer, Kilpatrick recalled how difficult traditional security-camera systems made tracking a person from one camera frame to another and maintaining object permanence. Vision-language models perform this kind of understanding extremely well, while rigid domain-specific systems may not be fault-tolerant in many cases.

  • He had not spoken with anyone running this exact scenario in production, preserving an important hedge, but considered it an obvious opportunity. Gemini’s spatial-understanding demo can already identify objects and draw bounding boxes in a way comparable to custom bounding-box models.

6. The 2.0 portfolio pairs production scale with experimental breadth

  • The launch made an updated Gemini 2.0 Flash production-ready at 10 cents per million input tokens and 40 cents per million output tokens. Google also previewed the smaller Flash-Lite, released experimental 2.0 Pro, and rounded out the lineup with a Flash reasoning model.

  • On Flash’s free API tier, Kilpatrick cited roughly 10 or 15 requests per minute and 4 million tokens per minute. Paid usage removes the daily request limit, defaults around 2,000 requests per minute, and was gaining quota tiers up to 10,000 requests and 10 million tokens per minute.

  • That capacity requires “lots of TPUs,” but model proliferation complicates allocation. Google cannot productionize every experimental variant when Cursor, Bolt and other scaled customers need dependable compute, so general availability reflects infrastructure commitment as much as model quality.

7. Flash-Lite protects pricing continuity while Pro chases coding

  • Standard 2.0 Flash costs more than 1.5 Flash, so Flash-Lite preserves the earlier 7.5-cent-per-million-token point while offering a better model. It can remain cheaper partly because it will not support higher-end capabilities such as native image or audio generation.

  • Kilpatrick cited the earlier Flash-8B model’s leading token volume on OpenRouter as a reasonable proxy for usage in some contexts and evidence that developers value low-cost models. Labenz remained skeptical that 7.5 versus 10 cents could disrupt a viable business; Kilpatrick agreed the decision was partly about avoiding a price-increase narrative.

  • Pro follows the opposite strategy: start with the best model, prove the application, then reduce costs through smaller models, optimization or fine-tuning. Its price had not yet been released because it was experimental, but Kilpatrick expected it to follow prior Pro models and be much more expensive. Coding is its strongest relative domain, it retains a 2 million-token context, and Kilpatrick still believes Google will produce “the world’s best coding model.”

  • Ultra’s future is framed around cost and infrastructure tradeoffs, not abandonment of pre-training scale. Kilpatrick said Google is still scaling pre-training, while asking whether a hypothetical model that scored 5% better but was five times larger and costlier would deserve the infrastructure—especially when reasoning offers substantial “low-hanging fruit.” Releasing another Ultra remains an open research question.

8. Model choice is an eval problem, not a benchmark answer

  • Kilpatrick’s rounded-table-corner debugging story punctured his own model intuition. Claude solved the task in one shot after Gemini had frustrated him—but when he reran the exact prompt with differently formatted question and context, Gemini also solved it. The episode exposed the “vibe-based, incredibly unstructured, almost scientific” way developers compare models.

  • Public evidence is fragmented across “20 random benchmarks here and 50 random ones here.” Kilpatrick called for one platform collecting benchmarks and leaderboards, while Kaggle was exploring personal evals that automatically run against new releases and alert developers when a model fits their stated needs.

  • Labenz’s pushback is worth keeping: many valuable outputs have no objectively correct answer. For video creation, his team can detect structural failures, excessive length and other “clear thou-shalt-nots,” but quality beyond those guardrails remains taste, expert demonstrations and “vibes.”

  • Letting users choose models can at least give developers data and users more agency. Copilot’s move from GPT-only service to a model dropdown exemplifies the trend; similar benchmark scores can still conceal consequential behavioral differences, as Labenz’s comparison of DeepSeek with more behaviorally refined Western models illustrated.

9. Fine-tuning, reasoning and an agentic web define the startup map

  • Kilpatrick remains “incredibly bullish” on personalized fine-tuning: everyone could eventually use a version of a model carrying needed context without that context excessively distorting its priors. But 2.0 Flash tuning was not yet available, and 1.5 Flash still had rough edges such as no image fine-tuning.

  • Asked directly about reinforcement learning, he confirmed that RL is part of building Google’s reasoning models. The company had taken its customary low-key approach with experimental releases, but he expected Google to explain more of that work as reasoning attracts broader attention.

  • His strongest forecast was that “reasoning is going to make agents work.” Many agent companies currently tackle valid problems with products that simply do not function reliably; over the next two years, he expects reasoning to unlock more of them than any other model advance.

  • Agents also challenge the internet’s old social contract that visitors are humans, apart from indexing crawlers. Kilpatrick sees room for companies handling agent access, website protections, attribution and value capture. He also sees an opportunity in evaluations, while acknowledging that he does not know how the industry will turn human taste into a programmatic test—“most of the problems in life end up being eval problems.”

Logan Kilpatrick

We released the experimental first iteration of Gemini 2.0 Flash back in December. Today, we brought an updated version of Gemini 2.0 Flash into production so that developers can actually continue to build with it. We announced pricing of $0.10 per million input tokens and $0.40 per million output tokens, which I think is a huge accomplishment for us to pull off.

We're going to have the world's best coding model at Google. I still believe this deeply, and I think Pro is going to be that model. A bunch of the reasoning work that we're doing is going to be that model that continues to push the frontier for us.

Nathan Labenz

Logan Kilpatrick from Google, product manager of the Gemini API and AI Studio, welcome back to The Cognitive Revolution.

Logan Kilpatrick

Thank you for having me, Nathan. I'm excited. I'm hopeful that I'm getting close to the record for the most times on your podcast, so I appreciate you.

Nathan Labenz

I think this might be setting the record at 4, if my count is correct. Congratulations. That's rare air and well deserved.

It's launch day, so we'll get to everything that you've launched and what we should be thinking about building with it. A quick detour before we get there: you're not part of DeepMind, so Google obviously is a vast company and is continuing to align, restructure, and streamline itself to focus more and more on AI. What's the story from the inside on what it's like to be at DeepMind now, specifically?

Logan Kilpatrick

I'm super excited about this. I joined Google 10 or 11 months ago, and literally from day 1, it's been a deep collaboration with DeepMind. DeepMind has gone through all these evolutions over the last few years, transitioning from an organization doing fundamental research to actually productionizing models. Then, within the last 3 months, with the Gemini app moving over, and now with AI Studio and the Gemini API, it's actually an organization that does the research end to end, creates models, and then brings them to products inside of Google.

I think that's been a shift for them, but from my personal vantage point, I think this is the thing that makes the most sense. Being really close to research—we already were really close to research through this collaboration we've had—but removing as much friction as possible to bring together the researchers who actually know how to get the most capabilities out of the models makes a lot of sense, and it's going to be a ton of fun.

As an external person who doesn't care about Google reorganizations, which is most of the world, the thing that you'll hopefully see is an acceleration of model progress, but also an acceleration of product progress, because we're bringing these 2 teams together.

Nathan Labenz

It sure seems, from my vantage point on the outside, that everything is accelerating. We've had previews of some of the things that are now going into general availability today over the last few months, and of course there has been 1 advance after another from DeepMind and others over the last few months.

Looking back a little bit, what would you say are the customer success stories, or just the coolest apps, that you've seen come online that have been built with the Gemini API recently?

Logan Kilpatrick

I think the thing that I'm most excited about, and that also feels like we have the biggest opportunity in, is all these text-to-app creation software products. There are a bunch of examples: Bolt.new just went live, I think yesterday, with Gemini support; Cursor has Gemini support now and is using 2.0 Flash; and hopefully we'll see others, like Lovable and v0, have that support as well.

If you look at the economics of running those products, it's extremely cost-intensive. The number of people who know that you can put in a text prompt and get a basically working app or website for free is still very small. It feels like this is a new frontier use case and product paradigm that will be picked up across all the big players, but I also think there's going to be a ton of startup activity around how you build domain-specific software for people without those people actually having to know how to code.

I'm really excited about that use case, and I think we have a lot more model progress to do to become the world's best model at coding. But even for 2.0 Flash and 2.0 Pro, compared with where we were 6 months ago, it's just an incredible amount of progress. Trying to keep pushing progress in the context of Flash without increasing the price in any dramatic way has been the biggest win for us.

I've seen tweets about the LLM usage costs for some of these startups. You can imagine what those costs would be with Flash: probably $1,000 or less a month, instead of $40,000 or $50,000 a month. That's a 40x cost reduction, which is crazy.

We also previewed the Multimodal Live API in December, which allows you to have this collaborative, real-time, conversational video and text interface with the models. That feels like it's getting us closer to the future we've all been promised with AI, which is this co-presence that's able to see the things that you do and interact with the services you use.

I've gotten the entire spectrum of outreach from people using it to help them code to people who aren't developers and are just using the product experience. Some of those people are blind and are trying to navigate their daily lives using this tool. It's crazy that the product we're building as a demo for developers to drive adoption of the API is actually helping people who are just trying to live their daily lives.

I think it speaks to where we are in the adoption of this technology. Everyone is waking up and trying to find the best way to use these tools, which is really interesting to see happen.

Nathan Labenz

Let's put a pin in the Lovable and Bolt discussion, because you teed me up perfectly. I think immediately after this episode, the next 2 episodes are going to be with the founders of those 2 companies.

We're calling it “Software Supernova,” because it does seem like we're in the middle of a transition and a tipping point right now. It's becoming pretty realistic for people who don't know how to code to create at least basic full-stack applications. Obviously, they're not going to create enterprise platforms just yet, but the models continue to advance, so it feels like we're going to be in a very different world quite soon with products and paradigms like that.

In terms of living my daily life, I've tried both the AI Studio version, which is the desktop experience where I can share my screen with the multimodal Gemini, and the mobile version. I don't know why, but OpenAI only has that experience enabled on the mobile app, as far as I know, so I've been trying both of them.

This is too stupid, but it's also too good of an example. I got my kids a Nintendo Switch for the holidays, and we're going through historical video games. I feel like they're young, and the modern games are too overstimulating. Plus, let's burn through the old catalog first and work our way up.

We're doing the original Nintendo games and Nintendo 64. We're playing Super Mario 64 for the original Nintendo 64, which is this open-world game where you go around and hunt stars and whatever. The frustration for me is that I often don't know what to do, so I've been sitting there with Advanced Voice Mode on, sometimes showing it the screen, telling it what level I'm in, and having it tell me what to do in the game: What is my objective? Where should I go to find the star?

My kids are really getting used to this. Even my 1-year-old sometimes comes over and wants to take the phone. He's trying to talk to AI. For seniors, too—I think about my grandmother all the time—this switch from typing to it and getting text back, to it really being co-present with you everywhere, feels like a very different world for people who aren't anchored to their desk all day doing computer work.

Imagine putting it in glasses, too. I'm sure that, coming from the company that made Google Glass, there's a lot of thought going into that sort of thing.

Logan Kilpatrick

I feel like we're actually going to see a lot of this. If folks have been watching this closely, it takes time for this progression to happen. Text was the best example: it was kind of a toy demo a year and a half ago, and now, at large scale, text LLM applications are broadly being used throughout the world at billion-user scale.

With co-presence, we have to be humble about where we are today. We're going to have to iterate on the API, bring the cost down, and make it something that developers can actually build. You could imagine that the cost of having AI with you all the time is probably pretty expensive. There's a reason we limit sessions to only 10 minutes right now, because there are a whole lot of challenges to scaling this up beyond that and having it maintain the memory, state, and context of everything you were just talking to it about.

It's very clear to me that this is the direction we're going in. Everyone, 2 years ago, was talking about how everything was a wrapper and there wasn't a lot of value to be created. That continues to be wrong. It feels like there's a new thing, and then all of a sudden all these new things that weren't possible before can be created.

That continues to get me excited about the future, and about making sure we're enabling those next things so that people show up and create the experiences that don't yet exist today. Hopefully, send over feedback as you keep using the real-time mode.

Nathan Labenz

Taking 1 more beat on the cool stuff with the Gemini API: last time we talked about a couple of things I wanted to check in on. One was really long context, another was insane affordability, as you mentioned—a huge drop compared with other options—and then there's the question of exactly how natively the video is being consumed. You can feed video and audio multimodal inputs into the Gemini API as well.

What would you say is the status of those things right now? Are you seeing cases where people say, “200,000 tokens is not enough. I really need to provide more,” and therefore Gemini is the 1 thing they can use? Or, cost-wise, this simply wouldn't be affordable if they didn't have the $0.10-per-million-input-token pricing?

Logan Kilpatrick

This is a great question. I had a long conversation with Jack Rae yesterday, who is 1 of the co-leads for the reasoning effort. He previously worked on the long-context breakthroughs with Gemini, enabling that from a research perspective, and is now 1 of the co-leads with Noam on the reasoning models.

We opined for a long time about how funny it is that the real unlock for long context might end up being reasoning. Long context is extremely impressive and useful, and we do see people using it in production. I think there are cases where 200,000 tokens makes a lot of difference compared with 1 million or 2 million, but 1 inherent challenge is how many things the model can attend to in the context window.

It works really well if you're asking questions about a couple of things in the context window, but if you're trying to put together 1,000 different things in a 2-million-token context window, it gets really hard to do that because of the inherent nature of how the models are trained and set up.

I think reasoning is where this starts to change. Having a really long context window where the models can think through things, and perhaps in the future use tools and bring information in and out of the context window, is where long context is going to start to make a lot more sense and truly become an enabler.

Video and audio are still happening natively, and images are still happening natively inside the models with the 2.0 release. We showcased some of the next steps of this native multimodality: the models actually being able to output images and audio. That's available to early testers. I don't know if you're in the Early Access program, Nathan, or if you've played around with it yourself, but it's pretty good. There is still more quality work that needs to happen.

We're also rolling out Imagen 3 in the Gemini API tomorrow, at the time of this recording, which I'm super excited about. If you look at why we still have state-of-the-art image-generation models when we know the models can natively have this capability themselves, there's definitely a quality tradeoff in some domains where you trade quality for world knowledge.

There are a ton of models that generate really pretty pictures and can do all kinds of cool things, but they lack the world knowledge that the Gemini models have because this is a native capability coming as part of the training process. I think there will be a whole new onslaught of use cases that didn't work before because the models weren't smart; they were just good at generating images. We'll see that happen with native image generation.

Nathan Labenz

That's interesting. Your point about the need to have a long chain of thought in order to really take advantage of super-long context is quite interesting. In application development, I do find that I want the hardest thought, so I'll go to the model that's going to think for me the longest. Then I sometimes have to contort myself, or contort my inputs, to get everything to fit into the context.

I wouldn't 10x-dump context into a model that's going to immediately jump to the answer. I hadn't really put together why that might be, but I think that's a pretty interesting hypothesis. I look forward to dumping my full-million-token codebases into Flash Thinking sooner rather than later.

Logan Kilpatrick

I'd love to see if you have use cases that haven't worked well historically for long context. I'm curious whether people in the audience have examples of what they can do with it.

You can use the compare mode in AI Studio: try it with 2.0 Pro with long context and try it with reasoning with long context, and see whether the extra reasoning steps actually make a difference. My intuition is that they will, which is exciting.

Nathan Labenz

Have you seen anything passive? This is another area where Flash has been famously inexpensive. You said that more people should try to spend a dollar a day on Flash, and that's a lot of tokens. You need passive applications for most people to get there.

I've been thinking about a couple of use cases. Could I use a vision-language model to monitor my factory floor for safety incidents or policy violations? My grandmother lives in a senior-living community, and the seniors hate wearing their fall monitors, so that's a constant battle. To the degree that they have the ability to make their own decisions about it, a lot of them would probably accept not having to wear that thing if there were another visual monitor that could keep track of their state at any given time.

Have you seen anything like that, where people are truly using the API passively—almost like the Internet of Things—to send signals into the API?

Logan Kilpatrick

This is so core to my thesis about what's going to happen in a lot of these domain-specific areas. If you take a step back, how would people solve that problem today? They would need to buy custom software that does it, which is probably expensive and might not generalize well.

In my past life, I was a machine-learning engineer, and we did a bunch of work with security cameras. You had to figure out how to track someone moving from 1 frame, where a camera is visible, to another frame, and how to maintain object permanence for that person. It's incredibly hard. It's not an easy problem to solve with traditional computer-vision technologies.

I think vision-language models do this task incredibly well, and the cost basis with Flash is now so low. I haven't talked to anyone who actually has this in production, but I have to imagine this is an opportunity that people are going after.

It's not just the bounding-box and image-understanding use cases, although those are really powerful. Being able to know where an object is—we have a good demo of this in AI Studio. If folks haven't tried the bounding-box capabilities, go to Starter Apps. There's an example in there, and you can put in images, ask it to identify the objects, and it will throw bounding boxes pretty much like you would get from 1 of those custom bounding-box models that you could find in open source.

These use cases just take time, but my guess is that as vision becomes more prominent, we're going to see the whole YC startup wave go after these ecosystems and industries that are using domain-specific vision models instead of a general-purpose model.

The cost basis is going to be wildly different, and you'll unlock all these use cases that those models simply aren't capable of doing. They're very rigid and can't be fault-tolerant in a lot of those cases.

Nathan Labenz

Let's get into what you're launching today. I saw an interesting tweet from YC, and the idea was basically that every time the frontier advances, all the YC companies go and see whether it can work for their use case. The report was that every time a new model comes out, some subset of the current YC batch's products start to work, while the others are waiting and continuing to build everything else with the expectation that they're going to get a model that tips them from not quite working to working.

What are you launching? To the degree that you can speculate, what is it going to make work that wasn't previously working?

Logan Kilpatrick

It's a great question. To draw a broader point, because of how much excitement there is about AI, the resource constraint hasn't made people think as deeply as they need to about this problem. Maybe this is because of Gemini models being at the frontier of cost per intelligence if you look at that as a ratio, but the YC companies are so well funded that when we bring the cost of intelligence down by some reasonable factor, it doesn't move the needle for these startups in a lot of cases. They have millions of dollars.

I think it will be interesting to see the outcome of that trend. If I had to guess, I think it gives a lot of power to individual developers who don't have this large amount of financial backing from tier-1 VCs and can actually push the frontier of some of these use cases and capabilities. That's a really interesting and cool phenomenon.

To answer your question specifically, we're launching a whole suite of Gemini 2.0 models. We released the experimental first iteration of Gemini 2.0 Flash back in December. Today, we brought Gemini 2.0 Flash—an updated version of it—into production so that developers can continue to build with it.

We announced pricing of $0.10 per million input tokens and $0.40 per million output tokens, which I think is a huge accomplishment for us to pull off. We also announced a preview of Flash-Lite, the smaller variant of the Flash model, which we intend to make available for production use very soon, along with its pricing.

Then we released the experimental variant of 2.0 Pro, which is the most capable frontier model we have, rounding out the full offering with the Flash reasoning model. Now we have the reasoning Flash model, Flash-Lite, the smallest model; Flash, which is the best performance-to-cost tradeoff; and Pro, which is the most capable model.

Nathan Labenz

Let's go through availability, too. You guys have high limits for my typical work, which is usually focused on proof-of-concept-type stuff. I honestly love my life these days. I never have to worry about the harder work of making something production-ready. I get to focus on that easier, faster ascent of making the proof of concept work.

When I go to AI Studio, grab the code, and go do things with it—with 1 exception, which I'll mention in a minute—I basically never hit rate limits. Even when something is still in preview, for my purposes, there's enough headroom for me to do all the testing that I want to do.

But if I'm Bolt, Lovable, Cursor, or any of these other products, I would be hitting those limits. What's highly scalable now versus what's still in experimental access? Give us the concrete details on how much of these different models we can use.

Logan Kilpatrick

For 2.0 Flash during the experimental period, I think the API has 10 or 15 requests per minute on the free tier, and 4 million tokens per minute, which is a lot of tokens per minute. That's probably why a lot of people don't hit the limits. The use cases where people reach out because they're getting rate-limited are usually when they have users, are trying to run internal evaluations, or are running leaderboards.

With production availability, if you're on the paid tier, there's no daily requests-per-day limit. You can send as many requests every day as you want. The default is 2,000 requests per minute, and it still stays at 4 million tokens per minute.

At midnight tonight, we're rolling out new quota tiers. As you continue to scale usage, you can unlock 10 million tokens per minute and 10,000 requests per minute to help those who need to keep scaling up.

Nathan Labenz

The infrastructure behind that is truly an incredible accomplishment.

Logan Kilpatrick

There are a lot of TPUs to make all this happen, and a lot of complexity in how many models there are. I see a lot of memes online about the bad naming conventions we have for our models, and I appreciate them, but I think they point to something else: there are so many different model variants.

It's really difficult. A lot of the feedback about our experimental model-release chain has been, “We love these models. Let us use them in production.” The challenge is that we have to be picky about which model we use in production. The compute footprint required to make sure that Cursor, Bolt, other YC startups, and developers trying to scale and build companies can get the compute they need is substantial.

We have to be more intentional about how we do it. Ideally, we would make every model generally available, everyone would take every model to production, and we wouldn't need to worry about it. But there are a lot of constraints involved in doing that.

Nathan Labenz

Help me develop my intuition for how I should think about Flash-Lite as it relates to Flash. My totally candid initial reaction was that Flash is already so cheap and pretty fast. Flash-Lite is 25% cheaper and presumably faster, but also slightly weaker. You've got the table of benchmarks, and it's a little lower on most of them.

Have you been getting demand for an even cheaper and faster model than Flash? That's hard for me to wrap my head around.

Logan Kilpatrick

I think the positioning is 2-fold. By default, the price of the 2.0 Flash model is more expensive than the price of the 1.5 Flash model. Given how much we had leaned into the low cost per intelligence of the models historically, we wanted to give people a direct option. If $0.075 per million tokens was what enabled your business and we showed up and said, “By the way, now it's $0.10,” that wouldn't feel like a great story for developers.

We wanted to provide an option that was not only a better model, but the same exact cost as what people were getting before. For the 2.0 Flash model, because of a bunch of constraints, it wasn't going to be possible to keep that same price. This was about making sure we didn't mislead customers into thinking they could continue to push the frontier of cost per intelligence at the same price.

There are also some features that aren't supported in the Flash-Lite models, particularly more high-end things. For example, Flash-Lite won't be able to do native image generation or native audio generation. There are a lot of things we can do to keep the cost down when serving those models at scale that the default 2.0 Flash model supports.

It's also similar to the Gemini 1.5 Flash-8B model that we released. You can think of Flash-Lite as another version of the small-model track we've previously done with Flash and Flash-8B.

If people have looked at OpenRouter before, Flash-8B was the model with the highest token-volume usage. That is a reasonable proxy for model usage in some contexts. The clear feedback was that developers love low-cost models, and there are a huge number of new use cases you can unlock by continuing to reduce the cost.

Nathan Labenz

I would love to hear from anybody listening who fits that description—where a $0.075-to-$0.10 price change per million tokens would have made a meaningful difference to what you're trying to do in the world. Reach out to me. I want to hear that story.

I can easily imagine people looking at a menu and choosing the cheapest one because they're processing something and extracting addresses from a stream of data. I have a hard time imagining how a business model gets disrupted by that kind of change, but I'd welcome that story.

Logan Kilpatrick

I think you're right. A lot of this is about how we tell the story to the world and how we show up. It would be easy for the narrative to become, “Google is raising the price for developers.” That's not the narrative we want, especially given how much we've pushed to reduce costs for developers.

It was important to preserve the continuity of having that low price point available for developers who care about it. Generally, though, I agree with you. The other signal we've gotten is that if you have better models, people will pay for them. That's not the limiter in a lot of cases.

Nathan Labenz

They're all cheap compared with human labor. That's such a striking fact that is often glossed over.

Let's do the Pro side. How should I think about Pro? You could compare it with Flash or with the other frontier models that are out there, but how should I understand it in the increasingly busy constellation of available models?

Logan Kilpatrick

Pro is for when you're not bound by costs in a lot of ways. We haven't released the price of the Pro model yet because it's still experimental, but it's probably going to roughly follow the pattern of the Pro models we've had in the past, which is to be much more expensive.

There's a traditional piece of advice for developers building frontier applications: go with the best model, even if it's a premium. Make your use case work, and then figure out how to bring the cost down over time by switching to a smaller model, optimizing, fine-tuning, or whatever it is.

It's important for us to continue honoring that flow, which I think actually works. That's the default experimentation path developers follow today.

Specifically, we're seeing the best performance relative to other domains in coding. I had a glib tweet—I don't remember when it was, maybe the day before o3 was announced—about how we're going to have the world's best coding model at Google. I still believe that deeply, and I think Pro is going to be that model.

A bunch of the reasoning work we're doing is going to be part of the model that continues to push the frontier for us in coding. It's a domain we need to win, especially if you believe in the continued growth of text-to-app creation and the acceleration of developers from an internal software-engineering productivity standpoint.

That's probably the best use case for it. It still has a 2-million-token context window, and maybe it will have a longer context window in the future.

Nathan Labenz

We've heard whispers of infinite context. I'm pretty sure you and I were sitting in that room together when Jeff Dean said we were getting infinite context at some point in the future. He didn't put a date on it at I/O last year, in that session we were in together.

Hopefully, that's going to be a huge unlock for people once we land it.

How would you guide people to think about the non-release of Ultra? Ultra was the largest scale we'd seen. We saw something similar from Anthropic, where people asked, “Wait a second, what happened to Opus?” Is there anything you can share to help people understand why we seem to have gone from small, medium, and large to just small and medium?

Logan Kilpatrick

The historical context on Ultra was basically about proving out the research direction that scaling would continue. They proved with Gemini 1.0, when the original Ultra model candidate came out, that this was the case. But there was also continual, rapid innovation in making models better. All of a sudden the Pro model was better, and now I'm pretty sure Flash-Lite is better than the original Ultra model.

It becomes a question of the cost tradeoff and the infrastructure equation. You could imagine a world where we have an Ultra model that's 5% better on every benchmark, is 5 times larger, and costs 5 times more than Pro. You start to do the math and think about, from a research perspective, where it makes sense to spend our time and energy, and from an infrastructure-footprint perspective, where it makes sense to spend our time and energy.

We continue to see gains from making models much higher quality at the same size, or even a fraction of the size, of previous models. Now, especially with reasoning, there are even more question marks in my mind. If we could get a model that was 10% better on every benchmark, does that make sense in a world where reasoning has so much scaling and there is so much low-hanging-fruit work we can do there?

I had a conversation with Jack yesterday about pre-training scaling and reasoning scaling. Google is still scaling pre-training, so this doesn't mean we're not going to keep scaling pre-training. Whether we officially release an Ultra model is an open technical research question because of all those constraints.

It's also somewhat tongue-in-cheek. We could rename all the models and say that we have an Ultra model if we wanted to. Maybe that would have been the right thing to do historically—to make Pro Ultra and remap all the other names—but I'm trying to honor the essence of what was intended through the Ultra model.

Nathan Labenz

I'm in favor of anything that maintains clarity in naming schemes, so I'm with you on consistency. That's the only reason I even knew to ask this question. If everything had been renamed, I'd be even more confused.

Let's go back to coding and think about Bolt, Lovable, Cursor, and Devin. There are increasingly many options, and I can't keep up with all my AI coding paradigms. The models probably can't keep up with all the models, either.

We have all these benchmarks. The new model wins on some and loses on others. An interesting thread that has gained some weight over the last month or 2 is that o1 is great, o3-mini is great, and Gemini Pro is potentially great, but for some reason people still seem to think Claude 3.5 Sonnet is the best coder and the best coding assistant. In fairness, people haven't had time to compare it with everything.

How would you suggest people think about this? The simplest thing would be to swap in a model in exactly the same situation where another model is performing best. But you could easily say that you're leaving performance on the table because different models may not have maximum performance under the exact same conditions. Each model has different conditions that elicit the best from it, and you have to find those conditions.

How do you know how much time to invest in that? It seems very difficult, even for me as an obsessive hobbyist. What's the best practice for absorbing a new Gemini Pro and comparing it with Claude 3.5 Sonnet and o3-mini in today's world?

Logan Kilpatrick

This is such a tough problem space, and I have an incredible amount of empathy for developers, founders, and people building things. There is no silver bullet.

I'll give a couple of reactions. My own personal example underscores the second point I'm going to make. I was doing the normal web-developer thing the other day, trying to get the corners of a table rounded, and I was smacking my head against the problem because of a bunch of weird constraints in the environment I was working in.

I was using the Gemini models, and at 1 point I got fed up. I thought, “Maybe the Gemini models just aren't that good. I'm going to try Claude and see what it does.” I went out of AI Studio, went into Cursor, tried the same simple prompt, and it worked in a single shot. Everything just worked.

I started messaging people, frustrated, saying that we needed to keep making coding better. Then, for my own sanity, I reran the exact same prompt with the Gemini models. It also worked.

This was a good example of how I was doing a bad job prompting. I had pulled myself out of an iterative loop with the model, then started over from scratch and formatted the question and context in a different way. It worked with both models.

That underscores the vibe-based, incredibly unstructured, almost scientific way in which people make these decisions today. I work closely with a lot of teams that have LLMs in production, and you would be surprised, with many of your favorite LLM products, how few evaluations people actually have to understand the metrics that matter for their product and service.

I think there are 2 sides to this coin. The world needs a platform that hosts all the publicly available benchmarks, leaderboards, and other resources. I find it incredibly difficult to navigate and get a snapshot of how good a model is. There are 20 random benchmarks here and 50 random benchmarks there, and I have to look at Minecraft because that benchmark is really cool. I love that benchmark, but everything is split up all over the place, and it's hard to keep track as a developer.

This is my job. It's what I spend every waking moment of my life doing, and it's still difficult to keep track of all of it. For people who have much less time and are doing other things, it's just hard to keep up.

Someone needs to build this platform. Maybe this is the Y Combinator call for startups that you and I should do: build a platform to help people bring all of this together.

Separately, the folks at Kaggle are pushing on the idea of having personal evaluations. You could build a platform where, as new models are released and made available to the world, your personal evaluation runs behind the scenes on those new models. You would get an email saying, “Based on what you've told us, this model might be one you should spend time checking out because it's really good at the things you indicated you care about.”

That type of platform and product experience—taking the burden off developers—is going to be awesome. The challenge is that you still have to create your personal evaluation to begin with, but that 1-time cost is a lot less than having to do all the setup every time a new model hits the market.

Nathan Labenz

If I were going to take 1 practical recommendation from all that for developers, I might say, “Let your users choose.” At least that way, you can get some data and they can feel a little more agency. There might be something good to find there.

I like the idea of personal benchmarks, too, but the things I care about typically don't have a right answer. At Waymark, it's always the same challenge: What makes a good video? We can have a language model judge, but now we're in hall-of-mirrors territory. We don't have evaluations that work for us.

We can detect outright failures to follow the structure or that something is too long, so we have some guardrails. We can detect the clear “thou shalt nots” of the task, but beyond that, it's still really just vibes.

Logan Kilpatrick

I think that's a trend you hit the nail on the head. I don't even know how much of this is a conscious decision that founders are making, but across many of your favorite LLM products—with the exception of Waymark—developers and end users have this choice today.

You go in, there's a model dropdown, and you can choose from most or all of the models available. A timely example is Copilot. Historically, Copilot was powered only by GPT models, but the developer community and the world have moved on. Now it has a model dropdown, and you can choose the Gemini models, the Anthropic models, or the OpenAI models.

More products are going to go in that direction. I think this is also a corollary to a point that people make about the commoditization or contraction of the delta between these models. I don't think that's actually true.

There's a lot of weird nuance that will continue to sprawl over time. You will get substantively different answers from different model providers, even when their capabilities are similar on paper. There is still all this stuff that will be different, and I think it will be important for people to continue having that choice, because those subtle differences have a big impact on the end-product experience.

Nathan Labenz

No doubt about that. Even DeepSeek over the last couple of weeks has shown a very different profile. Perhaps not surprisingly, given its source, its closest description is that its base model is probably closest to the base models from other providers. It can write in ways that are extremely compelling, but it's also much less behaviorally refined than the top-tier models created by Western providers.

A lot remains to be unpacked beneath the surface of any major new model release.

Let's talk for a minute. I know you have to go before too long. The last things I wanted to cover were fine-tuning—because as a developer, that's something I'm always interested in and would love to know the status of—and then what's happening with reinforcement learning at Google.

We've seen the thinking model, and it struck me that there wasn't much actual mention of reinforcement learning, whereas other developers have said, “We did a reinforcement-learning model.” DeepMind seems to have positioned it differently, although I assume a lot of the same techniques are happening under the hood.

I was also going to ask about your call for startups, because I know you've recently raised a solo venture fund. We'll see if we can get you some deal flow. Fine-tuning, reinforcement learning, and your fund—take as long as you have.

Logan Kilpatrick

I continue to be incredibly bullish on fine-tuning. I think the future is one where everyone in the world is using their own fine-tuned model—a version of the model that has the context it needs without that context overly impacting the model's priors and how it makes decisions. That's the rough explanation of how I think about fine-tuning.

We don't have it yet for 2.0 Flash. I think we need to. We've been having a lot of internal discussions about the size of investment we want to make, and personally, I think this is 1 of the biggest opportunities. I completely agree with your point about developers wanting this.

We'll keep pushing on it. It's not available yet, but hopefully it will be available soon. There are also a lot of limitations in how we do fine-tuning with 1.5 Flash today. You can't use images, and there are a bunch of rough edges, so we need to solve all those things and make it a first-class experience.

Nathan Labenz

Your second question was reinforcement learning.

Logan Kilpatrick

Reinforcement learning is part of making those reasoning models. Historically, why haven't we talked about it much? We took the normal, low-key approach to doing the release, which is what we've historically done for our experimental models.

As the world gets more and more excited about what's happening with reasoning models, we're going to start talking a lot more about that work. I'm excited for us to tell more of that story to the world.

The reasoning narrative inside Google is the thing that gets me most excited about the direction we're going in. There has been so much progress and so many breakthroughs, so hopefully we'll tell that story soon. We'll get you some people on the podcast to talk about it.

As for startups, this is still the moment to build impactful, interesting companies. Vision is 1 area I'm really excited about. Take the entire ecosystem built on domain-specific computer-vision models. I think all of that is up for grabs with vision-language models, and there will be a huge amount of startup activity in that space.

I think reasoning is going to make agents work. I haven't been an agent investor in a lot of ways, but there are so many companies trying to tackle problems with agents, and a lot of them don't work today. I think reasoning is going to be the biggest unlock.

We were talking about new models enabling some percentage of Y Combinator startups to suddenly work. I think reasoning is going to be the continual breakthrough that, more than anything else that happens in the next 2 years, makes those companies' products actually work. That's exciting for them and for the world as we figure these things out.

I think there's a whole class of startups that doesn't yet exist. I've started talking to some of the people building these companies about how the fundamental nature of the internet changes in a world where you can't assume the only thing visiting your website is a human, with the exception of index crawlers that make your content more discoverable.

The social contract of the internet isn't set up for that. It's going to be very interesting to see how fundamentally things change, whether websites get locked down, what protections are put in place, and all of that.

There's a very basic new experience around how we engage with the internet that's going to change over the next few years. People with businesses, websites, and companies are going to have to solve that problem, and they're not going to be able to do it themselves. They'll need someone else to solve it for them.

I think there are interesting companies that can be built to enable how your website talks to agents, how attribution happens, and how you're able to capture your share of the value creation that's taking place. I'm interested to see what happens in that space.

Nathan Labenz

That seems like a really good candidate for what will change the world most in the not-too-distant future. It seems very likely to work with a reinforcement-learning paradigm, because you'll get a pretty clear signal from many tasks as to whether they succeeded.

I've been watching Payman a little bit recently. Are there any specific companies you would suggest people look at, or specific problems you think are most worth pitching to you?

Logan Kilpatrick

Evaluations are another one. Going back to this point, you made the comment that even someone I would classify as incredibly sophisticated in this space—someone who understands the ecosystem and why evaluations might be important—can't articulate what the taste is in a way that can happen programmatically.

I think that's really interesting. I don't know how that problem gets solved, but if you've spent too much time in the evaluation rabbit hole, as I have, one of the big realizations is that most problems in life end up being evaluation problems.

If you follow that chain of thought, it's very interesting how things end up happening.

Nathan Labenz

For now, I use expert demonstrations. At Waymark, we have the creative team write a bunch of good stuff, fine-tune on that, or put it into a few-shot prompt and hope for the best. From there, it becomes vibes, but it would be nice to have something better.

Where can people find you if they want to pitch you or if they have questions? You're incredible, as everybody knows, at responding to questions, concerns, and issues around the API, so I don't want to bring you more of that than you already have. Where should people find you if they want to point out an issue or pitch you a startup?

Logan Kilpatrick

I saw someone respond to 1 of my tweets the other day and say, “I miss the old Logan. He used to reply to all of his replies on Twitter.” I looked at how much time I spend on Twitter every week and realized that my time doesn't scale. I'm already putting in way too many hours.

But I'm on Twitter, LinkedIn, and everywhere the internet exists. Hopefully I'm there helping people with Gemini.

Nathan Labenz

Gemini 2.0 Flash is out today in general availability. It's good, fast, and cheap, so it's definitely 1 to check out.

Logan Kilpatrick from Google DeepMind, thank you again for being part of The Cognitive Revolution.

Logan Kilpatrick

Thanks for having me, Nathan. This was fun. Always a pleasure.