Speaker 2
Hey everyone, welcome to the Latent Space podcast. This is Allesio, partner and CTO at Desible, and I'm joined by my co-host Swix, founder of Small AI. Today, we have a returning guest as well as a new friend. Welcome, Michelle and Josh.
Speaker 1
Both of you work on post-training. Michelle, I think you used to introduce yourself as a manager on the API team. It seems like you've changed your role since we last talked on the podcast.
Michelle Pokrass
Now I lead a team on the research side, specifically in post-training.
Speaker 1
And Josh, you are also in post-training?
Josh McGrath
Yep. I'm a researcher on Michelle's team.
Speaker 1
I just found an interesting commonality you guys have: you're also both from Waterloo, continuing the tradition of extremely correct engineers.
Michelle Pokrass
Oh, yeah. We talked about that last time.
Speaker 1
That's right. We're gathering to talk about GPT-4.1. You launched it. We got a little preview, and it was a little bit rumored, right? It was pre-released, I guess, with OpenRouter as Quasar Alpha, and then it was also an Optimus version. I think people are trying to figure out why we're going back from 4.5 to 4.1. What are the headline facts you guys want to emphasize about 4.1?
Michelle Pokrass
We released 3 new models today: GPT-4.1, GPT-4.1 Mini, and GPT-4.1 Nano. The real focus was making models that were great for developers. We improved instruction following and coding, and shipped our first 1-million-token-context models.
Speaker 1
Josh, anything to add? Is there anything else that people should know that's sort of in the fine print?
Josh McGrath
I think the only thing I would touch on is that there's actually a new model in the lineup, Nano, which is even faster and cheaper for developers making low-latency applications.
Speaker 1
What's the fun story behind the code names? We got the Strawberry hat as another fun time in the lore of OpenAI.
Michelle Pokrass
We really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world, so we tested it through OpenRouter. It was super cool to see people latch onto the names and get the theories going, but the feedback we got from there was super helpful.
Josh McGrath
It's not even about the name. It's more about the API shape. Once we saw Chat Completions, it was very obviously OpenAI.
Speaker 1
That's a good note. But is there an emphasis on stars? What inference were we supposed to draw from “supermassive black holes”?
Michelle Pokrass
I don't think there's anything to draw from there. I think they're just cool—just fun names. They make you think of cool concepts.
Speaker 1
The vibes are good.
Josh McGrath
The vibes are good.
Speaker 1
The other thing about the examples—we're just mining for lore here—is tapirs. They're an interesting animal. They come up a few times on the livestream and in the blog post. What's up with tapirs? Who likes tapirs here?
Michelle Pokrass
Our team is just a super-big fan of tapirs. They just happen to work their way into a lot of our content.
Speaker 1
Okay, cool. Awesome.
Speaker 2
I think the first thing we want to run through is obviously the move from 4.1 to 4.5. That's the first thing that everybody was maybe confused about, and I know you're deprecating 4.5. It sounds like 4.1 is just a kickass model, and the 4.5 size maybe isn't as good of a fit. That was just a research preview. Whatever you want to say to address that, I think it's something we've seen come up in the Discord as well.
Michelle Pokrass
Naming is really hard, and we've tried to make this as unconfusing as we can, but nothing's perfect. Basically, the way we got here is that GPT-4.1 is a pretty big improvement over the 4o line, and we really wanted to signify that. However, it's a model that's much smaller and cheaper than GPT-4.5. As a result, it doesn't achieve the same AIME or other intelligence-evaluation results, so it doesn't beat 4.5 on all of the evals. We didn't think it made sense to increment beyond 4.5.
For most developers, though, they can replace a lot of their 4.5 usage with 4.1, and the Mini is strictly better than 4o Mini.
Speaker 1
With Nano, we don't know if 4.1 is a distillation of 4.5, or whether there's a relationship there. What can we say about the shared lineage?
Michelle Pokrass
We're always using various research techniques to improve our models, and distillation is something we've talked about before. It's especially meaningful for the small models. We've pulled out some of the things that made 4.5 really good—it has a lot of instruction-following strengths—and rolled those into 4.1.
Speaker 2
I strongly remember that, at the GPT-4 launch, the communication was that we're moving to a new model architecture that is an omni model. That's what the “o” in 4o means. GPT-4.1 is part of this subsequent trend of trying to merge everything—the reasoning model, the omni model, everything.
There's some doubt about whether 4.1 is being sold as a strict replacement for 4o. Is it going to be fully omni-modal? Is it roughly the same architecture that we think 4o has?
Michelle Pokrass
We already have different slugs on the Realtime API and the Responses API, so they're already somewhat different checkpoints. We don't currently have plans to release 4.1 in the Realtime API, but things may change.
Speaker 1
And then there's image generation and all that, right? As far as we know, there are no plans—maybe nothing announced?
Michelle Pokrass
Not right now. The focus for 4.1 was these 3 core capabilities for developers.
Speaker 2
Our Discord also did a launch watch party for the recent 4.5 podcast that Sam Altman did. I think, for the first time, it was basically confirmed—something that people already knew, since Andrej Karpathy was already talking about it—that 4.5 was 10x the size of 4.
I think there's a question about whether we do a linear interpolation for 4.1. Is 0.1, I don't know, 2x the size or something?
Michelle Pokrass
That's not really how we think about naming the models. There are a whole bunch of different parts that go into the recipe, so our version numbering doesn't really reflect just the pre-training recipe. GPT-4.1 is named that way because of the large jump in coding capabilities, long context, and so on. It's more about what it's like for the end user than anything about the training recipe.
Josh McGrath
We can go a little under the hood on training, though. Nano is obviously a new pre-training run. We also have a new pre-training run for Mini, and then the larger version is a new mid-training run.
We find that a significant amount of the gains actually come from new post-training techniques. In the past, the narrative was that you needed to pre-train these larger and larger models to get better performance. We're finding that we're able to squeeze a lot more out of post-training now.
Speaker 1
The other side of how big a model is is the context window. You have a 1-million-token context window. Sam said that day last year that 1 million was months away, so you're right on time. Can you talk about how hard it was to get to 1 million, and then maybe where the end game is in your mind? Is it 10 million, 100 million, or infinite? What really matters as you start to scale this?
Michelle Pokrass
Josh worked a lot on long context, so he's the right person to ask.
Josh McGrath
The first thing I thought was really interesting when we were working on long context was that some of the evals you see as headlines on other blogs—needle in a haystack—are actually things that most models do really well right out of the box. A single needle in a haystack was easy to saturate.
We first had to get a lot of measurement on longer-context, long-context reasoning. We just open-sourced 2 new evaluations that are about using the context in a more complex way. In one of them, you have to reason a lot about ordering, and the other involves walking through graphs.
There's a lot of reasoning that you have to do in those data sets, and that's where long-context work is actually much harder. But single-needle-in-a-haystack tasks were pretty easy to saturate.
Josh McGrath
Yeah, I think the mental model that I have is maybe actually has some more variables in it. So there's single, there's the needle in a haystack where you have some amount of distractors and some needles that you're trying to find. And I think that it's more so about how dense of the context do you need to use. So like summarization, you're actually just using the entirety of the context, whereas needle in a haystack it's very sparse. And then I also generally think about orderedness. If you're going to make some sort of inference on this, are you just looking sort of front to back, or do you need to move around in the context in order to generate a good answer while the model is sampling?
Speaker 2
Yeah. Is that something that you worked on with Graphwalks? Is that the thing?
Michelle Pokrass
Yeah, that was sort of the most synthetic and clean way to measure the model. Then we worked on a lot of other training techniques and data to sort of test and train the model's ability to reason throughout the context in a sort of shuffled way.
Josh McGrath
Yeah. I actually like to give people a little bit of visual aid with these things. I went into your Hugging Face release and got an example of the graph task. There are a few versions of this, right? There's the BFS and DFS version, and it's also, I guess, very character-specific. Could you tell us about the design choices around this, what was surprisingly hard, or anything like that?
Michelle Pokrass
Yeah. The idea here is that you take a graph and encode it into the context by looking at the edge lists and just putting that into the context, then asking the model to do an operation. Under the hood, we're actually just executing the real operation and using that to evaluate the model's ability to work.
One of the things that I found surprising at first was what the model would do when it wasn't sure how to use its context. Early versions of the model would just loop, saying, “Oh no, I can't find this edge that I think needs to be there.” I was actually very surprised that all models seemed to have more difficulty than I would have expected on a task that we would find very simple, or that maybe an undergrad could write a Python script to run in a couple of minutes.
Josh McGrath
Okay, what is the real-life task that this is meant to model? I feel like the other one, MRCR, seems a little more intuitive, where you have 4 different stories and pick out the 2nd one. That's a real task that people have, but people don't really traverse graphs like this. This is a bit more theoretical. Was there any sort of correlation study done?
Michelle Pokrass
Yeah. This is actually meant to be the idealized version of a multihop reasoning benchmark. We have a lot of things where you're putting 100s of documents into the context, and then you might ask a question that you actually have to traverse 10 documents to answer. But there, the edges are implicit, right? There is some underlying graph connecting all of these documents that you need to traverse in order to answer the question.
The question was, if I just give you all of the IDs of the things that you need to traverse, can the model even do that? It's actually just a lower bound on how well the model can do, and I think that's somewhat well reflected in some of the internal benchmarks we have that are using more natural data. You can imagine something like a tax return, where you upload the entire tax code and, to figure out what to put into this box, you'll need to reference all of these boxes. This is a similar level of multihop reasoning, but again, like Josh said, all of the references are implicit.
Speaker 2
Yeah, I think that some kind of backtracking, if it's needed, is also super interesting, especially for agent work. For listeners who've been listening to us for a while, we actually covered this paper in NeurIPS, where they modeled graphs for graph traversals for agent planning, and it reminds me closely of that. It's just that they never came up with this exact format that you have here, which is basically the same thing.
I also like that you included blank answers, because sometimes people—or models—do hallucinate answers, and you have a fair amount of blank ones.
Michelle Pokrass
Yeah, I think that's thanks to the random sampling over graphs I did, I guess.
Speaker 2
Is this tied also to the File Search API that you released recently? How should people think about how everything comes together in the API?
Michelle Pokrass
Yeah. Oftentimes with retrieval, you might be using RAG to fill the context, and a lot of this is to get around the limitation of a short context window. We do expect a lot of developers to start uploading their full context more directly to the model, so for smaller tasks, maybe you don't need the whole vector store.
We do anticipate this to play well with that paradigm as well. Maybe you can just insert way more chunks into the context, so we think it'll play nicely.
Speaker 2
Any relationship to the memory upgrades in ChatGPT that we recently got? Is long context just directly usable for memory, or should we always have a separate memory system?
Michelle Pokrass
Yeah, it's a good question. Right now, the dreaming feature has some of these memories embedded in the context, but they are separate features. GPT-4.1 is powering the API, whereas the enhanced memory is ChatGPT only.
Speaker 2
Yeah, I think that's interesting. I guess the 1 last thing I'll call out on long context, which is kind of unintuitive—or maybe there's an explanation—is that you had 2 needles for MRCR, and then we had 4 and 8, and everything kind of just regresses to some kind of baseline of, let's say, 30% or 20%. But it's interesting to see where the smaller models sometimes match or outperform the larger models. I was wondering if there's anything unusual there, or do you think it was just a bad roll of the dice?
Michelle Pokrass
I think it's probably just a bad roll of the dice. I would probably look more at how these things regress as you increase the number of needles, because there's sort of more complex reasoning that has to do with the order of different things in its context.
Speaker 2
Awesome. Cool. Happy to move on from there. We have a whole bunch of other evals that we can go over, so I had in my notes that we could talk over anything that you want. There was also, like, Collie[?] from Shyu[?], whom we had on the podcast for instruction following. I realized that he joined OpenAI, and I wondered if he had a role to play in that one.
Michelle Pokrass
No, we did not collaborate on it. Honestly, I think it's best when eval authors and model developers don't collaborate too much, because you want things as objective as possible without trying to game any evals.
Speaker 2
And then I think there was also, for the first time, the announcement—or shout-out—of the internal instruction-following benchmark from API data. People have had the ability to opt in to share data for a while. I posted a tweet because I found it in the dashboard: You can just opt in, and there are basically 16 days left for this program where you can just get free inference. I'm curious what you found from that kind of eval that might be different from the normal instruction-following evals that people have.
Michelle Pokrass
Yeah, totally. A lot of the instruction-following evals that are open-sourced are crafted in a way that makes them easy to verify. For example, Graph Walks is somewhat easy to craft: You can create this graph and verify it easily, but it is not exactly aligned with what users are doing. This is true for some of the instruction-following evals where you ask the model to output exactly 4 words or 3 paragraphs—things that you can verify easily in code.
These are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade, and so the open-source evals often don't have them. Getting this real-world, diverse set of data helps us find the commonalities in what developers are doing, what is a really good example of a negative instruction, and then we can go from there and figure out how to evaluate it.
Josh McGrath
Yeah, I think there's also an interesting question of what domains people use you in. I wonder if there's a way to tell you, because sometimes it can be very confusing—especially if I'm building an app and letting people use my key, but other people are building apps on top of me. You just have a lot of chaos from multiple degrees of abstraction, where you have to parse through the prompts.
Michelle Pokrass
Yeah, it's true. I will say that we do use our own products internally where we can, so we're not manually reading every prompt. After they're anonymized, we scrub them of any identifying data, and then we use our models to make passes over them and categorize them.
If we get feedback that we're not doing well on ordered instructions, then we can do a pass over all of our data and find some good examples of those.
Speaker 2
So there's an instruction-following section in this great guide to prompting GPT-4.1 models. I think maybe we can go through some of these examples. The first one that caught my mind was: It's not necessary to use all caps and other incentives, like bribes or tips, but developers can experiment with this for extra emphasis.
That second part leaves me confused. Are you saying that people should still try to do this, and sometimes the model responds positively to it? Do you feel like it's still just part of the lore? I'm curious why—I would have loved for you to say either, “Yes, it works,” or, “No, you should stop.”
Speaker 2
It looks silly.
Michelle Pokrass
I guess the truth is somewhere in the middle. The truth is always messy. The reality is that our models have gotten a lot better at following instructions that are stated once and clearly. But we find, honestly, that developers often become the best experts at prompting our models because they’re building their livelihood on this thing and get to know its details really intimately. So I will say that stuff like that won’t hurt the model’s performance, but we always want to leave it open to people to figure out what works best.
Speaker 2
Yep. And then you had to always start with a response rules or instructions section. Are those keywords meant to be taken kind of verbatim? Are those the tokens that work the best, or is it just an example?
Michelle Pokrass
More of an example.
Speaker 2
Yeah. Okay, cool. This is great. I feel like, until today, we did an episode with the Prompt Report on all these prompting techniques, but it’s also unclear which ones work best for which model. So it’s super useful.
And then in the agentic workflows one, you have a persistence thing: “Yeah, please keep going.” How much? I think I read that it improves SWE-bench by like 20% just by having the persistence.
Michelle Pokrass
I wouldn’t say that this one prompt improves SWE-bench by 20%. It’s that we found this is the most effective harness for our model, and combined with all the post-training improvements, it results in the big improvement.
But yeah, the model is trying hard to be helpful, and often it wants to check back in with the user and say, “Should I keep doing this? Am I on the right track?” A prompt like this makes sure it keeps going, doesn’t bother you again, and just gets the test done.
Josh McGrath
Yeah. I think there’s this interesting trade-off between persistence and yielding back to the user. The more agentic a model wants to be, the more persistent it should be, but then sometimes it just goes off the rails.
There have been criticisms of Claude Sonnet trying to rewrite too many files at once when I just wanted to make one thing, for example. That’s a form of bad persistence. I wonder how you solve this trade-off, because sometimes it just goes too far. What are the axes here in which you think about it?
Michelle Pokrass
I think one interesting thing that comes to mind here is that we had an extraneous edits eval, where you ask the model to make an edit and classify whether all its changes were related to what it was asked to do, or whether it went off and did a little too much.
We found that GPT-4o got 9%, which is pretty crazy. Making an extraneous edit 9% of the time is a lot. GPT-4.1 is at 2%, so it’s a pretty big improvement.
I’ll just say that, focusing on this, we’ve heard feedback about it, made an eval, and made sure to track it and improve it during training, too.
Speaker 2
Yeah. I mean, everything comes down to evals, as is no surprise to anybody. There’s another interesting eval that I think is causing some noise. You, being the master of structured output, should know that JSON is bad now and we should all use XML.
I want to say that I don’t know which eval you’re talking about, but it’s in the prompting guide, which maybe you guys didn’t write, so we’re kind of springing this on you.
Michelle Pokrass
Noah and Julian on our team wrote the prompting guide and did a great job. I do think XML is very helpful for structuring prompts, whereas for parsing outputs, maybe the story is a bit different. Sometimes it’s really useful to get outputs in JSON so you can plug them directly into your application, but I do think the models work particularly well with XML as inputs.
But curious—anything to add?
Josh McGrath
No. Cool. I mean, I think people always care a lot about code tool calls and structured outputs, as you all know, so any updates to instructions over there are good.
People are also interested in this concept that apparently putting the instructions and user query at the top and the bottom in the context—duplicating them at the top and bottom—is much better. That is better than putting them at the top only and much better than putting them at the bottom only. Again, this is from the prompting guide, so I don’t know how aware you guys are of this.
Michelle Pokrass
I think part of that was just empirical. We tried all 3 when we were evaluating the model, and having that redundancy is definitely the best. But then, using those instructions at the beginning, the model’s going to be able to take that into account as it does processing.
Speaker 2
I think a lot of people would see this as running counter to prompt caching, because obviously you want to put the things that change a lot at the bottom. Basically, is this fixable in post-training? Can we just tell models to take instructions or user queries only at the bottom because we want to optimize for prompt caching?
Michelle Pokrass
When we figure it out, we will do that. I mean, it seems doable. Well, it seems like a post-training thing. I don’t know—maybe my mental model of post-training is wrong.
I think actually having things at the beginning of the prompt would still get you prompt caching there. If you’re putting in, for example, a big needle-in-a-haystack prompt and you have the data changing each time, like for each user, there are still different ways that you can put the prompt at the beginning and get a lot of cache hits. It sort of just depends on your use case.
Josh McGrath
Awesome. The other thing I noticed—I know you made a note of this, too—is the chain of thought and reasoning, and how people should think about this model versus a reasoning model. Should I just use GPT-4.1 and prompt it to do chain of thought? Should I use o1 and make a plan, and then use GPT-4.1 to implement the plan? How should people think about composability?
Michelle Pokrass
It’s a great question. We have found that GPT-4.1 is a lot better at doing planning and thinking through its steps in CoT when prompted than our previous non-reasoning models. But our reasoning models are designed to have more coherent plans and be able to reason over longer horizons than these non-reasoning models.
You can see that reflected in intelligence benchmarks. AIME, GPQA, and stuff like that—you’ll see the reasoning models do much better. In general, I would say the question you’re really getting at is: “I’m a developer. Which model should I be using?”
I think the answer is always going to be the fastest model that accomplishes your task. Maybe you start prompting GPT-4.1. If it does your task super well, then maybe drop down to GPT-4.1 mini and save latency, or even GPT-4.1 nano. Whereas if GPT-4.1 is struggling a little bit and needs more coherent reasoning over longer time horizons, then maybe you upgrade to a reasoning model.
Speaker 2
Is there a quick way to get through these heuristics? I know one thing that a lot of people do is use o1 for a plan, put that plan in Cursor, and then have the plan applied to their codebase. It sounds like there’s maybe not a rule for when to do which; it’s just task-dependent.
Michelle Pokrass
Yeah, I would say we’re all kind of figuring out the best way to use these models together. I do think reasoning models for planning and using more targeted models to execute is definitely a good architecture.
Speaker 2
Cool. If there’s nothing else on that side, I’d love to go into coding, which is something that we’re emphasizing a lot. It’s doing super well. It’s better than o1 on SWE-bench. Was that expected?
Michelle Pokrass
Not really.
Josh McGrath
Yeah, what’s the story there? There’s also SWE-Lancer, which is a newer one that attaches a money value to things. What should people understand is going on here? Is it a better coding base model or just a coding-agent model?
And I think there’s also a question about how important coding is if I’m not using a coding use case.
Michelle Pokrass
I’ll start by saying we just set out to make a model that was great at coding, both in your terminal or in your editor or wherever you want to use it. So we kind of broke that down into the problems that it encompasses.
Developers want the model to produce better diffs, for example, or they want the model to explore the codebase correctly, produce code that compiles, or produce code that writes tests. Our approach was teaching the model all of these various facets. There’s just a bunch of work streams that all coalesced around GPT-4.1.
Josh McGrath
Yeah, I think much-improved post-training all over makes for a better coding model. I think there are different kinds of coding, right? It’s interesting for me to observe that. I’m just going to pull it up on the chart here because I always like to show people visuals.
You’re at 55 on SWE-bench and o1 gets like a 41, but on Aider, it is less—it is not at that level. I struggle to get some kind of intuition of when this applies. What are the different elements of coding? I guess there are single-file edits, whether it’s a diff or a whole file, and then there are entire-project edits. Is that a reasonable split? Are there more to this?
Michelle Pokrass
Yeah, that’s one way to think about it. Basically, GPT-4.1 can explore and go through a repo; it’s been trained to do that particularly well. Whereas, to just get some code and produce a change, a reasoning model might do better because it can reason over the entire file.
And so that's one good way to think about it.
Josh McGrath
Yeah, that's fair. Do you have any understanding of the smaller models? For coding, should I only use GPT-4.1 and forget the rest?
Michelle Pokrass
You might want to use the smaller models if you have an IDE where you need an autocomplete feature, for example. Or if you want something super fast—if you're building, I don't know, a text-to-SQL application, you might want the first version to populate instantly. You can see that GPT-4.1 mini is actually quite significantly better than GPT-4o mini, and not that far away from the old GPT-4o. I do think that model will find use cases in a bunch of these coding niches.
Speaker 2
I know you might not be able to talk about this, but the clip of an AI CFO talking about agentics has been going viral, I think, today. It seems like every lab is putting a lot of emphasis on coding, so I'm curious if there's anything you can share about how people should think about OpenAI in coding. Obviously, today you don't have Claude Code. You don't have anything related to coding, and I think the Windsurf partnership today—they're giving GPT-4.1 for free for a couple of weeks—is maybe one of the first OpenAI endorsements, I guess, on the livestream. I know there might not be an answer the PR team would approve, but I'm curious if you have any takes or thoughts.
Michelle Pokrass
I think just stay tuned. Coding is an important use case for our users, and that's why we focused on it a lot for GPT-4.1. We also love to use our own products internally, so making GPT-4.1 better selfishly helps us move faster as a company. That's where the real focus has been for this model.
Speaker 2
Do you track what percentage of code is written by GPT-4.1 internally?
Michelle Pokrass
We do have some metrics like that. I don't have them off the top of my head, but I was actually just talking to one of the researchers on the team who worked on something over the weekend. He said that GPT-4.1 was able to get 49 out of 50 of his commits on this massive PR done, so we were pretty happy to hear that.
Speaker 2
I'm excited to use it. I think coding is a super exciting use case, and OpenAI has always been very developer-first, as you've seen, Michelle. It's great to see the convergence.
The other capability that I kind of zeroed in on was vision, or just multimodality in general. It is a lot better. I really like these niche benchmarks, like Math Vista and Chartive. Is there any extra color on the vision side that you wanted to talk about but maybe couldn't fit into the blog post?
Michelle Pokrass
One small nugget there is that GPT-4.1 mini is really exciting on that front. As we were talking about, it's a different pretraining base, and I think that really shows up in some of the vision evaluations. We talked about coding, instruction following, and long context—a lot of gains coming from post-training—but in particular, for multimodal capabilities, basically everything you're seeing is gains from pretraining. Kudos to the pretraining teams; they've done incredible work on perception and multimodality.
Speaker 2
Something that we've been exploring for a while, and I'm curious if there are any takes on your side, is whether there's a strong split between what I call screen vision and embodied vision. Are you taking pictures of, or training on, snapshots of a computer for computer use? Anything with charts or a PDF is very similar to that, whereas pictures from the real world are more embodied—something a robot might be able to use. People have argued back and forth, so I'm curious where the movement or emphasis is.
Michelle Pokrass
First off, I think GPT-4.1 is better at both of those things. Regardless of how it was actually trained, I would probably defer somewhat to the pretraining team when it comes to which one you should be using. We're using a mixture of both, but we've improved our results across evaluations.
Speaker 2
That's something that people should definitely explore—the more embodied stuff as well—because the benchmarks tend to focus on the screen-vision stuff, more chat evaluations that are easy to grade.
Michelle Pokrass
Yeah, exactly. Those are the things that get looked at the most, for sure. One of the things that was really funny with both GPT-4.1 mini and nano is that we had some strange internal evaluation results. It turns out that these new vision capabilities were able to read signs in the background and stuff, which was actually changing the validity of our results. We were running into different evaluation problems as we improved the models.
Speaker 2
Is there a GPT-4.1 image-generation feature, or is that a completely different part of vision? In some sense, vision is image-to-text, and the other way around is image generation. Is it that simple, or is it something else?
Michelle Pokrass
There is no place right now to get GPT-4.1 image generation.
Josh McGrath
Okay. Well, it's very popular. It's like melting your GPUs. Part of this whole deprecation of GPT-4.5 and moving people to GPT-4.1 is to get back your GPUs. That's a message that both Shuky and Kevin Weil have mentioned. But you're running all these models concurrently for the next 3 months, so I don't know if you get back those GPUs and then just grow their usage even more.
Michelle Pokrass
I do think that people get the message on deprecation and start moving over. As developers use this model a little less, we can reclaim that compute. You're right that it takes a while, and the trade-off there is really our commitment to developers: if we have something in the API, we won't take it away without sufficient notice.
Josh McGrath
Okay, awesome. Then a couple of other smaller announcements: fine-tuning is available on day 1, which is new for OpenAI. Usually, you have to wait like a month or 2 for the fine-tuning capability—for GPT-4.1 only and GPT-4.1 mini only, with nano in the future? Any specific callouts for fine-tuning? Fine-tuning is a general discipline that always applies, but are there any wins that you can talk about?
Michelle Pokrass
First off, shout-out to the fine-tuning team; they've worked really hard to get this ready on day 1. One thing I will say is that I think people have slept on the preference fine-tuning offering, or whatever we call the product. SFT is pretty well known; it's the original fine-tuning we had. Preference fine-tuning is super helpful for steering in a particular style, so I think not enough people are using that.
Speaker 2
Isn't that only for reasoning models, or is that for everything?
Michelle Pokrass
No, reinforcement fine-tuning is only for reasoning models, right? Preference fine-tuning offers the pairs.
Speaker 2
Exactly. I thought it was in alpha, which is why I haven't looked into it. I thought RFT was still in alpha.
Michelle Pokrass
That's a lot of confusion that we just cleared up.
Speaker 2
I'm doing my conference again in June, and I think we're going to do a workshop on all the general fine-tuning options. I think that will clear up a lot of things, which is good. I know we can't talk about many of the new models. Noam Brown from your reasoning team just said that there should be a follow-up on reasoning models soon. What can we say about that?
Michelle Pokrass
Sounds like we're not the right people to ask, but stay tuned.
Speaker 2
GPT-4.1 is a good basis for whatever comes next, right?
Michelle Pokrass
Not all of our models necessarily build on each other, but we think GPT-4.1 is a great standalone offering for developers. We also think reasoning models are a good tool in the toolbox.
Speaker 2
More generally, I always want to explore the relationship between non-reasoners and reasoners, and also how we merge them. Are we doing routing? You obviously have a lot of secret sauce. The other thing that a lot of people are demanding or asking about is the creative-writing model. Will that ever see the light of day?
Michelle Pokrass
We're working on incorporating those improvements into the models more generally, rather than doing a separate release. People loved the humor, the green text, and the nuance of GPT-4.5, so we've heard that feedback. I know there are lots of folks working on that and trying to bring it into our next models.
Speaker 2
Awesome. Alessio, anything else?
Speaker 1
No, this was great. Any requests for the developer community? Are there things you want them to try out that maybe people are not doing, or things you want them to build for you using the new APIs?
Michelle Pokrass
First off, send us feedback. It was really useful to look at different partners and customers who are using our models and get this nice, rapid feedback from them. It allows us to iterate a lot faster. On that vein, opt in to data sharing; this just helps us make the model better for you. One kind of slept-on way to do this is the evals product. You can upload an eval such that we'll pay for the inference costs if we can also use the eval. This is just another great way for us to use those evals to make sure our models are getting better for people over time.
Speaker 0
Yeah, I think the evals are permanent. There's no end date announced, but the opt-in for the API is at least until April 30. I think a lot of people still don't know about it. We might want to extend that so that people can do more.
Speaker 1
Good flag. I'll raise it with the team.
Speaker 0
Yeah. Awesome. I think the last question I had was just on pricing. I think pricing is generally cheaper than GPT-4o—not by a ton, but it is cheaper. Then you're also introducing this concept of blended pricing for the first time that I've seen. Maybe it's just been out there for a while because you have caching and all that. Generally, what is the cached-to-uncached ratio that we should be thinking about for workloads? Is there a general rule of thumb?
Michelle Pokrass
One clarification: GPT-4.1 mini is not cheaper than GPT-4o. It's not just a blanket decrease in all the models. However, 4.1 mini is cheaper than 4.1. I'm also not sure if this is widely reported, but we've increased our prompt-caching discount from 50% to 75% on these models.
Speaker 0
Yeah, I saw that. So that's a big input into figuring out what kind of application you build. And then your question was about blended pricing, right? I think there's this question of comparability of prices across models and across providers, because some people are 3:1 in terms of context to output, and some part of that is cached. I selfishly make a chart that just plots all the model labs versus all the prices, and I'm sure you guys have seen it. I don't know what numbers to plug in there. So what are people seeing in real life? What's the median caching rate?
Michelle Pokrass
I don't think we have that off the top. The blended pricing is more to just make it easier to compare. You could say something like, “GPT-4.1 is 25% cheaper than GPT-4.”
Speaker 0
Yeah. You want 1 number.
Speaker 1
Yeah. Yeah. Yeah. No, all right, we'll all have to figure it out.
Speaker 0
But thank you so much. That was fantastic. Thanks for all the work. I think people are very excited to get to work testing this out and giving you feedback. I'm sure we'll be back again for the next one—probably the reasoner.
Speaker 1
Nice. Thank you guys.
Speaker 0
Thank you.