[BidClub_]
Latent Space · · 39 min

Inside OpenAI DevDay: Superhuman Computer Use, Decisions API, and the AI Cloud — Ari & Nikunj

swyxVibhuAri WeinsteinNikunj Handa

AI & SoftwareTechnicalCompany Building
YouTube ↗
TL;DR
  • OpenAI argues computer use has crossed from brittle demos into useful labor, responding to a claim from a leading AI podcast that it had not advanced in two years. Ari Weinstein says agents are now “faster at accomplishing tasks than the average human probably in most cases,” in part because they can debug, retry, and introspect instead of merely starting tasks successfully.
  • The economic unlock is improving faster than headline model pricing suggests. GPT-6.1 Soul was presented as one-fifth the cost of Astra overall and one-seventh the cost for computer use specifically. The host also cited a keynote claim of a 7× improvement in computer-use speed. That shifts the investable bottleneck from raw intelligence toward inference, harness design, and software response times.
  • Dots turns computer use into general-purpose delegation by giving each Dot its own Linux cloud computer. The addressable surface is effectively any human-facing software, including services with poor APIs: Ari’s meal order fell from two hours manually to 15 minutes, while the hosts highlighted flight booking, customer support, YouTube operations, games, and shopping.
  • OpenAI’s Decisions API responds to Jev and shows how quickly a new AI product category can be absorbed by a frontier lab. Nikunj Handa said it was “not at all” on the table four weeks earlier; user requests and internal teams prompted inference and infrastructure engineers to build a prototype, after which the project moved into latency hill-climbing. The first version uses Luna’s existing weights rather than a new model.
  • Computer use is increasingly a multimodal systems problem, not a screenshot-clicking benchmark. Agents can combine screenshots, accessibility data, Playwright, direct DOM access, and generated JavaScript that performs multiple actions at once; AppShot similarly exposes underlying metadata rather than pixels alone. “A lot of little paper cuts” now matter as much as isolated model gains.
  • The Agents API creates a distribution advantage because OpenAI trains models on the same computer-use harness developers can use. Ari says that may improve “speed and cost and accuracy,” while autonomous testing closes the software-development loop: an agent can build and test an application before a human becomes QA.
  • The platform roadmap is converging on an AI-native cloud built around asynchronous execution, caching, and context management. WebSockets enable mid-turn steering and async tool calls; cache guarantees are available within 30 minutes, while a 12-hour guarantee is in preview and not yet public. Pre-warming and cheaper reads are also being developed. The unresolved question is how much memory, storage, and agent structure OpenAI should abstract versus leave flexible.
Digest · the substance, structured for research

1. Computer use has advanced 180 degrees, not stood still

  • The hosts invoke a claim from a leading AI podcast that computer use had not advanced in two years; Ari’s blunt rebuttal is that it is now “180 degrees different.” The meaningful comparison is less with last week’s release than with the cumulative jump over the past one or two months.

  • A year ago, models could “reliably start tasks” but would derail when something unexpected happened. Ari identifies the decisive capability gain as debugging: current models try again, inspect what failed, and alter their approach, allowing them to complete longer workflows rather than merely demonstrate a plausible first action.

  • The keynote’s stack included Dots, GPT-6.1 Soul, computer use inside the Agents API, AppShot, native Mac control, and the Decisions API. GPT-6.1 Soul was described as one-fifth the cost of Astra generally and one-seventh the cost on computer-use workloads.

2. The harness now sees structure and writes actions in batches

  • Ari’s framing: modern computer use does not depend on a single interface. Depending on the task, the model may combine screenshots, accessibility representations, Playwright, and direct DOM access, then write JavaScript that executes many operations instead of clicking through one action at a time.

  • The concrete breakthrough is eliminating repeated screenshot-scroll cycles. When an agent can inspect an entire page or application structurally, it can locate content beyond the viewport and manipulate multiple elements at once; Ari calls added modalities one of the field’s biggest “aha moments.”

  • AppShot illustrates the distinction. Double-tapping the Command keys may appear to attach a screenshot, but it also supplies token-efficient accessibility metadata: link destinations that pixels cannot reveal, full calendar titles hidden by truncation, and the raw text representation available behind the attachment.

  • Ari resists attributing gains to one magic change: production products use different safety configurations, while improvements arrive from models, inference, representations, and harness engineering. The speed curve is increasingly moved by “a lot of little paper cuts” found through detailed introspection.

3. Dots makes every piece of human software addressable

  • Each Dot has its own Linux virtual computer in the cloud, capable of running desktop applications as well as a browser. Ari’s core thesis is universality: because “all the software in the world was designed for humans,” an agent that can operate human interfaces can accept nearly any computer-bound task.

  • His sharpest personal example is a customizable meal-prep service. Specifying exact grams of chicken and rice took him two hours; computer use completed the same order in 15 minutes—“eight times faster”—on GPT-6.1 Soul.

  • The host extends that logic to surfaces neglected by APIs. YouTube workflows such as A/B testing and community posts still require UI operation, while the host described customer-service workflows that can wait, research with subagents, supply reference numbers, and continue negotiating without consuming the user’s attention.

  • This lowers what Ari calls the “activation energy” of delegation: once agents become fast enough, users may default to assigning work they currently do manually. What proves valuable will still depend heavily on the individual user rather than a universal killer workflow.

4. Superhuman speed now depends partly on the software being operated

  • Ari says computer use is already “faster at accomplishing tasks than the average human probably in most cases.” The next frontier is narrower: become as fast or faster than expert computer users, enabling genuinely real-time products rather than merely asynchronous task completion.

  • As inference accelerates, external applications become the bottleneck. A meaningful share of a DoorDash benchmark can be spent waiting for DoorDash itself to load; the system must then minimize the delay between completion of that load and the model’s next action.

  • Event-driven triggers solve some waits, such as browser navigation events, but not every state transition can expose a clean event. Customer-support replies that arrive anywhere from 30 seconds to three minutes exemplify the scheduling problem still hiding beneath apparently simple automation.

5. Trust and autonomous QA determine whether universality compounds

  • The host describes how quickly his own trust frontier moved: three or four years ago, connecting LLMs to devices and the web felt dangerous; now he lets Codex configure DNS, pay bills, and handle “tens of thousands of dollars” of activity while “YOLOing with computer use.”

  • Ari’s counterweight is progressive trust, not unrestricted autonomy. Reliable systems should request consent before consequential actions such as payments and, depending on the application, restrict access to only the websites and applications needed for the assigned task.

  • For developers, OpenAI’s computer-use harness offers more than convenience because the models train against it. Ari says staying in-distribution “might be a speed and cost and accuracy advantage,” though developers can still build their own harnesses when they need specialized control.

  • His favorite coding workflow lets an agent test what it built. Computer use closes the lifecycle from implementation through testing and inspection, removing the pattern where “you are now QA for the agent”; Ari sometimes even has one computer-use agent test another.

6. Jev inspired a category, and OpenAI responded with Luna

  • Nikunj gives Jev explicit credit for “inspiring a whole segment in the market.” He says that four weeks earlier the Decisions API was not on the table; user requests and internal teams prompted inference and infrastructure engineers to build a prototype, after which the project moved into latency hill-climbing.

  • The initial Decisions API is not a newly trained model. It constrains the existing Luna weights through structured outputs, runs multiple questions in parallel as a batch, and optimizes the inference stack for fast time to first token and time between tokens; Luna’s vision capability comes along “for free.”

  • Its clearest use is fast classification, including support-ticket classification. Low-latency LLM-as-a-judge evaluation and computer-use or GPT Live demos were also discussed, but Nikunj preserves the limitation: a non-reasoning Luna choosing one action at a time will not match Astra writing a JavaScript program for sophisticated, long-horizon computer use.

  • Vibhu frames an unresolved differentiator as calibrated confidence, arguing that RLHF can push a model toward what users want to hear without necessarily preserving the right confidence level. Nikunj agrees calibration may be an area OpenAI must “hill climb” through future model releases.

7. OpenAI is assembling the primitives of an AI-native cloud

  • GPT-6’s API additions include asynchronous function calling and mid-turn steering: a model can launch a slow tool, keep reasoning, and accept new instructions when results arrive. WebSockets supply the bidirectional channel, reducing tool-call overhead for agentic execution and Ultrafast inference.

  • Nikunj credits inference work first with cutting Luna’s price by roughly 80%, then with pushing Astra toward “frontier intelligence at extreme speeds.” GPT Live can act as a fast talker and delegator with Astra behind it, while Decisions API and Luna make rapid tool-calling demos feel snappier.

  • Responses API performance now extends beyond token generation. OpenAI guarantees cache hits within 30 minutes, is previewing a 12-hour guarantee for one user, and supports pre-warming so developers can pay the cache-write cost before traffic arrives and reuse the prompt across multiple thread instances.

  • Context still requires compression despite caching. Agents API includes proprietary compaction; Responses API offers threshold-based server compaction or manual /compact, while file-based techniques are implemented and being experimented with in the Codex harness. Nikunj’s open design question is whether future memory vaults and storage concepts should become high-level primitives or remain examples developers assemble themselves.

Full transcript
Ari Weinstein

Computer use is now faster at accomplishing tasks than the average human, probably in most cases. I think the next frontier is to have computer use be literally superhuman in its performance, where it’s as fast or faster at using software than expert computer users like us.

I think that’ll be really consequential and exciting when it happens, because we’ll be able to build products that provide much more real-time experiences.

swyx

4 weeks ago, this was not at all the case. No, this is Jevons-inspired. I think you’re officially the first frontier lab to clone and adopt this.

Ari Weinstein

Yeah. I feel like OpenAI has such a strong hacker culture, and people just get excited about things. A guy from inference and an awesome guy from the infrastructure team were like, “This is amazing. We’re going to hack on it.” They built a prototype, it worked, and now we’re just hill-climbing on latency and trying to make this as fast as possible.

We want to launch it in the coming days, so as soon as we hit our latency target, we’ll try to get this out.

swyx

1. OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API

Okay, we’re very excited to be here. Today is the OpenAI Dev Day special Aftershow podcast. We have Ari here, who leads the product and engineering team for computer-use agents. Before we kick in and dive deep on computer use, do you want to give a quick recap of what was announced? What was the quick slew of announcements you guys had today?

Ari Weinstein

Yeah, it was a super exciting day. We just got out of the keynote. It was really sick. There were a bunch of computer-use announcements that I think are worth thinking about.

We have Dots, which is the new personal assistant product, and that has some really exciting computer-use features. There’s GPT-6.1 Soul, which is an amazing new model that I think is particularly great for computer use because of its cost and speed advantages. I think we shared that it’s a fifth the cost of Astra and a seventh the cost if you’re looking at computer use specifically, which is really amazing. Sorry, there were so many things—I’m trying to sort through it.

swyx

And the Agents API, which now has computer use in it, is really cool because developers can build on the same computer use that is part of Codex and ChatGPT. Then there were demos of our existing computer-use features, like AppShot, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast.

There was also native computer use on your Mac, where Romain had it taking screenshots of his app automatically, and he could do other things on his computer while computer use was using his applications. So, yeah, it was a really exciting keynote.

And not to mention the Decisions API.

Are they all the same model, or the same dataset distilled to different models? Basically, is computer use in the Decisions API, or are they separate?

Ari Weinstein

What’s really cool about the Decisions API is that it has all these new capabilities. It does inference in parallel, and it doesn’t have reasoning. It’s a smaller model than the ones we use for computer use, and those capabilities make it really fast.

They also make it a little bit less good at doing long-horizon, sophisticated tasks. I would say it’s still an open area of research for how we bring those approaches together. But I’m really excited to see what people build with the Decisions API.

swyx

One of the interesting things is that Dots now have attached personal computers, so they seem much more persistent. You’ve been using them for a while. How should people push the bounds? What should people aim for? What should they try?

Personally, right now I use it for a lot of customer service: “This was wrong. I don’t want to sign in. I don’t want to authenticate. Find whatever and just get it fixed.” How should we push further? What should people try?

2. Dots and Personal Cloud Computers

Ari Weinstein

Dots are a really cool product because each Dot has access to its own Linux virtual computer in the cloud, which is different from our other products. Traditionally, we’ve had access to a browser in the cloud, or it has access to your own computer, but now you get your own entire Linux computer in the cloud, so it can run full desktop applications.

It can also use a web browser. I think the powerful thing about computer use, and the reason why I think it’s so exciting, is that it makes it possible for the agent to do anything you as a person can do. All the software in the world was designed for humans, and now agents can use that same software, so you can delegate to the agent.

Anything that you would do on a computer, you can ask a Dot to do. What’s particularly useful is really going to depend on who the end user is and what’s valuable in their life. I would start by thinking about the things that you spend time on and how you could delegate those to an agent.

swyx

Yeah, a lot of flight booking and shopping, and honestly even playing a game or whatever.

Ari Weinstein

Totally. For me, something I did recently: I subscribed to a meal-prep service because I was trying to eat healthy. I really like this meal-prep service because it lets me customize the meals I order to a high degree of granularity. I can say, “I want this many grams of chicken and this many grams of rice.”

But it was so complicated that it took me 2 hours to place an order. I found that I could ask computer use to do it for me, and it did it in 15 minutes. I actually saved 2 hours. It did it 8 times faster than I could, and it saved me 2 hours on GPT-6.1 Soul. Those are the kinds of tasks that I feel are really powerful.

swyx

As a creator, I can tell you immediately that my number 1 use case is automating YouTube. YouTube doesn’t expose a lot of things through its API, and you have to put it in a VM and run it—for example, its A/B testing feature or making community posts. None of this is available through the API because they hate developers.

3. Why Computer Use Is “180 Degrees Different”

Ari Weinstein

I’ve heard that from our developer-experience team. They use it with YouTube a lot. It’s really awesome.

swyx

I want to draw on that. One of the leading AI podcasts—our friends—is famous for saying that computer use hasn’t advanced in the last 2 years.

Ari Weinstein

Yeah, they said that a few months ago, and I hope they have a different perspective now, because computer use is 180 degrees different.

swyx

He’s a tough guy to impress.

Ari Weinstein

Okay.

swyx

You’ve worked on—basically, you’ve spent your whole career working on some kind of computer automation, right? Shortcuts at Apple, then Sky, and then joining OpenAI. Can you draw what your through line is for what is driving you, what wasn’t possible back then, and what your milestones were?

4. From Sky to Self-Debugging Computer Use Agents

I guess, to add on to that as a follow-up question, what’s the major change from using Codex computer use last week through to today? Is it the model? Is it Dots? Is it the harness? What’s the history, and what really just changed in today’s announcements?

Ari Weinstein

On the through line, I’ve always been excited about automation and helping people automate tasks, because then you can save time in your life and focus on things that are more important to you than operating a computer very intricately. That’s why we worked on some of those products. I was at Apple before. We made a company called Sky, and we ended up joining OpenAI, which is really exciting.

swyx

It was almost like you had to hack around Apple until Apple was like, “Fine, we’ll just hire you and you can work on the inside,” right?

Ari Weinstein

It was a cool place to get to work. What was really interesting looking back at Sky was that we were working on computer use there as well, but the models were so much less capable. Now, in just the last 1 year, the models have become extraordinarily capable at computer use.

I think the biggest delta I see is that before, they could reliably start tasks, but then they would run into problems. Now they’re really good at debugging. They’re really good at trying again and introspecting on what is and isn’t working.

I also think the computer-use field itself has moved forward. We’re using more techniques now. Computer use often writes code, so if you look at it in Codex and expand the tool calls manually, you can see that it’s not just doing 1 action at a time. It’s actually writing JavaScript code that the computer executes to perform, sometimes, many actions at once, which is a great speed-up and a great capability.

We use more accessibility and multimodal interfaces. The model may use screenshots, accessibility, or Playwright. It can use a lot of different mechanisms based on the task at hand.

The model acceleration has just been amazing. What’s different today? I think we’re making computer use better all the time.

So I think just 1 day’s difference is probably a little less consequential than even the past month or the past 2 months. But, yeah, I think computer use in Dots is really exciting, as is the new model that we came out with.

swyx

In the keynote, Tijall[?] was mentioning 7× improvements in computer-use speed and a lot better performance on a few benchmarks. How do you think about measuring it? Computer use is one of those things where, as you say, there are improvements over time. Is it the harness? Is it the model? Is it post-training? How do you look at it internally when measuring how good it is, and what were the changes with the new model?

Ari Weinstein

We actually have a bunch of different ways of measuring it, some of which involve different permutations and configurations of the harness. It’s a bit of a complicated story because our production products have more safety checks, and those are configured differently based on the needs of the task at hand.

There are a lot of ways to measure it, but regardless of how we measure it, we find pretty consistent gains. Those gains are sometimes in the harness and sometimes in the model. I was really excited by the result that GPT-6.1 is even more cost-effective for computer use than its baseline cost improvement as compared to Astra. It’s really cool to see.

swyx

5. How Computer Use Sees and Operates Software

Yeah. One of the visuals I really liked from the livestream was that you’re improving the Pareto frontier of your curve. There was a lot of talk about how you’re improving it together with the harness. Can you give some examples of aha moments you’ve had, whether it’s the model driving the harness, the harness driving the model, or something else?

Ari Weinstein

I don’t mean to repeat myself, but introducing more modalities has been really powerful. One more specific example is that, in the past, a lot of computer-use products had to spend a lot of time scrolling. They would take a screenshot, try to do something, realize they had to scroll down to the next page of results, take another screenshot, try to do something, and scroll down again.

With accessibility, direct access to the DOM, and other things like that, the language model can now see an entire page or an entire application. It can write code that performs multiple steps at once. Those have probably been the biggest single aha moments.

There are a lot of tiny ones that are less exciting in comparison, but we also find that many speed improvements are driven by little paper cuts that we have to go in and introspect.

swyx

A lot of really hard engineering. I mean, app snapshots in general, right? I think people don’t quite get the difference because there’s a nice visual in Codex when we take an app snapshot, but they may not understand that you’re able to actually drive each button and that you have each piece of text in a very optimal representation.

Ari Weinstein

Yeah, exactly. It’s kind of fun, actually, if you want to be really nerdy about it. You can go into Codex and take an app snapshot by hitting the 2 Command keys. You grab the content from whatever app you’re working with and bring it into Codex or ChatGPT. If you click on the attachment and then click on the tiny button in the top right, you can see the raw text and the raw accessibility representation.

swyx

Dumping everything out—

Ari Weinstein

Dumping it out, but also making it token-efficient and doing it efficiently. There’s a bit of an art to it. It turns out that the same technology invented for humans who may have accessibility needs and want to use screen-reader technology is really helpful for them to use computers. It’s also really helpful for an LLM to use computers. That’s been really fun to work on.

swyx

For context, I feel like a lot of people don’t understand app snapshots. They don’t even know it’s a feature. When you double-hit Command, it pulls in what looks like a screenshot, and you think, “Why have I just opened a screenshot and thrown it in?” No, it’s actually pulling in all the metadata, all the code, everything.

Ari Weinstein

Yeah, exactly. If you take a screenshot of a web page that has a link, the screenshot doesn’t include where the link goes. It doesn’t include, for example, the full event titles if you take a screenshot of your calendar and the titles are truncated. But when you take an app snapshot, it gives the language model full context about everything, and that lets it do much more.

swyx

For those who want to see more, Jason Louu—I invited him to do a full workshop on this at AI Engineer. Amazing.

6. From Faster Than Humans to Superhuman Computer Use

He did a great job. I have a broader vision question on computer-use agents. Your example of taking a screenshot, scrolling the page, and taking another screenshot shows where we were. They can automate a lot. What are the bottlenecks? Is it the models? Is it the harnesses? Where do you see it going in 2 years? Do you see it just running for hours? How do we get there? Any predictions on where computer use goes?

Ari Weinstein

Yeah. I think what’s really crazy—and what the team has accomplished over the past couple of months—is that computer use is now faster at accomplishing tasks than the average human, probably in most cases. The next frontier is for computer use to be literally superhuman in its performance, where it’s as fast or faster at using software than expert computer users like us.

That will be really consequential and exciting because we’ll be able to build products that provide much more real-time experiences. Lowering the barrier to entry—or the activation energy, I suppose—of using computer use will make us start to default to doing certain things with agents that we’ve become accustomed to doing manually. That’s exciting because it will save us a ton of time.

There are a lot of little paper cuts and bottlenecks standing in the way. There are things on the model side, the inference side, the harness side, and in the representation. We find that as computer use gets faster, we’re increasingly bottlenecked by the speed of performing an operation. For example, in our computer-use benchmarks, a nontrivial amount of time is spent waiting for DoorDash.com itself to load.

swyx

Yeah. Then you just write a wait and execute the wait.

Ari Weinstein

Totally. You want to get as little delay as possible between when it finishes loading and when you trigger the LLM to do the next action. That’s actually a statistical science.

swyx

An event-driven way, maybe, to do that?

Ari Weinstein

When possible, you want it to be event-driven. JavaScript has load events, and the web browser has load events for web navigation, but there are other types of events that really can’t be event-driven. There’s a lot of complexity there.

swyx

The one that comes to mind is chatting with customer service. Replies could take 30 seconds or 3 minutes.

Ari Weinstein

Oh, right. I’ve dealt with so many bots with Codex. It’s great, but I also wonder if the other side knows they’re talking to a bot, because I’m answering in complete sentences, capitalizing correctly, and giving full reference numbers. It’s too good. It’s clearly too good. I don’t care; I’m just trying to get my—

swyx

I prompted it not to pretend it’s a bot, but to be a very annoyed human. Short one-liners, pushing it, all of that. I also tell it that while it’s waiting for responses, it should use subagents to research better ways to figure out what we need.

Ari Weinstein

That’s awesome. I also feel like half the time it’s a bot on the other end, so now you have bots talking to each other.

swyx

Yeah. I will also say that one milestone of computer use we’re at now is that 3 or 4 years ago, we were scared of hooking LLMs up to the web and to our devices. Now I’m having it configure DNS for me. I’m having it pay my bills. Tens of thousands of dollars’ worth of things, I’m just sending over and YOLOing with computer use. What’s the worst thing that can happen?

7. Agents API: Trust, Permissions, and Safety

Ari Weinstein

First of all, I’m just really excited that we brought computer use into the Agents API. This is great because a lot of developers are building applications that need to work with third-party websites and services. Computer use has this universality to it: it can work with anything. Now developers can build using the same computer-use implementation that we’re building on.

I think there’s great work to be done if you want to build your own computer-use harness, but it’s hard. We also train our models on our computer-use harness, so there’s an advantage to using the one that’s in distribution for the model. There might actually be a speed, cost, and accuracy advantage. So I think it’s really great for people to build on top of that.

To the point that you were making, I think we’re all still in the process, and maybe some of us are ahead of many people in the world of getting comfortable with this technology and trusting it. I think it’s incumbent on us to build that trust over time by making sure we’re building things that are reliable, building the right kinds of safety checks, asking for the user’s consent before doing something consequential like making a payment, and, depending on the application, making sure you’re only letting it access the websites or applications that it actually needs for the task. That’s something important to think about.

I’d really encourage people to try the new Agents API and build all kinds of cool stuff on it. We’d love to hear your feedback, depending on how it goes.

swyx

8. Computer Use for Coding, Testing, and QA

Have you seen any changes in the way it affects developer workflows? With Dots, you’re seeing it in Slack, and you’re seeing people use voice and build. The example Romain showed was, “Change this app and send me screenshots along the way,” and all this. Is there anything you’re seeing in adoption about how people are using computer use for coding workflows? Are there any tips people should take from that?

Ari Weinstein

One of my favorite use cases for computer use, and one that we see a lot in the wild, is computer use letting the agent actually test the software that the agent has built. That’s far more consequential than it sounds because, traditionally, you’d build something in Codex, Codex would build it for you, and then you’d have to test it. You are now QA for the agent.

With computer use, you can complete the software development life cycle: the agent can build software, and it can test it. I have a lot of fun building stuff and having the agent test it. By the time it comes to me, it’s already working.

I have extra fun because sometimes I’m developing computer use itself, and now I have a computer-use agent that’s using my computer-use agent, which is using something else. I really think this is a super-powerful use case.

Vibhu

I have a visual playtest skill that I’ve developed that really catches a lot of design issues that, normally, when you just look at code, you wouldn’t pick up. It’s also really good for cloning apps. If you’re using a shitty SaaS and you want to kill the SaaS, you just clone it screen by screen by screen. Obviously, computer use can completely drive everything, take screenshots, note it down, and then clone everything with Codex.

Thanks for all your progress. I think that is our time. This is not the last time we’re going to talk.

9. GPT-6 APIs, Async Tool Calling, and UltraFast Inference

Ari Weinstein

Yeah, cool. This has been really fun. Thank you guys for having me.

swyx

All right. Okay, we’re strict on the cutoff, so we’re just going to dive right in. Let’s do it. Yeah. Okay, Nikunj, we’re very excited to have you. You’ve shipped a lot on the API side, like we just talked about with Ari. You can now build with computer-use agents. Anything you want to highlight on the API side of the changes, and introduce yourself a little bit and explain what you do?

Nikunj Handa

Yeah, for sure. My name is Nikunj. I lead product for the API team. I’ve been here for roughly 3 years. I’ve been working on launching models, and I feel like that’s just been a constant thing throughout my time here at OpenAI. With every new model, we try to work super closely with the post-training team and the research team to figure out what’s new in it, and then we expose those capabilities in the API. That’s the basic way of putting it.

If you look at everything that’s new with GPT-6, the cool new capabilities that we launched were, firstly, async function calling. A lot of the things that you’re seeing in Codex and ChatGPT involve tool calls that take so long that you don’t have to pause the model’s execution while the tool is running. You can kick off a tool call, keep running, keep reasoning, and then check back in. We launched async tool calling and mid-turn steering. Now you can inject messages while the model is reasoning in the middle. As your tool call finishes, you can put in those instructions.

swyx

And that’s also partially a model-alignment capability, right?

Nikunj Handa

I feel like we’ve had it in the app. You could always steer it as it was reasoning. It wasn’t the best, but it’s gotten much better.

swyx

I’m excited to see how it does this.

Nikunj Handa

Our main goal in the API is to put things in the API once they’re trained into the harness. We wait for that moment until it’s good enough, and a lot of that is actually being powered by WebSockets, which we launched a few months ago. WebSockets open up this whole bidirectional communication channel with the model.

This is not the GPT Live thing; I’m just talking about GPT-6. You can do async tool calling, async reasoning, and inject messages, so it’s a really fun API to work on. I’m really enjoying it.

swyx

This is why we are the engineering podcast: we get to talk about WebSockets. This also pairs very well with Ultrafast, right? That is now, I think, for the first time ever, available in the API.

Nikunj Handa

Yes.

swyx

Which is basically the theoretical fastest speed you can ever get: frontier intelligence.

Nikunj Handa

Yeah. It’s been so exciting to work on that project. Before I go into the API, the most fun part of Ultrafast has been watching the inference team cook with Astra. They’re constantly having these Codex agents running, trying to squeeze out more performance.

At least for a couple of months, a lot of it was focused on efficiency and driving the cost down, which is how we were able to cut the Luna price by 80%. A lot of that was driven by all the inference improvements they landed. Now they’ve shifted gears toward how we can make this run as fast as possible.

Ultrafast has just been amazing to see on a model like Astra. To go that fast has been really cool. WebSockets was actually first launched for GPT-5.3-Codex-Spark, which—I can’t believe we named a model that—but that’s what we launched it for. Obviously, it helps so much because you’ve got to have the tool calls. You’ve got to really reduce the overhead of going back and forth with tools, and WebSockets is awesome for that.

swyx

Yeah. It’s always cute to see: I have my reset usage limit, and then I have my Spark usage limit that I never use.

Like, it’s there if I want it.

Nikunj Handa

I think it’s gone. Finally, it’s gone.

Vibhu

Yeah, you’re slowly killing off all the old ones.

Nikunj Handa

It was a great week. I mean, it was the first time we had frontier intelligence at extreme speeds. People really liked it.

swyx

Yeah.

Nikunj Handa

So, the first time it comes back.

swyx

10. The Rapid Story Behind Decisions API

And for GPT-5.3-Codex-Spark, it’s explicitly attributed to Cerebras. You guys are not confirming or denying that Ultrafast is related to Cerebras, but people do care and are wondering about it. You have your own silicon as well.

The elephant in the room: decision models.

Nikunj Handa

Oh, yeah.

swyx

Decisions API. We were the first podcast to do a big Jev[?] deep dive with Dogo[?], and I also featured him at AI Engineer. How quickly did you see Jev[?] and go, “...”?

Nikunj Handa

Firstly, huge props to Jev[?] and the Jev[?] team for inspiring a whole segment in the market. Obviously, Jev[?] comes out and everyone’s losing their minds over it. Our users are hitting us up, but our internal teams are also saying, “We need a much faster classification system.”

I don’t want to get ahead of some of the Dots features that are going to come, but you’re going to see some cool, really snappy, fast things built on top of the Decisions API. Props to Jev[?] for inspiring this whole thing. Obviously, a bunch of people at OpenAI get nerd-sniped by that, and they’re like, “How can we make this work?” We’re not going to train a new model, but—

swyx

Four weeks ago, this was not on.

Nikunj Handa

Not at all. No, this was Jev[?]-inspired.

swyx

I think you’re officially the first frontier lab to clone and adopt this.

Nikunj Handa

Yeah. I feel like OpenAI has such a strong hacker culture, and people just get excited about things. A guy from inference and one awesome guy from the infra team were like, “This is amazing. We’re going to hack on it.” They built a prototype, it worked, and now we’re just hill-climbing on latency and trying to make this as fast as possible. We want to launch it in the coming days, so as soon as we hit our latency target, we’ll try to get this out.

swyx

It’s interesting: at the same time as hacker culture, you also have, as Sam said, a 99% reliable API—probably the one with the most usage—which is your team directly.

11. What Decisions API Is and How It Works

How should people see Decisions API? I feel like a lot of people saw Jev, heard the buzz, and haven't built with it. You're making it very mainstream. What should people see it as? How should they use it?

Nikunj Handa

Yeah, I think the main use cases we've seen are really fast classification. All the computer-use demos have been amazing and really cool. I think there will be limitations, of course, in terms of having Astra write a JavaScript script to control your computer versus having Luna pick one action at a time. I don't think it's going to be at the same intelligence level, but maybe there are some computer-use tasks where this is good enough. I'm excited to see that come through.

The other cool prototype I've seen internally is people hooking it up with GPT Live. GPT Live is our bidirectional real-time API, and it's built on this model of front-end models and back-end models. GPT Live is this super-fast thinker-talker thing. GPT Live is the talker: super-fast and really good at delegation. You have something like Astra sitting at the back, but tool calling has always felt really slow in GPT Realtime. People have been putting together these tool-calling demos of GPT Realtime controlling a computer, and it just feels so much snappier and more natural. I'm excited to see what people do with GPT Live and with Luna on Decisions API.

swyx

So I want to iron this out for people, especially from the product side, because a lot of people have been putting out Jev clones. There have been about 100 in the last 2 weeks.

Nikunj Handa

Oh, really? That's amazing. They can clone the Jev API, which is honestly structured outputs, something OpenAI was first to do, too.

swyx

Yeah. Right. So let's iron out for people what a decision model is and what is important. It is not just latency. It's not just structured output, right? I could just have Luna. Is the decision model priced the same as Luna?

Nikunj Handa

If I have turned off reasoning and then have structured output, do I have a Jev? No, right? That's the real distinction.

Vibhu

There's a confidence there.

12. Memory, Higher-Level APIs, and the AI Cloud

Nikunj Handa

Yeah. Yeah, totally. We haven't trained a new model for this. We're building this purely on top of the same Luna weights that we have. This is really just Luna, and on top of that, what you're doing is constraining it. Structured output is a big part of it. You're really optimizing the inference stack to get very fast on TTFT and TBT, and because you can have multiple questions, you basically run those in parallel. You run them as a batch. There are all sorts of inference techniques people are working on to try to make it as fast as possible.

I'd say that at least our implementation of it at the start, in this first version, is zeroing in on top of Luna to see how it goes. Obviously, you want to put it out there. This is classic iterative deployment: put it out there, see what people think, and then make more model improvements as needed. So, yeah, that's Decisions API.

swyx

Obviously, as a benefit, you have vision and they don't have vision, right?

Nikunj Handa

True. We get it for free with Luna.

Vibhu

Yeah. I do think some of the innovations are still to come, if it's still the same Luna weights—the confidence stuff and calibration. This is something we've talked about on the podcast with benchmarking calibration, because the whole point is that RLHF kind of collapses you toward what you want to hear, but not necessarily what the amount of confidence is.

Nikunj Handa

Yeah. Yeah, totally. I'm eager to see how it pans out. Maybe these are going to be the key areas where we may have to hill-climb with a future model release.

swyx

And architecture-wise, the other thing that's in the debate—obviously, nobody knows because Jev doesn't talk about it—but the 2 speculations are, 1, maybe it's a diffusion model instead of autoregressive—

Nikunj Handa

But you're able to achieve the parallel generation in your way.

swyx

And the other one is some mechanistic-interpretability-type thing, where you're analyzing the activations and then just outputting the weights. You guys have all done the research on this, and people have speculated. There have been demos on both of these as well. I think Google shared Gemini Diffusion and Gemma Diffusion on a Jev-style output, and interp people have also pulled out interp from a middle layer, but this is all speculation. It's just about what you're trying to aim for, right? Because you can achieve the API. Everyone can achieve the API. It's actually pretty trivial. But then there's the speed, the accuracy, and the other calibration features. I don't know what else.

Nikunj Handa

Yeah. Yeah. No, totally. It's so cool that this whole space has been kicked off now. People are going to do so much cool stuff, everyone's going to learn from each other, and, yeah, I'm excited.

swyx

13. What OpenAI Is Building With the New APIs

I feel like, being on the platform team, a lot of your job is to empower builders. What do you think people should build with Decisions API and computer-use agents? Is there anything you've been building internally that you think really opens up after the new changes?

Nikunj Handa

Yeah. Decisions API use cases internally have been pretty obvious. The user-ops team was jumping on it: “We've got to classify all of our support tickets.” What else came up? Obviously, there were the really cool GPT Live demos. I'm sure the Codex app team might pick this up and try to do something cool with it. This whole thing started about a week ago, so it's very early. I'm excited.

Vibhu

There was a big push in evals for LLM-as-a-judge to have really low latency there, right?

Nikunj Handa

Yeah. Yeah, that'll be interesting to see.

With the Agents API, we're basically having a bunch of first-party products at OpenAI built fully on top of it. We've had the Codex security stuff that just went out, and that's fully built on top of the Agents SDK. We're having a meetings-type thing launching today. I think there was a demo. Do you remember the plugin extensions when Sam was showing them? There was a demo where you're in a calendar and you can have your meeting notes drop into something like Granola.

swyx

Yeah.

Nikunj Handa

All of that stuff is fully built on top of the Agents API. I'm just excited that we're getting this out. Let's see what people build on top of it.

swyx

I think you showed it off very well. The whole thing—edit spaces, pages, collaborate, add in your doc—that's a lot.

Vibhu

There's a lot of inspiration people can build from.

Nikunj Handa

Yeah. Yeah. All possible with Astra. Things just move so fast now that people go from idea to execution so quickly. It's amazing.

swyx

14. Prompt Caching, Pre-Warming, and API Performance

Is there something you want people to focus on to give you feedback? Maybe you're putting this out there and there's a fork in the road, and you want developers to help you decide.

Nikunj Handa

I think Agents API and Decisions API are our newest products. We'd love any and all feedback on those to figure out where to take them.

Over here, we're very open about Responses API, which is sort of our workhorse. We're really focused on performance right now, and that comes in 2 main ways. The first is latency. We've been rewriting the whole Responses API stack to make it as fast as possible from a TTFT and TBT perspective, so that continues to be a main area of focus for us.

The second thing we've been trying to do is go really deep on caching, particularly with these personal agents that are basically a single thread that just goes on and on forever. We've been trying to up our game on caching. We now provide guarantees of cache hits within 30 minutes. For one of our users, we just launched a much longer cache window, so we have a 12-hour caching guarantee that we offer. That gives you guaranteed cache hits for—

swyx

Is there a public API?

Nikunj Handa

Not yet. That's in preview. We're going to try to get that out to everyone as soon as possible. You pay a little bit more for the cache, and we guarantee cache reads for a much longer period. Even if your agent thread, for example, does something and then you come back to it 3 or 4 hours later, you're still getting the caching performance out of it.

swyx

You cut the cost there quite a bit with the new model, too, right?

Nikunj Handa

Oh, yeah. We're driving down cache reads.

swyx

For builders, they should implement it because it's significantly cheaper.

Nikunj Handa

Yeah. Just build your apps to be very cache-aware, and use our prompt diagnostics or cache diagnostics tools to figure out where things are dropping off. The caching part is really important.

I also want to talk about pre-warming. We have that in the API now. If you know, “Hey, I’m going to get this prompt. I just want to pre-warm the cache, pay the cache-write fee right now, and then have it ready to go for the next 30 minutes.”

Vibhu

And I can spawn many instances of that thread.

Nikunj Handa

Exactly. You can just keep going. I’m very excited about getting feedback on the low-latency performance so that we can keep making the Responses API the most performant and reliable way to build on top of an LLM. Then you basically have our new products, and I’m just looking for any and all feedback.

swyx

Yeah, just use it. Define our roadmap for us, please.

15. Context Compaction for Long-Running Agents

Vibhu

I think, for me, the caching thing is great. Obviously, it’s very needed, but at the end of the day, you’re still bumping up against a 1-million-token context, and that’s probably not going to change for the foreseeable future. You still need good compression, right? What is the best practice there?

Nikunj Handa

Yeah, totally. Firstly, OpenAI has its own proprietary compression, which is called compaction. It’s in the Agents API.

swyx

You decide for us, right?

Nikunj Handa

Yeah, exactly. In the Agents API, it comes built into the harness. If you’re in the Responses API, there are 2 ways of doing it. One is what we call server-side compaction: you basically tell the Responses API that if you ever hit this threshold of tokens, just auto-compact it and reduce the context being used. The second way is /compact, if you want full control, so you can /compact at any time and have your own logic for when to—

swyx

Yeah, but I mean, it is the manual override.

Nikunj Handa

Yeah, it is the manual way. A lot of the big coding agents like to do it manually. If you look at the Codex implementation in the open-source Codex harness, you can see that they use /compact.

There are also new compaction techniques that we’re working on. Some of them you’ll be able to see in the Codex harness; they’re already implemented there. They’re file-based systems that we’re experimenting with. Lots of cool stuff going on around compaction as well.

Vibhu

Cool. We’re running out of time. I think you’ve talked a lot about performance and a lot about the new APIs that you’re launching.

swyx

Can you give us any other hints as to things that you’re interested in as far as the future of the platform is concerned?

Nikunj Handa

We’re obviously very low-level. I used to work at Stripe before this, and at Stripe, a lot of the game was building these higher-level primitives and products on top of the core payments primitives. I’m always curious about what the best way of doing that is in AI. I think we’ve had a couple of attempts at that. We launched the Assistants API way back in the day, and it wasn’t really the right fit.

We’re sort of going off with this Agents API, and it gives you the Codex harness, but what’s the right amount of flexibility to give in that? That’s an open question. How should we have memory vaults and all of these higher-level API objects to abstract away more storage concepts? This is a whole space that I’m very curious about figuring out how we design.

I think a lot of things in AI are just: have a low-level API primitive, see an example harness, and then have your coding agent implement that. But how much of that should be built into the API is a constant question that I’m thinking about.

Vibhu

I don’t know if folks have thoughts on that. If anyone has ideas, it’ll be super interesting to hear.

swyx

Yeah, the analogy I always bring back to—and, to end there—is that you’re building an AI cloud, right? That’s something that Sam said a year ago, and you’re almost doing the AWS invention. You have to say, okay, this is EC2, this is S3, but you’re doing the AI-native versions of each of these. There are a lot of analogies.

You’re pre-warming caches for stuff that you know—

Vibhu

And it’s nice that it’s all exposed to builders because it just opens up ways that you can build new things.

Nikunj Handa

Yeah, absolutely.

swyx

Okay, awesome. Thank you guys. Thank you.