(Preview) All About Agents: Dots Arrive, Google and Apple Questions, Muse Follow-Ups, Super Intelligence, and the Mandate of Heaven
- OpenAI’s Dots launch exposes a positioning mismatch: it looks like a friendly mass-market product but behaves like a narrowly useful, unfinished productivity agent. Andrew calls the result “confused and confusing”; Ben sees the missing coordination layer Codex needs, yet finds Dots “weirdly neutered” and unaware of his actual projects.
- Ben had largely written off OpenAI’s consumer prospects after Gemini 3, but Amazon may give the company another route to monetization. OpenAI could become an agentic surface for Amazon products and ads: “Access our products, no problem. They also show our ads. That’s the trade.”
- Consumer agents should be sold as pain relief, not as accessories for the “grind lifestyle.” Ben’s pitch is the “modern-day Tylenol, the modern-day Advil”: technology has become painful for normal people, from roughly 700 apps to logins, two-factor authentication, and device migration.
- Effective agents need fresh, task-specific context and durable external records—not one enormous conversation. Ben argues that million-token context “sucks” when old instructions become buried; his preferred system continually writes things down, retrieves the relevant material, and starts a new thread when a task changes direction.
- Claude and Codex currently occupy complementary roles in Ben’s workflow. Claude is the context-preserving “chief of staff,” while Codex is better for executing a defined project—Andrew’s analogy is that “Claude is the barrel, and Codex is the ammo.”
- OpenAI’s larger mistake was forcing ChatGPT, ChatGPT Work, and Codex into a super app before their architectures converged. Ben thinks the destination is right but “they had no patience”; agents will ultimately beat apps because users can state what they want without knowing how to do it, and “it’s going to win because it’s going to be better.”
1. Dots arrived amid a welcome return to the AI Wild West
Andrew’s emblem for the moment: after OpenAI announced Dots, SpaceX AI bought dot.com and redirected it to the Grok download page. Ben welcomed the prank as a revival of the era when tech companies publicly mocked competitors—behavior harder to sustain once the industry became “a bunch of monopolies who have their own little fiefdoms.”
Listener Steven supplied the sharper launch critique: OpenAI has 90% consumer mindshare, yet showcased “San Francisco-coded techno-parents” building startups and side hustles instead of relieving family scheduling, kids’ activities, catching up with friends, groceries, or caring for parents. Andrew found the branding similarly incoherent: cuddly characters and a lowercase “d” signal a mass-market consumer product, while the demonstrated product remains productivity-heavy.
2. Consumer agents need to remove pain, not glorify work
Ben admitted he had “given up on OpenAI in the consumer market ages ago,” particularly after Gemini 3 looked strong and OpenAI still lacked an advertising model. He softened that conclusion because Amazon benefits from OpenAI as a surface for its external agentic opportunity: OpenAI can access Amazon’s products while also showing Amazon’s ads.
Andrew’s household offered a small counterpoint to OpenAI’s positioning. His wife initially asked why she needed Muse when ChatGPT already handled her personal organization, but after trying it for two days, “she’s now a Muse user.”
Ben’s preferred framing is unequivocal: “market these as the modern-day Tylenol, the modern-day Advil.” Agents matter because ordinary technology is painful, not because agents are novel; apps improved on what came before, but that never made them the final interface.
The concrete pain is cumulative: roughly 700 apps, forgotten locations, recurring logins, two-factor authentication, and device migrations. Ben abandoned moving to a Max halfway through because reversing the migration would mean enduring the same ordeal again: “It’s awful.”
3. Agent quality depends on context architecture
Codex’s project-and-thread structure maps cleanly to folders on Ben’s computer, making it effective when he knows exactly what he is building. Its willingness to discard a failed direction and start a fresh thread is a feature: carrying every mistake forward only confounds later reasoning.
Ben’s reversal on context windows is notable. The early assumption was that “larger and larger context” would solve everything; his current view is, “Actually, no, that sucks.” Long threads bury initial instructions, while capable agents continually write state down, keep active context fresh, and reload only what becomes relevant.
Claude handles continuity better, so Ben uses it as a “chief of staff” spanning servers, cameras, and other projects, then shifts defined execution into Codex. He has his agents write to ordinary files outside
.codexor.claudebecause he is “allergic to locking.” Andrew’s summary landed: “Claude is the barrel, and Codex is the ammo.”
4. Dots identifies a real coordination gap but does not fill it
Conceptually, Dots looked like the overarching liaison Ben wanted for his Codex agents. In practice, “Dots kinda sucks”: it ignored his projects, discovered a calendar plugin, and fixated on an October 14 conflict—prompting both “How did you know that?” and “That’s not what I need help with.”
Dots also failed to recognize Ben’s custom Codex permissions file or where it was until explicitly directed to it. That made it simultaneously intrusive about irrelevant information and “unknowledgeable” about the work it was supposed to coordinate.
Spaces, by contrast, resembles a system Ben independently built for Gecko: cards hold an item’s history and attachments, including an interview transcript ready for editing. Even the interface can remain lightweight—clicking Done merely sends Gecko a prompt, and the agent performs the bookkeeping.
The frustration is that many constituent ideas “make sense and are super cool,” but no organizing principle makes Dots, Spaces, Pages, scheduled tasks, ChatGPT Work, and local versus cloud execution legible. Compute constraints deepen the mismatch: Dots could help broader audiences, yet initially reaches only Pro-tier subscribers.
5. OpenAI’s super-app vision arrived before its foundations
Ben calls the “cardinal sin” the super app. ChatGPT was stateless across fresh conversations, ChatGPT Work ran inside a VM, and Codex operated on the local computer; combining them while they still had different architectures “didn’t make sense architecturally” or conceptually. “They had the right vision, but they had no patience.”
The intended destination still convinces him: users should state what they need rather than learn which app, website, or workflow performs it. Ben says agents will win because the experience will be better—not because nerds favor them or the economy demands them. Andrew frames that advantage as a consequence of the rest of tech sucking. The preview ends just as a listener challenges whether consumer agents can be winner-take-all when demand does not improve the product and switching costs remain structurally low.
Full transcript
Hello, and welcome to a free preview of Sharp Tech. Hello, and welcome back to another episode of Sharp Tech. I'm Andrew Sharp and on the other line, Ben Thompson. Ben, how you doing?
I just had an amusing moment. You had to drop a couple of hellos. One of them was my fault—
Mm-hmm.
I wasn't quite ready. The second one was a little too enthusiastic. It was more a greatest-of-all-time hello. Which, by the way—
Yep.
I always laugh because my kids would imitate it. Whenever I'd say your name, they would just drop into—
Mm-hmm.
A long, extended hello. And it turns out all your listeners do this. You've gotten multiple emails—
Yeah.
Of kids imitating the Andrew hello. Anyhow, you said after you did it too enthusiastically that you had too much vim.
Too much vim, yes.
Did you know that?
I had to crank it down a little bit.
Did you know that this is very funny because this is reestablishing our normie-versus-nerd credentials? Did you know—
Okay.
That there is a Vim—Vim as a tech reference?
Oh, no. What does Vim stand for? I'm still a little bit ashamed of the VPN-VPS mistake—
Oh, I think that's fine.
Mm-hmm.
That was also very funny. No, Vim is a text editor. There are famous wars in Unix-type systems. There are Vim people and there are Emacs people. Emacs is the ultimate nerd-out sort of—
Mm.
You can configure literally everything.
Okay.
And people have these massive configuration files that they carry around their whole lives, and they can't function without them.
So the nerd Chads, they're all Emacs people, is that right?
I think Emacs is going to be more on the nerd side of the spectrum, and Vim a little less. I don't actually use any of these. If I have to edit text in a terminal, I use nano, which is just the bare-bones sort of thing. So I'm actually out of my depth as well, but it did amuse me. Vim.
Okay.
There you go.
Well—
That's your word of the day.
The more you know, you know? The Emacs Chads versus the Vim virgins out there—
You know what? If you know what Vim is, you're already off the reservation, so it doesn't matter.
Okay.
Everyone is a Chad nerd in that scenario.
Well, Ben, we are going to be talking agents today, and I want to start with a note that I saw earlier this week. OpenAI famously announced Dots on Tuesday, and at some point either before the announcement or immediately afterward, SpaceX AI bought dot.com, and dot.com now redirects straight to the Grok Bot download page. I highlight that sequence first because I find it delightful, and second just to step back and celebrate the Wild West nature of the ecosystem right now.
Oh, no, 100%.
It's the best.
That was going to be my response to this. I saw the rundown. This sort of thing used to happen all the time, like companies doing practical jokes on each other. There was a whole thing where I think Salesforce would do all these crazy stunts at its keynotes. They would make fun of their competitors in ways that were way over the top and arguably over the line—actually, almost certainly over the line.
Mm-hmm.
That used to be the way tech operated all the time. And you're right, it's good to have it back. Dot.com goes to a different spot.
I feel like—
It's great.
We completely missed that generation of technology. I certainly did.
Yeah, there was no podcast, period, back then, which is a real shame—
Mm-hmm.
Because there wasn't enough documentation going on. But then again, maybe that's why it all happened: there wasn't enough documentation going on, you know?
That's right. You could get away with—
What's done is done.
Line-stepping here and there. Well, this is all a thousand times more fun than dealing with a bunch of monopolies who have their own little fiefdoms. So thank you to all of our participants in the AI wars—
That's right.
And the frontier wars.
That's right.
As for Dots and OpenAI DevDay, I will say at the top that my feelings after consuming the content from the OpenAI event are close to the exact opposite of my feelings after watching Meta Connect last week. So I'm curious about your thoughts. We'll start with an email from Steven. He says:
“Ben and Andrew, I can't wait for Ben's rant on how OpenAI has 90% of consumer mind share when it comes to AI, but somehow completely fumbled the Dots launch to feature San Francisco-coded techno-parents launching their startups and side-hustle companies as the primary use case, rather than going hard into family scheduling, kids' activities, catching up with friends, taking care of parents, groceries, et cetera. It even came with dystopian fake-house aesthetics that feel like a Black Mirror episode, presumably to ensure we really get the point that the real world and real people's lives are an afterthought, and all that matters is the techno-grind lifestyle.”
What do you think of those sentiments from Steven?
It's interesting because I think Steven made me realize something that's just been in the back of my head, which is that I've given up on OpenAI in the consumer market ages ago. I think it was when Gemini 3 looked amazing, and I thought, they don't have an advertising model. Even if they launch it now, they're swimming upstream. It's going to be tough.
Mm-hmm.
They had the chance, and they missed it. Now it sounds like their advertising product is doing okay, by all accounts. I think this relationship with Amazon is very interesting. There's a bit where it's very much in Amazon's interest that OpenAI wins because OpenAI is a surface for Amazon Ads.
Mm-hmm.
That's a way for Amazon to deliver on the sort of external agentic opportunity. Yes, you can have a great experience. Look at OpenAI. It accesses our products, no problem. They also show our ads.
Right. It doesn't kill our ads business.
That's the trade.
Sure.
So maybe I've jumped to conclusions too quickly. This didn't really bother me because it's what I expect, and I think the big change, if anything, is the extent to which Meta and Muse have walked into the opportunity and Google hasn't.
Right. I feel like Meta has highlighted a consumer opportunity that other companies—and including on this podcast, we haven't really talked about the consumer space for nine months on Sharp Tech—
Well, the other thing was the ChatGPT app reboot—
Mm-hmm.
Which made it unbelievably more complicated, I think in an attempt to introduce people to what is possible and capable.
Yeah.
But along those lines, it was still kind of jarring for them to introduce Dots by basically mocking the use cases. “Of course you could use it to check your school email, but you could also do this.”
Right.
And it's funny because I actually think there is a lot to mock about the use cases being put forward for consumer agents, to be clear.
Mm-hmm.
Everyone is terrible at this. But my prescription is not that you highlight how you can embrace the grind lifestyle.
Yeah.
My prescription is actually that you have to go the exact opposite direction and market these as the modern-day Tylenol, the modern-day Advil.
Mm-hmm.
It's pain relief. There's all this stuff in your life that is painful, and agents can make it easier and better.
Yeah.
And it helps that I actually think that is true. The thing I didn't really put in my Monday article—I mentioned it in a follow-up on Tuesday—is that the reason I'm optimistic about agents is not because I'm a nerd who thinks agents are fun and cool.
Mm-hmm.
That part is all true, to be clear. The reason I'm optimistic is because I think tech sucks. It sucks—
Yeah.
For normal people. It's so hard to use. Apps on an iPhone were way better than what came before, but there is a mindset that says apps were so much better than before, therefore apps are the best experience.
Yeah.
They're not. And actually, if anything, apps have gotten way harder to use.
I was going to say, there was probably a point in time when they were a much better experience. At this point, I think you mentioned you have around 700 apps on your phone.
That's true for everybody.
They have so many apps that they don't even know where they are. You just search and find one, and they're all kind of a pain in the ass to use these days.
Well, especially things like security. Just keeping track of logins, then two-factor authentication, and stuff doesn't stay logged in, so you have to do it all over again. Migrating—
Yeah.
To a new phone—one of the reasons I bailed on the Max was that I thought, “I should try it for at least a week to see if I can get used to the size.”
Mm-hmm.
I got halfway through migrating, and it was like, “This is so painful that if I go through with this and I don't like it, I'm going to have to do the same thing all over again in the opposite direction.” No, I'm done.
I don't want to grind through it. Yeah, fair enough.
It's funny because I've been like that with a Mac for a long time. Look, I knew the big price hikes were coming, and I was telling everyone, “If you need a computer, buy a computer.” I did not buy a computer because—
Mm-hmm.
I'm like, “Mine's fine.”
Don't put me through that.
I just—
I don't absolutely—
It's awful.
—have to go through that.
Yes.
Well, as far as Dots is concerned, it's interesting because I had the experience of introducing my wife to Muse last weekend. I thought it would be more useful for her than it has been for me, and her first response was, “Why do I need that? I use ChatGPT for all of my personal organization and AI work.” I said, “Just try it.” She really liked it the first couple of days, so she's now a Muse user.
I understand that, from OpenAI's perspective, the best market for them is still enterprise work. Pointing Dots in the productivity direction actually makes sense. The consumer-facing agent category is still mostly theoretical, so this could be a reflection of continued strategic coherence on OpenAI's part.
However, watching the event, the visual language around Dots—and, of course, the lowercase D for Dots—
Oh.
—an OpenAI trademark.
Really?
Very groan-worthy—
Oh.
—stuff from them, as ever.
Seriously.
But the cuddly little creatures that are just a straight rip-off of the Muse character—
I think that's unfair. Muse—I mean, these are, I'm assuming, designed a while ago, but—
Maybe. Look, all I'm saying is that visually, it is completely incoherent to me because it looks like a product that's designed exclusively for mass-market consumers, but that's not really what the product is. That was part of what was confusing to me this week.
I'm not sure what the product is. The way I see it, I'm going to have to update my view.
Okay.
There is a clear need in Code, or at least that I've found. This is probably one of those things where I was overextrapolating from my experience. But from what I find—
Mm-hmm.
—I like the Codex sort of organization, where you basically—
Yeah.
—have projects that align to files on your computer. You can see that it grew up connected to the local computer, which is a big part of why smashing all these different products together feels so weird. But you have—
Mm-hmm.
—projects that align to folders, and in those you have different threads. Part of it is that I think people don't realize that, in general, one massive chat thread is bad. The bigger—
Mm-hmm.
—it gets, the more context it gets, the—
The more confounded the reasoning becomes.
It just gets super confused because—
Yeah.
—you load instructions into it at the beginning, and those are so far back in context. There was a big push early on in the first ChatGPT era: “Oh, we need larger and larger context. Context is the problem. Oh, it's a 1 million-token context. This is incredible.” Actually, no, that sucks. This is where I was trying to get at with the write things down. The way they actually—
Mm-hmm.
—function well is by continually writing things down. And the way—
Yeah.
—they keep track of stuff that's written down is by having relatively fresh context, and then loading stuff in as it's appropriate. And—
Okay.
—managing this is actually really, really interesting. This has been part of the whole thing I've done with this Gecko bot and just getting a much more fine-tuned understanding of how this works. But by and large, doing a thread per task is better than—
Mm-hmm.
—having this massive, ongoing thing. Anyhow, Codex actually aligns with this pretty well, and Codex is really good about spinning up new threads. There's one where I'm like, “Oh, let's redo this with this direction. Let's throw this away.” Instead of trying to continue and remember all the stuff it did that was wrong, which I talked about—
Mm-hmm.
—this is an issue here—it just starts a new thread. It can do it itself. Anyhow, this is all great. If I know exactly what I'm working on, the Codex organization makes much more sense to me than Claude, which is much less intuitive to me as far as—
Mm-hmm.
—organizing a project I'm working on.
Okay.
Claude, on the other hand, manages context much better. One long Claude thread just keeps track of itself, I feel, better than Codex does. And the way—
I see, yeah.
—this worked out for me is that I have my agents write everything down because I want to be able to use any agent, and I don't want stuff inside their weird folders—.codex, .claude, or whatever it might be. I want it out. I don't want to be locked in. I'm just allergic to locking—
Mm-hmm.
—the way I've always operated, the way I use text files. And I basically use Claude as a sort of chief of staff. It knows everything that I'm working on. I have—
Mm-hmm.
—this server that just came in, and I also have a bunch of cameras that I'm trying to set up, along with all these different things. I can bounce around, and I use this Claude thread to keep track of it, and it's writing it all down. Then, when I'm actually working specifically on a project, I find it easier working in Codex.
So Claude is the barrel, and Codex is the ammo as far as executing on your side.
That's a great analogy. I think that that—
Mm-hmm.
—is sort of the case. So I see Dots come out, and I'm like, “Yes, that is so needed in Codex.” An overarching—
Mm-hmm.
—sort of agent that can be my liaison to all my other agents that are working.
Okay.
Conceptually, this makes a ton of sense. Even starting there, that was my assumption around Dots. That is so heavily nerd-coded, right?
Right.
It's like, “Let me manage all my AI-building projects.”
How big is the market really for that sort of thing? Sure.
So here's the problem: Dots kind of sucks. The reason it sucks—
Mm.
—for me in this use case, it seems not to know anything about what I'm working on. It immediately started looking for my calendar. I have a plugin to pull my calendar in for Gecko. It found that, and it's like—
Mm-hmm.
—“I notice you have a conflict on October 14. Do you want me to help you resolve that?” Number 1, I'm like, “How did you know that? That was weird.” Of course, I figured out how it did. But number 2, this is so far—no. That's not what I need help with.
That's not what I'm signing up for.
I need help managing all these projects that I have. It starts asking me for permissions to everything. I made my own custom permissions file for Codex so it knows when it has to prompt me to approve something and when it doesn't. I want it a little bit—I want fairly loose, because I have a dedicated computer just for this. I don't want total full access, so I actually made my—
Okay.
—my whole thing. It didn't know where that was. It didn't know about it.
Mm.
I had to direct it to the file, and then it found it. I'm like, “Oh, I'm working on this in this file.” It's like, “I don't know what you're talking about.” It's weirdly neutered and unknowledgeable about stuff. And so—
Hmm.
—but again, like I said, it latched onto this calendar thing. It's weird. It's just a weird—I see a hole here in the Codex offering, but then—
Okay.
—to your point, everyone in the media is looking at it. It's like, “Oh, they're a Muse competitor.” I guess so, but it's just—
It's not really that.
I know.
It's not as good as Muse. It doesn't sound like it's solving your problems either. It's just sort of a confused and confusing product for the time being. And, of course, now there's Dots, Spaces, Pages, local/cloud, ChatGPT Work—
So the Spaces thing is a great idea—
Scheduled tasks.
By the way. The way I backed into Gecko bot is I actually created my own page that just listed all the stuff I was working on. I developed that more. The biggest thing is basically cards. Any item that I get reminded about by Gecko, I click on it, and it has all the history of that item and any attachments to it. For example, I’m editing an interview now. Devin does the initial transcript, and then he posts it.
I don’t have to find where the link is or whatever; it’s just in the card. I click it, download the text file, and edit it. It’s great. The idea of having some sort of substantiation place where you can actually see the stuff you’re working on in a way that makes sense—
Yeah.
I actually have a UI on that page, and it’s all fake UI because the UI is just prompts. You click it—say you click Done. It’s not actually—there’s no machinery in the page that’s going to the database and clicking Done. It’s firing off a prompt to Gecko saying, “This is done,” and then Gecko actually does all the bookkeeping. So this whole concept of a space—
Mm.
Now, obviously, they have way more resources, way more capabilities. What it looks like you can do with those spaces looks great. So I think—
Mm-hmm.
And this is the frustration: There are all these ideas that make sense and are super cool.
And should exist.
Right.
Yeah.
There’s no organizing principle here. And I think it is all downstream—to bring this full circle—to the emailer: They had ChatGPT. They decided they had to get money from the enterprise, and we’re still living in sort of the shipwreck that is downstream from that. Everything’s great, and it all sucks because you don’t know what’s going on. Yeah, I think the cardinal sin was the super app.
Okay.
Pushing this stuff together—in the long run, you can see why all this should be together. Like, Dot should be the ultimate manifestation of ChatGPT. You shouldn’t have a bunch of different threads. Even in ChatGPT, it’s kind of weird that you have a bunch of different threads. It’d be easier just to have one thread, right?
Yeah.
You just talk to your agent, and it answers your question.
That’s true. Sure.
That’s what Dot is, I think, ultimately going for. But this app wasn’t ready. And then you can imagine it going off and spinning off new projects. For normal people, that means you have to have cloud compute. Just like Meta did with Muse, you have to have the VM in there. And they were so eager. I think they had the right vision, but they had no patience.
Yeah.
It wasn’t ready to be shoved together. When they shoved it together, you had Codex, which worked on your local computer; you had ChatGPT Work, which had its VM; and you had ChatGPT, which was sort of a stateless thing that synced between devices, but every conversation is fresh. There was no ongoing environment.
Putting those 3 things together didn’t make sense architecturally; it didn’t make sense conceptually, even if in the long run those should be 1 product. It’s a lack of patience. It really, really is. Where they’re going makes sense, but they’re not there yet, and by accelerating it, they’re making everything harder to use and confusing and a mess.
Including with Dots. They don’t have enough compute to serve Dots to more than just the Pro-tier subscribers right now, but Dots could be a product that’s actually really useful to more people.
It might be telling that Dots is exposed in the same part of the ChatGPT app as the rest of ChatGPT, whereas Codex is its own tab. By the way, why do both Anthropic and OpenAI not have a dedicated Claude Code app and a dedicated Codex app? This having to go into a little thing to get to your thing—it just drives me crazy. There’s no reason for it. It’s weird. It’s like everything has to be in 1 app. Does it?
Mm.
I understand that lots of people have this app already installed.
Take advantage and put more capability in there for the people who might not discover it otherwise.
Yeah, but it just makes for a really crappy user experience.
Mm-hmm.
To go super meta, this is all of tech. The user experience sucks everywhere. It really does.
Yeah.
And this is why, ultimately, the actual answer for all this—and again, I think this is where Dot is going in the long run—
Okay.
—is you don’t need to know any of the complexity. You just ask your agent. You say, “I want to…” Instead of operating at a level of “I need to know how to do something,” you just need to know what you need to do, and the agent will take care of the details.
Mm-hmm.
And I absolutely do think that in the long run, and probably sooner than people think, the experience of using an agent is going to be so much better than using a bunch of different apps and websites that this paradigm is going to win. It’s not going to win because nerds want it to win. It’s not going to win because the economy needs it to win. It’s going to win because it’s going to be better.
It’s such a good take. Well, it’s going to win because the rest of tech sucks, as you said 20 minutes ago.
Yes.
And that’s something that millions and millions of people identify with. Well, speaking of the agent future here, Andrew says, not me: “Hi, Andrew and Ben. I’m halfway through Thursday’s episode, and I can’t wait to listen to the rest. But you’ve been discussing why AI agents for consumers will be or should be a winner-take-all market. But I feel like you aren’t addressing my main concern with that argument. More consumer demand doesn’t make the product itself any better, and the switching costs aren’t all that significant structurally. So why do you think this is a winner-take-all market?” Any thoughts on that, Ben?
All right, and that is the end of the free preview. If you'd like to hear more from Ben and I, there are links to subscribe in the show notes, or you can also go to sharptech.fm. Either option will get you access to a personalized feed that has all the shows we do every week, plus lots more great content from Stratechery and the Stratechery Plus bundle. Check it out, and if you've got feedback, please email us at email@sharptech.fm.