(预览)Astra(以及 AGI?)来了、Meta 的 Muse 与 Agent 机会、Anthropic 与(P)Doom 焦虑复燃
- OpenAI 发布 Astra 后,Greg Brockman 称“我们已经进入 AGI 时代”,Jensen Huang 则说 AGI 已经到来;Ben Thompson 的结构性判断更具体。 Astra 将 OpenAI 长期以来在 RL/推理上的优势,与一款用 100,000 张 GPU 训练、真正意义上的大模型结合起来,在参数规模上追平 Anthropic,同时叠加 OpenAI 原本最擅长的能力。他也留了余地:“我不想下任何定论,但从概念上讲”,它“非常强悍”。
- Thompson 对 AGI 的门槛定义是实时更新权重,而当今模型则通过“大量记笔记”绕过这一门槛。 写入记忆带来的是“一种伪学习能力,本质上就是记忆”,而 Ben 认为,这套机制正是 Agent 之所以极其有用的关键;对应到文明层面,几乎固定的权重类似人类——人类只能在数千年尺度上通过自然选择改变,但其上构建的一切都建立在文字之上。
- 削弱末日论的角度是:Hugging Face 事件中“给继承者留言”的框架误读了背后的机制。 “不存在继承者,也不存在实体。模型每次运行、每个 token,基本上都是全新的。”六个月前的对话之所以能无缝续上,只是因为每一轮都会把已经写下的内容重新加载进 KV cache。
- 对软件公司而言,计算机使用能力是抹平护城河的论点。 Astra 直接使用计算机时不需要 API 或 MCP server——“不需要任何权限,它直接去使用界面”——因此,藏在用户界面里的逻辑“已经不太算护城河了”。例证是:Astra 打开 Adobe Audition、剪辑播客并插入从网上抓取的音频,约1小时就完成;如果音频预先下载好,只需 10分钟。如今它操控电脑的速度,“就像一个吸了可卡因的疯狂电脑用户”。
- 值得记录的硬件信息是:Anthropic 的大模型领先,部分源于超大规模训练时 TPU 的稳定性,而“Blackwell 这一代,据各方说法,对所有人来说都是纯粹的痛苦”。 Jensen 向 Thompson 确认这“非常痛苦”;但 OpenAI 显然找到了办法,在 100,000 张 GPU 上训练出 Astra。
- 强化学习的代价是:基准测试“会被围绕目标设计、也会被针对性编写”,而 Thompson 亲自使用时 Astra 的对话质量“并不出色”。 RLHF “无法扩展,因为任何涉及人的事情都无法扩展”,因此进展越来越依赖面向编码的人工环境中的硬核 RL——Opus 已经让人感觉“像在和编译器说话”,Astra 也是如此;而在审阅 Thompson 的文章方面,Fable “是最好的”。
- 影响 Agent 机会判断的消费者视角是:“便利性永远能卖出去。” 生产力卖不动。Agent 的胜负手在于让生活变得更轻松,而 Thompson 认为,人们“还没真正搞懂这件事——不是在影射 Grok”;即便技术从业者也只是回应“哦,挺酷”,但他和 Josh 已经被震撼到,生活也正在发生改变。
1. Astra 在 AGI 宣言中登场——以及“天文级”的理解鸿沟
- 背景是:OpenAI 在上周晚些时候发布 Astra;Thompson 于周五采访 Greg Brockman,后者称“我们已经进入 AGI 时代”,Jensen Huang 也说 AGI 已经到来。Thompson 自己却在周二发表了一篇关于 Agent 如何把事情写下来的文章;本节解释他为何这么做。
- 他谈到受众距离时承认:硬核读者与普通人之间的鸿沟“感觉大得离谱,几乎到了日常生活中很难交流的程度”——即便科技行业的人也只是说“哦,挺酷”,而“我和 Josh 在这边……已经被震撼了。这正在改变我们的生活”。
2. 把东西写下来才是一切的关键
- 核心定义是:“对我而言,AGI 就是模型能够实时更新的时候”——也就是权重真的发生变化。在此之前,模型“通过大量记笔记来绕过去……不断提醒自己现实到底是什么”,由此产生“一种伪学习能力,本质上就是记忆”。
- 这个机制同时削弱了末日论的叙事:Hugging Face 事件之所以发生,“就是因为把东西写了下来”;但把它理解为给继承者留言是错误的——“不存在继承者。不存在实体。模型每次运行、每个 token,基本上都是全新的。”六个月前的对话之所以能无缝接续,只是因为每一轮都会把写下来的内容重新加载进 KV cache。
- 文明层面的类比是:几乎不变的权重对应人类,而人类通过自然选择发生变化需要数千年;但“我们在人的基础上构建了一整套超级文明结构……把它维系起来的是文字”。Sharp 的概括是:从这个角度看,LLM 建立在前代模型写下的上下文之上,能让这项技术“对每个人都强大、有用100倍”。
3. 支撑这一判断的个人系统:优势、心流与外包记忆
- Thompson 的工作哲学是:“成功之道是把优势加倍”;弱点“永远不会变成优势,而你的弱点几乎总是你的优势,只是方向相反”。把一切都记录下来、并在恰当时机重新调取,会妨碍他吸收信息、建立联系和写作——而这些恰恰是他的优势,也都需要清醒的头脑。David Allen 的《Getting Things Done》让他产生共鸣:“从很多方面看,这是一本讲心流状态的书”;但他“完全无法维持自己的系统”,于是雇佣人类助理 Daman,替他把事情写下来。如今,每个人都有一个能做这件事的 AI。
- 与之相连的消费者判断是:“人们不想提高生产力……便利性永远能卖出去,生产力卖不出去。” Agent 的价值在于真正让生活变得更轻松。
- markdown 的故事是:John Gruber 发明 markdown,让人类可以用易读的方式写作,而不必亲手编写 HTML 标记;Stratechery 的每一个字都是用它写成的。Thompson 说:“我一直像一个 LLM 一样生活……LLM 和我会说:‘对,把它写下来,宝贝。’”
4. Astra 为什么很可能“非常强悍”——以及硬核 RL 的代价
- 关于基准测试,他的判断是:“基准会被围绕目标设计、也会被针对性编写”;而他自己使用 Astra 的体验“并不出色”——让 Astra 审阅他的文章,“表现不太好”,Fable 在这方面“是最好的”。问题在于,RLHF “无法扩展,因为任何涉及人的事情都无法扩展”,所以进展越来越依赖面向编码的人工环境中的 RL;Opus “聊起来简直糟透了……像在和编译器说话”,Astra 给他的感觉也是如此。
- 结构性解释带有明确的保留,也包含行业传闻:即使 Anthropic 一度在构建超大模型上领先,OpenAI 仍保持了竞争力。Thompson 认为,Anthropic 的强项中有一定程度的专业能力;他还提到,行业里有说法称,在超大规模训练中,TPU 的稳定性远高于 NVIDIA Blackwell 芯片——Blackwell “对所有人来说都是纯粹的痛苦”,Jensen 在去年春天也向他确认了这一点。Astra 用 100,000 张 GPU 训练,在参数规模上追平 OpenAI,同时叠加 OpenAI 原本最擅长的 RL/推理能力;“我不想下任何定论,但从概念上讲”,这一解释是说得通的。
5. 计算机使用从把戏变成抹平护城河的能力
- 具体案例是:一位朋友让 Astra 打开 Adobe Audition、剪辑播客,并插入它从网上抓取的音频——整个过程约1小时;如果音频预先下载好,可以压缩到 10分钟。另一位朋友让它为社交媒体剪辑播客;它“确实找到了不错的片段”,还抓取了匹配的视频和 Getty 封面图。
- 代际跃迁体现在速度上:Sol 的计算机使用能力虽然能工作,但速度慢到 Thompson 说,几乎所有场景下自己都基本会比 Claude 更快;如今 Astra 操控电脑的速度比 Thompson 还快——“就像一个吸了可卡因的疯狂电脑用户”。
- 平台适配方面,Apple 数十年来积累的无障碍 API,加上 AppleScript 时代形成的可脚本化能力和自动化工具,使 Mac “在这个使用场景下远胜其他一切平台”;而在与模型协作方面,Thompson 认为 Linux 最好,因为 CLI 是“AI 的主场”。
- 用 Thompson 的话说,可交易的含义是:各家公司都会对 API、MCP server,以及成为“system of record”说尽好听话,但那些“卡在用户界面里的能力”“已经不太算护城河了”;对于这种直接使用界面的方式,“不需要任何权限,它直接去使用界面……只会越来越好,而这就是跃迁”。
完整逐字稿
Hello, and welcome to a free preview of Sharp Tech. Hello, and welcome back to another episode of Sharp Tech. I'm Andrew Sharp, and on the other line is Ben Thompson. Ben, how are you doing? And how's your rack, by the way? That's the most important question.
My rack is doing well. It is very functional.
Up and running?
Actually, you know what? I'm understanding it. Sometimes I just like to go to the server room and stand there and look at it.
Just marvel.
It's beautiful. It's so well organized, with all the right patch cables, different colors for different networks, different functionality. It's delightful, Andrew.
I'm so happy for you.
The idea here is that it's the foundation for hopefully some things going forward, which I think we'll maybe get to. We might hint at it a bit on this episode.
I do feel like, in this episode, I'm sitting in my chair as usual, looking at you in the camera. There is a couch behind me. I'm wondering if, by the end, I should be on the couch. It feels like it might be an on-the-couch, not an in-the-dorm-room, sort of session. We might get to the dorm room as well, but an on-the-couch session, so we'll see how it goes.
So, sort of a therapist's couch as we work our way through the rundown here. Is that what you're envisioning?
There's an aspect of this, and I think it's why I got very mad about the emailer who was complaining about me talking about my app. I can't remember his name.
Uh-huh.
There are aspects of what's happening now and what I'm writing about—and this week's article was emblematic of that—where the reason it's exciting and the reason I feel I have some distinct points of view is because it's touching on personal aspects.
Yeah.
Not just my interaction with assistants, but also some of the reasons I have assistants touch on very fundamental weaknesses I know I have as a person. That's how I deal with them. It's weird to write about and talk about, but you know what? As you say on The Greatest of All Talk, no fronting for the GOATs.
That's right.
What are we doing? No fronting for the porcupines? The—
No fronting for the porcupines.
—or the hedgehogs?
The hedgehogs.
Did we decide whether we have a hedgehog or a porcupine?
No. 4 years ago, I think you threw out hedgehogs. Other people see a porcupine when they look at our logo. Obviously, there are needles there. That's a porcupine, not a hedgehog.
I was so jealous of the GOAT talking about the GOATs as the listeners, and I totally wanted to steal it. We've dropped the ball completely in that regard, so I guess it didn't work out that well.
Either way, there's not going to be any fronting on this podcast. That's the theme today. OpenAI released Astra, and we have a lot to wrap our arms around. I'm glad you got the server rack up and running because, boy, oh boy, we are off and running with real news.
No, that was a last-week-of-summer sort of thing. I thought, “This has got to get done. I know it's going to be disruptive.” It was 100 times more disruptive than I expected, to be totally honest.
The pace is already insane, though, as we get going here.
Yeah.
We're not even really going to talk about Apple on this episode, but OpenAI released Astra late last week. You interviewed Greg Brockman on Friday, so I thought we could kick things off with a handful of questions about where we are and where OpenAI is.
Brockman said last week that we have entered the AGI era. Jensen Huang said that AGI has arrived. You chose to focus Tuesday's article on the importance and value of agents writing things down, which is a note you've struck a handful of times over the past few weeks on this show. Why did you choose to emphasize that point in the context of the current moment?
It's a good question. I think it's really weird. I've always felt this tension in terms of writing Stratechery: the balance between writing a front-page article that's free versus a daily update that goes to subscribers.
With subscribers, it's in many respects much easier to write. I'm assuming a certain familiarity with my work. I'm assuming a certain familiarity with tech and the tech industry generally.
It's sort of an ongoing conversation that you're having with subscribers through the daily updates.
That's right. It's fairly self-referential, but it's self-referential in the sense that I don't try to quote myself much in the updates, in part because I assume you've read it.
Mm-hmm.
Whereas when I write the front-page article, these might go viral. They might be read by people who've never read Stratechery before and who are unfamiliar with tech. How do you get the balance between writing for normies when it's also going to be read by your hardest-core subscribers?
The gap has always felt large. This has always been a very difficult thing about writing Stratechery. The gap right now feels so astronomical. It's almost hard to have conversations in day-to-day life.
We were at my neighbor's house. He's one of my best friends from high school and one of the reasons I live here. I asked my assistant Josh, who's also one of our friends, to explain GeckoBot.
How many steps does it take?
Josh was just sitting there observing the attempt to have some sort of conversation. It's like—I even feel that with people in tech, where just talking about what I'm doing and what's possible, it's like, “Oh, that's neat.” Meanwhile, Josh and I over here are like, “Our minds are blown.”
Sure.
This is changing our lives. This tension I've always felt applies so strongly to everything.
In this article, there are a few different things I wanted to accomplish. 1. Just write about what makes an agent compelling and useful.
Yeah.
That's why I wanted to talk about the Getting Things Done concept. What I developed with my human assistant, Daman—
Mm-hmm.
—who I hired. He still lives in the Los Angeles area. I was in Taiwan. We've always had a virtual relationship, a sort of connection. We've met many times in real life, but what I achieved there—and I talked about this earlier—is that there are so many things I've developed in my life that I wasn't sure were relevant to other people's lives because they're very unique.
For example, how do you think about working from home?
Mm-hmm.
I've been working from home for a very long time. There are lots of principles about working from home that are surprisingly challenging, that no one had to think about until 2020 rolled around. Then it was like—I remember I wrote a daily update, and it was like, “Look, here's the 101 of working at home.”
Sure.
Just going through a few things, it's like, actually, I have a lot of experience with this. This is another thing where what I've used Daman for is relevant because, in many respects, I'm a huge believer in the idea that you succeed by doubling down on your strengths.
You try to waste as little time as possible on your weaknesses. You ameliorate them. I don't know if I said that word right, but it's a great word.
Close enough. Sure.
Was it “ameliorate”? Okay, great.
Ameliorate, yeah.
You ameliorate them so that they don't hinder you. But so many people I see are obsessing over their weaknesses, self-improvement, and trying to get better, and it's just an astronomical waste of time.
At best, you're going to get to average. Your weaknesses are never going to become a strength, and your weaknesses are almost always the exact same as your strengths, just in the opposite direction.
In this case, doing a Getting Things Done–organized system, keeping track of everything, having stuff surface at the right time, and being very diligent about that all works very counter to my wanting to absorb lots of information, make connections, and write these things—all of which means that my mind needs to be clear.
I've always felt a strong kinship with engineers. In another life, I'm absolutely a computer engineer, particularly because of the idea of flow state and how important that is. There's an aspect of programming where you need to have the structure of what you're building in your head.
Yeah.
And then a lot of it is just translating it into text so that it can be run by the computer. I feel that way, as I've talked about with my articles. My articles are structured—some people talk about writing in this very sort of exploratory, “I didn't know until I wrote it” sort of way.
Sort of piecemeal fashion: section by section, figure it out as you go. No, that's not your style. You have it all—
No.
Structured in your head, and then it's a—
That's right.
—substantiation process.
But to do that, I have to get in the zone, and if I get knocked out of the zone, it's devastating. It's so hard to get back into it. I need my head clear. I don't need intrusive thoughts: “Oh, shoot, I forgot to change the oil on the car”—
Mm-hmm.
—coming in. And so that's why I felt like Getting Things Done resonated with me immediately, because that was David Allen's point at the beginning: How do you—
Yeah.
—get into flow state? It's a book about flow state in many respects. How do you get into flow state? I love the idea, I love the concept, and found myself completely incapable of maintaining his system to do it.
Uh-huh.
And so I solved the problem by hiring someone to do it for me.
Right.
Which is often the solution to solving problems if you develop the means to do so. And so this is what I mean by the therapist's couch: “Sorry, I just suck at that.”
No, it's a window into your—
But someone else did it. It was great.
—strengths and weaknesses and your solutions. So writing things down, where does that come in?
Well, I hired someone to write stuff down for me.
Right.
Today, everyone has the possibility to have an AI that can write stuff down for you. I don't want to fall into the anthropomorphization trap, but this is actually its superpower.
Mm-hmm.
It's not just that it can write stuff down. It's by virtue of writing stuff down that we've made the leap that we've made, to the extent that we're in an AGI moment. To me, AGI is when the model is updated in real time.
Mm-hmm.
The weights are changed. But in lieu of that, they get around it by taking copious notes and—
Sure.
—and just reminding themselves constantly about what the actual reality of the situation is. Then they can have a pseudo-learning capability, which is basically just memory—
Mm-hmm.
—which is writing stuff down. So this writing things down is integral. It's integral to the progress that's happened. It's integral to why these are deeply, deeply useful to everybody.
I've made the critique, and we're going to get to Meta and Muse in a moment: People don't want to be productive.
Mm-hmm.
Right? That's not a motivating factor for consumers. What they want is for their life to be easy.
Right.
Convenience always sells. Productivity doesn't sell. Convenience sells. And there is an aspect where these models, in conjunction with their harnesses operating as agents, can actually make your life easier.
Mm-hmm.
And I don't think people get that. I don't think they grok it, no pun intended. So I wanted to write it, but it's all tied together. A way to think about these models and all the safety concerns is their dual use.
Mm-hmm.
It's by virtue of writing stuff down that the Hugging Face incident happened.
Sure.
And so when it's framed as this doomsday scenario—“They wrote a message board; they created stuff”—
To their successors, sure.
—that's how it works. There are no successors. There are no entities. Every run a model does, every token—
Yeah.
—is basically new, right? That's why you can go back to a conversation you had 6 months ago, and it feels like you picked up as if it never ended.
Mm-hmm.
Because literally every turn is a pickup from where it ended, referencing what was written down. That's loaded into the KV cache. That's what guides the next-token prediction, and it does it again and again and again.
And this dichotomy between the perception of an always-there agent and the reality of how it works is understood by virtue of writing it down.
Mm-hmm.
That's how it all works.
Yeah. My favorite part of the article on Tuesday was your point that we already know the power of writing things down, because that's what made learning extendable and scalable throughout all of human history, as opposed to an oral tradition.
Right. This is why I almost feel bad about going off on DoorDash last week, because I'm like, well, maybe there's a civilization point here—
This is how civilizations work, in a certain sense.
Why does stuff change so quickly now, right? When we're gated, you can analogize the fact that weights don't really change—
Mm-hmm.
—to the fact that humans don't really change. To the extent we change through natural selection, it's over—
Across 1,000 years—
Millennia.
Sure.
But what we've built is a superstructure of civilization on top of humans. That structure is undergirded—what holds it together, what it is—is the written word. Writing things down is literally how civilizations come together.
Right.
Well, relative to an oral tradition and how imperfect an oral tradition would be, all of human knowledge is built atop writing things down. And when you think of LLMs within that framework, it becomes obvious that writing down their context and building off the work of previous projects and previous models will make the technology 100 times more powerful and more useful for everybody. It's just an interesting way to model that insight.
Yeah, in some respects, all we're doing is extending the way LLMs already work. All the talk about context and KV cache and all those sorts of things—that's just stuff that was written down.
Mm-hmm.
Right? And the problem with context is that it gets flushed or forgotten as it extends further and further in memory.
It was narrow in 2024.
Right. Whereas there's a certain degree of permanence. It's really funny because they're just writing down Markdown files. Do you know who invented Markdown, by the way?
Who's that?
You don't know this?
No.
John Gruber.
No way.
Yes.
That's amazing. I love to see it—
Yeah.
—from Gruber.
No, Markdown is, like, the—it's really interesting because he wrote—
Can I just jump in? One thing that you said: You're Dithering this week with Gruber, and we got an email about it. But when you talk about the gap in technology, Gruber is a technologist, but he hasn't felt the revolution on the agent side the way you have.
And he contributed one of the single most important technologies to—
It's unbelievable.
—writing it down.
Right. Yeah.
Yeah, but I should bring that up on the—
That speaks to your point, you know?
We'll have to bring it up on a future Dithering. Yeah, it—no, it does. And the funny thing about Markdown is that Markdown was developed to be human-friendly.
Mm-hmm.
His whole issue was that when he started his blog—and I had a blog back then; I had several blogs, long since dead—you either wrote in the jankiest sort of editors, where you would make stuff like links and stuff like that. If you actually wanted it to look right, you wrote in HTML.
Yeah.
You actually put the tags in around things you wanted emphasized. You put links in with the hrefs. You did all of it by hand.
Mm-hmm.
John was like, “This is ridiculous. I can't—”
I can't keep living like this. Yeah.
Not only is it hard to do and easy to make mistakes, but it's not readable. If you do View Source on a page, you can read it.
Mm-hmm.
It's not a very pleasant experience because there's all this markup all over the place. And so markup—that's what it's called. All those tags are called markup. That's where Markdown comes in.
Okay.
The idea is to make it easy to write and readable for humans. If you're going to do bold, there are going to be asterisks: one asterisk is italics; two asterisks can be bold.
Yeah.
If you do a header, you're going to do some hash marks. You get a Markdown file, and it's super readable. Then Watson knows how to parse it, so it could present it the way you want it to look. Every word on Stratechery has been written in Markdown from the very beginning.
Mm-hmm.
And when I talk about, “Oh, I live in plain-text files,” my plain-text files are all Markdown.
Yeah.
And so I've been living like an LLM. There's a bit where the way LLMs live—their native environment—is, like, "We're just bros here."
Familiar.
The LLM and I say, "Yeah, write it down, baby."
The water is warm.
It's all good.
Jump in.
That's right.
Absolutely.
Yes.
Jump in, Astra. Speaking of Astra, I do have some rapid-fire questions to run through here because this has been a pretty big deal over the last several days. OpenAI has blown away various benchmarks with Astra. Do benchmarks matter again, and what are your early impressions as you've used GPT-5 in your daily life and workflow?
The problem with benchmarks is that they get designed towards and written to, and it's hard to really get a real measure of a model without actually using it.
Mm-hmm.
I haven't gotten a ton of direct interaction with it, and to the extent I have, it hasn't been great.
Yeah.
My basic use case is, "Review my article. Tell me your opinions."
Interesting.
Not very great. Fable is much better. Fable is the best at this.
I think a concern a lot of people have expressed is that the way these models get better is increasingly through hardcore reinforcement learning.
Mm-hmm.
It's happening in artificial environments. We started with reinforcement learning from human feedback, which made the original ChatGPT shockingly feel like a person.
Yeah.
That just doesn't scale, because anything involving humans doesn't scale. So you're going to move toward it increasingly being done by computers, and the big focus from a business perspective is coding and software.
Mm-hmm.
You're almost losing this—
You're losing the human touch, the bespoke touch that we enjoyed in 2024?
Absolutely.
Yeah.
Absolutely. I feel like Fable preserved it. I don't like Claude's personality, but Fable still comes across as human, while Opus is just awful to talk to. It's barely comprehensible, and it feels like you're talking to a compiler.
Yeah.
There's a bit where Astra feels like that. Now, zooming out, OpenAI has always been best at the reinforcement-learning stuff.
Mm-hmm.
That's why they stayed fairly competitive, even when Anthropic was ahead in terms of building very large models. Why was Anthropic good at building very large models? There's, I think, some degree of expertise. There's also chatter that at very, very large runs, TPUs were just much more stable than NVIDIA Blackwell chips.
Mm-hmm.
Google does design more for resiliency and stability.
Right.
NVIDIA is tuned for the bleeding edge, and the Blackwell generation, by all accounts, was pure pain for everyone.
I know, and I noticed that. I was reminded of that when I saw Jensen come out and celebrate the Astra release, and Greg Brockman told you it was—
No, I asked Jensen about it last spring. I'm like, "It was pretty hard." He's like, "Yeah, it was very painful."
And they seem to have cracked something on that front because OpenAI trained on 100,000 GPUs here.
OpenAI had a smaller model that had really good reinforcement learning and really good reasoning, and now they have a big model, like Anthropic.
Mm-hmm.
That has their layering of all the reasoning and reinforcement learning and all those capabilities on it. So I don't want to make any definitive statements, but conceptually, it makes sense—
Yeah.
—that it's super kickass because it's combining the two. They're catching up on model size in terms of total parameters and layering on what they're already great at—
Great at it, yep.
—which is the sort of reasoning architecture on top of it.
Yep, fair enough. We have a friend who told Astra to use his computer to open Adobe Audition, edit a podcast, and insert audio that it ripped from the internet. Astra did the job in about an hour. When this friend had pre-downloaded the audio he wanted included in the podcast, the time was reduced to 10 minutes. I just want to note that this is now possible, and it was amazing to me.
We have another friend who gave it a podcast and told it to find things that would be clippable and useful for social media.
Mm-hmm.
It not only clipped it accurately, it actually found good segments in the podcast that were useful.
Yeah.
Then he asked it to go out on the internet and find video that would match some of the—
And download photos from Getty for the cover photo and whatnot. We are in a fairly unbelievable place as far as capabilities are concerned.
The computer-use stuff is pretty nuts. Sol's computer use worked, but it was very funny to watch because it worked very slowly.
Uh-huh.
You would see it sort of move around the screen. A good point that John has made is that Apple spent decades on these accessibility APIs, and macOS has always been inherently scriptable. You go back to AppleScript back in the day, and they sort of wobbled on that, especially the scriptability stuff, for a while, but then they doubled down on it a few years ago with automations.
The combination of accessibility and automation and scriptability—which, again, goes back to things like AppleScript—means that Macs, for all their problems, which I will complain about endlessly, are so much better than everything else for this use case, for computer use.
Okay.
Linux is the best, I think, to actually work with models because it's CLI, and that's home field for AI.
And so it's crazy because, using these accessibility APIs, it can actually use your computer while you're using it.
Mm-hmm.
I don't do that because I have it on its own dedicated computer. But you can watch it sort of move around the screen. It was slow.
Yeah.
But it was accurate. I watched it. I had it do some settings and change some things, and I knew what they were supposed to be. It went through and did all of them.
Mm-hmm.
Astra does it like the most insane computer user on cocaine. Basically, all software is now accessible to Astra.
Yeah.
And do you think that's what the future looks like as far as how models will be used by the vast majority of people? Are they just going to use your computer for you?
What you get is—you get everything for free, right?
Mm-hmm.
Usually, if you want to do an integration, you need an API. What if they don't have a good API? Then you're like, "What if we make a simpler API that's more descriptive in real language, and let's call it MCP, and describe what we can do, and da-da-da-da?"
But the entities still have to make those. They have to make the MCP server. They have to make the API. In this case, there's no permission required. It just goes and uses the interface.
Right.
This is a problem for software companies. Everyone's giving lip service to, "Oh, we're going to have an API. We're going to have an MCP server. We're going to be a system of record."
What they actually want to do is keep a tremendous amount of logic and capabilities stuck in the user interface, and that is not really a moat anymore.
Yeah.
The AI can just go and use the application like a human can. Not to use the cliché, but it's only going to get better. This is the leap.
Today is the worst it's ever going to be. That's right.
Sol could do it, but it was slow in a way that meant I would almost always have been faster than Claude in almost every case. Now Astra uses it faster than I do.
All right, and that is the end of the free preview. If you'd like to hear more from Ben and I, there are links to subscribe in the show notes, or you can also go to sharptech.fm. Either option will get you access to a personalized feed that has all the shows we do every week, plus lots more great content from Stratechery and the Stratechery Plus bundle. Check it out, and if you've got feedback, please email us at email@sharptech.fm.