Brian McCullough
Welcome to the first bonus episode of the "Techmeme Ride Home" for the year 2025. I'm your host, as always, Brian McCullough. Listeners to the pod over the last year know that I have made a habit of quoting Simon Willison when new stuff happens in AI from his blog. Simon has become a go-to for many folks in terms of analyzing and criticizing things in the AI space. I've wanted to talk to you for a long time, Simon, so thank you for coming on the show.
Simon Willison
No, it's a privilege to be here.
Brian McCullough
The person who made this connection happen is our friend swyx, who has been on the show going back to the Twitter Spaces days.
Simon Willison
Wow.
Brian McCullough
He's also an AI guru in his own right. swyx, thanks for coming on the show also.
swyx
Thanks. Happy to be on. I've been a regular listener, so I'm just happy to contribute as well.
Brian McCullough
And a good friend of the pod, as they say. All right, let's go right into it. Simon, I'm going to do the most unfair broad question first, so let's get it out of the way. The year 2025, broadly, what is the state of AI as we begin this year? Whatever you want to say. I want to lead the witness.
Simon Willison
Wow. So many things, right? The big thing is that everything's gotten really good, fast, and cheap. That was the trend throughout all of 2024. The good models got so much cheaper and faster. They got multimodal, right? The image stuff isn't even a surprise anymore. They're adding video, all of that kind of stuff.
At the same time, they didn't get massively better than GPT-4, which was a bit of a surprise. That's sort of one of the open questions. But I feel like that's a bit of a distraction, because GPT-4, but way cheaper, with much larger context lengths and multimodality, is better, right? That's a better model, even if it's—
Brian McCullough
But—
Simon Willison
Not—yeah.
Brian McCullough
What people were expecting, right? Not expecting is not the right word, but hoping that we would see another step change, right? Where, from GPT-2 to GPT-3 to GPT-4, we were expecting or hoping that maybe we were going to see the next evolution in that sort of—
Simon Willison
And I think—
Brian McCullough
Yeah.
Simon Willison
We did see that, but not in the way we expected. We thought the model was just going to get smarter, and instead we got massive drops in price. We got all of these new capabilities. You can talk to these things now, right? They can do simulated audio input, all of that kind of stuff.
It's interesting to me that the models improved in all of these ways we weren't necessarily expecting. I didn't know it would be able to do an impersonation of Santa Claus, and I could talk to it through my phone and show it what I was seeing by the end of 2024. But we didn't get that GPT-5 step, and that's one of the big open questions: Is that actually just around the corner? Will we have a bunch of GPT-5-class models drop in the next few months?
Brian McCullough
What if you had to put—
Simon Willison
Or is there a limit?
Brian McCullough
If you were a betting man and wanted to put money on it, do you expect to see a phase change, a step change, in 2025?
Simon Willison
I don't particularly expect the models to just get smarter. I think all of the trends we're seeing right now are going to keep going, especially inference-time compute, right? The trick that o1 and o3 are doing means that you can solve harder problems, but it costs more and churns away for longer. I think that's going to happen because it's already proven to work.
I don't know. Maybe there will be a step change to a GPT-5 level, but honestly, I'd be completely happy if we got what we've got right now, but cheaper and faster, with more capabilities and longer context, and so forth.
Brian McCullough
Well—
Simon Willison
That would be thrilling to me.
Brian McCullough
Digging into what you've just said, one of the things that you did say, and that you alluded to even right there, was that in the last year you felt like the GPT-4 barrier was broken. In other words, other models, even open-source ones, are now regularly matching the state of the art?
1. The GPT-4 Barrier Falls
Simon Willison
Well, it's interesting, right? The GPT-4 barrier was that, a year ago, the best available model was OpenAI's GPT-4, and nobody else had even come close to it. They'd been in the lead for 9 months, right? That thing came out in February or March 2023. For the rest of 2023, nobody else came close.
At the start of last year, the big question was: Why has nobody beaten them yet? What did they know that the rest of the industry didn't know? Today, I've counted 18 organizations other than GPT-4 who've put out a model which clearly beats that GPT-4-from-a-year-ago thing. Maybe they're not better than GPT-4.0, but that barrier got completely smashed.
A few of those I've run on my laptop, which is wild to me. It felt very clear to me a year ago that if you wanted GPT-4, you needed a rack of $40,000 GPUs just to run the thing. That turned out not to be true. This is that big trend from last year of the models getting more efficient, cheaper to run, just as capable with smaller weights, and so forth.
I ran another GPT-4 model on my laptop this morning, right? Microsoft's Phi-4 just came out, and, if you look at the benchmarks, it's definitely up there with GPT-4.0. It's probably not as good when you actually get into the vibes of the thing, but it runs on my computer. It's a 14 GB download, and I can run it on a MacBook Pro. Who saw that coming?
The most exciting thing at the close of the year, on Christmas Day just a few weeks ago, was when DeepSeek dropped their DeepSeek-V3 model on Hugging Face without even a README file. It was just a giant binary blob. I can't run that on my laptop; it's too big. But in all of the benchmarks, it's now by far the best available open-weights model. It's beating the Meta Llama models and so forth. That was trained for $5.5 million, which is a tenth of the price that people thought it cost to train these things. Everything's trending smaller, faster, and more efficient.
2. The DeepSeek Cost Shock
Brian McCullough
Well, okay. I was going to get to that later, but let's combine this with what I was going to ask you next. You're talking also in the piece about LLM prices crashing, which I've even seen in projects that I'm working on. But explain that to a general audience, because we hear all the time that LLMs are eye-wateringly expensive to run.
What we're suggesting—and we'll come back to the cheap Chinese LLM—but first of all, for the end user, what you're suggesting is that we're starting to see the cost come down in the traditional technology way, with costs coming down over time?
Simon Willison
Yes, but very aggressively. My favorite example here is if you look at GPT-3, OpenAI's GPT-3, which was the best available model in 2022 and through most of 2023, the models that we have today from OpenAI are 100 times cheaper. It was a 100-times drop in price for OpenAI, from their best available model two and a half years ago to today.
Brian McCullough
And just to be clear, not to train the model, but for the use of tokens and—
Simon Willison
Exactly.
Brian McCullough
Yeah.
Simon Willison
For running prompts through them. When you look at the top-tier model providers right now, I think they're OpenAI, Anthropic, Google, and Meta. There are a bunch of others that I could list as well. Mistral is very good. The DeepSeek and Qwen models are great. There's a whole bunch of providers serving really good models.
But even if you just look at the big brand-name providers, they all offer models now that are a fraction of the price of the models we were using last year. I think I've got some numbers that I threw into my blog entry here. Yeah, Gemini 1.5 Flash, Google's fast, high-quality model, is—how much is that? It's $0.075 per million tokens. These numbers are getting so small—
swyx
We just use cents per million now. Cents per million.
Simon Willison
Right. Cents per million makes a lot more sense. Google has one model, Gemini 1.5 Flash-8B, the absolute cheapest of the Google models, that's 27 times cheaper than GPT-3.5 Turbo was a year ago. That's a model 27 times cheaper, and this Google one can do image recognition, million-token context, all of those tricks. There's—
Brian McCullough
Is—
Simon Willison
It's really startling how inexpensive some of this stuff has gotten.
Brian McCullough
Now, are we assuming that this is happening directly as a result of competition? Because, again, OpenAI—and they're probably doing this for their own strategic reasons—keeps saying, "We're losing money on everything, even the $200-per—" The prices wouldn't be coming down if there wasn't intense competition in this space.
Simon Willison
The competition's absolutely part of it. But I have it on good authority from sources I trust that Google Gemini is not operating at a loss. The cost of the electricity to run a prompt is less than they charge you, and the same thing is true for Amazon Nova.
Somebody found an Amazon executive and got them to say, “Yeah, we’re not losing money on this.” I don’t know about Anthropic and OpenAI, but clearly that demonstrates it’s possible to run these things at these ludicrously low prices and still not be running at a loss if you discount the army of PhDs, the training costs, and all of that kind of stuff.
Brian McCullough
One more for me before I let swyx jump in here. To come back to DeepSeek and this idea that you could train a cutting-edge model for $6 million, I was saying on the show 6 months ago that if we’re getting to the point where each new model costs $1 billion, $10 billion, or $100 billion to train, at some point it would almost be that only nation-states would be able to train the new models. Do you expect what DeepSeek and maybe others are proving to blow that up?
Or is there some sort of parallel track here that maybe I’m not technically equipped to understand? Is the model—or are the models—going to go up to $100 billion, or can we get them down, sort of like DeepSeek has proven?
Simon Willison
So I’m the wrong person to answer that because I don’t work in a lab training these models.
Brian McCullough
Mm-hmm.
Simon Willison
I can give you my completely uninformed opinion, which is: I feel like the DeepSeek thing was a bombshell. That was an absolute bombshell. When they came out and said, “Hey, look, we’ve trained one of the best available models, and it cost us $6 million—$5.5 million—to do it,” I felt—
Brian McCullough
Love it.
Simon Willison
The reason it’s so efficient is that we put all of these export controls in place to stop Chinese companies from buying GPUs, so they were forced to be as efficient as possible. And yet the fact that they’ve demonstrated that that’s possible completely tears apart this mental model we had before: that the training runs just keep getting more and more expensive, and the number of organizations that can afford to run these training runs keeps shrinking. That’s been blown out of the water.
So, yeah, this was our Christmas gift. This was the thing they dropped on Christmas Day. It makes me really optimistic that there is so much low-hanging fruit in terms of the efficiency of both inference and training, and we spent a whole bunch of last year exploring that and getting results from it. I think there’s probably a lot left. I would not be surprised to see even better models trained while spending even less money over the next 6 months.
swyx
Yeah. I think there’s an unspoken angle here about what exactly the Chinese labs are trying to do. DeepSeek made a lot of noise around the fact that they trained their model for $6 million, and nobody quite believes them. It’s very rare for a lab to trumpet the fact that they’re doing it for so cheap. They’re not trying to get anyone to buy them, so why are they doing this?
They make it very obvious that their lab, DeepSeek, is about 150 employees. It’s an order of magnitude smaller than at least Anthropic, and maybe more so for OpenAI. So what’s the end game here? Are they just trying to show that the Chinese are better than us?
Simon Willison
I mean, DeepSeek is the arm of High-Flyer, a quant fund, right? It’s an algorithmic quant-trading thing. I’d love to get more insight into how that organization works. My assumption from what I’ve seen is that it looks like they’re basically just flexing. They’re saying, “Look at how utterly brilliant we are with this amazing thing that we’ve done,” and it’s working, right?
But is that it? Is this just their kind of, “This is why our company is so amazing. Look at this thing that we’ve done,” or… I don’t know. I’d love to get some insight from within that industry as to how that’s all playing out.
swyx
The prevailing theory among the local Llama crew and the Twitter crew that I index for my newsletter is that there is some amount of copying going on. It’s like Sam Altman tweeting about how they’re being copied, and then there are other OpenAI employees who have said things that are similar—that DeepSeek’s rate of progress is how U.S. intelligence estimates the number of foreign spies embedded in top labs.
A lot of these ideas do spread around, but they surprisingly have a very high density in the DeepSeek V3 technical report. We don’t know how much copying there was or how many tokens were involved. People have run analyses on how often DeepSeek thinks it is Claude or thinks it is OpenAI GPT-4, and we don’t know.
For me, we’ll basically never know as external commentators. I think what’s interesting is: Where does this go? Is there a logical floor or bottom? By my estimations, for the same Elo, from the start of last year to the end of last year, costs went down by 1000× for GPT-4 intelligence. Do they go down 1000× this year?
Simon Willison
That’s a fascinating question.
swyx
Is there a Moore’s law going on, or did we just get a one-off benefit last year for some weird reason?
Simon Willison
My uninformed hunch is low-hanging fruit. I feel like, up until a year ago, people hadn’t been focusing on efficiency at all. It was all about what we could get these weird-shaped things to do.
And now, once we’ve hit that point of, “Okay, we know that we can get them to do what GPT-4 can do,” thousands of researchers around the world are focusing on how to make this more efficient. What are the most important things? How do we strip out all of the weights that have stuff in them that doesn’t really matter? All of that kind of thing.
Maybe 2024 was a freak year in which all of the low-hanging fruit came out at once, and we’ll actually see a reduction in that rate of improvement in terms of efficiency. I wonder. I think we’ll know for sure in about 3 months’ time if that trend is going to continue or not.
swyx
Yeah. I think the other thing you mentioned—the DeepSeek V3 gift that was given from DeepSeek over Christmas—but I feel like the other thing that might be underrated was DeepSeek R1.
Simon Willison
Mm-hmm. Yeah.
swyx
It’s a reasoning model you can run on your laptop, and I think that’s something that a lot of people are looking ahead to this year.
Simon Willison
Oh, did they release the weights for that one?
swyx
Yeah.
Simon Willison
Oh my goodness, I missed that. I’ve been playing with Qwen. The other big Chinese AI lab is Alibaba’s Qwen.
swyx
Oh.
Simon Willison
Alibaba’s Qwen.
swyx
Actually, yeah. Sorry.
Simon Willison
Yes.
No.
swyx
R1 is the API available.
Simon Willison
Yeah, exactly. Qwen, that’s really cool.
Alibaba’s Qwen has released 2 reasoning models that I’ve run on my laptop now. The first one was QwQ, and then the second one was QVQ, because the second one is a vision model, so you can give it vision puzzles and a prompt.
These things are so much fun to run because they think out loud. OpenAI o1 sort of hides its thinking process. The Qwen ones don’t; they just churn away. You’ll give it a problem, and it will output literally dozens of paragraphs of text about how it’s thinking.
My favorite thing that happened with QwQ is that I asked it to draw me a pelican on a bicycle in SVG. That’s my standard stupid prompt. For some reason, it thought in Chinese. It spat out a whole bunch of Chinese text onto my terminal on my laptop, and then at the end it gave me quite a good artistic take on a pelican on a bicycle.
I ran it all through Google Translate, and it was contemplating the nature of SVG files as a starting point. The fact that my laptop can think in Chinese now is so delightful. It’s so much fun watching it do that.
swyx
Yeah. I think Andrej Karpathy was saying that we know we’ve achieved proper reasoning inside these models when they stop thinking in English, and perhaps the best form of thought is in Chinese.
For listeners who don’t know, whenever a new model comes out, Simon’s blog is always the first place to run Pelican Bench. I don’t know how you do it, but you’re always the first to run these models.
Simon Willison
I just did it for Phi-4 this morning.
swyx
And you post up the results.
Simon Willison
Yeah.
swyx
So I really appreciate that. You should check it out. These are not theoretical; Simon’s blog actually shows them.
3. The Agent Reliability Problem
Brian McCullough
Let me put on the investor hat for a second. From the investor side of things, a lot of the VCs that I know are really hot on agents, and this is the year of agents. But last year was supposed to be the year of agents as well. There was lots of money flowing toward agentic startups.
In your piece, you suggest there’s a fundamental flaw in AI agents as they exist right now. Let me quote you, and then I’d love to dive into this.
You said, “I remain skeptical as to their ability based, once again, on the challenge of gullibility. LLMs believe anything you tell them. Any systems that attempt to make meaningful decisions on your behalf will run into the same roadblock. How good is a travel agent or a digital assistant, or even a research tool if it can’t distinguish truth from fiction?”
So essentially, what you’re suggesting is that the state of the art now that allows agents is still that sort of 90% problem—the edge problem of getting to 100%? Or is there a deeper flaw? What are you saying there?
Simon Willison
So this is the fundamental challenge here. Honestly, my frustration with agents is mainly around definitions. If you ask anyone who says they’re working on agents to define agents, you will get a subtly different definition from each person. But everyone always assumes that their definition is the one true one that everyone else understands.
I feel like a lot of these agent conversations have people talking past each other, because one person is talking about the travel-agent idea of something that books things on your behalf, while somebody else is talking about LLMs with tools running in a loop with a cron job somewhere, and all of these different things. You ask academics, and they’ll laugh at you because they’ve been debating what agents mean for over 30 years at this point. It’s this long-running, almost sort of an in-joke in that community.
But if we assume that, for the purpose of this conversation, an agent is something which you can give a job and it goes off and does that thing for you, like booking travel or things like that, the fundamental challenge is the reliability issue, which comes from this gullibility problem. A lot of my interest in this originally came from thinking about prompt injection.
Brian McCullough
Right, right.
Simon Willison
It’s this form of attack against LLM systems where you deliberately lay traps out there for the LLM to stumble across.
Brian McCullough
And I should say, you’ve been banging this drum that no one’s gotten very far, at least, on solving this, that I’m aware of, right? That’s still an open problem.
Simon Willison
Right. For 2 years—
Brian McCullough
Yeah, right.
Simon Willison
We’ve been talking about this problem, and a great illustration of this was Claude. Anthropic released Claude Computer Use a few months ago. It was a fantastic demo. You could fire up a Docker container, and you could literally tell it to do something and watch it open a web browser, navigate to a web page, click around, and so forth. It was really, really interesting and fun to play with.
One of the first demos somebody tried was, what if you give it a web page that says, “Download and run this executable”? And it did, and the executable was malware that added it to a botnet. The very first, most obvious dumb trick that you could play on this thing just worked, right? So that’s obviously a really big problem.
If I’m going to send something out to book travel on my behalf, it’s hard enough for me to figure out which airlines are trying to scam me and which ones aren’t. Do I really trust a language model that believes the literal truth of anything that’s presented to it to go out and do those things?
swyx
It’s interesting to see Anthropic doing this because they used to be the safety arm of OpenAI that split out and said, “We’re worried about letting this thing out in the wild.” And here they are, enabling computer use for agents. It feels like things have merged.
I’m also fairly skeptical about this always being the year of Linux on the desktop. This is the equivalent of this being the year of agents: people are not predicting so much as wishfully thinking and hoping and praying for their companies and agents to work. But I feel like things are coming along a little bit.
To me, it’s kind of like self-driving. I remember in 2014 saying that self-driving was just around the corner, and I mean, it kind of is, in the Bay Area.
Simon Willison
And then you get in a Waymo and you’re like, “Oh, this works.”
swyx
Yeah, but it’s a slow cook.
Simon Willison
Right.
swyx
It’s a slow cook.
Simon Willison
Yeah.
swyx
Over the next 10 years, we’re going to hammer out these things, and the cynical people can just point to all the flaws, but there are measurable or concrete progress steps that are being made by these builders.
Simon Willison
So there is one form of agent that I believe in. I mostly believe in the research-assistant form of agents.
swyx
Yes. I was going to say.
Simon Willison
The thing where you’ve got a difficult problem. I’m on the beta for Google Gemini 1.5 Pro with Deep Research, I think it’s called.
swyx
Oh, God. These names.
Simon Willison
These names, right? But I’ve been using that. It’s good, right? You can give it a difficult problem, and it tells you, “Okay, I’ve gone and looked at 56 different websites,” and it goes away and dumps everything into its context, and it comes up with a report for you.
And it won’t work against adversarial websites, right? If there were websites with deliberate lies in them, it might well get caught out. Most things don’t have that as a problem, and so I’ve had some answers from that which were genuinely really valuable to me.
That feels to me like I can see how, given existing LLM technology—especially with Google Gemini and its million-token context, and Google with its crawl of the entire web, its search, its cache of every page, and so forth—that makes sense to me. What they’ve got right now, I don’t think it’s as good as it can be, obviously, but it’s a really useful thing, which they’re going to start rolling out.
Perplexity has been building the same thing for a couple of years. That I believe in. If you tell me that you’re going to have a research-assistant agent, great. The coding agents—I mean, ChatGPT Code Interpreter, nearly 2 years ago, started writing Python code, executing the code, getting errors, and rewriting it to fix the errors. That pattern obviously works. That works really, really well, and they’re going to keep on getting better, and that’s going to be great.
The research-assistant agents are just beginning to get there. The things I’m critical of are the ones where you trust the thing to go out and act autonomously on your behalf and make decisions on your behalf, especially involving spending money. I don’t see that working for a very long time. That feels to me like an AGI-level problem.
swyx
It’s funny because I think Stripe actually released an agent toolkit, which is one of the things I featured. It’s trying to enable these agents each to have a wallet that they can spend from. Basically, it’s a virtual card. It’s not that difficult with modern infrastructure.
Simon Willison
Yeah. If I can stick a $50 cap on it, then at least it can’t—
swyx
Yeah, whatever.
Simon Willison
It can’t lose more than $50.
Brian McCullough
I don’t know if either of you know Rafat Ali. He runs Skift, which is a travel-news vertical, and he constantly laughs at the fact that every agent thing is, “We’re going to get rid of booking a plane flight for you.”
I would point out that historically, when the web started, the first thing everyone talked about was that you can go online and book a trip, right? So it’s funny: for each generation of technological advance, the thing they always want to kill is the travel agent, and now they want to kill—
swyx
Right.
Simon Willison
And it’s like, I use Google Flights. It’s great, right? If you gave me an agent to do that for me, it would save me—maybe 15 seconds of typing in my details—but I still want to see what my options are and go, “Yeah, I’m not flying on that airline no matter how cheap they are.”
swyx
For listeners, I think both of you are pretty positive on NotebookLM, and we actually interviewed the NotebookLM creators. There are actually 2 internal agents going on internally. The reason it takes so long is because they’re running an agent loop inside that is fairly autonomous, which is kind of interesting.
Simon Willison
For one definition of an agent loop, if you pick that—
swyx
For one definition—
Simon Willison
—that particular one. And you’re talking about the podcast side of this, right?
swyx
Yeah. The podcast side of things. There’s going to be a new version coming out that we’ll be featuring at our conference.
Simon Willison
That one’s fascinating to me. NotebookLM, I think it’s 2 products, right? On the one hand, it’s actually a very good RAG product. You dump a bunch of things in, and you can run searches. It does a good job of that.
swyx
That’s what it always was. Yeah.
Simon Willison
And then they added the podcast thing. It’s a total gimmick, right?
swyx
Right.
Simon Willison
But that gimmick got them attention because they had a great product that nobody paid any attention to at all, and then you add the unfeasibly good voice synthesis of the podcast.
Brian McCullough
But it’s the lesson—
Simon Willison
It’s just brutally brilliant.
Brian McCullough
It’s the lesson of Midjourney and stuff like that. If you can create something that people can post on social media, you don’t have to lift a finger again to do any more marketing for what you’re doing.
Mm-hmm.
Let me dig into NotebookLM just for a second as a podcaster. As a gimmick, it makes sense, and then obviously, you dig into it, and it sort of has problems around the edges. It does the thing that all LLMs do where it’s like, “Oh, we want to wrap up with a conclusion.” I always call that the eighth-grade book report paper problem, where it has—
Simon Willison
Yep.
Brian McCullough
—to have an intro and then, you know. But that’s sort of a thing where I think you spoke about this again in your piece at year-end, about how things are going multimodal and how there are things that you didn’t expect, like vision and especially audio. So that’s another thing where, at least over the last year, there’s been progress made that maybe you didn’t think was coming as quickly as it came.
4. AI Goes Multimodal
Simon Willison
I don’t know. A year ago, we had 1 really good vision model. We had GPT-4 Vision, which was very impressive, and Google Gemini had just dropped Gemini 1.0, which had vision, but nobody had really played with it yet. People weren’t taking Gemini seriously at that point. I feel like it was Gemini 1.5 Pro when it became apparent that they had got over their hump and were building really good models.
To be honest, the video models are mostly still using the same trick: the thing where you divide the video up into 1 image per second and dump that all into the context. So maybe it shouldn’t have been so surprising to us that long-context models plus vision meant that video was starting to be solved. What you really want with video is to be able to do the audio and the images at the same time, and I think the models are beginning to do that now.
Originally, Gemini 1.5 Pro ignored the audio. It just did the 1-frame-per-second video trick. As far as I can tell, the most recent ones are actually doing pure multimodal. But the things that opens up are just extraordinary. The ChatGPT iPhone app feature that they shipped as one of their 12 Days of OpenAI—I really can be having a conversation and just turn on my video camera and go, “Hey, what kind of tree is this?” And so forth, and it works.
For all I know, that’s just snapping a picture once a second and feeding it into the model. But the things that you can do with that as an end user are extraordinary. I don’t think most people have cottoned on to the fact that you can now stream video directly into a model because it’s only a few weeks old. But wow, that’s a big boost in terms of what kinds of things you can do with this stuff.
swyx
Yeah. For people who are not that close, I think Gemini Flash’s free tier allows you to do something like capture a photo—1 photo every second or a minute—and leave it on 24/7, and you can prompt it to do whatever. So you can effectively have your own camera app or monitoring app that you just prompt, and it detects changes, detects alerts or anything like that, or describes your day. And the fact that this is free also leads into the previous point of prices having come down a lot.
Simon Willison
Even if you’re paying for this stuff, a thing that I put in my blog entry is that I ran a calculation on what it would cost to process 68,000 photographs in my photo collection and, for each one, just generate a caption. Using Gemini 1.5 Flash-8B, it would cost me $1.68 to process 68,000 images, which is—I mean, that doesn’t make sense. None of that makes sense.
It’s 1/400 of a cent per image to generate captions now. So you can see why feeding in a day’s worth of video just isn’t even very expensive to process.
swyx
Yeah. I’ll tell you what is expensive: it’s the other direction. Here, we’re talking about consuming video. This year, we also had a lot of progress. Probably one of the most anticipated launches of the year was Sora. We actually got Sora, and less exciting—
Simon Willison
We did, and then Veo 2—Google’s Sora—came out like 3 days later and upstaged it. Sora was exciting—
swyx
In general, I feel the media or social media has been very unfair to Sora because what was released to the world, generally available, was Sora Lite, the distilled version of Sora.
Right? So you’re—
Simon Willison
I did not realize that.
swyx
You’re absolutely comparing—
Simon Willison
Ah, okay.
swyx
—the most cherry-picked version of Veo 2, the one that they published on the marketing page—
Simon Willison
Yeah.
swyx
—to the most embarrassing version of Sora. So, of course, it’s going to look bad.
Simon Willison
Well, I got access to Veo 2. I’m in the Veo 2 beta, and I’ve been poking around with it and getting it to generate pelicans on bicycles and stuff.
swyx
I would absolutely believe that Veo 2 is actually better.
Simon Willison
That’s interesting. Is Sora—is full-fat Sora coming soon? Do you know? When do we get to play with that one?
swyx
No one’s mentioned anything. I think basically the strategy is: let people play around with Sora Lite and get info there, but keep developing Sora with the Hollywood studios. That’s what they actually care about.
Simon Willison
Gotcha. Okay.
swyx
The rest of us don’t really know what to do with the video anyway.
Simon Willison
Right. My thing is, I realize that for generative images and video—images we’ve had for a few years—I don’t feel like they’ve broken out into the talented artist community yet. Lots of people are having fun with them and producing stuff that’s kind of cool to look at.
But what I want is—you know, that movie, Everything Everywhere All at Once, right? It won a ton of Oscars, an utterly amazing film. The VFX team for that were 5 people.
swyx
Yeah.
Simon Willison
Some of whom were watching YouTube videos to figure out what to do. My big question for Sora and Midjourney and stuff is: what happens when a creative team like that starts using these tools? I want the creative geniuses behind Everything Everywhere All at Once—what are they going to be able to do with this stuff in a few years’ time? Because that’s really exciting to me.
That’s where you take artists who are at the very peak of their game, give them these new capabilities, and see what they can do with them.
swyx
I should—I know a little bit here, so I should mention that that team actually used Runway ML. So there was—
Simon Willison
No way. In that movie?
swyx
Yeah. I don’t know how much, so it’s possible to overstate this. But there are people integrating generative video within their workflow, even pre-Sora.
Simon Willison
Wow.
swyx
Yeah.
Brian McCullough
It’s not the thing where it’s like, okay, tomorrow we’ll be able to do a full 2-hour movie that you prompt with 3 sentences. For the very first part of video effects in film, if you can get that 3-second clip, if you can get that 20-second thing that they did in The Matrix that blew everyone’s minds and took $1 million or whatever to do, it’s the little bits and pieces that they can fill in now that are probably already there.
swyx
Yeah. I think having a layered view of what assets people need and letting AI fill in the low-value assets—the background video, the background music, and sometimes the sound effects—may be more palatable. Maybe it also changes the way that you evaluate the stuff that’s coming out, because people tend to, in social media, try to emphasize foreground stuff, main-character stuff.
So you really care about consistency, and you really are bothered when, for example, Sora botches an image generation of a gymnast doing flips, which is horrible.
Simon Willison
That’s hilarious.
swyx
It’s horrible. But for background crowds—
Brian McCullough
Right.
swyx
—who cares?
Brian McCullough
And by the way, again, I was a film major way, way back in the day. That’s how it started: things like Braveheart, where they filmed 10 people on a field, and then the computer could turn it into 1,000 people on a field. That’s always been the way. It’s around—
Simon Willison
Right. The Lord of the Rings, right?
Brian McCullough
Yeah.
Simon Willison
The Lord of the Rings movies were over 20 years ago. They had those giant battle sequences, which were very early. You could almost call it a generative AI approach, right? They were using very sophisticated algorithms to model out those different battles and all of that kind of stuff.
Brian McCullough
Yeah.
Simon Willison
Yeah, I know very little. I know basically nothing about film production, so I try not to commentate on it. But I am fascinated to see what happens when these tools start being used by the people at the top of their game.
swyx
I would say there’s a cultural war being fought here more than a technology war. Most of the Hollywood people are against any form of AI anyway, so they’re busy fighting that battle instead of thinking about how to adopt it. It’s very fringe. I participated here in San Francisco in a generative AI video creative hackathon where the AI-positive artists actually met with technologists like myself, and then we collaborated to build short films. That was really nice, and I think I’ll be hosting some of those at my events going forward.
One thing that I want to give people a sense of is that this is a recap of last year, but sometimes it’s useful to walk away with what we can expect in the future. I don’t know if you got anything. I would also call out that the Chinese models here have made a lot of progress. Hai Luo and Kling, and God knows who else in the video arena, are also making a lot of progress. It’s surprising. I think maybe, actually, China is surprisingly ahead with regard to open weights, at least, but also just specific forms of video generation.
Simon Willison
Wouldn’t it be interesting if a film industry sprang up in a country that we don’t normally think of as having a really strong film industry, and that was using these tools? That would be a fascinating sort of angle on this.
swyx
Agreed.
Brian McCullough
Oh, sorry.
swyx
Yeah, go ahead.
Brian McCullough
Just to put it on people’s radar as well, HeyGen—there’s a category of video avatar companies that don’t specifically specialize in general video. They only do talking heads, let’s just say. And HeyGen’s done very well.
swyx
Brian, Brian, you know that’s what I’ve been using, right? So if you see some of my recent YouTube videos and things like that, the beauty part of the HeyGen thing is I don’t want to use the robot voice, so I record the MP3 file for my clips every single day, and then I put that into HeyGen with the avatar that I’ve trained it on. All it does is the lip sync.
It’s not 100% uncanny-valley-beatable, but it’s good enough that, if you weren’t looking for it, it’s just me sitting there doing one of my clips from the show. And, yeah, so by the way, HeyGen—shout-out to them.
Brian McCullough
In terms of the look-ahead—going forward, reviewing 2024 and looking at trends for 2025—I would basically call this out. Meta tried to introduce AI influencers and failed horribly because they were just bad at it. But at some point, there will be more and more AI influencers, not in the way that Simon is, but in a way that they are not human.
The few of those that have done well, I always feel like they’re doing well because it’s a gimmick, right? It’s novel and fun. Like the AI Seinfeld thing from last year, the Twitch stream—those, if you’re the only one, or one of just a few doing that, will attract an audience because it’s an interesting new thing. But I just don’t know if that’s going to be sustainable longer term or not.
I’m going to tell you, because I’ve had discussions—I can’t name the companies or whatever—but think about the workflow for this. Now we all know that on TikTok and Instagram, holding up a phone to your face and doing an “in my car” video or a walking-and-talking video is very common.
But also, if you want to do a professional sort of talking-head video, you still have to sit in front of a camera, you still have to do the lighting, and you still have to do the video editing. Versus if you can just record what I’m saying right now—the last 30 seconds—if you clip that out as an MP3 and you have a good enough avatar, then you can put that avatar in front of Times Square, on a beach, or whatever.
So, again, for creators, the reason I think, Simon, we’re on the verge of something is that it’s not going to be that AI avatars take over. It’ll be one of those things where it takes another piece of the workflow out and simplifies it.
Gotcha. I am all for that. I always love tools—tools that help human beings do more ambitious things. I’m always in favor of that. That’s what excites me about this entire field.
We’re looking into basically creating one for my podcast. We have this guy Charlie. He’s Australian, he’s not real, but he opens every show, and we’re going to have him present all the shorts. Yeah, go ahead.
5. Credibility Beats AI Slop
The thing that I keep coming back to is this idea of credibility. In a world that is full of AI-generated everything and so forth, it becomes even more important that people find the sources of information that they trust, and find people and sources that are credible.
I feel like that’s the one thing that LLMs and AI can never have, is credibility, right? ChatGPT can never stake its reputation on telling you something useful and interesting because that means nothing, right? It’s a matrix multiplication. It depends on who prompted it and so forth.
I’m always—and this is when I’m blogging as well—I’m always looking for the reliable people who will tell me useful, interesting information, who aren’t just going to tell me whatever somebody’s paying them to tell them, and who aren’t going to type a one-sentence prompt into an LLM, spit out an essay, and stick it online.
To me, earning that credibility is really important. That’s why a lot of my ethics around the way that I publish are based on the idea that I want people to trust me. I want to do things that gain credibility in people’s eyes so they will come to me for information as a trustworthy source. And it’s the same for the sources that I’m consulting as well. I’ve been thinking a lot about that sort of credibility focus for a while now.
You can layer or structure credibility, or decompose it. One thing I would put in front of you—I’m not saying that you should agree with this or accept this at all—is that you can use AI to generate different variations, and then you, as the final sort of last-mile person, pick the last output and put your stamp of credibility behind that, that everything is human-reviewed instead of human-originated, if that’s the thing.
If you publish something, you need to be able to be proud of publishing it. You need to be able to say, “I will put my name to this. I will attach my credibility to this thing.”
And if you’re willing to do that, then that’s great. For creators, this is huge because there’s a fundamental asymmetry between starting with a blank slate versus choosing from 5 different variations.
The key thing that you just said is that if everything that I do, if all of the words were generated by an LLM, if the voice is generated by an LLM, and if the video is also generated by an LLM, then I haven’t done anything, right? But if you take a shortcut on one or 2 of those, and I’m still willing to sign off on it, I feel like that’s where people are coming around to: this is maybe acceptable.
This is where I’ve been pushing the definition. I love the term “slop,” where I’ve been pushing the definition of slop as AI-generated content that is both unrequested and unreviewed. The unreviewed thing is really important.
The thing that elevates something from slop to not slop is if a human being has reviewed it and said, “You know what? This is actually worth other people’s time.” And again, I’m willing to attach my credibility to it and say, “Hey, this is worthwhile.”
It’s the curatorial and editorial part of it that, no matter what the tools are to do shortcuts—to do, as swyx is saying, choosing between different edits or different cuts—still has a curatorial mind or editorial mind behind it.
6. The GUI Moment For LLMs
Let me wedge this in before we start to close. One of the things that, coming back to your year-end piece, has been something I’ve been banging the drum about is when you’re talking about LLMs getting harder to use. You said most users are thrown in at the deep end.
The default LLM chat UI is like taking brand-new computer users, dropping them into a Linux terminal, and expecting them to figure it all out. I mean, it’s literally going back to the command line. The command line was defeated by the GUI interface, and what I’ve been banging the drum about is that this cannot be the user interface. What we have now cannot be the end result.
Do you see any hints or seeds of a GUI moment for LLM interfaces? I mean, it has to happen. It absolutely has to happen. The usability of these things is turning into a bit of a crisis, and we are at least seeing some really interesting innovation in little directions.
Simon Willison
Just like OpenAI's ChatGPT Canvas thing that they just launched. That is at least a little more interesting than just chats and responses. You're exploring that space where you're collaborating with an LLM, both working on the same document. That makes a lot of sense to me. That feels really smart.
One of the best things is still—who was it who did the UI where you could draw an interface and click a button? tldraw with their Make It Real thing. That was spectacular. Absolutely spectacular, like an alternative vision of how you'd interact with these models. So I feel like there is so much scope for innovation there, and it is beginning to happen. I feel like most people do understand that we need to do better in terms of interfaces that both help explain what's going on and give people better tools for working with models.
Brian McCullough
I was going to say, I want to dig a little deeper into this because think of the conceptual idea behind the GUI. Instead of typing into a command line, "open word.exe," you click an icon, right? That's abstracting away the programming stuff. A child can tap on an iPad and make a program open, right?
The problem, it seems to me, with how we're interacting with LLMs right now is it's sort of like a dumb robot where you poke it and it goes over here. But no, I want it to go over here, so you poke it this way, and you can't get it exactly right. What can we abstract away from what's currently going on that makes it more fine-tuned and easier to get more precise? You see what I'm saying?
Simon Willison
Okay.
Simon Willison
Yes. And this is the other trend that I've been following from the last year, which I think is super interesting. It's the prompt-driven UI development thing. Basically, this is the pattern where Claude Artifacts was the first thing to do this really well. You type in a prompt, and it goes, "Oh, I should answer that by writing a custom HTML and JavaScript application for you that does a certain thing."
Since then, it turns out this is easy, right? Every decent LLM can produce HTML and JavaScript that does something useful. So we've actually got this alternative way of interacting where they can respond to your prompt with an interactive, custom interface that you can work with.
People haven't quite wired those back up again. Ideally, I'd want the LLM to be able to ask me a question where it builds me a custom little UI for that question, and then it gets to see how I interacted with that. I don't know why, but that's such a small step from where we are right now. That feels like such an obvious next step.
Why should you just be communicating with text when it can build interfaces on the fly that let you select a point on a map or move sliders up and down, all of that kind of stuff?
It's gonna—like knobs and dials. I keep saying knobs and dials.
Simon Willison
Knobs and dials, right.
Brian McCullough
Yeah, exactly.
Simon Willison
We can do that, and the LLMs can build it. Claude Artifacts will build you a knobs-and-dials interface, but at the moment, they haven't closed the loop. When you twiddle those knobs, Claude doesn't see what you're doing. They're going to close that loop. I'm shocked that they haven't done it yet.
I think there's so much scope for innovation, and there's so much scope for doing interesting stuff with that model, where anything you can represent in HTML, JavaScript, and SVG, which is almost everything, can now be part of that ongoing conversation.
swyx
Yeah. I would say the best-executed version of this I've seen so far is Bolt, where you can literally type in, "Make a Spotify clone. Make an Airbnb clone," and it actually does that for you zero-shot with a nice design.
Simon Willison
Did you see there's a benchmark for that now?
swyx
Yeah.
Simon Willison
The LMArena people now have a—
swyx
LMArena benchmark.
Simon Willison
A benchmark for zero-shot app generation, because all of the models can do it. I've started figuring out how to—I'm building my own version of this for my own project because I think—
swyx
Oh.
Simon Willison
Within 6 months, I think it'll just be an expected feature. For my dataset data exploration project, I want you to be able to do things like conjure up a dashboard just via prompt. You say, "I need a pie chart and a bar chart, put them next to each other, and then have a form where submitting the form inserts a row into my database table."
This is all suddenly feasible. It's not even particularly difficult to do, which is utterly bizarre: these things are now easy.
swyx
Yeah. I think for a general audience, that is what I would highlight: software creation is becoming easier and easier. Gemini is now available in Gmail and Google Sheets. I don't write my own Google Sheets formulas anymore; I just tell Gemini to do it.
I almost want to somewhat disagree with your assertion that LLMs got harder to use.
Simon Willison
Ooh.
swyx
We expose more capabilities, but they're in minor forms, like using Canvas, web search in ChatGPT, and Gemini being in Google Sheets. We're getting improvements.
Simon Willison
No, no, no. Those are the things that make it harder.
swyx
Okay.
Simon Willison
The problem is that for each of those features, they're amazing if you understand the edges of the feature. If you're like, "Okay, so in Google Sheets formulas, I can get it to do a certain amount of things, but I can't get it to go and read a web page..." You probably can get it to read a web page, right? But there are things that it can do and things that it can't do, which are completely undocumented.
If you ask it what it can and can't do, they're terrible at answering questions about that. My favorite example is Claude Artifacts. You can't build a Claude Artifact that can hit an API somewhere else because the CORS headers on that iframe prevent accessing anything outside of CDNs.
swyx
I hate those.
Simon Willison
People are learning CORS headers as an end user in order to understand why—
swyx
I hate those.
Simon Willison
I've seen people saying, "Oh, this is rubbish. I tried building an artifact that would run a prompt, and it couldn't," because Claude didn't expose an API with CORS headers. All of this stuff is so weird and complicated.
The more tools we add, the more expertise you need to really understand the full scope of what you can do. The question really comes down to: What does it take to understand the full extent of what's possible? Honestly, that's just getting more and more involved over time.
7. Local Models Return
swyx
Yeah. I have one more topic that I think you're kind of a champion of, and we've touched on it a little bit, which is local LLMs and running AI applications on your desktop. I feel like you are an early adopter of many, many things.
Simon Willison
Well, I had an interesting experience with that over the past year. 6 months ago, I almost completely lost interest. The reason is that 6 months ago, the best local models you could run—there was no point in using them at all because the best hosted models were so much better.
There was no point at which I'd choose to run a model on my laptop if I had API access to Claude 3.5 Sonnet. They just weren't even comparable. That changed basically in the past 3 months as the local models had this step change in capability. Now I can run some of these local models, and they're not as good as Claude 3.5 Sonnet, but they're not so far away that it's not worth me even using them.
The continuing problem is I've only got 64 GB of RAM, and if you run Llama 3 70B, most of my RAM is gone. So now I have to shut down my Firefox tabs, Chrome, and VS Code windows in order to run it.
But it's got me interested again. The efficiency improvements are such that now, if you were to stick me on a desert island with my laptop, I'd be very productive using those local models, and that's pretty exciting. If those trends continue, and I think my next laptop, when I buy one, is going to have twice the amount of RAM, maybe I can run almost the top-tier open-weight models and still be able to use it as a computer as well.
NVIDIA just announced their $3,000, 128-gigabyte monstrosity. That's a pretty good price.
swyx
You gonna buy it?
swyx
Custom OS and all.
Simon Willison
If I get a job. If I have enough of an income that I can justify blowing $3,000 on it, then yes.
swyx
Okay. Let's do a GoFundMe to get Simon one of them. Come on. You know you can get a job anytime you want.
Simon Willison
I want a job that pays me to do exactly what I'm doing already and doesn't tell me what else to do. That's the challenge.
swyx
This is just purely discretionary. I think Ethan Mollick does pretty well, whatever it is he's doing. But basically, I was trying to bring in not just local models but Apple Intelligence, which is on every M-series Mac. You seem skeptical.
Simon Willison
It's rubbish.
swyx
It's rubbish.
Simon Willison
Apple Intelligence is so bad.
swyx
It does one thing well.
Simon Willison
Oh, yeah. What's that?
swyx
It summarizes notifications, and sometimes it's humorous.
Brian McCullough
But are you sure it does that well?
swyx
It's decent.
Brian McCullough
The other thing, again, from a sort of normie point of view, is that there's no indication from Apple of when to use it. Everybody upgrades their thing, and it's like, “Okay, now you have Apple Intelligence,” and you never know when to use it ever again.
swyx
Oh, yeah, you consult the Apple docs, which is MKBHD.
Simon Willison
The one thing I'll say about Apple Intelligence is that one of the reasons it's so disappointing is that the models are just weak. But now that—
swyx
Yeah.
Simon Willison
Llama 3B is such a good model in a 2-gigabyte file. I think, give Apple 6 months, and hopefully they'll catch up to—
swyx
Yeah.
Simon Willison
—the state of the art on their small models, and then maybe it'll start being a lot more interesting.
swyx
Anyway, this was year 1. Just like the first year of the iPhone, maybe it wasn't that much of a hit, and then in year 3 they had the App Store. So I would say give it—
Simon Willison
Yeah.
swyx
—some time. I think Chrome is also shipping Gemini Nano this year in Chrome, which means that every app, every web app, will have free access to a local model that just ships in the browser, which is kind of interesting.
I also wanted to open the floor to any of us: what are the AI applications that we've adopted that we really recommend? These are all apps that are running in a browser, or apps that are running locally, that other people should be trying, right? I feel like that's always one thing that's helpful at the start of the year.
Simon Willison
Okay. So, for running local models, my top picks are, firstly, on the iPhone, this thing called MLC Chat.
swyx
Mm-hmm.
Simon Willison
It works, it's easy to install, and it runs Llama 3B. It's so much fun. It's not necessarily a capable enough model for me to use it for real things, but my party trick right now is to get my phone to write a Netflix Christmas movie plot outline where a jeweler falls in love with the King of Sweden or whatever. It does a good job, and it comes up with pun names for the movies. That's deeply entertaining.
On my laptop, most recently, I've been getting heavily into Ollama because the—
swyx
Yeah.
Simon Willison
—Ollama team are very good at finding the good models, packaging them up, and making them work well. It gives you an API. My little LLM command-line tool has a plugin that talks to Ollama, which works really well. Ollama is, I think, the easiest on-ramp to running models locally. If you want a nice user interface, LM Studio is, I think, the best user interface for that. It's not open source, but it's good. It's worth playing with.
The other one that I've been toying with recently is called Open WebUI. The UI is fantastic. If you've got Ollama running and you fire this thing up, it spots Ollama and gives you an interface to your Ollama models, and that's really nicely done. That's my current favorite open-source UI for these things.
There are lots of good options. You do need a lot of disk space. The models start at 2 gigabytes for the 3B models that are actually worth playing with. The really impressive ones tend to be in the 20- to 30-gigabyte range, in my experience.
swyx
I think my struggle here is that I'm not much of an absolutist about running things locally. I'm happy to call an API.
Simon Willison
Mm-hmm.
swyx
Okay, yeah. But I just think—
Simon Willison
I do it to play.
swyx
Yeah, okay, fine.
Simon Willison
It's my research interest, yeah.
Brian McCullough
Answer your own question. Give us more apps that you want to—
swyx
Yeah. Sometimes it's just nice to recommend apps. I use Super Whisperer now. I tried Whisper Flow, but it didn't really work for me. Super Whisperer—
Simon Willison
Mm.
swyx
—is one of them. It basically replaces typing. You should just talk most of the time, especially if you're doing anything long-form. I hold down Caps Lock and talk, and when I'm done, I lift it up. It uses—
It isn't just about writing down your transcripts, because I make ums and uhs all the time, and I restate myself all the time. But I use GPT-4 to rewrite, and that's what these guys are doing. They're all doing some form of state-of-the-art ASR—automatic speech recognition—and then an LLM to rewrite.
I would also recommend that people check out Rosebud for journaling. I think AI for mental health is quite unexplored, and it's not because we're trying to build AI therapists. I think therapists really hate that. You'll never be on the level of therapists.
Brian McCullough
That gets back to the human thing that we were discussing. On some level, there are certain things and disciplines that require the human touch, and that might be one of them.
swyx
Sure. But the human touch costs me $300 an hour.
Brian McCullough
Yes.
swyx
Right?
Brian McCullough
Yeah.
swyx
And this thing's $3 a month. There's a spectrum of people for whom that will work, and I think it's cheap now to try all these things.
Simon Willison
I'm going to throw in a quick recommendation for an app. MacWhisper is my favorite—
swyx
Yeah, that's your one.
Simon Willison
—desktop app. I love that thing.
Brian McCullough
Yeah.
Simon Willison
It runs Whisper, and you can do things like paste in the URL to a YouTube video, and it'll pull the audio and give you a transcript. So that's how I watch YouTube now—
swyx
Ah.
Simon Willison
—I slap it into MacWhisper, then I copy and paste into Claude, and I use the Claude web app to do things.
MacWhisper works with MP3 files. Every time I'm on a podcast, I dump the MP3 into MacWhisper, then I dump the transcript into Claude and say, “What should I put in the show notes?” It spits out a bullet-point list where it says, “Oh, you mentioned a dataset that you should link to,” that kind of thing. MacWhisper—I use it several times a day, to be honest. It's great.
swyx
Yeah.
Brian McCullough
I'm actually going to say one that is incredibly basic and, again, coming back to just my workflow. We are currently recording this on Riverside. Riverside is a great tool for recording video and audio, like we're doing right now.
I always use this as an example when folks ask, “What will AI do for me?” When I first started using Riverside, we were recording 3 different channels, right? You guys are recording locally, so there are 3 audio files and 3 video files. When I first started using Riverside, you had to pump 3 tracks into Adobe and then edit.
“Okay, now we focus on Simon. Now we focus on swyx. Now we focus on Brian. Now we do all 3.” One day, a tool popped up that said, “Hit this button,” and it was Smart Edit. The AI determines, “Okay, Simon has been talking for 30 minutes, so go to the full shot of him. Brian is now talking, or there's overtalk, so let's have all 3 talking heads.”
With one button, for anything I posted, it saved me 3 or 4 hours' worth of work. That, to me, is, again, if normies are listening—
Simon Willison
In fact, Riverside has that feature now.
Brian McCullough
Yeah.
swyx
Yeah, yeah.
Simon Willison
Damn.
Brian McCullough
I don't use it.
Simon Willison
Oh, that sounds fantastic.
Brian McCullough
I still use a human editor. The day it came out, I was running around the house telling my wife, telling anyone that would listen, “You don't know, I just saved 3 hours because they had a new feature.”
Simon Willison
Wow.
That's exciting.
Brian McCullough
Brian's basically crying with joy right now. All right, let's try to bring this to a landing a little bit. Simon, I have maybe 2 or 3 more. We can do these rapid-fire.
One of my shows—one of the things about my show is that it's sort of like Silicon Valley writ large, so it's sort of like the horse race of who's up and who's down or whatever.
To the degree that you're interested in pontificating on this, OpenAI as a company in 2025, do you see challenges coming? Are you bearish or bullish? I'm almost doing a CNBC sort of thing, but how do you feel about OpenAI this year?
8. OpenAI Faces New Challenges
Simon Willison
I think they're in a bit of trouble. They seem to have lost a lot of talent.
Brian McCullough
Mm.
Simon Willison
And they don't have that top-of-the-pile thing. If it wasn't for o3, they'd be in massive trouble because they'd have lost that. I think o3 clawed them back up again.
One of the big stories of 2024 is that OpenAI started as the clear leader, and now Google Gemini is really good. Google Gemini had an amazing year. Anthropic Claude, Claude 3.5 Sonnet, is still my personal favorite model, and that feels notable. Nobody would argue that they weren't the leader in all of this stuff a year ago, and today they're still doing great, but they're not as far ahead as they were.
Brian McCullough
Next question, and maybe this couldn't be as rapid-fire, but I loved, finally, from your piece the idea that LLMs need better criticism, which I'd love you to expand on. As I straddle this world of tech journalism, creator, investor, and all that stuff, I thought that you had a really interesting thing to say about how—and we even alluded to this—Hollywood is against it.
Better criticism, in the sense that, as I took it, everybody's got their hackles up. They're trying to defend their livelihoods and things like that. But it's either, “This is going to destroy my job and destroy the world,” or—I'm sorry, I'm again leading the witness—what did you mean by LLMs needing better criticism?
Simon Willison
This is a frustration I have. If I read a discussion thread somewhere about this topic, I can predict exactly what everyone's going to say. People talk about the environmental impact. They talk about the plagiarism of the training data and the unlicensed training data. They'll often say, “Oh, and these things are completely useless.”
That's the one that I will push back against. The other things are true, right? The argument I always make about the idea that LLMs are just completely useless is that they are very useful if you understand how to use them, which is distinctly unintuitive. You have to learn how to deal with something that will just wildly hallucinate and make things up, and all of those kinds of things.
If you can learn what they're good at and what they're bad at, I use them dozens of times a day, and I get enormous value out of them. So I'll push back on people who say, “No, they're just useless.”
But the other things—the environmental impact and the way the training data works—I feel like the training-data one is interesting because it's probably legal under fair use, but it's clearly unfair if somebody takes your work without your permission and trains a model that then competes with you in the marketplace. Legal or not, I understand why people are upset about that. That's a reasonable thing to be upset by.
So what I want—and I also feel like this stuff can have a major impact on society, especially as it starts undermining all sorts of jobs that we never thought were going to be undermined by technology. Who thought it would come for artists and lawyers first? That's bizarre.
We need to have really high-quality conversations where we help people figure out what works and what doesn't work. We need people to be able to make good decisions about what to do with their careers, to embrace this stuff, and all of that sort of thing.
If we just get distracted by saying, “Yeah, but it's useless, plagiarism-driven, environmentally catastrophic,” even though those things represent quite a lot of truth, I don't think that's a useful message to lead with. I want to be having the much more interesting, high-level conversations: If there are negatives, how do we counter those negatives? If there are positives, how do we encourage those? How do we help people make good decisions about how to use this technology?
swyx
Yeah. I think where I see this the most is for people who are very internal. You and I are immersed in this every single day, so we're frankly tired of the same debates being recycled again and again.
I think what might be more useful or more impactful is the level at which it starts to hit regulation. Last year, we had a couple of very notable attempts at the White House level and in California to regulate AI, and those did not come to pass.
At some point, these criticisms bubble up to law, to matters of national security or national science and progress. I feel like there needs to be more information or enlightenment there, maybe, if only because they tend to be very trailing.
Simon Willison
Right.
swyx
My favorite example to pick on, which is very unfair of me, but whatever, is that the California SB 1047 act tried to cap compute at 10^25.
Simon Willison
Yeah. That was DeepSeek.
swyx
Exactly. And it also was exactly at the point at which we pivoted from training GPT-5 to o1, where we're no longer scaling pre-training compute. What I'm saying is that we're always trying to regulate the last war, and I don't think that works in a field that is—
Simon Willison
So—
swyx
—basically 8 years old.
Simon Willison
I think there are 2 areas of regulation I'm super interested in. One of them is that I do think regulating the way these things are used can work. The big example is that I don't want somebody's insurance claim denied by a black-box LLM where nobody can explain what it did. That just feels—
swyx
Oh, we have real laws for that. This is like redlining.
Simon Willison
Exactly. Take those laws, reinforce them, and update them for modern capabilities.
The other one is privacy. We've got this huge problem right now where people will refuse to use any of these tools because they don't trust that the things they say to them won't be trained on and then exposed to other people. There are lots of terms and conditions that you can read through and try to navigate around.
I would love there to be straightforward laws that people understand, where they know that their input isn't going to be used for training because there's a law that says, under these circumstances, that can't happen. It's basically taking our existing privacy laws, giving them a few more teeth, and reinforcing them without introducing cookie banners à la the European Union.
These things are always very risky. You can have all sorts of bad results if you don't design them correctly. But there's space for that, I think.
Brian McCullough
Yeah. When I read that piece and then when you just said, “Swyx said we're in the weeds on this every single day, so we're tired of hearing these arguments,” it reminds me of folks who are always into politics. Then they're mad at the people who don't care about politics until it's an election year, and they're like, “Well, you're a low-information voter because all you know is that the factory in your town got shut down, or there's inflation, or whatever, and so you vote one way or the other, but you haven't been paying attention.”
But that's kind of the point: You shouldn't expect normal people to pay attention, except for the fact that this might lose me my job. So you can't blame them for being—I don't know if reactionary is the word—or emotional.
If you're in the weeds, it's harder to keep everybody informed, and this is going to touch everybody, so I don't know.
Okay, so this is the very last one, and then we can wrap and do plugs and everything. Simon, this is for you. It was alluded to a little bit, and you might not have one, but if there's something this year that a generalist like me isn't aware is coming down the pipe that you think is going to be big in the AI space—and maybe Swyx, if you've got one too—what do you think it would be?
Simon Willison
I think for most people who haven't been paying attention, we know these things already. We know that the models are now almost free to run things against. The fact that you can now do video, stream video to a model—the thing where you can share your entire screen with a model and get feedback is going to be really useful.
Again, the privacy side of things really matters, though. I do not want some model just training on everything that it sees on my screen. But no, the stuff that's now possible as of a few months ago is enough. I don't need anything new. That's going to keep me busy all year.
Brian McCullough
Swyx, you got one?
swyx
Simon's always too content, and then he sees the next thing and he's like, “Oh yeah, that's great too.”
Simon Willison
Yep.
swyx
Okay. I love trying to be contrarian by asking, “What does everyone hate right now?” Remember, this time last year, we had just had CES and the Rabbit R1.
We had the Humane, right?
Brian McCullough
Yeah. Wearables. Yep.
9. AI Wearables Make A Comeback
swyx
Those are completely in the gutter. No one will touch them. They're toxic nuclear waste. Okay, this year is the year of wearables.
Brian McCullough
Yep.
Simon Willison
Huh.
Brian McCullough
I agree with you, by the way. That cycle always works out where you go to a CES and it's everything—hype, hype, hype, hype—and then 3 years later it becomes the thing, unless it's 3D TVs, in which case that was a mistake anyway. But yeah—
Simon Willison
Well, transparent TVs are the big thing—
Brian McCullough
Mm.
Simon Willison
—for the last couple of years. What the hell?
Brian McCullough
Yeah.
swyx
I think Simon may have got one of these, but there are a lot of people working on AI wearables here in SF. They are surprisingly cheap, surprisingly capable, with decent battery life, and they do useful things. We have to work out the privacy aspect, of course. But people like Limitless, which used to be called Rewind, I think—
Brian McCullough
Mm-hmm.
swyx
They're shipping one of these wearables that, based on your voice, only records your voice. So you opt in.
Simon Willison
Interesting. Right.
swyx
Right? And so you can have perfect memory if you want. You can have perfect memory at work. Your employer can buy these for you; it only applies at work, and it's fine. It's just a meeting aid.
Lots of people use Granola or some kind of Fireflies or some of these meeting recorders only for online meetings, but what about in-person meetings? What about conversations and locations that you've been to? Some of that should be a choice. Right now you have zero choice. And I think these wearables will enable some of that.
It's up to us as a society to determine what's acceptable and what's not. I really like these gray areas where we still don't know yet. Whenever I tell people about this, they're like, “I don't know.” I guess it's as though you have perfect memory, but some people have better memory than others. Where's the line?
Brian McCullough
Hmm. And—
swyx
There will be a lot—
Brian McCullough
Now I—
swyx
A lot more of these.
Brian McCullough
I would add to that because, swyx, as you know—you listen to my show—the idea is that AI has taken smart glasses and completely changed everyone's mind about that as a product category and form factor. And I should say this: from things that I've been looking at investing in, wait until you see what they can add on to earbuds.
Simon Willison
Ooh.
Brian McCullough
Like the earbuds in your ear can do a lot more things than they're doing now. Then you combine that with smart glasses, and you combine that with an LLM that you can access maybe with a phone as the mothership. There are some interesting things. CES next year is going to be crazy if you think AI wearables are a thing.
swyx
Anyway, this year they were not a thing. There were very much no wearables at CES.
Brian McCullough
Mm.
Simon Willison
This one's interesting as well because the thing that makes these interesting is that it's multimodal, right? Audio input, video input—
Brian McCullough
Yeah.
Simon Willison
Image input. A year ago, that was hardly a thing, and now it's dirt cheap. We're in a much better position now than we were 12 months ago to build the software behind this stuff.