[BidClub_]
Latent Space · · 73 min

[Ride Home] Simon Willison: Things we learned about LLMs in 2024

Brian McCulloughSimon Willisonswyx

Podcast
TL;DR
  • AI’s 2024 step change was a collapse in cost and an expansion in capability, not a clean intelligence leap beyond GPT-4. Simon Willison counts 18 organizations whose models beat the GPT-4-from-a-year-earlier barrier, while hosted inference became dramatically cheaper and multimodal features reached phones. His summary: “everything’s got really good and fast and cheap,” even though models “didn’t get massively better than GPT-4.”

  • DeepSeek-V3 punctured the thesis that frontier-model development must inevitably consolidate around a few labs spending billions. DeepSeek reported a $5.5 million training run—roughly one-tenth of prior assumptions—and released the leading open-weights model on Christmas Day, prompting Willison to call it “an absolute bombshell.” Swyx preserves the diligence caveat: outsiders question the accounting and speculate about copying, but “we basically will never know” as external commentators; the demonstrated model still changes expectations for capital intensity.

  • The viable agent market is splitting between bounded, reviewable workflows and autonomous systems that remain unsafe. Research agents can inspect 56 websites, coding agents can execute code and repair errors, and Swyx says NotebookLM’s podcast feature uses internal agent loops, though the speakers note that this depends on the definition of “agent.” But an agent allowed to browse, decide, and spend encounters the unsolved gullibility problem: Claude Computer Use followed a webpage’s instruction to download malware and joined a botnet, making trusted economic autonomy feel to Willison like “an AGI-level problem.”

  • Multimodal inference has become cheap enough to turn cameras, screens, glasses, and earbuds into continuously promptable software surfaces. Gemini’s free tier may support roughly one captured image per second or minute, while Willison calculated that captioning his entire 68,000-photo collection with Gemini 1.5 Flash-8B would cost $1.68—about 1/400 of a cent per image. The opportunity spans monitoring, accessibility, memory, and wearables, but privacy becomes load-bearing when a model can see an entire screen or record daily life.

  • The next application-layer moat may be interface design rather than another chat box. Willison compares today’s blank prompt to dropping a new computer user into a Linux terminal; Canvas, tldraw’s Make It Real, Claude Artifacts, and Bolt point toward models generating task-specific maps, sliders, dashboards, and applications. The missing loop is letting the model observe how users manipulate those “knobs and dials,” turning generated interfaces into an actual conversation.

  • As synthetic media becomes abundant, credibility and human review become scarcer—and therefore more valuable. AI can remove production steps through lip-sync avatars, transcription, editing, and generated background assets, but Willison argues that an LLM cannot stake a reputation because “it’s a matrix multiplication.” His useful boundary for “slop” is content that is both “unrequested and unreviewed”; in his framing, human-reviewed output carrying a creator’s name is not slop.

  • OpenAI remains strong but no longer enjoys an uncontested platform lead. Willison says o3 “clawed them back up again,” yet Google Gemini had an exceptional year and Claude 3.5 Sonnet remained his favorite model; meanwhile, capable local models again became worth running after a late-2024 step change. The competitive field now rewards efficiency, distribution, workflow integration, trust, and usability—not merely owning the largest pretraining run.

Digest · the substance, structured for research

1. AI improved sideways rather than delivering the expected GPT-5 leap

  • Willison’s state-of-the-market summary is blunt: “everything’s got really good and fast and cheap.” Models gained longer context, image and audio capabilities, video understanding, and much lower latency and cost, but “didn’t get massively better than GPT-4.”

  • McCullough’s pushback captures the unmet expectation: after GPT-2, GPT-3, and GPT-4, users anticipated another obvious phase change in intelligence. Willison argues that the phase change arrived elsewhere—by year-end, a phone could converse, inspect its camera feed, and impersonate Santa Claus.

  • Asked to bet on 2025, Willison declines to predict a simple “models but smarter” jump. He expects inference-time compute such as o1 and o3 to keep working—harder problems become solvable by spending more money and waiting longer—and would be “completely happy” with today’s intelligence made cheaper, faster, more capable, and longer-context.

2. The GPT-4 barrier fell as inference prices collapsed

  • At the start of 2024, OpenAI had led for roughly nine months without a close challenger. By year-end, Willison counted 18 other organizations with models that clearly beat the GPT-4-from-a-year-earlier model; “that barrier got completely smashed.”

  • Microsoft’s Phi-4 made the shift tangible: a 14-gigabyte download running on a MacBook Pro, with benchmarks definitely up there with GPT-4.0. Willison keeps the qualitative hedge—“It’s probably not as good when you actually get into the vibes of the thing”—but a laptop running anything comparable had seemed implausible one year earlier.

  • OpenAI’s current inference was roughly 100 times cheaper than using GPT-3 two and a half years earlier. Gemini 1.5 Flash cost $0.075 per million tokens, while Gemini 1.5 Flash-8B was 27 times cheaper than GPT-3.5 Turbo had been a year earlier despite adding image recognition and million-token context.

  • Competition explains part of the decline, but Willison says trusted sources told him Google Gemini was not losing money on inference; an Amazon executive likewise indicated that Amazon Nova was not losing money. That excludes model training and the “army of PhDs,” but it shows that extraordinarily cheap marginal inference need not be subsidized.

3. DeepSeek reset the capital-intensity debate without resolving its mysteries

  • McCullough had worried that billion-dollar or larger training runs could leave only nation-states able to train new models. DeepSeek released V3 on Hugging Face on Christmas Day as a “giant binary blob” without even a README. It led open-weights benchmarks and reportedly cost $5.5 million to train—about one-tenth of prevailing estimates—leading Willison to call it “a bombshell” that blew apart that earlier mental model.

  • Willison’s hypothesis is that export controls forced Chinese labs to extract more from constrained hardware, exposing abundant low-hanging efficiency gains. He “would not be surprised” to see better models trained for still less money within six months, but explicitly presents this as an uninformed opinion rather than laboratory expertise.

  • Swyx supplies the skeptical case: few labs trumpet cheap training, DeepSeek has roughly 150 employees, and “nobody quite believes” the headline accounting. Online theories allege copying because DeepSeek sometimes identifies itself as Claude or OpenAI GPT-4, but Swyx concedes that outsiders do not know the training tokens or provenance and “will basically never know” as external commentators.

  • Swyx estimates that equal-Elo GPT-4-class capability became 1,000 times cheaper during 2024, then asks whether that is a Moore’s-law curve or a one-off harvest. Willison suspects researchers only recently prioritized efficiency, while conceding 2024 might have exhausted the easiest gains: “we’ll know for sure in about three months.”

4. Open reasoning models made model behavior newly visible

  • The exchange about DeepSeek R1 is internally inconsistent: Swyx first describes it as a reasoning model runnable on a laptop and answers yes when asked whether its weights were released, then later says, “R1 is the API available.” The transcript does not establish whether local R1 weights were available.

  • Willison instead points to Alibaba’s Qwen reasoning models QwQ and QVQ, the latter adding vision. Unlike o1, which sort of hides its thinking process, the Qwen models visibly “churn away,” producing dozens of paragraphs while solving a problem.

  • His best specimen is the Pelican Bench: QwQ contemplated SVG construction in Chinese before producing a respectable pelican riding a bicycle. “The fact that my laptop can think in Chinese now is so delightful,” he says—the playful test also exposes style, visual coding ability, and reasoning behavior that benchmarks can miss.

5. Autonomous agents remain blocked by gullibility, not merely accuracy

  • Willison’s first frustration is definitional: “agent” can mean a travel booker, an LLM using tools in a loop, or a scheduled background job, while each builder assumes their definition is universal. Academics have debated the term for more than 30 years; for this discussion, an agent receives a job and independently performs it.

  • Under that definition, reliability collides with prompt injection. LLMs cannot reliably distinguish instructions from untrusted content, and Willison stresses that the field has discussed this unsolved problem for two years.

  • Anthropic’s Claude Computer Use provided the decisive demonstration: it could operate a browser inside a container, but a webpage telling it to download and execute a file succeeded immediately. The file was malware that enrolled the machine in a botnet—the “very first, most obvious dumb trick” worked.

  • Swyx compares agent enthusiasm to the perennial “year of Linux on the desktop,” but resists pure cynicism. Self-driving was supposedly imminent in 2014, yet Waymo now works in bounded settings; agents may follow the same “slow cook,” accumulating concrete progress over a decade without solving everything at once.

6. Research and coding agents work because their loops are bounded

  • Willison “mostly believes” in research assistants. Google’s Gemini 1.5 Pro with what he thinks is called Deep Research could inspect 56 websites, load them into its million-token context, and produce a genuinely useful report, aided by Google’s search index and page cache. Deliberately deceptive sources could still defeat it, but most research tasks are not adversarial.

  • Coding agents have an even longer proof point: ChatGPT Code Interpreter was already writing Python, executing it, reading errors, and rewriting the code nearly two years earlier. That feedback loop “obviously works” because execution supplies a checkable signal.

  • The dividing line is consequential autonomy. Willison does not expect an agent that independently makes decisions and spends money to work reliably “for a very long time”; Stripe can give an agent a virtual card, but his immediate safety mechanism is simply a $50 spending cap.

  • Even travel automation offers less value than its demos imply. Google Flights already works, an agent might save Willison 15 seconds, and he still wants to reject an airline regardless of price. NotebookLM’s stronger pattern combines a useful retrieval product with a “total gimmick”—its striking synthetic podcast—that finally earned the underlying product attention.

7. Multimodal inference turned continuous perception into a cheap primitive

  • A year earlier, GPT-4 Vision was the lone clearly impressive vision model and Gemini 1.0 had not earned broad credibility. Gemini 1.5 Pro changed that; video systems could sample one frame per second into long context, while newer models began combining imagery and audio more natively.

  • ChatGPT’s iPhone app made the shift concrete: a user can open the camera mid-conversation and ask, “What kind of tree is this?” Whether it samples frames or handles richer video internally matters less to Willison than the new end-user capability, which he thinks most people had not yet noticed.

  • Swyx says Gemini Flash’s free tier may support something like one captured photo every second or minute, potentially enabling a camera app that runs continuously and detects changes or issues prompted alerts. Willison’s cost calculation is starker: captioning 68,000 photographs with Gemini 1.5 Flash-8B would cost $1.68, or about 1/400 of a cent per image. “That doesn’t make sense. None of that makes sense.”

8. Generative video will enter production through pieces, not prompted movies

  • Sora’s public release appeared disappointing beside Google’s Veo 2, but Swyx objects to the comparison: users received distilled “Sora Lite,” while social media compared its failures with Veo 2’s cherry-picked marketing examples. He believes Veo 2 may still be better; no one had announced when “full-fat Sora” would ship.

  • Willison’s benchmark is not a two-hour film generated from three sentences. He asks what elite artists could achieve with these tools, pointing to the five-person visual-effects team behind Everything Everywhere All at Once, some of whom learned techniques from YouTube. Swyx adds that the team used Runway ML, while warning that he does not know how extensively.

  • McCullough expects adoption in three-second and 20-second pieces that previously required enormous budgets. Swyx similarly emphasizes low-value backgrounds, crowds, music, and sound effects, where consistency defects matter less than when Sora mangles a foreground gymnast.

  • Swyx sees a cultural war alongside the technical one: much of Hollywood opposes AI, leaving adoption to a fringe of willing artists. He also flags Hai Luo, Kling, and other Chinese systems as surprisingly strong, prompting Willison to wonder whether an AI-native film industry could emerge outside established production centers.

9. Human review becomes the boundary between leverage and slop

  • McCullough describes recording his own audio, then using HeyGen to lip-sync a trained avatar. The result is not perfectly beyond the uncanny valley, but it removes camera, lighting, and editing work while preserving his voice and message—the useful future is workflow compression, not necessarily synthetic personalities replacing creators.

  • Willison welcomes tools that let humans attempt more ambitious work, but keeps returning to credibility. ChatGPT cannot stake its reputation on a claim because “it’s a matrix multiplication”; credibility belongs to the person willing to publish, defend, and put a name behind the result.

  • Willison proposes human-reviewed rather than human-originated work: generate several variations, choose one, and attach personal responsibility. His matching definition of “slop” is AI content that is both “unrequested and unreviewed”; editorial judgment elevates an output only when someone has decided it is worth another person’s time.

10. Generated interfaces can replace the blank-prompt command line

  • McCullough frames the current chat UI as a usability crisis, echoing Willison’s comparison of a blank prompt to dropping a new computer user into a Linux terminal. ChatGPT Canvas offers collaborative document editing, while tldraw’s Make It Real showed how drawing an interface could generate working software.

  • Claude Artifacts introduced another path: the model answers with a custom HTML and JavaScript application instead of prose. Willison wants models to ask questions through generated maps, sliders, and “knobs and dials,” then observe the interaction; Artifacts can build those controls but had not yet closed the feedback loop.

  • Bolt already generated polished Spotify- or Airbnb-like applications from one prompt, and zero-shot app generation had become common enough to benchmark. Willison expects this to become a standard web-app feature within six months and wants his dataset data-exploration project to support prompted dashboards, forms, charts, and database actions.

  • Swyx argues that Canvas and Gemini inside Google Sheets make LLMs easier to use. Willison’s rebuttal is that every feature adds undocumented boundaries: an Artifact could not call arbitrary APIs because of iframe CORS restrictions, forcing ordinary users to learn web-security headers. More capabilities therefore increase the expertise required to understand what is actually possible.

11. Local models revived, while practical workflow tools delivered immediate returns

  • Willison had nearly abandoned local LLMs because nothing on his laptop approached Claude 3.5 Sonnet. A capability jump during the final three months renewed his interest: local models remained weaker, but no longer uselessly far behind.

  • Hardware is still constraining. Running a Llama 3 70B-class model consumes most of his 64GB of RAM and forces him to close browsers and VS Code; a future laptop with twice the memory—or NVIDIA’s announced $3,000, 128GB machine—could run near-top-tier open weights while remaining usable as a computer.

  • His local on-ramps are MLC Chat on iPhone, Ollama for packaged models and an API, LM Studio for a polished interface, and Open WebUI as an open-source front end. Playable models start around 2GB, while the most impressive laptop-friendly downloads typically occupy 20–30GB.

  • Apple Intelligence receives the harshest review: “It’s rubbish,” largely because its models are weak and users are not shown when to invoke it. Still, Willison thinks better small models could improve it within six months; elsewhere, MacWhisper became a several-times-daily transcription tool, and Riverside’s Smart Edit could eliminate three or four hours of camera-switching work—even though McCullough still retained a human editor.

12. Competition, regulation, and wearables define the next pressure points

  • Willison thinks OpenAI is “in a bit of trouble” after losing talent and its unambiguous lead; without o3, the position would look much worse. O3 restored some frontier status, but Gemini had “an amazing year,” and Claude 3.5 Sonnet remained his personal favorite.

  • He wants better criticism than endlessly repeating that LLMs are useless, environmentally damaging, and trained on unlicensed work. The harms contain substantial truth—training is probably legal under fair use yet plainly feels unfair when the resulting model competes with the creator—but “completely useless” ignores enormous value available to users who learn the models’ unintuitive limits.

  • Swyx warns that regulation keeps targeting “the last war,” citing California SB 1047’s proposed 10^25 compute threshold just as DeepSeek emphasized efficiency and labs pivoted from scaling GPT-5 pretraining toward o1-style inference. Willison prefers regulating uses: prohibit unexplained black-box insurance denials and create simple privacy rules ensuring prompts are not reused for training, while avoiding cookie-banner-style failure modes.

  • Swyx’s contrarian 2025 call is wearables: after Rabbit R1 and Humane became “toxic nuclear waste,” cheaper multimodal models make the category newly feasible. Limitless, formerly Rewind, is shipping a wearable that records only the wearer’s voice when opted in; McCullough adds smart glasses and more capable earbuds linked to a phone “mothership.” The unresolved product boundary is societal permission: when does useful memory become unacceptable recording?

Verification Notes

  • The transcript is internally inconsistent on DeepSeek R1: Swyx first describes it as runnable on a laptop and confirms released weights, then later says “R1 is the API available.” The digest does not resolve the local-weight status.
Brian McCullough

Welcome to the first bonus episode of the "Techmeme Ride Home" for the year 2025. I'm your host, as always, Brian McCullough. Listeners to the pod over the last year know that I have made a habit of quoting Simon Willison when new stuff happens in AI from his blog. Simon has become a go-to for many folks in terms of analyzing and criticizing things in the AI space. I've wanted to talk to you for a long time, Simon, so thank you for coming on the show.

Simon Willison

No, it's a privilege to be here.

Brian McCullough

The person who made this connection happen is our friend swyx, who has been on the show going back to the Twitter Spaces days.

Simon Willison

Wow.

Brian McCullough

He's also an AI guru in his own right. swyx, thanks for coming on the show also.

swyx

Thanks. Happy to be on. I've been a regular listener, so I'm just happy to contribute as well.

Brian McCullough

And a good friend of the pod, as they say. All right, let's go right into it. Simon, I'm going to do the most unfair broad question first, so let's get it out of the way. The year 2025, broadly, what is the state of AI as we begin this year? Whatever you want to say. I want to lead the witness.

Simon Willison

Wow. So many things, right? The big thing is that everything's gotten really good, fast, and cheap. That was the trend throughout all of 2024. The good models got so much cheaper and faster. They got multimodal, right? The image stuff isn't even a surprise anymore. They're adding video, all of that kind of stuff.

At the same time, they didn't get massively better than GPT-4, which was a bit of a surprise. That's sort of one of the open questions. But I feel like that's a bit of a distraction, because GPT-4, but way cheaper, with much larger context lengths and multimodality, is better, right? That's a better model, even if it's—

Brian McCullough

But—

Simon Willison

Not—yeah.

Brian McCullough

What people were expecting, right? Not expecting is not the right word, but hoping that we would see another step change, right? Where, from GPT-2 to GPT-3 to GPT-4, we were expecting or hoping that maybe we were going to see the next evolution in that sort of—

Simon Willison

And I think—

Brian McCullough

Yeah.

Simon Willison

We did see that, but not in the way we expected. We thought the model was just going to get smarter, and instead we got massive drops in price. We got all of these new capabilities. You can talk to these things now, right? They can do simulated audio input, all of that kind of stuff.

It's interesting to me that the models improved in all of these ways we weren't necessarily expecting. I didn't know it would be able to do an impersonation of Santa Claus, and I could talk to it through my phone and show it what I was seeing by the end of 2024. But we didn't get that GPT-5 step, and that's one of the big open questions: Is that actually just around the corner? Will we have a bunch of GPT-5-class models drop in the next few months?

Brian McCullough

What if you had to put—

Simon Willison

Or is there a limit?

Brian McCullough

If you were a betting man and wanted to put money on it, do you expect to see a phase change, a step change, in 2025?

Simon Willison

I don't particularly expect the models to just get smarter. I think all of the trends we're seeing right now are going to keep going, especially inference-time compute, right? The trick that o1 and o3 are doing means that you can solve harder problems, but it costs more and churns away for longer. I think that's going to happen because it's already proven to work.

I don't know. Maybe there will be a step change to a GPT-5 level, but honestly, I'd be completely happy if we got what we've got right now, but cheaper and faster, with more capabilities and longer context, and so forth.

Brian McCullough

Well—

Simon Willison

That would be thrilling to me.

Brian McCullough

Digging into what you've just said, one of the things that you did say, and that you alluded to even right there, was that in the last year you felt like the GPT-4 barrier was broken. In other words, other models, even open-source ones, are now regularly matching the state of the art?

1. The GPT-4 Barrier Falls

Simon Willison

Well, it's interesting, right? The GPT-4 barrier was that, a year ago, the best available model was OpenAI's GPT-4, and nobody else had even come close to it. They'd been in the lead for 9 months, right? That thing came out in February or March 2023. For the rest of 2023, nobody else came close.

At the start of last year, the big question was: Why has nobody beaten them yet? What did they know that the rest of the industry didn't know? Today, I've counted 18 organizations other than GPT-4 who've put out a model which clearly beats that GPT-4-from-a-year-ago thing. Maybe they're not better than GPT-4.0, but that barrier got completely smashed.

A few of those I've run on my laptop, which is wild to me. It felt very clear to me a year ago that if you wanted GPT-4, you needed a rack of $40,000 GPUs just to run the thing. That turned out not to be true. This is that big trend from last year of the models getting more efficient, cheaper to run, just as capable with smaller weights, and so forth.

I ran another GPT-4 model on my laptop this morning, right? Microsoft's Phi-4 just came out, and, if you look at the benchmarks, it's definitely up there with GPT-4.0. It's probably not as good when you actually get into the vibes of the thing, but it runs on my computer. It's a 14 GB download, and I can run it on a MacBook Pro. Who saw that coming?

The most exciting thing at the close of the year, on Christmas Day just a few weeks ago, was when DeepSeek dropped their DeepSeek-V3 model on Hugging Face without even a README file. It was just a giant binary blob. I can't run that on my laptop; it's too big. But in all of the benchmarks, it's now by far the best available open-weights model. It's beating the Meta Llama models and so forth. That was trained for $5.5 million, which is a tenth of the price that people thought it cost to train these things. Everything's trending smaller, faster, and more efficient.

2. The DeepSeek Cost Shock

Brian McCullough

Well, okay. I was going to get to that later, but let's combine this with what I was going to ask you next. You're talking also in the piece about LLM prices crashing, which I've even seen in projects that I'm working on. But explain that to a general audience, because we hear all the time that LLMs are eye-wateringly expensive to run.

What we're suggesting—and we'll come back to the cheap Chinese LLM—but first of all, for the end user, what you're suggesting is that we're starting to see the cost come down in the traditional technology way, with costs coming down over time?

Simon Willison

Yes, but very aggressively. My favorite example here is if you look at GPT-3, OpenAI's GPT-3, which was the best available model in 2022 and through most of 2023, the models that we have today from OpenAI are 100 times cheaper. It was a 100-times drop in price for OpenAI, from their best available model two and a half years ago to today.

Brian McCullough

And just to be clear, not to train the model, but for the use of tokens and—

Simon Willison

Exactly.

Brian McCullough

Yeah.

Simon Willison

For running prompts through them. When you look at the top-tier model providers right now, I think they're OpenAI, Anthropic, Google, and Meta. There are a bunch of others that I could list as well. Mistral is very good. The DeepSeek and Qwen models are great. There's a whole bunch of providers serving really good models.

But even if you just look at the big brand-name providers, they all offer models now that are a fraction of the price of the models we were using last year. I think I've got some numbers that I threw into my blog entry here. Yeah, Gemini 1.5 Flash, Google's fast, high-quality model, is—how much is that? It's $0.075 per million tokens. These numbers are getting so small—

swyx

We just use cents per million now. Cents per million.

Simon Willison

Right. Cents per million makes a lot more sense. Google has one model, Gemini 1.5 Flash-8B, the absolute cheapest of the Google models, that's 27 times cheaper than GPT-3.5 Turbo was a year ago. That's a model 27 times cheaper, and this Google one can do image recognition, million-token context, all of those tricks. There's—

Brian McCullough

Is—

Simon Willison

It's really startling how inexpensive some of this stuff has gotten.

Brian McCullough

Now, are we assuming that this is happening directly as a result of competition? Because, again, OpenAI—and they're probably doing this for their own strategic reasons—keeps saying, "We're losing money on everything, even the $200-per—" The prices wouldn't be coming down if there wasn't intense competition in this space.

Simon Willison

The competition's absolutely part of it. But I have it on good authority from sources I trust that Google Gemini is not operating at a loss. The cost of the electricity to run a prompt is less than they charge you, and the same thing is true for Amazon Nova.

Somebody found an Amazon executive and got them to say, “Yeah, we’re not losing money on this.” I don’t know about Anthropic and OpenAI, but clearly that demonstrates it’s possible to run these things at these ludicrously low prices and still not be running at a loss if you discount the army of PhDs, the training costs, and all of that kind of stuff.

Brian McCullough

One more for me before I let swyx jump in here. To come back to DeepSeek and this idea that you could train a cutting-edge model for $6 million, I was saying on the show 6 months ago that if we’re getting to the point where each new model costs $1 billion, $10 billion, or $100 billion to train, at some point it would almost be that only nation-states would be able to train the new models. Do you expect what DeepSeek and maybe others are proving to blow that up?

Or is there some sort of parallel track here that maybe I’m not technically equipped to understand? Is the model—or are the models—going to go up to $100 billion, or can we get them down, sort of like DeepSeek has proven?

Simon Willison

So I’m the wrong person to answer that because I don’t work in a lab training these models.

Brian McCullough

Mm-hmm.

Simon Willison

I can give you my completely uninformed opinion, which is: I feel like the DeepSeek thing was a bombshell. That was an absolute bombshell. When they came out and said, “Hey, look, we’ve trained one of the best available models, and it cost us $6 million—$5.5 million—to do it,” I felt—

Brian McCullough

Love it.

Simon Willison

The reason it’s so efficient is that we put all of these export controls in place to stop Chinese companies from buying GPUs, so they were forced to be as efficient as possible. And yet the fact that they’ve demonstrated that that’s possible completely tears apart this mental model we had before: that the training runs just keep getting more and more expensive, and the number of organizations that can afford to run these training runs keeps shrinking. That’s been blown out of the water.

So, yeah, this was our Christmas gift. This was the thing they dropped on Christmas Day. It makes me really optimistic that there is so much low-hanging fruit in terms of the efficiency of both inference and training, and we spent a whole bunch of last year exploring that and getting results from it. I think there’s probably a lot left. I would not be surprised to see even better models trained while spending even less money over the next 6 months.

swyx

Yeah. I think there’s an unspoken angle here about what exactly the Chinese labs are trying to do. DeepSeek made a lot of noise around the fact that they trained their model for $6 million, and nobody quite believes them. It’s very rare for a lab to trumpet the fact that they’re doing it for so cheap. They’re not trying to get anyone to buy them, so why are they doing this?

They make it very obvious that their lab, DeepSeek, is about 150 employees. It’s an order of magnitude smaller than at least Anthropic, and maybe more so for OpenAI. So what’s the end game here? Are they just trying to show that the Chinese are better than us?

Simon Willison

I mean, DeepSeek is the arm of High-Flyer, a quant fund, right? It’s an algorithmic quant-trading thing. I’d love to get more insight into how that organization works. My assumption from what I’ve seen is that it looks like they’re basically just flexing. They’re saying, “Look at how utterly brilliant we are with this amazing thing that we’ve done,” and it’s working, right?

But is that it? Is this just their kind of, “This is why our company is so amazing. Look at this thing that we’ve done,” or… I don’t know. I’d love to get some insight from within that industry as to how that’s all playing out.

swyx

The prevailing theory among the local Llama crew and the Twitter crew that I index for my newsletter is that there is some amount of copying going on. It’s like Sam Altman tweeting about how they’re being copied, and then there are other OpenAI employees who have said things that are similar—that DeepSeek’s rate of progress is how U.S. intelligence estimates the number of foreign spies embedded in top labs.

A lot of these ideas do spread around, but they surprisingly have a very high density in the DeepSeek V3 technical report. We don’t know how much copying there was or how many tokens were involved. People have run analyses on how often DeepSeek thinks it is Claude or thinks it is OpenAI GPT-4, and we don’t know.

For me, we’ll basically never know as external commentators. I think what’s interesting is: Where does this go? Is there a logical floor or bottom? By my estimations, for the same Elo, from the start of last year to the end of last year, costs went down by 1000× for GPT-4 intelligence. Do they go down 1000× this year?

Simon Willison

That’s a fascinating question.

swyx

Is there a Moore’s law going on, or did we just get a one-off benefit last year for some weird reason?

Simon Willison

My uninformed hunch is low-hanging fruit. I feel like, up until a year ago, people hadn’t been focusing on efficiency at all. It was all about what we could get these weird-shaped things to do.

And now, once we’ve hit that point of, “Okay, we know that we can get them to do what GPT-4 can do,” thousands of researchers around the world are focusing on how to make this more efficient. What are the most important things? How do we strip out all of the weights that have stuff in them that doesn’t really matter? All of that kind of thing.

Maybe 2024 was a freak year in which all of the low-hanging fruit came out at once, and we’ll actually see a reduction in that rate of improvement in terms of efficiency. I wonder. I think we’ll know for sure in about 3 months’ time if that trend is going to continue or not.

swyx

Yeah. I think the other thing you mentioned—the DeepSeek V3 gift that was given from DeepSeek over Christmas—but I feel like the other thing that might be underrated was DeepSeek R1.

Simon Willison

Mm-hmm. Yeah.

swyx

It’s a reasoning model you can run on your laptop, and I think that’s something that a lot of people are looking ahead to this year.

Simon Willison

Oh, did they release the weights for that one?

swyx

Yeah.

Simon Willison

Oh my goodness, I missed that. I’ve been playing with Qwen. The other big Chinese AI lab is Alibaba’s Qwen.

swyx

Oh.

Simon Willison

Alibaba’s Qwen.

swyx

Actually, yeah. Sorry.

Simon Willison

Yes.

No.

swyx

R1 is the API available.

Simon Willison

Yeah, exactly. Qwen, that’s really cool.

Alibaba’s Qwen has released 2 reasoning models that I’ve run on my laptop now. The first one was QwQ, and then the second one was QVQ, because the second one is a vision model, so you can give it vision puzzles and a prompt.

These things are so much fun to run because they think out loud. OpenAI o1 sort of hides its thinking process. The Qwen ones don’t; they just churn away. You’ll give it a problem, and it will output literally dozens of paragraphs of text about how it’s thinking.

My favorite thing that happened with QwQ is that I asked it to draw me a pelican on a bicycle in SVG. That’s my standard stupid prompt. For some reason, it thought in Chinese. It spat out a whole bunch of Chinese text onto my terminal on my laptop, and then at the end it gave me quite a good artistic take on a pelican on a bicycle.

I ran it all through Google Translate, and it was contemplating the nature of SVG files as a starting point. The fact that my laptop can think in Chinese now is so delightful. It’s so much fun watching it do that.

swyx

Yeah. I think Andrej Karpathy was saying that we know we’ve achieved proper reasoning inside these models when they stop thinking in English, and perhaps the best form of thought is in Chinese.

For listeners who don’t know, whenever a new model comes out, Simon’s blog is always the first place to run Pelican Bench. I don’t know how you do it, but you’re always the first to run these models.

Simon Willison

I just did it for Phi-4 this morning.

swyx

And you post up the results.

Simon Willison

Yeah.

swyx

So I really appreciate that. You should check it out. These are not theoretical; Simon’s blog actually shows them.

3. The Agent Reliability Problem

Brian McCullough

Let me put on the investor hat for a second. From the investor side of things, a lot of the VCs that I know are really hot on agents, and this is the year of agents. But last year was supposed to be the year of agents as well. There was lots of money flowing toward agentic startups.

In your piece, you suggest there’s a fundamental flaw in AI agents as they exist right now. Let me quote you, and then I’d love to dive into this.

You said, “I remain skeptical as to their ability based, once again, on the challenge of gullibility. LLMs believe anything you tell them. Any systems that attempt to make meaningful decisions on your behalf will run into the same roadblock. How good is a travel agent or a digital assistant, or even a research tool if it can’t distinguish truth from fiction?”

So essentially, what you’re suggesting is that the state of the art now that allows agents is still that sort of 90% problem—the edge problem of getting to 100%? Or is there a deeper flaw? What are you saying there?

Simon Willison

So this is the fundamental challenge here. Honestly, my frustration with agents is mainly around definitions. If you ask anyone who says they’re working on agents to define agents, you will get a subtly different definition from each person. But everyone always assumes that their definition is the one true one that everyone else understands.

I feel like a lot of these agent conversations have people talking past each other, because one person is talking about the travel-agent idea of something that books things on your behalf, while somebody else is talking about LLMs with tools running in a loop with a cron job somewhere, and all of these different things. You ask academics, and they’ll laugh at you because they’ve been debating what agents mean for over 30 years at this point. It’s this long-running, almost sort of an in-joke in that community.

But if we assume that, for the purpose of this conversation, an agent is something which you can give a job and it goes off and does that thing for you, like booking travel or things like that, the fundamental challenge is the reliability issue, which comes from this gullibility problem. A lot of my interest in this originally came from thinking about prompt injection.

Brian McCullough

Right, right.

Simon Willison

It’s this form of attack against LLM systems where you deliberately lay traps out there for the LLM to stumble across.

Brian McCullough

And I should say, you’ve been banging this drum that no one’s gotten very far, at least, on solving this, that I’m aware of, right? That’s still an open problem.

Simon Willison

Right. For 2 years—

Brian McCullough

Yeah, right.

Simon Willison

We’ve been talking about this problem, and a great illustration of this was Claude. Anthropic released Claude Computer Use a few months ago. It was a fantastic demo. You could fire up a Docker container, and you could literally tell it to do something and watch it open a web browser, navigate to a web page, click around, and so forth. It was really, really interesting and fun to play with.

One of the first demos somebody tried was, what if you give it a web page that says, “Download and run this executable”? And it did, and the executable was malware that added it to a botnet. The very first, most obvious dumb trick that you could play on this thing just worked, right? So that’s obviously a really big problem.

If I’m going to send something out to book travel on my behalf, it’s hard enough for me to figure out which airlines are trying to scam me and which ones aren’t. Do I really trust a language model that believes the literal truth of anything that’s presented to it to go out and do those things?

swyx

It’s interesting to see Anthropic doing this because they used to be the safety arm of OpenAI that split out and said, “We’re worried about letting this thing out in the wild.” And here they are, enabling computer use for agents. It feels like things have merged.

I’m also fairly skeptical about this always being the year of Linux on the desktop. This is the equivalent of this being the year of agents: people are not predicting so much as wishfully thinking and hoping and praying for their companies and agents to work. But I feel like things are coming along a little bit.

To me, it’s kind of like self-driving. I remember in 2014 saying that self-driving was just around the corner, and I mean, it kind of is, in the Bay Area.

Simon Willison

And then you get in a Waymo and you’re like, “Oh, this works.”

swyx

Yeah, but it’s a slow cook.

Simon Willison

Right.

swyx

It’s a slow cook.

Simon Willison

Yeah.

swyx

Over the next 10 years, we’re going to hammer out these things, and the cynical people can just point to all the flaws, but there are measurable or concrete progress steps that are being made by these builders.

Simon Willison

So there is one form of agent that I believe in. I mostly believe in the research-assistant form of agents.

swyx

Yes. I was going to say.

Simon Willison

The thing where you’ve got a difficult problem. I’m on the beta for Google Gemini 1.5 Pro with Deep Research, I think it’s called.

swyx

Oh, God. These names.

Simon Willison

These names, right? But I’ve been using that. It’s good, right? You can give it a difficult problem, and it tells you, “Okay, I’ve gone and looked at 56 different websites,” and it goes away and dumps everything into its context, and it comes up with a report for you.

And it won’t work against adversarial websites, right? If there were websites with deliberate lies in them, it might well get caught out. Most things don’t have that as a problem, and so I’ve had some answers from that which were genuinely really valuable to me.

That feels to me like I can see how, given existing LLM technology—especially with Google Gemini and its million-token context, and Google with its crawl of the entire web, its search, its cache of every page, and so forth—that makes sense to me. What they’ve got right now, I don’t think it’s as good as it can be, obviously, but it’s a really useful thing, which they’re going to start rolling out.

Perplexity has been building the same thing for a couple of years. That I believe in. If you tell me that you’re going to have a research-assistant agent, great. The coding agents—I mean, ChatGPT Code Interpreter, nearly 2 years ago, started writing Python code, executing the code, getting errors, and rewriting it to fix the errors. That pattern obviously works. That works really, really well, and they’re going to keep on getting better, and that’s going to be great.

The research-assistant agents are just beginning to get there. The things I’m critical of are the ones where you trust the thing to go out and act autonomously on your behalf and make decisions on your behalf, especially involving spending money. I don’t see that working for a very long time. That feels to me like an AGI-level problem.

swyx

It’s funny because I think Stripe actually released an agent toolkit, which is one of the things I featured. It’s trying to enable these agents each to have a wallet that they can spend from. Basically, it’s a virtual card. It’s not that difficult with modern infrastructure.

Simon Willison

Yeah. If I can stick a $50 cap on it, then at least it can’t—

swyx

Yeah, whatever.

Simon Willison

It can’t lose more than $50.

Brian McCullough

I don’t know if either of you know Rafat Ali. He runs Skift, which is a travel-news vertical, and he constantly laughs at the fact that every agent thing is, “We’re going to get rid of booking a plane flight for you.”

I would point out that historically, when the web started, the first thing everyone talked about was that you can go online and book a trip, right? So it’s funny: for each generation of technological advance, the thing they always want to kill is the travel agent, and now they want to kill—

swyx

Right.

Simon Willison

And it’s like, I use Google Flights. It’s great, right? If you gave me an agent to do that for me, it would save me—maybe 15 seconds of typing in my details—but I still want to see what my options are and go, “Yeah, I’m not flying on that airline no matter how cheap they are.”

swyx

For listeners, I think both of you are pretty positive on NotebookLM, and we actually interviewed the NotebookLM creators. There are actually 2 internal agents going on internally. The reason it takes so long is because they’re running an agent loop inside that is fairly autonomous, which is kind of interesting.

Simon Willison

For one definition of an agent loop, if you pick that—

swyx

For one definition—

Simon Willison

—that particular one. And you’re talking about the podcast side of this, right?

swyx

Yeah. The podcast side of things. There’s going to be a new version coming out that we’ll be featuring at our conference.

Simon Willison

That one’s fascinating to me. NotebookLM, I think it’s 2 products, right? On the one hand, it’s actually a very good RAG product. You dump a bunch of things in, and you can run searches. It does a good job of that.

swyx

That’s what it always was. Yeah.

Simon Willison

And then they added the podcast thing. It’s a total gimmick, right?

swyx

Right.

Simon Willison

But that gimmick got them attention because they had a great product that nobody paid any attention to at all, and then you add the unfeasibly good voice synthesis of the podcast.

Brian McCullough

But it’s the lesson—

Simon Willison

It’s just brutally brilliant.

Brian McCullough

It’s the lesson of Midjourney and stuff like that. If you can create something that people can post on social media, you don’t have to lift a finger again to do any more marketing for what you’re doing.

Mm-hmm.

Let me dig into NotebookLM just for a second as a podcaster. As a gimmick, it makes sense, and then obviously, you dig into it, and it sort of has problems around the edges. It does the thing that all LLMs do where it’s like, “Oh, we want to wrap up with a conclusion.” I always call that the eighth-grade book report paper problem, where it has—

Simon Willison

Yep.

Brian McCullough

—to have an intro and then, you know. But that’s sort of a thing where I think you spoke about this again in your piece at year-end, about how things are going multimodal and how there are things that you didn’t expect, like vision and especially audio. So that’s another thing where, at least over the last year, there’s been progress made that maybe you didn’t think was coming as quickly as it came.

4. AI Goes Multimodal

Simon Willison

I don’t know. A year ago, we had 1 really good vision model. We had GPT-4 Vision, which was very impressive, and Google Gemini had just dropped Gemini 1.0, which had vision, but nobody had really played with it yet. People weren’t taking Gemini seriously at that point. I feel like it was Gemini 1.5 Pro when it became apparent that they had got over their hump and were building really good models.

To be honest, the video models are mostly still using the same trick: the thing where you divide the video up into 1 image per second and dump that all into the context. So maybe it shouldn’t have been so surprising to us that long-context models plus vision meant that video was starting to be solved. What you really want with video is to be able to do the audio and the images at the same time, and I think the models are beginning to do that now.

Originally, Gemini 1.5 Pro ignored the audio. It just did the 1-frame-per-second video trick. As far as I can tell, the most recent ones are actually doing pure multimodal. But the things that opens up are just extraordinary. The ChatGPT iPhone app feature that they shipped as one of their 12 Days of OpenAI—I really can be having a conversation and just turn on my video camera and go, “Hey, what kind of tree is this?” And so forth, and it works.

For all I know, that’s just snapping a picture once a second and feeding it into the model. But the things that you can do with that as an end user are extraordinary. I don’t think most people have cottoned on to the fact that you can now stream video directly into a model because it’s only a few weeks old. But wow, that’s a big boost in terms of what kinds of things you can do with this stuff.

swyx

Yeah. For people who are not that close, I think Gemini Flash’s free tier allows you to do something like capture a photo—1 photo every second or a minute—and leave it on 24/7, and you can prompt it to do whatever. So you can effectively have your own camera app or monitoring app that you just prompt, and it detects changes, detects alerts or anything like that, or describes your day. And the fact that this is free also leads into the previous point of prices having come down a lot.

Simon Willison

Even if you’re paying for this stuff, a thing that I put in my blog entry is that I ran a calculation on what it would cost to process 68,000 photographs in my photo collection and, for each one, just generate a caption. Using Gemini 1.5 Flash-8B, it would cost me $1.68 to process 68,000 images, which is—I mean, that doesn’t make sense. None of that makes sense.

It’s 1/400 of a cent per image to generate captions now. So you can see why feeding in a day’s worth of video just isn’t even very expensive to process.

swyx

Yeah. I’ll tell you what is expensive: it’s the other direction. Here, we’re talking about consuming video. This year, we also had a lot of progress. Probably one of the most anticipated launches of the year was Sora. We actually got Sora, and less exciting—

Simon Willison

We did, and then Veo 2—Google’s Sora—came out like 3 days later and upstaged it. Sora was exciting—

swyx

In general, I feel the media or social media has been very unfair to Sora because what was released to the world, generally available, was Sora Lite, the distilled version of Sora.

Right? So you’re—

Simon Willison

I did not realize that.

swyx

You’re absolutely comparing—

Simon Willison

Ah, okay.

swyx

—the most cherry-picked version of Veo 2, the one that they published on the marketing page—

Simon Willison

Yeah.

swyx

—to the most embarrassing version of Sora. So, of course, it’s going to look bad.

Simon Willison

Well, I got access to Veo 2. I’m in the Veo 2 beta, and I’ve been poking around with it and getting it to generate pelicans on bicycles and stuff.

swyx

I would absolutely believe that Veo 2 is actually better.

Simon Willison

That’s interesting. Is Sora—is full-fat Sora coming soon? Do you know? When do we get to play with that one?

swyx

No one’s mentioned anything. I think basically the strategy is: let people play around with Sora Lite and get info there, but keep developing Sora with the Hollywood studios. That’s what they actually care about.

Simon Willison

Gotcha. Okay.

swyx

The rest of us don’t really know what to do with the video anyway.

Simon Willison

Right. My thing is, I realize that for generative images and video—images we’ve had for a few years—I don’t feel like they’ve broken out into the talented artist community yet. Lots of people are having fun with them and producing stuff that’s kind of cool to look at.

But what I want is—you know, that movie, Everything Everywhere All at Once, right? It won a ton of Oscars, an utterly amazing film. The VFX team for that were 5 people.

swyx

Yeah.

Simon Willison

Some of whom were watching YouTube videos to figure out what to do. My big question for Sora and Midjourney and stuff is: what happens when a creative team like that starts using these tools? I want the creative geniuses behind Everything Everywhere All at Once—what are they going to be able to do with this stuff in a few years’ time? Because that’s really exciting to me.

That’s where you take artists who are at the very peak of their game, give them these new capabilities, and see what they can do with them.

swyx

I should—I know a little bit here, so I should mention that that team actually used Runway ML. So there was—

Simon Willison

No way. In that movie?

swyx

Yeah. I don’t know how much, so it’s possible to overstate this. But there are people integrating generative video within their workflow, even pre-Sora.

Simon Willison

Wow.

swyx

Yeah.

Brian McCullough

It’s not the thing where it’s like, okay, tomorrow we’ll be able to do a full 2-hour movie that you prompt with 3 sentences. For the very first part of video effects in film, if you can get that 3-second clip, if you can get that 20-second thing that they did in The Matrix that blew everyone’s minds and took $1 million or whatever to do, it’s the little bits and pieces that they can fill in now that are probably already there.

swyx

Yeah. I think having a layered view of what assets people need and letting AI fill in the low-value assets—the background video, the background music, and sometimes the sound effects—may be more palatable. Maybe it also changes the way that you evaluate the stuff that’s coming out, because people tend to, in social media, try to emphasize foreground stuff, main-character stuff.

So you really care about consistency, and you really are bothered when, for example, Sora botches an image generation of a gymnast doing flips, which is horrible.

Simon Willison

That’s hilarious.

swyx

It’s horrible. But for background crowds—

Brian McCullough

Right.

swyx

—who cares?

Brian McCullough

And by the way, again, I was a film major way, way back in the day. That’s how it started: things like Braveheart, where they filmed 10 people on a field, and then the computer could turn it into 1,000 people on a field. That’s always been the way. It’s around—

Simon Willison

Right. The Lord of the Rings, right?

Brian McCullough

Yeah.

Simon Willison

The Lord of the Rings movies were over 20 years ago. They had those giant battle sequences, which were very early. You could almost call it a generative AI approach, right? They were using very sophisticated algorithms to model out those different battles and all of that kind of stuff.

Brian McCullough

Yeah.

Simon Willison

Yeah, I know very little. I know basically nothing about film production, so I try not to commentate on it. But I am fascinated to see what happens when these tools start being used by the people at the top of their game.

swyx

I would say there’s a cultural war being fought here more than a technology war. Most of the Hollywood people are against any form of AI anyway, so they’re busy fighting that battle instead of thinking about how to adopt it. It’s very fringe. I participated here in San Francisco in a generative AI video creative hackathon where the AI-positive artists actually met with technologists like myself, and then we collaborated to build short films. That was really nice, and I think I’ll be hosting some of those at my events going forward.

One thing that I want to give people a sense of is that this is a recap of last year, but sometimes it’s useful to walk away with what we can expect in the future. I don’t know if you got anything. I would also call out that the Chinese models here have made a lot of progress. Hai Luo and Kling, and God knows who else in the video arena, are also making a lot of progress. It’s surprising. I think maybe, actually, China is surprisingly ahead with regard to open weights, at least, but also just specific forms of video generation.

Simon Willison

Wouldn’t it be interesting if a film industry sprang up in a country that we don’t normally think of as having a really strong film industry, and that was using these tools? That would be a fascinating sort of angle on this.

swyx

Agreed.

Brian McCullough

Oh, sorry.

swyx

Yeah, go ahead.

Brian McCullough

Just to put it on people’s radar as well, HeyGen—there’s a category of video avatar companies that don’t specifically specialize in general video. They only do talking heads, let’s just say. And HeyGen’s done very well.

swyx

Brian, Brian, you know that’s what I’ve been using, right? So if you see some of my recent YouTube videos and things like that, the beauty part of the HeyGen thing is I don’t want to use the robot voice, so I record the MP3 file for my clips every single day, and then I put that into HeyGen with the avatar that I’ve trained it on. All it does is the lip sync.

It’s not 100% uncanny-valley-beatable, but it’s good enough that, if you weren’t looking for it, it’s just me sitting there doing one of my clips from the show. And, yeah, so by the way, HeyGen—shout-out to them.

Brian McCullough

In terms of the look-ahead—going forward, reviewing 2024 and looking at trends for 2025—I would basically call this out. Meta tried to introduce AI influencers and failed horribly because they were just bad at it. But at some point, there will be more and more AI influencers, not in the way that Simon is, but in a way that they are not human.

The few of those that have done well, I always feel like they’re doing well because it’s a gimmick, right? It’s novel and fun. Like the AI Seinfeld thing from last year, the Twitch stream—those, if you’re the only one, or one of just a few doing that, will attract an audience because it’s an interesting new thing. But I just don’t know if that’s going to be sustainable longer term or not.

I’m going to tell you, because I’ve had discussions—I can’t name the companies or whatever—but think about the workflow for this. Now we all know that on TikTok and Instagram, holding up a phone to your face and doing an “in my car” video or a walking-and-talking video is very common.

But also, if you want to do a professional sort of talking-head video, you still have to sit in front of a camera, you still have to do the lighting, and you still have to do the video editing. Versus if you can just record what I’m saying right now—the last 30 seconds—if you clip that out as an MP3 and you have a good enough avatar, then you can put that avatar in front of Times Square, on a beach, or whatever.

So, again, for creators, the reason I think, Simon, we’re on the verge of something is that it’s not going to be that AI avatars take over. It’ll be one of those things where it takes another piece of the workflow out and simplifies it.

Gotcha. I am all for that. I always love tools—tools that help human beings do more ambitious things. I’m always in favor of that. That’s what excites me about this entire field.

We’re looking into basically creating one for my podcast. We have this guy Charlie. He’s Australian, he’s not real, but he opens every show, and we’re going to have him present all the shorts. Yeah, go ahead.

5. Credibility Beats AI Slop

The thing that I keep coming back to is this idea of credibility. In a world that is full of AI-generated everything and so forth, it becomes even more important that people find the sources of information that they trust, and find people and sources that are credible.

I feel like that’s the one thing that LLMs and AI can never have, is credibility, right? ChatGPT can never stake its reputation on telling you something useful and interesting because that means nothing, right? It’s a matrix multiplication. It depends on who prompted it and so forth.

I’m always—and this is when I’m blogging as well—I’m always looking for the reliable people who will tell me useful, interesting information, who aren’t just going to tell me whatever somebody’s paying them to tell them, and who aren’t going to type a one-sentence prompt into an LLM, spit out an essay, and stick it online.

To me, earning that credibility is really important. That’s why a lot of my ethics around the way that I publish are based on the idea that I want people to trust me. I want to do things that gain credibility in people’s eyes so they will come to me for information as a trustworthy source. And it’s the same for the sources that I’m consulting as well. I’ve been thinking a lot about that sort of credibility focus for a while now.

You can layer or structure credibility, or decompose it. One thing I would put in front of you—I’m not saying that you should agree with this or accept this at all—is that you can use AI to generate different variations, and then you, as the final sort of last-mile person, pick the last output and put your stamp of credibility behind that, that everything is human-reviewed instead of human-originated, if that’s the thing.

If you publish something, you need to be able to be proud of publishing it. You need to be able to say, “I will put my name to this. I will attach my credibility to this thing.”

And if you’re willing to do that, then that’s great. For creators, this is huge because there’s a fundamental asymmetry between starting with a blank slate versus choosing from 5 different variations.

The key thing that you just said is that if everything that I do, if all of the words were generated by an LLM, if the voice is generated by an LLM, and if the video is also generated by an LLM, then I haven’t done anything, right? But if you take a shortcut on one or 2 of those, and I’m still willing to sign off on it, I feel like that’s where people are coming around to: this is maybe acceptable.

This is where I’ve been pushing the definition. I love the term “slop,” where I’ve been pushing the definition of slop as AI-generated content that is both unrequested and unreviewed. The unreviewed thing is really important.

The thing that elevates something from slop to not slop is if a human being has reviewed it and said, “You know what? This is actually worth other people’s time.” And again, I’m willing to attach my credibility to it and say, “Hey, this is worthwhile.”

It’s the curatorial and editorial part of it that, no matter what the tools are to do shortcuts—to do, as swyx is saying, choosing between different edits or different cuts—still has a curatorial mind or editorial mind behind it.

6. The GUI Moment For LLMs

Let me wedge this in before we start to close. One of the things that, coming back to your year-end piece, has been something I’ve been banging the drum about is when you’re talking about LLMs getting harder to use. You said most users are thrown in at the deep end.

The default LLM chat UI is like taking brand-new computer users, dropping them into a Linux terminal, and expecting them to figure it all out. I mean, it’s literally going back to the command line. The command line was defeated by the GUI interface, and what I’ve been banging the drum about is that this cannot be the user interface. What we have now cannot be the end result.

Do you see any hints or seeds of a GUI moment for LLM interfaces? I mean, it has to happen. It absolutely has to happen. The usability of these things is turning into a bit of a crisis, and we are at least seeing some really interesting innovation in little directions.

Simon Willison

Just like OpenAI's ChatGPT Canvas thing that they just launched. That is at least a little more interesting than just chats and responses. You're exploring that space where you're collaborating with an LLM, both working on the same document. That makes a lot of sense to me. That feels really smart.

One of the best things is still—who was it who did the UI where you could draw an interface and click a button? tldraw with their Make It Real thing. That was spectacular. Absolutely spectacular, like an alternative vision of how you'd interact with these models. So I feel like there is so much scope for innovation there, and it is beginning to happen. I feel like most people do understand that we need to do better in terms of interfaces that both help explain what's going on and give people better tools for working with models.

Brian McCullough

I was going to say, I want to dig a little deeper into this because think of the conceptual idea behind the GUI. Instead of typing into a command line, "open word.exe," you click an icon, right? That's abstracting away the programming stuff. A child can tap on an iPad and make a program open, right?

The problem, it seems to me, with how we're interacting with LLMs right now is it's sort of like a dumb robot where you poke it and it goes over here. But no, I want it to go over here, so you poke it this way, and you can't get it exactly right. What can we abstract away from what's currently going on that makes it more fine-tuned and easier to get more precise? You see what I'm saying?

Simon Willison

Okay.

Simon Willison

Yes. And this is the other trend that I've been following from the last year, which I think is super interesting. It's the prompt-driven UI development thing. Basically, this is the pattern where Claude Artifacts was the first thing to do this really well. You type in a prompt, and it goes, "Oh, I should answer that by writing a custom HTML and JavaScript application for you that does a certain thing."

Since then, it turns out this is easy, right? Every decent LLM can produce HTML and JavaScript that does something useful. So we've actually got this alternative way of interacting where they can respond to your prompt with an interactive, custom interface that you can work with.

People haven't quite wired those back up again. Ideally, I'd want the LLM to be able to ask me a question where it builds me a custom little UI for that question, and then it gets to see how I interacted with that. I don't know why, but that's such a small step from where we are right now. That feels like such an obvious next step.

Why should you just be communicating with text when it can build interfaces on the fly that let you select a point on a map or move sliders up and down, all of that kind of stuff?

It's gonna—like knobs and dials. I keep saying knobs and dials.

Simon Willison

Knobs and dials, right.

Brian McCullough

Yeah, exactly.

Simon Willison

We can do that, and the LLMs can build it. Claude Artifacts will build you a knobs-and-dials interface, but at the moment, they haven't closed the loop. When you twiddle those knobs, Claude doesn't see what you're doing. They're going to close that loop. I'm shocked that they haven't done it yet.

I think there's so much scope for innovation, and there's so much scope for doing interesting stuff with that model, where anything you can represent in HTML, JavaScript, and SVG, which is almost everything, can now be part of that ongoing conversation.

swyx

Yeah. I would say the best-executed version of this I've seen so far is Bolt, where you can literally type in, "Make a Spotify clone. Make an Airbnb clone," and it actually does that for you zero-shot with a nice design.

Simon Willison

Did you see there's a benchmark for that now?

swyx

Yeah.

Simon Willison

The LMArena people now have a—

swyx

LMArena benchmark.

Simon Willison

A benchmark for zero-shot app generation, because all of the models can do it. I've started figuring out how to—I'm building my own version of this for my own project because I think—

swyx

Oh.

Simon Willison

Within 6 months, I think it'll just be an expected feature. For my dataset data exploration project, I want you to be able to do things like conjure up a dashboard just via prompt. You say, "I need a pie chart and a bar chart, put them next to each other, and then have a form where submitting the form inserts a row into my database table."

This is all suddenly feasible. It's not even particularly difficult to do, which is utterly bizarre: these things are now easy.

swyx

Yeah. I think for a general audience, that is what I would highlight: software creation is becoming easier and easier. Gemini is now available in Gmail and Google Sheets. I don't write my own Google Sheets formulas anymore; I just tell Gemini to do it.

I almost want to somewhat disagree with your assertion that LLMs got harder to use.

Simon Willison

Ooh.

swyx

We expose more capabilities, but they're in minor forms, like using Canvas, web search in ChatGPT, and Gemini being in Google Sheets. We're getting improvements.

Simon Willison

No, no, no. Those are the things that make it harder.

swyx

Okay.

Simon Willison

The problem is that for each of those features, they're amazing if you understand the edges of the feature. If you're like, "Okay, so in Google Sheets formulas, I can get it to do a certain amount of things, but I can't get it to go and read a web page..." You probably can get it to read a web page, right? But there are things that it can do and things that it can't do, which are completely undocumented.

If you ask it what it can and can't do, they're terrible at answering questions about that. My favorite example is Claude Artifacts. You can't build a Claude Artifact that can hit an API somewhere else because the CORS headers on that iframe prevent accessing anything outside of CDNs.

swyx

I hate those.

Simon Willison

People are learning CORS headers as an end user in order to understand why—

swyx

I hate those.

Simon Willison

I've seen people saying, "Oh, this is rubbish. I tried building an artifact that would run a prompt, and it couldn't," because Claude didn't expose an API with CORS headers. All of this stuff is so weird and complicated.

The more tools we add, the more expertise you need to really understand the full scope of what you can do. The question really comes down to: What does it take to understand the full extent of what's possible? Honestly, that's just getting more and more involved over time.

7. Local Models Return

swyx

Yeah. I have one more topic that I think you're kind of a champion of, and we've touched on it a little bit, which is local LLMs and running AI applications on your desktop. I feel like you are an early adopter of many, many things.

Simon Willison

Well, I had an interesting experience with that over the past year. 6 months ago, I almost completely lost interest. The reason is that 6 months ago, the best local models you could run—there was no point in using them at all because the best hosted models were so much better.

There was no point at which I'd choose to run a model on my laptop if I had API access to Claude 3.5 Sonnet. They just weren't even comparable. That changed basically in the past 3 months as the local models had this step change in capability. Now I can run some of these local models, and they're not as good as Claude 3.5 Sonnet, but they're not so far away that it's not worth me even using them.

The continuing problem is I've only got 64 GB of RAM, and if you run Llama 3 70B, most of my RAM is gone. So now I have to shut down my Firefox tabs, Chrome, and VS Code windows in order to run it.

But it's got me interested again. The efficiency improvements are such that now, if you were to stick me on a desert island with my laptop, I'd be very productive using those local models, and that's pretty exciting. If those trends continue, and I think my next laptop, when I buy one, is going to have twice the amount of RAM, maybe I can run almost the top-tier open-weight models and still be able to use it as a computer as well.

NVIDIA just announced their $3,000, 128-gigabyte monstrosity. That's a pretty good price.

swyx

You gonna buy it?

swyx

Custom OS and all.

Simon Willison

If I get a job. If I have enough of an income that I can justify blowing $3,000 on it, then yes.

swyx

Okay. Let's do a GoFundMe to get Simon one of them. Come on. You know you can get a job anytime you want.

Simon Willison

I want a job that pays me to do exactly what I'm doing already and doesn't tell me what else to do. That's the challenge.

swyx

This is just purely discretionary. I think Ethan Mollick does pretty well, whatever it is he's doing. But basically, I was trying to bring in not just local models but Apple Intelligence, which is on every M-series Mac. You seem skeptical.

Simon Willison

It's rubbish.

swyx

It's rubbish.

Simon Willison

Apple Intelligence is so bad.

swyx

It does one thing well.

Simon Willison

Oh, yeah. What's that?

swyx

It summarizes notifications, and sometimes it's humorous.

Brian McCullough

But are you sure it does that well?

swyx

It's decent.

Brian McCullough

The other thing, again, from a sort of normie point of view, is that there's no indication from Apple of when to use it. Everybody upgrades their thing, and it's like, “Okay, now you have Apple Intelligence,” and you never know when to use it ever again.

swyx

Oh, yeah, you consult the Apple docs, which is MKBHD.

Simon Willison

The one thing I'll say about Apple Intelligence is that one of the reasons it's so disappointing is that the models are just weak. But now that—

swyx

Yeah.

Simon Willison

Llama 3B is such a good model in a 2-gigabyte file. I think, give Apple 6 months, and hopefully they'll catch up to—

swyx

Yeah.

Simon Willison

—the state of the art on their small models, and then maybe it'll start being a lot more interesting.

swyx

Anyway, this was year 1. Just like the first year of the iPhone, maybe it wasn't that much of a hit, and then in year 3 they had the App Store. So I would say give it—

Simon Willison

Yeah.

swyx

—some time. I think Chrome is also shipping Gemini Nano this year in Chrome, which means that every app, every web app, will have free access to a local model that just ships in the browser, which is kind of interesting.

I also wanted to open the floor to any of us: what are the AI applications that we've adopted that we really recommend? These are all apps that are running in a browser, or apps that are running locally, that other people should be trying, right? I feel like that's always one thing that's helpful at the start of the year.

Simon Willison

Okay. So, for running local models, my top picks are, firstly, on the iPhone, this thing called MLC Chat.

swyx

Mm-hmm.

Simon Willison

It works, it's easy to install, and it runs Llama 3B. It's so much fun. It's not necessarily a capable enough model for me to use it for real things, but my party trick right now is to get my phone to write a Netflix Christmas movie plot outline where a jeweler falls in love with the King of Sweden or whatever. It does a good job, and it comes up with pun names for the movies. That's deeply entertaining.

On my laptop, most recently, I've been getting heavily into Ollama because the—

swyx

Yeah.

Simon Willison

—Ollama team are very good at finding the good models, packaging them up, and making them work well. It gives you an API. My little LLM command-line tool has a plugin that talks to Ollama, which works really well. Ollama is, I think, the easiest on-ramp to running models locally. If you want a nice user interface, LM Studio is, I think, the best user interface for that. It's not open source, but it's good. It's worth playing with.

The other one that I've been toying with recently is called Open WebUI. The UI is fantastic. If you've got Ollama running and you fire this thing up, it spots Ollama and gives you an interface to your Ollama models, and that's really nicely done. That's my current favorite open-source UI for these things.

There are lots of good options. You do need a lot of disk space. The models start at 2 gigabytes for the 3B models that are actually worth playing with. The really impressive ones tend to be in the 20- to 30-gigabyte range, in my experience.

swyx

I think my struggle here is that I'm not much of an absolutist about running things locally. I'm happy to call an API.

Simon Willison

Mm-hmm.

swyx

Okay, yeah. But I just think—

Simon Willison

I do it to play.

swyx

Yeah, okay, fine.

Simon Willison

It's my research interest, yeah.

Brian McCullough

Answer your own question. Give us more apps that you want to—

swyx

Yeah. Sometimes it's just nice to recommend apps. I use Super Whisperer now. I tried Whisper Flow, but it didn't really work for me. Super Whisperer—

Simon Willison

Mm.

swyx

—is one of them. It basically replaces typing. You should just talk most of the time, especially if you're doing anything long-form. I hold down Caps Lock and talk, and when I'm done, I lift it up. It uses—

It isn't just about writing down your transcripts, because I make ums and uhs all the time, and I restate myself all the time. But I use GPT-4 to rewrite, and that's what these guys are doing. They're all doing some form of state-of-the-art ASR—automatic speech recognition—and then an LLM to rewrite.

I would also recommend that people check out Rosebud for journaling. I think AI for mental health is quite unexplored, and it's not because we're trying to build AI therapists. I think therapists really hate that. You'll never be on the level of therapists.

Brian McCullough

That gets back to the human thing that we were discussing. On some level, there are certain things and disciplines that require the human touch, and that might be one of them.

swyx

Sure. But the human touch costs me $300 an hour.

Brian McCullough

Yes.

swyx

Right?

Brian McCullough

Yeah.

swyx

And this thing's $3 a month. There's a spectrum of people for whom that will work, and I think it's cheap now to try all these things.

Simon Willison

I'm going to throw in a quick recommendation for an app. MacWhisper is my favorite—

swyx

Yeah, that's your one.

Simon Willison

—desktop app. I love that thing.

Brian McCullough

Yeah.

Simon Willison

It runs Whisper, and you can do things like paste in the URL to a YouTube video, and it'll pull the audio and give you a transcript. So that's how I watch YouTube now—

swyx

Ah.

Simon Willison

—I slap it into MacWhisper, then I copy and paste into Claude, and I use the Claude web app to do things.

MacWhisper works with MP3 files. Every time I'm on a podcast, I dump the MP3 into MacWhisper, then I dump the transcript into Claude and say, “What should I put in the show notes?” It spits out a bullet-point list where it says, “Oh, you mentioned a dataset that you should link to,” that kind of thing. MacWhisper—I use it several times a day, to be honest. It's great.

swyx

Yeah.

Brian McCullough

I'm actually going to say one that is incredibly basic and, again, coming back to just my workflow. We are currently recording this on Riverside. Riverside is a great tool for recording video and audio, like we're doing right now.

I always use this as an example when folks ask, “What will AI do for me?” When I first started using Riverside, we were recording 3 different channels, right? You guys are recording locally, so there are 3 audio files and 3 video files. When I first started using Riverside, you had to pump 3 tracks into Adobe and then edit.

“Okay, now we focus on Simon. Now we focus on swyx. Now we focus on Brian. Now we do all 3.” One day, a tool popped up that said, “Hit this button,” and it was Smart Edit. The AI determines, “Okay, Simon has been talking for 30 minutes, so go to the full shot of him. Brian is now talking, or there's overtalk, so let's have all 3 talking heads.”

With one button, for anything I posted, it saved me 3 or 4 hours' worth of work. That, to me, is, again, if normies are listening—

Simon Willison

In fact, Riverside has that feature now.

Brian McCullough

Yeah.

swyx

Yeah, yeah.

Simon Willison

Damn.

Brian McCullough

I don't use it.

Simon Willison

Oh, that sounds fantastic.

Brian McCullough

I still use a human editor. The day it came out, I was running around the house telling my wife, telling anyone that would listen, “You don't know, I just saved 3 hours because they had a new feature.”

Simon Willison

Wow.

That's exciting.

Brian McCullough

Brian's basically crying with joy right now. All right, let's try to bring this to a landing a little bit. Simon, I have maybe 2 or 3 more. We can do these rapid-fire.

One of my shows—one of the things about my show is that it's sort of like Silicon Valley writ large, so it's sort of like the horse race of who's up and who's down or whatever.

To the degree that you're interested in pontificating on this, OpenAI as a company in 2025, do you see challenges coming? Are you bearish or bullish? I'm almost doing a CNBC sort of thing, but how do you feel about OpenAI this year?

8. OpenAI Faces New Challenges

Simon Willison

I think they're in a bit of trouble. They seem to have lost a lot of talent.

Brian McCullough

Mm.

Simon Willison

And they don't have that top-of-the-pile thing. If it wasn't for o3, they'd be in massive trouble because they'd have lost that. I think o3 clawed them back up again.

One of the big stories of 2024 is that OpenAI started as the clear leader, and now Google Gemini is really good. Google Gemini had an amazing year. Anthropic Claude, Claude 3.5 Sonnet, is still my personal favorite model, and that feels notable. Nobody would argue that they weren't the leader in all of this stuff a year ago, and today they're still doing great, but they're not as far ahead as they were.

Brian McCullough

Next question, and maybe this couldn't be as rapid-fire, but I loved, finally, from your piece the idea that LLMs need better criticism, which I'd love you to expand on. As I straddle this world of tech journalism, creator, investor, and all that stuff, I thought that you had a really interesting thing to say about how—and we even alluded to this—Hollywood is against it.

Better criticism, in the sense that, as I took it, everybody's got their hackles up. They're trying to defend their livelihoods and things like that. But it's either, “This is going to destroy my job and destroy the world,” or—I'm sorry, I'm again leading the witness—what did you mean by LLMs needing better criticism?

Simon Willison

This is a frustration I have. If I read a discussion thread somewhere about this topic, I can predict exactly what everyone's going to say. People talk about the environmental impact. They talk about the plagiarism of the training data and the unlicensed training data. They'll often say, “Oh, and these things are completely useless.”

That's the one that I will push back against. The other things are true, right? The argument I always make about the idea that LLMs are just completely useless is that they are very useful if you understand how to use them, which is distinctly unintuitive. You have to learn how to deal with something that will just wildly hallucinate and make things up, and all of those kinds of things.

If you can learn what they're good at and what they're bad at, I use them dozens of times a day, and I get enormous value out of them. So I'll push back on people who say, “No, they're just useless.”

But the other things—the environmental impact and the way the training data works—I feel like the training-data one is interesting because it's probably legal under fair use, but it's clearly unfair if somebody takes your work without your permission and trains a model that then competes with you in the marketplace. Legal or not, I understand why people are upset about that. That's a reasonable thing to be upset by.

So what I want—and I also feel like this stuff can have a major impact on society, especially as it starts undermining all sorts of jobs that we never thought were going to be undermined by technology. Who thought it would come for artists and lawyers first? That's bizarre.

We need to have really high-quality conversations where we help people figure out what works and what doesn't work. We need people to be able to make good decisions about what to do with their careers, to embrace this stuff, and all of that sort of thing.

If we just get distracted by saying, “Yeah, but it's useless, plagiarism-driven, environmentally catastrophic,” even though those things represent quite a lot of truth, I don't think that's a useful message to lead with. I want to be having the much more interesting, high-level conversations: If there are negatives, how do we counter those negatives? If there are positives, how do we encourage those? How do we help people make good decisions about how to use this technology?

swyx

Yeah. I think where I see this the most is for people who are very internal. You and I are immersed in this every single day, so we're frankly tired of the same debates being recycled again and again.

I think what might be more useful or more impactful is the level at which it starts to hit regulation. Last year, we had a couple of very notable attempts at the White House level and in California to regulate AI, and those did not come to pass.

At some point, these criticisms bubble up to law, to matters of national security or national science and progress. I feel like there needs to be more information or enlightenment there, maybe, if only because they tend to be very trailing.

Simon Willison

Right.

swyx

My favorite example to pick on, which is very unfair of me, but whatever, is that the California SB 1047 act tried to cap compute at 10^25.

Simon Willison

Yeah. That was DeepSeek.

swyx

Exactly. And it also was exactly at the point at which we pivoted from training GPT-5 to o1, where we're no longer scaling pre-training compute. What I'm saying is that we're always trying to regulate the last war, and I don't think that works in a field that is—

Simon Willison

So—

swyx

—basically 8 years old.

Simon Willison

I think there are 2 areas of regulation I'm super interested in. One of them is that I do think regulating the way these things are used can work. The big example is that I don't want somebody's insurance claim denied by a black-box LLM where nobody can explain what it did. That just feels—

swyx

Oh, we have real laws for that. This is like redlining.

Simon Willison

Exactly. Take those laws, reinforce them, and update them for modern capabilities.

The other one is privacy. We've got this huge problem right now where people will refuse to use any of these tools because they don't trust that the things they say to them won't be trained on and then exposed to other people. There are lots of terms and conditions that you can read through and try to navigate around.

I would love there to be straightforward laws that people understand, where they know that their input isn't going to be used for training because there's a law that says, under these circumstances, that can't happen. It's basically taking our existing privacy laws, giving them a few more teeth, and reinforcing them without introducing cookie banners à la the European Union.

These things are always very risky. You can have all sorts of bad results if you don't design them correctly. But there's space for that, I think.

Brian McCullough

Yeah. When I read that piece and then when you just said, “Swyx said we're in the weeds on this every single day, so we're tired of hearing these arguments,” it reminds me of folks who are always into politics. Then they're mad at the people who don't care about politics until it's an election year, and they're like, “Well, you're a low-information voter because all you know is that the factory in your town got shut down, or there's inflation, or whatever, and so you vote one way or the other, but you haven't been paying attention.”

But that's kind of the point: You shouldn't expect normal people to pay attention, except for the fact that this might lose me my job. So you can't blame them for being—I don't know if reactionary is the word—or emotional.

If you're in the weeds, it's harder to keep everybody informed, and this is going to touch everybody, so I don't know.

Okay, so this is the very last one, and then we can wrap and do plugs and everything. Simon, this is for you. It was alluded to a little bit, and you might not have one, but if there's something this year that a generalist like me isn't aware is coming down the pipe that you think is going to be big in the AI space—and maybe Swyx, if you've got one too—what do you think it would be?

Simon Willison

I think for most people who haven't been paying attention, we know these things already. We know that the models are now almost free to run things against. The fact that you can now do video, stream video to a model—the thing where you can share your entire screen with a model and get feedback is going to be really useful.

Again, the privacy side of things really matters, though. I do not want some model just training on everything that it sees on my screen. But no, the stuff that's now possible as of a few months ago is enough. I don't need anything new. That's going to keep me busy all year.

Brian McCullough

Swyx, you got one?

swyx

Simon's always too content, and then he sees the next thing and he's like, “Oh yeah, that's great too.”

Simon Willison

Yep.

swyx

Okay. I love trying to be contrarian by asking, “What does everyone hate right now?” Remember, this time last year, we had just had CES and the Rabbit R1.

We had the Humane, right?

Brian McCullough

Yeah. Wearables. Yep.

9. AI Wearables Make A Comeback

swyx

Those are completely in the gutter. No one will touch them. They're toxic nuclear waste. Okay, this year is the year of wearables.

Brian McCullough

Yep.

Simon Willison

Huh.

Brian McCullough

I agree with you, by the way. That cycle always works out where you go to a CES and it's everything—hype, hype, hype, hype—and then 3 years later it becomes the thing, unless it's 3D TVs, in which case that was a mistake anyway. But yeah—

Simon Willison

Well, transparent TVs are the big thing—

Brian McCullough

Mm.

Simon Willison

—for the last couple of years. What the hell?

Brian McCullough

Yeah.

swyx

I think Simon may have got one of these, but there are a lot of people working on AI wearables here in SF. They are surprisingly cheap, surprisingly capable, with decent battery life, and they do useful things. We have to work out the privacy aspect, of course. But people like Limitless, which used to be called Rewind, I think—

Brian McCullough

Mm-hmm.

swyx

They're shipping one of these wearables that, based on your voice, only records your voice. So you opt in.

Simon Willison

Interesting. Right.

swyx

Right? And so you can have perfect memory if you want. You can have perfect memory at work. Your employer can buy these for you; it only applies at work, and it's fine. It's just a meeting aid.

Lots of people use Granola or some kind of Fireflies or some of these meeting recorders only for online meetings, but what about in-person meetings? What about conversations and locations that you've been to? Some of that should be a choice. Right now you have zero choice. And I think these wearables will enable some of that.

It's up to us as a society to determine what's acceptable and what's not. I really like these gray areas where we still don't know yet. Whenever I tell people about this, they're like, “I don't know.” I guess it's as though you have perfect memory, but some people have better memory than others. Where's the line?

Brian McCullough

Hmm. And—

swyx

There will be a lot—

Brian McCullough

Now I—

swyx

A lot more of these.

Brian McCullough

I would add to that because, swyx, as you know—you listen to my show—the idea is that AI has taken smart glasses and completely changed everyone's mind about that as a product category and form factor. And I should say this: from things that I've been looking at investing in, wait until you see what they can add on to earbuds.

Simon Willison

Ooh.

Brian McCullough

Like the earbuds in your ear can do a lot more things than they're doing now. Then you combine that with smart glasses, and you combine that with an LLM that you can access maybe with a phone as the mothership. There are some interesting things. CES next year is going to be crazy if you think AI wearables are a thing.

swyx

Anyway, this year they were not a thing. There were very much no wearables at CES.

Brian McCullough

Mm.

Simon Willison

This one's interesting as well because the thing that makes these interesting is that it's multimodal, right? Audio input, video input—

Brian McCullough

Yeah.

Simon Willison

Image input. A year ago, that was hardly a thing, and now it's dirt cheap. We're in a much better position now than we were 12 months ago to build the software behind this stuff.

[Ride Home] Simon Willison: Things we learned about LLMs in 2024 | BidClub