Kevin Roose
I had a first this week.
Casey Newton
What’s that?
Kevin Roose
I had my first experience with smelling salts.
Casey Newton
Wait, did you faint?
Kevin Roose
Yes. I had to get a blood draw at the doctor, and I am a big baby when it comes to getting blood taken. Half the time when I get blood taken, I pass out, and this time I not only passed out, but I vomited—
Casey Newton
Oh, no.
Kevin Roose
—and had to be brought back with smelling salts.
Casey Newton
Kevin—
Kevin Roose
Casey, if you have never experienced smelling salts, they’re not messing around.
Casey Newton
They are not. I cannot believe that I’m just learning this information about you because I am also a fainter.
Kevin Roose
You’re a fainter.
Casey Newton
I am a fainter.
Kevin Roose
We are legion.
Casey Newton
In 12th grade, we went to see a cadaver for my AP Biology class, and intellectually I was fascinated by all the systems of the body. All the other kids and I were standing around the cadaver, and the person was explaining, “Well, this is the liver and this is the spleen.”
Then I got a whiff of something. I don’t know if it was embalming fluid or formaldehyde or something, but it was like something triggered in my brain and said, “This is against nature. You should not be this close to an opened-up dead body.” I spun around, took a header off a whiteboard that was against the wall, and woke up staring at the ceiling. The first thing I heard was my AP Biology teacher, Ms. Oliver, saying, “Do we have an emergency contact for this kid?”
Obviously, I don’t want to tell people that they should faint, but it is one of the most amazing, crazy experiences you can have. Do you know what I mean? The moment when your consciousness just leaves you.
Kevin Roose
Yes.
Casey Newton
Crazy.
Kevin Roose
When I was brought back with the smelling salts, it felt very Victorian. It was like I was on my fainting couch.
Casey Newton
The vapors.
Kevin Roose
Yes. Call Mr. Darcy.
I’m Kevin Roose, a tech columnist at The New York Times.
Casey Newton
I’m Casey Newton from Platformer.
Kevin Roose
And this is Hard Fork.
Casey Newton
This week, give me five.
Kevin Roose
GPT-5. We’ll tell you all about OpenAI’s latest frontier model. Then Kevin and I get access to the new Alexa+. We found a few minuses, and we’re bringing in Amazon’s VP of Alexa to talk about it.
Alexa, prepare my interview questions.
Well, Casey, it has been another busy week in the world of AI.
Casey Newton
Boy, has it.
Kevin Roose
Because we’re going to talk about OpenAI, I should add my disclosure that The New York Times Company is suing them and Microsoft for copyright violations.
Casey Newton
And my boyfriend works at Anthropic.
Kevin Roose
We’ve gotten a bunch of new AI releases and announcements this week. We’re not going to go through all of them, but some of the highlights include something called a world model from Google DeepMind. Genie 3 has an interactive game engine where you can describe a game that you want to play, and it can build it in real time. Pretty cool. We can’t use that yet, so that was just a demo or research preview.
That was early in the week, and then we got a new Claude version. Opus 4.1 is out, so I’ve been playing around with that. It’s not too different, but it’s a newer update from them. We also got open-source models from OpenAI, putting the open back in their name. They released 2 open-source models this week. Casey, have you played around with either of those?
Casey Newton
I have not yet downloaded them. Have you?
Kevin Roose
I have not. One of them is apparently small enough that you can run it on a MacBook. For another one, you need a dedicated GPU. These are basically OpenAI’s first open-source models since GPT-2, many years ago.
People have been hounding them, saying, “You guys are betraying the founding spirit of OpenAI by not making these things open and accessible through open source.” They said, “Well, here you go. Here are some models.” They’re not their top-of-the-line models, but people are finding various uses for them. This is designed to compete with the open-source models coming out of China from companies like DeepSeek.
Casey Newton
The early word on these is that they’re pretty good and competitive with o3-mini and o4-mini, which are more proprietary models. The early reviews I was reading of the open-source models were that they were powerful and good.
1. OpenAI Launches GPT-5
Kevin Roose
Those are some of the announcements that we got earlier in the week, but the big one is that OpenAI released GPT-5, its long-awaited flagship model. People have been asking Sam Altman about this, including us, for many months now. This was long-awaited, and there was lots of hype and rumors flying around about it.
We just got off a press briefing, a Zoom call with Sam Altman and some of the other leaders of OpenAI. Casey, what did we learn?
Casey Newton
It probably won’t surprise people to learn that what they told us during this briefing was that GPT-5 is their best model ever. Sam Altman said in his remarks that this is a major upgrade. He called it a significant step along the path to AGI, but he also said that we’re not at AGI yet.
Among other things, he said, “Look, this model does not continuously learn,” and in his view, AGI will continuously learn. I thought it was cool that he said that, because now we have one thing to hang onto: Maybe when a model can do that, we’ll feel like we really are getting close to AGI.
The other thing he said that struck me, and that I thought was kind of funny, was that after they had put GPT-5 together, he went back to using GPT-4 and said, quote, “It was quite miserable.” He said he never wants to go back to using GPT-4 ever again. That’s how good he says GPT-5 is, Kevin.
Kevin Roose
He compared it to the previous models. He said GPT-3 felt like talking to a high school student, GPT-4 felt like talking to a college student, and GPT-5 was the first time it felt like talking to an expert—someone who has a Ph.D. in a subject.
I think we should caveat this all by saying that, as of our taping this week, GPT-5 had not yet been rolled out, and we hadn’t been able to put it through its paces. But it will be rolled out this week, including to free users of ChatGPT who have not previously had access to their top-of-the-line models.
Casey Newton
I think that’s important because OpenAI’s best models at the moment have been reserved for paying users. The chatbot that I use the most is o3, which is a reasoning model that OpenAI makes. That’s not accessible to people who are on the free plan.
I do think it’s really notable that even free users—which I think is going to include a lot of high school and college students out there—are now going to have access to what at least they are saying is PhD-level intelligence and reasoning.
Kevin Roose
One of the most annoying features of ChatGPT for years now has been this model selector. You go in, and it gives you a little drop-down menu. It defaults to GPT-4o now, but you can pick your own model if you want something more powerful than that and you’re a paying user.
For GPT-5, OpenAI is doing away with the model picker, or at least making it less necessary, because it has built what it calls a router. That will essentially analyze your request and how much computation it needs to answer that request—whether it’s a simple query or something more involved—and direct it to the correct model.
For a lot of people, this is going to be their first experience with a reasoning model. OpenAI does not make that the default right now in ChatGPT, so I think that will be a big update for people, regardless of whether GPT-5 is actually better than previous models. The ability to use these reasoning models for free seems like a pretty big deal.
Casey Newton
Getting rid of the model picker could cut both ways. We should say that all of the big labs have a model picker. Gemini has one, and Anthropic has one in Claude.
Sometimes I’ll ask an easy question of one of these models that is set to reasoning mode, and then I’ll think, “I probably didn’t need that much computational power.” On the other hand, I feel like it sets up an incentive for OpenAI, which wants to save as much money as it can, to always try to route you to the absolute least compute that you need.
I’ll be curious to see whether I feel like that’s affecting the quality of my experience now that I maybe can’t go in and say, “Hey, let me use the good stuff.”
Kevin Roose
On this briefing, OpenAI said all the expected things about how GPT-5 is better at everything than previous models.
Casey Newton
But they also spent a lot of time talking about what they called the vibes of the model, which they believe are quite good.
2. GPT-5 Builds Software On Demand
Casey Newton
They also gave a series of demos, and one that I thought was interesting introduced this concept of what I believe Sam Altman called “software on demand.” GPT-5 can instantaneously create a piece of software for you. In the demo that we saw, one of the employees there built a tool to let his girlfriend learn French, and it did this in some fun ways.
In one case, it created a series of flashcards for her. In another, it created a little snake game with a mouse and a piece of cheese, so every time the mouse caught a piece of cheese, it would show her a new word to learn. He was able to do all of that just via a text-based prompt, and it actually looked pretty good. Five years ago, if you turned that in in an intro to computer programming class, you probably would have gotten an A.
Kevin Roose
Yeah. It's pretty impressive, but those are also things that other models can do today. So I'm going to need to really drive this thing myself to figure out what it can do. I'm going to put it through my usual bevy of tests, known as RustBench, and see how it does.
Casey Newton
Yeah.
Kevin Roose
I confess, during this briefing, I zoned out a little bit. I've been to a bunch of these. Everyone says their model is the latest and greatest, and it's so good at coding, and it's got all these agentic capabilities, and it all starts to sound a little bit like marketing hype to me.
For me, the interesting question to ask about new models these days is not, “How much better is it?” or “How does it score on these benchmarks?” It's, “What is possible for me now that wasn't before?”
Casey Newton
Right.
Kevin Roose
And I still don't have a really good answer to that from GPT-5, although I'm going to investigate.
3. GPT-5 Faces The Hard Questions
Casey Newton
Yeah, and that might be a good point to bring up 2 of the questions that got asked of the GPT-5 team during the briefing that I think would be of interest to our listeners.
One was somebody asked, “Are you starting to run into the limits here? Are the scaling laws holding?” Sam Altman said, quote, “They absolutely still hold, and we keep finding new dimensions to scale on.” He said that they're still finding new paradigms that will let them scale in new ways.
So he very much tried to give the impression that, no, they are not struggling at all to figure out how to build better models. I suspect, though, that might get some pushback as people start to use this thing and observe that, yes, it is clearly better in a handful of ways, but to your point, Kevin, can you really do anything that you couldn't do before? That doesn't seem to be the case. It's just that it can do what it used to do a little bit better. I'm curious what you made of that.
Kevin Roose
Yeah. I think that's reasonable. I wonder if they are starting to finesse their definitions of the scaling laws to account for the new reasoning models, because people for months now have been saying, “Well, the models in the pre-training phase may have gotten as good as they're going to get, or about as good as they're going to get. But the way to get them to be more intelligent is through this post-training phase, through these reinforcement-learning cycles, these reasoning environments that they're trained on and put into.”
So I suspect that when they say that the scaling laws have not broken, they're also referring to this kind of reinforcement-learning reasoning approach as well, and I believe them. I've talked to people who say they think there's still a long way to go on that.
But this was a really big model. We don't know exactly how big. We don't know exactly how many GPUs it was trained on or how much data it was fed, but it's safe to assume that they did everything they could to max out the scale of the model, at least in pre-training.
From what we saw in the demos, it doesn't look like it's that much smarter. Maybe it's a little better at some things, but it did not come out of the box superintelligent or anything like that.
Casey Newton
Yeah. The other big question that I'm always interested in when these big new models come out is: What was the safety-testing experience like? Is this model going to be sycophantic, and what sorts of very intense relationships are people going to form with it?
Nick Turley addressed that one. He noted that earlier this week, OpenAI put out a blog post, which I actually wrote about in Platformer this week, that is all about their approach here. They say they're working with physicians and really trying to bring in a lot of outside expertise to help them understand how people are interacting with these models and make them safer.
He said that they are absolutely not optimizing for engagement here. They just want to make a useful tool that sends you on your way, and that essentially they're going to have more to communicate about this soon. So we didn't get a ton of detail there, but they have said that, at least in some ways, they think that they have improved these models to make them less sycophantic.
In addition to that, they said they did 5,000 hours of red-teaming. They shared, I believe, these models with some external experts for advice on that. They did say that they rate this model as high on the scale of whether it could be used to create novel bio risks, so they're building in a bunch of protections around that. That doesn't seem great. But anyway, that was the sort of safety report that we got in advance of the launch.
Kevin Roose
They also said that, in addition to all these new capabilities, GPT-5 is much more reliable than previous models. They claim it hallucinates less, and it does this interesting thing called safe completions, where basically, if a model doesn't want to accommodate some request or carry out some task because it's against the guidelines, instead of just refusing it, it will make up a safer version of the request and complete that instead.
It'll be interesting to see how people use that. But yes, this is the claim they make: It's more reliable, less deceptive, and gives these kinds of safe completions.
Casey Newton
Well, and that actually gets into something interesting, though, Casey, which is: What is it that makes OpenAI say this is GPT-5?
The big-number releases have this amazing marketing power now, I think, because the leap from GPT-2 to GPT-3 was so big, and the leap from 3 to 4 was also pretty big. That creates a lot of expectations for 5.
But in the background, OpenAI is just trying a bunch of things, building a bunch of new models, and then stapling various things together. Eventually, they get to something and say, “We're going to call this one 5.” But it's not quite as linear as it looks from the outside, right?
Kevin Roose
Yes, and all of the labs have had experiences where they thought they were training a new model, and then it didn't quite work out as well as they wanted it to, so they assigned it some lower number.
That's happened at a number of big labs that I know about. It's happened at OpenAI. They had a previous big model that they were building, which ended up becoming 4.5. I believe it was supposed to be GPT-5 at one point, and it just didn't turn out as well as they'd wanted it to.
So yes, they are playing games with the numbering of the models and the marketing around that. But I think calling this GPT-5 signals that they want this to be viewed as a similar step in capability to what people saw from GPT-3 to GPT-4.
Casey Newton
Yeah. To me, that is one of the most interesting things about this release. Whatever GPT-5 turns out to be, this is the thing that they thought was the next big step forward, and I think we should evaluate it on those grounds.
Kevin Roose
I think the big picture here is that OpenAI is trying really, really hard to stay at the head of the pack. This is a company that has been racing very hard toward AGI, or something that they can claim is AGI, and they are still going.
I find their execution to be quite impressive. This is now a very large company. They've got a lot of different competing teams and priorities. They had all this board drama, and I think it was reasonable to expect—and I certainly expected—that in the wake of all that, they would slow down and maybe allow some competitors to catch up.
But they showed this week that they are not slowing down. They are, in fact, accelerating, and they want to get there before anyone else.
Casey Newton
True, although they have also experienced a lot of poaching in recent weeks and months. I think one thing I'll have my eye on over the next several months is whether they're able to continue iterating very quickly, or whether some of the losses that they've experienced over the past few weeks have really hurt them.
Casey Newton
Yeah.
Kevin Roose
Incidentally, Casey, I'm told that in response to the GPT-5 launch, inside Meta headquarters, the superintelligence researchers have moved their desks even closer to Mark Zuckerberg. So that's how seriously they're starting to take this over there.
Casey Newton
They are now sitting on top of Mark Zuckerberg.
Kevin Roose
There are now 2 researchers who are sitting at Mark Zuckerberg's desk with him.
Casey Newton
And we'll have to see how that plays out.
Casey Newton
Yeah.
Kevin Roose
Yeah.
Casey Newton
Those are some of our initial impressions. But we are going to come back tomorrow, after we've had a little time to play with the model, and give some first impressions there, too.
Kevin Roose
Let's travel to the future now, Kevin.
4. GPT-5 Gets A Vibe Check
Kevin Roose
All right, Casey. It is now Thursday. GPT-5 has been officially released for a few hours. I still do not have access to it for some reason, but I gather that you do, so give me your day-one vibe check. What are you seeing? What is GPT-5 like, and what do you make of the reaction to it?
Casey Newton
Well, this is a very significant moment in the history of Hard Fork, Kevin, because for the first time, I'm having a conversation with you while vibe coding something.
Kevin Roose
What are you vibe coding?
Casey Newton
Well, ChatGPT is currently hard at work building a to-do list app for me. I said I wanted it to have the aesthetic of the Fantastic Four: First Steps movie that just came out. I didn't love the movie, but I did love the production design, so I was like, "Make me a to-do list app that looks like that." Let's see how it goes.
Kevin Roose
I can't believe we're building giant gigawatt data centers for your stupid to-do apps. This is so wasteful. God.
Casey Newton
Listen, you need to send over RoosBench, your proprietary suite of evals, so I could really put this thing through its paces. But look, let me give you some high-level notes, Kevin, on what I'm seeing and on what others are seeing.
The headline here is that this does seem like a really meaningful improvement to ChatGPT. I think, in particular, if you are a free user of ChatGPT, you're going to have a great day, right? Because for the first time now, in addition to the standard ChatGPT model, you're going to have some reasoning capability. So, essentially, if you're cheating your way through high school, you're just going to have a lot easier time of it now because this thing can do some really extended work on long problems.
Kevin Roose
Yes, you can now cheat your way through an entire semester with just one press of a button.
Casey Newton
Exactly.
Kevin Roose
This is good. The way I saw people talking about it online was that they thought OpenAI had not raised the ceiling of the AI frontier by a lot with GPT-5, but they had raised the floor. Essentially, all the free users who previously got defaulted into the less powerful models are now going to be using the more powerful models, which could be a big perceptual shift, if not a shift in frontier capabilities.
Casey Newton
For sure. I do think it has some things that are not quite capabilities but still will meaningfully affect how people use these AI systems. For example, this thing really is just a lot faster than its predecessor. Over the past couple of hours, I took some editing work that I sometimes ask ChatGPT to do. I know about how long it takes using the o3 model. I put it through GPT-5, and sure enough, it blazed through it. It did just as good a job as it had done before. So if you're the sort of person who's using ChatGPT a lot, I think that's really going to stand out to you.
Kevin Roose
Yeah. What about the pricing? I saw some people saying that GPT-5 was much cheaper than they expected it to be. Not cheaper for the sort of ChatGPT subscriber—the subscription prices are staying the same—but for developers who are building on top of it, my understanding is that it's a lot cheaper than other models from other AI labs.
Casey Newton
That's right. It came in at $1.25 per 1 million input tokens, which is the same as Google's Gemini 2.5 Pro. Google has, of course, also been pricing really aggressively to try to box out the competition.
What makes that figure interesting, I think, Kevin, is that that number is a lot smaller than Anthropic's Claude Opus 4 API, which comes in at $15 per million input tokens. So I think some of these really well-capitalized AI labs are taking this moment to say, "Hey, we're going to put a lot of pricing pressure on some of our competitors."
Kevin Roose
Yeah, it very much reminds me of the moment, like, 10 years ago when Ubers were $4 because venture capitalists were just subsidizing the artificially cheap prices. We're in sort of that moment for AI tokens now.
Casey Newton
Yeah.
Kevin Roose
What else can we say about GPT-5 in the couple of hours now that it has been out?
Casey Newton
Yeah. I sometimes like to joke that the worst insult you can make to anyone who has just released a new AI model is, "My timelines are now longer." And it does seem like that is something that people are saying about the new GPT-5. What I mean by that is I now think it's going to take a little bit longer until we reach AGI, some sort of very powerful AI system.
In fact, some people are posting online screenshots of some prediction markets that, until today, when asked, "Who do you think will have the most powerful AI model at the end of August?" were showing OpenAI in the lead. Almost instantaneously after the livestream on Thursday, OpenAI collapsed, and Google has now ascended and is assumed to have the best model by the end of this month.
I don't want to overstate what that means necessarily, but it does seem like there was a huge contingent of people who thought that GPT-5 was going to be this revolutionary new model, and it seems instead like a more evolutionary one.
Kevin Roose
Yeah. That makes a lot of sense to me. One other thing that stuck out to me, and I wonder if it stuck out to you, too, was that OpenAI released some benchmarks and some data about GPT-5. One of the things they showed was that hallucinations—the rate of GPT-5 just making stuff up while answering questions—has gone way down. For some types of questions, it's now sort of around a 1% hallucination rate.
I think that was interesting to me because this is clearly something that was a problem with earlier versions of this. In fact, there was some speculation and some indication that these newer reasoning models were hallucinating at higher rates than the previous generation of models, and there was a lot of concern about that.
It seems like they have figured out a way to get the hallucinations under control with GPT-5, although, with everything, I don't totally trust these benchmarks. I'm going to have to see this for myself.
Casey Newton
Yeah. Everyone's mileage is going to vary on this one. I will say I have already caught it hallucinating a couple of times, somewhat disappointingly. So as always, don't trust these things for anything mission-critical. You're always going to want to double-check your facts.
Kevin Roose
Yeah. Only use it to build stupid to-do apps with the Fantastic Four aesthetic on them.
Casey Newton
The Fantastic Four have a very cool aesthetic, and I think you need to open up your mind a little bit.
Kevin Roose
Okay. That is our day-one vibe check of GPT-5, and we will continue to play around with this and tell you anything cool or interesting or strange or upsetting that we find.
Casey Newton
Sounds good.
Kevin Roose
All right. That's enough about GPT-5. When we come back, we'll talk about another AI system we got our hands on this week, Alexa+.
Now, Casey, are you an Alexa user?
Casey Newton
I have been an Alexa user for a long time. I still have one of the original Amazon Echos in my house, and to Amazon's credit, it still works.
Kevin Roose
The Pringles can, they call it.
Casey Newton
Yeah, I have the big old sort of Pringles-can Echo.
Kevin Roose
Yeah, me too.
Casey Newton
Yeah.
Kevin Roose
So I am a heavy user of this product. I have probably 5 of them in my house—
Casey Newton
Okay.
Kevin Roose
—in various rooms. And so I'm very excited for our conversation today, which is going to be about the new AI-ified Alexa+. And before we get into our experiences using this thing and our interview with the guy who runs it, we should make a couple of disclosures.
One of them is that The New York Times Company has recently agreed to a licensing deal with Amazon that will allow Amazon access to Times content for its AI platforms, including Alexa. We just thought you should know that. We have nothing to do with that, obviously, but that is going on in the background in another part of the company.
The second thing we should say is that if you have an Alexa device, it is going to be going off constantly during this segment unless you go over right now and hit the little button that mutes it. Sorry in advance to Alexa owners, but we'll give you a little bit of time right now to pause this, go over, hit the mute button on your Alexa, and come back.
Casey Newton
Or alternatively, just find the circuit breaker in whatever house you're in right now. Shut them all off. Run on battery power for the rest of this episode.
Alexa, order 14 bags of dog food.
Kevin Roose
I wonder if that actually works.
Casey Newton
Wait, and I should also probably disclose that my boyfriend works for Anthropic, because I'm pretty sure that Anthropic is providing APIs that are being used in Alexa+.
Kevin Roose
Wow, we've got so many disclosures today.
Casey Newton
Yeah, yeah.
Kevin Roose
Okay.
Casey Newton
All right.
Kevin Roose
Let's get started.
Casey Newton
Okay.
5. Alexa Finally Gets Generative AI
Kevin Roose
So, Alexa. Alexa is one of the most puzzling technology products that I have ever encountered. Like you, I have been an Alexa user since the very early days. People don't realize this product was released in 2014. Alexa is 11 years old.
Casey Newton
Yeah.
Kevin Roose
And when it came out, I was very excited. I thought, “I'm going to put this smart speaker in my house, and I'm going to ask it to do things for me, and it's going to be like having a little assistant right there on my kitchen counter.” Alexa has added dozens of features, maybe hundreds of features, since 2014, and I use zero of them, because the 3 things that I use Alexa for are setting timers, choosing music to play in my house, and telling me the weather before I leave for the day.
Casey Newton
Absolutely.
Kevin Roose
Are those similar to what you use this for?
Casey Newton
Those are the exact 3 things that I use Alexa for. Have I tried to use it for other things? Yes, but the experience, frankly, has just never been that great, so I always come back to those 3.
Kevin Roose
Yes, those are the big 3 in my house. Same use cases, same limitations. But when generative AI started to get good a couple years ago, I think people naturally started to ask, “Well, when is Alexa going to start using this new generative AI technology?” It's sort of built on this older, more deterministic kind of system, but it seemed like a natural thing to expect that Alexa would start to incorporate some of this technology to be able to answer maybe more open-ended questions, to give longer, more detailed responses, and to do more than just set timers and tell you the weather.
Casey Newton
Yeah, I mean, once OpenAI released Voice Mode for ChatGPT, it immediately seemed so much more interesting and powerful than Alexa and Siri, which is Apple's very similar system that it makes for its devices. So, yeah, I think both of us were like, “Okay, well, when are we going to get that OpenAI-style voice mode in these smart devices that we have in our homes?”
Kevin Roose
Yeah, and so it's taken a while. We should say that. It has not been a smooth or simple process, and part of what I'm so excited to talk with Daniel Rausch, the vice president of Alexa, about later in the show is just why it's been so hard to shove an LLM-based generative AI technology into this preexisting assistant product.
But we should just talk briefly about what Alexa+ is, and then our experiences with it, because both you and I have gotten to try this over the past few days.
Casey Newton
Yeah, so Kevin, tell us a little bit about Alexa+.
Kevin Roose
So Alexa+ is the name for the most recent overhaul of the Alexa virtual assistant. It's powered by generative AI. We don't know exactly which model or models, but it seems to be a mix of Amazon's proprietary AI models and then maybe some of Claude, which it has a deal with Anthropic for. Amazon has been using Claude inside of its AI products for a number of months now.
Amazon claims that the new Alexa+ is able to do much more. It's able to be much more conversational and more personalized. It can do things like book reservations at a restaurant or order you an Uber. It can answer questions that aren't just pure lookups, where you're looking for, you know, what time is the baseball game tonight. It can actually do more complex things for you. It can control the smart devices and appliances in your house, and it can purchase things for you online.
This new Alexa+ is not out to everyone yet. They've been rolling it out slowly. They are now in what they call the early-access period, but we were able to get this on some new devices that we ordered. It also doesn't work on every kind of Echo device. You have to have one of the newer ones to be able to run it.
Casey Newton
And Kevin, when you say that this has been rolling out slowly, it has been rolling out extremely slowly. It was only on June 23 that Amazon said that 1 million people had Alexa+, across presumably hundreds of millions of Echo devices out there.
Kevin Roose
Yeah, so you and I both got the new Echos that can run the Alexa+ early-access program, and turned it on and set it up. A few things stick out to me right away. One is that the voice on this new Alexa is just way better than the old Alexa.
Casey Newton
Yes, I would agree with that.
Kevin Roose
It is way more fluid. It sounds more like something you'd hear out of ChatGPT Voice Mode. They have managed to overhaul the actual voice part of the voice assistant, so it sounds much more like a human.
Casey Newton
And there are a bunch of different voices. I think I saw 8 of them inside the app. Half of them are masculine, half of them are feminine, so, yeah, you can change that to your liking.
Kevin Roose
Yeah. The other big difference I noticed right away is that the new Alexa+ does not require you to say the wake word, like “Alexa,” between every question-and-answer pair. With the old Alexa, if you wanted to ask a follow-up question, you had to say Alexa again. With the new one, you can just kind of leave it, and it will intuit or pick up that you have a follow-up question, and it will listen for a while longer. So you can actually have these more extended, multi-turn conversations.
Casey Newton
Yeah, and that lets it do different kinds of things. One of the first things that I did with Alexa+ was that it said, “Hey, would you like to try to solve a riddle?” And I thought, “What are you, the Sphinx?” But I said, “Sure. What the heck?” It gave me a series of clues, and within 3 clues, Kevin, I was actually able to solve the riddle.
Kevin Roose
Wow.
Casey Newton
Yeah.
Kevin Roose
Good for you.
Casey Newton
Yeah, I feel really smart.
Kevin Roose
I'm so proud of you.
Casey Newton
Thank you. Thank you so much. So, yeah, what else were you doing with this thing?
Kevin Roose
So another thing it can do is just give you longer answers. The original Alexa was limited to a sentence or 2. Maybe you could ask it to look something up on Wikipedia, and it would spit out a few sentences, but it was really limited beyond that.
I can now ask it to make up a story and read it to my kids, so we had some fun doing that the other night as a family. You can ask it to suggest a recipe for dinner based on what's in your fridge, and it will help you with that. I used that last night. So these are some of the new features that I was excited to try. I also tried some of their integrations. They have an integration with OpenTable and with Uber and a bunch of other companies.
Casey Newton
Oh, yeah, tell me about this, because I set this up, but I did not actually use it. So how did that work?
Kevin Roose
Basically, you scan a little QR code on your phone and link your Uber account or your OpenTable account to your Alexa account. It takes a minute or so, and then you can just say, “Order me an Uber from this place to this place,” or, “I want a table at a restaurant in downtown San Francisco near the Ferry Building for 2 people at 6:30 tomorrow,” and it will pull up a couple of options. You choose what you want, and then it can go book the table for you. I thought that was cool.
Casey Newton
And that actually worked when you tried it?
Kevin Roose
So I did not actually follow through with the booking, but I did order an Uber for myself, and it did work.
Casey Newton
Okay, cool.
Kevin Roose
Yeah.
Casey Newton
I mean, that actually seems truly useful. Just say to your thing on your desk, “Hey, I need an Uber to the airport,” and it pulls one up. That's great.
Kevin Roose
Yeah, and it can do other cool, multistep things, too. I was able to say I needed a new thing for my kitchen, like a box grater, and I was able to go to Alexa and say, “Hey, look up on Wirecutter what the best-rated box grater is and add it to my Amazon cart.”
Casey Newton
Now, can I guess why you needed a new box grater?
Kevin Roose
Why is that?
Casey Newton
You used it to grate ginger and it dulled the edges.
Kevin Roose
No.
Casey Newton
Okay. What was the reason?
Kevin Roose
I left it in an Airbnb.
Casey Newton
Okay. I should have seen that coming.
Kevin Roose
Yeah.
Casey Newton
Anyways, go ahead.
Kevin Roose
So anyway, those were some of the good things about this product, but we have to talk about some of the limitations as well. Casey, what was your experience with Alexa+?
6. Alexa Breaks The Basics
Casey Newton
I have to say, I did not have a good experience with this thing. First of all, I bought an Echo Show 5. There's a big banner on the page that says it works with Alexa+.
The thing shows up at my house, and basically what I've come to understand is that an Echo Show is a device that just constantly invites you to spend money with Amazon. I found it honestly infuriating, because I plugged this thing in, and when you set it up, it's like, “What kind of background do you want?” I was like, “Show me some art.” That's one of the options.
I would say for about 4 seconds per minute, it would show me some Renaissance masterpiece or something, and then it would be like, “Hey, do you want aspirin? Do you want paper towels? You want to buy paper towels? You can actually buy paper towels right now. Just say, ‘Hey, Alexa, buy paper towels.’”
It was just sort of this forever. So I eventually just unplugged the thing, because I was like, “Why did I just spend $90 to have a permanent rotating advertisement for household products on my desk?” That is so weird.
It put such a bad taste in my mouth about the whole thing. Then, a day later, I got the Echo Show 15. For some reason, Amazon sent me 2 of them. I truly don't know why. I did not need 2 of them.
I unboxed the thing, and the thing is meant to be mounted on a wall. Now, there are a lot of things I'm willing to do for a podcast, but mount an Echo Show on my wall—
Kevin Roose
You're not willing to do a construction project.
Casey Newton
It's not one of them. No, I was not going to do that. Also, the thing can't stand up on its own, so I just had a 15-inch screen sitting on my desk for a day while I was talking to it. This whole thing was very silly.
So that's the hardware side of it. You may have a better experience because you like mounting things to your wall, and so you did that and you're having a good time. But that was all of the precursor steps I needed to take to even be able to engage with this thing.
Then I finally had it set up and started to try to put it through its paces. I went through the little riddle game, and it's like, “Hey, I could help you with a personalized meal plan.” I was like, “All right, great. Set me up with a personalized meal plan.”
It's like, “Well, we could do this or that.” It showed me a row of recipes that it could cook for me. I swiped through with my finger and saw a lemon pasta. I said, “Okay, show me the lemon pasta,” and it said, “Sorry, I didn't get that.”
I said, “Alexa, the lemon pasta right there. Could you make me this lemon pasta from this website that you're showing me right now?” Dead silence. I was like, “Oh my God.”
Right here, we have just landed in the exact spot that has been bedeviling Apple for the last year and that is bedeviling Alexa right now. These systems are just very hard to make reliable.
Now, I will say the device was sort of having trouble connecting to my internet. Everything else in my house was connected to the internet and was working fine, but this was just, every once in a while, saying, “You're not connected to the internet.” Was that an issue with me? Was that an issue with the hardware? I'm not totally sure. Maybe that was why it wasn't able to perfectly answer my question. I do want to say that in case this was not actually an AI issue.
But, oh man, within 5 minutes I was like, “Get this thing out of my house.” Again, I wanted to like it. I was excited about it, and after 2 days of ads for paper towels and 1 day of it refusing to show me the lemon pasta, I thought, “What am I doing with my life?”
Kevin Roose
Yeah. I should say, I have also had a bunch of very bizarre and frustrating experiences with this thing.
Casey Newton
Okay, let's get into it.
Kevin Roose
Okay. We've said what we like about this thing.
Casey Newton
Yeah. Which, remind me what that is again.
Kevin Roose
Many of the new capabilities are quite cool.
Casey Newton
Yeah.
Kevin Roose
Unfortunately, many of the old capabilities I relied on as the reason I used Alexa at all have become broken—
Casey Newton
Okay.
Kevin Roose
—as a result of this update.
Casey Newton
Okay, so tell me about this.
Kevin Roose
One of the things you also notice very quickly when you're using this thing is that the latency is just a problem. It's a little slow—
Casey Newton
Yeah.
Kevin Roose
—to respond to questions. It's not as zippy as the old, pre-LLM Alexa. I understand that these things have to go to the cloud, they're processing more complex instructions, and it's all going to take a little time. I assume that will get better.
The basic things that it gets wrong now include alarms, which is actually a thing that I use Alexa for every day.
Casey Newton
Wait, so tell me how it got it wrong.
Kevin Roose
The new Alexa Plus update seems to have broken Alexa's ability to reliably set and cancel alarms—
Casey Newton
Oh my goodness.
Kevin Roose
—which is a core thing that I use this product for. For example, this morning I woke up on my own a little bit earlier than my alarm, about 10 minutes before it was supposed to go off. I said to Alexa, “Alexa, cancel the alarm.” Silence. Nothing. This is a command that I have issued probably 1,000 times.
Casey Newton
And Alexa Plus is a little smarter now, and she's giving you the cold shoulder.
Kevin Roose
Yes. She's saying, “Actually, I'm going to wake you up anyway in 10 minutes.” So that was not good.
I also experienced some hallucinations when I would ask it questions about things happening in the world, things happening in the news. I asked it about a tennis tournament that's going on right now. I said, “Who's the top seed in this tennis tournament?” It gave me the name of a player who's not even playing in this tournament.
It also has trouble orchestrating the different tasks. One of the things that would happen is I gave it a research project for a dinner playlist. I was looking for some new music—
Casey Newton
Mm-hmm.
Kevin Roose
—to put on our dinner playlist, and instead of doing that research project, it just started searching on Spotify. It routed—
Casey Newton
Ooh.
Kevin Roose
—the query to Spotify within the Alexa interface and started playing the music—
Casey Newton
Okay.
Kevin Roose
—when what I had asked was, “Do some research for me.” So it seems to have a little trouble figuring out exactly what the user wants and orchestrating the commands.
Casey Newton
That case seems a little borderline to me. I can imagine some people asking for that and maybe being happy if it played some music.
But I had this almost opposite issue where, again, I'm going through, “Okay, what can this thing actually do?” It says, “Ask me what I can do.” So I asked it, and one of the things it said was, “I can help you explore Gen Z music trends.” There was just something funny about the way it said it to me.
I was like, “Yeah, sure. Why don't you help me explore Gen Z music trends?” It thinks for a second, and then it goes, “Well, I found some podcasts about it on Amazon Music.”
I was like, “I sort of assumed you were either going to tell me something about Gen Z music or you were going to play Gen Z music, but now you're trying to sell me Amazon Music,” which I feel like is very consistent with how Alexa Plus handles everything: “Can we sell you a service right now? Could we sell you a product?”
Kevin, I want to say 2 things. One is, I have not used this product all that long, and so I don't want people to think about anything I'm saying as anything other than first impressions. I have not truly had a chance to do the amount of reviewing that I would like to do.
Two, I'm very confident that lots of other people are probably having much better experiences with this thing, because I think if most people were having experiences as bad as mine, I would have heard about this before now.
But all of that said, Alexa Plus did not make a great first impression on me. The Echo family of devices that are just little windows that let you send money to Amazon.com are not for me.
Kevin Roose
Yeah. I had a slightly more positive experience than you. I did actually enjoy some of my interactions with Alexa Plus, but it just seems like it is not quite there yet.
Casey Newton
Yeah.
Kevin Roose
I think Amazon knows this, which is why it's in this early access program. If you open it up, it says, “Alexa may make mistakes,” so they're doing all of the careful rollout that you would expect from a product that is not fully baked.
But some of the features just don't seem to work. There's another feature that I tried where you can email a document to this email address, and it will ingest it into your Alexa. Then you can have it summarize it.
I was very excited. I was like, “I can learn about new papers in AI while I'm doing the dishes.”
Casey Newton
Mm-hmm.
Kevin Roose
And so I email the paper to the Alexa email address, and I say, “Summarize the paper I just sent you,” and it says, “I did not receive a document.” So I think they need to spend a little more time in the kitchen cooking this one. But I think my overall impression is that the Alexa+ that you have now in this early access program is a little like having a GPT-3.5-class model inside of a smart speaker.
Casey Newton
Hmm.
Kevin Roose
Which I think is a valuable thing and one that I would like them to continue to build on. But it is not state-of-the-art in either the language model or the basic tasks. And actually, it seems to be regressing on some of the basic tasks. So I would say this is 2 steps forward, 1 step back.
Casey Newton
I think the most powerful thing that the new Alexa+ has done for me is that it has made me forgive Apple for not shipping anything with the new Siri. I get it now, Apple. I talked a lot of mess about you on this podcast about not shipping this thing, but now, having used one of your close rivals’ attempts to do the same thing that you’re doing, I get it now. I think the finest minds in the world who are working on this stuff actually don’t know how to do this yet. That’s my big takeaway.
Kevin Roose
Yeah. I think what’s happening with Alexa and Siri right now is a symbol of what’s happening in the American economy writ large, which is that we are trying to jam these new AI technologies into these legacy systems and processes, and it’s just kind of a messy fit. These things are weird. They are not deterministic. They are not reliable in the ways that an older, more rule-based thing could be. And they have these amazing capabilities, but when you try to make these hybrid Frankenstein things with the old system and the new brain, it just doesn’t really work. I think that’s happening not just in these virtual assistants, but in a lot of places throughout the economy.
Casey Newton
Absolutely. I also just think that when I’m using a chatbot on my laptop and it gives me something that’s 80 or 85% right, that’s much more useful to me than an Alexa response that’s 85% right. Because in a chatbot setting, I can just take what I need. I can edit or modify it. I can maybe ask the same question of another chatbot and see if I get a slightly different or better result. I feel much more in control of my own destiny. I can take the stuff that works and leave behind the stuff that doesn’t. When you’re doing this with a smart speaker, if it doesn’t work, you say, “God, why’d I spend 90 bucks on this piece of junk?” You know?
Kevin Roose
Totally.
Casey Newton
And I think what I learned about myself was that I have so much less patience for this sort of thing when it is a piece of hardware in my home that has made some really big promises about how it’s gonna help me with all my routines and everything. If it’s kind of hard to set up and it doesn’t work the vast majority of the time, it all just feels like a waste.
Kevin Roose
Daniel Rausch, welcome to Hard Fork.
Daniel Rausch
Thanks so much for having me.
7. Alexa Gets A New Architecture
Kevin Roose
So Casey and I have both spent the past few days playing around with the new Alexa+. I’d like to just start by asking about the technology that powers this thing.
Daniel Rausch
Yeah.
Kevin Roose
How much of it is a new LLM-based system versus the old, more deterministic model that powered the old Alexa?
Daniel Rausch
Yeah. From an AI and model perspective, everything is entirely new.
Kevin Roose
Hmm.
Daniel Rausch
There are some legacy deterministic systems downstream, but really, it’s a complete rearchitecture of everything that you would say Alexa is, from the way you have a conversation and engage with the experience at a very basic level, all the way through Alexa acknowledging you or just maintaining a chat. So there’s a lot of new under the hood.
Kevin Roose
Yeah. Talk about the challenge of moving from this deterministic system to something that is very powerful but also much less reliable.
Daniel Rausch
Yeah. I would say, well, hopefully you’re not seeing it as much less reliable. We’ve got some edges to sand, and we’re in early access. I’m sure we’ll get to talk about the nature of the rollout—
Kevin Roose
Yeah.
Daniel Rausch
—of the rollout, but—
Kevin Roose
I just mean, in general, LLMs—
Daniel Rausch
In general—
Kevin Roose
—are not as reliable—
Daniel Rausch
I see.
Kevin Roose
—as a deterministic system.
Daniel Rausch
I get it. So we want to capture all the benefits of that nondeterministic—we call it stochastic—system in this space. It has the elegance of really engaging in human conversation, but we want the predictable outcomes. Now, large language models don’t support interfaces out of the box to classic systems, so getting those capabilities to interface—we would talk about it as APIs across these interfaces for other systems—is quite hard. They speak natural language. APIs don’t speak natural language. They speak clunky computer science language, but it’s very predictable and it gets a lot of things done. So I would say, if you had to list the technical challenges, the many millions of things—we stopped counting at some point—that the original Alexa could do, marrying that with the power of LLMs is definitely the first and most prominent on the list.
Kevin Roose
So take us back to when LLMs first started coming out. You guys are starting to play around with them, and it’s sparking ideas for you: “Gosh, if we could marry this to Alexa, we could have something really cool.” What are some of the uses that you’re thinking about? What are the kinds of dreams that you have for this model that you’re hoping you can bring into reality?
Daniel Rausch
I think we think of the capabilities in 2 buckets, I would say. Take everything that Alexa, the original Alexa, can do and just make it way better. Just picking up from what customers are already doing with Alexa. Then you start brainstorming, and I think where you were really headed was: What are all the new things that we can do? And the depth of conversation that you can have with the new Alexa experience just opens whole vistas of new kinds of things we can get done. We can help you plan a trip and then follow through on it. We can watch for concert tickets for you. We can not just help you brainstorm about cuisine, but either pick a recipe, get some groceries, and invite the neighbors, or let your partner know it’s date night, that we’re going out, and book a table. So I think the kinds of journeys and the kinds of tasks we can get done for customers are just so much more expansive.
Kevin Roose
Hmm. So Casey and I have spent the past couple of days trying out Alexa+, and we have some feedback, which we can share with you now or later. We’ve talked about it on the show just before this. I think it’s fair to say we both had some things that impressed us about the new Alexa+ and some things that were challenging, including some of the basic stuff that Alexa seemed to be very good at before—or at least that I knew how to get reliable performance out of Alexa on before—which no longer seems to work as well. But what I actually want to know is: Why has it been so hard to do this? Because back in 2023, when Amazon announced that it was going to revamp Alexa, sort of give it this brain upgrade with these new AI capabilities, they said this was going to be ready in 2024, and then that got pushed back a couple of times. So walk us through the journey that you all have been on over there, trying to shoehorn this new technology into this existing product, and maybe some of the challenges that you encountered along the way.
Daniel Rausch
Well, I’ll tell you, we should definitely get some of the feedback. We can cover as much as you like here on the show. If you rewind the tape, you were asking about this too: As we’re starting to experiment, what can we imagine doing? If you go back to 2023 and the models that were available then, the state of the art had very little instruction-following, reasoning, or ability to execute on interfaces with other systems.
We announced something called “Let’s Chat,” which was a mode of Alexa. Think about flipping a switch on Alexa and turning on a chat interface so that you can do some basic question-and-answer and have a discussion about a topic, mostly about knowledge native to the model’s training data versus bringing something in at runtime, the way modern chatbots answer questions by going out on the internet.
I think what we mostly learned from that announcement and the customers to whom we rolled it out was that we had to increase our vision and do something more audacious. Customers really wanted, and we all really wanted, to pick up from where Alexa is and was and extend all of those capabilities. That is many millions of things that Alexa can do, and when you count the tens of thousands of services and devices that are integrated with Alexa, as well as the interfaces and systems that you need to integrate with, it’s incredibly large.
That’s the first technical challenge I mentioned before, the first and probably most important bucket. The second is really grounding it in authoritative sources. As all of us know, you can sit there and fiddle with a chatbot long enough to press it into being smarmy or responding in ways that we don’t believe are the way Alexa might act, for example. You can press it to give you wrong information from some unauthoritative source or from a mistake in its training data. Alexa shifts back to its native training.
Getting Alexa to speak confidently in her personality, with authority, and answer questions correctly is another key challenge. Personalizing an experience of this depth so that Alexa is always learning from her interactions with you and extending your interactions so they get more delightful over time is something you probably wouldn’t have seen in a weekend’s worth of fiddling with the experience. You’ll see it get more personalized. That’s another big technical challenge because the surface area is so much bigger.
Those are a few of the reasons why it took so long. If you rewind the tape to 2023, it’s really about learning how big a project Alexa+ would be and then starting to put one foot in front of the other, really inventing the space of creating those integrations because it just hasn’t been done.
8. Alexa Confronts Early Failures
Kevin Roose
What’s an example of some early failure mode that you all had to overcome? I’ve heard some stories from folks who have worked on Alexa or worked with suppliers that provide models to Alexa. They would tell me stories about—you’d ask Alexa to set a timer for you, and it would write you an essay about the history of timers. It was just misunderstanding the request in the way that a large language model might. So tell us some of those stories.
Daniel Rausch
Verbosity was definitely an early issue.
Kevin Roose
And it continues to be an issue on our podcast, by the way. We still haven’t solved it.
Daniel Rausch
Yeah. I’ve got some training ideas.
Kevin Roose
Okay, good.
Daniel Rausch
Verbosity: These models want to give you an extensive answer. Customers don’t want an extensive answer read out, and they certainly don’t want a disquisition on the nature of timers. What they want is an interface that sets a spaghetti timer.
Kevin Roose
And how do you get them to do that? Is it as simple as putting in the system prompt, “If a customer asks for a timer, don’t give them an essay on the history of timers. Be concise”? Or how do you actually solve that problem?
Daniel Rausch
I would love it if it were that easy. You need a set of models. There are over 70 models in Alexa+.
It’s a vast space. There are different models specialized in different tasks. There are different corpuses of training data we use on different models to get them to complete instruction sets for us and really follow the rules of the road in interfacing with something. You always need to loop back to central systems that are maintaining context in the conversation and picking up on references and pronouns that you’ve used to refer back in time, and cascade those forward.
The amount of work that went into just the interface between a large language model and the downstream systems that complete tasks is the biggest body of work that we’ve put in. Without a whiteboard here, it would be too much to even try to explain to you and your listeners the technical depth that went into it. We’ve got a great team working on it, and it’s hard.
Kevin Roose
Of those 70 models in Alexa+, how many are Amazon’s own in-house models versus models like Claude that you get from external companies?
Daniel Rausch
There’s a mix. The best way to know what models are in Alexa+ is just to go to the Amazon Bedrock webpage and look at the latest update there. We use the best tools that we have available to us for the job, and we’ve got great partners over in AWS helping make sure we’ve got the right, best tools for the job.
Most of our traffic does flow through Amazon Nova models. We have the most control over how those get trained, tuned, and post-trained. I think it’s over 80 percent of traffic on the main, big inferences within the system that flows through Nova models. But there are many different reasons to use many different models. I think you guys know better than most that models specialize in different things, so we use the best tool for the job.
Casey Newton
Can you give us a sense of how big the team is that’s working on Alexa? How big of a priority is this within Amazon?
Daniel Rausch
It’s thousands of people.
Casey Newton
Okay.
Daniel Rausch
That’s building hardware, building Alexa+, integrating with all those systems, and adding new integrations and new things that Alexa can do. It’s a pretty vast scope, so it takes a big team.
Casey Newton
Yeah.
Kevin Roose
There was a former machine-learning scientist at Alexa AI, Mihail Eric, who did a long post on X last year—his version of a postmortem or retrospective on what was happening with Alexa. He wrote that Amazon had, quote, “All the resources, talent, and momentum to become the unequivocal market leader in conversational AI.” But then he said that Amazon and Alexa had fumbled the ball because Alexa was, quote, “Riddled with technical and bureaucratic problems.”
It made it seem like the problem was not just that the technology was an uneasy fit, but that there were also organizational and bureaucratic problems that had to be solved. Can you talk a little bit about that?
Daniel Rausch
I won’t comment on that post in particular. Honestly, I don’t remember it, but there is definitely a startup-culture transformation happening within the Alexa team. The life cycle of any product that’s been around for 10 years has ups and downs. But I think our rate of innovation had slowed down, and coming through for customers on integrating these new, powerful tools is something that’s really quickened and inspired the team.
I don’t identify with the bureaucratic comment. Maybe it’s a comment about me, so maybe I won’t identify with it. But I do think the team is inspired by the vision, executing at an unbelievable pace, and really creating a lot of invention because there are a lot of really hard problems.
Kevin Roose
I’m curious where the new Alexa sits in relation to Amazon’s overall AI ambitions. This is a company that has offered a lot of AI models through AWS, has a big market share in cloud-based AI, and also recently started an AGI lab at Amazon that is going to be pushing toward something like artificial general intelligence. Is Alexa part of that overall effort to create and serve more capable AI systems, or is this a consumer-targeted spinoff of those efforts?
Daniel Rausch
I would say we do believe, and I share this belief, that the leadership team at Amazon has this generation of generative AI is going to transform every customer experience we have, and that means...
We have a lot of different types of customers. You mentioned AWS. We have enterprise business customers. We have consumer customers. We offer a very big landscape of services. At some point within the last year, we counted and there were over 1,000 different AI efforts going on with consumer applications alone.
If you look at the scale and scope of what Amazon does and assume our belief that every experience will be transformed with generative AI, it’s as big as Amazon is at that point. I would also say that internally, it’s part of how we work now. To be as productive as you can be in this day and age and get as much done for customers as we aspire to, you have to build AI into how you’re working. You both do this, I know, and I’m sure many of your listeners do too, but it’s certainly part of what’s going on at Amazon as well.
Kevin Roose
Hmm.
Casey Newton
Yeah.
9. Alexa Faces Product Feedback
Kevin Roose
Okay. Well, Daniel, we have some product feedback for you—
Daniel Rausch
Let's do it.
Kevin Roose
—because, as they say, feedback is a gift.
Daniel Rausch
Always.
Kevin Roose
So we'd like to give you some gifts.
Casey Newton
And it's Christmas.
Kevin Roose
Casey, why don't you start?
Casey Newton
All right. Most of my feedback is less about Alexa+ as an AI than about Alexa+ and the actual hardware I got. I first started with the Echo Show 5, which does say on the website that it is Alexa+ enabled, but then some of your folks told me, “No, to get the full experience, you should get the Echo Show 15.” So I had the 2 experiences.
Daniel Rausch
Okay.
Casey Newton
On the Echo Show 5, my first observation was that after I told it I would like to see art, every time I looked over at it, it was asking me if I wanted to buy paper towels or Advil or something. That was a little less the case once I got the Echo Show 15. I don't know why that might have been, but I felt like the Alexa+ AI thought of me primarily as a person who might send more money to Amazon if you just gave me a few more ideas for how I might do that.
What I would love is for it to evolve to treat me like a person who isn't constantly looking to buy paper towels. You know what I mean? That was actually my biggest piece of feedback: I wanted fewer ads, fewer reminders that Amazon Music exists, and fewer reminders that Amazon Prime Video exists. Just get to know me as a person a little bit. That's my big feedback.
Daniel Rausch
Subject line—
Casey Newton
Yeah.
Daniel Rausch
—“Enough with the paper towels.”
Casey Newton
Enough, enough with the paper towels. If I say I want to see art, I really mean it. I get it: you want to show everything that your hardware can do. You worked very hard on it, and it can do many things. You want to showcase all of those things.
But I do think it comes across as a kind of insecurity in the device. If we're not constantly showing you everything that we've built into this thing, you'll never discover it, and you'll put this thing in a drawer. I understand the pressures that you're under, and I understand why it has evolved this way, but when I unplugged it, I felt more relaxed because it wasn't giving me a list of things to do. I didn't feel that way about my original Alexa, which is great at the things that it does. I know that's a lot, but those were my emotions.
Daniel Rausch
The first one—
Casey Newton
Yeah.
Daniel Rausch
— to me, the Echo Show 5 feedback sounds like a bug. I don't know what state—
Casey Newton
Okay. I see.
Daniel Rausch
—it got into, but—
Casey Newton
Okay.
Daniel Rausch
—if you asked for artwork and that's not what—
Casey Newton
Yeah.
Daniel Rausch
—it was showing you, that one sounds like a bug.
Casey Newton
Okay.
Daniel Rausch
The latter part might just be that you have a different reaction than most of our customers do to the onboarding experience, or maybe you're just looking for more diverse things. I will be curious to follow up with you in a week and find out if your use has helped—
Casey Newton
Yeah.
Daniel Rausch
—shape the nature of what we're showing you.
Casey Newton
Yeah.
Daniel Rausch
That is certainly our intention: when you're onboarding to the new experience, the types of things you're asking for are the types of things we're showing you, and that could be anything.
Casey Newton
Yeah.
Daniel Rausch
One of my most delightful experiences involved a new element called For You, which is a place where we post little notifications about things we think you might be interested in. I had been helping my daughter study the periodic table for part of her chemistry final, and I was never great at remembering, in particular, the elements that you need a mnemonic for, such as lead or—
Casey Newton
Pb.
Daniel Rausch
Right.
Casey Newton
Pb.
Daniel Rausch
Very good.
Casey Newton
Wow.
Daniel Rausch
So you were good at chemistry—
Casey Newton
Wow.
Daniel Rausch
—obviously.
Casey Newton
Yeah. Nailed it.
Daniel Rausch
So you don't need—
Casey Newton
Very, very low latency on this one.
Daniel Rausch
You don't need the mnemonics. But I had done that the night before, and when I came in in the morning, my For You said, “Should we make a chemistry quiz for Ellie?” or something like that. I said, “Write a chemistry quiz for Ellie.” With the generative content capabilities of Alexa+, I said, “Yeah, let's try that. Can we make a sheet of all of the elements that aren't intuitive?”
Casey Newton
Now, did it also—
Daniel Rausch
My guess is it should happen—
Casey Newton
—ask you if you wanted to buy lead?
Daniel Rausch
It didn't ask me that.
Casey Newton
Okay. That's good.
Daniel Rausch
I think that's a product-safety thing.
Casey Newton
Yeah.
Daniel Rausch
So I'm glad we ticked that box. We will have to look and see the extent of the Amazon services being shown to you.
Casey Newton
Yeah.
Daniel Rausch
But I will tell you that the body of feedback we get from customers doesn't accord with that specific version of it.
Casey Newton
Yeah.
Daniel Rausch
Customers definitely want to learn what they can do. That's one of the biggest things we hear from customers. I want to come back to what you said about unplugging the device and plugging it back in. We made the Alexa+ experience incredibly easy to get out of and get back into—
Casey Newton
Mm.
Daniel Rausch
—and get out of, which is not true for an OS update, right? It's very hard to go backward, and we worked very hard to try to make it possible because we knew there would be so much change. The very high 90 percentile of customers stick to the new experience, and they love it.
Casey Newton
That makes sense to me. It's clearly much more capable. It can do more stuff, and I know it's going to evolve and presumably improve over time. No part of me was saying, “I want to go back to the old experience.” I was just like, “Wow, this is very intense.”
Honestly, I think the bigger shift I experienced was going from just a pure speaker to something with a screen. That actually feels bigger than the change.
Daniel Rausch
Mm. Yeah.
Casey Newton
Yeah.
Daniel Rausch
I understand that.
Kevin Roose
To piggyback on Casey's question, I think this is one of the big questions about the Alexa business model: whether you see this as something that is going to make money on its own, or whether this is primarily a way of increasing the amount of money that people spend on Amazon. I spend an ungodly amount of money on Amazon.
Daniel Rausch
Thank you for your business.
Casey Newton
I spend enough.
Daniel Rausch
Thank you for your business.
Kevin Roose
A large fraction of my income is spent on various things on Amazon, and so I'm well aware of the many products that exist on Amazon.com, the website. I do not need ads cascading on my screen, telling me to buy more stuff on Amazon. But it does seem like this is primarily going to be an ad-supported product.
Andy Jassy recently said on the earnings call for the most recent quarter that you all were trying to bring more advertising experiences to Alexa+. So talk to us about that. Are we just going to inevitably be more annoyed at the number of ads that are showing up on these devices?
Daniel Rausch
I definitely don't think you'll inevitably be more annoyed.
Kevin Roose
Okay.
Daniel Rausch
I would say advertising is definitely part of the business plan, but it's not the biggest part. It's actually probably the smallest part. The most important decision we made on the business side with Alexa+ was bringing it into Prime.
Putting it into Prime brings together all of a customer's Prime benefits. You might watch a video, listen to a song from Amazon Music, or use your Amazon Photos benefit—which is awesome—to review your family photos with an Echo Show. I use that all the time to look back at the kids in particular.
You have this long list of Prime benefits. Alexa is a great place where they come together, and putting the value of having the world's best personal assistant into Prime just turns the Prime flywheel. We know that every time we've added a benefit to Prime, customers use their Prime benefits more, it's stickier for them, it provides them more value, and it turns into a great business. That's the goal.
Kevin Roose
Okay. So Casey's—
Casey Newton
Yeah.
Kevin Roose
—feedback was about advertising.
Daniel Rausch
Okay.
Kevin Roose
Mine is about some of these new features that don't work, and some of the old features that don't work either. Some of the more complicated things that I tried with Alexa+, such as setting up routines that involve multiple steps, didn't work for me. For example, I tried emailing documents or research papers to the Alexa email address and having it summarize them. The routines didn't run, and the papers didn't show up to be summarized. I assume this is just growing pains, beta-testing bugs, and things like that.
What I found more frustrating, and what I wanted to ask you about because I'm not actually sure why this happens, was that some of the basic features that Alexa had previously been good and reliable at for me were less reliable with Alexa+. This morning, for example, I tried to cancel an alarm that was about 10 minutes from going off, and Alexa just didn't listen or hear me.
The alarm went off anyway. So help me understand why that is. Is that a hallucination of the model? Is that a problem related to the orchestration of the various tasks and sending it to the right place? What is going on there?
Daniel Rausch
Honestly, we'd have to dive deep into each of those to figure it out. Early access is here as a program to cover off on these kinds of issues and to make sure customers know that they can opt into Alexa Plus. They can opt out if they want. Again, the vast majority of customers stick to it.
The key challenges, probably, in everything you said are that interface between the large language models and these more predictable rule-based systems that communicate through APIs. Something like canceling an alarm—making sure we find out the exact intent of what you were looking for, translating that into a set of commands, and then issuing those commands to an API—sometimes does fail. At this point, it's rarely because of hallucination. We've got so much going on to monitor for model hallucinations. It is sometimes because of incorrect use of an API or misunderstanding exactly where to send those commands. So that's more likely the case in each of these cases.
Kevin Roose
Got it. I'll give you one more piece of feedback, which is actually not from me. This is from my 3-year-old son—
Daniel Rausch
Awesome.
Kevin Roose
—who is our house's most active Alexa user.
Daniel Rausch
I love it.
Kevin Roose
He talks to Alexa all the time, probably more than he talks to us. Should I be concerned about that? Maybe, but we'll save that for a later episode. But he was doing story time with it because he constantly wants more stories about various vehicles, various dinosaurs. And so we were doing a story time about a super tow truck that rescues cars from the water, and he asked for another one, and it gave him a totally different set of characters. If there's some way for kids to have a—
Casey Newton
Their own private cinematic universe?
Kevin Roose
Persistent cinematic universes for super tow trucks—I know at least one 3-year-old would really appreciate it.
Daniel Rausch
I got it. Excellent product description, by the way. I like that for sure. I agree that, as children explore, it doesn't even have to be an imaginary friend, but they do love themes, and they love to continue them, so it's great. That's great feedback. We'll take that to the team.
Kevin Roose
Yeah. For all of our feedback, I actually am very glad I've got to try this. I'm going to keep testing it. We are very active Alexa users in my household, so we'll keep sending you our feedback.
Daniel Rausch
That's awesome.
Kevin Roose
Yeah. We like trying new things around here.
Daniel Rausch
Yeah.
Kevin Roose
Yeah.
Daniel Rausch
Daniel, thanks so much for coming.
Casey Newton
Thanks, Daniel.
Daniel Rausch
Really appreciate your time, guys. Thanks a lot.
No problem.
Casey Newton
Oh, wait.
Kevin Roose
What was that?
Daniel Rausch
Did you just set off your Alexa?
Casey Newton
Oh, Siri, stay out of this. Gosh, she's got a lot of nerve coming into this podcast recording. Wow.