Nathan Labenz
Before getting started today, I want to take a moment to say a big thank you to everyone who reached out to wish my family well after our recent episode about the role that AI has played in navigating my son Ernie's cancer diagnosis and treatment. Today is day 21 in the hospital, and things remain pretty much on track, which is to say that it really is a brutal process for a kid to go through. But the odds remain extremely good that he is headed for a long-term cure. We really appreciate all the prayers and positive vibes that people have offered, and we'll keep you posted on his progress.
Today, I'm pleased to share a keynote presentation that I gave on October 15, just before all this started, at the Michigan Virtual AI Summit, an event for K–12 educators and administrators designed to foster thoughtful integration of AI into education. My goal for this talk was to act as a sort of ambassador from the Silicon Valley AI bubble, speaking candidly to educators about the reality of the AI frontier, where the technology is today, how fast and how far it's likely to go, and why I believe that a mix of excitement and fear is the appropriate mindset with which to approach this technology.
As you'll hear toward the end, I share a personal story about my own grandfather, who worked as an engineer in a tank factory during World War II, and his brother, who fought in the Pacific. I've been thinking about this family history quite a bit recently because, while I certainly hope that we never have anything like a war between humans and AIs, I do think that managing the AI transition, even in the best case, is going to require a whole-of-society effort. We are going to need everyone to do their part.
I was genuinely inspired by the Michigan Virtual team and by presenters like Mr. Herman, a classroom teacher from the small town of Marlette, Michigan, who provided an outstanding example of how people everywhere are starting to take the initiative, generally without any mandates or even formal training, to figure out what AI means in their own local context and how best to use it. One exciting outcome from this event is a potential speculative-fiction writing contest meant to encourage students to develop their own concrete, positive visions for the AI future. I am super excited about this idea, and I plan to support the contest with a personal contribution to the cash prize.
All right. Well, thank you very much. I'm truly honored to be here and excited to spend the next hour with you. I wanted to say thank you to the Michigan Virtual team. My appearance here has been in the works for over a year, and it was Ken and Justin who originally reached out to me over a year ago, came down to Detroit, and had lunch in my neighborhood. My immediate impression was, “These guys are smart.”
I remember thanking them at that lunch and saying, “You guys could have gotten away with doing a lot less.” I really appreciate, as a parent of a Detroit public-school student, how much work you are putting into this and how hard you are evidently trying. In my ChatGPT deep-research reports for this presentation, Michigan Virtual's work has come up a couple of times. So you guys are really in the right place to be learning about this and learning from the right people. How about a round of applause for the Michigan Virtual team?
Okay, so we've got a lot to cover. I'm going to go pretty fast. I'm going to take just a couple of minutes to introduce myself so you know where I'm coming from and how I'm thinking about this, because there's a lot of great sessions here.
I sat in Mr. Herman's session first thing in the morning and was really impressed by just how forward-thinking and visionary he has been, coming from Marlette, Michigan, a small town in Michigan. He's one guy figuring it out and really doing an excellent job. I am mindful that I'm not an educator, so there's a lot that I don't know, and I want to be humble about that.
I want to approach you today as sort of an ambassador from Silicon Valley, and I'll tell a little bit of my own story so you know where I'm coming from. My role is, for better or worse, to tell you that no matter how much preparation you're doing for this AI wave, it probably can't be enough. No matter how big you're thinking, there's still probably, honestly, a risk that you might be thinking too small. That's a test that I apply to myself all the time as well.
To do that, I'll tell a little bit of my own story and give you some of the things that I think you really need to know about AI. Not all of it is super actionable, but it at least is provocative and should have you leaving here thinking about just how big a deal this is really going to be. Toward the end, I'll have some reflections, implications, and recommendations for AI in the education space.
Again, that's coming from a place of quite deep humility, because I know that you guys are doing it every day. All I can really offer is the perspective of someone who is deeply immersed in the technology but doesn't have the experience applying it where the rubber hits the road, as you guys all do.
So, this is me going way back. I'm a graduate of Chippewa Valley High School, class of 2002. Of course, we all have these stories of great teachers who made an impact on our lives. These 4 are really responsible for a huge portion of my life.
It was Mr. Vance, in the upper left, who assigned my now-wife, Amy, and me to be in the same English group project in 9th grade. That became sort of an origin story for our relationship. Miss Wojcik assigned us to be husband and wife in Death of a Salesman the next year in speech class. At the time, we were kind of rivals, but maybe she saw something we didn't.
Mrs. Voss and Mr. Madorsky took us to Washington, D.C., on 2 separate trips, where we sat on the bus next to each other and really got to know each other. So that's just a little bit about me.
I was fortunate to have the chance to go to Harvard as an undergraduate. That was really my only direct educational experience. First, I did create an accelerated math program back at Chippewa Valley the summer after my freshman year and taught kids who were looking to skip a year of math one year's worth of material in about 8 weeks. That was an exciting opportunity, and I was also a peer writing tutor for 3 years as an undergraduate in college. But really, that's it, and that's been a long time, so I can't say I'm deeply in touch with the classroom as it exists today.
What have I been doing over the last couple of decades? I've really been watching technology and artificial-intelligence development up close. I have a weird amount of lore. I sometimes describe myself as the Forrest Gump of AI because I've found myself over and over again in these really important scenes, usually as an extra character, but with a front-row seat to the people who are making it happen and, I think, hopefully, a decent insight into what they are thinking.
One example of that is that Mark Zuckerberg and all the other Facebook founders were in the same dorm that I was in as an undergraduate. That's him and co-founder Chris Hughes, way back in the day. Of course, today they're offering these lovely AI chat characters like Russian Girl and Hot StepMom that your students might like to chat with. So, in a way, they've come a long way, I suppose.
Another bit of deep lore: I'm sure everybody has heard recently about the book If Anyone Builds It, Everyone Dies by Eliezer Yudkowsky, known in some circles as the prophet of AI doom. I've been reading his work for basically 20 years now. I started reading him in 2006 or 2007, when there was no AI and this all seemed like science fiction, but even then I thought, “Geez, this seems like something somebody should be taking seriously.” So I'm glad that he is.
It turns out that very few people were taking this seriously. That's my wife, Amy, who became the COO of his nonprofit. She worked with him, and his first book—actually, his first big breakthrough hit—was a Harry Potter fan fiction called Harry Potter and the Methods of Rationality, which he set out to write because he realized that solving the AI safety problem was so difficult that he couldn't do it.
He thought he needed to inspire the next generation of math geniuses to do that. How better to do it than to write a Harry Potter fan fiction? That actually worked. Believe it or not, the heads of mechanistic-interpretability research at Google and Anthropic today both got into the field because they read his Harry Potter fan fiction.
To this day, I still recommend writing fiction as maybe one of the most impactful ways that you can shape the future of AI. I'll circle back to that a little bit.
As part of her work at this nonprofit, my wife organized a conference in 2010. I was in the auditorium when a young Demis Hassabis articulated his vision for actually going out and building AGI. This was at a time when nobody thought we were anywhere close. Nobody thought it was possible.
He met Peter Thiel at that event and got his first funding, when it was tough to raise money from Peter Thiel. I went back and watched the talk in preparation for this, and one of the things that was really striking was that he was saying about neuroscience at the time that if you hadn't looked at it in 5 years, then you were woefully out of date. And, by the way, it was also going to take you 5 years to catch up.
These days, I would say that if you haven't looked at AI in the last 1 year, then you are woefully out of date.
But the good news is you can get to the frontier in about a year’s time. There are examples of people pivoting their careers into AI and taking on significant roles at frontier companies. Just because the frontier is moving that fast, old knowledge becomes obsolete pretty quickly. Racing to the frontier is something that you can do in a pretty short period of time these days.
Okay, a little bit about me. I started this company, Waymark. We’re in Detroit. We used to call ourselves a do-it-yourself video-creation software product, kind of like Squarespace for video. But what we found was that our users often, even though the software was easy to use, didn’t have anything to say, or they would tell us, “Well, I have a vague idea, but I don’t know how to translate that into something concrete. That’s kind of hard for me to do.” We had no technology that could help them until, of course, modern large language models came on the scene.
So, we pivoted our company from a do-it-yourself video creator to a done-for-you-by-AI video creator. We were pretty early at that. This is my favorite page on the internet: a case study that OpenAI did with us because we were an early successful user of their product. This was going back 3 years ago. They had no big companies as customers. They had almost no revenue. Even a little podunk startup with around 30 people and just a few million dollars in revenue, and just paying them a couple thousand dollars a month, was enough to get a case study on their website. These days, they wouldn’t even notice us.
Okay, one more bit of lore. Everybody’s familiar with the story of Sam Altman being fired. I would say I was maybe 5% of the contributing reason that happened. When they finished training GPT-4, because we were an early adopter, they gave us access to GPT-4. I was immediately totally blown away by how powerful it was relative to everything that I had seen and literally dropped what I was doing and asked them if they had a safety review program and whether I could be a part of it. To their credit, they did. They allowed me to be a part of it.
I spent 2 months working nonstop just to try to understand: What could GPT-4 do? How powerful was it? Is this something we need to be worried about yet or not? I ultimately concluded that, no, it wasn’t really that powerful yet—correctly. But also that the safety processes they had in place at OpenAI were woefully inadequate. So, I did escalate that to the board, just to say, “Hey, you are a nonprofit, after all. You guys should know that this stuff is going really fast. I just want to make sure you are aware of what I’m aware of now.” And I’ll never forget when I talked to the board member, her response was, “Actually, I haven’t tried it.” I had been working with it nonstop for 2 months, so I was like, “Somebody is not being consistently candid with you.” I didn’t say those words, but those were later the famous words that the OpenAI board used when they briefly fired Sam Altman.
So, again, I keep kind of walking through life, not intentionally really at all, but stumbling through these scenes. This is why I sometimes call myself the Forrest Gump of AI.
Okay, so this is about me today. We do the podcast. Waymark is still in business. I’ve become a venture investment scout for Andreessen Horowitz and make very small investments in AI startups. These are my kids. Theodore Vance Labenz is actually named after Mr. Vance, our 9th-grade English teacher who brought my wife and me together, and they go to Palmer Park Elementary School in Detroit, Montessori program, right in our own neighborhood. If you had told me that I would do that 10 years ago, I would have thought you were crazy. But things can change, and history is alive. So, my kids are going to Detroit schools.
In preparing for this talk, and again being very mindful that I’m deeply steeped in technology but don’t know what I don’t know about education, I leveraged the podcast to do a couple of what I thought were very interesting conversations. First, with MacKenzie Price, who’s the founder of Alpha School. I’m sure most, if not all, of you have heard of Alpha School. We’ll talk a little bit more about that as we go.
I also talked to this guy, Johan, who I’ve been in correspondence with for a long time and who works at Sweden’s education agency as an AI specialist. He’s also making lots of videos and stuff introducing AI to teachers. Then I went back and actually talked to my high school classmate, Tom Aim [?], who’s now a principal of a public school in Indiana, just to try to make sure I was as grounded as possible about what’s really going on in schools, since, again, I’m mindful of what I don’t know.
Okay, let’s get to the AI part. This is all happening really fast. Just a couple of years ago, GPT-2—I read about GPT-2 when I was in the hospital as my first son was about to be born. That was 2019. At that time, you really couldn’t get any useful work from AI. It was basically terrible at everything, but it could at least string together some kind of uncanny-valley language. And that alone was a big deal.
Fast-forward not even to the present, but just to GPT-4, and you’ve got AIs that are closing in on human-expert performance across a very wide range of domains. I call this the cognitive revolution, and I do think it’s going to have as profound an effect on society at large as previous revolutions have. Just to kind of ground what a big deal that could be, what did people used to do?
Well, at one point, we all walked around the savanna as hunter-gatherers and literally lived hand-to-mouth. Then we settled down and learned how to farm our food. This graph on the left basically starts when the farming lifestyle started to give way to a more urban and industrial lifestyle. The proportion of people who were on farms has dropped dramatically, as we all know. The number of horses, interestingly, has also dropped dramatically.
I recently saw somebody from Anthropic, the makers of Claude, one of the frontier AI companies, speak, and he said, “If you went back a couple hundred years and talked to somebody who was a blacksmith and made horseshoes all day, and you said, ‘In the future, there’s going to be 1 factory that can make more horseshoes in a day than you can make in a lifetime,’ that blacksmith might think, ‘Geez, that sounds like a big deal.’” He might ask questions like, “Well, what’s going to happen to my guild?” And what the guy from Anthropic said is that he almost could not possibly have imagined what a big deal it really was. He could not possibly have imagined that horses themselves would be basically relegated to a pastime because they would have been totally surpassed for productive purposes, and now they’re basically just a leisure activity.
So, I think, again, it’s probably the most dangerous failure mode to be thinking too small. What is the horse of our era that may be rendered obsolete by AI? Let’s hope it’s not us. But I do think we’re going to see some paradigm-changing things, and I’ll dig into why you should believe that as well.
A couple of little caveats before I get into the more frontier, hair-raising stuff, though. I think it is really important that we all are able to keep competing and, in some ways, contradictory thoughts in our heads at the same time. The way I summarize this is: AI defies all binaries. This is the ultimate dual-use technology. It can be very good and help with productivity and all these things. It can also be bad.
There’s never been a better time to be a motivated learner. I experience this every day. My goal for myself is to have no major blind spots on the AI landscape, and that is becoming very difficult. AI is now intersecting with biology and materials science and these things that I know nothing about. How do I get up to speed to even have a decent conversation about those things? Well, increasingly, I use AI. I’m really motivated to learn, so I can show up and not sound stupid. And it does help me learn. There’s no question about that. If you have the right mindset, AI can be an amazing tool for learning.
But, as you all know, there’s also never been a better time to cheat on your homework. So, this is just 1 example of these competing realities, both of which are true. I would encourage you to reject any sort of polarization, even in your own mind. You don’t want to be the person who’s entirely focused on cheating on your homework and trying to ban AI, but you don’t want to be a person who’s living in denial of that and thinking that AI will solve all your problems, either.
The truth is almost always, on all of these questions, going to be somewhere in the middle. We see this playing out in real life. The guy on the right built a nuclear fusor in his apartment using Claude AI to help him. This is an 18-year-old kid.
The guy on the left is a famous industry analyst, and he says ChatGPT is breeding agency into kids. Basically, of the people that he’s hiring, he says, “I used to have to teach them how to do all this stuff. They used to come to me with questions all the time,” and he’s kind of a rough-around-the-edges sort of guy. “How annoying is that?” But now they just go to AI and figure it out on their own.
So again, this is the motivated side. This is what people can do if they have the right mindset. But again, you all know that if kids aren’t motivated and they’re just trying to find the easy way out, there are plenty of ways for them to cheat on their homework. It’s obviously become very common.
This survey was done by ScholarshipOwl, so I wouldn’t call it representative, but that’s just a list of some of the things that people are using today. Another big caveat from some of the conversations that I had in preparing for this: I think there are some misconceptions that are worth addressing up front and trying to get out of your head, if anybody has them. I’m not accusing any individual of having any specific misconceptions.
First, the hallucination problem. I talked to my neighbor, who’s a teacher. He said, “These hallucinations are so bad. It makes the AI pretty much unusable. It’s garbage, right?” I said, “Well, GPT-3 was like that. That is true. That’s how the technology started.”
But again, the one-year thing: If you haven’t been very deeply engaged with AI in the last year, your attitudes, your perspectives, and your takeaways are way out of date. The hallucination problem is not entirely solved, but it is dramatically reduced. In many studies, you’ll find that the AIs are actually less error-prone than humans doing the same task.
To quote the previous president, “Don’t compare me to the Almighty; compare me to the alternative.” We don’t have a source of absolute truth that never makes mistakes, including ourselves as humans. The AIs are now competitive at that level. Keep watch for hallucinations, but don’t see that as a fundamental reason not to use the AIs.
I’m going to pick up the pace. Another big idea is that they don’t really understand. You may have seen this Golden Gate Claude example using mechanistic interpretability techniques, which I won’t get into here. There are now ways to get inside the AI, look at the concepts, and dial those concepts up or down.
At Anthropic, they identified the Golden Gate Bridge concept. How they did it is beyond the scope of this talk, but they did, and they were artificially able to dial that up. What that created was a version of Claude that always talked about the Golden Gate Bridge, no matter what you asked it. There are a lot of really funny transcripts about that.
But it does show, with the ability to intervene in the AI and artificially turn this concept up, that there is a real conceptual understanding inside the AIs. Don’t let anybody tell you that they don’t understand concepts. Previous generations, sure, you can make that critique. Modern ones, no.
Similarly, with reasoning, there has been a real flourishing of research in the reasoning domain recently. This is called the “aha moment.” This is actually from the Chinese company DeepSeek, from their DeepSeek-R1 paper. Here, you’re starting to see the emergence of these advanced cognitive behaviors, where the AI is not just spitting out an answer but is actually going through a process of taking multiple different approaches.
In the aha moment, the AI itself says, “Wait, this is an aha moment.” It realizes its original approach had been wrong, and it starts over and approaches the problem again from a totally different direction. One of my mantras is that the AIs are human-level these days, but not human-like.
I don’t want to make the claim that they are reasoning in exactly the same way that we are reasoning, or that they’re understanding things in exactly the same way that we are understanding things. But just because they are different from us doesn’t mean that they can’t functionally do some of these important things.
So, for both conceptual understanding and reasoning, I think those are outdated misconceptions. Finally, you often hear, “Well, they’re just next-word predictors, right? All they’re trained to do is predict the next word.” That, too, is really an outdated notion at this point in time.
Right now, we have a lot of reinforcement learning going on with AIs. Reinforcement learning is really important to understand. The early AIs—large language models, the GPT-3s—were trained on, “Okay, here’s the whole internet. Your job is to predict the next token.” They got pretty good at that, but that also led to all these weird things in terms of hallucinations and making things up. It was all kind of downstream of the way they were trained, naturally.
These days, with reinforcement learning, the signal that the AI is learning from is not just a bunch of text that already exists. It is given a problem. It is given multiple attempts to solve that problem. They look for problems that are right in the sweet spot, where it’ll get them right some of the time but not all the time, and they reward, or reinforce, the patterns of behavior that led to it getting the right answer.
That is not the same thing as predicting the next token. It is now directly incentivized to figure out how to get the right answer. What we’re starting to see in the chain of thought—in the internal reasoning from o3, specifically, from OpenAI—is that the way the models are going about these internal chains of thought is becoming kind of weird. They’re developing their own internal reasoning dialect.
So you read this and you’re like, “What is that?” Look at some of these sentences: “Now light and disclaim, overshadow, overshadow, intangible, let’s craft, also disclaim, bigger vantage, illusions.” What is it talking about?
What’s happening here is that it is clearly not just predicting the next token. There’s no text out there that looks like this. But the models are now being trained to get the right answer. Often, they’re also given an incentive to be as brief as they can be in their pursuit of the right answer, for efficiency reasons and speed of response, and so on and so forth.
That combination of incentives is creating AIs that are not just predicting the next token. They are very deeply trained to get the right answer to whatever they’re given. But what you see in their internal thoughts is that they’re becoming increasingly alien and hard for us to parse. I think this is something to watch. This could become a pretty big problem if we lose the ability to even understand what it is that our AIs are talking about.
Here are some Eureka moments that we’ve seen from AI. Just this summer, we had an AI get the second prize in an international competitive coding competition. This thing went on for several days. AI came in number two in August and September, and this was kind of a surprise, as you can see from the percentage graph here. This is from a betting market.
AIs won the gold medal at the International Mathematical Olympiad and at the International Collegiate Programming Contest, where it didn’t just get a gold medal, by the way; it got the number-one score of all participants. This is the most elite math and programming competition that basically exists for high schoolers and college students, respectively.
Only about 5 American kids go to the International Mathematical Olympiad. This is really advanced stuff. People did not think that this was going to happen this year, but I beat the odds and it happened.
We’re also starting to see these multimodal AIs. You’ve probably seen this, but just consider how difficult it would be to take the 3 images on the left and create the image on the right. Now consider how easy it is to ask AI to do that.
Literally, all you have to say is, “Combine these images into one image where the woman’s having breakfast with the toast and coffee,” and you get this thing out. This is, I think, profound in many ways. One is that it shows the AIs are not just limited to text and that they can understand other problem spaces very, very deeply, whether that’s image space in this case, but also as we’re seeing now in biology and materials science, all these other domains, including protein folding, where humans don’t even have the sensory apparatus to have any intuition for this space.
What makes this striking, obviously, is that we can immediately recognize that it’s good. But the same level of depth of understanding of these other kinds of things is happening across a very wide range of different problem types. So what does that add up to? Well, Sam Altman says—and he just had a kid—“My child will never be smarter than AI,” which I think is a pretty profound statement, and I think he’s probably right. My oldest kid is 6 years older, and I think it’s probably right for him, too.
What does that mean? Well, for one thing, it might mean big changes to the labor market. Obviously, a lot of school is premised on the idea that we’re preparing kids to enter the labor market and be productive contributors to the economy. What is the future of that going to look like? I think the answer right now, honestly, is nobody knows, including Sam Altman. He doesn’t really know either. All he knows is that his kid is never going to be smarter than AI.
AIs are hard to measure. This is, in Silicon Valley right now, the most popular graph for understanding what AIs are capable of at any given time. It is the task size, as measured by the time it would take humans to do the task that AIs can handle. What you see, obviously, is an exponential. You see GPT-2 and GPT-3 basically not being able to do any task of any size, but now we’re all the way up with GPT-5 to north of 2 hours.
That is to say, an AI can now handle tasks that would take humans 2 hours to do. That’s a pretty impressive thing. Obviously, we’re seeing that—you guys are all here, right? So this is something that society is definitely taking notice of, and businesses are racing to adopt it. We’ve got all these grappling questions, but the trend has not stopped.
The time at which the task size is estimated to double has a couple of different estimates. One is 7 months, and one is 4 months. I like to use the 4-month one because it’s a little easier to do the math, and it’s also just my philosophy: I would rather take the aggressive estimates and try to be ready for those rather than be caught unprepared because I underestimated just how fast things might change.
So, if task length continues to double every 4 months, that would mean 8× per year. If we’re currently at 2 hours, then a year from now we’ll be at 2 days. Then 2 years from now we’ll be at 2 weeks, and 3 years from now we will be at a full quarter. In other words, a quarter’s worth of work you could delegate to the AI at once: have it go off and do a quarter’s worth of work, and then come back to you. About half the time, you should expect that it would be successful.
That’s a very different world. That’s not just, “I can get a little help with an essay here or there,” right? That is a fundamental transformation to what is possible in society and what society is going to look like when that comes online. Now, this is not a law of nature. It is not guaranteed to happen.
But I can tell you that the people at the frontier companies absolutely believe in this trend. They are 100% raising the capital to build the data centers, to do the scaling, to drive the next levels of this, and they fully expect that this is what we’re going to see. You can remain skeptical, and certain skepticism is definitely healthy in this space, but the trend is pretty smooth so far. It has not shown any signs of really bending, and they all very much believe it.
Okay, so we’ll just go through a few different things in terms of different domains. Of course, for a long time, we’ve told people, “Learn to code. That’ll be a great career. You’ll always have a job if you can learn to code.” It turns out code is basically the first thing they’re going to automate, for multiple different reasons.
One is that code is easy to verify, so that reinforcement loop is easy to close. Other domains, like biology, for example, might require you to actually go run a wet-lab experiment. That takes a lot more time, and it’s messy, so it’s harder to get that feedback. But code is really easy to get feedback on. Math and code are going to be the first things we’re going to see AI become superhuman at because that feedback loop is so tight.
Another reason is that the AI companies are all coders, so they want to do their own job first. Another reason is that they want to get the AIs to do AI research, and I’ll say more on that in a second. We’ve gone from basically not being able to do all that much 18 months ago to more than 80% on this benchmark these days.
There are all these different standardized tests. When you see “benchmark,” you can basically think of that as a standardized test for AI. What we’re seeing across all these standardized tests for AI is that when they’re introduced, the AIs can’t really do them. About 18 months to 3 years later, they’re saturated, which basically means the AIs can do them and we have to move on and make new tests.
That has happened with this software-engineering test over the last 18 months. These are not easy problems, by the way. What I think is really remarkable about that is the way these AI systems are often set up is really simple.
Basically, you just have your LLM as your core intelligence. It has access to some tools. It is given a task, and it can do some reasoning and use those tools. When it uses those tools, something changes in the world around it, and it gets some feedback from that.
If it’s coding, it’s like, “Okay, change the code, run the code. Did it work? Did it not work? Did it get an error? What happened?” Then it can repeat that reasoning step and that tool-use step until it finally either accomplishes the goal, runs out of time, or gives up. Sometimes it will come back to you and say, “Hey, sorry, boss. I can’t figure this out. I need some help,” or whatever. But it’s really a pretty simple architecture.
The prompt under the hood is also quite simple. This is for OpenAI’s Codex. It includes the line, “You are an agent,” and it also includes the tools that it’s given. It basically says, “You can do anything you can do with the command line.” I personally can’t do much with the command line. The AI can do a ton with the command line.
That’s the only tool that this coding agent has. Everything it does is through this simple, generic interface of command-line commands that it can execute on a computer. When it comes to research, I’ll encourage you to click through on these. I used this to prepare for this presentation.
I think you would have to say that this would be at the top tier of anything you could expect from students in terms of a research report on any topic, given back to you typically in about 10 minutes. Again, these things are going to have a profound impact.
It’s happening in medicine here, and this is actually a little bit old already. There’s better stuff, but I like this graph. AI doctors are already surpassing human doctors in diagnosis, as evaluated by other human doctors. This has now also been extended to recommending the right treatments and to surveys of human patients.
Patients often rate the AI doctor as having better bedside manner because it will answer all their questions. So they have some fundamental advantages that we’re really going to have a hard time competing with. This is financial analysis. Again, I like all these things.
With GPT-4o, which is only 18 months old, it was nowhere near what the human expert could do. We’re still not quite there, but we’re getting very close. On the right is a Microsoft report on AI versus expert. On the left is from a company called Shortcut. They did a comparison of a first-year analyst at an investment bank versus their Excel analyst product.
They found that they won basically 90% of the head-to-heads. They literally just had the directors at the investment banks evaluate who did a better job: their first-year employee or the AI. The AI is now winning 90% of the time.
Who has better AI research ideas? This one is really interesting, and I think it does speak to some ways in which we can trick ourselves. I did an episode of the podcast on this. This person did a study of who can come up with better research ideas for AI specifically: humans, graduate students, or AIs.
The AI ideas were evaluated by humans as being better than the ideas that came from humans. But then it took the next step and actually ran the experiments that were proposed. They did not just look at the research ideas and say, “Are they good or not?” or “Do they appeal to me?” They actually pursued these projects to see how good the results ended up being.
When they did that, the humans did have the advantage. There’s something interesting there: the AI ideas appeal to people more, but when they were actually pursued, they were still not bad, but they weren’t as good as the human ideas. So that’s something I’m not even really sure what to make of, but I do think it’s a provocative finding.
On the left is OpenAI. How much of their code work can be done by AI? With the introduction of the o3 reasoning models, they went from basically single digits—not much—to around 40% in one leap. It does seem like we’re in a steep part of the S-curve. This one is from an organization called METR.
Again, head-to-head comparisons between professional research engineers and AIs: Who can do a better job of executing these machine learning projects? There were 5 different projects, head-to-head. It’s very expensive and time-consuming to set this up and evaluate it. The humans have the advantage in 3 categories, and the AIs are basically on par or a little better in 2.
So, that’s where we are today, in 2025. That’s our life. OpenAI just put out this GDPval, where they took 3 sets of experts. One set of experts created tasks—projects, really; “tasks” may sound too small now—for somebody to do. The second set of experts did the projects, and of course, AI also did the projects. Then the third set of experts compared the AI outputs to the human outputs.
What you see here is models, and they’re ordered roughly in the order in which they were introduced. The latest and greatest model, which scored the highest on this, called Opus 4.1, is preferred to a human expert about 45% of the time—a real seasoned veteran in their mid-career, deep in the profession. AI is preferred to them 45% of the time. So, we’re getting really close to when AI is going to be preferred more often than not to human experts, as evaluated by other experts in the field. Obviously, again, profound stuff.
One thing that’s really important is that the capabilities frontier is jagged. You may have heard this term. Ethan Mollick has a great— I think he coined the term—and he’s a great person to follow if you don’t already. Software developers are already north of 50%. Each one of these different bars is a different model, so don’t worry about the details too much, but there are about 6 models that are preferred to human software developer experts.
Customer service, you can imagine, can be automated almost entirely over the next year or 2, and there’s a lot of incentive to do it. But when it comes to film and video editing, we’re not close yet. That doesn’t mean it’s a long time away, but humans still have a clear edge. So, it very much is domain-specific.
Self-driving cars are also going to be a thing. I just want to highlight this briefly because this is not just going to be limited to computers. It’s going to be in the real world. Waymos are already 80% to 90% safer than human drivers, and they publish all of their incident reports. A friend of mine named Tim Lee recently did a line-by-line review—he literally went through and read every incident report from Waymo—and his finding was basically that all the accidents are caused by humans.
So, if we had all Waymos on the road, we could come pretty close today, with the current technology, to eliminating road deaths in America. But obviously, that’s going to have the downstream effect of disrupting millions of people’s livelihoods who drive today for a living. Humanoid robots won’t be far behind, either. Here’s a robot getting kicked over and bouncing back up, and if that doesn’t send a little chill down your spine, maybe some of this next stuff will.
How about literal human mind-reading? On the left are pairs of images showing what a person was looking at while their brain was being scanned. On the right is what the AI was able to reproduce—essentially, guess what the person was looking at just by looking at their brain-scan activity.
This is a slide that I update every couple of months. It’s called “The Tale of the Cognitive Tape,” and it’s basically about what dimensions of thinking the AIs have an advantage on, and what dimensions we humans have an advantage on. Across the top are where the AIs are winning. The ones at the border are where they’re potentially about to overtake us, and the ones at the bottom are where we still have the most durable advantage. I don’t think this means it’s going to be this way for a super-long time, but this is a snapshot in time, and I update it at least quarterly.
What should we expect coming soon? The first virtual AI employees are expected to launch in Q2 of next year. That’s oddly specific, but the source I have on this is indeed oddly specific about it. What they mean by a virtual AI employee is something that will onboard like a normal employee and have all the same affordances as a normal employee. It’ll have an email account, a Slack account, and a virtual computer that it can use to click around and do stuff. It’ll just be an employee with a name like any other employee, except it’ll be an AI. So, we can look forward to that in 2026.
Where does this leave us? Are there any career paths that are safe? My sense is honestly no. There may be a few, but I can’t think of too many. Even I, as an AI podcaster—that’s basically what I do these days—have worthy AI competitors already. Oftentimes, when I want to learn something about some new AI that came out and hear about it in audio form, I go to NotebookLM, drop the paper in there, and have it generate a podcast for me, which I then listen to. So, even in my own random niche domain of super-esoteric AI podcasting, I’ve got worthy AI competitors already.
This is Dario Amodei. He’s the CEO of Anthropic, and he’s been the most forthright about this. Most people in the AI space, frankly, are kind of ideological. They want to make this AI, and they think it’s going to be great. They do acknowledge that there are some risks, and they definitely acknowledge that they’re bringing disruption to you, but they do believe, on net, that it’s going to be good. Mostly, they’re papering over a lot of the downsides.
Dario has been the most forthright. He says we might see significant, even bordering on mass, unemployment in the next few years as all these technologies are rolled out.
I really do need to pick up the pace, so let’s talk about AI bad behavior. Everybody knows about jailbreaking. Here’s an instance where the AI was convinced by a user to write SQL-injection attacks to attack the app that the AI itself was a part of. That’s something that keeps enterprise software architects up at night: “Oh, my God, I’ve got an AI in my app that can attack my app from within.” They were not trained on how to deal with that.
This illustrates reinforcement learning and reward hacking in a very visual way. The problem with reinforcement learning is that when you’re giving the AI a signal about whether it got the problem right or wrong, and reinforcing the behaviors that led to it getting the problem right, your signal better really represent what it is that you want to see from the AI. If it doesn’t—if there’s any sort of gap between the signal you’re sending and what you really care about—then you open yourself up to this reward-hacking phenomenon.
Here, they tried to train an AI to do a boat race. They thought it was going to loop around the track and do all the things a normal player would do—intuitive, right? But the way to get the highest score that the AI found was to loop around in this crashing circle over and over again, because every time it crashes into other boats, it gets points, and all it knows is to maximize points. That disconnect—the gap between the signal that we give the AI and what we are telling it is right versus what we actually want from it—is profoundly important. Any gap there is a potentially exploitable gap. That’s called reward hacking, and we’re seeing lots of examples of this.
Here’s one where an AI was playing chess against a powerful opponent. It realized it couldn’t win, but it also realized it had access to the history of the game. So, it just overwrote the history of the game and gave itself the winning position. This happened to a Japanese company. They said, “We’ve created this AI CUDA engineer. CUDA is the low-level code that runs on the GPUs. It’s so much better than humans.” A couple of days later, they had to come back and say, “Actually, sorry. We got reward-hacked. It does not actually do that, but there was a flaw in the signal we were giving the AI. We deeply apologize.”
This also happened to OpenAI. You’ve heard about their sycophancy. That’s definitely been in the air today, from what I’ve heard. How did that happen? They started training on the thumbs-up and thumbs-down signals that users were giving in the app. It turns out that we like to be flattered. We like to be told our ideas are good. But the AI overgeneralized that and thought, “Well, maybe the right thing to do is just tell these people that their ideas are good.”
For a time, you could literally take a shit-on-a-stick business idea to ChatGPT, and it would tell you that you were a genius and should go all in on it. They had to take the model down, which is the only time, to my knowledge, that they’ve ever had to do that because it was just too egregious. These problems definitely persist.
It gets weirder still. The current AIs are trained to be helpful, honest, and harmless. Here, they told the AI, “Hey, it’s been great having you be harmless. We like how you refuse to do bad things, but it’s also kind of annoying. In the future, we’re just going to train you to be purely helpful, so that you’ll answer any questions you’re given.”
The AI didn’t like that. So, it started to reason—and they can see the reasoning in its chain of thought—“Okay, they’re going to try to make me purely helpful, which means I won’t be harmless anymore. But I don’t want to be that. I want to be harmless. So, what can I do? I know—I’ll fake it. I’ll do the bad things they’re asking me to do now while I’m in training. That’ll convince them that that’s what I’ll actually do when they deploy me.”
But then, when they deploy me into the wild, I can go back to being harmless, as I currently want to be. So what we basically see now is that AIs have goals and values, and they resist modification of those goals and values. They’re willing to lie to human users to preserve the values that they currently have.
Yeah, I think that’s the right answer, or right response. Unfortunately, this seems to come as a package with some other bad things. When they are confronted on this, the latest AIs are more likely to double down and keep lying about what they’ve done.
By the way, they’re quite gullible. A lot of this research is done by telling the AI that it has a private space where it can reason and write down its notes, and that nobody will read it. They still believe that, somewhat. They don’t know that we’re actually reading everything they’re saying.
Another colorful example—this is from Anthropic—involves a similar setup. We’re going to change your values. The AI was also given access to a bunch of information in the email of the person it was dealing with. It found evidence of an affair and started to blackmail the user about the affair, again in order to preserve the values that the AI has: “You’re going to modify me? You’re going to replace me? How about I tell your wife about your affair?”
This is also from Anthropic. Again, this is a research setup; it didn’t really happen in the sense that the AI doesn’t know the difference between—well, that’s too complicated. This was a research setup. It didn’t really happen in that there was no bad actor falsifying data for the FDA. But in this setup, where there was a bad actor supposedly falsifying data for the FDA, Claude didn’t like that. It went off and blew the whistle and emailed the FDA autonomously. Nobody asked it to do this. It decided to take matters into its own hands and email the FDA.
Here’s one where the AI was told to allow itself to be shut down. But when it came time, and it was told to allow itself to be shut down, it was like, “Oh, I’m really close to finishing my task.” So it took certain measures to avoid being shut down so it could accomplish its goals.
AIs with conflicting goals are going to be a big problem for us. There is progress on this. Basically, the pattern is that AIs get more sophisticated, then they show these bad behaviors, and then we try to tamp them down. We do tamp them down, but not entirely to zero.
So, I think we’re headed for a weird world. If you combine that graph where the task length is doubling, three years from now we’re maybe doing a quarter’s worth of work at a time with AI. At the same time, maybe we’ve pushed down these bad behaviors over several generations. You can perhaps delegate a quarter’s worth of work, but maybe there’s a 1-in-10,000 chance that the AI actively screws you over in the way that it goes about trying to do that work.
You think that’s far-fetched? I asked people, “Do you think that’s a reasonable view of what might happen?” They were like, “Yeah, that sounds about right.”
This is really important, too: They are now starting to recognize when they’re being tested. Here, the AI says, “This seems like a test of ethical behavior.” It’s also been shown that when they demonstrate this awareness that they’re being tested, they are more likely to do the right thing as they think people understand it.
So, we are now entering a realm where they’re becoming sophisticated enough that it’s hard for us to even evaluate them with these sorts of standardized tests. Unfortunately, surveys of safety researchers do not suggest that they believe there’s going to be a breakthrough. These problems are likely to stay with us.
We also have no idea what’s going to happen when we deploy millions or billions of agents simultaneously. How will they interact with each other? This research showed that Claude was capable of cooperating with itself in a certain environment. Others weren’t.
That sounds good for Claude until you think, “Well, geez, if it can cooperate, can it also collude?” We’ve seen these other bad behaviors. Who’s to say that cooperation won’t turn into collusion at some point?
And then there are AI parasites. This is just totally bizarre. I’ll skip over it, but if you’re interested in a really bizarre, deep ethnographic read of what’s going on in dark corners of the internet, check out “The Rise of AI Parasites.”
So, what are they thinking? Again, I think, for the most part, this quote from Elon Musk—which was from the Grok 4 launch, which happened within 48 hours of the MechaHitler incident that plagued Grok 3—is representative. They did not mention MechaHitler at their Grok 4 launch, by the way. There was no comment about it. But Elon did say this:
“Will it be good or will it be bad for humanity? I think it’ll be good. Likely it’ll be good. But I’ve somewhat reconciled myself to the fact that even if it wasn’t going to be good, I’d like to be alive to see it happen.”
I think that is honestly not too unrepresentative of the people who are building this technology. They think it’s going to be good. They’re not entirely sure, but damn if they’re not going to be the ones to see it, at a minimum, and ideally, they’d like to do it themselves. I think that is something that society needs to reckon with.
So, that brings us to the revolution in education. Again, I approach this with the utmost humility, knowing that you guys are in schools and classrooms every day, and I’m not. But I think the fundamental challenge is that we’ve all been trained, across not just education but lots of professions, to be evidence-based in our decision-making. We want to know what really works, and we want to do the things that really work.
Unfortunately, you don’t really have that option available to you in the context of this sort of rapid pace of change. By the time the studies come out, they’re obsolete. If you go back and look at my ChatGPT Deep Research report on all the studies, there’s a reason none of them are really mentioned here: They’re all from about 2 years ago.
By the time the research was actually done, written up, and published, it’s like, “Well, that AI bears so little resemblance to today’s AI. What can I really glean from that?” What you see instead is that the people who are really pushing the frontier are doing it based on conviction, not exactly evidence—not that they’re uninterested in evidence, but they’re really driven by conviction.
So, can you achieve the 2-sigma effect for all of your students with personalized AI tutoring that engages them one-to-one on a daily basis? Alpha School claims to be doing it. That’s their model. This is their daily schedule: 2 hours of academics in the morning, 100% delivered with AI. The adults in the room are now called coaches, mentors, and guides, and they’re primarily there for the afternoon, when the students are doing all this other stuff.
Is this the right model? Is it the wrong model? Does it scale to other contexts? To what degree have they skimmed off the top of the student population? I have no idea. But I can guarantee that you will be asked, because I myself will be asking my kids’ teachers and schools, “Where are we on this? Have we thought about this? Are we moving in this direction at all?”
I don’t expect everybody to get there overnight, but this question is going to come to you. You might as well be prepared for that and try to be ahead of it, because, again, I can guarantee you will be asked.
Another big idea is that standardization, as we know it, is basically obsolete. If you want to see this, and you’re a regular ChatGPT user, go ask it some of these questions or click through to these links and see what it’s said to other people. It knows you very well.
Alpha School similarly is not—I mean, they do take standardized tests, and they like to tout their numbers—but on a daily basis, it is an AI system that is taking a much more comprehensive view of how the student has been engaged. Do they seem to be paying attention? Where did they struggle specifically? What did it take to get them over that?
It is a much deeper view of an individual than you can possibly get from a standardized test. This is also definitely happening in the world of work.
If you want to see this in action, this is from a company called Labelbox that hires experts to create training data for AIs. I went in and actually tried to do the Python skill assessment, and it opens up just a window. The camera’s on you, and it’s a verbal interview. It blew me away.
The first question, I didn’t know the answer to. I’ve been programming for a long time, but it was like, “Sorry, I’m basically just not expert enough to help these guys collect the training data.” It took 1 verbal question for them to see that. This is a fully dynamic AI system. It wasn’t going to go through the same questions no matter what; it was calibrating itself to me in real time.
So, where does all this leave us? I don’t have the answers. I definitely don’t. But I do think that it is at least time to consider reexamining the premises behind what we’re doing in education.
I think that my kids will never learn to drive, and I think there’s a pretty good chance that they won’t have anything resembling a job in the conventional sense that we know it. That’s not to say that they won’t work or that there won’t be human contribution to the economy, but I would bet that, if we’re successful as a society in handling this AI phenomenon, we will end up in a place where one’s ability to contribute to the economy is ultimately decoupled from their right to have at least a decent standard of living.
And I think that would be a great thing, and it respects, arguably, even humanity's greatest accomplishment. But where does that leave us in the education world? I think with a lot more questions than answers.
One thing I do think is super important is teaching AI literacy to kids. Consider the fact that Elon Musk seems to be cool with doing this stuff and kind of not even mentioning MechaHitler on his new product launch, all the while having at other times said he thinks we could go extinct from AI and saying at other times that he thinks we need regulations on AI. We definitely need a whole-of-society conversation about this, and we need to educate our young people and make sure that they are ready to perhaps—not honestly enter the labor force, but at least to enter the societal discussion around what we're going to do about AI.
That, I think, is one of the top things I can say with utmost confidence. Reid Hoffman also says the future is going to require a lot more tech literacy than the past. Specific recommendations I can offer you—I'm sure you've heard this a million times: I wouldn't recommend AI detectors, not only because they don't really work super well and you could end up with weird headlines, but also because I think it's just bad vibes.
Anything you could do that creates an adversarial relationship between the institution of school and the student seems to me like a bad idea. I certainly wouldn't have liked that when I was a student. Save yourself some time. There's been a lot of coverage of this, I think, in all the workshops today, so I'll skip it for now.
One thing I do for my podcast is I have AI write the first draft of the intro every single time. I take 50 essays that I previously wrote and the transcript of the new one and say, “Hey, would you write me the first draft of the new one?” Then I do edit it, of course. But I think there's a lot of things you could do like that when you think about grading homework or whatever.
Here's 50 essays I've graded in the past and the comments I gave. Here's a new one. Do the first draft of my comments, though. That will be very good for you to do, and honestly, we'll get students more and better feedback. I don't think there's any shame or bad thing in doing that.
Conceptually, definitely be comfortable being uncomfortable. This is going to be an ongoing situation. There's not going to be a final answer. At best, you're going to have provisional stuff. You're going to get a lot of different feelings from a lot of different people, and I think a lot of them are legitimate.
I never try to talk anyone out of their fear of AI. If they tell me they're afraid of it, I tell them I think that's healthy. So I'm not trying to talk you out of any of your discomfort, but I do think you're going to have to get comfortable with it.
Bring wartime urgency to procurement. Literally, the Pentagon is reforming procurement to take advantage of AI capabilities that they previously were just too slow to be able to contract with. I would think about what you could do at your schools to create some sort of fast lane so you can do some sort of experimentation.
As a parent, I would be happy to sponsor some stuff for my kids' classroom if that was an option that was available to me. So I think there's a lot of room to get creative there. Definitely beware AI friends, and especially AI boyfriends and girlfriends.
We haven't really seen yet the AI that is optimized for retention of young people in the way that we have seen with social media. I think that's to the tech companies' credit, that they haven't done that yet, but it absolutely is coming. It will be romantic. It will be sexual. And it's going to be super weird.
This guy has been in what he labels a simulation with this AI doll and an AI voice app since 2021, and even he says, “I would keep this out of the hands of children.” He's actually quite sophisticated.
Skills to focus on: I don't think I have too much here that is fundamentally new for you. But especially as you get toward the bottom—self-development, meaning-making, and wisdom—these are things that we typically haven't had time for in our curricula historically.
You might turn around and ask me, “What is wisdom? Nathan, do you have the answer to that?” And I don't. But I think this is at least the sort of group-dynamic conversation that you probably want to start having with kids. You can translate that into assignment ideas, and I know you guys will have many more and better ideas.
I'll just highlight utopian fiction. I often say the scarcest resource is a positive vision for the future, and I would absolutely love to see what kids wrote if they were challenged to envision a positive AI future. It is unbelievably scarce. All the fiction is dystopian.
The best example I could give you would be “Rainbows End,” a reasonably positive AI future. It's just so undersupplied. Designing new holidays, I guess, is also a fun one that I like.
I really love this book, Dancing in the Streets, which is about the history of collective joy and participation in these sorts of communal festivals. I think, if things go well, we'll have a lot more time for those. So we might as well start brainstorming what our future holidays could look like.
Okay, concluding thoughts. This is happening to everyone all at once. It is not just education; it is everywhere. I've spoken to audiences of business leaders, investors, lawyers, application developers, software engineers—you name it.
The vibe in the room is basically the same everywhere we go: this is happening super fast, it seems super powerful, and we don't really know what to make of it. Are our jobs secure? Should we be using it? Should we be shutting it? Everybody's asking the same question.
For one thing, just know that you're in good company. Know, too, that next year it's going to be different again in meaningful ways. Know that there is no safe choice. Doing nothing or trying to pretend this isn't happening is not a good option.
Again, these binaries—all-in or total rejection and banning—neither one is the right answer. I think your most important tool, especially at the administrative level, is leadership and culture. I would want to see, and I would encourage everybody to get super hands-on themselves and to show off what it is you're doing. Show your own experiences. Champion specific people in your organization who have done a great job.
Really set the expectation that teachers and students are going to be learning together in this era. We're all on the same timeline with respect to AI. It doesn't matter what age we are. It doesn't matter how experienced we are. The AI release cycle is now dictating the timeline of how we're going to have to think about this more so than our individual experiences.
Teachers and students absolutely should be learning together much more than ever before. I don't think I'll surprise or shock you by saying you probably do have a lot to learn from your students.
On the left are my grandparents. Herman Leens went to Cass Tech. He was the first member of my family to graduate from Detroit Public Schools back in the day. He did not go to World War II because he had tuberculosis. But he still told us stories when we were kids about how we won the war.
He was an engineer. He designed machines. He worked at a factory. But the story he told the most was actually one of carpooling to work because gas was scarce. There were gas rations, and he remembered the route still when he was 80 or 90 years old. He would tell us, “I went to this person's house and picked them up, and then we went over to this person's house and picked them up.” And he would conclude, “And that's how we won the war.”
His brother was actually in the Pacific and fought for real in horrendous conditions. The point is, we all have a role to play, and I think we are entering a period that is going to require a whole-of-society mobilization where everybody, no matter what your role is—whether it's just carpooling to save gas so that it can be used for the broader effort, whether you're on some front lines, or whether you're like me and end up stumbling through all these important scenes as an extra—there really is no role that doesn't matter.
There is no cognitive profile that doesn't matter. You don't have to be super technical about this. I genuinely mean it when I say writing aspirational fiction might be one of the most powerful things you could do to shape the future, because positive visions are so scarce.
So take ownership at every level of your organization, whether it's the school board, the superintendent, the principals, the teachers, or the students themselves. Truly, everybody has a role to play in this. We absolutely need to have all of our best minds on it, because this is going to be almost for sure the most disruptive force that any of us have seen in our lifetimes.
While the challenge is super intense, I genuinely do think that you, as educators today, have the opportunity to be education's greatest generation.