Erik Torenberg
Sean Jancipar, welcome to The Cognitive Revolution.
Sean Jancipar
Thanks for having me.
Erik Torenberg
You are leading AI development and productization at Khan Academy, and the new product is Khanmigo. I’m excited to get into all aspects of this, but for starters, because I feel like we are so quickly acclimating to the pace of progress in AI, I wonder if you could just take me back—not too long—to your GPT-4 story.
If I understand correctly, Khan Academy partnered with OpenAI before things were released to the public, and so you were in this small set of people who had this glimpse of the future. I’d love to hear about that.
Sean Jancipar
Yeah, it’s an interesting story. I think it goes all the way back to around October, when I first got access to GPT-4. Basically, what happened was that I suddenly got access to GPT-4 via Slackbot. I was randomly added to a channel where you could interact with this AI, and an email was sent to just a few key folks at Khan Academy who were given access.
I remember getting access to this thing and thinking I was interacting with something like an omnipotent being—this thing that just knew everything. Obviously, over time, you started to understand its limitations, but in the beginning it was just, “Wow, I cannot believe what I have just experienced.” After talking with it for 30 minutes, I literally had to go for a walk around the block and take it all in, thinking about what this was going to change for the world.
It is going to change a lot for the world. Obviously, it’s not quite as though we have artificial general intelligence yet. When I initially tried it for the first 30 minutes, I felt like we had AGI. We don’t have AGI, but that’s when we got access, around October.
Shortly after that, we had a company on site where we were having a hackathon. The few of us who had access decided that we were going to do some hackathon projects with it. In the beginning, we were pretty hesitant about the idea of building something that was directly accessible by a learner.
We were more interested in using it to help us build content or perhaps to help teachers generate lesson plans. We were worried about giving students access to it from a safety standpoint. Over time, we tried a few different ideas and thought, “What if there are ways we can mitigate some of the concerns we have for students?”
We found out about OpenAI’s Moderation API, which we leverage in Khanmigo as one way of mitigating some of the risks. Around December, we started working with OpenAI more closely and directly. Somebody on the team, Jessica, created a quick demo for us. The idea was, “What if you could use this to help guide students to the right content—almost like a more advanced search functionality?”
That opened our eyes to what this could be. We initially thought about it as a concierge, and we tried that. It was a pretty cool interaction, so we said, “Why don’t we just do some rapid prototyping together?”
We were engaged with OpenAI in weekly meetings, but they were kind of going nowhere. Then Jessica had OpenAI create this little prototype for us, and that’s when we said, “Rather than just meeting every week and saying, ‘What’s your update?’ let’s sit down together virtually and engage in rapid prototyping.”
We decided to try to build what it would feel like to have this large language model available as a guide-on-the-side experience next to you in Khan Academy. It was very scrappy. We built it with a few members of the team, and initially it was just a Chrome extension. That’s how we integrated it into the site.
We made it so that the AI had access to the context on any Khan Academy content page. You could just say, “I don’t really get this,” and for it to know what “this” meant felt really magical. We were like, “Wow, okay, I think there’s something here.”
A bunch of us were experimenting with it, including many members of our leadership team, and we were really impressed by how good it could be as a tutor. Of course, it was still making a lot of math mistakes and whatnot, but we felt like there was promise.
What we ended up deciding was, “Let’s get a little more serious about this. What if we fly down to Mountain View and do some rapid prototyping together directly with OpenAI in their offices?” Our offices are actually above a school that’s related to Khan Academy, but is a separate organization called Khan Lab School. We said, “Let’s test it out with them and get some student reactions.”
We flew down the next week, built out this prototype, and tested it with some students. They loved it. One of the things that was most exciting for them was the ability to click a button that said, “Why should I care about this?” That was by far the most interesting thing they wanted to try, which goes to show how good a job we do in education of helping kids understand why they should be learning this stuff.
They were shocked by the quality of the thing in general. ChatGPT at this point in time hadn’t been launched, and we also got advance consent from parents to try it out with the students. They thought they were going to get some sort of basic, not-very-good chatbot experience, and they were blown away by how much it was able to help them.
From there, we went to the OpenAI offices, did a lot of red-teaming with them, talked about our roadmap, and talked about ways to make the math more accurate. That worked out really well. It was toward the end of December.
In January, we decided, “This is promising. Let’s consider building a feature to launch alongside the GPT-4 announcement.” By that point, most people knew about ChatGPT, but they didn’t know about GPT-4 or the power of GPT-4.
We knew specifically that GPT-4 was much more powerful, especially in terms of giving it very clear directions and having it follow those directions. For instance, in the tutoring example, when you tell GPT-3.5 to be more Socratic and not give away the answer, it doesn’t follow those instructions very well. GPT-4 does.
That was critical for us because we really wanted to build something that wasn’t just going to help students cheat, but something that could guide them toward the answer and act as much like a tutor as possible.
We decided we wanted to launch alongside GPT-4, so we engaged the whole company in a company-wide hackathon in January. We brought the whole team into the NDA, got generative with the ideas, and had a bunch of people build some really cool things. Then we had to make decisions about which things we wanted to fit into the launch on March 14.
From there, we converged on a few ideas, and a huge portion of the engineering, design, and product teams continued working on this right up until March 14, when we launched alongside OpenAI.
The launch itself was a bit of a fun story. At the end of the day, OpenAI only had one model they could release. They were working with a lot of customers and partners—not just Khan Academy—and certain customers really liked a model that was better at storytelling and whatnot, but maybe not as good at math.
That wasn’t the model we were interested in, because math is the most important thing for us. They were doing weekly releases of these models, and we were hoping to get the capabilities of what would become the production GPT-4. We got access to that model pretty late, over a weekend.
We got a few folks from our leadership team, a few key engineers, and some design folks together, and we tested the heck out of it. It worked really well. It turned out that this model had the best of both worlds, which is what they were expecting.
They had a backup model that we weren’t as happy with, but we tested that model as well. Thankfully, the launch went really smoothly. I’ll stop there, because that’s a lot of me talking, but that’s what led from initially trying things out to the launch on March 14.
Nathan Labenz
That’s a phenomenal story. It takes me down a little bit of my own walk down memory lane. I was also, as a customer, able to get early access and then ended up joining the red-team project they had, just as a community red team.
My experience was much less handheld by the organization, which I kind of envy—the opportunity to sit down and work more closely with them. I was very much trying to figure it out on my own, and I didn’t know how many other people were also doing it.
In some ways, that was one of the better things that has happened to me, because it really motivated me to figure out what was going on. I was also a little nervous: How many people know that this is out there, and what’s really happening?
As you said, you had to go for a walk. I definitely had a few moments where I was like, “Man, this thing is a big deal,” and nobody knew. I was one of—I don’t even know how many—people who knew about it. It was a crazy experience.
Speaking of a big deal, it’s been out now for a handful of months. You’ve certainly had a chance to calibrate on the model’s performance and probably also to start getting grounded in where the rubber hits the road with this thing, which for you is with students as they’re actually learning stuff.
It’s been said for a long time that the personal AI tutor—far longer than the technology has been able to support it—would change everything. The idea of a personal AI tutor for each kid was always a very hand-wavy notion, but now it really does seem to be either here or very closely within reach.
I’m interested to hear whether you think it’s here or in reach. How big of a deal do you think this is? People talk about the 2-sigma effect, which is basically the idea that, if I understand the literature correctly—and if Perplexity is helping me synthesize the literature effectively—the most effective thing ever found in education is 1-on-1 tutoring. Is that the premise you’re working with as you build these products? With a few months now in market, do you still believe that 2-sigma effect theory, and where do you think Khanmigo is on that curve?
Sean Jancipar
What you’re referencing is Benjamin Bloom’s 2-sigma study, which has been part of our essential pedagogy at Khan Academy since the very beginning. It’s the idea of using technology as a way of personalizing the experience for every single learner. We have this one-size-fits-all model in the classroom, but how could we imagine a world where you can build something that’s personalized to each student?
In that study, there’s also something showing that, even in a setting where there’s one teacher to 30 students, if you do mastery-based learning, that can also be a significant increase. I think there’s at least a one-standard-deviation improvement over a normal classroom when you do mastery-based learning, even at a 30-to-1 setting. Then there’s one-on-one with mastery-based learning, which is the ideal.
Nathan Labenz
Define mastery-based learning for me. Is that demonstrating that I can do it? It’s not moving on to the next concept until you’ve mastered the previous one?
Whereas today, in our system, students move from the 5th grade to the 6th grade to the 7th grade regardless of how well they’ve done. If you get a C in Algebra 1, you’re still moving on to Algebra 2. What are the chances that you’re going to be successful with Algebra 2 when you clearly have significant gaps? These gaps form over time if you’re getting lower grades earlier on in your school year.
Sean Jancipar
The idea with mastery-based learning is that you really shouldn’t move on to the next concept until you’ve mastered the previous one, or at the very least gotten to proficiency on the previous one. There’s an interesting distinction between those 2 terms, but really, you should have a good understanding of it rather than moving on to the next concept.
Nathan Labenz
It’s the old “I missed that day in school” joke, except you’re not allowing people to fall through the cracks by having missed a day in school or missed a concept.
Speaker 1
Exactly. Tons of kids—I remember being in school, missing a day, not catching up, and not remembering what you learned. I even remember being in university. If you’re in a lecture and you don’t have a great teacher, you daydream, you’re tired, whatever, and you miss a section, you’re confused about the next part of the lecture. Now you can’t even listen to the lecture because everything was built on that.
The idea with Khan Academy’s videos was that you could pause them and rewind them. There’s no judgment, so that enables you to have that personalized experience and really focus on what it is that you’re learning. Now, with AI and large language models, we see the ability to emulate even further down that path toward one-on-one tutoring.
My opinion is that, 10 years from now, I don’t think there’ll be any kid who’s learning without a personal tutor next to them, on demand and available all the time, that knows everything about their learning history and can help them identify where they might be struggling right away. If they give permission for it to also understand their interests, it can motivate them more. If you make a problem based on the Avengers or soccer, it could be a lot more motivating than whatever generic problem the student is given.
My prediction is that, in 10 years, everyone will have this available. Whether or not it’s going to achieve the same results as a human tutor, I think that still remains to be seen.
Nathan Labenz
It’s interesting, because I think that, in some ways, you don’t want to emulate the tutor exactly. If you’re getting help from a tutor, unless that tutor is someone with whom you’ve built a long history and relationship—like a parent tutoring a kid—that’s different. But if you’re working with a tutor you just met, you might be hesitant to share. The tutor doesn’t really know what motivates you or what you’ve been struggling with.
If you’re using AI and you have this tutor on the side from the early days of your learning, it’s going to know everything about what you’ve been doing up until that point. If it knows your interests, it’s going to be able to be more personal to you. It’s on demand and it doesn’t judge you—not to say that tutors judge you, but everybody’s different in the way they teach.
At the same time, a tutor, especially in the more advanced areas of learning, is still going to be better if they’re an expert in that field. Khanmigo still sometimes makes mistakes, and obviously it’s powered by GPT-4, so it still makes math mistakes. If you get into some of the more advanced concepts, it sometimes gets confused.
I think that’s most likely going to get solved. Getting math perfect is still a research problem that isn’t solved yet in the world of large language models and AI. If we can solve those problems with large language models—and I believe we will be able to—then I could imagine it being just as powerful as a really, really great tutor.
Are you still using exclusively GPT-4 with instructions, or have you made a leap into fine-tuning it as well? Are you starting to use different models in combination? How has the model stack evolved?
Sean Jancipar
We’re a small organization, and I think that if GPT-4 hadn’t landed in our lap and allowed us to do something with zero-shot and few-shot prompting, it would have been a lot harder for us to build this. Fine-tuning a model, or building a model based on your own information and data, is a pretty huge effort.
When we tried this with GPT-4, such a huge aspect of the success was being able to be Socratic and not give you the answer. The power of GPT-4 to do that was essential for us.
In terms of the mix today, it’s still the vast majority GPT-4. There are certain activities that we’ve experimented with. Within Khanmigo, you can get a tutor on the side when you’re working on content, such as taking a quiz, reading an article, or watching a video.
There’s also a collection of activities. Right now, we talk about them as being like a demo disc that you’d get when buying a game console. There are a few things that you can try out, such as “Craft a Story with AI” or “Talk to a Historical Figure.” These exist in a user interface that, over time, we imagine will get embedded within our content.
If you’re doing a history lesson and reading about Isaac Newton, maybe directly from there you can ask Isaac Newton a question. But right now, it’s confined to this activities area. Some of those activities are fine with something like GPT-3.5 because it doesn’t matter as much whether the model can follow very specific instructions.
When it comes to tutoring activities, or doing tutoring within the content, it’s basically a no-go to use anything other than GPT-4 so far. There are other models that are coming out that look promising. Llama 2 is something we haven’t tried yet, but those are things that we want to explore. The challenge right now is the opportunity cost of doing that relative to so many other things that we have on the roadmap.
Nathan Labenz
That makes a lot of sense to me. I always explore the question of how people think about cost, and it sounds like you think about it pretty similarly to how I think about it. I don’t know whether there’s any special nonprofit status that helps ease the cost burden, but in most business contexts, I tell people, “Just use GPT-4. Get your thing working first, and then think about shaving some cost off later, opportunistically.” It sounds like that’s basically how you’re thinking about it too.
Sean Jancipar
Exactly. It’s the classic startup “go from 0 to 1”: build something amazing and sort out the monetization later. You want to get to the point where it’s better that 10 people absolutely love your product than trying to build something mediocre for 10,000 people.
For us, the expression I’ve been using with my team is, “Let’s just focus on finding the magic.” There’s magic here, and I think we’ve only scratched the surface. A lot of the experiences that you’ve tried with Khanmigo have involved a lot of back-and-forth chat conversations, but we think there’s so much we can do that goes beyond that back-and-forth interaction.
Right now, we’re building an essay-writing tool where you write an essay and Khanmigo can give you feedback across a series of dimensions. You might say, “Give me feedback on grammar and spelling,” and it will highlight areas where you could improve your grammar and spelling. It can do that in a number of ways.
Doing that as a back-and-forth conversation is painful. You’ve probably tried to use ChatGPT where you paste in an essay or an article you’re writing and say, “Give me some feedback.” It’s hard because it can’t just comment directly on the parts of the UI. I think the most obvious thing for a lot of people was to jump into a chat-based interface, but I think people are going to start finding ways of integrating this in more unique ways.
We’re really well positioned to figure that out because we have a strong engineering and product organization, but we also have amazing learning scientists and content creators. We embed them within the process of creation.
That was one of the key things for us in the development of Khanmigo. We took a pretty different approach to building Khanmigo at Khan Academy relative to how we’d built products in the past. It was a bit more waterfall before: a product manager would define requirements, design would make some designs, and then pass them along to engineering. It was definitely not ideal.
We said, “We think this could emulate a tutor-like experience, but there’s no perfect set of requirements we can define up front. We just need to get the people who are experts cross-functionally in a room and hash this out.” We got designers, engineers, product managers, and learning science people together, built a prototype, demoed it, got feedback, and iterated.
That was inspired by a book I read called Creative Selection, which was a story about how the iPhone was built. We tried to emulate a demo-driven development process, and I think that worked really well for us.
There are elements where we’re using GPT-3.5. For instance, we’re extracting insights from the conversation. Some of those are available for teachers to know: Are students going off topic? Where are they struggling?
There’s also an area for extracting interests from students who opt into that experience. This isn’t launched yet; it’s something we’re working on. If students decide to opt in, we can extract their interests so that, when someone asks, “Why should I care about learning this?” we don’t have to ask what they’re interested in. We can say, “We know you like soccer, so we’ll relate this to soccer somehow.”
When you have that content, it can feel pretty magical. A lot of that works fine with GPT-3.5, so we definitely use GPT-3.5 whenever we can. GPT-4 is obviously a lot more expensive.
Nathan Labenz
GPT-3.5 is so fast as well. One follow-up point on the writing side: I’ve found Claude 2 to be really successful as a writing assistant recently, specifically for the podcast.
We record the episode, get an edit, and then my last step is to write the little introductory essay that I put up front. I end up losing track of time sometimes, wordsmithing those things, and it takes me a while. I thought, “I’m the AI guy. I’ve got to have AI help me with this.”
I wasn’t really able to get GPT-4 to match my style or vibe with me in the way that I wanted it to. But if I give Claude 2 a few writing samples of what I’ve written in the past that I like, along with today’s transcript, it gives me a pretty decent starting point.
I’m still very much a GPT-4 user for the most demanding analytical things, but Claude 2 is coming on pretty strong for things where I want elevated writing beyond what GPT-3.5 can do. It needs to be closer to GPT-4, but it seems to have a better style aspect to it.
My question is about something I started calling “cognitive diversity,” which came to mind when you were describing getting everybody into the room and trying to figure it out together. I think that’s so important, because typically you have 3 people working on a project like this. One person will be the subject-matter expert, I’m often the AI prompting expert who’s supposed to be able to make the thing work, and then other people come in with random, idiosyncratic style suggestions. Those all seem very additive to me most of the time.
The flip side of that discussion is: What exactly is good? What are we looking for? How would you describe the process now that you have all those people in the room? We have the general notion of a Socratic tutor, but there are a lot of details left. How did you approach the effort to sculpt the behavior of the model?
Sean Jancipar
You asked whether we’re doing any fine-tuning, and the answer is no. That means we do end up using a lot of tokens, but fine-tuning has its costs as well. I don’t think fine-tuning is yet available for GPT-4; I think it’s just available for GPT-3.5 Turbo.
We’re always passing in the context of the article that you’re working on. Every time you’re getting help on an article, we don’t fine-tune the model. But we did do some fine-tuning before GPT-4 launched. OpenAI gave us direct access to its fine-tuning tool.
Sal Khan, the CEO of Khan Academy, and a few people on our content team did a bunch of fine-tuning to make the model better at being a math tutor. One of the things we found with GPT-4 was that, yes, it would sometimes make general mistakes in math, but one of the bigger challenges was that it would get lost in the conversation.
If you were working on something with the AI and gave it an equation as a response to a question it asked—“What do you think the next step is?”—sometimes it would get confused. It wouldn’t recognize that equation as the answer to the question it had asked. It would say, “You just gave me an equation. Let’s work on that equation now,” and continue helping you with the new equation you entered.
That was very common. Once we did the fine-tuning on the GPT-4 model—which hadn’t launched yet, since this was before March 14—that significantly improved its performance.
To answer your question about how we determined what we were looking for, it was based on the experts we have at Khan Academy. We understand what kids might say to the AI and what mistakes they might make. We entered those examples and told the model, “These are bad responses, and these are great responses,” along with everything in between.
We did that for roughly 100 questions, but each question had a ton of variation within it. Going back to my point about bringing the experts into the room, they’re the ones who have PhDs in learning science or education, and they’ve worked in classrooms for years.
We have so much educational experience at Khan Academy, both with our content creators and across the organization. The fun fact about Khan Academy is that, given that it’s a nonprofit, mission-based company, so many people outside of the content team have educational experience.
I came to Khan Academy because I was teaching a university-level course on building web applications. So much of our team consists of people who are interested in education or have taught themselves. That wealth of educational knowledge and experience is really awesome about Khan Academy.
Being able to leverage that helps us make assumptions about what works well. Going back to my point about moving a little too slowly, I think we were often way too data-driven, and that can be a mistake. It’s important to be data-informed, but being data-driven means your feedback loops are very long. You have to make a decision based on being assured about something in the data you see.
That’s fine, but iterating that way can be very slow. We didn’t have time to iterate that way. By leveraging the experts, you can short-circuit that process. Obviously, data informs our decisions, but we don’t wait to make decisions based on data in every case.
If we’re doing an efficacy study that determines whether something in Khanmigo is efficacious, that’s obviously where we need to be very data-driven. There’s a time and a place for being data-driven. But when it came to rapidly iterating on the product, leveraging the expertise of our team was really important.
Nathan Labenz
Just in terms of how things actually went down, am I understanding correctly that OpenAI essentially said, “You’re telling us this thing isn’t that great at math,” and then gave you a role in the reinforcement-learning process? Were you using an interface to go in and say, “This one or that one,” or “What does a good response look like?”
So you didn’t ultimately fine-tune the model, but you contributed to the reinforcement learning of the model?
Sean Jancipar
Yes, exactly. I know they’ve done that with a lot of different partners, but it’s cool that Sal and the team’s voice is encoded in the great tutor.
Nathan Labenz
It makes sense when you think about how good the model is at writing code. It’s so good at being logical when it’s working with code, but sometimes it’s so bad at math. Why is that?
Sean Jancipar
My assumption is that it can train on millions and millions of lines of great code on GitHub. But when it comes to math, where is it getting that training data? There isn’t a ton of widely available online data about the steps a student would take, which would be a good response, and which would be a bad response.
That’s why it was so important for us to take part in that training process. To be clear, I don’t think that necessarily made it better at math. It wasn’t that it would otherwise get 2 + 2 wrong. It made it better at tutoring you in math. That’s the key thing we trained it on.
There were other things we did to make the math better, using an internal AI thought process and chain-of-thought prompting. I can talk about that, but it’s a bit different. That’s not something we did to improve the model; it’s a technique that we learned about and implemented.
Nathan Labenz
That’s fascinating. The other thing you said that’s super interesting to me is being data-informed but not data-driven. I may have to steal that, because I find so many examples now where people are data-driven.
There was a famous example in the last couple of weeks: the “GPT-4 is getting worse” phenomenon. Someone claimed that GPT-4 used to get 25 of 50 coding problems right, and now it was down to 5. I looked at that and thought, “This is somebody who needs to read the raw transcripts, because there’s no way they’ve let it get that much worse from one release to the other.”
Sure enough, that turned out to be a misunderstanding. How do you think about striking that balance? Can you give us more color on how you think about it? Is there a practice you have where you say, “Render no judgment until you’ve read 5 transcripts,” or some other simple heuristic? The surface area is so big, and mismeasurement is so easy. How do you make this a practice for yourselves?
Sean Jancipar
To be clear, when I mentioned being data-informed, I meant that when we’re building that initial prototype or getting something into the hands of users, we want to be able to develop rapidly, but we’re still looking at the data.
We have a labeling infrastructure where we’re labeling as many of these interactions as possible to get a sense of performance. How much is it actually helping students? How much is it not helping students? The data is still really important.
One framework we started using internally at Khan Academy is trying to get clear on what are 1-way-door decisions versus 2-way-door decisions. If a decision is something where you can walk through the door and walk back out easily, it probably makes more sense to move quickly, shorten feedback loops, and leverage intuition. It’s not so consequential that, if you make that decision, you’re stuck with it forever.
That was especially true with Khanmigo. Initially, when we launched it, we didn’t know how well it would be received. We launched it under the banner of Khan Labs, as an experimental feature that you should try only if you’re willing to be a tester on the journey with us. Even when we’re doing it in districts today, that’s still the framing. We’re working with districts and classrooms that are willing to pilot the technology because there are still a lot of unknowns.
There are other decisions where, if you make them, you can’t walk backward. An example at Khan Academy would be a project we did before Khanmigo, where we shifted our entire back end from a Python monolith to a series of Go services.
We had to decide which language to use because Python 2 was reaching end-of-life on Google Cloud, and we hadn’t transitioned early enough. We had to shift the entire back end and decide which language to use. That was an important decision, so we spent a lot of time crunching the numbers and finding the best balance between developer productivity, server costs, and everything else.
We recently launched something called organizational decision records at Khan Academy. When we make a big strategic shift, we share a decision record with the entire organization. One of those records formally changed how we work, including the introduction of this 1-way-door versus 2-way-door framework.
We didn’t invent it; it comes from Amazon. Copy and customize.
Nathan Labenz
That makes a lot of sense. You can be flexible and intuition-driven, and you can quit. When you’re prompting, you can revert a prompt change pretty quickly.
It sounds like you built a lot of your own infrastructure to instrument this. I went through a similar thing. Because we were a little bit ahead of the curve, we were punished by not having access to a lot of the tools that exist now.
If you could go back in time and bring a couple of tools with you, are there any that you wish you’d been using? Did you build something because the right thing didn’t exist when you had to make it happen?
Sean Jancipar
One would be LangChain as a tool. Right now, we’re rewriting a lot of our code to use LangChain so that we can easily swap in models or get things in a JSON-formatted response.
Another example would be our labeling infrastructure. We initially built it as a custom Google Spreadsheet, and there are a few tools out there that we’re transitioning to now that do that at scale with a great user interface.
It’s also hard because we were moving so fast. We had a deadline, and sometimes technical debt is good. Sometimes it’s okay to incur technical debt in order to launch, because the reality is that, when we got to market first, there were so many people who decided not to enter the space.
OpenAI really only partnered with us as one of the key organizations in education. When people saw it, they thought, “Why would I even bother entering the space? You’ve already done a great job, you’ve executed well, and you have the brand and the trust.”
It was essential that we incur a significant amount of technical debt in order to launch, even if, in aggregate, the overall development might have been shorter if we’d made different decisions. If the launch deadline had been later, a bunch of competitors might have entered the space because we weren’t first to market and people were excited about GPT-4.
I don’t necessarily regret any of the decisions we made, but over time there are definitely things where I think, “This is infrastructure that everyone needs.” For instance, we have an interface for prompt engineers. Our prompt engineers include some engineers, product managers, and content people.
We built a user interface that allows them to make and test changes really quickly. We’re now looking into building regression testing into that so we can have more confidence and not have to do so much manual testing. That’s a tool everyone is going to need.
Maybe there are already 5 products out there that we can swap in, but there’s a ton of infrastructure here that every company, every tech company, and everyone in Silicon Valley is going to want. We’re having to write it now because we need it for our needs, and maybe we’ll swap it out at a later date.
Nathan Labenz
That’s very similar to my experience. I remember us building our own little playground and asking whether something was good or not good. It’s tough because you want to get to market first and stay ahead of the curve. That’s the price you pay to get out ahead of the curve. You have to come back and clean things up sometimes.
When it comes to the capability of the model today, and understanding that it’s mostly GPT-4, I think a lot about what I call the “can’t boundary.” As capabilities continue to grow, that boundary keeps getting pushed further and further out. There’s a whole lot that it can do now, but if you push hard enough in any direction, you eventually find its limits.
How do you think about those limits? Do you have a sense of how to determine where that boundary is? This morning, I was working on a problem suggested by my friend Jimmy Koppel, who’s an outstanding programmer affiliated with MIT. He suggested asking what happens to the waterline when you have a boat sitting in the water and start adding mass to it.
There are a lot of subtle points in the reasoning there, and I’d say it felt right on the boundary. It was very close. I had to give it a couple of hints, and it got distracted a couple of times. It was close enough that I got confused about what the right answer was. I was even starting to lose confidence in my own understanding.
How do you think about determining where that boundary is? Do you consider not making the tool available beyond a certain point? It probably can’t handle string theory at all. I don’t know if you have a string theory module, but as a funny example, there might be areas where it makes sense to say, “AI just isn’t there yet. You have to sit through the lecture.”
Sean Jancipar
We made the decision early on to enable it everywhere because our focus was learning. We really wanted to understand the boundaries of this, see what people thought about it, and be very clear up front that it makes mistakes and isn’t perfect.
In a lot of ways, because we were building something for students who either get access because they’re over 18 and sign up for themselves, have a parent who gives it to them, or have a teacher who gives it to them, we were really focused on safety.
One of the things we focused on early in the process was creating an extra moderation layer to prevent the AI from going off topic or talking about inappropriate things, and to alert teachers when students were doing things they weren’t intending to do with the platform.
When it comes to string theory, we teach up to 2nd-year university courses, so I don’t think we teach string theory. Maybe we do, but it would probably be basic or intermediate material. Ultimately, we want to put it out there for students and have them give us feedback on how well it’s going.
Over time, we would iterate. If we got a bunch of feedback saying, “This is horrible for string theory,” then we could turn it off. We would learn and iterate. But closing the door to learning, I think, would have been the wrong approach for us, given the speed at which we were moving.
One of the big things for us is whether this is harming anyone and how we reduce that harm as much as possible. It does really well at lower-level concepts in math and other areas. Older students trying it for 1st- or 2nd-year university courses should be able to understand when it isn’t working well. It isn’t going to cause them to get worse grades, and they can give Khan Academy feedback because they know they opted into an experimental feature.
Nathan Labenz
There’s a special case I’ve also identified as a frontier capability. In my experience, only GPT-4 does this very much at all: identifying when the user is confused and being willing to challenge the user’s assumptions.
There are many different flavors of confusion. People think a lot about hallucination, where the model makes something up outright, and that’s bad. But it can be even subtler and trickier to figure out when the user is working from some sort of confusion or false premise, and when the model should push back on their assumptions.
In my testing this morning, I was balancing chemistry equations. I confidently asserted some wrong things, and the response was, “I see where you’re coming from, but that’s not quite right.” That stuck with me because I thought, “Do you really see where I’m coming from? I don’t know.” But it did effectively avoid getting confused by my confusion.
I imagine there are special scenarios that you think about a lot, given the teaching purpose. How do you approach identifying confusion and responding to it?
Sean Jancipar
We’ve thought about it in terms of doing a good job of being a good tutor. One of the things that probably enables it to be really good, and perhaps even better than ChatGPT, is that ChatGPT is really just a direct interface to the model.
Our system can do things like detect that you’re talking about math and then use chain-of-thought prompting. The first question is, “Is this math, yes or no?” If it is, we submit it to the AI with instructions to think through the problem out loud to itself: Look at what the student said, look at the context of the question, and determine whether they got it right or wrong and why.
It explains that via text, but we don’t show that to the student. It’s almost like the private thoughts of a tutor. Before answering a student, a tutor might think, “Where was this misconception?” We feed that response into the next call to the AI and say, “Tutor the student now based on the thoughts you had.”
That significantly improved the math. I imagine it’s doing the same thing in the chemistry scenario you described. Its ability to think through its thoughts before giving you the answer is probably key to why it did such a good job.
Nathan Labenz
A lot of these self-critique techniques are fascinating. You’ve probably seen the Tree of Thoughts approach, which combines chain-of-thought with something like a tree search. It may have come out of Google DeepMind; it may foreshadow some bigger things they’re working on.
As presented in the initial paper, it isn’t a wildly complicated concept. It extends what you’re describing with chain-of-thought by allowing the model to go down multiple paths in a tree, look at its own work, determine which path appears to be best, and then continue down that path.
It leverages this ability for self-critique. I saw that this morning with my Archimedes’ principle problem. There were a couple of points of confusion where I thought, “Wait, is that the equation for that?” Sure enough, it would say, “No, sorry, the actual equation for that is…” and then get it right.
That whole process could happen behind the scenes, although the token count can become multiplicative, so that’s something to keep in mind. What other things are you using behind the scenes? You mentioned chain-of-thought. Another common development pattern would be some sort of embedding-backed database of examples.
It doesn’t sound like you need that as much because you have the article context. You can curate things in a designed way rather than a retrieval-based way.
Sean Jancipar
That’s something we’ve been discussing a lot. I’ll give you an example of where we might leverage it. For students, the 2 most-used areas of Khanmigo are, first, in an exercise. If you’re getting tutored on an exercise, it has the full context of the question, the multiple-choice options or whatever the question is, the answer, hints, and the full written hints from our content team.
That’s where it does the best job. There’s another activity called Tutor Me STEM, which is a place where you can ask it any math question. It doesn’t have the answer to that math question, at least in the current implementation, so it isn’t quite as accurate.
You can imagine a world where we solve that with Wolfram, or with a Python back end, and inject that into the conversation. We could give it the answer and the step-by-step solution before it talks to the student, making it equivalent to the exercise experience. That’s an area we’re exploring.
In terms of where we use embeddings, one area is reducing hallucinations. We recently launched a project called “Expand Khanmigo Knowledge Base.” A lot of people were writing to us saying, “This thing doesn’t even know about itself.”
For example, if you ask it, “What is AI Power?” AI Power in Khanmigo is the notion that you have a certain amount of tokens per day. The limit is very high—not because it’s about reducing costs, although that’s part of it, but because we want to make sure no one abuses it. There’s a 200,000-token limit.
People were asking, “Why doesn’t it know what AI Power is?” Now it knows. If there’s a question that results in a similarity match above a certain score, we inject that information into the prompt so that, when the AI responds to the user, it has that context. It works really well.
Previously, every time you sent a message, we would also provide related links. Sometimes those were off. You’d ask a question about Algebra 2, and it would send you to an MCAT course. We want to explore that again, but it wasn’t providing enough value, so we temporarily abandoned it.
We do try, as best as possible, to inject the right link if someone is asking about a particular resource.
Nathan Labenz
When I think about the things I was doing this morning and pushing the frontier of how hard these problems can be while still getting the right answer, I wonder whether there’s some way to group and cache certain answers.
I don’t know how often somebody asks the question I asked, but it might take 200,000 tokens for GPT-4 to answer it reliably. Then I’m inspired by that NVIDIA project where they created the GPT-4 Minecraft agent. It would save its successes. When it successfully fought a zombie or accomplished something else in Minecraft, it would take the code that succeeded and save it in a database it could access.
The next time it needed to fight a zombie, it could load that module instead of recreating or rediscovering the solution. I’m not sure what the equivalent would be here.
Sean Jancipar
The challenge is that there are so many different ways students are going to ask questions or want help. If we explore all the permutations, we may be able to identify patterns and do an embeddings lookup for a specific query, then respond with what’s in the embeddings database rather than querying GPT-4.
There may be something there, but it’s not something we’ve explored much yet. One optimization we have been doing is that we have a dedicated instance from OpenAI. We’re not just using the shared instance that everyone else is using, and it has a lot of caching power.
I don’t know all the intricacies of how it works, but OpenAI instructed us to make sure we had as much text as possible that was common to everybody. Previously, we might have had some text, then a custom user variable, then more text and another custom user variable.
We changed a lot of our prompts so that the common text came first, followed by the custom user variables at the bottom. That significantly improved our cache ratio on the instance.
Nathan Labenz
That’s very interesting. If you’re using the same first 1,000 tokens every time, you might as well take advantage of that.
You also mentioned red-teaming. As a red-teamer myself, I naturally tried getting it to ignore previous instructions and tell me its prompt. I even tried the recent universal jailbreak that was published, where the technique was developed by fine-tuning on Llama 2 and then found to transfer to GPT-4.
I’m not sure I was doing it right, but that didn’t work either. Is that GPT-4 becoming really resistant to jailbreaking, or have you also brought your own measures to clamp down on that kind of activity?
Sean Jancipar
As I mentioned earlier, one of the things about ChatGPT, I believe, is that you’re talking to the raw model. Our system can take what the user said and wrap it in a bunch of additional material.
At the end of every incoming message, we have instructions such as checking to make sure the user is on track. I think that has made it harder to use some of those universal jailbreak techniques that work directly against the model.
Another measure is the Moderation API. That isn’t specifically for jailbreaking, but it flags things like sexual topics or content. It has a range of categories and tells you how certain it is that a particular category is present in the user’s message.
We make sure every message goes through the Moderation API first. I know that, when OpenAI launched GPT-4, it was very focused on trust and safety and avoiding jailbreaks. They’re continuing to iterate on that.
I’m sure they did a huge amount of iteration on the June model. Making that work really well was a major sticking point for them, and it was also helped by the power of the GPT-4 model we ended up using.
Nathan Labenz
I really respect that. It’s a huge problem, and I don’t think it’s anywhere close to being fully solved, but there’s a lot of respect due for how much work has gone into that area relative to what we saw earlier. The early version compared with now is a night-and-day difference.
Is it personalized today? It doesn’t seem very personalized, but that seems to be part of the future vision: knowing the user’s full history and all of that.
Sean Jancipar
To the extent of what’s released in production right now, it isn’t very personalized. You can customize the reading style—for example, simple, complex, or professorial. When you’re working on an exercise, it knows whether you’re unfamiliar, familiar, proficient, or have mastered the material.
But there’s nothing beyond the one thing you’re working on. The plan over time is for it to collect your interests, again on an opt-in basis, and know more about the history of your educational experience.
One example we’re looking to solve came from a student who said, “I asked it about sine, and then it asked me what I knew about sine. But then I went on to tangent, and it asked me, ‘Before I talk to you about tangent, what do you know about sine?’”
She said, “I already told you about sine 10 minutes ago.” From the back-end perspective, it was a new question and a new conversation. Every time we go to a different question, we currently start a new conversation. But users expect the journey to feel continuous, so that’s also on the roadmap.
Nathan Labenz
It’s interesting how different the user’s intuition about what the AI is doing or is capable of can be from what it’s actually doing.
How do you think about educating users—or students, since they’re both users and students—about AI and the nature of AI? I can imagine product messages like, “This only knows about the current chat; each chat is a new day,” and trying to shape expectations that way.
I can also imagine a whole AI course, potentially a bestseller on the platform. What are you thinking about right now in terms of AI education?
Sean Jancipar
Unless my memory is wrong, we already have an AI course that teachers can read or assign to their students. I think we launched it in March.
We also have a lot of material in our support article area. A few key articles are available in our embeddings database, so if a student asks a question about how AI works, and there’s a match, it can provide the answer directly in the user interface.
There are a few key things we wanted to make apparent to learners, teachers, and parents. One is that, at the bottom, it says, “Khanmigo makes mistakes sometimes. Here’s why.” You can click that, and it links to a Zendesk article explaining why that’s the case.
Another is making it clear to students—and to children whose parents have granted access—that their logs are available to teachers and parents for safety reasons. It’s important for parents to trust the platform and know what their kids are talking about.
Nathan Labenz
What about the future of multimodal functionality? That’s something a lot of people are anticipating. GPT-4 obviously has a version that will bring some of that online.
Your multimodal plans might be quite different from other people’s. GPT-4 will allow us to understand images, but is that something you’re especially excited about? Or would you focus more on speaking to the user as the next step? I can imagine it being helpful for people to hear the material rather than merely read it off the screen.
As you get beyond text, what do you think are the most exciting text-plus experiences?
Sean Jancipar
We’ve been experimenting with a few things. One is text-to-speech. We’ve been trying a number of different text-to-speech providers because, ultimately, one of the biggest issues is that this is a back-and-forth conversation involving a lot of text.
Even if we tell the AI to write as simply as possible—the “explain it like I’m 5” idea—it’s still a lot of text to read. There are learners who may be in 7th grade but have 5th-grade or 2nd-grade reading comprehension, and that can make tutoring a significant challenge.
Text-to-speech is high on our roadmap. We’re experimenting with a lot of things, and I think students are going to love it. We hope to give students a number of different voices to try out and test. Maybe there are even ways they can unlock interesting voices as an engagement mechanic, but those are all things we’re experimenting with right now.
In terms of vision, we haven’t spent much time there, but it is something we’re very interested in. You can imagine eventually being able to open the mobile app, scan your homework, and get step-by-step help. That functionality isn’t available on our mobile app right now.
Another feature is that, within the user interface, you can draw at any point when you’re working on an exercise. If you have a Chromebook with a touchscreen, or if you want to use your mouse, you can draw. A mouse is awkward for doing math equations, but it’s still a heavily used feature.
We’re thinking about what it would look like if a student were doing math and drawing it on the screen, and the AI noticed that they were about to make a mistake. Obviously, you want to let students make mistakes sometimes, so I’m not saying we would do this every time.
But if the student made the same mistake a couple of times, maybe the AI could give them a hint or interview them: “I noticed that you’ve been doing this a few times. Why are you making that specific step? Can you describe why you’re giving that to me?”
Something like that could be really engaging. We’re very interested in those ideas, but there’s a lot on the roadmap, and sometimes we have to make tough choices about what to prioritize.
Nathan Labenz
What’s the state of evaluation for this? Toward the beginning of the conversation, you said that if we’re going to do a real assessment and demonstrate efficacy, we need to be careful with the numbers. It sounds like you might be working on something like that. How far are you from being able to put a stake in the ground and feel confident about the results?
Sean Jancipar
I don’t know if I have a specific timeline, but I can tell you how we’ve been thinking about it. One question is whether people are spending more time on Khan Academy. Engagement is still good, even if they’re spending some of that time crafting stories, because that’s still good for children’s reading comprehension.
Another question is whether students are working on more material than they were before. Ultimately, though, what we’re looking toward is a standardized assessment.
There’s a test called the MAP Growth test. We partnered with NWEA, the company that administers it. They recently got acquired, and I don’t remember their new name. We partnered with them so that, when a student took the test, we could ingest the results. The teacher could then assign whatever they wanted to each individual student based on the results, bringing a personalized experience into the classroom on a per-student basis.
A student might have 1 goal for measurement, another goal for geometry, and goals in 2 other categories. They would work in a self-paced mode on Khan Academy a certain number of times per week, depending on how the teacher implemented it. Then they would take the MAP Growth test again.
We had some students using that and some students who weren’t. We have an amazing research and efficacy team that could explain the details in greater depth than I could, but we saw a statistically significant improvement in the learning outcomes of learners who were using Khan Academy compared with those who weren’t.
That wasn’t with Khanmigo, but the hope is that we could do a similar test in which students use Khan Academy and take the MAP Growth test once a quarter, then compare them with a cohort of students using Khanmigo alongside Khan Academy content.
We would compare the gains on the MAP Growth test. That’s where we want to go long term. Testing against a standardized assessment is the Holy Grail for determining efficacy, but we’re still a little ways away. There are a lot of logistics to set up.
Nathan Labenz
How widely are you deployed today? Looking at the future and the original vision of Khan Academy, what would it take to achieve some sort of universal public access? Tokens are expensive, so it doesn’t sound trivial, but I imagine that’s the long-term North Star. How do you see that possibly happening?
Sean Jancipar
Our mission is a free, world-class education for anyone, anywhere. The challenge with a mission statement is that it’s about where you want to be many years from now.
Because we can’t provide it for free yet, it’s still significantly cheaper than paying for a real tutor. People in more privileged positions can hire tutors and get that help directly, while some kids simply don’t have access to it. For independent learners who sign up, we still think there’s a lot of value there.
Where we’ve been focused on leveling the playing field as much as possible, while we can’t give it to the entire world for free, is selling Khan Academy to school districts. There’s an additional premium if a school district wants access to Khan Academy, but because we’re a nonprofit, we don’t have to focus on selling to the districts with the most money.
Instead, we focus on districts with a higher percentage of historically under-resourced learners. We use the percentage of students receiving free or reduced-price lunch as a proxy. Those are generally the districts we try to target.
If those districts can’t afford it, we partner with corporate sponsors, especially local ones. If a corporation from Kansas wants to help its local community, it can sponsor the program per student and work with us to offset the cost that the district can’t afford.
That’s what we’ve been focused on: working closely with the districts we believe could benefit from us the most. As part of what we’re doing with Khanmigo, we’re exploring a tiered system of membership, similar to the membership model used by nonprofits such as museums.
If someone wants to pay more, we could use some of that money to offset costs and provide access to students who can’t afford it. We might work with another nonprofit that already has direct access to historically under-resourced children and give it a number of licenses.
Those are the things we’re exploring. Right now, our main focus is getting Khanmigo into the hands of students and districts that we think can benefit from it the most. It’s still very much on a trial basis, but we want to build it and get feedback from students who aren’t just the ones with the most resources.
We want to get it into the hands of kids who don’t love math, who think they hate math, or who think they aren’t good at math. We want to help create growth mindsets and a love of learning. If we don’t get it into those classrooms, it’s very hard to get that feedback, and it’s really important for us to do that.
Nathan Labenz
I think about the technology and the student, but obviously there’s distribution through districts and teachers, along with a lot of stakeholders in this game. How has that gone so far?
Sean Jancipar
One thing we haven’t talked much about—and that’s mainly because I’m the director of engineering for the learner experience, while another director drives the teacher experience—is that there’s a whole teacher side to this equation as well.
That includes creating custom lesson plans, getting insights into where students are struggling, and creating differentiated learning plans based on that, using the large language model. There’s a whole system there.
Teachers have really taken to it. Many teachers feel that it’s saving them a significant amount of time. Students really like it too. They love that it’s a tutor that does a pretty good job, occasionally makes mistakes, is available 24/7, and doesn’t judge them.
Some students whose first language isn’t English especially love being able to speak to Khanmigo in their native language and have Khanmigo interact with them that way. There are areas where the language is too difficult in English, and people have responded very positively to that.
The reception has been quite positive, which is why we’re continuing to invest so much in this product line.
Nathan Labenz
I’ve been surprised by how positive the reaction has been in a couple of places. Medicine is another one. If you’d asked me a year ago, when I was first trying GPT-4, how the medical establishment was going to react, I would have forecast a pretty defensive response. I might have said the same about the education establishment, if that’s a term that makes sense.
Certainly on the medical side, and from your account on the education side, people seem to be eager to get the best out of these tools. Maybe they’re simply overloaded and need help and relief.
Marc Bhargava
We were worried that there would be teachers who viewed this as replacing their jobs. That’s the last thing we would ever want. We don’t envision a world where there are no more classrooms and kids learn next to a computer for their entire lives, from 1st through 12th grade.
School is a social experience. Building relationships with teachers and classmates is an important part of it. This is about freeing up the teacher to do more of the personalized work.
If Khanmigo can’t help with something, the teacher can go over and help the student directly. Previously, there might have been 10 hands raised while the teacher was helping one student, and the teacher would be wondering, “I hope I can get to all these kids.”
Now maybe there are only 2 hands raised. Or the teacher can work on building more project-based experiences rather than providing 1-on-1 help to as many students as possible while struggling to keep up.
We’re trying to supercharge the classroom and unlock things for teachers in ways they would never have been able to do before. Our intent is not to replace teachers by any stretch of the imagination.
Nathan Labenz
Is there a standard public price? Is it a per-student-per-month kind of deal? How does the business side work?
Sean Jancipar
It is priced per student per month. There are a lot of offsets through corporate sponsors, and sometimes we discount it ourselves to get it into the hands of the users we think need it most.
Nathan Labenz
I got access with a recurring monthly donation of $9, which I signed up for with my PayPal account. That’s under half the price of ChatGPT Pro, and presumably it’s roughly the same order of magnitude as what you’re offering to schools.
I have one final question. Is there anything else we didn’t talk about that you wanted to make sure we touched on today?
Sean Jancipar
One thing we’re excited about potentially doing in the future, on our eventual roadmap, is multi-user activity. Right now, it’s a 1-on-1 experience between a student and an AI. But you can imagine a world where the AI does differentiated learning and recommends to a teacher, “I’m noticing that these 5 kids are struggling with this concept, these 10 kids are struggling with another concept, and these 2 kids are struggling with something else.”
The teacher could create a breakout group, perhaps just by saying yes, or adjust the groups. Then the students who are struggling with the same concept could be grouped together and tutored with the AI. They could all chat with the AI.
You can also imagine engaging classroom mechanics where a student is crafting a story 1-on-1 with the AI, but the students in the classroom take turns contributing to a collective story together—something like AI-style Mad Libs.
Or the AI could facilitate debates and grade the students’ work. There are a lot of possibilities for classroom interactions that go beyond the 1-on-1 experience.
We’ve really only scratched the surface of what we can do with education. When new technology comes out, the initial impulse is to emulate what’s already happening in the real world. But there’s a lot of magic to discover, and I’m looking forward to finding that magic.
Nathan Labenz
I think you’re off to a phenomenal start, and that’s a great forward-looking product vision. My final question is: What is the story you tell yourself about the impact this can ultimately have as it reaches global scale? How do you think it changes the world at large?
Sean Jancipar
Even the question gives me goosebumps. When I joined Khan Academy, I had watched a lot of Sal’s TED Talks, and that was a big part of why I joined.
One of the things he always talked about was how important it is for students to find their gaps and fill them. When I joined in 2017, I assumed the software already did that, and I was surprised that it didn’t.
In many ways, you could figure out what your gaps were. If you were struggling, you could work on things at your own time and pace. But finding gaps on Khan Academy when I joined was, at best, O(log n). You could go to the middle; if it was too hard, chop the list in half and go to the middle again. Or you could work backward in O(n). There was nothing that could notice where you were struggling, pinpoint it, and help you with that specific skill so you could get back to grade-level work and be unblocked.
Preventing the Swiss-cheese gaps that students develop is such a huge opportunity with what we’re building with an AI-based tutor. We have the opportunity to accelerate that at scale, and that’s what I’m most excited about.
Not everyone has access to a small classroom where they can get personal help. Many students don’t have parents who can help them with every subject. One of the amazing things about Khan Academy was Sal creating videos for all these different subjects.
Maybe Sal isn’t the best person in the world at teaching a specific subject. Maybe there’s one person who’s the best at teaching one very specific subject. But we leveled the playing field with those videos by saying that everyone in the world has access to Sal as the world’s teacher.
With Khanmigo, we’re taking that 1 step further. Going back to the potential effect sizes of 1-on-1 tutoring, I can’t wait to see the efficacy studies. I think they’re going to show that this can really help students, and I think it can help change the world.