Andrew Gordon
Formula One cars are the absolute pinnacle of engineering, right? Everything is perfect. They have a huge top speed. If you used one as your daily commuting car, you’d have an absolute nightmare, right? I think the same can be said for these models, right? A model that is incredibly good on Humanity’s Last Exam or MMLU might be an absolute nightmare to use day to day.
Most reporting on benchmarks is done on technical benchmarks these days, right? That is where you get the model, give it a set of evaluations—maybe on one theme or maybe an exam—and then get a score. Humans aren’t really involved in that loop.
My name is Andrew Gordon. I’m a staff researcher in behavioral science at Prolific, so I work on the sciences team, tackling questions related to humans in research, specifically online research.
Nora Petrova
My name is Nora Petrova. I’m an AI researcher at Prolific. I’m tackling questions around how we include humans in the development and evaluation of AI models, how we align them with human values, and how we fully understand what they’re capable of.
Andrew Gordon
How helpful do people find models? What’s the communication like? How adaptive do they find them? What do they think of the model’s personality? When you get ratings on those kinds of factors, what you actually get is an actionable set of results that say, “Okay, your model is struggling with trust,” or, “Your model is struggling with personality.”
We had a first go at this with what’s called the Prolific User Experience Leaderboard. That was a proof of concept with 500 participants from the US, a representative set of participants who evaluated a single model at a time and gave their feedback in a Likert scale format. How helpful did you find the model, from 1 to 7, and so on?
We’ve now taken the learnings from that initial leaderboard and built them into what we’re calling HUMANE, which is our main leaderboard. It uses a similar approach to Chatbot Arena, in the sense that we have comparative battles between models, which allows us to much more clearly differentiate which model is performing better.
Tim Scarfe
What if we could have a fairer approach where we actually diversely sampled and stratified folks based on how old they are, where they live, and what their values are? What would be a fairer approach to understanding the behavior of models and conducting evaluation?
This is what Andrew and Nora have been doing at Prolific. How do we know whether these models are actually good for the humans who are using them? How can we make evaluation metrics fairer? This is Andrew and Nora.
1. Benchmarking Misses Human Experience
Andrew Gordon
The problem is that, at the moment, the field of evaluating and benchmarking these models is incredibly nascent, right? It’s only been around as long as LLMs have been around, in the last couple of years. Because of that, it’s a fractured field. There’s no standard playing field for how these labs report data on benchmarking.
Some may emphasize, like with Grok 4 recently, a huge amount of emphasis on Humanity’s Last Exam and less so on other benchmarks. Some models will come out without any benchmarking data at all. Essentially, there’s a lot of heterogeneity in how these labs report results, and it leads to a situation where I think we’re at risk of struggling to compare the models on an even playing field.
There are, of course, bigger questions as well. Models are often lauded for getting the highest score on Humanity’s Last Exam, so we know from a technical perspective that the model has advanced above its rivals. But my core argument in this space is that if you just rely on those technical metrics, you miss half the point. These models are designed for humans to use. At the end of the day, most of the users are humans, and simple performance on these exams doesn’t necessarily correlate to a good user experience.
All the frontier labs really need to start having human-preference leaderboards more front of mind, alongside all the technical metrics.
2. Safety Needs Its Own Metric
Nora Petrova
People are increasingly using these models for very sensitive topics and questions—for mental health and for how they should navigate problems in their lives—and there is no oversight on that. In any other area where these topics are discussed, there is a lot of regulation and ethical conduct built into it.
Here, it’s the Wild West at the moment. Some companies are taking it more seriously than others and trying to study the ways in which humans are using the models for more personal topics and problems. We’ve seen some pretty startling examples recently with Grok 3 and Mecha Hitler, and it does raise questions about how thin a veneer the safety training is on top of some of these models.
Andrew Gordon
There is no leaderboard for safety, right? There’s no metric. We don’t grade LLMs by how safe they are. In fact, safety isn’t really even part of the question for some researchers. I would argue that it should be just as important as how fast or smart the model is. How safe is it for people to use?
Nora Petrova
There’s been a lot of interesting research coming from Anthropic in that direction, with regards to safety and alignment of the models, using Constitutional AI and various approaches that they’ve explored. There’s also been research around mechanistic interpretability, peering behind the curtains of the models and understanding how an input produces a certain output: which features and concepts, and which circuits, get activated along the way.
It’s about tracing the thoughts, essentially, of these models and trying to isolate where potential problems may emerge. Work of this kind is very important in raising confidence that these models will be able to handle novel situations in safe ways.
3. The Leaderboard Illusion
Andrew Gordon
The one other thing I’d say is that, given Chatbot Arena is the No. 1—and frankly, pretty much the only—human-preference leaderboard out there for LLMs, it’s really important that we understand what’s going on behind the scenes.
Obviously, Chatbot Arena is entirely open source. People go in, put in a prompt, get a response from 2 different models, and then say which one is better. The Leaderboard Illusion paper found that some companies are getting access to a lot more private testing in the background than others.
For instance, before Llama 4 launched, Meta released 27 models on the Arena, but of course only 1 was actually reported in the end. That undermines the integrity of the Arena because the more comparisons you have for your model, the more access to prompts you have, and the more data you have to refine a better model that’s better at the Arena. It adds an element of bias into the data that’s very, very hard to get around.
There are other issues that were called out, and issues that we’ve seen ourselves, which we think dictate the need for a more rigorous and methodologically sound approach to creating these kinds of human-preference datasets. Beyond the criticisms in The Leaderboard Illusion paper, we think that paper didn’t really touch on some of the other things we think should be considered when you’re doing human-preference evaluation.
4. HUMANE Makes Feedback Actionable
For me, there are 3 big areas where we’ve sought to improve. First of all, as you mentioned, the sample. The sample for Chatbot Arena is anybody, right? We don’t know anything about them, and we don’t collect any demographic data. They’re just people going there anonymously, prompting the models, and giving their preference data.
Obviously, that’s great. You get a huge amount of data, which is fantastic, but you know nothing about the people giving the data, which is fairly suboptimal. In terms of specificity, for anybody who’s used Chatbot Arena, all you’re doing is saying, “I like this response more,” or, “I like this response more.”
In the real world, that kind of data is useless, in a sense. It gives you a really nice way to make a nice leaderboard of AI models, but it tells the companies nothing about why that preference has been given. In our approach, we sought to mitigate that by splitting preference down into its constituent parts.
Things like: How helpful do people find models? What’s the communication like? How adaptive do they find them? What do they think of the model’s personality? When you get ratings on those kinds of factors, what you actually get is an actionable set of results that say, “Okay, your model is struggling with trust,” or, “Your model is struggling with personality.” That’s where you need to be focusing to actually build a model that is good for real users in the real world.
But there’s no QA in the sense that I could go in and just say hello, or say absolutely nothing, or have a multi-turn conversation and completely wander from “How big is the sun?” to “How long is a snake?”—just topic wandering. I don’t think that gives a really good, nuanced view of models. So we’ve built into our structure that participants come in and have multi-step conversations with models.
We built in QA that actually says, if you put low effort into your question or you start wandering, we're going to penalize you. Three of those and you're out. Those are the kinds of principles we built the leaderboard around.
5. TrueSkill Guides Model Battles
Nora Petrova
I would touch upon the methodology that we've used, which is TrueSkill. It's a framework developed by Microsoft for estimating the skill levels of players on Xbox Live. They take into account things like randomness in games, changing skill levels across time, and whether someone is having a fluky win streak versus a seasoned player who consistently performs well. All of these things were factors that we thought would be good to take into account.
It's a very flexible system that estimates probabilities with Bayesian distributions, with a mean and a variance that get narrower and narrower over time as the system learns about the outcomes of these battles or comparisons. Most importantly, it's based on information gain. The way we pick the next pair that should occur in the tournament is based on how much we will learn from these models going head-to-head. How much information are they giving us? How much are they reducing the uncertainty?
We order the queue of pairs according to that, and that gets us to a place of minimized uncertainty as fast as possible. It's a really flexible approach. We can run separate tournaments, as we've done with our demographic groups. We have around 20 demographic groups, and we've run separate tournaments for them.
We can consolidate the findings for each tournament to obtain an overall leaderboard that is much less uncertain than any of the individual tournaments or leaderboards that we can produce from any of the demographic groups. It really allows us to slice and dice the data in any way we want, and we can easily add more demographic groups and more models over time. We're developing in the open and welcoming feedback.
Andrew Gordon
I think one of the things they pointed out in “The Leaderboard Illusion” paper was that some models are sampled considerably more than other models. I believe the folks behind Chatbot Arena said that's because people come to the Arena to play with the latest models. People want to play with the state of the art, which is all well and good, but it doesn't lead to an efficient sampling method.
It essentially means that some models get a lot more battles than others. They therefore get a lot more data and get a lot better in the Arena. There's a pretty strong relationship between the number of battles and the place on the leaderboard. We only ever conduct battles based on the need from the data. If the uncertainty is high for a specific model against another specific model, we conduct a battle to lower that uncertainty.
So it's all driven by the data. It's very computationally sound because we don't make any more comparisons than we need to. It allows us to get to a point where models are strongly differentiated based on uncertainty.
Nora Petrova
If we have a certain goal with regard to uncertainty in order to fully differentiate the models at the confidence interval we're interested in, we can conduct more battles until we get there. The control is in our hands. We just need to recruit more participants in order to get to that level of certainty.
Andrew Gordon
The way we've sampled for this study, we're obviously using our own participants from the Prolific platform, but we've sampled effectively based on the census data that we have for both the US and the UK. The long-term vision would obviously be a more global product, but at the moment, we stratify our sample—that is, our participants who are giving us this feedback—by demographics like their age, ethnicity, and political alignment.
We have an awful lot of data from censuses that tells us each country is made up of certain proportions of these demographics. That essentially allows us to say that when we've amalgamated all these findings and found that leading model, we can very confidently say that model is preferred by as representative a set of the general public as we can possibly get.
Hopefully, in that sense, it's more related to the real-world preferences of people in the world rather than a very potentially skewed and biased subset that might be responding to Chatbot Arena. We ran our first one as an MVP, a proof of concept that was more about proving that we could do this in a rigorous and methodologically sound way. When we actually ran it, we only ran it with 500 participants.
It gave us a lot of insights about how we build Humane, which is our leaderboard that we're working on at the moment. That leaderboard is actually running as we speak in the background. We're still having battles, so we expect to have more data from that.
6. Models Struggle With Personality
What I can say about the first round that we did is that, across the board, the 6 models we tested, which were leading models at the time, tended to perform a lot worse on personality metrics and background and culture metrics than on things like helpfulness, communication, and adaptiveness. What that really signals is that there's some more subjective aspect of these models that people are less impressed by.
Potentially, they were doing tasks that don't elicit a personality in the model, or they don't elicit the model talking about background and culture. Also, the model doesn't know their background and culture, so it's very hard to align with them. But the other possibility is that models are just not very good at that. That would potentially be an effect of the data they've been trained on, right?
We know very little. Obviously, models are trained on the entire internet, but when you train a model on the entire internet, do you get a personality that really represents what people want? From this testing, we found that generally people were less impressed with model personality or its ability to have an understanding of their background and culture than with more objective measures.
Nora Petrova
Obviously, a lot of these models have undergone extensive fine-tuning to tailor their personalities or tailor how they approach answering questions, which are different across the different companies. We've observed recently that there has been an increase in sycophancy, or this people-pleasing behavior of models, and people generally don't seem to like it.
One thing that the results of this experiment and the later leaderboard datasets will allow us to answer is what the correlation is between telltale signs of sycophancy and downvotes in the personality metric, and whether that influences people's decisions about which model they prefer. We can perform various types of post-processing and analysis of the data to identify the levels of sycophancy observed in the datasets and try to identify the interesting relationships between the feedback that people gave, more model-driven, LLM-as-a-judge–oriented analysis of the conversations, and the classification of various patterns and what the models exhibit.
It's quite interesting to see what we'll find.